The best AI Model in Africa and the middle east
▲ 5 r/developer+2 crossposts

The best AI Model in Africa and the middle east

Today, we are officially announcing Early Access for our latest and most advanced model, Horus Cyper Nano 1.0 BETA.

We are making Horus Cyper Nano 1.0 BETA available to developers, researchers, and students through our Early Access program.

You can apply through the official Early Access portal. Once you meet the required eligibility criteria and your application is approved, you will receive your personal Access Token, which can be used through our NeuralNode Framework to access and integrate the model.

Apply for Early Access:
https://tokenai.llc/horus-cyper-nano-access

Horus Cyper Nano is a specialized cybersecurity model designed for offensive security and cybersecurity research workflows.

Its core use cases include:

Offensive security and red teaming, including penetration testing workflow support, vulnerability analysis, and exploitation path building.

Capture The Flag challenges and cybersecurity training.

Active Directory security, including enumeration and lateral movement planning within authorized engagements.

Authorized security testing labs and controlled environments.

Safe and scoped cybersecurity research within authorized environments.

Red team report drafting and attack chain structure planning.

Horus Cyper Nano 1.0 will be the first release in the Horus Cyper series, a family of specialized cybersecurity models developed by TokenAI, an AI startup based in Egypt.

The Open Weights of Horus Cyper Nano 1.0 will be released on September 3, 2026, which also happens to be my 19th birthday.

What a way to celebrate.

Our vision is to build Horus Cyper Nano into one of the strongest cybersecurity AI models to emerge from Egypt, the Arab world, the Middle East, and Africa, and to establish it as one of the leading openly available cybersecurity models across the region.

This is only the beginning of the Horus Cyper series.

Horus Cyper Nano 1.0 BETA
Developed by TokenAI
Built in Egypt

u/assemsabryy — 12 days ago

Introducing Horus Hiero | A Hieroglyphic Language Translation Model

Our new open-source AI model for Ancient Egyptian hieroglyph translation.

Available in two versions:

  • Horus Hiero 9B
  • Horus Hiero Mini 4B (optimized for CPUs and mobile devices)

Built on Qwen 3.5, Horus Hiero supports ~150 languages, understands text, images, and video, and is the first model of its kind to combine large-scale multimodal capabilities with dedicated hieroglyph translation.

It also delivers strong general reasoning and coding performance:

  • 79% on MMLU-Pro
  • 63% on LiveCodeBench
  • 84% on HumanEval

With a 512K context window (expandable up to 1M tokens), it offers one of the largest context windows available in the Arab AI ecosystem.

We hope Horus Hiero helps make Ancient Egyptian heritage more accessible, supports tourism, and encourages the study and understanding of hieroglyphs.

The models are fully open source on Hugging Face with full support through the NeuralNode framework.

https://huggingface.co/collections/tokenaii/horus-hiero

u/assemsabryy — 1 month ago

Horus Image Generation is here! 🤩📷

https://preview.redd.it/n55ohr6wrd5h1.png?width=1537&format=png&auto=webp&s=991397299a33b91459c9b33597ea920bf43abc28

I'm not here to promote my work or make money from what I'm about to say.

I'm here to say that Egypt is already part of the AI race.

Today, at TokenAI, we announced our first image generation model and the first release in the Horus Lens family: Horus Lens 1.0.

Horus Lens is a family of models specialized in text-to-image generation, forming a dedicated branch of the broader Horus model family developed and owned by TokenAI.

This launch marks an important step forward for Egypt's AI ecosystem and highlights the growing role of the region in advancing artificial intelligence technologies.

reddit.com
u/assemsabryy — 3 months ago

نموذج حورس لتوليد الصور! 🤩📷

https://preview.redd.it/8fpou4gird5h1.png?width=1537&format=png&auto=webp&s=abfb6d187f369b946d19479bee896876981a05bd

https://preview.redd.it/c4lhunmjrd5h1.png?width=1254&format=png&auto=webp&s=51137e986a1b2ee6a5b69211069d65a8d8fd739a

بفضل الله انهاردة بنعلن عن نموذج Horus Lens 1.0
و هو اول نموذج من مجموعة نماذج Horus Lens المتخصصة في توليد الصور بالذكاء الاصطناعي

خطوة كبيرة و عظيمة في Tokenai تعزز صناعة الذكاء الاصطناعي في جمهورية مصر و في الوطن العربي بأكمله
و نحط في الاعتبار ان نماذج توليد الصور هي اعقد و اصعب نوع نماذج ذكاء اصطناعي في الوجود و الاكثر تكلفة ماديا و Computing

و بالرغم من كل ده انهاردة بعلن عن اول نموذج توليد صور من TokenAI و السلسلة الاولي من نوعها مفتوحة المصدر في الوطن العربي
حيث ان سلسلة Horus Lens دخلت خطتنا المستقبلية و الي هنتابعها بتحديثات كبيرة لنماذج Horus Lens و لعائلة نماذج Horus بشكل عام

بعد بحث طويل تأكدت من ان مشروع Horus Lens هو الاول من نوعه في جمهورية مصر يعني صناعة مصرية خالصة 100% 🇪🇬

و الاول من نوعه في الوطن العربي بأكمله من بعد نموذج Fanar Image Generation الي تم الاعلان عنه لكن النموذج عبارة عن LoRA Adapter ليس الا بيشتغل علي نموذج تاني ملوش علاقة
بنموذج Fanar Image Generation

فا نقدر نقول ان Horus Lens انجاز جديد و بيتقدم علي طبق من ذهب للمطورين و الباحثين و لكل الناس لأن كما ذكرت النموذج مفتوح المصدر تماما

مش محتاج اقولك اكيد صورة البوست دي عملتها ازاي🫠🦅

و زي ما قولت في ابريل الي فات و بعيد و بكرر تاني

إحنا قدام مشروع قادر يحط مصر على خريطة الذكاء الاصطناعي عالميًا و هنا اكيد بتكلم عن نماذج حورس للذكاء الاصطناعي.

نموذج Horus Lens 1.0 مفتوح المصدر تحت رخصة Apache license 2.0

و كمان النموذج ليه خمس نسخ مختلفة من حيث الضغط و حجم النموذج عشان تناسب كل الاجهزة و كل المستخدمين علي حسب قدرات اجهزتهم

النموذج متوفر عبر اطار عملنا Neuralnode و تقدر تشوف كل تفاصيل النموذج عل
الموقع الرسمي ل Tokenai

https://lnkd.in/dmZ7m-cU

متحمس اشوف الناس هتعمل ايه بالموديل و الصور الي هتتولد بنموذج Horus Lens 1.0 و اكيد بأنتظار رأيكم.🩶

Enjoy 📸🦅

reddit.com
u/assemsabryy — 3 months ago

Horus Image Generation is here! 🤩📷

https://preview.redd.it/57kqog9iqd5h1.png?width=1537&format=png&auto=webp&s=85b3ec32b0797bdeb2a0210881164f8806f54bf1

I'm not here to promote my work or make money from what I'm about to say.

I'm here to say that Egypt is already part of the AI race.

Today, at TokenAI, we announced our first image generation model and the first release in the Horus Lens family: Horus Lens 1.0.

Horus Lens is a family of models specialized in text-to-image generation, forming a dedicated branch of the broader Horus model family developed and owned by TokenAI.

This launch marks an important step forward for Egypt's AI ecosystem and highlights the growing role of the region in advancing artificial intelligence technologies.

Horus Lens 1.0, the first model in the Horus Lens family, a specialized series of AI models focused on image generation.

This is a major milestone for TokenAI and a significant step forward for the AI industry in Egypt and across the Arab world.

It's important to recognize that image generation models are among the most complex, computationally demanding, and expensive types of AI systems to develop. Despite these challenges, today we are proud to introduce TokenAI's first image generation model and what we believe is the first open-source image generation model series of its kind in the Arab world.

Horus Lens has become a core part of our long-term vision, and we plan to continue expanding it with major updates and improvements, both for the Horus Lens family and the broader Horus AI ecosystem.

After extensive research, I confirmed that Horus Lens is the first project of its kind developed entirely in Egypt — a truly 100% Egyptian-made AI initiative. 🇪🇬

It is also the first open-source image generation model family of its kind in the Arab world following the announcement of Fanar Image Generation. However, Fanar was released as a LoRA adapter that relies on an existing base model rather than being a standalone image generation model.

For that reason, we can confidently say that Horus Lens represents a new achievement, offered openly to developers, researchers, and the wider community, as the model is fully open source.

I probably don't need to explain how the cover image of this post was created. 🫠🦅

As I said back in April, and I will say it again today:

We are building a project capable of putting Egypt on the global AI map — and I'm talking about the Horus family of AI models.

Horus Lens 1.0 is open source under the Apache License 2.0.

The model is also available in five different quantized versions, providing multiple size and performance options to suit different hardware capabilities and user requirements.

It is available through our Neuralnode framework, and you can explore the full model details on the official TokenAI website:

https://tokenai.cloud/models/horus-lens-1-0

I'm excited to see what developers, creators, and researchers will build with Horus Lens 1.0, and I'm looking forward to seeing the images generated by the community.

Enjoy. 📸🦅

reddit.com
u/assemsabryy — 3 months ago

افضل مرجع تعليمي ليك لو انت بتدرس AI 👇🏼🧑‍💻

https://preview.redd.it/058oj6u6nt2h1.jpg?width=1076&format=pjpg&auto=webp&s=a27d33f8173889c2e5bb62106d2d84bcd8d832e6

فاكرين البوست الي نزلته و فيه افضل كورسات لل AI ؟ انا بقا جمعتلك محتوي الكورسات دا كله بال Source codes في الريبو دا

📌 ريبو تعليمي كامل شامل كل حاجة عن الذكاء الاصطناعي بيعلمك ازاي تبني اول AI Model لوحدك

الريبو متقسم لخمس اجزاء:

  1. الأساسيات والرياضيات
  2. تعلم الآلة والتعلم العميق
  3. هندسة النماذج اللغوية الضخمة
  4. التطبيق العملي وبناء الموديل
  5. أهم 53 ورقة بحثية في تاريخ الـ AI

هتطلع من الريبو دا ملم بكل اساسيات و مصطلحات ال AI خصوصا تطوير النماذج الغوية ال LLMs زي GPT و Claude و غيره طبعا

الريبو مش بس تعليمي لا دا فيه Source code لمشاريع حقيقية
فيه 53 ورقة بحثية كاملة في مجال الذكاء الاصطناعي و شرحهم و كل بحث فيهم ليه شرحه و ليه انت محتاج تقرأ البحث دا و تدرسه

الريبو مش مركز علي ال LLMs بس لا دا كمان مجمع كل مفاهيم ال

Machine Learning
Deep Learning
Neural Networks

يعني انا بكلمك علي واحد من اضخم ال Repos التعليمية علي جيت هاب كله

الريبو فيه Source code لتدريب LLMs و Fine Tuning LLMs و كمان Deploying الموديل بتاعك

ريبو كامل متكامل لهندسة الذكاء الاصطناعي اوبن سورس بالكامل

و بعيد و بكرر تاني يا جماعة ابعد عن اي حد بيحاول يبيعلك كورس علي السوشيال ميديا خصوصا منصات زي فيسبوك
انت مش غ/بي لدرجة انك تروح تدفع فلوس لحد عشان يقولك تستخدم ChatGPT ازاي ولا ايه ال AI Tools الي تستخدمها
خصوصا ان اليومين دول هتبدأ تلاقي بقا بوستات خصم مش عارف كام في المية علي الكورس بمناسبة العيد الاضحي و انتو عارفين ان النصابين دول بيستغلو الاعياد و المناسبات للتسويق لمنتجهم الفاضي عن طريق خصم كذاب اصلا

استغل المصادر مفتوحة المصدر
انا جايبلك ريبو والله يخليك تعرف تدرب اول AI Model ليك مستريح
- يخليك ملم بكل المصطلحات المهمة في ال AI
- يخليك عارف ازاي تحط رجلك علي سلم ال AI و ال LLMs
- يخليك عارف تفرق بين ال Deep Learning و ال Machine Learning و ال Neural Networks

و الاهم يخليك عارف ان ال AI دا بحر كبير جدا و النماذج انواع مختلفة و مفيش حاجة اسمها تعالي هعلمك AI كله دا كدا كلام فاضي

الريبو دا بيركز علي اهم شئ و هو تطوير النماذج اللغوية LLMs

دا لينك الريبو

https://github.com/assemsabry/LLMs-from-scratch

الريبو كله طبعا مفتوح المصدر و دي اول مرة هطلب منكم شير كتير للبوست و للريبو في كل حتة
خلي اقصي عدد من الناس يستفيد بالكنز دا
و توعية للناس عن النصب و الكورسات الكدابة الي بيتم التسويق ليها
و "للاسف من الزملاء في نفس المجال" و كلها محتوي فاضي و كذاب

استعن بالله و خليك صبور و انت بتتعلم
ال AI واحد من اصعب الدراسات في العالم فا متفكرش انك في يوم و ليلة هتقدر تكتب في
البايو AI Engineer ادخل شوف الريبو و ادرسه كويس جدا

والله محتوي الريبو دا بيتباع في كورسات ب ارقام عالية و انا مقدمهولك open source علي طبق من ذهب

انا مقسملك كل حاجة بالترتيب و عامل جدولة كاملة لكل حاجة بالترتيب و كل جزء منفصل عن التاني
الي يحتاج حاجة او استفسار ال DM بتاعي مفتوح و اعذروني لو اتأخرت في الرد ساعات

https://github.com/assemsabry/LLMs-from-scratch

توكل علي الله و متنساش شير كتير. 🤗🩶

reddit.com
u/assemsabryy — 3 months ago

افضل مرجع تعليمي ليك لو انت بتدرس AI 👇🏼🧑‍💻

https://preview.redd.it/potcki15nt2h1.jpg?width=1076&format=pjpg&auto=webp&s=4479bcce206a9e597e848553dae00f3b2e0a5e99

فاكرين البوست الي نزلته و فيه افضل كورسات لل AI ؟ انا بقا جمعتلك محتوي الكورسات دا كله بال Source codes في الريبو دا

📌 ريبو تعليمي كامل شامل كل حاجة عن الذكاء الاصطناعي بيعلمك ازاي تبني اول AI Model لوحدك

الريبو متقسم لخمس اجزاء:

  1. الأساسيات والرياضيات
  2. تعلم الآلة والتعلم العميق
  3. هندسة النماذج اللغوية الضخمة
  4. التطبيق العملي وبناء الموديل
  5. أهم 53 ورقة بحثية في تاريخ الـ AI

هتطلع من الريبو دا ملم بكل اساسيات و مصطلحات ال AI خصوصا تطوير النماذج الغوية ال LLMs زي GPT و Claude و غيره طبعا

الريبو مش بس تعليمي لا دا فيه Source code لمشاريع حقيقية
فيه 53 ورقة بحثية كاملة في مجال الذكاء الاصطناعي و شرحهم و كل بحث فيهم ليه شرحه و ليه انت محتاج تقرأ البحث دا و تدرسه

الريبو مش مركز علي ال LLMs بس لا دا كمان مجمع كل مفاهيم ال

Machine Learning
Deep Learning
Neural Networks

يعني انا بكلمك علي واحد من اضخم ال Repos التعليمية علي جيت هاب كله

الريبو فيه Source code لتدريب LLMs و Fine Tuning LLMs و كمان Deploying الموديل بتاعك

ريبو كامل متكامل لهندسة الذكاء الاصطناعي اوبن سورس بالكامل

و بعيد و بكرر تاني يا جماعة ابعد عن اي حد بيحاول يبيعلك كورس علي السوشيال ميديا خصوصا منصات زي فيسبوك
انت مش غ/بي لدرجة انك تروح تدفع فلوس لحد عشان يقولك تستخدم ChatGPT ازاي ولا ايه ال AI Tools الي تستخدمها
خصوصا ان اليومين دول هتبدأ تلاقي بقا بوستات خصم مش عارف كام في المية علي الكورس بمناسبة العيد الاضحي و انتو عارفين ان النصابين دول بيستغلو الاعياد و المناسبات للتسويق لمنتجهم الفاضي عن طريق خصم كذاب اصلا

استغل المصادر مفتوحة المصدر
انا جايبلك ريبو والله يخليك تعرف تدرب اول AI Model ليك مستريح
- يخليك ملم بكل المصطلحات المهمة في ال AI
- يخليك عارف ازاي تحط رجلك علي سلم ال AI و ال LLMs
- يخليك عارف تفرق بين ال Deep Learning و ال Machine Learning و ال Neural Networks

و الاهم يخليك عارف ان ال AI دا بحر كبير جدا و النماذج انواع مختلفة و مفيش حاجة اسمها تعالي هعلمك AI كله دا كدا كلام فاضي

الريبو دا بيركز علي اهم شئ و هو تطوير النماذج اللغوية LLMs

دا لينك الريبو

https://github.com/assemsabry/LLMs-from-scratch

الريبو كله طبعا مفتوح المصدر و دي اول مرة هطلب منكم شير كتير للبوست و للريبو في كل حتة
خلي اقصي عدد من الناس يستفيد بالكنز دا
و توعية للناس عن النصب و الكورسات الكدابة الي بيتم التسويق ليها
و "للاسف من الزملاء في نفس المجال" و كلها محتوي فاضي و كذاب

استعن بالله و خليك صبور و انت بتتعلم
ال AI واحد من اصعب الدراسات في العالم فا متفكرش انك في يوم و ليلة هتقدر تكتب في
البايو AI Engineer ادخل شوف الريبو و ادرسه كويس جدا

والله محتوي الريبو دا بيتباع في كورسات ب ارقام عالية و انا مقدمهولك open source علي طبق من ذهب

انا مقسملك كل حاجة بالترتيب و عامل جدولة كاملة لكل حاجة بالترتيب و كل جزء منفصل عن التاني
الي يحتاج حاجة او استفسار ال DM بتاعي مفتوح و اعذروني لو اتأخرت في الرد ساعات

https://github.com/assemsabry/LLMs-from-scratch

توكل علي الله و متنساش شير كتير. 🤗🩶

reddit.com
u/assemsabryy — 3 months ago

A very important milestone for me in the AI field.

https://preview.redd.it/fstlfyxl8i1h1.png?width=1254&format=png&auto=webp&s=7c8ea60a4b508293705e9bea70ef36f4986bf1d6

Three days ago, I officially released my first AI research paper: STAM (Stable Training with Adaptive Momentum).

STAM introduces a new optimizer for deep learning and focuses on:

  • improving training stability,
  • reducing resource consumption during training,
  • and addressing several limitations found in optimizers like Adam, AdamW, and Muon.

The paper explains what makes STAM different, the problems it aims to solve, and includes comparisons with existing optimizers and training results.

The research paper is currently available on SSRN, and it has reached a ranking of around 646K so far. ()

What matters most to me is not numbers, but having AI engineers, researchers, and specialists read the paper and share honest technical feedback and criticism.

I consider STAM one of the biggest projects I’ve ever worked on, and I plan to continue improving and developing it further. I would genuinely appreciate hearing opinions from researchers and experienced people in the AI community about the paper, the optimizer design, and the reported results compared to other optimizers.

Research paper:
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6699059

https://preview.redd.it/kjvurqun8i1h1.png?width=1672&format=png&auto=webp&s=0a8c16ee59e7424901f56852a9aa2d7708a4404a

reddit.com
u/assemsabryy — 3 months ago

اول ورقة بحثية ليا ك AI Researcher

https://preview.redd.it/q6ncj02w131h1.png?width=1672&format=png&auto=webp&s=ee65e2d76c649b1807530b0f840a8229f5f566b0

ازيكم يشباب
انا عاصم صبري AI Engineer & Researcher
عملت انجاز جميل اوي فا حبيت اشاركو معاكم و اني نشرت اول Research paper ليا و الي بتتكلم عن Optimizer جديد في عالم الذكاء الاصطناعي

محتاج رأيكم عشان افرح اكتر
شوفو البيبر من هنا و قولولي رأيكم
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6699059

reddit.com
u/assemsabryy — 3 months ago

أول ورقة بحثية رسمية ليا في مجال الAI اتقبلت على SSRN

https://preview.redd.it/pubbe861hs0h1.jpg?width=910&format=pjpg&auto=webp&s=8a17b56e707e790bb6f81544cc21db1ba4bf8f30

النهارده ورقتي البحثية بعنوان
“Stable Training with Adaptive Momentum (STAM)”
اتقبلت رسميًا على منصة SSRN، وده يعتبر أول نشر بحثي رسمي وموثق ليا كـ AI Researcher.

البحث بيقدم خوارزمية Optimization جديدة لتدريب نماذج الـ Deep Learning، وقدرت تتفوق على شوية من أشهر الـ optimizers في بعض الـ benchmarks، وكمان بتعالج مشاكل مختلفة في استقرار التدريب، بالإضافة إنها قدرت تقلل التكلفة الحاسوبية أثناء التدريب بنسبة وصلت لـ 50% في بعض التجارب.

دي خطوة مهمة جدًا بالنسبالي في رحلة البحث العلمي، ومتحمس أكمل أشتغل على تطوير تقنيات تساعد في جعل تدريب نماذج الذكاء الاصطناعي أكثر كفاءة واستقرار.

تقدروا تشوفوا البحث من هنا:
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6699059

reddit.com
u/assemsabryy — 3 months ago

My First Official AI Research Paper Accepted on SSRN

https://preview.redd.it/oz4vpoxdfs0h1.jpg?width=910&format=pjpg&auto=webp&s=fa4c91aad0e3c56850fbfc06099e9c4095712bbd

Today, my research paper “Stable Training with Adaptive Momentum (STAM)” was officially accepted on SSRN — marking my first documented and official publication as an AI Researcher.

The paper introduces a new optimization algorithm for deep learning training that outperformed several popular optimizers in selected benchmarks, addressed multiple training stability challenges, and achieved up to 50% reduction in computational training cost in some experiments.

This is an important milestone in my research journey, and I’m excited to continue exploring optimization techniques for efficient and stable AI training.

You can read the paper here:
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6699059

reddit.com
u/assemsabryy — 3 months ago

https://preview.redd.it/3ccm5gd1puzg1.png?width=1179&format=png&auto=webp&s=c940d2e6ef1d61288ac214eae4679a7c910b7917

Today, I’m talking about a new research paper from Token AI:
"Stable Training with Adaptive Momentum"

It introduces what could be one of the strongest optimizers, both in theory and in results.

For years, we’ve relied on well-known optimizers like Adam, AdamW, LAMB, and others. No doubt, they’ve been the go-to choices when training AI models.

If you’re not familiar with what an optimizer is, in simple terms: it’s a core part of training any AI model. It’s the algorithm responsible for updating the model’s weights during training to reduce the loss.

That said, these optimizers come with limitations that affect training.

For example, Adam uses a fixed beta1 throughout training, which can carry outdated momentum and keep pushing the model in the wrong direction.

STAM addresses this by measuring the difference between the current gradient and previous momentum (g - m). When the difference is large, it reduces beta1, leading to more stable training during noisy phases.

Another issue appears when there’s a shift or noise in training. Old momentum can become harmful. STAM handles this with an adaptive beta1 based on residual variance.

A major issue in SGD is that if the direction becomes wrong, it keeps going due to fixed momentum. STAM solves this by allowing the first momentum to self-correct.

Now let’s talk about STAMLite, the lighter version.

It’s designed to replace AdamW as a default choice in many cases. The key difference is that beta1 is dynamic instead of fixed:

  • If gradients are noisy, it reduces momentum
  • If gradients are stable, it keeps momentum high

It also improves efficiency in terms of optimizer state memory:

  • AdamW requires about 2× the parameter size
  • STAM Full is close to AdamW
  • STAMLite requires about 1× the parameter size

In practice, STAMLite saves around 50% of the resources compared to AdamW and STAM, meaning significantly less GPU usage during training.

Looking at benchmarks, the results speak for themselves.

In Hyperparameter Sweep, STAMLite achieved:
Accuracy: 0.61
Loss: 0.91

In Long-Horizon Non-Stationary MLP, STAM ranked first alongside NAdam with nearly identical results:
Accuracy: 0.97
Loss: 0.09

More benchmarks are available on the website and in the research paper.

This is an important step from TokenAI, breaking the long-standing reliance on a limited set of optimizers that come with known issues.

Even as an early release, it proves strong and promising. Personally, I’ve already shifted to STAM and I’m currently training my first full LLM from scratch using it. I’ll be sharing the results soon.

Research paper:
https://tokenai.cloud/research/stam

Let me know what you think.

reddit.com
u/assemsabryy — 3 months ago

https://preview.redd.it/59ji98udnuzg1.png?width=1179&format=png&auto=webp&s=319e083bb75c3ff8e4cd4b55d5c2230b86c57531

انهاردة هتكلم عن ال Research paper الجديدة من Token AI

"Stable Training with Adaptive Momentum"

ورقة بحثية لأفضل Optimizer علي الورق و بالارقام

بعد اعتمادية استمرت لسنوات علي اشهر الـ Optimizers زي Adam و AdamW و LAMP و غيرهم كثير من ال Optimizers الي مفيش شك بانهم هما افضل اختيارات تختار منهم الـ Optimizer لتدريب الـ AI Model بتاعك

و عزيزي لو انت شخص بسيط و مش فاهم ايه هو الـ Optimizer فا بكل اختصار دا جزء لا يتجزء من تدريب اي نموذج ذكاء اصطناعي في الدنيا و دا بيكون ال Algorithm المسؤولة عن تعديل weights النموذج أثناء التدريب عشان تقلل قيمة ال loss

الحقيقة ان الـ Optimizers الي ذكرتهم دول بيواجهم مشاكل بتنعكس علي الموديل بتاعك اثناء التدريب زي مشكلة ان Adam بيستخدم beta1 ثابتة طول التدريب و دا الي بيشيل momentum قديم غلط ويفضل ماشي في اتجاه مش مناسب

و كان حلها في STAM انه بيقيس فرق الجريدينت الحالي عن المومنتوم القديم g - m ويقلل beta1 لو الفرق كبير و دا يديك تدريب أهدى في المراحل اللي فيها noise

مشكلة تانية زي لو حصل shift أو noise، الـ momentum القديم ممكن يتضر و دا الي كان حلها Adaptive beta1 حسب residual variance في STAM

مشكلة كبيرة تانية في SGD و ان لو الاتجاه القديم بقى غلط، بيكمل فيه بسبب ال momentum الثابت و المشكلة دي اتحلت ب First momentum بيعالج نفسه في STAM

طيب خلينا نتكلم عن النسخة التانية من STAM و هي STAMLite و مصمم عشان يكون اختيار افتراضي بدل AdamW في حالات كتير، لأنه بيضيف ميزة مهمة مش في AdamW إن beta1 بيتغير حسب حالة الـ gradients بدل ما يفضل ثابت طول التدريب

لو الـ gradients noisy بيقلل momentum

لو الـ gradients مستقرة فا بيحافظ علي momentum عالي

توفير موارد أثناء التدريب من ناحية الـ optimizer state memory

AdamW تقريبًا بيحتاج حوالي 2× حجم parameters

STAM Full قريب من AdamW في الذاكرة

STAMLite بيحتاج تقريبًا 1× حجم parameters

يعني عمليًا STAMLite يوفر حوالي 50% من الموارد المستهلكة من AdamW و STAM

يعني انت حرفيا بتوفر نص الموارد و ال GPUs في تدريب الموديل بتاعك باستخدام STAMLite

نيجي بقا لجزء ال Benchmarks و الارقام لا تكذب

نيجي ل Hyperparameter Sweep و الي تفوق فيها STAMLite ب Accuracy 0.61 و Loss 0.91

و في Long-Horizon Non-Stationary MLP كان في المركز الاول STAM و NAdam بتقريبا نفس النتيجة

Accuracy 0.97 و Loss 0.09

و بقيت ال Benchamarks كلها علي الويبسايت و في ال Research paper

انجاز جديد و مهم في TokenAI، فك احتكار و اعتمادية كاملة علي خيارات محدودة بتواجه مشاكل

نسخة اولية بتثبت جدارتها و قوتها فعلا و انا شخصيا وجهت اعتمادي حاليا لـ STAM و شغال علي تدريب اول LLM كامل from scratch علي STAM و هشارك معاكم النتائج في الايام الجاية

رابط الـ Research paper

https://tokenai.cloud/research/stam

و شاركوني رأيكم

Assem Sabry.

reddit.com
u/assemsabryy — 3 months ago

Following up on the Horus project — the first fully built-from-scratch language model in Egypt.

If this is your first time hearing about Horus: it’s a fully built-from-scratch language model, and it’s open-source.

https://preview.redd.it/v0lw20vuh5zg1.jpg?width=3267&format=pjpg&auto=webp&s=10af499b2c5aab925c549a64cd6a6149217c490a

https://preview.redd.it/3blbewtuh5zg1.jpg?width=1459&format=pjpg&auto=webp&s=fc7ce3c706ba94bc776305f8f172169a69c00818

Hugging Face repo: https://huggingface.co/tokenaii/horus

About a week ago, the source code used to train the model was also released, making it available for developers to explore, use, and build on.

https://github.com/tokenaii/horus-1.0

This makes Horus the first fully trained-from-scratch LLM in Egypt, developed by Assem Sabry and TokenAI.

Today, I’m sharing some early details about the next version: Horus 1.5 Instruct.

This new version is expected to be 5x better than Horus 1.0, with a 64K context length, which is 8x larger than the 8K context in Horus 1.0 4B.

But it’s not just about scaling — Horus 1.5 comes with major improvements in architecture and overall capability, pushing the model to a completely different level.

Also, there are updates about a new cybersecurity model from TokenAI.

A specialized model designed to detect vulnerabilities and fix them instantly. It’s planned to be a large-scale model, trained on trillions of highly specialized security-related data, which puts us in front of something extremely powerful.

All of this is fully built in Egypt, in the field of AI.

TokenAI is starting to seriously shift the AI scene in Egypt and the Arab world, and what we're building is honestly something exceptional.

More official announcements are coming soon about the next Horus models

bigger, stronger, and significantly more efficient.

reddit.com
u/assemsabryy — 4 months ago