Trained from scratch on the lexicons, the grammarians, and fourteen centuries of poetry and prose. No English. No modern web Arabic. No translations.
لغة الضاد، بلا لغة أخرى
Every large model speaks Arabic as a translation of English. It was trained on web text, subtitles, and machine-translated pages, so its Arabic carries English sentence shapes, borrowed idioms, and broken grammar. Ask it for a line of verse and it will invent one and attribute it to al-Mutanabbī.
The classical language is different. Sībawayh formalized its grammar in the 8th century, Lisān al-ʿArab catalogued its vocabulary root by root, and its poetry is measured to the syllable. That language is fully documented, machine-readable, and almost absent from the data these models learned from.
النماذج الكبيرة اليوم تتكلم العربية كأنها ترجمة عن الإنجليزية. تدرّبت على نصوص الويب والترجمة الآلية، فجاءت عربيتها بتراكيب أعجمية وأخطاء نحوية. واللغة الأصيلة اللي قعّدها سيبويه وجمعها ابن منظور موجودة كاملة ومرقمنة، بس شبه غائبة عن بيانات التدريب.
Pre-trained from scratch on lexicons, grammar and morphology treatises, poetry anthologies, and classical prose. Nothing written after the 19th century. Nothing translated.
Trained on the corpus itself, so roots, patterns, prefixes, and suffixes become the units the model thinks in, instead of the byte fragments English-first tokenizers produce.
Saudi Arabia's ALLaM understands the request and arranges the answer. It never rewrites a word. Every sentence the user reads was produced by Dhad.
Arabic is called lughat al-ḍād, the language of the letter ḍād, a sound Arabs held to be theirs alone. The model is named for it.
Dhad is a venture of TRIO, developed by Datoura Analytics, its data and AI arm. The first model release is planned for Q4 2026. Research partners, publishers, and institutions with classical corpora are welcome to write.
info@datoura.com