Classical Arabic language model

The Arabic model that speaks only classical Arabic.

Trained from scratch on the lexicons, the grammarians, and fourteen centuries of poetry and prose. No English. No modern web Arabic. No translations.

لغة الضاد، بلا لغة أخرى

ḍād · the 15th letterضاد
Al-Ṭawīl, scannedIn the 8th century Al-Khalīl ibn Aḥmad reduced Arabic verse to two symbols: a moving letter and a still one. Dhad learns from the poetry he measured.
فعولن
مفاعيلن
فعولن
مفاعيلن
The problem

Today's models learned Arabic second-hand.

تعلّمت العربية من الترجمة

Every large model speaks Arabic as a translation of English. It was trained on web text, subtitles, and machine-translated pages, so its Arabic carries English sentence shapes, borrowed idioms, and broken grammar. Ask it for a line of verse and it will invent one and attribute it to al-Mutanabbī.

The classical language is different. Sībawayh formalized its grammar in the 8th century, Lisān al-ʿArab catalogued its vocabulary root by root, and its poetry is measured to the syllable. That language is fully documented, machine-readable, and almost absent from the data these models learned from.

النماذج الكبيرة اليوم تتكلم العربية كأنها ترجمة عن الإنجليزية. تدرّبت على نصوص الويب والترجمة الآلية، فجاءت عربيتها بتراكيب أعجمية وأخطاء نحوية. واللغة الأصيلة اللي قعّدها سيبويه وجمعها ابن منظور موجودة كاملة ومرقمنة، بس شبه غائبة عن بيانات التدريب.

The approach

A model that has never read anything else.

نموذج ما قرأ غير العربية الأصيلة
Corpus

Classical sources only

مدوّنة تراثية فقط

Pre-trained from scratch on lexicons, grammar and morphology treatises, poetry anthologies, and classical prose. Nothing written after the 19th century. Nothing translated.

Tokenizer

Built for Arabic roots

توكنايزر عربي

Trained on the corpus itself, so roots, patterns, prefixes, and suffixes become the units the model thinks in, instead of the byte fragments English-first tokenizers produce.

Orchestration

ALLaM alongside, not inside

ALLaM إلى جانبه

Saudi Arabia's ALLaM understands the request and arranges the answer. It never rewrites a word. Every sentence the user reads was produced by Dhad.

The corpus

What Dhad reads.

ما يقرأه ضاد
6source families
~150Mwords of classical Arabic
14centuries, pre-Islamic to 1300 AH
0words of English or modern Arabic
المعاجمLisān al-ʿArab, al-Qāmūs, Tāj al-ʿArūs40–60M
النحوSībawayh, al-Zamakhsharī, Ibn ʿAqīl5–8M
الصرفIbn ʿUṣfūr, Ibn al-Ḥājib1–2M
الشعرMuʿallaqāt, Mufaḍḍaliyyāt, the dīwāns10–20M
النثرKalīla wa-Dimna, al-Maqāmāt, al-Aghānī20–30M
الأدبal-ʿIqd al-Farīd, ʿUyūn al-Akhbār10–15M
Who it is for

Where authentic Arabic is the product.

لمن تكون الفصحى هي المنتج
Linguists and lexicographersRoot derivation, attested meanings, parsing with citations from the source texts.
Publishers and editorsClassical register drafting and review that does not drift into translated Arabic.
EducationGrammar and morphology explanations grounded in the treatises students are tested on.
Brands and namingNames drawn from the rare vocabulary that no web-trained model remembers.
وليل كموج البحر أرخى سدوله · علي بأنواع الهموم ليبتلي
Imruʾ al-Qays · Muʿallaqa · meter al-Ṭawīl

Arabic is called lughat al-ḍād, the language of the letter ḍād, a sound Arabs held to be theirs alone. The model is named for it.

Contact

Building in Riyadh.

نبني من الرياض

Dhad is a venture of TRIO, developed by Datoura Analytics, its data and AI arm. The first model release is planned for Q4 2026. Research partners, publishers, and institutions with classical corpora are welcome to write.

info@datoura.com