Insights 6 min read
Arabic dialects and Franco-Arabic in enterprise AI: why MSA-only systems fail and what to test
Why AI built for Modern Standard Arabic misreads Gulf, Egyptian and Levantine dialects and Franco-Arabic, and what to test and index before launch.
Most Arabic AI is built and tested on Modern Standard Arabic (MSA). Your customers and staff rarely write in it. They write "عايز أحجز ميعاد بكرة", "ابغى اعرف وين طلبي", or "3ayez a7gez" in Latin letters and numbers, often with English words in the middle. A system that only reads MSA misreads these messages, and you only find out before launch if you test on real ones.
What does real Arabic text look like?
Customer chats, emails, support tickets and staff questions mix every form of written Arabic. A single inbox can hold all of these in one morning.
Spelling variants
أحمد · احمد · إحمد
Mixed languages
تم approve الـ invoice
Franco-Arabic
3ayez a3raf el mo3ad
Dialects
أبغى · عايز · بدي
Scanned PDFs
OCR on stamped pages
RTL tables
Columns read right to left
Here is one request, to book an appointment for tomorrow, written six ways.
| Message | Where an MSA-only system can slip | |
|---|---|---|
| MSA | أريد حجز موعد غدًا | Usually reads it correctly |
| Egyptian | عايز أحجز ميعاد بكرة | May not know عايز, ميعاد or بكرة |
| Gulf | ابغى احجز موعد باكر | May read باكر as "early" instead of "tomorrow" |
| Levantine | بدي احجز موعد بكرا | May not know بدي or بكرا |
| Franco-Arabic | 3ayez a7gez maw3ed bokra | Sees Latin text with numbers it cannot place |
| Mixed | عايز أعمل reschedule للـ appointment | Splits the sentence and loses the intent |
How does Franco-Arabic work?
Franco-Arabic, also called Arabizi, writes Arabic in Latin letters. Numbers stand in for sounds Latin has no letter for: 2 for ء, 3 for ع, 5 for خ, 7 for ح. So "a7gez" is أحجز and "3ayez" is عايز.
There is no standard spelling. The same word shows up as "3ayez", "3ayz" or "ayez", and "maw3ed" as "mo3ad" or "ma3ad". People also switch scripts halfway through a message. A system that expects one spelling per word will miss most of it.
Why does Arabic break before the model sees it?
Search runs first. If retrieval cannot match the question to the right passage, even a strong model answers from the wrong page or not at all. Five things trip it up:
- Alef forms: أحمد, احمد and إحمد are the same name.
- Final letters: على and علي, or مدرسة and مدرسه, are often typed interchangeably.
- Diacritics: مُدَرِّسَة in a policy will not match مدرسة in a question unless the marks are stripped.
- Digits: an invoice number typed as ١٢٣ will not match 123.
- Scans: stamps, signatures and right-to-left tables confuse OCR that was not built for Arabic.
Normalisation fixes most of this, but it has a cost. Merge ي and ى and the name علي becomes the word على. That is why the normalised text is used for matching only, while answers and citations quote the original.
How should retrieval index Arabic?
- Scanned PDFعقد التوريد
- Read (OCR)Text and tables extracted
- Normaliseأ، إ، آ become ا
- SearchFinds clause 7.2
- Answer + sourceCited · page 14
A retrieval assistant (RAG) finds the right passages first, then writes an answer from them. For Arabic, the index needs a few extra steps.
Read scans
Arabic OCR that keeps page numbers and table order.
Normalise a copy
Unify alef, yaa, taa marbuta and digits; strip diacritics and tatweel.
Keep the original
Show the exact wording in answers and citations.
Cover both scripts
Convert Franco-Arabic queries to Arabic script before search.
Search two ways
Arabic-aware keywords plus search by meaning.
Add a short glossary that maps dialect words to the terms your documents use: عايز, ابغى and بدي to أريد, and ميعاد to موعد. Keep English terms your staff use, such as "SLA" or "invoice", linked to their Arabic equivalents. Your own team writes this list, and it grows with every failure you find.
What should you test before launch?
Test on your own messages. A general Arabic benchmark tells you little about how your customers in Riyadh, Cairo or Amman write to you.
- Real questionsWritten with your team
- Run the systemSame questions, every release
- Score each answerCorrect, cited, safe
- Go / no-goAgreed threshold
- 01Test setIs it built from real messages?Take anonymised chats, emails and tickets, not questions your team made up.
- 02Test setIs each message labelled by dialect and script?Mark MSA, Gulf, Egyptian, Levantine, Franco-Arabic and mixed, so you can sort the results.
- 03ScoringDo you report accuracy per dialect?A good overall average can hide a dialect that fails most of the time.
- 04ScoringDoes each citation point to the right page?Check that the quoted passage contains the answer, not a nearby one.
- 05ScoringDoes it decline when the documents are silent?Include questions with no answer and count how often it guesses.
- 06RepliesDoes it reply in the right register?Customers may write in dialect, but most companies reply in clear, professional Arabic.
A correct citation names the file, the clause and the page, and quotes the passage, so the reader can open it and check.
ما هي غرامة التأخير في عقد التوريد؟
0.5% of the order value per day of delay, capped at 10%.
Supply contract.pdf · Clause 7.2 · page 14
«تُفرض غرامة تأخير قدرها 0.5% عن كل يوم…»
Every answer shows its source
Agree the go-live threshold with the teams who will use it. Then run the same set after every change to the model, the documents or the glossary.
When should a person step in?
Some messages should never get an automatic reply. Route these to a person:
- The assistant cannot find a source, or its confidence is low.
- The message is a complaint, a legal threat or plainly angry.
- The reply would approve a refund or a payment, or give medical or legal advice.
- The customer asks for a human.
- AI draftsReply, memo, or update
- Human approvesThe right role signs off
- Action runsIn your systems
- LoggedWho, what, when
The assistant drafts, a person approves, and every handover is logged. Review those handovers each month. Many of them belong in your test set.
Where should you start?
Start with a Readiness Sprint of 2 to 3 weeks. We take a sample of your real messages and documents, build a dialect-labelled test set with your team, and test OCR, normalisation and retrieval on it, inside your own environment. You leave knowing which dialects work and which need more work before anyone relies on them.
The pilot then runs 4 to 8 weeks as part of the Private AI Launchpad. For more on how we handle Arabic, see Arabic AI.