Skip to content

Insights 6 min read

Arabic dialects and Franco-Arabic in enterprise AI: why MSA-only systems fail and what to test

Why AI built for Modern Standard Arabic misreads Gulf, Egyptian and Levantine dialects and Franco-Arabic, and what to test and index before launch.

Most Arabic AI is built and tested on Modern Standard Arabic (MSA). Your customers and staff rarely write in it. They write "عايز أحجز ميعاد بكرة", "ابغى اعرف وين طلبي", or "3ayez a7gez" in Latin letters and numbers, often with English words in the middle. A system that only reads MSA misreads these messages, and you only find out before launch if you test on real ones.

What does real Arabic text look like?

Customer chats, emails, support tickets and staff questions mix every form of written Arabic. A single inbox can hold all of these in one morning.

  • Spelling variants

    أحمد · احمد · إحمد

  • Mixed languages

    تم approve الـ invoice

  • Franco-Arabic

    3ayez a3raf el mo3ad

  • Dialects

    أبغى · عايز · بدي

  • Scanned PDFs

    OCR on stamped pages

  • RTL tables

    Columns read right to left

Here is one request, to book an appointment for tomorrow, written six ways.

One request written six ways
MessageWhere an MSA-only system can slip
MSAأريد حجز موعد غدًاUsually reads it correctly
Egyptianعايز أحجز ميعاد بكرةMay not know عايز, ميعاد or بكرة
Gulfابغى احجز موعد باكرMay read باكر as "early" instead of "tomorrow"
Levantineبدي احجز موعد بكراMay not know بدي or بكرا
Franco-Arabic3ayez a7gez maw3ed bokraSees Latin text with numbers it cannot place
Mixedعايز أعمل reschedule للـ appointmentSplits the sentence and loses the intent

How does Franco-Arabic work?

Franco-Arabic, also called Arabizi, writes Arabic in Latin letters. Numbers stand in for sounds Latin has no letter for: 2 for ء, 3 for ع, 5 for خ, 7 for ح. So "a7gez" is أحجز and "3ayez" is عايز.

There is no standard spelling. The same word shows up as "3ayez", "3ayz" or "ayez", and "maw3ed" as "mo3ad" or "ma3ad". People also switch scripts halfway through a message. A system that expects one spelling per word will miss most of it.

Why does Arabic break before the model sees it?

Search runs first. If retrieval cannot match the question to the right passage, even a strong model answers from the wrong page or not at all. Five things trip it up:

  • Alef forms: أحمد, احمد and إحمد are the same name.
  • Final letters: على and علي, or مدرسة and مدرسه, are often typed interchangeably.
  • Diacritics: مُدَرِّسَة in a policy will not match مدرسة in a question unless the marks are stripped.
  • Digits: an invoice number typed as ١٢٣ will not match 123.
  • Scans: stamps, signatures and right-to-left tables confuse OCR that was not built for Arabic.

Normalisation fixes most of this, but it has a cost. Merge ي and ى and the name علي becomes the word على. That is why the normalised text is used for matching only, while answers and citations quote the original.

How should retrieval index Arabic?

  1. Scanned PDFعقد التوريد
  2. Read (OCR)Text and tables extracted
  3. Normaliseأ، إ، آ become ا
  4. SearchFinds clause 7.2
  5. Answer + sourceCited · page 14

A retrieval assistant (RAG) finds the right passages first, then writes an answer from them. For Arabic, the index needs a few extra steps.

  1. Read scans

    Arabic OCR that keeps page numbers and table order.

  2. Normalise a copy

    Unify alef, yaa, taa marbuta and digits; strip diacritics and tatweel.

  3. Keep the original

    Show the exact wording in answers and citations.

  4. Cover both scripts

    Convert Franco-Arabic queries to Arabic script before search.

  5. Search two ways

    Arabic-aware keywords plus search by meaning.

Add a short glossary that maps dialect words to the terms your documents use: عايز, ابغى and بدي to أريد, and ميعاد to موعد. Keep English terms your staff use, such as "SLA" or "invoice", linked to their Arabic equivalents. Your own team writes this list, and it grows with every failure you find.

What should you test before launch?

Test on your own messages. A general Arabic benchmark tells you little about how your customers in Riyadh, Cairo or Amman write to you.

  1. Real questionsWritten with your team
  2. Run the systemSame questions, every release
  3. Score each answerCorrect, cited, safe
  4. Go / no-goAgreed threshold
  1. 01Test setIs it built from real messages?Take anonymised chats, emails and tickets, not questions your team made up.
  2. 02Test setIs each message labelled by dialect and script?Mark MSA, Gulf, Egyptian, Levantine, Franco-Arabic and mixed, so you can sort the results.
  3. 03ScoringDo you report accuracy per dialect?A good overall average can hide a dialect that fails most of the time.
  4. 04ScoringDoes each citation point to the right page?Check that the quoted passage contains the answer, not a nearby one.
  5. 05ScoringDoes it decline when the documents are silent?Include questions with no answer and count how often it guesses.
  6. 06RepliesDoes it reply in the right register?Customers may write in dialect, but most companies reply in clear, professional Arabic.

A correct citation names the file, the clause and the page, and quotes the passage, so the reader can open it and check.

ما هي غرامة التأخير في عقد التوريد؟

0.5% of the order value per day of delay, capped at 10%.

Supply contract.pdf · Clause 7.2 · page 14

«تُفرض غرامة تأخير قدرها 0.5% عن كل يوم…»

Every answer shows its source

Agree the go-live threshold with the teams who will use it. Then run the same set after every change to the model, the documents or the glossary.

When should a person step in?

Some messages should never get an automatic reply. Route these to a person:

  • The assistant cannot find a source, or its confidence is low.
  • The message is a complaint, a legal threat or plainly angry.
  • The reply would approve a refund or a payment, or give medical or legal advice.
  • The customer asks for a human.
  1. AI draftsReply, memo, or update
  2. Human approvesThe right role signs off
  3. Action runsIn your systems
  4. LoggedWho, what, when

The assistant drafts, a person approves, and every handover is logged. Review those handovers each month. Many of them belong in your test set.

Where should you start?

Start with a Readiness Sprint of 2 to 3 weeks. We take a sample of your real messages and documents, build a dialect-labelled test set with your team, and test OCR, normalisation and retrieval on it, inside your own environment. You leave knowing which dialects work and which need more work before anyone relies on them.

The pilot then runs 4 to 8 weeks as part of the Private AI Launchpad. For more on how we handle Arabic, see Arabic AI.