Skip to content

Insights 6 min read

How to build an Arabic knowledge assistant your staff will actually trust

Why Arabic search fails, how retrieval with cited sources fixes it, and how to measure accuracy on your own documents before a single employee relies on it.

Staff trust an Arabic knowledge assistant when they can check every answer. That means it answers only from your documents, shows the clause and page it used, says "I don't know" when the documents are silent, and only shows people what they are already allowed to see. You prove all of this on your own questions before go-live.

Why does staff search fail on Arabic documents?

Most enterprise search was built for English. It assumes a word has one spelling and clear edges. Arabic breaks both.

The same word appears as أ, إ or ا, with ة or ه at the end. Prefixes attach directly, so وبالعقد and العقود rarely match. Then real content adds its own problems.

  • Spelling variants

    أحمد · احمد · إحمد

  • Mixed languages

    تم approve الـ invoice

  • Franco-Arabic

    3ayez a3raf el mo3ad

  • Dialects

    أبغى · عايز · بدي

  • Scanned PDFs

    OCR on stamped pages

  • RTL tables

    Columns read right to left

Mixed language is common. A staff member types "ما هو الـ SLA في عقد الصيانة؟" and expects an English clause from an Arabic contract. A customer writes in Franco-Arabic, such as "3ayez a3raf el ma3ad". Much of the archive is scanned PDFs, and the fee tables read right to left. Keyword search misses most of this, and people stop using it.

What does a knowledge assistant do differently?

It searches by meaning and by Arabic-aware keywords at the same time. Then a private language model writes a short answer from the passages it found, and only from those.

  1. Scanned PDFعقد التوريد
  2. Read (OCR)Text and tables extracted
  3. Normaliseأ، إ، آ become ا
  4. SearchFinds clause 7.2
  5. Answer + sourceCited · page 14

Scanned files go through Arabic OCR first. Text is normalised for search, but the original wording is kept for display. Every step runs inside your environment: on-premise, private cloud or a sovereign GCC cloud.

Keyword search compared with an Arabic RAG assistant
Keyword searchArabic RAG assistant
Spelling variantsMisses أ / إ / ا and ة / هNormalised before search
Mixed Arabic and EnglishMatches one language at a timeFinds passages in either language
Scanned PDFsOften not searchableArabic OCR, page numbers kept
What you getA list of files to readA short answer with its source
When nothing matchesEmpty or irrelevant resultsSays it cannot find an answer

What makes staff trust the answers?

Staff trust what they can check. A few rules make that possible, and users see them working every day.

Does every answer cite its source?

Yes. The answer links to the document, page and passage. The user opens it, sees the highlighted text and checks it before acting.

ما هي غرامة التأخير في عقد التوريد؟

0.5% of the order value per day of delay, capped at 10%.

Supply contract.pdf · Clause 7.2 · page 14

«تُفرض غرامة تأخير قدرها 0.5% عن كل يوم…»

Every answer shows its source

What happens when the documents don't cover it?

The assistant says "I don't know" and lists the closest documents. Staff can work with a clear "not found". A fluent guess sends them the wrong way.

Who sees what?

Retrieval is filtered by what each user can already open. An HR assistant does not show a salary policy draft to someone who could not see it in the document system.

Governance is part of the build. Every question, retrieved passage and answer is logged for audit. Anything sensitive, such as a reply to a citizen or a legal opinion, goes to a person for approval before it is sent.

  1. AI draftsReply, memo, or update
  2. Human approvesThe right role signs off
  3. Action runsIn your systems
  4. LoggedWho, what, when

How do you know it is accurate before go-live?

You test it on your own questions. A generic benchmark tells you little about your policies, your scans or the way your staff write.

  1. Real questionsWritten with your team
  2. Run the systemSame questions, every release
  3. Score each answerCorrect, cited, safe
  4. Go / no-goAgreed threshold

Build a test set with the teams who will use it. Collect real questions in Arabic, English and both mixed together. Write down the correct answer and the page it comes from. Add hard cases on purpose: dialect, typos, poor scans, and questions the documents cannot answer.

Run the assistant against the set and score four things. Did it find the right passage? Is the answer correct? Does the citation point to the right page? Did it decline when it should have?

Your team reviews the failures and agrees the go-live threshold. Run the same set again after every model or document update, so a drop in quality shows up before users notice it.

Who should be involved?

Operations and knowledge management own the document set and the questions. IT owns access, hosting and integration with the document system. Legal and HR decide which answers need a human to approve them, and which topics are out of scope.

Keep the first pilot narrow. One document set, such as HR policies or procurement procedures, and one team that asks about it every week.

How should you roll it out?

  1. Pick one document set

    A bounded collection with clear owners and frequent questions.

  2. Build the test set

    Real questions, correct answers and source pages, including hard cases.

  3. Pilot with one team

    Real users, real questions, every answer logged.

  4. Measure

    Score against the test set and review failures together.

  5. Expand

    Add documents and teams only after the threshold holds.

The assistant runs on Seekers Core, our stack for connecting to your systems, understanding Arabic content, reasoning over it and acting, all inside a governance layer.

◇ Governance wraps every step

  1. 01ConnectDocuments, ERP, CRM, email
  2. 02UnderstandArabic + English search, cited
  3. 03ReasonPrivate models plan the task
  4. 04ActDraft, file, route, update

For how this applies to ministries and public bodies, see AI for government. For the Arabic details, see Arabic AI.

Where should you start?

Start with a Readiness Sprint of 2 to 3 weeks. We take a sample of your documents and real questions, test OCR and retrieval on them, and agree the test set and go-live threshold with your team.

You leave with a pilot plan for one use case. The pilot itself runs 4 to 8 weeks as part of the Private AI Launchpad.