Insights 6 min read
How to build an Arabic knowledge assistant your staff will actually trust
Why Arabic search fails, how retrieval with cited sources fixes it, and how to measure accuracy on your own documents before a single employee relies on it.
Staff trust an Arabic knowledge assistant when they can check every answer. That means it answers only from your documents, shows the clause and page it used, says "I don't know" when the documents are silent, and only shows people what they are already allowed to see. You prove all of this on your own questions before go-live.
Why does staff search fail on Arabic documents?
Most enterprise search was built for English. It assumes a word has one spelling and clear edges. Arabic breaks both.
The same word appears as أ, إ or ا, with ة or ه at the end. Prefixes attach directly, so وبالعقد and العقود rarely match. Then real content adds its own problems.
Spelling variants
أحمد · احمد · إحمد
Mixed languages
تم approve الـ invoice
Franco-Arabic
3ayez a3raf el mo3ad
Dialects
أبغى · عايز · بدي
Scanned PDFs
OCR on stamped pages
RTL tables
Columns read right to left
Mixed language is common. A staff member types "ما هو الـ SLA في عقد الصيانة؟" and expects an English clause from an Arabic contract. A customer writes in Franco-Arabic, such as "3ayez a3raf el ma3ad". Much of the archive is scanned PDFs, and the fee tables read right to left. Keyword search misses most of this, and people stop using it.
What does a knowledge assistant do differently?
It searches by meaning and by Arabic-aware keywords at the same time. Then a private language model writes a short answer from the passages it found, and only from those.
- Scanned PDFعقد التوريد
- Read (OCR)Text and tables extracted
- Normaliseأ، إ، آ become ا
- SearchFinds clause 7.2
- Answer + sourceCited · page 14
Scanned files go through Arabic OCR first. Text is normalised for search, but the original wording is kept for display. Every step runs inside your environment: on-premise, private cloud or a sovereign GCC cloud.
| Keyword search | Arabic RAG assistant | |
|---|---|---|
| Spelling variants | Misses أ / إ / ا and ة / ه | Normalised before search |
| Mixed Arabic and English | Matches one language at a time | Finds passages in either language |
| Scanned PDFs | Often not searchable | Arabic OCR, page numbers kept |
| What you get | A list of files to read | A short answer with its source |
| When nothing matches | Empty or irrelevant results | Says it cannot find an answer |
What makes staff trust the answers?
Staff trust what they can check. A few rules make that possible, and users see them working every day.
Does every answer cite its source?
Yes. The answer links to the document, page and passage. The user opens it, sees the highlighted text and checks it before acting.
ما هي غرامة التأخير في عقد التوريد؟
0.5% of the order value per day of delay, capped at 10%.
Supply contract.pdf · Clause 7.2 · page 14
«تُفرض غرامة تأخير قدرها 0.5% عن كل يوم…»
Every answer shows its source
What happens when the documents don't cover it?
The assistant says "I don't know" and lists the closest documents. Staff can work with a clear "not found". A fluent guess sends them the wrong way.
Who sees what?
Retrieval is filtered by what each user can already open. An HR assistant does not show a salary policy draft to someone who could not see it in the document system.
Governance is part of the build. Every question, retrieved passage and answer is logged for audit. Anything sensitive, such as a reply to a citizen or a legal opinion, goes to a person for approval before it is sent.
- AI draftsReply, memo, or update
- Human approvesThe right role signs off
- Action runsIn your systems
- LoggedWho, what, when
How do you know it is accurate before go-live?
You test it on your own questions. A generic benchmark tells you little about your policies, your scans or the way your staff write.
- Real questionsWritten with your team
- Run the systemSame questions, every release
- Score each answerCorrect, cited, safe
- Go / no-goAgreed threshold
Build a test set with the teams who will use it. Collect real questions in Arabic, English and both mixed together. Write down the correct answer and the page it comes from. Add hard cases on purpose: dialect, typos, poor scans, and questions the documents cannot answer.
Run the assistant against the set and score four things. Did it find the right passage? Is the answer correct? Does the citation point to the right page? Did it decline when it should have?
Your team reviews the failures and agrees the go-live threshold. Run the same set again after every model or document update, so a drop in quality shows up before users notice it.
Who should be involved?
Operations and knowledge management own the document set and the questions. IT owns access, hosting and integration with the document system. Legal and HR decide which answers need a human to approve them, and which topics are out of scope.
Keep the first pilot narrow. One document set, such as HR policies or procurement procedures, and one team that asks about it every week.
How should you roll it out?
Pick one document set
A bounded collection with clear owners and frequent questions.
Build the test set
Real questions, correct answers and source pages, including hard cases.
Pilot with one team
Real users, real questions, every answer logged.
Measure
Score against the test set and review failures together.
Expand
Add documents and teams only after the threshold holds.
The assistant runs on Seekers Core, our stack for connecting to your systems, understanding Arabic content, reasoning over it and acting, all inside a governance layer.
◇ Governance wraps every step
- 01ConnectDocuments, ERP, CRM, email
- 02UnderstandArabic + English search, cited
- 03ReasonPrivate models plan the task
- 04ActDraft, file, route, update
For how this applies to ministries and public bodies, see AI for government. For the Arabic details, see Arabic AI.
Where should you start?
Start with a Readiness Sprint of 2 to 3 weeks. We take a sample of your documents and real questions, test OCR and retrieval on them, and agree the test set and go-live threshold with your team.
You leave with a pilot plan for one use case. The pilot itself runs 4 to 8 weeks as part of the Private AI Launchpad.