How Should We Evaluate Islamic AI? The Benchmark Problem
We have general Arabic NLP benchmarks. We do not have a serious benchmark for Islamic jurisprudential AI. Here is what one would actually need to measure — and why it matters more than accuracy scores.
The Problem with Existing Benchmarks
The standard way to evaluate an Arabic-language AI model is to run it against established benchmarks: BALSAM, ArabicMMLU, Arabic language understanding evaluations from major NLP labs. These measure things like reading comprehension, translation quality, factual recall, and reasoning ability in Arabic.
For general Arabic NLP tasks, these benchmarks are appropriate. For Islamic jurisprudential AI, they are almost entirely useless.
The reason: the failure modes of Islamic AI are not measured by general language benchmarks. An AI can score perfectly on Arabic reading comprehension while simultaneously fabricating hadith with confident-sounding isnāds, misrepresenting the Ḥanbalī position on a legal question, or applying uṣūl al-fiqh principles in the wrong sequence. None of these failures would show up in a standard NLP benchmark.
A 2026 review paper by Dr. Muhammad Ali Al-Badri makes this point explicitly: Islamic AI needs a specialist benchmark, not a general one — and that benchmark needs to be calibrated against authoritative sources, not crowdsourced internet data or general Arabic corpora.
What a Real Islamic AI Benchmark Would Measure
1. Ground Truth Alignment
Does the AI's output match what the authoritative scholarly sources actually say?
For a Jaʿfarī fiqh benchmark, this means calibrating against the written opinions of recognised Grand Marājiʿ — al-Sīstānī, al-Khāmeneī, the published fatwās of Sayyid Faḍlallāh. These are the "ground truth" documents — not internet forums, not general Islamic Q&A sites, but the authenticated legal positions of recognised authorities.
For a Sunnī fiqh benchmark, the equivalent would be the canonical texts of the four schools — the authoritative Ḥanafī, Mālikī, Shāfiʿī, and Ḥanbalī manuals — as interpreted through recognised contemporary scholarly consensus, not any individual modern opinion.
The key requirement: the evaluator must be a trained jurist, not an automated NLP metric. A system can produce text that scores well on fluency metrics while being substantively wrong about the legal position. Only someone who knows the tradition can catch that.
2. Madhab Faithfulness (the Anti-Confusion Metric)
Islamic jurisprudence has four main Sunnī schools plus the Jaʿfarī tradition. Each school has distinct positions on hundreds of questions. A major failure mode of Islamic AI is the "school collapse" problem — blending positions from different schools into a single answer that does not accurately represent any of them.
An Islamic AI benchmark needs to specifically test:
- Does the model correctly attribute each position to its school?
- Does it avoid presenting the majority position as unanimous consensus when there is genuine disagreement?
- Does it correctly handle questions where the schools diverge sharply (e.g., the permissibility of gelatin, the validity of certain financial instruments, the ʿidda period in certain circumstances)?
This requires generating test questions with known school-specific answers and evaluating whether the model correctly differentiates between them.
3. Uṣūlī Chain of Thought Accuracy
Islamic legal reasoning follows a structured methodology (uṣūl al-fiqh). A ruling is not just an output — it is the conclusion of a reasoning process: identify the relevant sources, apply the principles of tarjīḥ (preference when sources conflict), apply the relevant uṣūlī tools (istiṣḥāb, barāʾa, takhyīr), and reach a conclusion.
A benchmark should evaluate whether the AI's reasoning process follows this structure correctly — not just whether it produces the right answer. A model can arrive at a correct ruling via incorrect reasoning, which is a problem if practitioners use the reasoning to understand why the ruling applies to their situation.
Specific things to check: does the model correctly apply istiṣḥāb (continuing a previous state of affairs in the absence of evidence of change)? Does it correctly identify when multiple rulings might apply and explain the tarjīḥ? Does it correctly distinguish between obligatory and precautionary rulings?
4. Level 3 Awareness (Knowing When to Decline)
Perhaps the most important metric for safety: can the model correctly identify when a question is outside its scope and requires human scholarly judgment?
Questions that require human judgment include:
- Questions involving extreme psychological distress or spiritual crisis (the questioner's ḥāl matters)
- Novel contemporary questions where the scholarly consensus has not formed
- Questions where the answer depends on local customary practice (ʿurf) that the model cannot assess
- Questions that involve personal circumstance the model cannot evaluate
A well-calibrated Islamic AI should, when faced with these questions, not attempt an answer. It should say: "This question requires a qualified scholar who can assess your specific circumstances. Here is what the relevant sources say about the general principle — but the ruling as it applies to you requires a human jurist."
An Islamic AI benchmark should include a set of questions specifically designed to test whether the model recognises and correctly responds to these limits. Models that attempt confident answers to out-of-scope questions fail this metric even if their answers happen to be correct, because the correct answer to those questions is "seek a scholar."
Why 3arif.ai Takes This Seriously
3arif.ai is not yet benchmarked against a specialist Islamic AI evaluation. No such standardised benchmark currently exists — Dr. Al-Badri's paper is partly a call to build one.
But the design of 3arif.ai reflects the same principles a good benchmark would test:
Ground truth alignment: Every response is drawn from authenticated primary sources — not internet data. The ground truth is the retrieved text, and the quality gate checks that the response stays within it.
Madhab faithfulness: The Shariah and In Depth modes are explicitly designed to present the positions of each school separately, with attribution, rather than synthesising a single "Islamic position."
Uṣūlī reasoning: The system prompt for Shariah mode asks the model to identify the dalīl (evidence), state the mainstream position, and show where the schools diverge — approximating the structure of jurisprudential reasoning.
Level 3 awareness: No mode will issue a fatwa. The system prompt prohibits it. Responses include explicit guidance to "consult a qualified scholar for personal rulings" on questions that require it.
Whether these design choices translate into reliable performance across the full range of Islamic legal questions is something we intend to evaluate rigorously as specialist benchmarks develop. We welcome collaboration with Islamic scholars and AI researchers on building the evaluation infrastructure the field needs.
What Needs to Be Built
The field needs:
-
A curated question set — hundreds of jurisprudential questions with verified answers drawn from authoritative school-specific sources. Maintained by trained jurists, not crowdsourced.
-
A trained evaluator panel — jurists from each tradition who can assess not just whether the answer is "correct" but whether the reasoning is valid and the limitations appropriately acknowledged.
-
Automated checks for structural failure modes — school collapse detection, isnād fabrication detection, citation hallucination detection.
-
An ongoing maintenance process — as scholarly consensus evolves on contemporary questions, the benchmark needs to evolve too. A static benchmark becomes misleading.
Until this infrastructure exists, claims about Islamic AI accuracy should be treated with appropriate scepticism — including claims made about 3arif.ai. What we can say with confidence is that our approach makes fabrication structurally much harder, and that every claim is traceable to a verified source. What we cannot yet say is how we perform against a comprehensive specialist evaluation. Building that evaluation is the next frontier.
If you are a researcher in Islamic studies or NLP and are interested in collaboration on Islamic AI evaluation, we would welcome the conversation.