Chain of Thought Meets Uṣūl al-Fiqh: Teaching AI to Reason Like a Jurist
The classical Islamic methodology for deriving legal rulings — uṣūl al-fiqh — is one of the most sophisticated reasoning frameworks ever developed. Modern AI's Chain of Thought technique maps onto it more closely than you might expect.
Two Reasoning Frameworks Separated by Twelve Centuries
In 820 CE, Imam al-Shāfiʿī completed Al-Risāla — the founding document of Islamic legal methodology. In it, he laid out a systematic framework for how a jurist should reason from the Quran and Sunnah to a legal ruling. Every step of the process was made explicit: what counts as evidence, how to weigh conflicting evidence, what to do when the primary sources are silent, and when analogical reasoning can extend a ruling to a new case.
In 2022 CE, researchers at Google published "Chain of Thought Prompting Elicits Reasoning in Large Language Models." The paper showed that AI models perform dramatically better on complex reasoning tasks when prompted to make their reasoning steps explicit — rather than jumping from question to answer.
These two developments are separated by twelve centuries and two entirely different intellectual traditions. They are also solving recognisably similar problems.
What Uṣūl al-Fiqh Actually Is
Uṣūl al-fiqh is the methodology of Islamic law — the rules about how to derive rules. It is, in a sense, the jurisprudence of jurisprudence.
A practising jurist (faqīh) who receives a question about a contemporary matter — say, whether a specific financial instrument is permissible — does not simply consult memory. They follow a structured process:
Step 1 — Source identification. Does the Quran speak directly to this? Does the Sunnah? What have earlier scholars said? Which ḥadīth collections contain relevant narrations, and what are their grades?
Step 2 — Tarjīḥ (weighing and preferring). When sources appear to conflict, how are they reconciled? A Quranic verse takes precedence over a ḥadīth. A mutawātir (multiply-transmitted) ḥadīth takes precedence over an āḥād (singly-transmitted) one. A specific ruling takes precedence over a general one.
Step 3 — Qiyās (analogical extension). If the new case is not directly addressed, is there an established case that shares the same underlying cause (ʿilla)? The ruling on wine establishes a principle (intoxication is prohibited); qiyās extends it to all intoxicants.
Step 4 — Uṣūlī tools for uncertainty. When the evidence is genuinely unclear, specific tools apply. Istiṣḥāb (continuation): assume the previous state of affairs continues until proven otherwise. Barāʾa aṣliyya (original permissibility): assume something is permitted unless specific evidence prohibits it. Iḥtiyāṭ (precaution): when consequences are severe, err on the side of caution.
Step 5 — Conclusion with disclosure. State the ruling, identify which school's methodology produced it, and flag where scholars genuinely disagree.
This is not a set of arbitrary steps. It is a rigorous framework developed over centuries to ensure that legal rulings are traceable, reproducible, and intellectually honest about their evidential basis.
What Chain of Thought Is
Chain of Thought (CoT) prompting is a technique in which an AI model is prompted to work through a problem step by step rather than producing a direct answer. The 2022 Google paper showed that for complex reasoning tasks — mathematics, logical puzzles, multi-step word problems — models that showed their work dramatically outperformed models that gave direct answers.
The intuition: when a model is forced to articulate intermediate reasoning steps, it is more likely to catch errors, apply relevant constraints, and produce logically consistent conclusions. The explicit reasoning trace also makes it possible to identify where a model went wrong when it does.
In simple applications, CoT looks like: "Let me think through this step by step. First... Second... Therefore..."
The Convergence
The parallel between uṣūl al-fiqh and Chain of Thought is not superficial. Both frameworks are solutions to the same underlying problem: how do you ensure that a complex reasoning process produces reliable, traceable, auditable conclusions?
Both require:
- Making intermediate steps explicit rather than jumping to conclusions
- Applying consistent rules for weighting evidence
- Having explicit procedures for handling uncertainty and conflict
- Producing conclusions that can be traced back to their premises
The difference is that uṣūl al-fiqh was developed by human scholars to constrain and guide human jurists — and has been refined over twelve centuries of application to real cases. Chain of Thought is a technique for constraining and guiding AI language models, developed over the last few years and still maturing.
A February 2026 paper by Dr. Muhammad Ali Al-Badri proposes something interesting: that the correct prompt architecture for Islamic jurisprudential AI should mirror the uṣūlī reasoning structure explicitly. Rather than asking a model "is X permissible?", the prompt should walk the model through the uṣūlī steps: "First identify the relevant Quranic evidence. Then identify the relevant hadith with grades. Then identify how earlier scholars have ruled. Then apply the relevant uṣūlī principles. Then state the conclusion with disclosure of any scholarly disagreement."
This is precisely what 3arif.ai's Shariah mode does.
How 3arif.ai Implements This
The system prompt for Shariah mode structures the response as a series of explicit steps that mirror uṣūlī methodology:
- Ruling — state the mainstream position clearly
- Evidence — identify the Quranic verse or hadith that establishes the ruling (this is the dalīl — the proof)
- Madhab positions — show where the four Sunnī schools and the Jaʿfarī tradition differ, with attribution
- Notes — flag important conditions, exceptions, or scholarly qualifications
The retrieval system ensures that the "evidence" step draws from authenticated hadith with verified grades. The quality gate checks that every cited source appears in the retrieved context.
This is Chain of Thought applied to Islamic legal reasoning — and the "chain" in question is structured to mirror the chain of uṣūlī methodology.
The Multi-Agent Extension
The most sophisticated application of this idea is the multi-agent debate architecture proposed in Al-Badri's paper.
Rather than a single model working through the uṣūlī steps, imagine specialised agents for each step:
- Agent A (Isnad and Hadith): searches the hadith databases, evaluates chains, produces a ranked list of relevant narrations with grades
- Agent B (Uṣūlī principles): applies the relevant principles — identifying when qiyās is appropriate, which uṣūlī tool applies under uncertainty
- Agent C (Tafsīr and ethics): reviews the relevant Quranic verses with tafsīr commentary and the broader ethical framework
- Agent D (Rational coherence): checks that the conclusion is logically consistent with the evidence and does not contradict established scholarly consensus
- Coordinating agent: synthesises the outputs, identifies any conflicts between agents, and produces a structured summary for human review
This architecture does something powerful: it externalises the internal deliberation that a trained jurist conducts mentally. Each step is executed by a specialised system and the outputs are visible, auditable, and subject to review.
The critical caveat — and this is where Al-Badri is emphatic — is that this architecture produces input for a human jurist, not a ruling itself. The coordinating agent's output is research material. The human jurist reads it, exercises independent judgment, and issues the ruling. Without that final step, the system has crossed from Level 2 to Level 3 — from research assistant to authority claimant.
What Remains Genuinely Hard
The parallel between CoT and uṣūl al-fiqh is illuminating, but it has real limits:
Context-dependence. A trained jurist assessing a question about, say, a financial transaction also considers the questioner's circumstances: their financial state, the local economic context, the customary practices of their community. Uṣūl al-fiqh operates on universal principles but applies them to particular human situations. AI models, working from text alone, cannot assess particular situations.
The ḥāl (condition) of the questioner. Classical jurists were trained to assess not just the legal question but the psychological and spiritual state of the person asking. Certain answers might be technically valid but spiritually harmful for a person in a specific condition. This requires the kind of contextual human wisdom that no prompt architecture can substitute for.
Ijtihād. For genuinely novel questions where the classical sources provide no clear guidance, a mujtahid (a jurist qualified to exercise independent legal reasoning) must apply the full weight of their training, character, and judgment. This is not a reasoning process that can be replicated by structured prompting. It is an exercise of scholarly authority that requires scholarly formation.
These limits are not reasons to avoid using AI for Islamic scholarship. They are reasons to be clear about where AI ends and human scholarship begins — and to build systems that are transparent about that boundary rather than papering over it.
The Twelve-Century Head Start
Islamic legal methodology is one of the most sophisticated reasoning frameworks in human intellectual history. Twelve centuries of brilliant scholars applied, debated, refined, and extended it — working through thousands of real cases, developing increasingly precise tools for handling uncertainty, and building an epistemological tradition that takes the traceability of conclusions with extraordinary seriousness.
AI reasoning research is approximately five years old.
This asymmetry suggests an obvious strategy: rather than building Islamic AI from scratch using only modern machine learning techniques, the most promising direction is to encode the uṣūlī framework explicitly into the prompt architecture and evaluation criteria.
3arif.ai is a step in this direction. The structured Shariah mode prompt, the retrieval-first architecture, the citation requirements, the quality gate — all of these are attempts to build a system whose reasoning process respects the epistemological standards of the tradition it is drawing on.
It is, at best, a beginning. But it is a beginning that starts from the right place: taking the classical methodology seriously as a framework for thinking about how AI should reason about Islamic law — rather than treating Islamic scholarship as just another corpus of text to be statistically modelled.
The Shariah mode (⚖️) in 3arif.ai attempts to implement uṣūlī reasoning structure in its responses. Research mode (🔬) applies the full depth of classical scholarship to complex questions. Both are research tools — not replacements for qualified scholarly guidance.