The 1,200-year-old science of grading who said what — rebuilt as software for AI
ISNAD grades every transmitter in your AI chain — model, retriever, tool, agent — carries the weakest link forward, and seals the result in a signed, tamper-detecting audit record.
When record-keeping obligations land, "show me the evidence" has an answer your auditor can independently inspect.
EU AI Act: high-risk (Annex III) from 2 Dec 2027 · more systems (Annex I) from 2 Aug 2028
weakest link
DAIF · review
one weak transmitter caps the chain
integrity
SHA-256 + HMAC
append-only Merkle log · exportable
DAIF = Arabic for "weak" — a verdict, not a typo.
Apache-2.0 · arXiv:2607.24117 · 1,000+ tests in public CI · benchmarked on 575,060 chains
Verifiable, not vibes — every figure traces to a primary source
Measured on a live LLM stack
κ = 0.575
On RAGTruth (17,790 responses · 6 models), ISNAD's grounding critic catches 97.6% of hallucinations — on the 75.3% that parsed (24.7% excluded, disclosed) — at 79.3% accuracy vs a 71.4% baseline, and grades all six models' reliability in exactly the right order.
Recall ≠ precision — the precision / false-positive sheet is on request. Read the case study ↗
The method, on its home turf
κ = 0.87 (3-way)
Agreement with a rule-based convention derived from Ibn Hajar's 12 narrator tiers — not ground truth — across 575,060 hadith chains. 5-way 0.8667 · lenient 0.761. Narrator-grade agreement κ = 0.33 (published on purpose). Shuffled control ≈ 0.
arXiv:2607.24117
single-author · cs.AI · 25pp
1,000+ tests
public CI, every commit
pip install isnad
Apache-2.0 · Python 3.11+
npm verifier
0.1.x · JS-only
Adapters
LangChain·LangGraph·CrewAI·LlamaIndex·MCP·OTel
Zenodo DOI
10.5281/zenodo.21216873
In plain language: agrees with expert graders 87% of the time across 575,060 chains; on live AI answers, labels them correctly 79.3% vs 71.4% — and tells you the 24.7% it couldn't grade, instead of hiding them.
κ = 0.87 is agreement with a rule-based convention we derived — not ground truth. The number that matters for your stack is κ = 0.575, with limits disclosed. We publish the unflattering parts. That is the whole point.
The problem
An LLM gives you an answer. Who handled it, in what order, and how much do you trust each one? Observability tools record what happened — they don't grade who transformed the claim.
A claim that survives five hand-offs isn't more reliable — it's more obscured. Correlated agents that share one bad source don't cancel the error; they amplify it.
EU AI Act high-risk obligations land 2 Dec 2027 (Annex III) and 2 Aug 2028 (Annex I). The cost of not being ready isn't "we need a log" — it's "we need evidence that survived unchanged, traceable for as long as our policy requires."
How it works
Register every model, tool, corpus, and retriever, and give each a grade — reliable / acceptable / weak / ungraded — on two axes (integrity & precision). Grades are operator-assigned priors — stated openly.
Every claim carries its full chain. ISNAD propagates the weakest link to one of four decisions: serve · caveat · review · quarantine. Corroboration discounts transmitters that share a source.
Every decision produces an AuditRecord: SHA-256 self-hash, HMAC/Ed25519 signature, append-only Merkle log. Export it and prove, later, that nothing changed.
Why ISNAD
Most tools ask about the output — "is this grounded?" or "what happened?" ISNAD asks the question they skip: "who handled this claim, and how much do we trust each transmitter?"
Cleanlab TLM
Grades the trustworthiness of the LLM response — but doesn't record the chain or grade the transmitters.
Galileo · Patronus · TruLens · RAGAS
Measure output faithfulness / groundedness vs context — but skip the chain and the transmitters.
LangSmith · Langfuse · Arize
Record traces (what happened) — but don't grade who handled the claim.
OpenLineage · Marquez · DataHub
Track data lineage (what touched what) — but don't grade trust or gate on it.
ISNAD
Grades the transmitters and the chain — weakest-link grading, madār-discounted corroboration, tamper-detecting audit. The only one that does all three.
Hallucination detectors and tracers are complementary — ISNAD composes with them, it doesn't replace them.
Compliance & security
We say this plainly: ISNAD issues no certificates. Certification belongs to the auditor; producing the record belongs to us. ISNAD produces evidence artefacts, not compliance. Framework references are informational — any "maps onto / supports" statement requires your counsel's sign-off.
Art. 12 → the logging requirement · Art. 13 → the chain · Art. 14 → fail-closed events · Annex IV → technical documentation. Art. 19: providers keep logs at least six months. Retention is set by your policy — self-host keeps it on your infrastructure.
Versioned registry, per-link hashes, signed records — the controlled, documented evidence an audit collects.
Govern/Map provenance & traceability — every output traces back through its chain to its source.
Transparency, accountability, and traceability — a direct implementation of the principles.
KMS/HSM-backed signing keys · 30-day rotation or BYO key · device-bound key ID · keys never stored in chain data.
Revocation is a signed event inside the log — not a rewrite. Tamper-detecting audit + encryption in transit/at rest + admin-action audit.
Sleeper narrators, Sybil, grade-import, tail-truncation, forger limits — documented in the repo ↗. Detects tampering; does not certify truth or compliance.
Who it's for
CISO
You own the blast radius when an agent fabricates a source. ISNAD gives you a signed, hash-chained custody record for every claim — and a compromise playbook for signing keys. You stop answering "where did that come from?" with a shrug.
Head of AI
You ship multi-agent and RAG pipelines and are accountable for what they emit. ISNAD grades the transmitters you already run, flags the weakest link before the claim ships, and treats a model version bump as a new transmitter.
Compliance
You answer the auditor. ISNAD gives you the one-click PDF evidence package — chains, grades, signatures, log proofs — in an open format you can verify offline. Verifiable artifacts, not a vendor's assertion.
Pricing
Today, by bank transfer — email [email protected] for details and an invoice.
$0
Apache-2.0, forever
$5,000
2 weeks · one pipeline
Talk to us
custom · annual
The Trust Assessment is a technical readiness review — not an Art. 43 conformity assessment, not a notified body.
FAQ
ISNAD puts a signed, tamper-detecting label on everything your AI says — showing who handled it and how much to trust each step — so you can prove it to a regulator.
No — and we will not claim it is. ISNAD is evidence infrastructure: it produces the signed, hash-chained records that Art. 12/13/14 and Annex IV ask for. Whether your deployment is compliant is a legal question for your counsel. We give you the evidence to make that argument.
No. ISNAD grades who handled a claim, not whether it is true. Grades are operator-assigned priors. ISNAD makes the provenance, the grades, and the reasoning verifiable and reproducible, so a human can adjudicate. That refusal to fake a confidence number is the brand.
κ = 0.87 is agreement with a rule-based convention derived from Ibn Hajar's 12 narrator tiers, across 575,060 hadith chains — not ground truth. The number that matters for your stack is κ = 0.575 on RAGTruth (97.6% recall, on the 75.3% that parsed; 24.7% disclosed). Both are published with their limits and negative controls.
The engine is Apache-2.0 and free forever. What you pay for is the Trust Assessment (a human maps and grades your pipeline, in two weeks) or the managed cloud: operation, isolation, retention, and evidence. You can run the core yourself; the paid path is for teams that need it done for them.
Yes — the Apache-2.0 core runs from a one-command Docker compose, on your infra, with your data. Your evidence exports in an open format (JSON + HMAC + Merkle proof) you can verify offline, with no ISNAD dependency. That continuity is part of the design.
ISNAD is built by Ali Zahid Raja, building in public (Apache-2.0, arXiv:2607.24117, 1,000+ tests). We lead with that, not hide it: the way we answer procurement questions is published adversarial docs — threat model, failure modes, benchmark controls — and an open repo your security team can inspect. Contact: [email protected].
Today, by bank transfer — email [email protected] for details and an invoice. The fastest start is the Trust Assessment: $5,000, two weeks, credited to your first annual contract.
Plain words
Attach a signed, hash-chained custody record to every claim your AI makes — before your auditor asks.