ISNAD

The 1,200-year-old science of grading who said what — rebuilt as software for AI

Every AI claim, graded and signed — from source to output.

ISNAD grades every transmitter in your AI chain — model, retriever, tool, agent — carries the weakest link forward, and seals the result in a signed, tamper-detecting audit record.

When record-keeping obligations land, "show me the evidence" has an answer your auditor can independently inspect.

EU AI Act: high-risk (Annex III) from 2 Dec 2027 · more systems (Annex I) from 2 Aug 2028

AuditRecord · claim_0x8f3a signed
chain: source retriever model · weak claim

weakest link

DAIF · review

one weak transmitter caps the chain

integrity

SHA-256 + HMAC

append-only Merkle log · exportable

DAIF = Arabic for "weak" — a verdict, not a typo.

Apache-2.0 · arXiv:2607.24117 · 1,000+ tests in public CI · benchmarked on 575,060 chains

Verifiable, not vibes — every figure traces to a primary source

Measured on a live LLM stack

κ = 0.575

On RAGTruth (17,790 responses · 6 models), ISNAD's grounding critic catches 97.6% of hallucinations — on the 75.3% that parsed (24.7% excluded, disclosed) — at 79.3% accuracy vs a 71.4% baseline, and grades all six models' reliability in exactly the right order.

Recall ≠ precision — the precision / false-positive sheet is on request. Read the case study ↗

The method, on its home turf

κ = 0.87 (3-way)

Agreement with a rule-based convention derived from Ibn Hajar's 12 narrator tiers — not ground truth — across 575,060 hadith chains. 5-way 0.8667 · lenient 0.761. Narrator-grade agreement κ = 0.33 (published on purpose). Shuffled control ≈ 0.

arXiv:2607.24117

single-author · cs.AI · 25pp

1,000+ tests

public CI, every commit

pip install isnad

Apache-2.0 · Python 3.11+

npm verifier

0.1.x · JS-only

Adapters

LangChain·LangGraph·CrewAI·LlamaIndex·MCP·OTel

Zenodo DOI

10.5281/zenodo.21216873

In plain language: agrees with expert graders 87% of the time across 575,060 chains; on live AI answers, labels them correctly 79.3% vs 71.4% — and tells you the 24.7% it couldn't grade, instead of hiding them.

κ = 0.87 is agreement with a rule-based convention we derived — not ground truth. The number that matters for your stack is κ = 0.575, with limits disclosed. We publish the unflattering parts. That is the whole point.

The problem

Your AI makes claims. You can't prove who handled them.

01

Every output is an unprovenanced assertion

An LLM gives you an answer. Who handled it, in what order, and how much do you trust each one? Observability tools record what happened — they don't grade who transformed the claim.

02

Multi-agent chains obscure, they don't reassure

A claim that survives five hand-offs isn't more reliable — it's more obscured. Correlated agents that share one bad source don't cancel the error; they amplify it.

03

The deadline is real

EU AI Act high-risk obligations land 2 Dec 2027 (Annex III) and 2 Aug 2028 (Annex I). The cost of not being ready isn't "we need a log" — it's "we need evidence that survived unchanged, traceable for as long as our policy requires."

How it works

Three steps from "who said that?" to signed evidence.

Step 1

Grade your transmitters

Register every model, tool, corpus, and retriever, and give each a grade — reliable / acceptable / weak / ungraded — on two axes (integrity & precision). Grades are operator-assigned priors — stated openly.

Step 2

Route via the weakest link

Every claim carries its full chain. ISNAD propagates the weakest link to one of four decisions: serve · caveat · review · quarantine. Corroboration discounts transmitters that share a source.

Step 3

Emit signed, tamper-detecting evidence

Every decision produces an AuditRecord: SHA-256 self-hash, HMAC/Ed25519 signature, append-only Merkle log. Export it and prove, later, that nothing changed.

Fail-closed by default: ambiguous or ungradeable chains are flagged for review, never silently passed. That's the 24.7%.
The hard part is done: 575,060-chain validation + the evidence layer (Merkle + SHA-256 + HMAC + fail-closed signing).
Integrates in ~a day: pip install isnad · Python 3.11+ · one keypair + one log append per claim.

Why ISNAD

Not another hallucination detector.

Most tools ask about the output — "is this grounded?" or "what happened?" ISNAD asks the question they skip: "who handled this claim, and how much do we trust each transmitter?"

Cleanlab TLM

Grades the trustworthiness of the LLM response — but doesn't record the chain or grade the transmitters.

Galileo · Patronus · TruLens · RAGAS

Measure output faithfulness / groundedness vs context — but skip the chain and the transmitters.

LangSmith · Langfuse · Arize

Record traces (what happened) — but don't grade who handled the claim.

OpenLineage · Marquez · DataHub

Track data lineage (what touched what) — but don't grade trust or gate on it.

ISNAD

Grades the transmitters and the chain — weakest-link grading, madār-discounted corroboration, tamper-detecting audit. The only one that does all three.

Hallucination detectors and tracers are complementary — ISNAD composes with them, it doesn't replace them.

Compliance & security

Evidence infrastructure. Not a certification.

We say this plainly: ISNAD issues no certificates. Certification belongs to the auditor; producing the record belongs to us. ISNAD produces evidence artefacts, not compliance. Framework references are informational — any "maps onto / supports" statement requires your counsel's sign-off.

EU AI Act

Art. 12 → the logging requirement · Art. 13 → the chain · Art. 14 → fail-closed events · Annex IV → technical documentation. Art. 19: providers keep logs at least six months. Retention is set by your policy — self-host keeps it on your infrastructure.

ISO/IEC 42001

Versioned registry, per-link hashes, signed records — the controlled, documented evidence an audit collects.

NIST AI RMF

Govern/Map provenance & traceability — every output traces back through its chain to its source.

SDAIA

Transparency, accountability, and traceability — a direct implementation of the principles.

Key custody

KMS/HSM-backed signing keys · 30-day rotation or BYO key · device-bound key ID · keys never stored in chain data.

Compromise playbook

Revocation is a signed event inside the log — not a rewrite. Tamper-detecting audit + encryption in transit/at rest + admin-action audit.

Published threat model

Sleeper narrators, Sybil, grade-import, tail-truncation, forger limits — documented in the repo ↗. Detects tampering; does not certify truth or compliance.

Who it's for

Built for the people who sign off on AI.

CISO

Chief Information Security Officer

You own the blast radius when an agent fabricates a source. ISNAD gives you a signed, hash-chained custody record for every claim — and a compromise playbook for signing keys. You stop answering "where did that come from?" with a shrug.

Head of AI

Head of AI / Platform Lead

You ship multi-agent and RAG pipelines and are accountable for what they emit. ISNAD grades the transmitters you already run, flags the weakest link before the claim ships, and treats a model version bump as a new transmitter.

Compliance

Compliance / DPO

You answer the auditor. ISNAD gives you the one-click PDF evidence package — chains, grades, signatures, log proofs — in an open format you can verify offline. Verifiable artifacts, not a vendor's assertion.

Pricing

Start free. Pay for the evidence you need.

Today, by bank transfer — email [email protected] for details and an invoice.

Self-host

$0

Apache-2.0, forever

  • pip install isnad
  • Your infra, your data
  • Verify offline, leave anytime
  • Community support
Run the open source ↗
Most popular

Trust Assessment

$5,000

2 weeks · one pipeline

  • Map your pipeline & grade transmitters
  • Measured accuracy report
  • Signed evidence package
  • Credited to your first annual contract
Email to book ↗

Enterprise

Talk to us

custom · annual

  • Managed cloud (hosted for you)
  • VPC / on-prem / private cloud
  • Retention set by your policy
  • Custom DPA & terms · SLA
Email to talk ↗

The Trust Assessment is a technical readiness review — not an Art. 43 conformity assessment, not a notified body.

FAQ

Honest answers to the hard questions.

What is ISNAD, in one sentence?

ISNAD puts a signed, tamper-detecting label on everything your AI says — showing who handled it and how much to trust each step — so you can prove it to a regulator.

Is ISNAD "EU AI Act compliant"?

No — and we will not claim it is. ISNAD is evidence infrastructure: it produces the signed, hash-chained records that Art. 12/13/14 and Annex IV ask for. Whether your deployment is compliant is a legal question for your counsel. We give you the evidence to make that argument.

Does ISNAD tell me whether a claim is true?

No. ISNAD grades who handled a claim, not whether it is true. Grades are operator-assigned priors. ISNAD makes the provenance, the grades, and the reasoning verifiable and reproducible, so a human can adjudicate. That refusal to fake a confidence number is the brand.

Where do the benchmark numbers come from?

κ = 0.87 is agreement with a rule-based convention derived from Ibn Hajar's 12 narrator tiers, across 575,060 hadith chains — not ground truth. The number that matters for your stack is κ = 0.575 on RAGTruth (97.6% recall, on the 75.3% that parsed; 24.7% disclosed). Both are published with their limits and negative controls.

Why pay for something that's open source?

The engine is Apache-2.0 and free forever. What you pay for is the Trust Assessment (a human maps and grades your pipeline, in two weeks) or the managed cloud: operation, isolation, retention, and evidence. You can run the core yourself; the paid path is for teams that need it done for them.

Can I self-host? What if you disappear?

Yes — the Apache-2.0 core runs from a one-command Docker compose, on your infra, with your data. Your evidence exports in an open format (JSON + HMAC + Merkle proof) you can verify offline, with no ISNAD dependency. That continuity is part of the design.

Who is behind this?

ISNAD is built by Ali Zahid Raja, building in public (Apache-2.0, arXiv:2607.24117, 1,000+ tests). We lead with that, not hide it: the way we answer procurement questions is published adversarial docs — threat model, failure modes, benchmark controls — and an open repo your security team can inspect. Contact: [email protected].

How do I pay, and how do I start?

Today, by bank transfer — email [email protected] for details and an invoice. The fastest start is the Trust Assessment: $5,000, two weeks, credited to your first annual contract.

Plain words

The terms, in one line each.

ISNAD — Arabic for "chain of narration." The graded chain your claim traveled.
Transmitter — anything that handled the claim: a model, tool, retriever, corpus, or agent.
κ (kappa) — agreement score: 1 = perfect, 0 = chance. We report ours with limits.
DAIF — Arabic for "weak." A chain verdict, not a typo.
HMAC / Ed25519 — the cryptographic signatures that seal each record.
Merkle log — an append-only ledger: if anyone edits history, the hashes break.
Weakest link — the least-trusted step caps the whole chain. One bad hop is one too many.
Retriever — the component that fetches source documents for a RAG pipeline.
Tamper-detecting — the record detects tampering; it does not certify that a claim is true.

Prove who handled it.

Attach a signed, hash-chained custody record to every claim your AI makes — before your auditor asks.

[email protected]