Open-source provenance for AI
ISNAD grades every transmitter in a claim's chain — source, retriever, model, tool, agent — propagates the weakest link, and seals the result in a signed, tamper-detecting audit record your auditor can inspect.
Annex III scoping is happening now. Art 12 & 19 logging duties bind high-risk systems from 2 Dec 2027 — and six months of logs can't be retrofitted. Coverage begins at install.
Apache-2.0 · arXiv:2607.24117 · 1,000+ tests · 575,060 graded chains
AuditRecord · claim_0x8f3a · isnad v3.0.3
transmission chain · weakest link caps the grade
Grade
DAIF · review
Retention
Art 19 · ≥6 mo
Integrity
SHA-256 + HMAC
DAIF = Arabic for "weak" — a verdict, not a typo. Grade is ordinal trust, not a truth claim.
Verifiable, not vibes — every figure traces to a primary source
Measured on a live LLM stack
κ = 0.575
On RAGTruth (17,790 responses · 6 models), ISNAD flags 97.6% of hallucinations — on the 75.3% that parsed (24.7% excluded, disclosed) — at 79.3% accuracy vs a 71.4% baseline, ranking all six models in the correct order.
Recall ≠ precision — the precision / false-positive sheet is on request. Read the case study ↗
The method, on its home turf
κ = 0.87 (3-way)
Agreement with a rule-based convention derived from Ibn Hajar's 12 narrator tiers — not ground truth — across 575,060 hadith chains. 5-way 0.8667 · lenient 0.761 · narrator-grade agreement κ = 0.33 (published on purpose). Shuffled control ≈ 0.
arXiv:2607.24117
single-author · cs.AI · 25pp
1,000+ tests
public CI, every commit
pip install isnad
v3.0.3 · Python 3.11+
LangChain
community middleware
Adapters
LangGraph·CrewAI·LlamaIndex·MCP·OTel
Zenodo DOI
10.5281/zenodo.21216873
In plain language: agrees with expert graders 87% of the time across 575,060 chains; on live AI answers, labels them correctly 79.3% vs 71.4% — and discloses the 24.7% it couldn't grade instead of hiding them.
The problem
An LLM gives you an answer. Who handled it, in what order, and how much do you trust each one? Observability logs what happened — not who transformed the claim.
A claim that survives five hand-offs isn't more reliable — it's more obscured. Correlated agents that share one bad source amplify the error.
Art 12 & 19 logging duties bind high-risk systems from 2 Dec 2027. Six months of logs can't be backfilled. Coverage begins at install.
How it works
Register every model, tool, corpus and retriever, each with a grade — reliable / acceptable / weak / ungraded — on two axes (integrity & precision). Grades are operator-assigned priors — stated openly.
Every claim carries its full chain. ISNAD propagates the weakest link to one of four decisions: serve · caveat · review · quarantine.
Every decision produces an AuditRecord: SHA-256 self-hash, HMAC/Ed25519 signature, append-only Merkle log. Export it and prove, later, that nothing changed.
Why ISNAD
Most tools ask about the output. ISNAD asks: "who handled this claim, and how much do we trust each transmitter?"
Cleanlab TLM
Grades the trustworthiness of the LLM response — doesn't record the chain or grade the transmitters.
Galileo · Patronus · TruLens · RAGAS
Measure output faithfulness vs context — skip the chain and the transmitters.
LangSmith · Langfuse · Arize
Record traces (what happened) — don't grade who handled the claim.
OpenLineage · Marquez · DataHub
Track data lineage — don't grade trust or gate on it.
ISNAD
Grades the transmitters and the chain — weakest-link grading, madār-discounted corroboration, tamper-detecting audit. The only one that does all three.
Complementary, not a replacement — ISNAD composes with these tools.
Compliance & security
ISNAD issues no certificates. Certification belongs to the auditor; producing the record belongs to us. ISNAD produces evidence artefacts, not compliance. Framework references are informational — any "maps onto / supports" statement requires your counsel's sign-off.
Art. 12 → logging capability (no retention) · Art. 13 → the chain · Art. 14 → fail-closed events · Annex IV → inputs for technical documentation. Art. 19: logs kept ≥6 months. Retention set by your policy — self-host keeps it on your infrastructure.
Versioned registry, per-link hashes, signed records — the controlled, documented evidence an audit collects.
Govern/Map provenance & traceability — every output traces back through its chain to its source.
Transparency, accountability and traceability — a direct implementation of the principles.
KMS/HSM-backed signing keys · 30-day rotation or BYO key · device-bound key ID · keys never stored in chain data.
Revocation is a signed event inside the log — not a rewrite. Tamper-detecting audit + encryption in transit/at rest + admin-action audit.
Sleeper narrators, Sybil, grade-import, tail-truncation, forger limits — documented in the repo ↗.
Signed = developer-held key, not third-party audited. We state exactly what's done vs in progress.
Who it's for
Head of AI / CTO
You ship RAG and agent pipelines and are accountable for what they emit. ISNAD grades the transmitters you already run, flags the weakest link before the claim ships, and treats a model version bump as a new transmitter.
CISO
You own the blast radius when an agent fabricates a source. ISNAD gives you a signed, hash-chained custody record for every claim — and a compromise playbook for signing keys.
Compliance / DPO
You answer the auditor. ISNAD gives you the evidence package — chains, grades, signatures, log proofs — in an open format you can verify offline.
Pricing
Today, by bank transfer — email [email protected] for details and an invoice.
$0
Apache-2.0, forever
$5,000
2 weeks · one pipeline
Talk to us
custom · annual
The Trust Assessment is a technical readiness review — not an Art. 43 conformity assessment, not a notified body.
FAQ
ISNAD puts a signed, tamper-detecting label on everything your AI says — who handled it and how much to trust each step — so you can show an auditor where every answer came from.
No — and we won't claim it is. ISNAD produces the signed, hash-chained evidence supporting Art 12 logging and Art 19 retention. Whether your deployment is compliant is your counsel's call.
No. ISNAD grades who handled a claim, not whether it is true. Grades are operator-assigned priors. Refusing to fake a confidence number is the brand.
κ = 0.87 is agreement with a rule-based convention derived from Ibn Hajar's 12 narrator tiers, across 575,060 hadith chains — not ground truth. The number that matters for your stack is κ = 0.575 on RAGTruth (97.6% recall, on the 75.3% that parsed; 24.7% disclosed).
Yes — the Apache-2.0 core runs on your infra. Your evidence exports in an open format (JSON + HMAC + Merkle proof) you can verify offline. That continuity is part of the design.
ISNAD is built by Ali Zahid Raja, building in public (Apache-2.0, arXiv:2607.24117, 1,000+ tests). We lead with that, not hide it: published threat model, failure modes, benchmark controls. Contact: [email protected].
Plain words
Attach a signed, hash-chained custody record to every claim your AI makes — before your auditor asks.