Open-source provenance for AI

Grade the chain.
Sign the claim.

ISNAD grades every transmitter in a claim's chain — source, retriever, model, tool, agent — propagates the weakest link, and seals the result in a signed, tamper-detecting audit record your auditor can inspect.

Annex III scoping is happening now. Art 12 & 19 logging duties bind high-risk systems from 2 Dec 2027 — and six months of logs can't be retrofitted. Coverage begins at install.

Apache-2.0 · arXiv:2607.24117 · 1,000+ tests · 575,060 graded chains

Sealed Evidence Record

AuditRecord · claim_0x8f3a · isnad v3.0.3

sha256 9f2c1b7e…3a61f0c2e95b8d1f4a7c0e2b9d6f8a13c5e7b2d9f1a4c7

transmission chain · weakest link caps the grade

source → retriever → model → tool → agent · QUARANTINE

Grade

DAIF · review

Retention

Art 19 · ≥6 mo

Integrity

SHA-256 + HMAC

DAIF = Arabic for "weak" — a verdict, not a typo. Grade is ordinal trust, not a truth claim.

Verifiable, not vibes — every figure traces to a primary source

Measured on a live LLM stack

κ = 0.575

On RAGTruth (17,790 responses · 6 models), ISNAD flags 97.6% of hallucinations — on the 75.3% that parsed (24.7% excluded, disclosed) — at 79.3% accuracy vs a 71.4% baseline, ranking all six models in the correct order.

Recall ≠ precision — the precision / false-positive sheet is on request. Read the case study ↗

The method, on its home turf

κ = 0.87 (3-way)

Agreement with a rule-based convention derived from Ibn Hajar's 12 narrator tiers — not ground truth — across 575,060 hadith chains. 5-way 0.8667 · lenient 0.761 · narrator-grade agreement κ = 0.33 (published on purpose). Shuffled control ≈ 0.

arXiv:2607.24117

single-author · cs.AI · 25pp

1,000+ tests

public CI, every commit

pip install isnad

v3.0.3 · Python 3.11+

LangChain

community middleware

Adapters

LangGraph·CrewAI·LlamaIndex·MCP·OTel

Zenodo DOI

10.5281/zenodo.21216873

In plain language: agrees with expert graders 87% of the time across 575,060 chains; on live AI answers, labels them correctly 79.3% vs 71.4% — and discloses the 24.7% it couldn't grade instead of hiding them.

The problem

Your AI makes claims. You can't prove who handled them.

01

Every output is an unprovenanced assertion

An LLM gives you an answer. Who handled it, in what order, and how much do you trust each one? Observability logs what happened — not who transformed the claim.

02

Multi-agent chains obscure, they don't reassure

A claim that survives five hand-offs isn't more reliable — it's more obscured. Correlated agents that share one bad source amplify the error.

03

The deadline can't be retrofitted

Art 12 & 19 logging duties bind high-risk systems from 2 Dec 2027. Six months of logs can't be backfilled. Coverage begins at install.

How it works

Three steps from "who said that?" to signed evidence.

Step 1

Grade your transmitters

Register every model, tool, corpus and retriever, each with a grade — reliable / acceptable / weak / ungraded — on two axes (integrity & precision). Grades are operator-assigned priors — stated openly.

Step 2

Route via the weakest link

Every claim carries its full chain. ISNAD propagates the weakest link to one of four decisions: serve · caveat · review · quarantine.

Step 3

Emit signed, tamper-detecting evidence

Every decision produces an AuditRecord: SHA-256 self-hash, HMAC/Ed25519 signature, append-only Merkle log. Export it and prove, later, that nothing changed.

Fail-closed by default: ambiguous or ungradeable chains are flagged for review, never silently passed. That's the 24.7%.
The hard part is done: 575,060-chain validation + the evidence layer (Merkle + SHA-256 + HMAC + fail-closed signing).
Integrates in ~a day: pip install isnad · Python 3.11+ · one keypair + one log append per claim.

Why ISNAD

Not another hallucination detector.

Most tools ask about the output. ISNAD asks: "who handled this claim, and how much do we trust each transmitter?"

Cleanlab TLM

Grades the trustworthiness of the LLM response — doesn't record the chain or grade the transmitters.

Galileo · Patronus · TruLens · RAGAS

Measure output faithfulness vs context — skip the chain and the transmitters.

LangSmith · Langfuse · Arize

Record traces (what happened) — don't grade who handled the claim.

OpenLineage · Marquez · DataHub

Track data lineage — don't grade trust or gate on it.

ISNAD

Grades the transmitters and the chain — weakest-link grading, madār-discounted corroboration, tamper-detecting audit. The only one that does all three.

Complementary, not a replacement — ISNAD composes with these tools.

Compliance & security

Evidence infrastructure. Not a certification.

ISNAD issues no certificates. Certification belongs to the auditor; producing the record belongs to us. ISNAD produces evidence artefacts, not compliance. Framework references are informational — any "maps onto / supports" statement requires your counsel's sign-off.

EU AI Act

Art. 12 → logging capability (no retention) · Art. 13 → the chain · Art. 14 → fail-closed events · Annex IV → inputs for technical documentation. Art. 19: logs kept ≥6 months. Retention set by your policy — self-host keeps it on your infrastructure.

ISO/IEC 42001

Versioned registry, per-link hashes, signed records — the controlled, documented evidence an audit collects.

NIST AI RMF

Govern/Map provenance & traceability — every output traces back through its chain to its source.

SDAIA

Transparency, accountability and traceability — a direct implementation of the principles.

Key custody

KMS/HSM-backed signing keys · 30-day rotation or BYO key · device-bound key ID · keys never stored in chain data.

Compromise playbook

Revocation is a signed event inside the log — not a rewrite. Tamper-detecting audit + encryption in transit/at rest + admin-action audit.

Published threat model

Sleeper narrators, Sybil, grade-import, tail-truncation, forger limits — documented in the repo ↗.

Signed = developer-held key, not third-party audited. We state exactly what's done vs in progress.

Who it's for

Built for the people who sign off on AI.

Head of AI / CTO

The builder who signs

You ship RAG and agent pipelines and are accountable for what they emit. ISNAD grades the transmitters you already run, flags the weakest link before the claim ships, and treats a model version bump as a new transmitter.

CISO

Chief Information Security Officer

You own the blast radius when an agent fabricates a source. ISNAD gives you a signed, hash-chained custody record for every claim — and a compromise playbook for signing keys.

Compliance / DPO

The person who answers the auditor

You answer the auditor. ISNAD gives you the evidence package — chains, grades, signatures, log proofs — in an open format you can verify offline.

Pricing

Start free. Pay for the evidence you need.

Today, by bank transfer — email [email protected] for details and an invoice.

Self-host

$0

Apache-2.0, forever

  • pip install isnad
  • Your infra, your data
  • Verify offline, leave anytime
  • Community support
Run the open source ↗
Most popular

Trust Assessment

$5,000

2 weeks · one pipeline

  • Map your pipeline & grade transmitters
  • Measured accuracy report
  • Signed evidence package
  • Credited to your first annual contract
Email to book ↗

Enterprise

Talk to us

custom · annual

  • Managed cloud (hosted for you)
  • VPC / on-prem / private cloud
  • Retention set by your policy
  • Custom DPA & terms · SLA
Email to talk ↗

The Trust Assessment is a technical readiness review — not an Art. 43 conformity assessment, not a notified body.

FAQ

Honest answers to the hard questions.

What is ISNAD, in one sentence?

ISNAD puts a signed, tamper-detecting label on everything your AI says — who handled it and how much to trust each step — so you can show an auditor where every answer came from.

Is ISNAD "EU AI Act compliant"?

No — and we won't claim it is. ISNAD produces the signed, hash-chained evidence supporting Art 12 logging and Art 19 retention. Whether your deployment is compliant is your counsel's call.

Does ISNAD tell me whether a claim is true?

No. ISNAD grades who handled a claim, not whether it is true. Grades are operator-assigned priors. Refusing to fake a confidence number is the brand.

Where do the benchmark numbers come from?

κ = 0.87 is agreement with a rule-based convention derived from Ibn Hajar's 12 narrator tiers, across 575,060 hadith chains — not ground truth. The number that matters for your stack is κ = 0.575 on RAGTruth (97.6% recall, on the 75.3% that parsed; 24.7% disclosed).

Can I self-host? What if you disappear?

Yes — the Apache-2.0 core runs on your infra. Your evidence exports in an open format (JSON + HMAC + Merkle proof) you can verify offline. That continuity is part of the design.

Who is behind this?

ISNAD is built by Ali Zahid Raja, building in public (Apache-2.0, arXiv:2607.24117, 1,000+ tests). We lead with that, not hide it: published threat model, failure modes, benchmark controls. Contact: [email protected].

Plain words

The terms, in one line each.

ISNAD — Arabic for "chain of narration." The graded chain your claim traveled.
Transmitter — anything that handled the claim: a model, tool, retriever, corpus, or agent.
κ (kappa) — agreement score: 1 = perfect, 0 = chance. Reported with limits.
DAIF — Arabic for "weak." A chain verdict, not a typo.
HMAC / Ed25519 — the cryptographic signatures that seal each record.
Merkle log — an append-only ledger: if anyone edits history, the hashes break.
Weakest link — the least-trusted step caps the whole chain.
Retriever — the component that fetches source documents for a RAG pipeline.
Tamper-detecting — the record detects tampering; it does not certify that a claim is true.

Prove who handled it.

Attach a signed, hash-chained custody record to every claim your AI makes — before your auditor asks.

[email protected]