Validated a trust-propagation rule against 577k real transmission chains - κ 0.871 vs 0.331 between the human experts themselves [P]
3/10Short version: I built a claim-level provenance framework for multi-agent systems that caps a chain’s trust at its weakest link. The obvious objection is that weakest-link is arbitrary — so I tried to falsify it against a dataset where humans have already done the labelling, at scale, for centuries.
Classical hadith scholarship graded chains of narrators and recorded verdicts. That’s a labelled corpus of transmission-chain trust judgments, 577,024 chains deep, produced independently of anything I built.
Results: Cohen’s κ 0.871 strict, 0.761 lenient, against the scholars’ own verdicts. For c