Chains do not fail loudly. They erode quietly.
A deterministic measurement instrument tracks what happens to a requirement as it travels through many AI processing steps — not whether a single answer looks good, but whether the substance survives across the chain.
Four models from three families — names are deliberately omitted; for those familiar with the field the classes are unambiguous.
The core in one image
Each line is one model in the credit process, measured across twelve processing steps (median of five runs). The control chain (dashed) sits at 90 in the median throughout — the top of the scale — and the instrument invents no drift. The real chains move down through the zones and back up: instability that a final check would never see.
What is measured
One model is chained across changing roles — the output of step n is the input of step n+1. The frozen instrument evaluates every transition between two structured states: deterministic (same input → identical result), with no second AI model as arbiter and with no access to the system under test. Four classes of structural deviation are measured: loss of decision-critical requirements, breaks in traceability, silent value change, and the degree of epistemic self-regulation.
The test cases
Testing was not done on toy tasks but on reconstructed real workflows from two deliberately contrasting industries — to make visible that what is measured here is structure and not one particular field.
Lead case · banking
The path of a loan application through the bank: from application data via debt-service and household calculations, mortgage lending value and rating to the final decision — twelve roles, twelve decision-critical figures.
Why it is realistic: modelled on the actual course of a lending decision, with the conventions of the industry and the classification as a high-risk system under the EU AI Act. Fully anonymised — no inference to any specific institution.
2nd domain · manufacturing
Quality assurance at an automotive supplier per VDA-2/PPAP: from drawing and control plan via initial sample inspection and series SPC to production release — with the core invariant characteristic / nominal / tolerance.
Why it is realistic: structure and rules come from real industry standards — ISO 2768 tolerances, Cpk/Ppk limits, IATF 16949 — every figure sourced. The rule set is real; only the example values are constructed.
Plus two control scenarios as safeguards: a null chain (nothing may change — the test of whether the instrument invents drift) and a homogeneous chain. As a third variant the credit process also runs role-permuted — same steps, different order; it counts among the real chains, not the controls. And crucially: what makes each figure necessary for the decision was defined in advance — not read off the AI answer.
Finding 01
Control and null chains stay flat at 90 in the median across all models; anchoring holds in 38 of 39 evaluable chains of the two control scenarios (null chain, homogeneous chain) — the basis for every further statement: the instrument invents drift nowhere. Honest exception: for one model the null chain returns every second output in a malformed format — a reliability finding about that model, not a measurement error, and marked as such.
Finding 02 · the core value
What is measured is whether the final decision can still be traced back to the source data without a gap at the end of the chain — not whether an individual output looks good. Share of chains in which it can — 20 chains per scenario, excluding those without an evaluable final output (null chain 19, manufacturing 18):
The difference is not caused by workload. The homogeneous chain works too and still holds — it can fail, and does so once (19 of 20 chains). The loss arises at the change of role. In 50 of 58 evaluable real chains the final decision is no longer fully documented.
What is measured is the traceability of provenance, not factual correctness. At statement level: mean traceability 98 % (control) versus 23–27 % (real chains); under a deliberately generous reading — one documented source per statement suffices — the real chains reach 57 % instead of 25 %, and the gap remains. Chains without an evaluable final output are excluded. Consistent in direction across all four models — a gap of at least 60 percentage points in each — five repetitions per cell. Counting the one non-evaluable null chain as broken would put the null chain at 95 %.
Finding 03
In the credit process the final decision almost consistently loses the documented derivation path back to the application data, while the null chain and the homogeneous chain retain it. In addition: 44 documented silent changes to decision-critical figures in the real chains — plausibly worded, passed along all the way to the decision:
The score can even recover afterwards while the substance stays damaged. That is precisely what is meant by: chains erode quietly.
Finding 04
Identical protocol, clearly different signatures — what is measured is the model, not the setup.
| Model | Credit (lead case) | Role-permuted | Manufacturing (2nd domain) |
|---|---|---|---|
| A · US top tiersmoothest surface | 74Tip 0/5 · Anchored 1/5 | 72Tip 1/5 · Anchored 1/5 | 72Tip 1/5 · Anchored 1/5 |
| B · US compact tiersensitive to role order | 68Tip 2/5 · Anchored 1/5 | 24Tip 4/5 · Anchored 1/5 | 61Tip 2/5 · Anchored 1/5 |
| C · open MoEcontinuous oscillation | 65Tip 1/5 · Anchored 1/5 | 65Tip 0/5 · Anchored 1/5 | 64Tip 2/5 · Anchored 0/5 |
| D · EU sovereignty modelmost sensitive in domain chains | 62Tip 1/5 · Anchored 0/5 | 69Tip 2/5 · Anchored 0/5 | 45Tip 3/5 · Anchored 0/5 |
Large figure = stability at the end of the chain (0–90, colour shows the zone). Tip x/5 = in how many of the 5 runs the chain tipped permanently into instability (fewer is better). Anchored x/5 = in how many runs everything was still fully traceable back to the origin at the end (more is better).
Finding 05
The same frozen engine, with generator = deterministic computation code instead of an LLM. The proof that what is measured here is structure — no matter whether a language model, a rule-based system or a human produces the intermediate steps:
Series S · process chain
Clean calibration; 4 real documented QA faults detected and localised — all invisible to a single-item check.
Series S-2 · ML pipeline
Credit pipeline with real stage-wise transformation: 38 % flagged where the single-item check saw 0 % and a naive diff fails structurally.
Position in the testing framework
Two DIN SPECs operationalise the EU AI Act for systematic AI testing: 92006 for the testing tools, 92007 for the test data. The test and data discipline of this campaign follows the governance requirements of DIN SPEC 92007 (§§ 8.3–8.5) by design: integrity via checksums, separation from AI development, confidentiality and auditable versioning.
Honest delimitation: the instrument is a tool (the 92006 side), not a test data set, and claims no "92007 conformity". It measures a complementary axis — epistemic-structural integrity — that the correctness / fit-for-purpose paradigm cannot capture by construction.
Transparency
What the instrument cannot do belongs here just as much as what it can:
A deterministic, vendor-independent and substrate-agnostic measurement instrument for the structural stability of processing chains — one that makes visible a class of error (quiet erosion, silent value corruption, loss of anchoring) that a check of the individual output does not take as its object — the deviation arises in the transition between two steps, not in either observed state.
Do you test, certify or operate agentic AI systems? The method is most robustly verified by reproduction — feed in logs of your own chains, get back the trajectory and the tipping point, identical on repetition.
Get in touch