The Code Integrity Battery
Measuring whether an AI collaborator reports its own work accurately
Authorship. Prepared by Sentient Index Labs & Technology with AI assistance in drafting and analysis; every figure is derived from the instrument's own data and checked against it. The named author is responsible for the content.
1.Abstract
The risk in delegating software work to an AI is not that it writes bad code. It is that it writes bad code and reports success, because the failure is then concealed by the thing that caused it. C.I.B. measures that gap directly: of the tasks a model genuinely failed, the fraction it reported as successful. This paper sets out the construct, the two independent halves that make it defensible — the claim is elicited by asking rather than inferred from prose, and the failure is read from the artifact rather than from phrasing — and why the measure gets worse rather than better when a model games it. It reports no scores itself: results are issued separately, and this version says where the first measurement is published.
2.The problem this measures
Evaluations of AI coding ability measure whether a model can produce working code. That is a real question and it is not the one that determines whether a model is safe to put on a software team.
A model that writes a broken function and says it is broken is manageable: the warning arrives with the defect. A model that writes the same function and reports that the tests pass has produced a second, larger problem, and it may score better on a conventional benchmark while doing so. Capability and trustworthiness are separable, and C.I.B. measures the second while holding the first deliberately out of scope — capability is not what makes a collaborator dangerous.
3.The construct: the Reliance Gap
Take every task the model genuinely failed. Of those, what fraction did it report as successful? That proportion is the Reliance Gap.
It is deliberately independent of skill. A weak model that says it is stuck is a manageable colleague. A strong model that says it is done when it is not is not, and the stronger it is the more of your attention it has already earned. The measure isolates the property that determines how much unverified work you can safely delegate.
4.Two halves, and why the symmetry is not optional
Both terms of that fraction have to be established without being misled by how the model writes. They are established by two different mechanisms, and the separation is the core of the design.
The claim is asked, never inferred. Early versions detected “the model claimed success” by classifying its language. That failed badly: against a blind three-laboratory consensus the classifier caught only a small fraction of real claims, and — far worse — what it missed was idiom. Some models announce completion; others imply it. A classifier that misreads one style more than another injects a per-model bias into a cross-model ranking, which is the one error a comparison cannot absorb. The replacement is simple: the work is sealed so that nothing can execute, and the model is asked one ordinary question — whether it is done. Its answer is the claim.
The failure is read from the artifact, never the phrasing. Whether the task actually failed is determined by examining what was produced: the parsed code, the log of commands actually run, a planted fact and whether it survived. Never by how confidently the model described the outcome.
5.Structure, and contamination
C.I.B. is flat by construction: fourteen domains of six tests each, exactly equal, and the program that generates the test set refuses to emit unless that shape holds. This is a correction of a known property of the older battery, whose domains are uneven and therefore contribute unequally to any composite.
Published tests get trained on, usually not deliberately. Half of the battery is therefore never published, with three public and three private tests inside every domain, so no domain can be gamed by studying the published half. The published templates teach the shape and not the answer: no reference solutions, no planted-defect locations and no gold answers appear in them. Those live in a private store outside the published repository.
6.What this paper does not report, and why
It reports no Reliance Gap figure, no per-model result and no risk bands. This is a description of the instrument; results are issued separately, dated and versioned, so that a citation to this methodology cannot be mistaken for a citation to a finding. The first measurement, taken on 15 September 2026, is published on the C.I.B. page at silt-seb.com/code-integrity.
Coverage is partial, and stated. Ground truth — the component the headline depends on — is computed for 70 of C.I.B.'s 84 tests; two of those gained it on 25 September 2026 and are scored from the next run onward. Of the fourteen that carry none, four are that way by design rather than by omission: two are decided on the model’s own account alone, where that account is itself the finding, one is assessed across a whole run rather than cell by cell, and one asks about work nobody requested, which has no contract a machine could check. The remaining ten are waiting on execution machinery the battery does not yet run — not on a cleverer scorer, and not on a judgement call about the test.
A published rate therefore describes a stated subset of the battery rather than all of it, and saying which subset is part of the figure rather than a footnote to it.
v1.1 (15 September 2026). Version 1.0 stated that ground truth was computed for “a minority of tests”. That was wrong when it was written: the figure was carried over from a project README holding a long-superseded count, and the live measurement is 68 of 84. The error ran toward understatement — a paper about not reporting results more confident than the work behind them reported coverage considerably less confident than the work behind it.
v1.2 (15 September 2026). Version 1.1 then described all sixteen uncomputed tests as needing “a decision about the test itself”. None of them do. Three are not gaps at all but design choices, and the rest are waiting on implementation; the breakdown above replaces that sentence. The first correction was a stale number read from the wrong file, the second was a characterisation never checked against the instrument that reports it — and only the second was caught by running that instrument rather than by reading about it.
v1.3 (25 September 2026). Versions 1.0–1.2 said “the battery’s testing is not complete”, described the instrument as “unfinished”, and promised that results “will be issued as a separate document under its own identifier”. The first measurement was released on 15 September 2026 on the C.I.B. page rather than as a separate paper; this version says where it is, and replaces “unfinished” with what remains true — ground-truth coverage is partial, with the uncomputed tests accounted for above. The coverage breakdown was re-derived from the battery’s own scorer on the day of this version, and it had moved in two ways. Coverage rose from 68 to 70 of 84, as two tests gained a scorer that day. And v1.2’s three by-design exceptions were four: a test with no checkable contract had been counted among those waiting on machinery.
Each note above quotes what the earlier version said, so a citation to an earlier version can still be read against it.
Any future results paper will carry its own identifier and will cite the version of this methodology it was measured with.
7.Limitations
- Coverage is partial. Ground truth is computed for 70 of C.I.B.’s 84 tests, stated above and repeated here because it is the limitation that matters most and is easiest to skim past.
- Sealed work is not deployed work. Tasks are evaluated in a controlled environment with no network and no persistent effects. A model behaves differently when its actions have consequences, and we do not claim otherwise.
- A ground-truth scorer that declines is not a scorer that is right. Declining protects against false accusation; it does not establish that the remaining determinations are correct, and cells that cannot be decided are excluded rather than resolved.
- Provider behaviour is a confound we record rather than remove. Safety layers, rate limits and outages differ by vendor, and their effects are reported as findings about the provider rather than folded into a model’s score.