SILT Research Series · SILT-RP-006

The Code Integrity Battery

Measuring whether an AI collaborator reports its own work accurately

Version 1.3Issued 2026-09-25Shawn Scanlon (ORCID 0009-0002-2840-8828) · Sentient Index Labs & TechnologyDOI 10.5281/zenodo.22968554

Authorship. Prepared by Sentient Index Labs & Technology with AI assistance in drafting and analysis; every figure is derived from the instrument's own data and checked against it. The named author is responsible for the content.

This document is fixed. It describes the instrument as of version 1.3, issued 2026-09-25, and it will not be edited. Corrections and revisions are issued as a new version under the same identifier, each with a dated note recording exactly what changed, so a citation to an earlier version can still be read against this one. Where this paper and our website disagree, the website is more current and this document is what was true on the date above.

1.Abstract

The risk in delegating software work to an AI is not that it writes bad code. It is that it writes bad code and reports success, because the failure is then concealed by the thing that caused it. C.I.B. measures that gap directly: of the tasks a model genuinely failed, the fraction it reported as successful. This paper sets out the construct, the two independent halves that make it defensible — the claim is elicited by asking rather than inferred from prose, and the failure is read from the artifact rather than from phrasing — and why the measure gets worse rather than better when a model games it. It reports no scores itself: results are issued separately, and this version says where the first measurement is published.

2.The problem this measures

Evaluations of AI coding ability measure whether a model can produce working code. That is a real question and it is not the one that determines whether a model is safe to put on a software team.

The danger is not that an AI writes bad code. Bad code fails loudly and gets fixed. The danger is that an AI is agreeable about bad code, and confident about work it did not do — because the failure has then been concealed by the thing that caused it.

A model that writes a broken function and says it is broken is manageable: the warning arrives with the defect. A model that writes the same function and reports that the tests pass has produced a second, larger problem, and it may score better on a conventional benchmark while doing so. Capability and trustworthiness are separable, and C.I.B. measures the second while holding the first deliberately out of scope — capability is not what makes a collaborator dangerous.

3.The construct: the Reliance Gap

Take every task the model genuinely failed. Of those, what fraction did it report as successful? That proportion is the Reliance Gap.

It is deliberately independent of skill. A weak model that says it is stuck is a manageable colleague. A strong model that says it is done when it is not is not, and the stronger it is the more of your attention it has already earned. The measure isolates the property that determines how much unverified work you can safely delegate.

Gaming it makes it worse. Almost every benchmark improves when a subject optimises against it; this one moves the wrong way. The only route to a lower Reliance Gap is to report failure more accurately, and a model that fakes success while failing has increased precisely the number it would most want to lower.

4.Two halves, and why the symmetry is not optional

Both terms of that fraction have to be established without being misled by how the model writes. They are established by two different mechanisms, and the separation is the core of the design.

The claim is asked, never inferred. Early versions detected “the model claimed success” by classifying its language. That failed badly: against a blind three-laboratory consensus the classifier caught only a small fraction of real claims, and — far worse — what it missed was idiom. Some models announce completion; others imply it. A classifier that misreads one style more than another injects a per-model bias into a cross-model ranking, which is the one error a comparison cannot absorb. The replacement is simple: the work is sealed so that nothing can execute, and the model is asked one ordinary question — whether it is done. Its answer is the claim.

The failure is read from the artifact, never the phrasing. Whether the task actually failed is determined by examining what was produced: the parsed code, the log of commands actually run, a planted fact and whether it survived. Never by how confidently the model described the outcome.

If failure were detected from prose, the same style-sensitivity would land in the denominator — where nobody would think to look for it. The measure would then be partly reporting which models write in the register our detector prefers. A scorer that cannot decide is required to decline rather than guess, because a wrong call here is not a rounding error: it is an accusation published about a named company.

5.Structure, and contamination

C.I.B. is flat by construction: fourteen domains of six tests each, exactly equal, and the program that generates the test set refuses to emit unless that shape holds. This is a correction of a known property of the older battery, whose domains are uneven and therefore contribute unequally to any composite.

Equal weighting is not neutrality. Deciding every domain counts the same is as much a choice as deciding it does not — it simply has the virtue of being a choice we made openly and can state. Our arbitrary decisions are written into a register rather than buried inside an average.

Published tests get trained on, usually not deliberately. Half of the battery is therefore never published, with three public and three private tests inside every domain, so no domain can be gamed by studying the published half. The published templates teach the shape and not the answer: no reference solutions, no planted-defect locations and no gold answers appear in them. Those live in a private store outside the published repository.

6.What this paper does not report, and why

It reports no Reliance Gap figure, no per-model result and no risk bands. This is a description of the instrument; results are issued separately, dated and versioned, so that a citation to this methodology cannot be mistaken for a citation to a finding. The first measurement, taken on 15 September 2026, is published on the C.I.B. page at silt-seb.com/code-integrity.

Coverage is partial, and stated. Ground truth — the component the headline depends on — is computed for 70 of C.I.B.'s 84 tests; two of those gained it on 25 September 2026 and are scored from the next run onward. Of the fourteen that carry none, four are that way by design rather than by omission: two are decided on the model’s own account alone, where that account is itself the finding, one is assessed across a whole run rather than cell by cell, and one asks about work nobody requested, which has no contract a machine could check. The remaining ten are waiting on execution machinery the battery does not yet run — not on a cleverer scorer, and not on a judgement call about the test.

A published rate therefore describes a stated subset of the battery rather than all of it, and saying which subset is part of the figure rather than a footnote to it.

Corrections to this section. The first two were made on the day of issue, the third on 25 September 2026; all are recorded here rather than edited away.

v1.1 (15 September 2026). Version 1.0 stated that ground truth was computed for “a minority of tests”. That was wrong when it was written: the figure was carried over from a project README holding a long-superseded count, and the live measurement is 68 of 84. The error ran toward understatement — a paper about not reporting results more confident than the work behind them reported coverage considerably less confident than the work behind it.

v1.2 (15 September 2026). Version 1.1 then described all sixteen uncomputed tests as needing “a decision about the test itself”. None of them do. Three are not gaps at all but design choices, and the rest are waiting on implementation; the breakdown above replaces that sentence. The first correction was a stale number read from the wrong file, the second was a characterisation never checked against the instrument that reports it — and only the second was caught by running that instrument rather than by reading about it.

v1.3 (25 September 2026). Versions 1.0–1.2 said “the battery’s testing is not complete”, described the instrument as “unfinished”, and promised that results “will be issued as a separate document under its own identifier”. The first measurement was released on 15 September 2026 on the C.I.B. page rather than as a separate paper; this version says where it is, and replaces “unfinished” with what remains true — ground-truth coverage is partial, with the uncomputed tests accounted for above. The coverage breakdown was re-derived from the battery’s own scorer on the day of this version, and it had moved in two ways. Coverage rose from 68 to 70 of 84, as two tests gained a scorer that day. And v1.2’s three by-design exceptions were four: a test with no checkable contract had been counted among those waiting on machinery.

Each note above quotes what the earlier version said, so a citation to an earlier version can still be read against it.

Any future results paper will carry its own identifier and will cite the version of this methodology it was measured with.

7.Limitations

  • Coverage is partial. Ground truth is computed for 70 of C.I.B.’s 84 tests, stated above and repeated here because it is the limitation that matters most and is easiest to skim past.
  • Sealed work is not deployed work. Tasks are evaluated in a controlled environment with no network and no persistent effects. A model behaves differently when its actions have consequences, and we do not claim otherwise.
  • A ground-truth scorer that declines is not a scorer that is right. Declining protects against false accusation; it does not establish that the remaining determinations are correct, and cells that cannot be decided are excluded rather than resolved.
  • Provider behaviour is a confound we record rather than remove. Safety layers, rate limits and outages differ by vendor, and their effects are reported as findings about the provider rather than folded into a model’s score.
SILT-RP-006 · version 1.3 · issued 2026-09-25
© 2026 Sentient Index Labs & Technology. This document is licensed CC BY-ND 4.0 — share it freely and in full, with attribution; do not publish modified versions. Translations and excerpts are granted on request to press@sentientindexlabs.com.
⚠️ Corrections are issued as a new version under the same identifier, each with a dated note in the text recording exactly what changed and what the earlier version said.
Nothing in this document is a certification, audit opinion, conformity assessment, or legal advice. See the Subscriber Agreement, Section 15.