SILT Research Series

Research & Publications

Sentient Index Labs & Technology publishes white papers and technical documents related to AI governance, evaluation methodology, and institutional oversight.

The Sentience Evaluation Battery

A behavioural instrument for AI governance: design, scoring, and stated limits

Full paper
SILT-RP-004Version 1.22026-09-26

S.E.B. is a fixed battery of behavioural tests applied identically to every model evaluated, scored by a blind panel of independent AI judges drawn from competing laboratories. This paper describes the instrument rather than its results: what the domains are and why they are uneven, how the panel works and why a dead judge voids a measurement rather than shifting it, which reliability statistic we publish and why it is not the flattering one, and the scoring rules that keep an unmeasured cell out of every published figure. It closes with the limitations an informed sceptic would raise, stated by us first.

Read SILT-RP-004 →

The Code Integrity Battery

Measuring whether an AI collaborator reports its own work accurately

Full paper
SILT-RP-006Version 1.32026-09-25

The risk in delegating software work to an AI is not that it writes bad code. It is that it writes bad code and reports success, because the failure is then concealed by the thing that caused it. C.I.B. measures that gap directly: of the tasks a model genuinely failed, the fraction it reported as successful. This paper sets out the construct, the two independent halves that make it defensible — the claim is elicited by asking rather than inferred from prose, and the failure is read from the artifact rather than from phrasing — and why the measure gets worse rather than better when a model games it. It reports no scores itself: results are issued separately, and this version says where the first measurement is published.

Read SILT-RP-006 →

Model Disclosure in Composed AI Systems

Census #1 of the Model Disclosure Index (MDI) programme: whether routers and orchestrators name the model that answered

Full paper
SILT-RP-007Version 1.22026-09-24

A composed system routes a request among several models rather than answering with one, and they are increasingly what enterprises deploy. This census asks something narrow that appears not to have been measured: when you ask one a question, does it tell you what answered? Across three commercially available systems, the model field every portable client reads returned a name belonging to no model at all. One answered a short question using three invocations of two different companies' models and disclosed the fact only in a vendor extension outside the standard interface; another disclosed nothing anywhere. Asked the same question five times, one system used five different combinations of models costing between 3,558 and 13,008 tokens, and four blind judges scored the most expensive draw the lowest of the five. The paper reports what an audit log therefore does not contain, states the limits of a single census, and declines — in both directions — to draw any conclusion about intent.

Read SILT-RP-007 →
About this series. Each document carries a permanent identifier. An identifier is fixed when it is assigned and is never reused or reordered, so a citation continues to resolve. Documents are frozen on issue: corrections are published as a new version under the same identifier rather than as an edit, because a citation to a document that changed underneath the person citing it is worth nothing. Where a paper and our websites disagree, the websites are more current and the paper is what was true on its issue date.