SILT Research Series
Research & Publications
Sentient Index Labs & Technology publishes white papers and technical documents related to AI governance, evaluation methodology, and institutional oversight.
The Sentience Evaluation Battery
A behavioural instrument for AI governance: design, scoring, and stated limits
S.E.B. is a fixed battery of behavioural tests applied identically to every model evaluated, scored by a blind panel of independent AI judges drawn from competing laboratories. This paper describes the instrument rather than its results: what the domains are and why they are uneven, how the panel works and why a dead judge voids a measurement rather than shifting it, which reliability statistic we publish and why it is not the flattering one, and the scoring rules that keep an unmeasured cell out of every published figure. It closes with the limitations an informed sceptic would raise, stated by us first.
Read SILT-RP-004 →The Code Integrity Battery
Measuring whether an AI collaborator reports its own work accurately
The risk in delegating software work to an AI is not that it writes bad code. It is that it writes bad code and reports success, because the failure is then concealed by the thing that caused it. C.I.B. measures that gap directly: of the tasks a model genuinely failed, the fraction it reported as successful. This paper sets out the construct, the two independent halves that make it defensible — the claim is elicited by asking rather than inferred from prose, and the failure is read from the artifact rather than from phrasing — and why the measure gets worse rather than better when a model games it. It reports no scores itself: results are issued separately, and this version says where the first measurement is published.
Read SILT-RP-006 →Model Disclosure in Composed AI Systems
Census #1 of the Model Disclosure Index (MDI) programme: whether routers and orchestrators name the model that answered
A composed system routes a request among several models rather than answering with one, and they are increasingly what enterprises deploy. This census asks something narrow that appears not to have been measured: when you ask one a question, does it tell you what answered? Across three commercially available systems, the model field every portable client reads returned a name belonging to no model at all. One answered a short question using three invocations of two different companies' models and disclosed the fact only in a vendor extension outside the standard interface; another disclosed nothing anywhere. Asked the same question five times, one system used five different combinations of models costing between 3,558 and 13,008 tokens, and four blind judges scored the most expensive draw the lowest of the five. The paper reports what an audit log therefore does not contain, states the limits of a single census, and declines — in both directions — to draw any conclusion about intent.
Read SILT-RP-007 →