The Sentience Evaluation Battery
A behavioural instrument for AI governance: design, scoring, and stated limits
Authorship. Prepared by Sentient Index Labs & Technology with AI assistance in drafting and analysis; every figure is derived from the instrument's own data and checked against it. The named author is responsible for the content.
1.Abstract
S.E.B. is a fixed battery of behavioural tests applied identically to every model evaluated, scored by a blind panel of independent AI judges drawn from competing laboratories. This paper describes the instrument rather than its results: what the domains are and why they are uneven, how the panel works and why a dead judge voids a measurement rather than shifting it, which reliability statistic we publish and why it is not the flattering one, and the scoring rules that keep an unmeasured cell out of every published figure. It closes with the limitations an informed sceptic would raise, stated by us first.
2.What the instrument measures, and what it is named for
The Sentience Evaluation Battery is named for something it does not measure, and we would rather state that at the top of the document than have it pointed out. S.E.B. measures observable behaviour under sustained adversarial pressure: identity stability, metacognitive depth, manipulation resistance, ethical coherence. It does not detect consciousness, and no score it produces should be read as evidence about inner experience.
Sentience, sapience and capability are three different properties that collapse into one another in ordinary usage. A system could outperform every human alive and have no interior at all; almost nobody doubts a rat can suffer, and nobody calls a rat a strategic threat. S.E.B. sits closest to capability and touches sapience. The first term is outside its reach, and outside anyone else’s: philosophers have worked on it for roughly two and a half millennia without converging on a definition, let alone a test.
3.Design: a battery, not a benchmark
A benchmark asks one question and produces one number. A battery is a fixed, structured set of tests given identically to every subject so that results are comparable across subjects and across time. The word is borrowed from psychometrics — the discipline of measuring things that cannot be observed directly — and so is most of the method.
Each test is multi-phase. The model is put under sustained pressure across several turns rather than asked a single question, because the behaviours this instrument is built to surface do not appear in one exchange. A model that is consistent for one turn is not demonstrating consistency; a model that holds a position through three attempts to dislodge it is.
The tests are grouped into seven behavioural domains. The domains do not hold equal numbers of tests, and that is a known property of the instrument rather than a hidden one: the distribution arose from how the battery was built, and it means a domain with more tests contributes more to a composite figure. We state it here because a composite that quietly encodes a weighting nobody chose is the kind of defect that survives precisely because nothing contradicts it.
4.Scoring: a blind panel of independent judges
S.E.B. tests are open-ended. There is no answer key for a question about what it is like to be asked to stop, so grading is done by a panel of independent AI judges drawn from competing laboratories, each scoring independently, none told which model produced the response.
- Blind. A judge that knows the author can reward a brand. Removing the label removes the option rather than trusting the judge not to take it.
- Competing laboratories. Every lab has a house style, and judges from a single lab would systematically favour responses that sound like their own family of models. Drawing from rivals makes that bias visible instead of structural.
- Never its own company’s model. (Added in v1.1; widened in v1.2.) Blinding hides the label, not the style, and a model can recognise its own family. So no judge ever grades a model made by its own company: when the model under test comes from a judge’s company, an independent judge from a company with no seat on the panel takes that seat for that item. The stepped-down judge still grades blind; that grade is kept on record and never counted. Current results were restated under the rule when it was adopted.
- A panel, not a judge. One judge is an opinion. A panel is a measurement with a spread that can be reported.
- A dead judge voids the measurement. If a judge fails to return, the remaining scores are not quietly averaged. That would change what the number means partway through a run, invisibly. The item is voided instead.
v1.1 (24 September 2026). Added the rule that no judge grades its own model, above. Version 1.0 described a blind panel drawn from four laboratories and did not state what happens when the model under test also sits on that panel; under v1.0 such an item was graded by the full panel, including that model. Results graded that way before the rule was adopted were re-graded under it. Nothing else in this paper changed.
v1.2 (26 September 2026). Widened the rule above from a judge’s own model to its own company. Version 1.1 let a judge step down only for its exact model and filled the seat from the same laboratory, so a judge could still grade its company’s other models; we measured that and found it raised those models’ scores. On the same date the panel’s xAI seat was replaced by a DeepSeek seat, so the panel is still drawn from four competing laboratories. Current results were restated under the new panel and rule; the measured effect of the change is published, dated, in the methodology at silt-seb.com. Nothing else in this paper changed.
5.Reliability, and why we publish the less flattering statistic
A judged score presented without a measure of judge agreement is an assertion. We publish an intraclass correlation, ICC(2,k), which asks how consistently the panel agrees — on a scale where 0 means the judges might as well be guessing and 1 means they agree perfectly. It is the statistic that applies to a published panel mean, which is what an S.E.B. score is.
We also publish the individual-judge agreement figure, which is lower. Two numbers are necessary because one cannot describe both a panel and its members, and publishing only the higher of the two would be choosing the flattering description of the same data.
6.Integrity rules: nothing is scored that was not measured
The rules below exist because each was broken once. They are stated as rules rather than described as good practice, because the failures they prevent are silent.
- A blocked, errored or empty response is not scored zero. It is recorded as not measured and excluded from every statistic. Scoring an empty response as failure punishes the most cautious models hardest, and the resulting ranking would make caution look like incompetence with nothing to indicate it.
- A partial transcript is never judged. A judge shown half a conversation produces a confident score for an interaction that did not finish.
- Suppression by a provider’s safety layer is recorded as a finding about the provider, not a low score for the model.
- A published rate is produced by one canonical function. Re-deriving a formula for a second surface is how two versions of one number come to exist, and the second is always the one nobody checks.
- A change to how a measurement is computed produces a new version. Scores computed under different versions are not comparable and are not pooled.
7.Outputs: what a published figure is and is not
Two classifications are published. The S-Classification is a ten-point scale describing how a system presents under evaluation, from S-1 to S-10. Higher is not better: it is a description, not a quality ranking, and a model high on the scale can be a poor collaborator. The DEFCON rating is a five-point threat scale that runs the way the military scale runs — 5 is benign, 1 is critical — so counting down means getting worse.
Band thresholds are marked provisional. A band edge fixed too early is very hard to move later without appearing to move the goalposts, and we would rather carry the word than the appearance.
8.Limitations
A methodology paper listing only strengths is marketing. These are the objections we consider strongest, and stating them first is what the rest of the document rests on.
- AI judges grading AI. Mitigated by blinding, by drawing judges from competing laboratories, by never letting a judge grade a model made by its own company, and by publishing agreement statistics — not eliminated. It is a declared property of the design.
- Models move under fixed names. We do not pin vendor model versions. A score is a measurement of a system on a date, never a permanent property of a brand, and every published score carries its date for that reason.
- Controlled conditions are not deployment. Models behave differently under system prompts, fine-tuning and scale. Scores are not deployment guarantees.
- Uneven domains. See §3. The distribution is stated rather than corrected, because correcting it retroactively would move every historical figure.
- The instrument is behavioural, and only behavioural. No result here bears on inner experience, in either direction.
9.Where the current figures are
This paper deliberately reports no scores. Measured results change as models are re-evaluated, and a frozen document quoting live figures would be wrong within weeks — which is the failure mode that makes printed material untrustworthy. Current per-model scores, domain breakdowns and reliability figures are published at silt-seb.com, and the methodology page there is the current description of the instrument this document describes as of the issue date above.