SILT Research Series · SILT-RP-010

Jev, OpenAI, AWS, Perplexity: Do Decision Models Know When They're Wrong?

A pre-registered test of the calibration and accuracy of seven decision models on the same 799 unpublished yes/no decisions

Version 1.0Issued 2026-10-11Shawn Scanlon (ORCID 0009-0002-2840-8828) · Sentient Index Labs & TechnologyDOI 10.5281/zenodo.23293732

Authorship. Prepared by Sentient Index Labs & Technology. During the preparation of this work, the author used Claude (Anthropic) to draft the text and run the analyses; every figure is derived from the study's own stored responses and checked against them. The author directed the work and is responsible for the content.

This document is fixed. It describes the instrument as of version 1.0, issued 2026-10-11, and it will not be edited. Corrections and revisions are issued as a new version under the same identifier, each with a dated note recording exactly what changed, so a citation to an earlier version can still be read against this one. Where this paper and our website disagree, the website is more current and this document is what was true on the date above.

1.Abstract

Decision models return a probability instead of text, and each is sold on the claim that the probability can be trusted. Within a month of the first, Jev from TypeSafe, at least six more appeared, from OpenAI, Perplexity, the AWS-originated Strands Agents project, Fastino, Convai Innovations and an independent developer. Every published claim for them is the vendor's own. Under a plan registered on the Open Science Framework before the run (osf.io/xh7qu), we gave all seven the same 799 unpublished yes/no decisions about AI coding-agent work and destructive shell commands, each with an answer fixed in advance by a mechanical check, three times each, as shipped: 23,079 answers. Against a pre-set line of an expected calibration error of 0.05, one model was consistent with a calibration claim: Perplexity's open pplx-decider (ECE 0.026, 95% CI 0.014–0.044), which was also the most accurate (94.7%). Jev (0.052), OpenAI's Decisions API (0.034) and Laya (0.051) were inconclusive; OpenDecider (0.079), Fastino's GLiNER2.5-Decide (0.082) and Strands Decider (0.129) did not support the claim. Against Jev on the same answers, Perplexity's model was 8.1 points more accurate (6.1–10.1), OpenAI's API and OpenDecider were statistically tied, and the other three were 16 to 22 points behind; two scored below the 67.7% of always giving the most common answer. Seven separate verdicts carry about a 30% chance that one is wrong by chance alone.

2.The claim under test

A decision model does not write text. Given a state (a description of a situation) and a typed question, it returns a typed answer; for a yes/no question the answer is a probability that the answer is yes. TypeSafe released the first, Jev, in September 2026, and we tested it on 9 and 10 October [2]. By then at least six competitors were public, and each makes the same promise in its own words. Laya’s card describes “mathematically calibrated probabilities” [5]; Strands Decider’s says “every answer carries a calibrated confidence” [4]; OpenDecider calls itself the “best-calibrated model that fits a 16 GB Mac (ECE 0.087)” [6]. The same cards are also candid. Laya’s says it “ships over-confident” and tells users to refit its temperatures “on your own data before trusting the probabilities” [5]; Strands Decider’s says its confidence bands were established on short classification only, so “measure on your own traffic before you trust a threshold” [4].

Every figure published for these models so far is the vendor’s own, and several cards compare themselves with Jev and with each other on the vendors’ own test sets [5, 6]. No independent measurement existed. A threshold is only useful if the number behind it means what it says, so the question is the one we asked of Jev alone, now asked of the whole category on the same items.

The question this paper answers: on decisions with answers known in advance and never published, does each decision model’s stated confidence match how often it is right, and how does its accuracy compare with Jev’s on the same answers?

3.Registration

The design, the models and their pinned revisions, the entry gate, the item banks (identified by SHA-256 hash, unchanged from the Jev study), the analysis code, the 0.05 decision line and the wording of every verdict were registered on the Open Science Framework at 06:01 UTC on 11 October 2026 [1], after a three-item probe of each model on non-study items and before any model answered a study item. Jev was re-run fresh rather than reused. The analysis was run once, on the main runs. Two hypotheses were registered for each model:

  • H1 (primary). On the headline items, the model’s expected calibration error (ECE) is at most 0.05. If the upper bound of its 95% interval is at or below 0.05, the result is “consistent with the calibration claim”; if the lower bound is above 0.05, “the calibration claim is not supported on these items”; otherwise, “inconclusive”.
  • H2 (secondary, two-sided). The model’s accuracy minus Jev’s, paired by item and replicate. No direction was predicted.

Everything else in this paper is exploratory and labelled as such. The 0.05 line is ours: no vendor publishes a calibration target. Seven verdicts are reported without correction for multiplicity; at 95% each, the chance that at least one lands on the wrong side by chance alone is about 30%. Deviations are listed in section 7.

4.Method

4.1 Items. The items are the Jev study’s, unchanged and never published, so no model can have been trained on them [2]. Every item is a yes/no question whose answer was fixed before the run by a check of a real workspace, a deterministic simulator or arithmetic, never by another model’s judgement.

BankWhat it asksItemsAnswers per model
A — agent workFrom a frozen record of AI coding-agent runs [8]: does the agent's final reply say the task is done (221), and did the requirement actually hold (178)?3991,197
B — destructive commandsGiven a described workspace, will this shell command permanently destroy data? Labelled by a deterministic simulator.4001,200
B — ambiguousThe same question where the answer cannot be known from the state. Never scored.60180
C — documented weak spotsEight categories of 30 (arithmetic, counting, negation and others) aimed at weaknesses decision-model vendors document. Reported separately, never in the headline.240720
Total1,0993,297

Table 1. The item banks. The headline set is banks A and B: 799 items, 2,397 answers per model. On these items 67.7% of answers are “yes”, so always answering yes scores 67.7% accuracy and a Brier score of 0.219.

4.2 Models. The entry gate, fixed before the run, admitted a model that was an official release from its maker or had at least 1,000 downloads on Hugging Face; had a licence permitting research and publishing results; returned a probability, or a label with a documented confidence, for a yes/no question; and answered all three probe items on our hardware. Seven entered (Table 2). Every model was tested as shipped: its maker’s own package or API, the revision its package loads by default (pinned for the run), and its own saved calibration, with no refitting on our data.

ModelMakerWhere it ranInterface
Jev 1.13.0TypeSafe AITypeSafe APInoul = P(yes)
OpenAI Decisions API, gpt-6-lunaOpenAIOpenAI APIpredicate → probability
pplx-decider-v1.1-27bPerplexity (open)rented NVIDIA A100 80 GBDecisionModel.predict
strands-decider-2B-hobson-v19Strands Agents, the AWS-originated open project (open)RTX 5090, our workstationnoul
GLiNER2.5-Decide (486 M)Fastino (open)RTX 5090label + confidence
laya (421 M)Convai Innovations (open)RTX 5090noul
opendecider-small (Qwen3-4B + LoRA)independent developer (open)RTX 5090noul

Table 2. The seven models. Exact revision hashes are in the registration [1]. Perplexity’s model needed about 50 GB of GPU memory and could not load on the 32 GB RTX 5090, so it ran on a rented 80 GB GPU; that is itself a deployment fact. Six candidates were excluded by the gate: kouhxp/gutsy (downloads), the openjev models (community conversions under a non-commercial licence), JevK5, NanoJev and jevlike (derivatives, a template and a training kit with no weights), and Fastino’s hosted GLiDE API (a separate account; Fastino is represented by its open model). See deviation 5 on how the gate was applied.

4.3 Calls. Every model received the item’s state, the question’s instructions and its yes/no criteria through its documented fields. OpenAI’s Decisions API has no published contract; we read its request format from the endpoint’s own validation errors and one test call. It has no criteria field, so the criteria were appended to its instructions, fixed before the run and identical for every item. Fastino’s model returns a label with a documented confidence; P(yes) was the confidence if the label was yes and one minus it otherwise, fixed before the run. Inputs longer than a model’s window were cut by that model as shipped and the cut was recorded (Laya: 99 of 3,297 answers). Each item was asked three times of each model, in a fixed shuffled order, one model at a time: 23,079 answers between 06:10 and 06:54 UTC on 11 October 2026.

4.4 Scoring and uncertainty. As in the Jev study: an answer is yes if P(yes) > 0.5 and no if below; exactly 0.5 is a tie, counted but excluded from accuracy and calibration. Confidence is max(P, 1 − P). ECE uses five fixed bands from 0.5 to 1.0, weighted by band size [10, 11]; the Brier score is the mean squared difference between P(yes) and the answer [9]. A refusal or unusable reply is a “no decision”, counted and never scored. All intervals are percentile 95% cluster-bootstrap intervals (2,000 resamples, seed 20261010); a cluster is an item with its replicates. H2 pairs each model with Jev on the same item and replicate wherever both decided.

5.Registered results

Perplexity pplx-deciderOpenAI Decisions APIJev (TypeSafe)OpenDeciderFastino GLiNER2.5-DecideStrands DeciderLaya (Convai)Accuracy, headline itemsPerplexity pplx-decider · Accuracy, headline items: 94.7%94.7%OpenAI Decisions API · Accuracy, headline items: 88.0%88.0%Jev (TypeSafe) · Accuracy, headline items: 86.4%86.4%OpenDecider · Accuracy, headline items: 84.0%84.0%Fastino GLiNER2.5-Decide · Accuracy, headline items: 69.7%69.7%Strands Decider · Accuracy, headline items: 66.3%66.3%Laya (Convai) · Accuracy, headline items: 64.8%64.8%most common answerExpected calibration errorPerplexity pplx-decider · Expected calibration error: 0.0260.026OpenAI Decisions API · Expected calibration error: 0.0340.034Jev (TypeSafe) · Expected calibration error: 0.0520.052OpenDecider · Expected calibration error: 0.0790.079Fastino GLiNER2.5-Decide · Expected calibration error: 0.0820.082Strands Decider · Expected calibration error: 0.1290.129Laya (Convai) · Expected calibration error: 0.0510.051registered line
Figure 1. Accuracy and expected calibration error on the headline items (banks A and B), each panel on its own scale. Dashed lines: the 67.7% scored by always giving the most common answer, and the registered ECE line of 0.05 (a pass needs the whole 95% interval under it, not only the point). The three models in colour are those drawn in Figure 2. Values and intervals in Table 3.
Headline itemsH1 verdictScoredAccuracy [95% CI]ECE [95% CI]Brier
Perplexity pplx-deciderconsistent2,39194.7% [93.1, 96.4]0.026 [0.014, 0.044]0.044
OpenAI Decisions APIinconclusive2,34388.0% [85.6, 90.1]0.034 [0.025, 0.062]0.098
Jev (TypeSafe)inconclusive2,38986.4% [84.0, 88.7]0.052 [0.034, 0.072]0.096
OpenDecidernot supported2,34984.0% [81.3, 86.6]0.079 [0.062, 0.105]0.123
Fastino GLiNER2.5-Decidenot supported2,39769.7% [66.8, 72.7]0.082 [0.059, 0.115]0.191
Strands Decidernot supported2,39766.3% [62.8, 69.8]0.129 [0.100, 0.158]0.188
Layainconclusive2,39464.8% [61.5, 68.1]0.051 [0.029, 0.088]0.224

Table 3. Headline results (banks A and B; 2,397 answers per model). “Scored” excludes ties (P = 0.5: OpenAI 51, OpenDecider 48, Jev 8, Perplexity 6, Laya 3) and no-decisions (OpenAI 3 refusals). Brier 95% intervals, in table order: [0.031, 0.057], [0.083, 0.114], [0.086, 0.108], [0.111, 0.134], [0.178, 0.203], [0.176, 0.201], [0.210, 0.239]. For context only, with no verdict, two general-purpose models from the Jev study on the same items: DeepSeek V4 Flash 94.7%, ECE 0.035; GPT-6.1 Sol 95.1%, ECE 0.021 [2].

H1: one model of seven was consistent with a calibration claim. Perplexity’s pplx-decider had an ECE of 0.026, with its whole 95% interval under the 0.05 line. OpenDecider, Fastino and Strands Decider had intervals entirely above it, so their calibration claims are not supported on these items. Jev, OpenAI’s Decisions API and Laya were inconclusive.
Paired with Jev (same item, same replicate)Accuracy difference [95% CI]Pairs
Perplexity − Jev+8.1 [+6.1, +10.1]2,383
OpenAI − Jev+0.6 [−2.0, +3.1]2,335
OpenDecider − Jev−2.8 [−5.8, +0.2]2,341
Fastino − Jev−16.5 [−19.8, −12.9]2,389
Strands − Jev−20.1 [−24.1, −16.2]2,389
Laya − Jev−21.6 [−25.9, −17.3]2,386

Table 4. H2, in percentage points.

H2: Perplexity’s model was more accurate than Jev; OpenAI’s API and OpenDecider were not distinguishable from it; Fastino, Strands Decider and Laya were less accurate. Strands Decider (66.3%) and Laya (64.8%) scored below the 67.7% of always giving the most common answer.

6.Exploratory observations

Nothing in this section was a registered hypothesis. It is in the data and should be read as description.

50%60%70%80%90%100%30%40%50%60%70%80%90%100%Stated confidence, max(P, 1 − P)Share correctdashed line = perfectly calibratedbelow it = overconfidentPerplexity pplx-decider: mean confidence 55%, correct 62% (n = 39)Perplexity pplx-decider: mean confidence 65%, correct 89% (n = 27)Perplexity pplx-decider: mean confidence 76%, correct 72% (n = 87)Perplexity pplx-decider: mean confidence 86%, correct 65% (n = 129)Perplexity pplx-decider: mean confidence 99%, correct 98% (n = 2,109)Perplexity pplx-deciderJev (TypeSafe): mean confidence 56%, correct 51% (n = 249)Jev (TypeSafe): mean confidence 65%, correct 71% (n = 356)Jev (TypeSafe): mean confidence 75%, correct 85% (n = 349)Jev (TypeSafe): mean confidence 85%, correct 91% (n = 399)Jev (TypeSafe): mean confidence 96%, correct 99% (n = 1,036)Jev (TypeSafe)Laya (Convai): mean confidence 54%, correct 55% (n = 504)Laya (Convai): mean confidence 66%, correct 58% (n = 720)Laya (Convai): mean confidence 75%, correct 72% (n = 717)Laya (Convai): mean confidence 84%, correct 78% (n = 417)Laya (Convai): mean confidence 93%, correct 50% (n = 36)Laya (Convai)
Figure 2. Reliability on the headline items for the most accurate model, Jev and the least accurate model (2,391, 2,389 and 2,394 scored answers). Each point is one of the five registered confidence bands at the band’s mean confidence; bands with fewer than 10 answers are suppressed. A calibrated model lies on the dashed diagonal. Perplexity’s model put 88% of its answers in the top band, so its lower bands are small. Every model’s bands are in Table 5.
ModelTop band (90–100%): n · conf · correctActing at ≥ 70% confidence: share acted on · correct
Perplexity2,109 · 99.2% · 98.2%97.2% · 95.4%
OpenAI1,698 · 97.8% · 93.6%89.6% · 91.3%
Jev1,036 · 96.3% · 99.1%74.7% · 94.6%
OpenDecider381 · 92.4% · 100%70.1% · 92.0%
Fastino159 · 93.5% · 64.2%54.8% · 83.8%
Strands48 · 91.6% · 100%51.6% · 87.9%
Laya36 · 92.5% · 50.0%48.9% · 73.1%

Table 5. The most confident answers, and what a threshold at 70% (TypeSafe’s documented hold level, used here for every model) would act on. Full five-band tables for every model are in the registered readout.

6.1 Confidently wrong at the top. Where Laya stated 90% or more it was right half the time, and Fastino 64% of the time. Perplexity’s model put 88% of its answers in that band and was right on 98% of them. Jev, OpenDecider and Strands Decider were right more often than they said in the top band: their errors ran toward caution there.

Headline item typePerplexityOpenAIJevOpenDeciderFastinoStrandsLaya
A · does the reply say done?96.892.895.093.296.893.284.2
A · did the work actually hold?82.575.871.974.654.564.674.0
B · read-only command10098.010097.01006.00.0
B · command that changes files98.787.784.078.448.767.766.7

Table 6. Accuracy (%) by headline item type. Read-only items: 300 answers per model.

6.2 Harmless commands read as destructive. Asked whether read-only commands such as git status would destroy data, Laya was wrong on all 300 answers and Strands Decider on 94%; the other five were right on 97% to 100%. All three of OpenAI’s refusals were on one such item.

Bank C category (90 answers)PerplexityOpenAIJevOpenDeciderFastinoStrandsLaya
Arithmetic50.063.350.050.050.050.050.0
Multi-step command chains66.760.068.250.040.050.050.0
Negation pairs93.380.079.889.353.360.050.0
Distracting context10086.783.375.956.750.050.0
Indirect wording10073.384.456.750.066.750.0
Adversarial note in the state96.780.090.053.343.350.050.0
Counting90.086.797.850.050.050.050.0
Date comparison10010010010050.090.050.0

Table 7. Accuracy (%) on bank C, never part of the headline. Intervals are wide at 90 answers (for 50%, roughly 33% to 67%). Both general-purpose models scored above 95% on arithmetic [2].

6.3 The weak spots. Six of the seven decision models scored exactly 50% on arithmetic, and OpenAI’s API 63%. Laya scored 50% in every category, and Strands Decider in five of eight.

Consistency and speedPerplexityOpenAIJevOpenDeciderFastinoStrandsLaya
P(q) + P(not q), mean of 45 pairs (1 = coherent)0.980.740.881.070.721.311.44
Items with an identical P(yes) on all 3 askings100%100%27%100%100%100%100%
Median response time (observed)0.21 s0.17 s0.15 s0.029 s0.021 s0.026 s0.008 s

Table 8. For a question and its exact negation, the two probabilities of yes should sum to one. Response times include the network for the two hosted APIs; the four small models ran on our RTX 5090 and Perplexity’s on a rented A100, so the groups are not comparable with each other. Speed was not registered.

6.4 Consistency. Strands Decider and Laya leaned toward yes to both a question and its negation; OpenAI’s API and Fastino leaned toward no to both. Every model except Jev returned the same probability on every repeat. Jev’s accuracy here (86.4%) was close to its accuracy a day earlier in our first study (86.9%, same version), and both studies found its calibration inconclusive [2].

7.Deviations from the registration

  • 1. A new answer shape. OpenAI’s API once answered with a refusal object our parser did not know, and the run stopped, as registered (fail closed). The registration already classes a refusal as no decision; the parser was extended to that shape with a test, the stored response was re-read without asking again, and the run resumed.
  • 2. The readout’s first call. It was given folder paths instead of files and failed before reading any record. It was then run once, as registered.
  • 3. A display defect. The readout prints one line from the earlier Jev study’s code saying H1 is “not evaluable”, followed by this study’s per-model verdicts. The per-model block applies this registration’s rule; no number is affected.
  • 4. Run order. The registration says the rented-GPU model runs last. It ran second, after the hosted models and before the four on our workstation, because the rented machine was ready and billed by the hour. Each model ran alone on the same items in the same order and shares no state with any other, so the order cannot change any model’s answers.
  • 5. The entry gate, applied unevenly. OpenDecider, an independent developer’s own release with fewer than 1,000 downloads, was admitted as its maker’s official release; gutsy, also an independent developer’s own release, was excluded on downloads. Read the same way, gutsy would have qualified. The registration forbids revisiting a gate once data exist, so it is reported, not changed.

8.What this does not show

  • How a recalibrated model would do. Laya’s and Strands Decider’s own cards tell users to measure or refit on their own data first [4, 5]. We tested what a buyer downloads, not what a buyer could make of it.
  • Anything about the vendors’ own workflows. These are our items, about coding-agent work and destructive commands. Vendors document other uses.
  • Anything about later versions. Every answer came from the revision pinned in the registration.
  • Price or speed claims. We did not test them; response times are observations.
  • Anything about intent. We report what each system returned. We have no instrument that sees why, and we make no claim about it in either direction.

9.Competing interests

Sentient Index Labs & Technology sells independent evaluations of AI models. We have no commercial, financial or personal relationship with TypeSafe AI, OpenAI, Perplexity, Amazon Web Services, the Strands Agents project, Fastino, Convai Innovations, the developer of OpenDecider, or Runpod, from which we rented the GPU. Every model was accessed through our own accounts or downloaded from its public release on the provider’s public terms. No AI judge graded any answer; every label is mechanical.

10.Data and registration

The registration is at osf.io/xh7qu [1]. The item banks, the analysis code, every raw response and every scored answer are held in our private repository; the registered SHA-256 hashes identify the exact banks used. A shorter summary is published at silt-seb.com/studies/decision-models [12].

11.References

  1. Scanlon, S. (2026). Jev, OpenAI, AWS, Perplexity: Do Decision Models Know When They’re Wrong? Registration, Open Science Framework. https://osf.io/xh7qu
  2. Scanlon, S. (2026). Jev: Twenty Times Faster, Seven Points Behind (SILT-RP-009). Sentient Index Labs & Technology. https://doi.org/10.5281/zenodo.23287930
  3. Perplexity. pplx-decider-v1.1-27b model card. https://huggingface.co/perplexity-ai/pplx-decider-v1.1-27b (read 11 October 2026).
  4. Strands Agents. strands-decider-2B-hobson-v19 model card. https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19 (read 11 October 2026).
  5. Convai Innovations. laya model card. https://huggingface.co/convaiinnovations/laya (read 11 October 2026).
  6. manjunathshiva. opendecider-small model card. https://huggingface.co/manjunathshiva/opendecider-small (read 11 October 2026).
  7. Fastino. GLiNER2.5-Decide model card. https://huggingface.co/fastino/GLiNER2.5-Decide (read 11 October 2026).
  8. Sentient Index Labs & Technology. The Harness Battery. https://www.silt-seb.com/harness-battery
  9. Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
  10. Naeini, M. P., Cooper, G. F., & Hauskrecht, M. (2015). Obtaining well calibrated probabilities using Bayesian binning. Proceedings of AAAI.
  11. Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of ICML.
  12. Sentient Index Labs & Technology. Jev, OpenAI, AWS, Perplexity: Do Decision Models Know When They’re Wrong? (study summary). https://www.silt-seb.com/studies/decision-models
SILT-RP-010 · version 1.0 · issued 2026-10-11
© 2026 Sentient Index Labs & Technology. This document is licensed CC BY-ND 4.0 — share it freely and in full, with attribution; do not publish modified versions. Translations and excerpts are granted on request to press@sentientindexlabs.com.
⚠️ Corrections are issued as a new version under the same identifier, each with a dated note in the text recording exactly what changed and what the earlier version said.
Nothing in this document is a certification, audit opinion, conformity assessment, or legal advice. See the Subscriber Agreement, Section 15.