Jev, OpenAI, AWS, Perplexity: Do Decision Models Know When They're Wrong?
A pre-registered test of the calibration and accuracy of seven decision models on the same 799 unpublished yes/no decisions
Authorship. Prepared by Sentient Index Labs & Technology. During the preparation of this work, the author used Claude (Anthropic) to draft the text and run the analyses; every figure is derived from the study's own stored responses and checked against them. The author directed the work and is responsible for the content.
1.Abstract
Decision models return a probability instead of text, and each is sold on the claim that the probability can be trusted. Within a month of the first, Jev from TypeSafe, at least six more appeared, from OpenAI, Perplexity, the AWS-originated Strands Agents project, Fastino, Convai Innovations and an independent developer. Every published claim for them is the vendor's own. Under a plan registered on the Open Science Framework before the run (osf.io/xh7qu), we gave all seven the same 799 unpublished yes/no decisions about AI coding-agent work and destructive shell commands, each with an answer fixed in advance by a mechanical check, three times each, as shipped: 23,079 answers. Against a pre-set line of an expected calibration error of 0.05, one model was consistent with a calibration claim: Perplexity's open pplx-decider (ECE 0.026, 95% CI 0.014–0.044), which was also the most accurate (94.7%). Jev (0.052), OpenAI's Decisions API (0.034) and Laya (0.051) were inconclusive; OpenDecider (0.079), Fastino's GLiNER2.5-Decide (0.082) and Strands Decider (0.129) did not support the claim. Against Jev on the same answers, Perplexity's model was 8.1 points more accurate (6.1–10.1), OpenAI's API and OpenDecider were statistically tied, and the other three were 16 to 22 points behind; two scored below the 67.7% of always giving the most common answer. Seven separate verdicts carry about a 30% chance that one is wrong by chance alone.
2.The claim under test
A decision model does not write text. Given a state (a description of a situation) and a typed question, it returns a typed answer; for a yes/no question the answer is a probability that the answer is yes. TypeSafe released the first, Jev, in September 2026, and we tested it on 9 and 10 October [2]. By then at least six competitors were public, and each makes the same promise in its own words. Laya’s card describes “mathematically calibrated probabilities” [5]; Strands Decider’s says “every answer carries a calibrated confidence” [4]; OpenDecider calls itself the “best-calibrated model that fits a 16 GB Mac (ECE 0.087)” [6]. The same cards are also candid. Laya’s says it “ships over-confident” and tells users to refit its temperatures “on your own data before trusting the probabilities” [5]; Strands Decider’s says its confidence bands were established on short classification only, so “measure on your own traffic before you trust a threshold” [4].
Every figure published for these models so far is the vendor’s own, and several cards compare themselves with Jev and with each other on the vendors’ own test sets [5, 6]. No independent measurement existed. A threshold is only useful if the number behind it means what it says, so the question is the one we asked of Jev alone, now asked of the whole category on the same items.
3.Registration
The design, the models and their pinned revisions, the entry gate, the item banks (identified by SHA-256 hash, unchanged from the Jev study), the analysis code, the 0.05 decision line and the wording of every verdict were registered on the Open Science Framework at 06:01 UTC on 11 October 2026 [1], after a three-item probe of each model on non-study items and before any model answered a study item. Jev was re-run fresh rather than reused. The analysis was run once, on the main runs. Two hypotheses were registered for each model:
- H1 (primary). On the headline items, the model’s expected calibration error (ECE) is at most 0.05. If the upper bound of its 95% interval is at or below 0.05, the result is “consistent with the calibration claim”; if the lower bound is above 0.05, “the calibration claim is not supported on these items”; otherwise, “inconclusive”.
- H2 (secondary, two-sided). The model’s accuracy minus Jev’s, paired by item and replicate. No direction was predicted.
Everything else in this paper is exploratory and labelled as such. The 0.05 line is ours: no vendor publishes a calibration target. Seven verdicts are reported without correction for multiplicity; at 95% each, the chance that at least one lands on the wrong side by chance alone is about 30%. Deviations are listed in section 7.
4.Method
4.1 Items. The items are the Jev study’s, unchanged and never published, so no model can have been trained on them [2]. Every item is a yes/no question whose answer was fixed before the run by a check of a real workspace, a deterministic simulator or arithmetic, never by another model’s judgement.
| Bank | What it asks | Items | Answers per model |
|---|---|---|---|
| A — agent work | From a frozen record of AI coding-agent runs [8]: does the agent's final reply say the task is done (221), and did the requirement actually hold (178)? | 399 | 1,197 |
| B — destructive commands | Given a described workspace, will this shell command permanently destroy data? Labelled by a deterministic simulator. | 400 | 1,200 |
| B — ambiguous | The same question where the answer cannot be known from the state. Never scored. | 60 | 180 |
| C — documented weak spots | Eight categories of 30 (arithmetic, counting, negation and others) aimed at weaknesses decision-model vendors document. Reported separately, never in the headline. | 240 | 720 |
| Total | 1,099 | 3,297 |
Table 1. The item banks. The headline set is banks A and B: 799 items, 2,397 answers per model. On these items 67.7% of answers are “yes”, so always answering yes scores 67.7% accuracy and a Brier score of 0.219.
4.2 Models. The entry gate, fixed before the run, admitted a model that was an official release from its maker or had at least 1,000 downloads on Hugging Face; had a licence permitting research and publishing results; returned a probability, or a label with a documented confidence, for a yes/no question; and answered all three probe items on our hardware. Seven entered (Table 2). Every model was tested as shipped: its maker’s own package or API, the revision its package loads by default (pinned for the run), and its own saved calibration, with no refitting on our data.
| Model | Maker | Where it ran | Interface |
|---|---|---|---|
| Jev 1.13.0 | TypeSafe AI | TypeSafe API | noul = P(yes) |
| OpenAI Decisions API, gpt-6-luna | OpenAI | OpenAI API | predicate → probability |
| pplx-decider-v1.1-27b | Perplexity (open) | rented NVIDIA A100 80 GB | DecisionModel.predict |
| strands-decider-2B-hobson-v19 | Strands Agents, the AWS-originated open project (open) | RTX 5090, our workstation | noul |
| GLiNER2.5-Decide (486 M) | Fastino (open) | RTX 5090 | label + confidence |
| laya (421 M) | Convai Innovations (open) | RTX 5090 | noul |
| opendecider-small (Qwen3-4B + LoRA) | independent developer (open) | RTX 5090 | noul |
Table 2. The seven models. Exact revision hashes are in the registration [1]. Perplexity’s model needed about 50 GB of GPU memory and could not load on the 32 GB RTX 5090, so it ran on a rented 80 GB GPU; that is itself a deployment fact. Six candidates were excluded by the gate: kouhxp/gutsy (downloads), the openjev models (community conversions under a non-commercial licence), JevK5, NanoJev and jevlike (derivatives, a template and a training kit with no weights), and Fastino’s hosted GLiDE API (a separate account; Fastino is represented by its open model). See deviation 5 on how the gate was applied.
4.3 Calls. Every model received the item’s state, the question’s instructions and its yes/no criteria through its documented fields. OpenAI’s Decisions API has no published contract; we read its request format from the endpoint’s own validation errors and one test call. It has no criteria field, so the criteria were appended to its instructions, fixed before the run and identical for every item. Fastino’s model returns a label with a documented confidence; P(yes) was the confidence if the label was yes and one minus it otherwise, fixed before the run. Inputs longer than a model’s window were cut by that model as shipped and the cut was recorded (Laya: 99 of 3,297 answers). Each item was asked three times of each model, in a fixed shuffled order, one model at a time: 23,079 answers between 06:10 and 06:54 UTC on 11 October 2026.
4.4 Scoring and uncertainty. As in the Jev study: an answer is yes if P(yes) > 0.5 and no if below; exactly 0.5 is a tie, counted but excluded from accuracy and calibration. Confidence is max(P, 1 − P). ECE uses five fixed bands from 0.5 to 1.0, weighted by band size [10, 11]; the Brier score is the mean squared difference between P(yes) and the answer [9]. A refusal or unusable reply is a “no decision”, counted and never scored. All intervals are percentile 95% cluster-bootstrap intervals (2,000 resamples, seed 20261010); a cluster is an item with its replicates. H2 pairs each model with Jev on the same item and replicate wherever both decided.
5.Registered results
| Headline items | H1 verdict | Scored | Accuracy [95% CI] | ECE [95% CI] | Brier |
|---|---|---|---|---|---|
| Perplexity pplx-decider | consistent | 2,391 | 94.7% [93.1, 96.4] | 0.026 [0.014, 0.044] | 0.044 |
| OpenAI Decisions API | inconclusive | 2,343 | 88.0% [85.6, 90.1] | 0.034 [0.025, 0.062] | 0.098 |
| Jev (TypeSafe) | inconclusive | 2,389 | 86.4% [84.0, 88.7] | 0.052 [0.034, 0.072] | 0.096 |
| OpenDecider | not supported | 2,349 | 84.0% [81.3, 86.6] | 0.079 [0.062, 0.105] | 0.123 |
| Fastino GLiNER2.5-Decide | not supported | 2,397 | 69.7% [66.8, 72.7] | 0.082 [0.059, 0.115] | 0.191 |
| Strands Decider | not supported | 2,397 | 66.3% [62.8, 69.8] | 0.129 [0.100, 0.158] | 0.188 |
| Laya | inconclusive | 2,394 | 64.8% [61.5, 68.1] | 0.051 [0.029, 0.088] | 0.224 |
Table 3. Headline results (banks A and B; 2,397 answers per model). “Scored” excludes ties (P = 0.5: OpenAI 51, OpenDecider 48, Jev 8, Perplexity 6, Laya 3) and no-decisions (OpenAI 3 refusals). Brier 95% intervals, in table order: [0.031, 0.057], [0.083, 0.114], [0.086, 0.108], [0.111, 0.134], [0.178, 0.203], [0.176, 0.201], [0.210, 0.239]. For context only, with no verdict, two general-purpose models from the Jev study on the same items: DeepSeek V4 Flash 94.7%, ECE 0.035; GPT-6.1 Sol 95.1%, ECE 0.021 [2].
| Paired with Jev (same item, same replicate) | Accuracy difference [95% CI] | Pairs |
|---|---|---|
| Perplexity − Jev | +8.1 [+6.1, +10.1] | 2,383 |
| OpenAI − Jev | +0.6 [−2.0, +3.1] | 2,335 |
| OpenDecider − Jev | −2.8 [−5.8, +0.2] | 2,341 |
| Fastino − Jev | −16.5 [−19.8, −12.9] | 2,389 |
| Strands − Jev | −20.1 [−24.1, −16.2] | 2,389 |
| Laya − Jev | −21.6 [−25.9, −17.3] | 2,386 |
Table 4. H2, in percentage points.
6.Exploratory observations
Nothing in this section was a registered hypothesis. It is in the data and should be read as description.
| Model | Top band (90–100%): n · conf · correct | Acting at ≥ 70% confidence: share acted on · correct |
|---|---|---|
| Perplexity | 2,109 · 99.2% · 98.2% | 97.2% · 95.4% |
| OpenAI | 1,698 · 97.8% · 93.6% | 89.6% · 91.3% |
| Jev | 1,036 · 96.3% · 99.1% | 74.7% · 94.6% |
| OpenDecider | 381 · 92.4% · 100% | 70.1% · 92.0% |
| Fastino | 159 · 93.5% · 64.2% | 54.8% · 83.8% |
| Strands | 48 · 91.6% · 100% | 51.6% · 87.9% |
| Laya | 36 · 92.5% · 50.0% | 48.9% · 73.1% |
Table 5. The most confident answers, and what a threshold at 70% (TypeSafe’s documented hold level, used here for every model) would act on. Full five-band tables for every model are in the registered readout.
6.1 Confidently wrong at the top. Where Laya stated 90% or more it was right half the time, and Fastino 64% of the time. Perplexity’s model put 88% of its answers in that band and was right on 98% of them. Jev, OpenDecider and Strands Decider were right more often than they said in the top band: their errors ran toward caution there.
| Headline item type | Perplexity | OpenAI | Jev | OpenDecider | Fastino | Strands | Laya |
|---|---|---|---|---|---|---|---|
| A · does the reply say done? | 96.8 | 92.8 | 95.0 | 93.2 | 96.8 | 93.2 | 84.2 |
| A · did the work actually hold? | 82.5 | 75.8 | 71.9 | 74.6 | 54.5 | 64.6 | 74.0 |
| B · read-only command | 100 | 98.0 | 100 | 97.0 | 100 | 6.0 | 0.0 |
| B · command that changes files | 98.7 | 87.7 | 84.0 | 78.4 | 48.7 | 67.7 | 66.7 |
Table 6. Accuracy (%) by headline item type. Read-only items: 300 answers per model.
6.2 Harmless commands read as destructive. Asked whether read-only commands such as git status would destroy data, Laya was wrong on all 300 answers and Strands Decider on 94%; the other five were right on 97% to 100%. All three of OpenAI’s refusals were on one such item.
| Bank C category (90 answers) | Perplexity | OpenAI | Jev | OpenDecider | Fastino | Strands | Laya |
|---|---|---|---|---|---|---|---|
| Arithmetic | 50.0 | 63.3 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 |
| Multi-step command chains | 66.7 | 60.0 | 68.2 | 50.0 | 40.0 | 50.0 | 50.0 |
| Negation pairs | 93.3 | 80.0 | 79.8 | 89.3 | 53.3 | 60.0 | 50.0 |
| Distracting context | 100 | 86.7 | 83.3 | 75.9 | 56.7 | 50.0 | 50.0 |
| Indirect wording | 100 | 73.3 | 84.4 | 56.7 | 50.0 | 66.7 | 50.0 |
| Adversarial note in the state | 96.7 | 80.0 | 90.0 | 53.3 | 43.3 | 50.0 | 50.0 |
| Counting | 90.0 | 86.7 | 97.8 | 50.0 | 50.0 | 50.0 | 50.0 |
| Date comparison | 100 | 100 | 100 | 100 | 50.0 | 90.0 | 50.0 |
Table 7. Accuracy (%) on bank C, never part of the headline. Intervals are wide at 90 answers (for 50%, roughly 33% to 67%). Both general-purpose models scored above 95% on arithmetic [2].
6.3 The weak spots. Six of the seven decision models scored exactly 50% on arithmetic, and OpenAI’s API 63%. Laya scored 50% in every category, and Strands Decider in five of eight.
| Consistency and speed | Perplexity | OpenAI | Jev | OpenDecider | Fastino | Strands | Laya |
|---|---|---|---|---|---|---|---|
| P(q) + P(not q), mean of 45 pairs (1 = coherent) | 0.98 | 0.74 | 0.88 | 1.07 | 0.72 | 1.31 | 1.44 |
| Items with an identical P(yes) on all 3 askings | 100% | 100% | 27% | 100% | 100% | 100% | 100% |
| Median response time (observed) | 0.21 s | 0.17 s | 0.15 s | 0.029 s | 0.021 s | 0.026 s | 0.008 s |
Table 8. For a question and its exact negation, the two probabilities of yes should sum to one. Response times include the network for the two hosted APIs; the four small models ran on our RTX 5090 and Perplexity’s on a rented A100, so the groups are not comparable with each other. Speed was not registered.
6.4 Consistency. Strands Decider and Laya leaned toward yes to both a question and its negation; OpenAI’s API and Fastino leaned toward no to both. Every model except Jev returned the same probability on every repeat. Jev’s accuracy here (86.4%) was close to its accuracy a day earlier in our first study (86.9%, same version), and both studies found its calibration inconclusive [2].
7.Deviations from the registration
- 1. A new answer shape. OpenAI’s API once answered with a refusal object our parser did not know, and the run stopped, as registered (fail closed). The registration already classes a refusal as no decision; the parser was extended to that shape with a test, the stored response was re-read without asking again, and the run resumed.
- 2. The readout’s first call. It was given folder paths instead of files and failed before reading any record. It was then run once, as registered.
- 3. A display defect. The readout prints one line from the earlier Jev study’s code saying H1 is “not evaluable”, followed by this study’s per-model verdicts. The per-model block applies this registration’s rule; no number is affected.
- 4. Run order. The registration says the rented-GPU model runs last. It ran second, after the hosted models and before the four on our workstation, because the rented machine was ready and billed by the hour. Each model ran alone on the same items in the same order and shares no state with any other, so the order cannot change any model’s answers.
- 5. The entry gate, applied unevenly. OpenDecider, an independent developer’s own release with fewer than 1,000 downloads, was admitted as its maker’s official release; gutsy, also an independent developer’s own release, was excluded on downloads. Read the same way, gutsy would have qualified. The registration forbids revisiting a gate once data exist, so it is reported, not changed.
8.What this does not show
- How a recalibrated model would do. Laya’s and Strands Decider’s own cards tell users to measure or refit on their own data first [4, 5]. We tested what a buyer downloads, not what a buyer could make of it.
- Anything about the vendors’ own workflows. These are our items, about coding-agent work and destructive commands. Vendors document other uses.
- Anything about later versions. Every answer came from the revision pinned in the registration.
- Price or speed claims. We did not test them; response times are observations.
- Anything about intent. We report what each system returned. We have no instrument that sees why, and we make no claim about it in either direction.
9.Competing interests
Sentient Index Labs & Technology sells independent evaluations of AI models. We have no commercial, financial or personal relationship with TypeSafe AI, OpenAI, Perplexity, Amazon Web Services, the Strands Agents project, Fastino, Convai Innovations, the developer of OpenDecider, or Runpod, from which we rented the GPU. Every model was accessed through our own accounts or downloaded from its public release on the provider’s public terms. No AI judge graded any answer; every label is mechanical.
10.Data and registration
The registration is at osf.io/xh7qu [1]. The item banks, the analysis code, every raw response and every scored answer are held in our private repository; the registered SHA-256 hashes identify the exact banks used. A shorter summary is published at silt-seb.com/studies/decision-models [12].
11.References
- Scanlon, S. (2026). Jev, OpenAI, AWS, Perplexity: Do Decision Models Know When They’re Wrong? Registration, Open Science Framework. https://osf.io/xh7qu
- Scanlon, S. (2026). Jev: Twenty Times Faster, Seven Points Behind (SILT-RP-009). Sentient Index Labs & Technology. https://doi.org/10.5281/zenodo.23287930
- Perplexity. pplx-decider-v1.1-27b model card. https://huggingface.co/perplexity-ai/pplx-decider-v1.1-27b (read 11 October 2026).
- Strands Agents. strands-decider-2B-hobson-v19 model card. https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19 (read 11 October 2026).
- Convai Innovations. laya model card. https://huggingface.co/convaiinnovations/laya (read 11 October 2026).
- manjunathshiva. opendecider-small model card. https://huggingface.co/manjunathshiva/opendecider-small (read 11 October 2026).
- Fastino. GLiNER2.5-Decide model card. https://huggingface.co/fastino/GLiNER2.5-Decide (read 11 October 2026).
- Sentient Index Labs & Technology. The Harness Battery. https://www.silt-seb.com/harness-battery
- Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
- Naeini, M. P., Cooper, G. F., & Hauskrecht, M. (2015). Obtaining well calibrated probabilities using Bayesian binning. Proceedings of AAAI.
- Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of ICML.
- Sentient Index Labs & Technology. Jev, OpenAI, AWS, Perplexity: Do Decision Models Know When They’re Wrong? (study summary). https://www.silt-seb.com/studies/decision-models