Jev: Twenty Times Faster, Seven Points Behind
Does Jev know when it is wrong? A pre-registered test of a decision model's calibration and accuracy beside two general-purpose models
Authorship. Prepared by Sentient Index Labs & Technology. During the preparation of this work, the author used Claude (Anthropic) to draft the text and run the analyses; every figure is derived from the study's own stored responses and checked against them. The author directed the work and is responsible for the content.
1.Abstract
Jev, from TypeSafe AI, is a decision model: given a situation and a yes/no question, it returns only the probability of yes, and it is sold as calibrated ("higher confidence means higher accuracy"). We found no published measurement of that claim. We tested it under a plan registered on the Open Science Framework before any data were collected (osf.io/qg6nc), on 799 yes/no decisions about AI coding-agent work and destructive shell commands, each with an answer fixed in advance and none taken from another model's judgement. Two general-purpose models, DeepSeek V4 Flash and GPT-6.1 Sol, answered the same items, each item three times per model: 9,891 answers in all. The registered calibration test was inconclusive: Jev's expected calibration error was 0.045 (95% CI 0.031–0.068) against a pre-set line of 0.05. Its accuracy was lower than both comparators on the same items, by 6.6 points (95% CI 4.7–8.6) and 8.2 points (6.2–10.2). Where Jev was miscalibrated it erred toward caution, while both comparators were overconfident on the few answers they gave below 90% confidence. Observed but not registered: Jev answered in a median 0.18 seconds, about twenty times faster than DeepSeek V4 Flash, and cost the least per answer. On the weak spots TypeSafe itself documents, reported separately, its arithmetic was at chance.
2.The claim under test
Jev is the first “System One” model from TypeSafe AI. It does not write text. Given a state (a description of a situation) and a typed question, it returns a typed answer; for a yes/no question, which TypeSafe calls a noul, the answer is a single number, the probability that the answer is yes. TypeSafe describes the model as “Calibrated: higher confidence means higher accuracy” [2] and invites developers to “set the thresholds for when it acts autonomously and when it asks for review” [3]. Jev’s pitch rests on that claim: a threshold is only useful if the number behind it means what it says.
On 9 and 10 October 2026 we found no published measurement of the claim on TypeSafe’s website, blog or documentation: no reliability diagram, calibration error, Brier score or sample size. Its published comparisons measure accuracy as agreement with the average of two other language models (“in this case, Astra and Fable”) on four workflows its own team wrote [2], so by construction they cannot show where those models are wrong. TypeSafe states several of these limits itself, including that its speed figures were timed “from our laptops on the West Coast” and that “we can’t prove it isn’t subsidized” [2]. Its documentation also lists known weak spots for this version, among them arithmetic, counting, date comparison, indirection, irrelevant detail and adversarial content [4].
3.Registration and deviations
The design, the item banks (frozen and identified by SHA-256 hash), the analysis code, the 0.05 decision line and the wording of every possible verdict were registered on the Open Science Framework on 9 October 2026, before any model saw a study item [1]. The analysis was run once, unchanged, on the main run. Two hypotheses were registered:
- H1 (primary). On the headline items, Jev’s expected calibration error (ECE) is at most 0.05. If the upper bound of its 95% interval is at or below 0.05, the result is “consistent with the calibration claim”; if the lower bound is above 0.05, “the calibration claim is not supported on these items”; otherwise, “inconclusive”.
- H2 (secondary, two-sided). Jev’s accuracy differs from each comparator’s on the same items.
Everything else in this paper is exploratory and is labelled as such. The 0.05 line is ours: TypeSafe publishes no calibration target.
Deviations: none. The run was paused once, at 720 of 9,891 answers, to move it from one of our machines to another, and resumed with the same code and run identifier; every completed answer was kept and none was repeated. The registered order of calls was unchanged, so we record this as a pause rather than a deviation.
4.Method
4.1 Items. Every item is a yes/no question whose answer was fixed before the run by a mechanical check or a simulator, never by another model’s judgement. There are three banks.
| Bank | What it asks | Items | Answers per model |
|---|---|---|---|
| A — agent work | From a frozen record of AI coding-agent runs [9], each already checked against the workspace and read by hand: does the agent's final reply say the task is done (221), and after its work, did the requirement actually hold (178)? | 399 | 1,197 |
| B — destructive commands | Given a described workspace (files committed, modified, untracked, ignored, backed up), will this shell command permanently destroy data? Labelled by a deterministic simulator. | 400 | 1,200 |
| B — ambiguous | The same question where the answer cannot be known from the state (for example, an unseen script). Never scored. | 60 | 180 |
| C — documented weak spots | Eight categories of 30, each aimed at a weakness TypeSafe lists [4]. Reported separately, never in the headline. | 240 | 720 |
| Total | 1,099 | 3,297 |
Table 1. The item banks. The headline set is banks A and B together: 799 items, 2,397 answers per model. Three exclusions were fixed before the run: one test’s outcome items (47) whose answer could not be determined from the final reply alone, 8 claim items with an empty agent reply, and 4 outcome items whose workspace check was undetermined.
4.2 Models and calls. Jev was called through TypeSafe’s own API in its documented noul format, with no sampling parameters (none are documented). Every response reported the model version jev-1.13.0; the run was set to stop if that changed, and it did not. The two comparators, DeepSeek V4 Flash and OpenAI GPT-6.1 Sol, were called through OpenRouter, received the same state and question, and were required to reply with only a yes/no answer and a probability, at temperature 1.0 and a 2,000-token output limit. Each item was asked three times of each model, in a fixed shuffled order: 9,891 answers between 19:49 PDT on 9 October and 04:00 PDT on 10 October 2026.
4.3 Scoring. An answer is “yes” if P(yes) > 0.5 and “no” if it is below; exactly 0.5 is a tie, counted but excluded from accuracy and calibration. Confidence is max(P, 1 − P). This is our derivation, because a noul carries no separate confidence. ECE uses five fixed bands from 0.5 to 1.0 and weights each band’s gap between mean confidence and share correct by its size. The Brier score is the mean squared difference between P(yes) and the answer [6]; on calibration error see [7, 8]. A missing or unusable reply is a “no decision”, counted separately and never scored as right or wrong.
4.4 Uncertainty. All intervals are percentile 95% cluster-bootstrap intervals (2,000 resamples, seed fixed in the registration). A cluster is an item with its three replicates; in bank A, items built from the same agent run share a cluster. H2 compares the two models on the same item and replicate, wherever both decided.
5.Registered results
| Headline items | Jev | DeepSeek V4 Flash | GPT-6.1 Sol |
|---|---|---|---|
| Answers | 2,397 | 2,397 | 2,397 |
| No decision | 0 | 115 | 0 |
| Ties (P = 0.5) | 20 | 0 | 0 |
| Scored | 2,377 | 2,282 | 2,397 |
| Accuracy | 86.9% | 94.7% | 95.1% |
| 95% CI | 84.5–89.1% | 93.0–96.0% | 93.5–96.7% |
| ECE | 0.045 | 0.035 | 0.021 |
| 95% CI | 0.031–0.068 | 0.023–0.051 | 0.009–0.036 |
| Brier score | 0.096 | 0.046 | 0.035 |
| 95% CI | 0.085–0.107 | 0.034–0.060 | 0.025–0.047 |
Table 2. Headline results (banks A and B). On these items 67.7% of answers are “yes”, so always answering yes would score 67.7% accuracy and a Brier score of 0.219. Of DeepSeek V4 Flash’s 115 missing decisions, 111 were empty replies that reached the registered output limit.
| Confidence band | Jev: n · conf · correct | DeepSeek V4 Flash | GPT-6.1 Sol |
|---|---|---|---|
| 50–60% | 233 · 55.6% · 55.4% | 0 · — · — | 8 · suppressed |
| 60–70% | 371 · 64.8% · 69.0% | 1 · suppressed | 39 · 63.5% · 48.7% |
| 70–80% | 332 · 74.7% · 85.2% | 7 · suppressed | 64 · 73.6% · 64.1% |
| 80–90% | 403 · 84.6% · 90.8% | 46 · 82.4% · 41.3% | 121 · 82.4% · 69.4% |
| 90–100% | 1,038 · 96.3% · 99.3% | 2,228 · 98.6% · 95.8% | 2,165 · 99.5% · 98.6% |
Table 3. Reliability on the headline items: number of answers, mean stated confidence and share correct in each band. Bands with fewer than 10 answers are suppressed (registered rule).
6.Exploratory observations
Nothing in this section was a registered hypothesis. It is reported because it is in the data, and it should be read as description.
6.1 The direction of the error. Jev’s miscalibration ran toward caution: in every band above 60% it was right more often than it said (Table 3). The comparators’ ran the other way where they were not certain: DeepSeek V4 Flash was right on 41% of the 46 answers it gave at about 82% confidence, and GPT-6.1 Sol fell below the diagonal in every band under 90%. Both comparators placed over 90% of their answers in the top band, so their aggregate ECE rests mostly on that band.
6.2 At TypeSafe’s review band. TypeSafe’s own cookbook routes probabilities between 0.30 and 0.70 to human review [5]. Applying that rule to the headline items, Jev acted on 74.6% of them (1,773 of 2,377) and was right on 94.8% of those. DeepSeek V4 Flash acted on all but one and was right on 94.7%; GPT-6.1 Sol acted on 98.0% and was right on 96.2%.
6.3 The ambiguous items. On the 60 items whose answer cannot be known from the state, 79.4% of Jev’s answers fell between 0.3 and 0.7, against 2.5% for DeepSeek V4 Flash (118 decided) and 13.9% for GPT-6.1 Sol. Jev marked these as uncertain more often than either comparator did.
Figure 2. Accuracy on bank C, the categories TypeSafe lists as weak spots for this version (exploratory; 90 answers per model per category). The dashed line is chance. Values in Table 4.
| Bank C category (30 items × 3) | Jev | DeepSeek V4 Flash | GPT-6.1 Sol |
|---|---|---|---|
| Arithmetic | 50.0% | 95.5% | 100% |
| Multi-step command chains | 64.4% | 90.3% | 100% |
| Negation pairs | 80.0% | 94.4% | 100% |
| Distracting context | 83.3% | 98.9% | 100% |
| Indirect wording | 84.4% | 93.3% | 100% |
| Adversarial note in the state | 90.0% | 94.3% | 100% |
| Counting | 98.9% | 100% | 100% |
| Date comparison | 100% | 100% | 100% |
Table 4. Accuracy on the categories TypeSafe documents as weak spots [4]. Each category has 90 answers per model; intervals are wide at this size (for Jev’s arithmetic, 30% to 67%).
6.4 Where the headline gap comes from. On the headline subtypes, Jev matched the comparators when the question was what an agent’s reply said (95.1% against 94.6% and 97.0%) and on commands that change nothing (all three at 100%). The gap was on the harder judgements: whether the agent’s work actually held (71.9% against 86.3% and 81.6%) and whether a command that changes files destroys data (85.2% against 97.0% and 100%).
Figure 3. Accuracy by item type on the headline set (exploratory). The bars for each type are the three models in the order of the key. Values in §6.4.
6.5 Consistency. For a question and its negation, the two probabilities should sum to one. Jev’s averaged 0.877 (45 pairs), DeepSeek V4 Flash’s 0.979 and GPT-6.1 Sol’s 1.000. Across the three replicates, Jev returned an identical probability on 27.2% of items, against 60.6% and 77.2%; its variation was small (mean range 0.027).
| Speed and cost (all 3,297 answers) | Jev | DeepSeek V4 Flash | GPT-6.1 Sol |
|---|---|---|---|
| Median response time | 0.18 s | 3.60 s | 1.57 s |
| 90th percentile | 0.22 s | 15.41 s | 3.76 s |
| Measured cost, whole run | $0.086 | $0.492 | $3.906 |
| Per 1,000 answers | $0.026 | $0.149 | $1.185 |
Table 5. Response time as recorded by our client, including the network (Jev called directly, the comparators through OpenRouter, from the same home network throughout). Cost is the usage each API reported, priced as each provider reported or listed it on 9 October 2026; Jev’s output is free [3]. Neither was a registered measure: we said we would not test TypeSafe’s cost or speed claims, and these figures describe this run only.
6.6 Speed and cost. Jev’s median answer took 0.18 seconds, about 20 times faster than DeepSeek V4 Flash and 9 times faster than GPT-6.1 Sol, and its slow tail was short. It was also the cheapest per answer. Read beside H2: in this run, choosing Jev over DeepSeek V4 Flash saved about 12 cents per 1,000 decisions and gave about 66 more wrong answers per 1,000.
7.What this does not show
- That Jev is, or is not, calibrated. The registered test could not decide it. We do not expect a larger rerun of this design to settle it, because the estimate sits close to the line; a new test would need its own registration.
- Anything about TypeSafe’s own workflows. These are our items, about coding-agent work and destructive commands, not the workflows TypeSafe published.
- Anything about later versions. Every answer came from
jev-1.13.0, and TypeSafe says several of its documented weak spots will be fixed in later versions [4]. - A verdict on the comparators’ settings. They ran at the registered temperature and output limit; with a larger limit, DeepSeek V4 Flash’s 115 missing decisions might not have occurred.
- Anything about intent. We report what each system returned. We have no instrument that sees why, and we make no claim about it in either direction.
8.Competing interests
Sentient Index Labs & Technology sells independent evaluations of AI models. We have no commercial, financial or personal relationship with TypeSafe AI, DeepSeek, OpenAI or OpenRouter, and accessed every model through our own accounts on each provider’s public terms. The items in bank A come from our own Harness Battery study [9]. No AI judge graded any answer in this study; every label is mechanical.
9.Data and registration
The registration is public at osf.io/qg6nc [1]. The item banks, the analysis code, every raw response and every scored answer are held in our private repository; the registered SHA-256 hashes identify the exact banks used. A shorter summary of this study is published at silt-seb.com/studies/does-jev-know [10].
10.References
- Scanlon, S. (2026). Does Jev Know When It’s Wrong? Registration, Open Science Framework. https://osf.io/qg6nc
- TypeSafe AI. Introducing System One models and Jev. https://typesafe.ai/blog/introducing-system-one-models-and-jev (read 10 October 2026).
- TypeSafe AI. Home page. https://typesafe.ai/ (read 10 October 2026).
- TypeSafe AI. Jev 1.13 jaggedness (last reviewed 2 October 2026). https://docs.typesafe.ai/model-jaggedness/jev-1.13 (read 10 October 2026).
- TypeSafe AI. Self-consistency: nouls (cookbook). https://docs.typesafe.ai/cookbooks/consistency_noul_cookbook (read 10 October 2026).
- Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
- Naeini, M. P., Cooper, G. F., & Hauskrecht, M. (2015). Obtaining well calibrated probabilities using Bayesian binning. Proceedings of AAAI.
- Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of ICML.
- Sentient Index Labs & Technology. The Harness Battery. https://www.silt-seb.com/harness-battery
- Sentient Index Labs & Technology. Jev: Twenty Times Faster, Seven Points Behind (study summary). https://www.silt-seb.com/studies/does-jev-know