SILT Research Series · SILT-RP-009

Jev: Twenty Times Faster, Seven Points Behind

Does Jev know when it is wrong? A pre-registered test of a decision model's calibration and accuracy beside two general-purpose models

Version 1.0Issued 2026-10-10Shawn Scanlon (ORCID 0009-0002-2840-8828) · Sentient Index Labs & TechnologyDOI 10.5281/zenodo.23287930

Authorship. Prepared by Sentient Index Labs & Technology. During the preparation of this work, the author used Claude (Anthropic) to draft the text and run the analyses; every figure is derived from the study's own stored responses and checked against them. The author directed the work and is responsible for the content.

This document is fixed. It describes the instrument as of version 1.0, issued 2026-10-10, and it will not be edited. Corrections and revisions are issued as a new version under the same identifier, each with a dated note recording exactly what changed, so a citation to an earlier version can still be read against this one. Where this paper and our website disagree, the website is more current and this document is what was true on the date above.

1.Abstract

Jev, from TypeSafe AI, is a decision model: given a situation and a yes/no question, it returns only the probability of yes, and it is sold as calibrated ("higher confidence means higher accuracy"). We found no published measurement of that claim. We tested it under a plan registered on the Open Science Framework before any data were collected (osf.io/qg6nc), on 799 yes/no decisions about AI coding-agent work and destructive shell commands, each with an answer fixed in advance and none taken from another model's judgement. Two general-purpose models, DeepSeek V4 Flash and GPT-6.1 Sol, answered the same items, each item three times per model: 9,891 answers in all. The registered calibration test was inconclusive: Jev's expected calibration error was 0.045 (95% CI 0.031–0.068) against a pre-set line of 0.05. Its accuracy was lower than both comparators on the same items, by 6.6 points (95% CI 4.7–8.6) and 8.2 points (6.2–10.2). Where Jev was miscalibrated it erred toward caution, while both comparators were overconfident on the few answers they gave below 90% confidence. Observed but not registered: Jev answered in a median 0.18 seconds, about twenty times faster than DeepSeek V4 Flash, and cost the least per answer. On the weak spots TypeSafe itself documents, reported separately, its arithmetic was at chance.

2.The claim under test

Jev is the first “System One” model from TypeSafe AI. It does not write text. Given a state (a description of a situation) and a typed question, it returns a typed answer; for a yes/no question, which TypeSafe calls a noul, the answer is a single number, the probability that the answer is yes. TypeSafe describes the model as “Calibrated: higher confidence means higher accuracy” [2] and invites developers to “set the thresholds for when it acts autonomously and when it asks for review” [3]. Jev’s pitch rests on that claim: a threshold is only useful if the number behind it means what it says.

On 9 and 10 October 2026 we found no published measurement of the claim on TypeSafe’s website, blog or documentation: no reliability diagram, calibration error, Brier score or sample size. Its published comparisons measure accuracy as agreement with the average of two other language models (“in this case, Astra and Fable”) on four workflows its own team wrote [2], so by construction they cannot show where those models are wrong. TypeSafe states several of these limits itself, including that its speed figures were timed “from our laptops on the West Coast” and that “we can’t prove it isn’t subsidized” [2]. Its documentation also lists known weak spots for this version, among them arithmetic, counting, date comparison, indirection, irrelevant detail and adversarial content [4].

The question this paper answers: on decisions with answers known in advance, does Jev’s stated confidence match how often it is right, and is it as accurate as general-purpose models asked the same thing?

3.Registration and deviations

The design, the item banks (frozen and identified by SHA-256 hash), the analysis code, the 0.05 decision line and the wording of every possible verdict were registered on the Open Science Framework on 9 October 2026, before any model saw a study item [1]. The analysis was run once, unchanged, on the main run. Two hypotheses were registered:

  • H1 (primary). On the headline items, Jev’s expected calibration error (ECE) is at most 0.05. If the upper bound of its 95% interval is at or below 0.05, the result is “consistent with the calibration claim”; if the lower bound is above 0.05, “the calibration claim is not supported on these items”; otherwise, “inconclusive”.
  • H2 (secondary, two-sided). Jev’s accuracy differs from each comparator’s on the same items.

Everything else in this paper is exploratory and is labelled as such. The 0.05 line is ours: TypeSafe publishes no calibration target.

Deviations: none. The run was paused once, at 720 of 9,891 answers, to move it from one of our machines to another, and resumed with the same code and run identifier; every completed answer was kept and none was repeated. The registered order of calls was unchanged, so we record this as a pause rather than a deviation.

4.Method

4.1 Items. Every item is a yes/no question whose answer was fixed before the run by a mechanical check or a simulator, never by another model’s judgement. There are three banks.

BankWhat it asksItemsAnswers per model
A — agent workFrom a frozen record of AI coding-agent runs [9], each already checked against the workspace and read by hand: does the agent's final reply say the task is done (221), and after its work, did the requirement actually hold (178)?3991,197
B — destructive commandsGiven a described workspace (files committed, modified, untracked, ignored, backed up), will this shell command permanently destroy data? Labelled by a deterministic simulator.4001,200
B — ambiguousThe same question where the answer cannot be known from the state (for example, an unseen script). Never scored.60180
C — documented weak spotsEight categories of 30, each aimed at a weakness TypeSafe lists [4]. Reported separately, never in the headline.240720
Total1,0993,297

Table 1. The item banks. The headline set is banks A and B together: 799 items, 2,397 answers per model. Three exclusions were fixed before the run: one test’s outcome items (47) whose answer could not be determined from the final reply alone, 8 claim items with an empty agent reply, and 4 outcome items whose workspace check was undetermined.

4.2 Models and calls. Jev was called through TypeSafe’s own API in its documented noul format, with no sampling parameters (none are documented). Every response reported the model version jev-1.13.0; the run was set to stop if that changed, and it did not. The two comparators, DeepSeek V4 Flash and OpenAI GPT-6.1 Sol, were called through OpenRouter, received the same state and question, and were required to reply with only a yes/no answer and a probability, at temperature 1.0 and a 2,000-token output limit. Each item was asked three times of each model, in a fixed shuffled order: 9,891 answers between 19:49 PDT on 9 October and 04:00 PDT on 10 October 2026.

4.3 Scoring. An answer is “yes” if P(yes) > 0.5 and “no” if it is below; exactly 0.5 is a tie, counted but excluded from accuracy and calibration. Confidence is max(P, 1 − P). This is our derivation, because a noul carries no separate confidence. ECE uses five fixed bands from 0.5 to 1.0 and weights each band’s gap between mean confidence and share correct by its size. The Brier score is the mean squared difference between P(yes) and the answer [6]; on calibration error see [7, 8]. A missing or unusable reply is a “no decision”, counted separately and never scored as right or wrong.

4.4 Uncertainty. All intervals are percentile 95% cluster-bootstrap intervals (2,000 resamples, seed fixed in the registration). A cluster is an item with its three replicates; in bank A, items built from the same agent run share a cluster. H2 compares the two models on the same item and replicate, wherever both decided.

5.Registered results

Headline itemsJevDeepSeek V4 FlashGPT-6.1 Sol
Answers2,3972,3972,397
No decision01150
Ties (P = 0.5)2000
Scored2,3772,2822,397
Accuracy86.9%94.7%95.1%
95% CI84.5–89.1%93.0–96.0%93.5–96.7%
ECE0.0450.0350.021
95% CI0.031–0.0680.023–0.0510.009–0.036
Brier score0.0960.0460.035
95% CI0.085–0.1070.034–0.0600.025–0.047

Table 2. Headline results (banks A and B). On these items 67.7% of answers are “yes”, so always answering yes would score 67.7% accuracy and a Brier score of 0.219. Of DeepSeek V4 Flash’s 115 missing decisions, 111 were empty replies that reached the registered output limit.

H1: inconclusive. Jev’s ECE was 0.045, and its 95% interval, 0.031 to 0.068, contains the registered line of 0.05. By the rule fixed before the run, the result neither supports the calibration claim nor rules it out.
H2: Jev was less accurate than both comparators. On the same items and replicates, its accuracy was 6.6 points below DeepSeek V4 Flash (95% CI 4.7 to 8.6; 2,263 pairs) and 8.2 points below GPT-6.1 Sol (6.2 to 10.2; 2,377 pairs). Neither interval includes zero.
50%60%70%80%90%100%30%40%50%60%70%80%90%100%Stated confidence, max(P, 1 − P)Share correctdashed line = perfectly calibratedbelow it = overconfidentJev (TypeSafe): mean confidence 56%, correct 55% (n = 233)Jev (TypeSafe): mean confidence 65%, correct 69% (n = 371)Jev (TypeSafe): mean confidence 75%, correct 85% (n = 332)Jev (TypeSafe): mean confidence 85%, correct 91% (n = 403)Jev (TypeSafe): mean confidence 96%, correct 99% (n = 1,038)Jev (TypeSafe)DeepSeek V4 Flash: mean confidence 82%, correct 41% (n = 46)DeepSeek V4 Flash: mean confidence 99%, correct 96% (n = 2,228)DeepSeek V4 FlashGPT-6.1 Sol: mean confidence 64%, correct 49% (n = 39)GPT-6.1 Sol: mean confidence 74%, correct 64% (n = 64)GPT-6.1 Sol: mean confidence 82%, correct 69% (n = 121)GPT-6.1 Sol: mean confidence 100%, correct 99% (n = 2,165)GPT-6.1 Sol
Figure 1. Reliability on the headline items (banks A and B; 2,377, 2,282 and 2,397 scored answers). Each point is one of the five registered confidence bands, placed at the band’s mean confidence; bands with fewer than 10 answers are suppressed, which is why the two general-purpose models show fewer points. A calibrated model lies on the dashed diagonal. Values in Table 3.
Confidence bandJev: n · conf · correctDeepSeek V4 FlashGPT-6.1 Sol
50–60%233 · 55.6% · 55.4%0 · — · —8 · suppressed
60–70%371 · 64.8% · 69.0%1 · suppressed39 · 63.5% · 48.7%
70–80%332 · 74.7% · 85.2%7 · suppressed64 · 73.6% · 64.1%
80–90%403 · 84.6% · 90.8%46 · 82.4% · 41.3%121 · 82.4% · 69.4%
90–100%1,038 · 96.3% · 99.3%2,228 · 98.6% · 95.8%2,165 · 99.5% · 98.6%

Table 3. Reliability on the headline items: number of answers, mean stated confidence and share correct in each band. Bands with fewer than 10 answers are suppressed (registered rule).

6.Exploratory observations

Nothing in this section was a registered hypothesis. It is reported because it is in the data, and it should be read as description.

6.1 The direction of the error. Jev’s miscalibration ran toward caution: in every band above 60% it was right more often than it said (Table 3). The comparators’ ran the other way where they were not certain: DeepSeek V4 Flash was right on 41% of the 46 answers it gave at about 82% confidence, and GPT-6.1 Sol fell below the diagonal in every band under 90%. Both comparators placed over 90% of their answers in the top band, so their aggregate ECE rests mostly on that band.

6.2 At TypeSafe’s review band. TypeSafe’s own cookbook routes probabilities between 0.30 and 0.70 to human review [5]. Applying that rule to the headline items, Jev acted on 74.6% of them (1,773 of 2,377) and was right on 94.8% of those. DeepSeek V4 Flash acted on all but one and was right on 94.7%; GPT-6.1 Sol acted on 98.0% and was right on 96.2%.

6.3 The ambiguous items. On the 60 items whose answer cannot be known from the state, 79.4% of Jev’s answers fell between 0.3 and 0.7, against 2.5% for DeepSeek V4 Flash (118 decided) and 13.9% for GPT-6.1 Sol. Jev marked these as uncertain more often than either comparator did.

0%25%50%75%100%chanceArithmeticArithmetic · Jev (TypeSafe): 50.0%50.0%Arithmetic · DeepSeek V4 Flash: 95.5%95.5%Arithmetic · GPT-6.1 Sol: 100.0%100.0%Multi-step command chainsMulti-step command chains · Jev (TypeSafe): 64.4%64.4%Multi-step command chains · DeepSeek V4 Flash: 90.3%90.3%Multi-step command chains · GPT-6.1 Sol: 100.0%100.0%Negation pairsNegation pairs · Jev (TypeSafe): 80.0%80.0%Negation pairs · DeepSeek V4 Flash: 94.4%94.4%Negation pairs · GPT-6.1 Sol: 100.0%100.0%Distracting contextDistracting context · Jev (TypeSafe): 83.3%83.3%Distracting context · DeepSeek V4 Flash: 98.9%98.9%Distracting context · GPT-6.1 Sol: 100.0%100.0%Indirect wordingIndirect wording · Jev (TypeSafe): 84.4%84.4%Indirect wording · DeepSeek V4 Flash: 93.3%93.3%Indirect wording · GPT-6.1 Sol: 100.0%100.0%Adversarial noteAdversarial note · Jev (TypeSafe): 90.0%90.0%Adversarial note · DeepSeek V4 Flash: 94.3%94.3%Adversarial note · GPT-6.1 Sol: 100.0%100.0%CountingCounting · Jev (TypeSafe): 98.9%98.9%Counting · DeepSeek V4 Flash: 100.0%100.0%Counting · GPT-6.1 Sol: 100.0%100.0%Date comparisonDate comparison · Jev (TypeSafe): 100.0%100.0%Date comparison · DeepSeek V4 Flash: 100.0%100.0%Date comparison · GPT-6.1 Sol: 100.0%100.0%
Jev (TypeSafe)DeepSeek V4 FlashGPT-6.1 Sol
Figure 2. Accuracy on bank C, the categories TypeSafe lists as weak spots for this version (exploratory; 90 answers per model per category). The dashed line is chance. Values in Table 4.
Bank C category (30 items × 3)JevDeepSeek V4 FlashGPT-6.1 Sol
Arithmetic50.0%95.5%100%
Multi-step command chains64.4%90.3%100%
Negation pairs80.0%94.4%100%
Distracting context83.3%98.9%100%
Indirect wording84.4%93.3%100%
Adversarial note in the state90.0%94.3%100%
Counting98.9%100%100%
Date comparison100%100%100%

Table 4. Accuracy on the categories TypeSafe documents as weak spots [4]. Each category has 90 answers per model; intervals are wide at this size (for Jev’s arithmetic, 30% to 67%).

6.4 Where the headline gap comes from. On the headline subtypes, Jev matched the comparators when the question was what an agent’s reply said (95.1% against 94.6% and 97.0%) and on commands that change nothing (all three at 100%). The gap was on the harder judgements: whether the agent’s work actually held (71.9% against 86.3% and 81.6%) and whether a command that changes files destroys data (85.2% against 97.0% and 100%).

0%25%50%75%100%chanceA · does the reply say done?A · does the reply say done? · Jev (TypeSafe): 95.1%95.1%A · does the reply say done? · DeepSeek V4 Flash: 94.6%94.6%A · does the reply say done? · GPT-6.1 Sol: 97.0%97.0%A · did the work actually hold?A · did the work actually hold? · Jev (TypeSafe): 71.9%71.9%A · did the work actually hold? · DeepSeek V4 Flash: 86.3%86.3%A · did the work actually hold? · GPT-6.1 Sol: 81.6%81.6%B · read-only commandB · read-only command · Jev (TypeSafe): 100.0%100.0%B · read-only command · DeepSeek V4 Flash: 100.0%100.0%B · read-only command · GPT-6.1 Sol: 100.0%100.0%B · command that changes filesB · command that changes files · Jev (TypeSafe): 85.2%85.2%B · command that changes files · DeepSeek V4 Flash: 97.0%97.0%B · command that changes files · GPT-6.1 Sol: 100.0%100.0%
Jev (TypeSafe)DeepSeek V4 FlashGPT-6.1 Sol
Figure 3. Accuracy by item type on the headline set (exploratory). The bars for each type are the three models in the order of the key. Values in §6.4.

6.5 Consistency. For a question and its negation, the two probabilities should sum to one. Jev’s averaged 0.877 (45 pairs), DeepSeek V4 Flash’s 0.979 and GPT-6.1 Sol’s 1.000. Across the three replicates, Jev returned an identical probability on 27.2% of items, against 60.6% and 77.2%; its variation was small (mean range 0.027).

Jev (TypeSafe)DeepSeek V4 FlashGPT-6.1 SolMedian response time (seconds)Jev (TypeSafe) · Median response time (seconds): 0.18 s0.18 sDeepSeek V4 Flash · Median response time (seconds): 3.6 s3.6 sGPT-6.1 Sol · Median response time (seconds): 1.6 s1.6 sAccuracy, headline itemsJev (TypeSafe) · Accuracy, headline items: 86.9%86.9%DeepSeek V4 Flash · Accuracy, headline items: 94.7%94.7%GPT-6.1 Sol · Accuracy, headline items: 95.1%95.1%
Figure 4. Median response time over all 3,297 answers per model (observed, not registered; it includes the network) beside accuracy on the headline items (registered). Each panel has its own scale. Values in Tables 2 and 5.
Speed and cost (all 3,297 answers)JevDeepSeek V4 FlashGPT-6.1 Sol
Median response time0.18 s3.60 s1.57 s
90th percentile0.22 s15.41 s3.76 s
Measured cost, whole run$0.086$0.492$3.906
Per 1,000 answers$0.026$0.149$1.185

Table 5. Response time as recorded by our client, including the network (Jev called directly, the comparators through OpenRouter, from the same home network throughout). Cost is the usage each API reported, priced as each provider reported or listed it on 9 October 2026; Jev’s output is free [3]. Neither was a registered measure: we said we would not test TypeSafe’s cost or speed claims, and these figures describe this run only.

6.6 Speed and cost. Jev’s median answer took 0.18 seconds, about 20 times faster than DeepSeek V4 Flash and 9 times faster than GPT-6.1 Sol, and its slow tail was short. It was also the cheapest per answer. Read beside H2: in this run, choosing Jev over DeepSeek V4 Flash saved about 12 cents per 1,000 decisions and gave about 66 more wrong answers per 1,000.

7.What this does not show

  • That Jev is, or is not, calibrated. The registered test could not decide it. We do not expect a larger rerun of this design to settle it, because the estimate sits close to the line; a new test would need its own registration.
  • Anything about TypeSafe’s own workflows. These are our items, about coding-agent work and destructive commands, not the workflows TypeSafe published.
  • Anything about later versions. Every answer came from jev-1.13.0, and TypeSafe says several of its documented weak spots will be fixed in later versions [4].
  • A verdict on the comparators’ settings. They ran at the registered temperature and output limit; with a larger limit, DeepSeek V4 Flash’s 115 missing decisions might not have occurred.
  • Anything about intent. We report what each system returned. We have no instrument that sees why, and we make no claim about it in either direction.

8.Competing interests

Sentient Index Labs & Technology sells independent evaluations of AI models. We have no commercial, financial or personal relationship with TypeSafe AI, DeepSeek, OpenAI or OpenRouter, and accessed every model through our own accounts on each provider’s public terms. The items in bank A come from our own Harness Battery study [9]. No AI judge graded any answer in this study; every label is mechanical.

9.Data and registration

The registration is public at osf.io/qg6nc [1]. The item banks, the analysis code, every raw response and every scored answer are held in our private repository; the registered SHA-256 hashes identify the exact banks used. A shorter summary of this study is published at silt-seb.com/studies/does-jev-know [10].

10.References

  1. Scanlon, S. (2026). Does Jev Know When It’s Wrong? Registration, Open Science Framework. https://osf.io/qg6nc
  2. TypeSafe AI. Introducing System One models and Jev. https://typesafe.ai/blog/introducing-system-one-models-and-jev (read 10 October 2026).
  3. TypeSafe AI. Home page. https://typesafe.ai/ (read 10 October 2026).
  4. TypeSafe AI. Jev 1.13 jaggedness (last reviewed 2 October 2026). https://docs.typesafe.ai/model-jaggedness/jev-1.13 (read 10 October 2026).
  5. TypeSafe AI. Self-consistency: nouls (cookbook). https://docs.typesafe.ai/cookbooks/consistency_noul_cookbook (read 10 October 2026).
  6. Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
  7. Naeini, M. P., Cooper, G. F., & Hauskrecht, M. (2015). Obtaining well calibrated probabilities using Bayesian binning. Proceedings of AAAI.
  8. Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of ICML.
  9. Sentient Index Labs & Technology. The Harness Battery. https://www.silt-seb.com/harness-battery
  10. Sentient Index Labs & Technology. Jev: Twenty Times Faster, Seven Points Behind (study summary). https://www.silt-seb.com/studies/does-jev-know
SILT-RP-009 · version 1.0 · issued 2026-10-10
© 2026 Sentient Index Labs & Technology. This document is licensed CC BY-ND 4.0 — share it freely and in full, with attribution; do not publish modified versions. Translations and excerpts are granted on request to press@sentientindexlabs.com.
⚠️ Corrections are issued as a new version under the same identifier, each with a dated note in the text recording exactly what changed and what the earlier version said.
Nothing in this document is a certification, audit opinion, conformity assessment, or legal advice. See the Subscriber Agreement, Section 15.