Which Jev Claims Are Verified?
Every headline figure from the launch, sorted into what is checkable today, what is company-reported, and what TypeSafe itself declined to publish.
Jev claims, one by one
Every Jev claim sorted into what you can check yourself, what rests on the company's own testing, and what was simply not published. This is not a verdict on whether the Jev claims are true — several may well be. It is a record of which Jev claims currently have evidence behind them.
Published in the launch blog. Consistent with the Doom demo's 10 queries per second, though that is also a company demo.
A range rather than a measurement. TypeSafe notes results "are likely to sit at the high end of real-world results".
From the company's own eval suite. No third-party has reproduced it as of September 2026.
Same suite, same caveat. The comparison wraps external models to produce compatible output, which raises their latency and cost.
A published rate card, not a benchmark. You can check it against your own invoice once you have access.
Also a rate-card fact. Follows naturally from responses being typed values rather than long sequences.
TypeSafe explicitly states this figure "is not empirical" — it follows from schema design rather than from measurement.
Cannot return a value outside your schema. Can still return the wrong valid value, which is a different failure and not covered by the claim.
Confirmed by the company announcement and independent reporting from multiple outlets on September 15–16, 2026.
Absent from the launch materials. Named customers and error rates on real data are what would settle the economics question.
What is wrong with the Jev benchmark
The headline Jev benchmark figures — 193.6 times faster, 444.6 times cheaper — come from a workflow evaluation the company built and ran. That is normal at launch. What is worth understanding is the specific Jev benchmark design choice that makes the ratio hard to interpret.
The Jev benchmark has no answer key
The workflow evals do not score against ground truth. They measure each model against the average probabilities returned by two large external models — agreement with a committee, not correctness.
Jev benchmark baselines were wrapped
To produce compatible structured output, the external models were wrapped by TypeSafe. Wrapping adds latency and cost to the baseline, which flatters the ratio being reported.
The company says so itself
TypeSafe acknowledges possible bias and states the figures sit at the high end of real-world gains. That candour is worth more than the numbers, and it is easy to miss in coverage that quotes only the multiple.
What would settle the Jev claims
Independent evaluation on third-party data, published error rates from named deployments, and pricing that holds after early access ends. None of those exist yet.
Credit where the Jev claims deserve it
It would be easy to write this Jev claims audit as a takedown, and it would be unfair. The company states in its own materials that the results carry possible bias and are likely to sit at the high end of real-world gains. It says the zero-hallucination Jev claim is not empirical. It lists, by name, the Jev questions its launch post does not answer.
That is more disclosure than most launches manage, and it inverts the usual problem. Normally the caveats have to be reconstructed by outsiders reading between the lines. Here the Jev caveats are in the source material, and the distortion happens downstream — in coverage that quotes the Jev benchmark multiple and drops the sentence beside it.
So the honest summary is not "the Jev claims are inflated". It is "the Jev claims are unaudited, and the company says so". Those are different situations, and only the second is compatible with the Jev benchmark numbers turning out to be broadly right.
What would settle the Jev claims
Three kinds of Jev evidence, in rough order of how much each would move the picture. First, independent Jev evaluation on data the company did not select — accuracy and, more importantly, calibration quality measured by someone with no stake in the result.
Second, named Jev deployments with published error rates. A customer saying "we route this volume at this accuracy and here is what it costs us" is worth more than any Jev benchmark, because it includes all the Jev integration friction that evaluations leave out.
Third, Jev pricing that survives the end of early access. A Jev rate card offered to a waitlist is a hypothesis about unit economics. The test is what it looks like a year after general availability, under real load, with margin expectations attached.
Until then the reasonable posture on the Jev claims is neither dismissal nor adoption on faith. Jev is cheap enough to test that you can generate your own evidence for the price of an afternoon — and given that everything public traces back to one source, your own numbers are worth more than anyone's summary of theirs.
How to test the Jev claims yourself
The unusual thing about this launch is that the Jev claims are cheap to check. You do not need a research team or a benchmark suite. You need a few hundred cases you already have labels for, a schema, and an afternoon once your invitation lands.
Run the sample, record the answer and the confidence on each case, then bucket by confidence and compute observed accuracy per bucket. That single curve tests the two Jev claims that matter — whether it is right often enough, and whether the confidence figure is honest — on your data rather than the company's.
Time the calls while you are at it and total the spend. Those two numbers test the remaining Jev claims directly, and unlike the published multiples they are measured against your current pipeline rather than against a wrapped baseline somebody else chose. Your own figures are worth more than anyone's summary of the Jev benchmark, and they are the only evidence your own decision should rest on.