TypeSafe launches Jev, claiming faster, cheaper alternative to LLMs for AI automation
TypeSafe's new Jev model returns structured decisions instead of text, with speed and cost claims that mostly rest on the company's own testing.
Estimated reading time: 6 minutes
TL;DR
TypeSafe AI has released Jev, an early-access model that returns structured decisions with probabilities instead of free-form text. The company says Jev responds in 70 to 500 milliseconds and, in its own workflow evaluations, was 193.6 times faster and 444.6 times cheaper than the language models it was compared with. Those evaluations were designed by TypeSafe and have not been independently reproduced — a second outlet reviewing the launch reached the same conclusion. In a separate test by the media company Every, Jev was about 25 times faster than Fable 5.1 and an estimated 580 times cheaper, and found six of seven deliberately planted writing problems; Fable found all seven. Investor DCVC says it led a $40 million seed round in the company.
What happened
TypeSafe announced Jev on September 15, 2026. Founder Diogo Almeida, who previously worked at OpenAI and co-authored the 2022 InstructGPT paper on training language models with human feedback, said the company had spent two years in stealth. Jev is available in early access. The same day, DCVC announced that it had led a $40 million seed round in TypeSafe, which is based in San Francisco.
TypeSafe calls Jev its first “System One Model,” a name taken from Daniel Kahneman’s distinction between fast, intuitive thinking and slow, deliberate reasoning. Chat models such as ChatGPT write their answers one token at a time. Jev does not write text: developers define the possible answers in advance, and Jev returns choices, scores and probabilities in that format.
For example, an application could pass Jev a customer message and ask for the probability that it concerns billing, technical support or another predefined category, then route the message based on the result.
According to TypeSafe, Jev combines a new model architecture, a sampler that produces all outputs in a single parallel step, and a training method the company calls Reinforcement Learning for Calibrated Decisions (RLCD). The company says it creates all of its training data internally and does not train on customer data. It also says Jev is not a language model.
Price and speed. TypeSafe charges $0.042 per million input tokens and does not currently charge for output. It lists end-to-end response times of 70 to 500 milliseconds and says this makes Jev 40 to 200 times faster than frontier language models on the kinds of structured tasks it is built for. The company says its latency measurements are generally taken from laptops on the US West Coast, where its service runs, and that it cannot yet prove its pricing is not subsidized.
TypeSafe’s own evaluations. The 193.6-times speed and 444.6-times cost figures come from four workflows TypeSafe built to test models inside software pipelines. The company has published the workflows, full queries, examples and the cases where models disagreed.
These evaluations do not check answers against independently verified correct answers. Instead, TypeSafe treats the average predictions of two large models, GPT-6 Astra and Fable 5.1, as the reference, and measures how closely other models match it. TypeSafe attaches several qualifications to the results:
- The gains are likely at the higher end of what users would see in practice.
- The workflows were written by TypeSafe’s own model-capabilities team, so bias is possible.
- Competing models ran through a TypeSafe wrapper that makes them return decisions with probabilities, which the company says tends to make them slower and more expensive.
- Using Astra and Fable as the reference favors OpenAI and Anthropic models; TypeSafe says this likely understates the performance of Jev and of DeepSeek’s models.
The hallucination claim. TypeSafe says Jev “can’t hallucinate.” The guarantee the company documents is narrower: Jev cannot return a value outside the structure the developer defined. TypeSafe says its 0% figure for this property is not a measurement but follows from that design. The comparison figures it shows for language models come from the routing service OpenRouter, which the company acknowledges may be biased.
A valid answer is not necessarily a correct one. A model asked to label a transaction as fraud or not fraud will always return one of those two labels, but it can still pick the wrong one.
Every CEO Dan Shipper compared Jev with Fable 5.1 on synthetic passages containing seven deliberately introduced writing problems. Jev took a median of 0.35 seconds per passage and Fable 8.83 seconds, making Jev about 25 times faster in that test. Every estimated Jev’s cost at about 580 times lower. Jev found six of the seven problems. Fable found all seven.
Every’s test used different tasks from TypeSafe’s evaluations and does not reproduce them.
What this means (and what it does not)
Jev offers developers a different kind of interface. Instead of receiving text that software must then parse and check, they receive values and probabilities in a format fixed in advance. That fits tasks such as classifying, routing, scoring and extracting information, where no written answer is needed. It is not designed for chat, writing or coding.
The speed and cost figures apply to specific tests. The 193.6-times and 444.6-times figures describe TypeSafe’s workflows and comparison setup. Every’s figures describe one writing-check task. Neither establishes a general advantage across AI tasks, and in Every’s test the faster, cheaper model was also the less thorough one.
The design rules out malformed output. It does not rule out wrong decisions.
What we still do not know
Independent reproduction. No source reviewed has rerun TypeSafe’s workflow evaluations with the same method. Accuracy and calibration. TypeSafe says Jev’s probabilities are calibrated, meaning higher confidence goes with higher accuracy. There is no independent evidence yet on how Jev’s accuracy and calibration compare with large language models across a broad set of real tasks with verified answers; Every’s six-of-seven result covers one task. Performance at scale. The published latency figures do not show how Jev performs under sustained production load or from other regions. Long-term price. TypeSafe says it cannot yet show that current pricing is not subsidized. Production use. The sources reviewed name no customer running Jev at large scale. Valuation. The $40 million figure and DCVC’s lead role come from DCVC and TypeSafe’s own announcement. The valuation was not disclosed.
Sources & Bylines
Every source cited in this article, gathered in one place.
- https://typesafe.ai/blog/introducing-system-one-models-and-jev
- https://arxiv.org/abs/2203.02155
- https://www.finsmes.com/2026/09/typesafe-ai-raises-40m-in-seed-funding.html
- https://www.theregister.com/ai-and-ml/2026/09/16/typesafe-ai-debuts-model-for-machines-that-plays-doom/5296711
- https://www.modemguides.com/blogs/ai-news/jev-typesafe-reality-check-run-locally
- https://docs.typesafe.ai/concepts/system-one
- https://every.to/also-true-for-humans/mini-vibe-check-typesafe-s-jev-judged-everything-i-ve-written-in-0-7-seconds
Editorial check, counted automatically
- 7 sources cited
- 36 inline-linked claims
- 7 unsourced claims found
- 0 banned words found
- 5 numbers without context
Also available in Portugues (BR)