A Model That Refuses to Write Sentences Is Now Handling a Trillion Tokens a Day

TypeSafe AI shipped Jev on September 15, 2026, and pulled the waitlist down on September 27. The pitch is narrow enough to sound like a typo at first: Jev does not generate text. You send it a piece of state and a set of typed questions, and it returns answers drawn from a schema you defined in advance, each one carrying a probability and a confidence value.
Three primitives cover what it will answer. Choice picks one option from a list you supply. Score rates the input against ordered descriptive levels. Noul answers a yes-or-no question with the probability that the answer is yes.
That is the entire product. No chat, no paragraphs, no markdown, no apologies for being an AI.
Why anyone would want that
Diogo Almeida founded TypeSafe AI and previously worked at OpenAI on the instruction-following methods that shaped how ChatGPT behaves. He left in 2024. His argument since then has been consistent: a chat interface is the wrong shape for software. Code does not want a sentence back. It wants a value it can branch on.
Anyone who has wired an LLM into a production system recognizes the pain he is describing. You ask a model to classify a support ticket, and you get back a friendly paragraph that starts with the classification and ends with an offer to help further. Then you write a regex. Then the model changes its formatting, and the regex breaks. Then you add a JSON mode, and occasionally the JSON is malformed. The model was never the fragile part. The interface was.
Jev sidesteps that by removing generation entirely. It uses a parallel sampler rather than autoregressive token generation, which means the possible outputs are enumerated before inference runs. Schema conformance stops being something you hope for and becomes something the architecture guarantees. TypeSafe's own numbers put end-to-end latency between 70 and 500 milliseconds.

The company trained it with a method it calls Reinforcement Learning for Calibrated Decisions. The emphasis falls on the word calibrated. A confidence score is only useful if 0.85 actually means roughly 85 percent of cases go that way, and most language models are notoriously optimistic about their own answers. If your application is going to act automatically above a threshold and escalate to a human below it, calibration stops being a bonus and becomes the whole design.
There is a second property that matters just as much in production, and it is easy to overlook. Because the output space is fixed before the call, two runs with the same state and the same questions return the same answer. Reproducibility is the thing most teams discover they need only after an incident review, when someone asks why the system routed a specific ticket a specific way three weeks ago and nobody can reconstruct it. A decision model makes that question answerable.
The adoption curve got strange fast
Vercel added Jev to its AI Gateway on day two and then reported that it became the fastest-adopted model in the gateway's history. Nearly 13 percent of paid teams were using it within 24 hours, roughly twice the rate of the GPT-5.6 family and more than six times that of Fable 5.1. Netlify followed. LangChain shipped TypeSafeClassifier integrations with model routing and middleware that screens tool calls before execution. Five independent Elixir clients appeared within days, which is the kind of detail that tells you developers were already looking for this shape of thing.
The scale numbers are harder to ignore. Almeida told the Wall Street Journal that Jev was processing around a trillion tokens a day within about a week of launch, and that roughly a quarter of Fortune 500 companies were using it. The Information reported that TypeSafe is negotiating a round targeting $1 billion or more, with some interested investors valuing the company above $10 billion. Almeida declined to comment on the funding report.
The cost structure explains part of the pull. Input runs $0.042 per million tokens with output free, and the context window is 32,000 tokens. Compared with routing a classification job through a frontier model, the arithmetic stops being close.
Where the thing actually holds up
TypeSafe's published workflow evaluations put Jev at parity with Sonnet-class models on accuracy across four internal tasks at a fraction of the latency and cost, while trailing the strongest frontier configuration by 6.3 points. Performance splits by task: 76.0 percent on customer service classification, 61.8 percent on invoice processing.
The company is also unusually direct about what Jev should not be asked to do. Its documentation warns against using it for arithmetic, date comparison, or counting across a large noisy state, and tells teams to keep that logic in ordinary code and pin to a versioned model id rather than a floating alias.
That guidance points at the honest boundary of the category. Jev is a classifier with a much better interface and language-level context understanding, and it is aimed at work that was already structured: routing a ticket to billing or technical, scoring an abuse signal, gating an agent's tool call before it spends money downstream, ranking search results. In those slots it replaces an expensive reranker or a prompt-engineered frontier call, and the replacement is faster and cheaper by a wide margin.
The competitor set arrived within three weeks
A category this obvious does not stay quiet. Cloudflare shipped Clef and Clef-flash on October 1, two open-weight decision models under Apache 2.0 built on Qwen 3.8-27B and a 9B backbone respectively. They accept the same request shape, add vision input and a 65,536-token context window, and come with an RL fine-tuning service. Cloudflare's own latency figures put Clef-flash at a 38.8 millisecond median against 524.1 milliseconds for Jev, which is a large enough gap to matter for anything in a request path. OpenAI's protocol work and Cloudflare's launch both leaned on the same idea, and Clef was designed so that code written against Jev keeps working.
Fastino Labs shipped GLiDE and GLiNER2.5-Decide inside the same window. Open replicas appeared too: Jeff v1.1 fine-tunes Qwen3.5 and Gemma 4 into 0.8B and 2B decision models that answer in 22 to 28 milliseconds on a workstation GPU. There is a caveat worth stating, and MarkTechPost made it: Fastino's comparisons run against its own test suites and an open reproduction of Jev rather than TypeSafe's hosted model, so cross-vendor numbers are not a clean ranking.
None of that erases what TypeSafe established. The company put a name on a split that had been forming slowly for a year, between models that generate language for people and models that return decisions for software. Once the second category has a request shape and a price per million tokens, it stops being a research curiosity and starts being a line item.
What to do with this
The practical question for a team is unglamorous. Look at what you are currently paying a frontier model to do, and count how many of those calls end with code reading a single value out of a response you spent tokens generating. Ticket routing, moderation gates, lead scoring, relevance ranking, tool-call approval. Each of those is a candidate to move.
The reason to move has nothing to do with a decision model being smarter. The interface simply stops fighting you. You send a state, you get back a bounded answer with a calibrated probability attached, and your code branches. Escalate below the threshold, act above it, and keep the arithmetic in code where it belongs.
There is a cost-side argument too, and it is the one finance teams will make. A decision model's output is a few hundred bytes rather than a paragraph, and output tokens are where frontier pricing hurts most. A routing call that costs fractions of a cent at $0.042 per million input tokens with free output turns a line item that needed justification into one that does not.
That is a smaller promise than "reasoning," and it maps more closely onto what most production systems were already trying to get out of a language model.
Related articles
The AI Video Price War: Luma Cut Seedance Rates by Up to 73%, and Runway Started Selling Rivals
The engines are close enough now that the invoice is a better guide than the leaderboard.
A 260M Image Model Beat a Rival 6.5x Its Size by Looping the Same Blocks
Adding parameters still works. The more active work is about making a given model do more with less.
Two Voice Models Just Reset the Bar: 50ms to First Audio, and a 99M Model on a Laptop CPU
Quality converged, and the competition moved to where the model runs, how fast it starts, and what it costs per call.
Moonshot's Kimi K2.6 Runs a Thousand Agents at Once and Built a Compiler in Ten Hours
A thousand agents that each need supervision multiply the supervision, not the capacity.