Cloudflare Shipped a 27B Decision Model That Beats Jev at Its Own Job

Cloudflare released Clef in the first days of October under an Apache 2.0 licence. It is a 27-billion-parameter model built on Qwen3.8-27B, and it does not write prose. It scores choices.
The category is small and specific. A decision model takes an input and a set of candidate schema choices, then returns a ranking rather than a generated sentence. Typesafe AI's Jev established the approach and became the reference point for it. Clef arrives as the second serious entrant, and on the numbers Cloudflare published, it beats the incumbent on both accuracy and latency.
How the architecture differs
Clef uses a prefill-only design. It reads the input and scores schema options in parallel rather than emitting tokens one at a time. That single choice explains most of the performance gap, because the expensive part of a conventional language model call is the decode loop.
On BANKING77, a standard intent-classification benchmark, Clef scores 94.20 macro-F1 against Jev's 79.74. Cloudflare also reports that Clef-flash, the smaller configuration, has a median latency of 38.8 milliseconds, which it describes as roughly 13 times faster than Jev's 524.1 millisecond baseline.
The interface matters as much as the score. Clef is Jev-API compatible, so a team already wired into Jev can swap the endpoint without rewriting integration code. Clef also accepts multimodal input inside a 64,000-token context window, where Jev is text-only and capped at 32,000.
Cloudflare's framing is blunt: for routing, classification and structured selection, generating text is wasted work. If the answer you need is one of twelve categories, a model that writes a paragraph and then post-processes it is doing something harder than the task requires.
Where this sits next to embeddings
It is worth separating decision models from the older approach to the same problem. Teams have routed and classified with embeddings for years: embed the input, compare it against labelled examples, pick the nearest one. That works, and it is cheap, but it needs a labelled set to compare against, and it struggles when the decision depends on instructions rather than on similarity.
A decision model takes a schema and a set of options as part of the request. It can answer a wider range of questions than a similarity check allows: which category a message belongs to, which of three tools to call, which model to route to, and which of several policies applies. The instructions live in the request rather than in a maintained index of examples, which makes it easier to change behaviour without re-labelling anything.
The trade is that a decision model is still a language model doing the scoring, so it inherits the failure modes of one. It can be wrong with high confidence, and its confidence scores are not calibrated to a probability a routing policy can trust without testing. The prefill-only design removes the decode cost, but it does not remove the problem of knowing when the model is guessing.
The cost math is what makes the category interesting anyway. Routing a request through a generative model costs a full generation on the critical path, often several hundred milliseconds and a few hundred tokens. Cloudflare reports Clef at under 40 milliseconds median for the flash configuration. For an agent that makes several routing decisions per user turn, that difference compounds, and it compounds on the step users actually feel.
The category is standardising fast
What makes this more than one vendor's launch is how quickly the tooling around decision models has thickened.
On September 30, SGLang added native /v1/decisions and /v1/systemone routes, letting language and vision models act as classifiers and scorers without token-by-token generation. The framework demonstrated the capability by running Qwen3.8-27B as a multimodal decision model and reported beating Pokemon FireRed in one attempt at sub-100 millisecond latency.
Days later, llama.cpp added decision model support through a new /v1/systemone endpoint, covering open models including Kev-4B and OpenJev, with a planned ship in version 0.6.0. Small models run on CPUs, and some support image inputs for classification.
Respan launched its own entry, Span-01, a hyper-parallel reasoning classifier for detecting prompt injection, hallucination and tool misuse in agent traces. It is trained with RLAIF and evaluates multiple unseen behaviour definitions in a single forward pass. Respan prices it at two cents per million input tokens, with a free Lite version.
That is four separate projects treating decision models as an interface worth implementing natively, inside two weeks. When a pattern shows up in an inference framework, a local runtime and a startup's product line at the same time, it has stopped being a curiosity.
Why this matters for agent routing
The people who stand to gain first are those building agents that call other models. An agent needs to decide which tool to use, which model to route to, whether a request is an injection attempt, and whether a response is grounded. Today many teams implement those steps by asking a large model to answer in JSON and hoping the format holds.
A classifier that returns a score in tens of milliseconds changes the cost of that decision. Routing a request through a language model costs a full generation, and it costs it on the critical path. Routing it through a prefill-only scorer costs a fraction of that and finishes sooner. If the accuracy holds up, the practical effect is that more of an agent's control flow can run at machine speed rather than at token speed.

Cloudflare's own argument is about infrastructure economics. The company says open-weight models now run 15 to 90 percent cheaper than closed equivalents for most production workloads, and that the era of open weights playing catch-up is over. Clef is offered as evidence: a downloadable, Apache 2.0 licensed model that outperforms a proprietary competitor on the competitor's own benchmark category and can be self-hosted.
What to check before believing the headline
Three caveats are worth holding onto.
First, the benchmarks are vendor-selected. BANKING77 is a real and widely used intent dataset, but a single macro-F1 figure does not describe how a model behaves across a production distribution with rare classes, ambiguous inputs and adversarial phrasing. Independent evaluation would settle more than a launch post can.
Second, Jev-API compatibility is a claim about request shape, not about behaviour. A drop-in endpoint that returns different confidence calibration can still break a routing policy tuned around the original.
Third, the comparison is measured against Jev as published. Typesafe AI has its own releases in flight, and the gap between a challenger and an incumbent rarely stays fixed for long.
None of that removes the more durable point. Decision models exist because routing, gating and classification are a large share of what agent infrastructure actually does, and doing them with a text generator was always an awkward fit. Clef, Span-01 and the new endpoints in SGLang and llama.cpp are converging on the same idea: some model calls should return a number rather than a sentence. The question now is which of these implementations ends up underneath everyone else's agent loop, and how long the category stays contested.
Related articles
The Gap Between Arena Leaderboards and Real Image Output Is Getting Wider
The infrastructure for ranking models has never been better, and the connection between rank and practical output has never been looser.
Google Flow and Adobe Firefly Move AI Video Out of the Chat Box
Base model quality has converged enough that the differentiator has moved to what surrounds the model.
Google Cut Nano Banana 2.1's Output Price in Half and Fixed Its Weakest Features
The most consequential detail sits outside the feature list, and it is the price.
Vida Wants To Bill for AI Agents by Results Rather Than Usage
Usage-based billing aligns the vendor's revenue with the agent taking longer. Outcome pricing inverts that.