Decision Models Are Quietly Replacing the LLM in the Agent Loop

Cloudflare shipped two models in the first days of October that do not write anything. Clef and Clef-flash read an input plus a set of typed questions and return a probability for every allowed answer. No prose, no chain of thought, no tokens generated one at a time. Just a structured choice.
That sounds like a downgrade until you look at the numbers. On BANKING77, a standard intent-classification benchmark, Clef scores a macro-F1 of 94.20 against 79.74 for Jev, the model that introduced the decision-model category. Clef-flash, the smaller 9B version, manages 90.93. The latency gap is wider still. Clef-flash returns a median answer in 38.8 milliseconds; Jev takes 524.1 milliseconds. Cloudflare says Clef beats Jev on seven of ten decision benchmarks.
Both models ship under Apache 2.0 on Hugging Face, and both are Jev-API compatible, so teams already built around Jev can switch without rewriting their integrations. The Hacker News thread hit 602 points and more than 200 comments within a day.
Why a smaller model can win at a narrower job
Decision models are a different shape from the chatbots most developers reach for. A large language model generates open-ended text. A decision model reads a state and a list of permitted answers, then scores each one. The category was pioneered by Typesafe AI with Jev, which showed that agent routing or classification tasks do not need a 400B general-purpose model. They need bounded, cheap, fast output.
Cloudflare's Clef is built on Qwen3.8-27B and uses a prefill-only architecture that scores schema choices in parallel instead of generating tokens sequentially. That is where the latency goes. A general model has to emit an answer token after token; a decision model only has to read the question and commit.
There is also a context difference worth noting. Clef accepts multimodal input inside a 64k-token window, covering text, JSON, images and video. Jev handles text only, at 32k. For an agent that routes on a screenshot or a PDF, that is not a small edge.
The community's first objection was the right one
The most useful pushback in the Hacker News discussion did not dispute the benchmarks. It disputed the label. "Open weights, not open source," one commenter wrote. The weights carry a permissive license, but the training data and pipeline were not published, so the model cannot be reproduced from scratch.
That distinction matters for anyone deciding where to run this. Clef was trained on a proprietary Qwen starting point. The weights are free to download and host, which is a real cost saving, but they are not auditable in the way open-source software is. Teams with strict supply-chain review policies should read the license and the model card before assuming the Apache 2.0 tag settles the question.
The same week, Amazon put one in your laptop
Cloudflare was not alone. AWS's Strands Labs published Strands Decider 2B, an open decision model with weights and training scripts, designed to run locally and return confidence-scored choices in tens to low hundreds of milliseconds. Cloudflare also added an RL fine-tuning service for its decision models on Workers AI.
The pattern is a small model doing one job well and running close to the data. If you can gate a tool call, verify that a request is grounded, or decide whether to escalate, with a 2B model in 40 milliseconds, the economics of routing an agent through a frontier model for every step start to look wasteful.

The paradigm Jev introduced, and why it took a year to catch on
Typesafe AI's Jev made a claim that was easy to dismiss when it launched: most of what agents ask a model to do is not generation, it is classification. Route this ticket, pick this tool, decide whether this action needs approval. For those jobs, a model that writes paragraphs is doing far more work than the task requires.
Jev proved the narrow approach could beat a general model on its own benchmarks. What it could not do was make the category feel urgent. A single vendor selling one decision model is a curiosity. Cloudflare shipping an Apache 2.0 version that is compatible with Jev's API is a different situation, because now anyone can host it, inspect the license, and swap it in without a rewrite. Amazon arriving in the same week with its own open decider turns a curiosity into a category.
The timing lines up with how agents are actually being built. The first generation of agent frameworks routed every step through one large model because it was simplest. As those systems move into production, the routing costs become the thing teams notice. Splitting the loop into a language model for the parts that need language and a decider for the parts that do not is the obvious fix, and it took the open release of a good decider to make it practical.
The fine-tuning angle
Cloudflare also added an RL fine-tuning service for its decision models on Workers AI, which points at where the category goes next. A generic decider is useful. A decider fine-tuned on your own routing history, your own escalation rules and your own notion of scope is more useful, because it learns the boundary that matters to your business.
That is also the part that should make teams careful. A fine-tuned decider encodes your past decisions, including the bad ones. If your team used to approve refunds it should have escalated, the model will learn that pattern and apply it faster than a human ever could. Calibration cuts both ways.
Where this fits in an agent pipeline
Picture a support agent handling a refund. Before it can call the refund tool, something has to answer questions like: is this request in scope, does the customer's account allow it, does this need a human. Those are decision questions with a fixed set of answers. A decision model can return the probabilities, and the agent acts on a threshold.
The cost profile is the point. Routing every one of those checks through a 400B model costs money per call and adds hundreds of milliseconds of latency. Offloading them to a local 2B or 9B decider keeps the frontier model for the parts that actually need language: writing the reply, summarizing a long thread, handling an ambiguous case. Budgets and latency budgets both shrink.
There is a catch. A decision model returns a probability, and probabilities can be wrong in confident ways. If you automate a refund approval on a 0.92 score, you have moved the risk from the model's wording to your threshold. Calibration, not raw accuracy, becomes the number to watch. Latency and cost are easy to measure; a well-calibrated 0.9 is harder.
What to watch next
The tooling is already following the models. llama.cpp added support for decision models, and Perplexity and Hugging Face are both placing bets in the same direction. If the category holds, the interesting question stops being which decision model wins a benchmark and becomes which decisions you are willing to hand to a probability.
For agent builders, the practical move is to audit the loop and find the steps that never needed fluent text. Classification, routing, tool selection and pre-action checks are the usual candidates. Those are the steps where a 40-millisecond answer over a fixed set of choices can replace a slow, expensive generation. The rest of the agent can stay on the model it already uses.
Related articles
A Wheeled Semi-Humanoid Finished an Hour of Laundry Without Help
Individual tasks can succeed while a workflow still fails. Dyna changed the metric.
LTX 2.5 Wants to Render Your Blocky Blender Draft Into a Finished Shot
You do not control what happens in text-to-video. This tries to fix that.
ServiceNow Turns Agent Failures Into Training Data
Generation without verification is noise. The gates are the product.
One Framework for Language and Vision: Horizon's 1.6B Open Model
A bet that the bridges between language and vision were never needed.