Frontier Agents Finished 30 Percent of a Research Workflow. That Is the Number.

Stanford has a new benchmark called Terminal-Bench-Science 0.1, and the headline result is a good deal less exciting than the marketing around agents usually is. The benchmark put 70 expert-written research workflows in front of frontier agents. The best one finished 30 percent of them.
Thirty percent is the number to hold onto. It reads as neither a failure nor a triumph. It is an honest measurement of where the technology sits, published by people who wrote the tasks by hand.
What the benchmark actually tests
The tasks are science workflows, written by domain experts, and the agents run in a terminal. That choice matters. A terminal is a working environment where a task is real: you have to actually install the dependency, write the file, run the analysis, and check the output. There is no partial credit for describing the right approach. The command either works or it does not.
The workflows are drawn from research practice, which is a much messier target than a coding task with a tidy test suite. A coding benchmark usually has a defined correct answer, so it can be scored automatically. Research work often does not, which is why building a benchmark like this required experts writing the tasks rather than pulling them from a repository.
Why 30 percent is the right framing
Benchmarks tend to get read as pass or fail. A model that scores 90 on a coding benchmark is treated as near-solved. A model that finishes 30 percent of research workflows can, in the same spirit, be read as failing most of the time.
That reading misses what the number is for. Thirty percent completion on hand-written expert tasks means the agents can handle the routine third of research work: setting up environments, running established pipelines, cleaning data, producing standard outputs. The other 70 percent is where the task requires judgement that the workflow did not spell out, or where a step breaks in a way that needs a human to decide what to do next.
That is a useful split. It tells a lab where to point an agent today, and it tells a tool builder where the gap is.
The terminal is the interesting part
Running agents in a terminal is a deliberate choice to test them where the work happens. Research software is largely command-line software. A model that can only operate a chat window does not help a lab. A model that can sit at a shell, read the docs, install a toolchain, and recover when a build fails is a different category of useful.
It is also where the failure modes are most visible. A terminal does not hide an error. When a command returns an ambiguous message or a script half-completes, the agent has to decide whether to retry, change approach, or stop. That decision point is where most of the 70 percent is lost, and it is exactly the kind of thing that a coding benchmark with a clean test suite will never surface.

How this differs from coding benchmarks
The obvious comparison is to coding benchmarks, which have become the standard way to advertise agent progress. A coding benchmark typically ships with a repository and a test suite, so the correct answer is defined and the score is automatic. That makes it cheap to run and easy to compare, and it is why those numbers dominate the conversation.
Research workflows resist that treatment. Often there is no single correct output, and the quality of a result depends on judgement about what to measure and how. That is why the tasks here were written by experts rather than scraped from a repository, and why the benchmark reports completion rather than a pass rate on tests.
The trade is coverage for realism. A coding benchmark can be run thousands of times a day and produce a tight number. A workflow benchmark is slower and noisier, but it tests the environment where a large share of technical work actually happens. Both are useful, and only one of them was widely available before.
Why the number is measured this way
Completion is a blunt metric, and the authors chose it on purpose. A workflow either finishes or it does not, and that is a fact an outside observer can verify without judging the quality of the result. Elegance is not scored. Correctness is not argued about. The agent either produced the output the workflow called for or it stopped somewhere short.
That bluntness is the point. Benchmarks that try to score quality on open-ended research tasks tend to collapse into the benchmark author's taste. By scoring completion, this one trades nuance for reproducibility. Two labs running the same workflow get the same answer, which is what makes a number worth citing.
The cost is that it cannot distinguish an agent that nearly finished from one that failed immediately. Both count as failures. That is a limitation, and a future version may address it, but a coarse number that everyone agrees on beats a fine one nobody trusts.
What the result says about scientific work
The benchmark is a quiet answer to a loud claim, which is that agents will soon do research. What it shows is that agents can run research, meaning they can execute a procedure a human has already worked out, across a growing share of routine steps. They cannot yet do the part where the procedure is unknown and someone has to invent it.
That gap is not a small engineering detail. The routine work takes up a large share of a scientist's time, and automating it saves real money and hours. It is also not the part that produces discoveries. The distinction matters for anyone forecasting what AI does to science over the next few years.
How to read benchmarks like this
Two cautions apply to any new benchmark. The first is version number. This is 0.1, which means the task set will change as the authors learn which tasks are well posed and which are ambiguous. Scores from different versions are not directly comparable.
The second is that a benchmark measures the workflows its authors chose. Seventy tasks drawn from scientific domains is a sample, not a census. The 30 percent figure is a good signal of the general shape of agent capability on research work. It is not a precise number that will survive the next revision.
What to watch
The progress to watch is not the headline percentage going up. It is which tasks start passing. If agents begin clearing workflows that require multi-step recovery, that is a real advance. If the gain comes only from easier setup and data-cleaning tasks, the ceiling is lower than the headline suggests.
A benchmark that reports a low number honestly is more useful than one that reports a high number people cannot reproduce. This one is doing the first thing, and the field needs more of it.
Related articles
Agility Digit 5 Ships With a Safety Case, Not Just a Spec Sheet
A warehouse floor is not a lab. Certification is the gate, not the demo.
Figure AI Locked In $3.5 Billion of Compute Before It Has a Product to Sell
The bet is that generalisation is a compute problem. The field has not settled that.
OpenAI Finally Put Transparent Backgrounds in the Image API
A small feature that deletes a whole step from the pipeline.
Dolphin AI Turns a Script Into a Multi-Shot Video Without Losing the Costume
Continuity across cuts is the part most video models still fail.