← Back to blog
AiAbout 6 min read

Step 5 Skipped Four Version Numbers, Landed Second Among Open Models, and Nobody Is Calling It

Published Oct 5, 2026
Step 5 Skipped Four Version Numbers, Landed Second Among Open Models, and Nobody Is Calling It

At the end of September 2026, Chinese lab StepFun published Step 5 Preview. The version number is not a typo. The previous model in the line was Step 3.7 Flash, and the company jumped over the entire 4 series to get here. The open-source release is scheduled for October 15, and the model is already reachable through StepFun's products and API.

The specification is aimed at one job. Step 5 Preview uses 600 billion total parameters with 27 billion active, supports a one-million-token context, and natively accepts both text and images. StepFun positions it for coding, software engineering, and long-horizon agent tasks rather than for general chat.

On the Artificial Analysis Intelligence Index it scores 44. That puts it level with Kimi K3 and Grok 4.6, one point behind GLM-5.3, and joint second among open-source models globally. Three months earlier, Step 3.7 Flash sat at roughly 19 on the same index.

A tall stack of flat rough stone blocks balanced on a dark surface in hard light

The jump is the easy part to explain

A 25-point climb in one release sounds implausible until you look at what changed underneath it. Step 3.7 Flash was a small, fast model built for vision work. It was good at reading screenshots and video frames, and it was, by the account of developers who wired it into coding tools, not strong enough to hand the complicated work to. Step 5 Preview is a different class of model with an activation budget around three times the size, a context window sized for whole codebases, and training aimed explicitly at agentic behavior.

Chinese labs have been running this playbook through 2026. Release a fast small model, learn what breaks, then spend the compute on a frontier attempt. What is unusual here is the public version arithmetic, and the arithmetic works in StepFun's favor. Nobody who used Step 3.7 Flash will confuse it with Step 5 Preview, and the gap between them makes the improvement legible without a benchmark chart.

Independent early testing supports the claim with a mundane observation. A developer who connected the model to a coding agent reported that the output was stable and that engineering code came back needing fewer rework cycles. That is the kind of praise that matters for a coding model, and it is unglamorous enough to be credible.

The adoption number is the real story

Here is where the release gets interesting. OpenRouter publishes a weekly ranking of model call volume, and Step 5 Preview does not appear in the top ten. It has the benchmark placement, the context window, and the open license, and it is not being called at volume on the busiest multi-model router.

That gap is worth sitting with, because it describes a pattern across Chinese open-weight releases rather than a failure specific to this model. Benchmark position is a signal producers use to judge a model. Call volume is a signal consumers produce by using it. The two have been diverging all year.

Part of the explanation is structural. OpenRouter's volume is dominated by developers in markets where hosting a Chinese model carries procurement or compliance questions, and by workloads that are already integrated against whatever the team picked first. Switching costs are real even when the model is free, because integration work is not. A model can be better and still lose, if the switching cost exceeds the perceived gain.

Part of it is distribution. StepFun is not Google or Meta, and it does not have a channel that funnels developers into its endpoint. The open-source release on October 15 will tell us more than any benchmark, because weights that people can download and run locally bypass the hosting question entirely. That is the moment when a model stops needing a vendor relationship to be adopted.

There is also a plain-language problem that Chinese labs have been slow to solve. Every Western developer who tries Step 5 Preview has to navigate an API console, documentation, and error messages that were written for a domestic audience first. A model that is one point behind the leader on an index can be fifteen minutes behind on the first request, and the first request is where adoption is actually decided.

What the local-running path could change

The open-weight release is the part of this story with the most upside, and it is worth being specific about why.

A developer who can download weights removes two obstacles at once. There is no vendor relationship to establish, and there is no question about where the inference happens or whose jurisdiction the data crosses. For a team in Europe or North America weighing a Chinese model against a domestic one, that difference is often the deciding factor rather than a footnote. It is the same reason Llama and Qwen derivatives show up inside enterprises that would never call a hosted endpoint in another country.

The catch is that open weights shift the cost from a per-token bill to hardware and operations. A 600-billion-parameter model with 27 billion active parameters is not running on a workstation. Serving it well enough to beat a hosted endpoint on cost requires either enterprise GPUs or an aggressive quantized build, and the quantized builds arrive from the community rather than from the lab. That timeline, not the license date, sets when the model becomes practically available to most teams.

StepFun's release schedule suggests the company understands this, since the weights land on October 15 while the model has been available through its own API since the end of September. That gap gives the lab a month of production feedback to incorporate before the version people can actually inspect becomes public.

What this says about the open-weight market

Step 5 Preview arrives into a crowded field that has shifted its competitive axis over the past year. In early 2026, the question was which model topped a leaderboard. By the fall, the questions that decide adoption had changed to something more practical: what does the license permit, which chip vendors have adapted their toolchains, how complete is the fine-tuning stack, and how active is the community around it.

That shift matters for anyone choosing a model today. A model that leads on a benchmark but has no clear deployment path, no quantized builds, and a thin community will cost more in integration time than a slightly weaker model with working tooling. The benchmark tells you what the model can do. The ecosystem tells you what it will cost you to find out.

StepFun's bet appears to be that the capability jump buys attention, and that the October 15 open-weight release converts some of that attention into use. That is a reasonable bet, and it is also the standard one. Almost every Chinese frontier release of 2026 has followed the same arc: strong benchmark placement, open weights, a cascade of adaptations from domestic chip vendors, then a quiet question about whether the weekly call volume moved at all.

What is genuinely notable about Step 5 Preview is the discipline in the specification. A one-million-token context with agentic training and a 27-billion active parameter budget is a coherent product for engineering work rather than a general-purpose contender. StepFun skipped four version numbers and did not overreach on the task list. Whether developers notice is a different matter, and the answer will show up in call volume rather than in a review.

Related articles