← Back to blog
AiAbout 6 min read

The Model Picker Is Turning Into an Orchestration Picker

Published Oct 3, 2026
The Model Picker Is Turning Into an Orchestration Picker

GitHub moved HydraFusion out of the command line and into the editor on September 30. The research preview now runs in VS Code 1.140 and later, and in the GitHub Copilot app, which puts a fairly unusual idea in front of ordinary developers: stop choosing a model, choose a workflow.

HydraFusion is not a model. It is a layer that decides, for each turn, how many model roles the task needs.

Three shapes for one request

Single is the familiar case. One model handles the request, the way Copilot Auto already works.

Cascade starts cheap. An efficient model writes a first draft, and a quality gate either accepts it or escalates the work to a stronger model. The point is to keep routine edits on inexpensive models and reserve the expensive ones for problems that actually need them.

Critique spends more to be safer. One model drafts, a second model from a different family acts as a read-only critic, and the drafting model revises once based on that feedback. The critic never touches the code, which keeps the loop from turning into two models arguing.

The difference from Auto is architectural. Auto decides which single model should receive your prompt. HydraFusion decides how many roles the turn needs and how they interact. That is a meaningful shift in where the intelligence control lives. Historically you picked a model because it was better at your kind of work. Here you pick an optimization policy and let the platform compose the models underneath.

A polished chrome junction splitting one path into three separate channels on a dark reflective surface, warm light pooling in the cleft

The economics are real and also self-reported

The catch is billing. GitHub charges by tokens at each selected model's standard rate. A Critique run calls at least two models, so it can cost more per request than a single call. Cascade is designed to cost less by pushing easy work downward. The saving is not a discount on any model. It is a bet that most turns do not need the best one.

GitHub's own controlled evaluations reported workflow-cost reductions of 36 to 67 percent against its Claude Opus 5 baseline across three coding-agent benchmarks, with quality 4.9 percentage points higher on TerminalBench 2.1, 1.5 points lower on DeepSWE, and 0.1 point lower on CheckpointBench. Those are vendor-produced numbers on GitHub's own test setup. They are a hypothesis about your repository, not a guarantee, and the model pool participating in HydraFusion has not been fully published. Anything built on the preview should treat routing behavior as something that can change without notice.

There is also a plain latency cost. A draft plus a review takes longer than one answer. Teams on Business and Enterprise plans have to watch usage dashboards closely, because the system decides when to escalate and that decision is not free.

An optimization you cannot inspect is hard to trust

The awkward part of routing is that the thing making the decision is the platform, not the developer. You can see which model answered after the fact, but you cannot easily see why the system chose a workflow, what the quality gate measured, or how close a Cascade run came to escalating. GitHub's benchmark claims describe an average across its own test set, and averages hide the cases that matter.

That gap turns orchestration into a governance problem rather than a purely technical one. An engineering team that has to explain why a particular pull request was reviewed by two model families and charged accordingly needs routing telemetry, and the preview does not yet expose it. The remedy is not to avoid the feature. It is to test it the way you would test any dependency whose internals are hidden: on a fixed, boring set of repository tasks, measuring completed-task cost against a single frontier model, and watching retries and review effort rather than the sticker price of one selection.

The comparison to Auto is worth holding on to. Auto answers a question about which model is best suited to a prompt. HydraFusion answers a question about how much process a turn deserves, and process has always been the expensive part of software work.

The curious side effect is that it makes model choice less emotional. Developers develop attachments to particular models, and those attachments are usually built on a handful of memorable successes. A system that routes by task type and escalates on failure quietly acknowledges that no single model wins every category, which has been true for a while and is rarely acted on.

The next pocket of the editor is being built for this

The Insiders build already shows where it goes next. A feature called Compare Agents, labeled Run Multiple Agents, sends one prompt to several agents in parallel, each in its own git worktree, then has a referee agent pick a winner or hand the shortlist to a person.

The interesting detail is what the referee looks at. It compares changed files and diff statistics, test results, build status, diagnostics, timing, and architectural differences. It runs the test suite. It does not ask another language model to read the code and guess. That is a small decision with large consequences, because it means the comparison is grounded in the actual state of the project rather than in a model's opinion of it.

The rest of the same release is quieter infrastructure that only matters once agents run in parallel. Multi-folder sessions let each conversation in one session use its own folder or worktree, so changes stop colliding. Remote delegation hands a task to a connected remote agent host. Shared worktree folders reuse ignored directories to avoid reinstalling dependencies on every branch. Dev Containers now stop after five minutes of idle time and restart on demand.

Governance keeps arriving in the same commits as features

The IT-facing half of the release is not an afterthought. When AI features are unavailable, the editor now states the minimum required version instead of a generic prompt to update. Administrators can set a default tier for the Auto model. A new OpenTelemetry setting maps Copilot usage back to individual developers.

Read those three together and the shape is clear. A platform that coordinates several models across several turns generates costs, telemetry, and permission questions that a single-model autocomplete never did. Cost attribution stops being a finance curiosity and becomes a prerequisite for letting the feature anywhere near production.

Multi-model review has been standard practice among careful engineering teams for a while. You ask one model to write and another to look for mistakes, or you try a fast model first and escalate when the output disappoints. HydraFusion takes that habit and makes it automatic, which is convenient and also removes a decision that some teams liked making by hand.

How long a preview feature should stay a preview is a fair question. The tooling is landing faster than the policy around it, and price predictability is the first casualty. A Critique run that quietly calls two frontier models on a large repository can spend real money before anyone notices. Whether HydraFusion graduates to a default depends less on whether the routing works and more on whether GitHub can make the bill legible enough for someone to sign it.

Related articles