← Back to blog
AiAbout 6 min read

Mistral Large 4: Europe Bets a Trillion Parameters on Open Weights

Published Oct 11, 2026
Mistral Large 4: Europe Bets a Trillion Parameters on Open Weights

On October 6, Mistral AI put a trillion-parameter model into public preview and gave it a nickname: Le Chonk. The French lab says Mistral Large 4 is the strongest open-weight model to come out of either the United States or Europe. The weights themselves will not ship until October 27, but the API is already live, and the benchmarks are already public.

The nickname is a wink, but the numbers behind it are not small. Large 4 carries 1.05 trillion total parameters in a sparse mixture-of-experts layout, and it activates about 49 billion of them on any given task. Mistral's own documentation rounds 52 billion active in one place and early reports cited 49 billion in another. Either way, the design goal is the same one every large lab now chases: get the capability of a very large model without paying to run all of it on every token.

What is actually inside

The model is natively multimodal. Mistral paired the language backbone with a separate 1.6 billion parameter image encoder, and it handles both text and images. The context window reaches one million tokens. Training data skewed heavily multilingual, covering more than 160 languages, including every official language of the European Union.

Mistral trained Large 4 from scratch on roughly 3,800 to 4,000 Nvidia Grace Blackwell GPUs inside its own European data centers, over about two months. The same hardware now serves the public preview. The company also says reinforcement learning is still running on the model, so it expects the numbers to move over the coming weeks.

A dense glowing core of many small nodes with only a thin band lighting up at once

The benchmark table, read carefully

On agentic coding, Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4, for a combined Coding Agent Index of 49.8%. That figure, by the company's account, edges past DeepSeek V4 Pro 0813 and Qwen3.8 Max. On AutomationBench, which runs 657 business workflows across tools like Gmail, Google Sheets, Slack and Salesforce, Large 4 scored 59.9%.

The most interesting claim is about visual grounding. On the Dense 200 benchmark, which measures how well a model can locate a specific object inside an image, Mistral puts Large 4 at 42%, one point above GPT-6 Astra at 41%. A one-point gap on a single evaluation is not a knockout, and it should not be read as one. But it is the kind of task where open models have trailed closed ones, so a tie is worth noting.

Cybersecurity is where the tone changes from marketing to argument. Mistral says Large 4 ranks in the top five globally on the Artificial Analysis Cyber Index. It reports an 82% success rate on reproducing and patching real vulnerabilities in open-source software, and 93% on Cybench, a set of 40 security competition problems. The company also makes a pointed observation: some closed models refuse to reproduce a vulnerability at all, because the request looks like offensive activity, which drives their scores near zero on the same tests.

A blind evaluation tells a more grounded story. In a test run with Surge AI, professional reviewers scored generated code on a five-point scale without knowing which model wrote it. Large 4 landed at 3.74. Claude Opus 5 scored 4.22, and GLM-5.3 scored 3.60. So on code quality as judged by humans, Large 4 sits second among the five models compared, ahead of the open competitors but behind the closed leader.

Every one of these figures comes from Mistral. They are preliminary, they are self-reported, and no independent party has reproduced them yet. Treat them as a strong starting claim, not a settled result.

Pricing and the open-weights bet

The API is available now through Mistral Studio and OpenRouter. List price runs at $1.36 per million input tokens and $4.18 per million output tokens, with a launch promotion at half those rates, $0.68 and $2.09. On October 27, the weights arrive, and anyone can download them and run Large 4 on their own hardware without renting access.

That last part is the whole strategy. Mistral cannot outspend OpenAI or Anthropic on compute, and it cannot match the release cadence of the Chinese open-weight labs that ship large models almost monthly. What it can offer is control. An enterprise with data-residency rules, or a security team that needs a model to help reproduce a vulnerability without a provider refusing the request, gets something a hosted closed model does not provide. Mistral leans hard into that pitch, describing Large 4 as forged in Europe end to end and deployable from Europe through its own cloud.

The competitive timing makes the stakes clear. Reflection AI, a Nvidia-backed lab, released a 501-billion-parameter open-weight model on October 5, a day before Large 4. Chinese labs including Alibaba, Z.ai, Moonshot and DeepSeek have spent the past two months shipping strong open models on a rolling basis. Analysts tracking benchmark convergence report that the gap between open-weight and proprietary models has narrowed on selected reasoning tasks, and that premium proprietary API prices have fallen as a result. Mistral is entering a market where being the biggest open model is worth less than being the most trusted one, and trust is built on the licenses and deployment options, not on parameter counts.

The cybersecurity red-teaming follows the same logic. Before the weights ship, Mistral is giving vetted partners and state authorities access to the same model "with reduced moderation and expanded cyber capabilities." The company argues that provider-level refusals can block legitimate incident response, and that losing access to a tool mid-incident is itself a risk.

The release also lands at a specific moment. Mistral recently closed a funding round that pushed its valuation past $24 billion, with Samsung among the new investors. Reflection AI shipped a 501-billion-parameter open model just a day earlier. The open-weight tier is getting crowded, and the differentiator is no longer raw size.

What this changes, and what it does not

Large 4 does not close the gap with the best closed models on every task, and Mistral does not quite claim it does. What it does is narrow the gap on a handful of jobs that matter to enterprises, and pair that with a licensing and deployment story those enterprises actually need. If the October 27 weights hold up under outside testing, the practical question stops being whether an open model is good enough in the abstract and becomes whether it is good enough for one specific, well-defined job.

The rarer claim to watch is the cyber one. A model that helps security teams patch real vulnerabilities, run under their own policies, on their own hardware, is a different kind of product from a chatbot that declines the request. If that holds, it may matter more than a few points on a coding benchmark. If external testers find the 82% figure does not travel outside Mistral's own setup, the pitch collapses back to price and portability, which is a crowded place to stand.

Either way, Europe now has a flagship of its own to argue about. That argument starts in earnest on October 27.

Related articles