Mistral opened a public preview of Mistral Large 4 this week, and it is the largest model the Paris company has shipped: a trillion total parameters with roughly 49 billion active, built as a natively multimodal Mixture-of-Experts. The company says the open weights will land by the end of the month. Until then the preview API runs on Mistral’s own infrastructure, priced at $1.36 per million input tokens and $4.18 per million output tokens, with a 1M token context window.

What the benchmarks actually show

The announcement leans hard on benchmark charts, so it is worth separating what is measured from what is marketed. On coding, ML4 scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4, for a combined Coding Agent Index of 49.8%. That puts it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max in Mistral’s own comparison. A blind human evaluation run with Surge AI rated its coding output 3.74 out of 5, second of five models, behind Claude Opus 5 at 4.22 and ahead of Kimi K3 and two GLM releases.

The agentic numbers follow the same shape. On AutomationBench, which scripts 657 business workflows across tools like Gmail, Sheets and Salesforce, ML4 scores 59.9%. On AA-Briefcase, a long-horizon knowledge work evaluation, it reaches 1,393 Elo, again ahead of DeepSeek V4 Pro.

Cybersecurity is the real headline

The most interesting part of the release is not the coding scores. It is the cybersecurity section, and specifically one test where ML4 scores 82%, the highest of any model measured. The test asks a model to reproduce a real vulnerability in open source software and then patch it. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test. Not because they cannot do the work, but because they refuse to.

Mistral’s argument is that defending software often starts with proving a flaw is real, and safety filters in closed models block exactly that work. There is a real tension in this position. The same capability that lets an incident responder reproduce an exploit lets an attacker automate one. Mistral’s answer is governance rather than gating: during the preview window it is red-teaming the model with cybersecurity firms, vetted partners and state authorities, and it says organizations that need auditable AI for security operations will be able to run the weights on private cloud or on premise.

It also reports 93% on Cybench, a set of 40 challenges drawn from security competitions, and says internal testing found the model useful for malware analysis, vulnerability prioritization and writing detection rules. Those claims are harder to verify independently, so treat them as directionally interesting rather than settled.

The European infrastructure bet

ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe, and the preview is served on the same hardware. The company positions this as sovereignty: a European deployment operated end to end under European law, independent of the big US clouds. Training data was heavily multilingual, spanning more than 160 languages including every official EU language.

The compute story has a concrete number behind it too. At the current scale of 3,000 GPUs, a single reinforcement learning run produces about 33 billion tokens per day, of which around 16 billion are trainable after filtering. Mistral says the RL run behind this preview is still in flight with no signs of saturation, which is why the preview model should keep improving before the weights ship.

All of this is funded by a 3 billion euro Series D, which Mistral calls the largest equity round ever raised by a European technology company.

What this means if you run models

Three practical takeaways. First, if your workloads include security research, red-teaming or incident response, the refusal behavior of closed models is a real operational constraint, and ML4 is explicitly targeting that gap. Whether that is a feature or a liability depends heavily on who holds the weights, and open weights mean the answer becomes “whoever deploys them”.

Second, the pricing matters. At $1.36/$4.18 per million tokens for a frontier-class model, ML4 undercuts most closed alternatives, and self-hosting the weights removes the API bill entirely once they drop.

One detail that got less attention than it deserves: the model card lists a 1.6B vision encoder alongside the language stack, and the announcement claims that on visual grounding tasks ML4 surpasses even frontier closed models. If that holds up under independent testing, it is a bigger deal than most of the coding numbers, because document and image understanding is where enterprise AI deployments actually make money today.

It is also worth being honest about what these comparisons are. Every number in the announcement comes from Mistral or from leaderboards it chose. Independent evaluations will land once the weights are out, and history says preview benchmark numbers drift downward when third parties get their hands on a model. The 1,649 points and 984 comments on the Hacker News thread tell you the community is paying attention, and community attention cuts both ways: scrutiny finds both the real wins and the soft spots.

Third, watch the weights release. Announcements are easy; a downloadable trillion-parameter MoE that actually matches the preview numbers is the part that has historically been where such claims soften. The end of the month deadline gives everyone a clear checkpoint to judge that on.

Leave a Reply

Your email address will not be published. Required fields are marked *