Mistral put Mistral Large 4 into public preview on 2026-10-06, and the concrete numbers are the interesting part: a 1-trillion-parameter natively multimodal model with 49 billion active parameters, priced at $1.36 per million input tokens and $4.18 per million output tokens, with the weights promised by the end of October. It is a hybrid instruct-and-reasoning MoE, natively fluent in 160+ languages, served today only through the preview API at docs.mistral.ai/models/mistral-large-4-0.

That is a model you cannot run on your own hardware yet — and when the weights land, 1T total parameters will still not fit on anything short of a multi-GPU node. The 49B active figure is what matters for the ones who do self-host: it sets the compute per token, not the memory footprint.

What The Benchmarks Actually Say

The coding numbers are the reference block worth keeping:

Benchmark ML4 score
DeepSWE v1.1 61.7%
SWE-Atlas-QnA 59.4%
Terminal-Bench 4.0 28.3%
Coding Agent Index (combined) 49.8%
AutomationBench (657 workflows) 59.9%
AA-Briefcase (long-horizon knowledge work) 1,393 Elo
Dense 200 (visual grounding) 42%

Mistral says the combined Coding Agent Index score of 49.8% places it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. On a blind human evaluation run with Surge AI, professional annotators rated ML4 Preview 3.74 on a 1–5 coding-quality scale — second of five, ahead of Kimi K3 (3.59) and GLM-5.3 (3.60), behind only Claude Opus 5 at 4.22.

The Cybersecurity Angle Is The Real Story

This is where ML4 diverges from the frontier closed models. On the Artificial Analysis Cyber Index it ranks among the top five globally and leads open-weight models developed outside China by a wide margin. On one index test — reproduce a real vulnerability in open-source software, then patch it — ML4 scores 82%, the highest of any model. It solves 93% of Cybench's 40 security-competition exercises, one of the highest scores reported for an open-weight model.

The reason is not raw capability. Mistral states plainly that several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on that same test because they refuse to perform the task. If your security work involves proving a flaw is real, provider-level refusals are not a nuisance — they are a blocker. ML4 pairs that with a surprisingly high refusal rate on malicious cyber prompts (higher than all OSS models across JailbreakBench, StrongREJECT, and AgentHarm), which is the balance you want if you are putting this in a SOC pipeline.

Safety-wise, ML4 resists 93.3% of attacks on Lakera's B3 AI Security Benchmark — Mistral reports no higher competitor score — and hits 1.691 on the KORA Benchmark, its highest measured OSS score.

Why The Deployment Geography Matters

ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, and the public preview runs on that same infrastructure. Mistral says the model will be available across multiple regions worldwide, including a European deployment it operates end-to-end under European law. If you have data residency constraints that keep you off US-hosted APIs, that is a real routing option — though it is still an API, with the same exit-clause and deprecation questions as any hosted model.

A significant share of the training data was multilingual, spanning more than 160 languages including every official EU language. Mistral also says the model was trained using the same training, customization, and RL environment it offers customers through Mistral Forge.

What I Would Watch Before Committing

The weights are not out. Mistral is still red-teaming with cybersecurity partners and state authorities, and says it will publish architecture details, additional benchmarks, and post-training methodology when the weights drop. Until then you are evaluating a preview API that can change under you.

A few things to watch:

  • The reasoning/instruction hybrid. ML4 is described as a hybrid instruct-and-reasoning MoE. How that mode is selected — and whether you can force reasoning on or off per request — is not stated in this announcement. That matters for token costs, since the $4.18 output rate is charged regardless of whether you wanted the thinking.
  • Terminal-Bench 4.0 at 28.3%. That is a contestable agentic-coding number. For comparison in the blog's recent coverage, Claude Sonnet 5.5 hit 70.6% on Terminal-Bench 4.0 and Beam put 80.1 on Terminal-Bench in 23B active parameters. ML4's strength is clearly not shell-heavy agent loops — it is document, vision, and cyber work. Route accordingly.
  • The schedule. Both the weights and any pricing change are still in the future. If your procurement needs a fixed number now, you have $1.36/$4.18 and a preview tag, and that is it.

For the workload ML4 was built for — analyzing engineering drawings, pulling evidence from PDFs, scanning geospatial imagery, working legal and finance agent tasks — the combination of 160+ language coverage, EU-hosted inference, and top-tier cyber scores is a genuinely different profile from the rest of the open-weight field. Whether that holds when the weights ship at the end of the month is the question I would not answer yet.

Sources: