Mistral Large 4 enters preview—and independent benchmarks add important context

Mistral’s trillion-parameter multimodal model enters API preview with weights promised later in October. Independent Vals results show strong finance and legal-agent scores but a much more mixed overall profile.

2 min read

Illustrative close-up of a circuit board representing AI agent software infrastructure

Mistral launched a public preview of Mistral Large 4 on October 6, positioning it as a multimodal mixture-of-experts model for coding, tool-using agents, cybersecurity and document work. It is available now through Mistral Studio and the API; downloadable weights are promised by the end of October, so this is not yet a downloadable release.

Mistral’s model documentation lists 1.05 trillion total parameters, 52 billion active parameters, a 1.6B vision encoder and a one-million-token context window. The launch post rounds those figures to 1T total and 49B active, an inconsistency developers should note while the preview is still changing. API pricing is listed at $1.36 per million input tokens and $4.18 per million output tokens.

The company highlights 61.7% on DeepSWE 1.1, 28.3% on Terminal-Bench 4.0 and 59.9% on AutomationBench. Those launch numbers are publisher-selected, and the post says the reinforcement-learning run is still continuing. Independent evaluator Vals AI presents a more mixed picture: with high reasoning effort, temperature 1 and up to 256,000 output tokens, ML4 scored 48.05% on its aggregate index, ranking 32nd of 44. It did narrowly exceed GPT-6 Astra on Finance Agent v2, 54.68% to 53.54%, and more substantially on Harvey’s Legal Agent benchmark, 15.83% to 5.42%. On Terminal-Bench 4.0, however, Vals measured 22.73% for ML4 versus 59.60% for Astra. These task-level differences are why no single benchmark establishes overall superiority.

Mistral says weights will arrive later this month, but it has not yet published the weight license, final architecture details or full post-training methodology. Until those appear, “open-weight” describes the planned distribution model, not a license available today.

Server hardware representing the compute infrastructure behind large AI models. Illustrative image, not a Mistral product image.

Sources: Mistral launch announcement · Mistral model documentation · Vals AI evaluation of Mistral Large 4 · Vals AI evaluation of GPT-6 Astra · Unsplash image license

benchmarksmultimodalMistral Large 4Mistral AImixture of expertsopen weights