Mistral Large 4 Takes on GPT-6 Astra With 1 Trillion Parameters Model
Mistral AI has opened access to Mistral Large 4, its most capable large language model to date.
The model is now available in public preview through Mistral’s cloud platform, while the company plans to release its weights later this month.
1 Trillion Parameters With Mixture-of-Experts Architecture
Mistral Large 4 uses a mixture-of-experts architecture with 1 trillion total parameters.
However, it activates only 49 billion parameters at a time, making it more hardware-efficient than models that activate their full parameter count for every request.
Mistral says the model can answer questions across more than 160 languages.
| Specification | Mistral Large 4 |
|---|---|
| Architecture | Mixture of Experts |
| Total Parameters | 1 trillion |
| Active Parameters | 49 billion |
| Languages | 160+ |
| Availability | Public preview |
| Model Weights | Planned for release later this month |
| Training Hardware | 3,800 Nvidia Grace Blackwell chips |
Strong Cybersecurity and Vision Performance
Mistral Large 4 earned a top-five score on the AA Cyber Index, which measures how well AI models can identify and fix software vulnerabilities.
The model performed particularly well at patching open-source software, scoring 82% in that category and outperforming open-source rivals.
Mistral also highlighted computer vision as a strength of Large 4. The model scored 1% higher than GPT-6 Astra on Dense200, a benchmark that measures how well models can identify objects of interest in images.
| Benchmark / Area | Performance |
|---|---|
| AA Cyber Index | Top-five overall |
| Open-Source Software Patching | 82% |
| Dense200 | 1% higher than GPT-6 Astra |
| Coding Benchmarks | Behind frontier models such as GPT-6 Astra |
| Qwen 3.8 Max | Mistral Large 4 scored higher on several coding benchmarks |
| DeepSeek V4 Pro | Mistral Large 4 scored higher on several coding benchmarks |
| AutomationBench | Higher than DeepSeek V4 Pro |
| AA-Briefcase | Higher than DeepSeek V4 Pro |
Mistral Large 4 still trails frontier models such as Astra on some of the industry’s most widely used coding benchmarks.
However, it outperformed several leading open-source models, including Qwen 3.8 Max and DeepSeek V4 Pro.
Large 4 also scored higher than DeepSeek V4 Pro on AutomationBench and AA-Briefcase. AutomationBench focuses on relatively simple tasks, while AA-Briefcase includes assignments that could take a human weeks to complete.
Trained on 3,800 Grace Blackwell Chips
Mistral trained Large 4 using 3,800 Nvidia Grace Blackwell chips.
Each accelerator combines two Blackwell graphics cards with one CPU.
The company did not disclose how long the training run took, but it shared details about the software infrastructure used during development.
Tens of Thousands of AI Rollouts at Once
Advanced AI models are partly trained through trial-and-error exercises known as rollouts.
During a rollout, the model receives a task and attempts to complete it without human guidance. A specialized AI model then analyzes the result and uses that feedback to improve the main model.
Mistral says it developed a software stack capable of running tens of thousands of rollouts in parallel.
The system assembles rollouts using components including coding sandboxes, a search engine for web access, and tests that check whether the model completed a task correctly.
This infrastructure was used to develop Mistral Large 4.
Training Generated 33 Billion Tokens Per Day
The model’s rollouts generated around 33 billion tokens per day.
Mistral used slightly less than half of those tokens in the training workflow responsible for refining Large 4.
The rollout and training processes operated asynchronously, meaning delays in one workload did not slow down the other.
Larger Versions Are Already in Development
Mistral did not stop the training process after producing the current version of Large 4.
The company expects the same workflow to produce larger and more capable versions of the model in the coming months.
Over the longer term, Mistral plans to use Large 4 as the foundation for a broader family of models optimized for specific use cases.
The post Mistral Large 4 Takes on GPT-6 Astra With 1 Trillion Parameters Model appeared first on ProPakistani.



