Mistral Large 4 Takes on GPT-6 Astra With 1 Trillion Parameters Model

Mistral AI has opened access to Mistral Large 4, its most capable large language model to date.

The model is now available in public preview through Mistral’s cloud platform, while the company plans to release its weights later this month.

1 Trillion Parameters With Mixture-of-Experts Architecture

Mistral Large 4 uses a mixture-of-experts architecture with 1 trillion total parameters.

However, it activates only 49 billion parameters at a time, making it more hardware-efficient than models that activate their full parameter count for every request.

Mistral says the model can answer questions across more than 160 languages.

Specification Mistral Large 4
Architecture Mixture of Experts
Total Parameters 1 trillion
Active Parameters 49 billion
Languages 160+
Availability Public preview
Model Weights Planned for release later this month
Training Hardware 3,800 Nvidia Grace Blackwell chips

Strong Cybersecurity and Vision Performance

Mistral Large 4 earned a top-five score on the AA Cyber Index, which measures how well AI models can identify and fix software vulnerabilities.

The model performed particularly well at patching open-source software, scoring 82% in that category and outperforming open-source rivals.

Mistral also highlighted computer vision as a strength of Large 4. The model scored 1% higher than GPT-6 Astra on Dense200, a benchmark that measures how well models can identify objects of interest in images.

Benchmark / Area Performance
AA Cyber Index Top-five overall
Open-Source Software Patching 82%
Dense200 1% higher than GPT-6 Astra
Coding Benchmarks Behind frontier models such as GPT-6 Astra
Qwen 3.8 Max Mistral Large 4 scored higher on several coding benchmarks
DeepSeek V4 Pro Mistral Large 4 scored higher on several coding benchmarks
AutomationBench Higher than DeepSeek V4 Pro
AA-Briefcase Higher than DeepSeek V4 Pro

Mistral Large 4 still trails frontier models such as Astra on some of the industry’s most widely used coding benchmarks.

However, it outperformed several leading open-source models, including Qwen 3.8 Max and DeepSeek V4 Pro.

Large 4 also scored higher than DeepSeek V4 Pro on AutomationBench and AA-Briefcase. AutomationBench focuses on relatively simple tasks, while AA-Briefcase includes assignments that could take a human weeks to complete.

Trained on 3,800 Grace Blackwell Chips

Mistral trained Large 4 using 3,800 Nvidia Grace Blackwell chips.

Each accelerator combines two Blackwell graphics cards with one CPU.

The company did not disclose how long the training run took, but it shared details about the software infrastructure used during development.

Tens of Thousands of AI Rollouts at Once

Advanced AI models are partly trained through trial-and-error exercises known as rollouts.

During a rollout, the model receives a task and attempts to complete it without human guidance. A specialized AI model then analyzes the result and uses that feedback to improve the main model.

Mistral says it developed a software stack capable of running tens of thousands of rollouts in parallel.

The system assembles rollouts using components including coding sandboxes, a search engine for web access, and tests that check whether the model completed a task correctly.

This infrastructure was used to develop Mistral Large 4.

Training Generated 33 Billion Tokens Per Day

The model’s rollouts generated around 33 billion tokens per day.

Mistral used slightly less than half of those tokens in the training workflow responsible for refining Large 4.

The rollout and training processes operated asynchronously, meaning delays in one workload did not slow down the other.

Larger Versions Are Already in Development

Mistral did not stop the training process after producing the current version of Large 4.

The company expects the same workflow to produce larger and more capable versions of the model in the coming months.

Over the longer term, Mistral plans to use Large 4 as the foundation for a broader family of models optimized for specific use cases.

The post Mistral Large 4 Takes on GPT-6 Astra With 1 Trillion Parameters Model appeared first on ProPakistani.

Exit mobile version