Xiaomi Is Suddenly Competing With the World’s Top AI Labs
Xiaomi’s MiMo-V2.6-Pro has climbed to the top of Artificial Analysis’ open-weight model rankings, putting it alongside some of the strongest proprietary AI models currently available.
Artificial Analysis gives MiMo-V2.6-Pro an Intelligence Index score of 46, the same score as Grok 4.7 High. It currently sits ahead of Kimi K3 Max at 44, GLM-5.3 Max at 45, and GLM 5.3 Flash at 42.
Artificial Analysis Ranking
| Model | Intelligence Index | Open Weight | Context Window |
|---|---|---|---|
| MiMo-V2.6-Pro | 46 | Yes | 1M tokens |
| Grok 4.7 High | 46 | No | 500K tokens |
| GLM-5.3 Max | 45 | Yes | 1M tokens |
| Kimi K3 Max | 44 | Yes | ~1.05M tokens |
| GLM 5.3 Flash | 42 | Yes | 1M tokens |
The result is notable because MiMo 2.6 Pro and Grok 4.7 arrived at almost the same time in late September, yet Xiaomi’s model reaches the same overall Artificial Analysis score while remaining open-weight.
MiMo 2.6 Pro vs Grok 4.7
The overall scores are identical, but the individual benchmarks show different strengths.
| Benchmark | MiMo 2.6 Pro | Grok 4.7 High |
|---|---|---|
| Artificial Analysis Intelligence Index | 46 | 46 |
| AA-Briefcase v1.1 | 1,515 | 1,632 |
| GDPval-AA v2.1 | 1,686 | 1,710 |
| AutomationBench-AA | 59% | 64% |
| Terminal-Bench 4.0 | 35% | 25% |
| SciCode | 61% | 58% |
| Humanity’s Last Exam | 49% | 42% |
| GDP.pdf | 19% | 23% |
| CritPt | 27% | 18% |
| AA-LCR v1.1 | 86% | 77% |
MiMo performs better on Terminal-Bench, SciCode, Humanity’s Last Exam, CritPt and long-context reasoning, while Grok leads on AA-Briefcase, GDPval, AutomationBench and GDP.pdf.
This means MiMo should not be treated as better than Grok in every area. Instead, the two models show different strengths despite having the same overall score.
Much Cheaper Than Grok 4.7
Cost is where MiMo has a much larger advantage.
| Metric | MiMo 2.6 Pro | Grok 4.7 High |
|---|---|---|
| Input / 1M tokens | $0.435 | $2.00 |
| Output / 1M tokens | $0.87 | $6.00 |
| Cached input / 1M tokens | $0.0036 | $0.50 |
| Average cost per Intelligence Index task | $0.13 | $2.73 |
| Cost to run full Intelligence Index | $207 | $3,881 |
Artificial Analysis estimates that an average Intelligence Index task costs about $0.13 on MiMo 2.6 Pro compared with $2.73 on Grok 4.7 High.
That makes MiMo roughly 21 times cheaper per benchmark task in this particular evaluation.
MiMo Is Not Faster
Xiaomi’s cost advantage does not extend to raw generation speed.
| Performance | MiMo 2.6 Pro | Grok 4.7 High |
|---|---|---|
| Output speed | ~46 tokens/sec | ~78 tokens/sec |
| Time to first token | 4.08 sec | 33.12 sec |
| Time to first answer token | 47.99 sec | 33.12 sec |
| End-to-end response time | 58.97 sec | 39.49 sec |
Grok produces tokens faster once generation begins, while MiMo starts processing sooner but spends longer reasoning before delivering its final answer.
Xiaomi also offers MiMo-V2.6-Pro-UltraSpeed, which is designed to improve output speed for latency-sensitive workloads.
MiMo 2.6 Pro vs Kimi K3
MiMo also leads Kimi’s current flagship on the overall index.
| Benchmark | MiMo 2.6 Pro | Kimi K3 Max |
|---|---|---|
| Intelligence Index | 46 | 44 |
| AutomationBench-AA | 59% | 58% |
| Terminal-Bench 4.0 | 35% | 13% |
| SciCode | 61% | 59% |
| Humanity’s Last Exam | 49% | 47% |
| CritPt | 27% | 23% |
| AA-LCR v1.1 | 86% | 89% |
| Cost per task | $0.13 | $2.00 |
Kimi K3 performs better in some long-context testing, but MiMo leads on the overall index and several coding, science and reasoning tests.
MiMo 2.6 Pro vs GLM-5.3 Max
The gap between MiMo and GLM-5.3 Max is much smaller.
| Benchmark | MiMo 2.6 Pro | GLM-5.3 Max |
|---|---|---|
| Intelligence Index | 46 | 45 |
| AutomationBench-AA | 59% | 62% |
| Terminal-Bench 4.0 | 35% | 42% |
| SciCode | 61% | 59% |
| Humanity’s Last Exam | 49% | 42% |
| CritPt | 27% | 19% |
| AA-LCR v1.1 | 86% | 80% |
| Cost per task | $0.13 | $2.01 |
GLM performs better on AutomationBench and Terminal-Bench, while MiMo scores higher on science, difficult reasoning, and long-context tasks.
A 1-Trillion-Parameter Model
MiMo-V2.6-Pro uses a Mixture-of-Experts architecture with 1 trillion total parameters and 42 billion active parameters during inference.
It supports a 1-million-token context window and up to 128,000 output tokens.
The model is released under the MIT license, allowing commercial use and modification.
Built for Agents and Long Tasks
Xiaomi positions MiMo 2.6 Pro for complex projects, long-running tasks, research, cybersecurity, and agentic workflows.
The model supports tool calling, web search, structured output, streaming, and context caching.
Xiaomi has also used large-scale reinforcement learning to improve software engineering performance. During recent training, MiMo 2.6 Pro reportedly improved from 58.4 to 72.6 on DeepSWE v1.1.
The training involved around 750,000 trajectories across MiMo 2.6 Flash and Pro, with Xiaomi saying the Pro model’s portion cost roughly $2.62 million.
The post Xiaomi Is Suddenly Competing With the World’s Top AI Labs appeared first on ProPakistani.



