All changelog entries
FeatureNew0.14.0ModelsVision
MiMo-V2.5 — one model for text, images, video and audio
September 18, 2026
MiMo-V2.5 from Xiaomi is live. It is natively omnimodal: it takes text, images, video and audio in a single model, with a 1M-token context window. It is a 310B mixture-of-experts model with 15B active per token, released under the MIT license. Call it as MiMo-V2.5.
Xiaomi's published scores:
| Benchmark | Score |
|---|---|
| MMMU-Pro | 77.9 |
| CharXiv RQ | 81.0 |
| OmniDocBench | 87.2 |
| Video-MME | 87.7 |
| DailyOmni | 83.5 |
| SWE-Bench Pro | 56.1 |
| Terminal-Bench 2.0 | 65.8 |
It costs 1.5 requests of quota per call and is available on every paid plan and the free tier, the same terms as the other flash vision models. Reach for it when the input is a document, chart, screen recording or audio clip rather than plain text.
Source: MiMo-V2.5 model card.
