All changelog entries
FeatureNew0.14.0ModelsVision

MiMo-V2.5 — one model for text, images, video and audio

September 18, 2026

MiMo-V2.5 from Xiaomi is live. It is natively omnimodal: it takes text, images, video and audio in a single model, with a 1M-token context window. It is a 310B mixture-of-experts model with 15B active per token, released under the MIT license. Call it as MiMo-V2.5.

Xiaomi's published scores:

BenchmarkScore
MMMU-Pro77.9
CharXiv RQ81.0
OmniDocBench87.2
Video-MME87.7
DailyOmni83.5
SWE-Bench Pro56.1
Terminal-Bench 2.065.8

It costs 1.5 requests of quota per call and is available on every paid plan and the free tier, the same terms as the other flash vision models. Reach for it when the input is a document, chart, screen recording or audio clip rather than plain text.

Source: MiMo-V2.5 model card.