Models

Xiaomi Launches MiMo-V2.6 Open-Source AI Models With Omnimodal Capabilities

Xiaomi has released its MiMo-V2.6 series of generative AI models as open-source software, featuring two natively omnimodal variants designed to handle text, image, video and audio inputs simultaneously at a fraction of the cost of proprietary competitors.

·3 min read
Xiaomi introduces Mimo-V2.6 series open-source AI model family
Xiaomi introduces Mimo-V2.6 series open-source AI model family

Xiaomi Corp. unveiled the MiMo-V2.6 series of generative artificial intelligence models today, making them available as open-source software. The lineup comprises two natively omnimodal models that aim to deliver a combination of performance, operational efficiency and affordability.

The series includes the flagship MiMo-V2.6-Pro and a more compact Flash variant, alongside a streamlined Pro-UltraSpeed edition capable of producing outputs up to 20 times faster than the Pro model while maintaining comparable quality levels.

Both the Pro and Flash models function as native omnimodal systems, accepting text, image, video and audio simultaneously within the same model family. They feature context windows of 1 million tokens. The Pro variant incorporates a vision encoder with 681 million parameters, an AudioTokenizer containing 308 million parameters and an audio patch encoder with 127 million parameters designed to process and distinguish speech.

This architecture addresses the expanding demand for agentic AI applications and complex task instructions, particularly those involving computer interaction. A computer-use agent operating within this framework could process a task instruction, examine a user interface screenshot or video, analyze the content and determine the appropriate next action all within a single model iteration.

Coding and design agents gain the ability to combine source code and extensive repository context alongside screenshots, UI mockups and reference images in a unified session. For enterprise applications, the model can process multiple documents containing multimodal elements such as call recordings, screenshots, system logs and notes without requiring separate tools to convert transcripts or manage individual data types.

Xiaomi demonstrated the model's capabilities through examples including game world construction, Blender-based three-dimensional modeling, embodied simulation and video-music production workflows.

On the Artificial Analysis Intelligence Index, V2.6-Pro achieved a score of 46.32, surpassing Kimi K3 and Qwen3.8 Max, which represents the highest-ranked open-source or open-weight model in that benchmark at its launch. The model trails the most sophisticated proprietary closed-source offerings available, including Claude Fable 5.1 and GPT-6 Astra, when measured by aggregate performance indicators.

When evaluated on agentic tasks, V2.6-Pro demonstrated competitive performance relative to leading closed systems. It scored 53.1 on AutomationBench compared to Claude Opus 5's 50.3, matched Opus 5 at 31.6 on Agents' Last Exam and achieved 89.9 on Terminal Bench 2.1 versus Opus 5's 89.1.

Cost to efficiency benefit

Xiaomi will maintain the standard application programming interface pricing structure established for MiMo-V2.5. The model cards indicate availability through Xiaomi's AI Studio, MiMo Desktop and MiMo Code platforms.

The Pro and Flash models are also accessible via OpenRouter, offering 1.05-million-token context through a unified API.

V2.6-Flash pricing stands at $0.14 per 1 million input tokens and $0.28 per output token, while Pro costs $0.435 and $0.87 respectively. The Pro-UltraSpeed variant reaches $4.35 and $8.70 for its 20-times faster generation capability.

Xiaomi asserts that Pro costs approximately one-20th to one-60th the price of comparable overseas models at equivalent intelligence levels, accounting for cached-token expenses, which are substantially lower. This contrasts with OpenAI Group PBC's GPT-6 Astra and Anthropic PBC's Claude Fable 5.1, which charge $10 to $50 per million input and output tokens. For cache-intensive agent applications, the pricing differential expands further, though the relevant business calculation focuses less on per-token cost and more on per-task cost, which depends on model reliability and the number of iterations required for successful completion.