Xiaomi's MiMo-V2.6-Pro entered the model market on September 21 with support for text, images, speech and video as input, text output and a context window of one million tokens. Early measurements from Artificial Analysis characterize the open-weight reasoning model as fast and relatively capable within its comparison class, while also showing that its input price sits above the median used by the evaluator.

The model recorded 46 on version 4.3.2 of the Artificial Analysis Intelligence Index. The evaluator says the median among comparable open-weight models of a similar size is 18, placing MiMo-V2.6-Pro well above that reference point. The composite index draws on a group of assessments covering analytical work, automation, coding, science, broad knowledge and long-context reasoning.

That score should be read as the result of one benchmark system rather than a universal ranking. Artificial Analysis combines tests including AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR. Different applications can assign very different importance to those capabilities, and a single aggregate can conceal strengths or weaknesses on individual tasks.

Speed is a prominent part of the initial profile. Artificial Analysis measured output at roughly 125 tokens per second, which it labels notably fast for the model's class. Generation speed is measured after the first response chunk arrives, so it does not by itself describe the complete wait experienced by a user. Time to first token, reasoning time and provider conditions can all affect end-to-end latency.

Listed pricing is 43 cents per million input tokens and 87 cents per million output tokens. In the comparison displayed by the evaluator, the input rate is higher than a 30-cent median, while the output rate is below a $1.13 median. Artificial Analysis says running its Intelligence Index evaluation cost $206.66. Those figures provide a standardized snapshot, although actual deployment cost will depend heavily on prompt length, output length, caching and workload design.

The million-token context window expands the amount of material that can be supplied in one request. That can be useful for retrieval-augmented systems, long documents and workflows that combine many records, but the headline capacity does not establish how reliably a model uses every part of a very long prompt. Context size, retrieval quality and reasoning accuracy are separate properties.

Artificial Analysis also describes MiMo-V2.6-Pro as somewhat verbose. Output length can raise cost and latency even when the per-token rate appears favorable, making response discipline an operational consideration for developers. The benchmark's cost metrics account for the tokens generated during its own test suite, but a production system may elicit a different pattern.

The initial numbers position MiMo-V2.6-Pro as a high-throughput, multimodal option with a large context allowance and competitive benchmark performance. They do not remove the need for application-specific evaluation. Teams considering the model will still need to test factual reliability, instruction following, long-context use and total response time on their own data before treating the composite score as a deployment decision.