Goldman Sachs: XIAOMI-W (01810) MiMo-V2.6 RL livestream shows better Flash returns; maintains "Buy" rating

date
10:41 22/09/2026
avatar
GMT Eight
Goldman Sachs believes that MiMo-V2.6-Flash appears to offer more favorable RL returns, achieving a DeepSWE score increase of over 19 points within 3.5 days at a total expenditure of US$850,000.
Goldman Sachs released a research report maintaining a "Buy" rating on XIAOMI-W (01810) with a 12-month target price of HK$39. Xiaomi's MiMo began publicly livestreaming the MiMo-V2.6 RL (reinforcement learning) process several days ago, with V2.6-Flash and V2.6-Pro stopping after 30 training steps. Goldman Sachs believes this is the first time a major AI lab has publicly disclosed post-training telemetry data in real time, and RL training has significantly improved long-horizon agent performance. Within 30 training steps, MiMo-V2.6-Flash/Pro achieved maximum scores of 67.86/72.57 (avg@3) on the DeepSWE benchmark. The bank noted that MiMo-V2.6-Flash appears to offer more favorable RL returns on investment, achieving a DeepSWE score improvement of over 19 points within 3.5 days at a total expenditure of US$850,000 (US$10 per million tokens or US$2.85 per second), while MiMo-V2.6-Pro spent US$2.62 million (US$35 per million tokens or US$5.71 per second, 2 to 3 times higher than Flash) to gain a DeepSWE score improvement of over 14 points. The bank estimates that MiMo-V2.6-Flash/Pro's RL training was based on approximately 4,000/8,000 H200-equivalent GPU clusters, and the mid-training interventions shown in the livestream, including infrastructure error reruns, bad pattern datasets, zero-gradient tasks, and GPU out-of-memory issues, indicate that the RL bottleneck has shifted to engineering optimization. Goldman Sachs expects that pricing at the official release of the MiMo-V2.6 series will draw attention, with V2.6 consolidating Xiaomi's focus on agentic coding and multimodal integration and offering higher API pricing potential, but Xiaomi may continue to disrupt the market pricing structure alongside DeepSeek. The bank also expects that MiMo-V3, which adopts the new HySparse architecture, may be released within a few months, around early 2027, with its sparsity ratio potentially increasing from 7:1 in V2/V2.5 to a more aggressive 11:1, further pushing the efficiency frontier and reducing inference costs.