Originally presented as a live talk on October 7, 2026
Paper · RL envs · RL envs explorer
Background
So, let’s talk about MiMo.
MiMo is the name for all the models from Xiaomi, a Chinese consumer electronics giant. In fact, the name MiMo just takes the “mi” from Xiaomi and adds “mo” on the end, which is short for “model”. The straightforward and perhaps uncreative name suits, as we’ll see in a bit.
Anyway, Xiaomi makes an incredible range of stuff: phones, computers, appliances, power tools, scooters, cars, even shoes. And they also make the software that runs it all. Well, not the shoes, but the rest of it yes.
So as many other big tech companies have done, Xiaomi has gotten into the LLM game. At minimum, they want their own AI to run their vast empire of stuff in the same way SpaceXAI does. But like SpaceXAI, they also want external users of their model, to make AI yet another line of business.
Xiaomi’s efforts started out modestly, with a 7B parameter model in April 2025. I recall no impact at all from their initial release, which is perfectly fine for a new lab catching up to the frontier. It’s just about proving you can do the thing.
Where it really starts to get interesting is later in 2025, when a major researcher from DeepSeek jumped ship to lead development of MiMo. After that the catchup got a lot quicker:
MiMo V2 Flash, at 309B parameters, in December 2025
MiMo V2 Pro, at 1T parameters, in March 2026
MiMo V2.5, at the Flash size but without the Flash name, in April 2026
V2.5 in particular was notable for being nearly identical in benchmarks and in price to DeepSeek V4, which also came out in April 2026. So at that point Xiaomi had made it crystal clear what their strategy was: copy the best Chinese lab and use their size and distribution channels to muscle them out.
As you would expect from corporate copycats, the MiMo series historically didn’t generate much notable research. The lab’s approach has started to change though, and the V2.6 development process has gained respect from the AI community for being remarkably open.
The screenshot here is part of the story. Xiaomi actually livestreamed their RL, with top-level metrics for accuracy, tokens, and cost. There are graphs for many granular metrics too, like the average staleness of rollouts and the KL divergence. It’s a remarkable view into what RL actually looks like from a researcher’s perspective, and a fair show of confidence to make your ups and downs so public. They even issued notices at the top about infra problems they encountered.
And of course, like DeepSeek, Xiaomi makes its model weights freely available under a permissive license.
Speaking of other model makers, this is where V2.6 Pro and Flash land on the Artificial Analysis Intelligence Index. It’s not the only way to index benchmarks into a ranking, but it is an industry standard.
A couple things to note. One is that the US models are clearly ahead. We’ll see later on that Xiaomi has chosen only US models for comparison, albeit the previous and not current set of frontier releases. As the heuristic dictates, the Chinese models remain about six months behind.
Within the Chinese models though, V2.6 Pro apparently leads the pack! It’s a lead of only one point, and I expect the next Pro-tier model from a leading Chinese lab to surpass it, but still impressive for a lab that has only been making models this big for about a year now.
In the Flash tier, MiMo is a bit behind DeepSeek, and here I credit the science of DeepSeek for keeping them ahead. Specifically, their Engram technique, which I covered in a previous talk. The reason I credit Engram is that the other Flash-tier model ahead of MiMo, Qwen3.8-Flash-Next, also implements Engram. So when DeepSeek V4.1 Pro comes out, I expect it to surpass MiMo-V2.6 Pro for at least partly the same reason. (The other big reason is that later releases always have more data to train on and so should generally be better, without any other changes.)
Now we move into the technical background. And in fact, we will be looking at the most important scientific contribution Xiaomi has made thus far: multi-teacher on-policy distillation (MOPD).
That’s quite a bit of jargon, so let me break down the meaning and lineage of each piece.
The base word, “distillation”, means taking the wisdom of one model and transferring it to another. Historically, distillation was from a bigger model to a smaller model, which makes sense with the term; the idea is to make the intelligence of the big model as pure and direct as possible so it fits in the smaller model. The original paper on distillation, from way back in 2015, is by Geoffrey Hinton and Jeff Dean, who are both celebrities in the ML world.
On-policy distillation is a much more recent contribution by the folks at Thinking Machines. In their post they describe the problem with distillation, which is that it fails to capture the range of good responses for a given prompt. Like if you take a big model, give it a bunch of prompts, capture its responses, and then do SFT on the student model using those prompts and responses, then you risk teaching the student to mimic the big model rather than learn from it. That’s because you have presented a single, complete response as the right one, when really there are endless variations on what a good response can be in most cases; language is highly flexible, and even math problems can admit multiple different solution pathways.
On-policy distillation solves this problem by teaching the student one token at a time, on its own tokens. So the teacher and the student get the same input, then you compare the next-token probabilities from the teacher and the student, and you minimize the difference between the two distributions. The feedback is much more granular now, and also the student gets credit over a wider set of tokens. Like if we analogize to chess, it’s helpful for the student to have in mind several good next moves that the teacher can each grade, not just the one the teacher picked at that step, because a slightly different board could make the second-best move into the best move.
The graphic above shows on-policy distillation on the right actually. You can see they have a student model that generates a probability distribution at each position - Token 1, Token 2 etc. Then the teacher gets the same input and generates its own distribution at each token position. You compare the distributions using this metric called “reverse KL”, and you optimize the student model to minimize that reverse KL across all the tokens.
By now you have probably figured out the final part of MOPD: “multi-teacher”. Instead of one smarter model teaching a less capable model, you have several domain-expert teachers all distilling into the one student model. Which makes sense right? When you’re in school, different teachers have different expertise and you learn from each of them.
Notice though that we still have to train all our teachers. So why train a bunch of teachers who then train a student - all originating from the same SFT model by the way - when you can just train the one model directly?
Well there are a couple reasons. One reason is it’s easier to train on one domain than it is to train on several domains, either concurrently or in sequence. Training domain-expert teachers allows you to optimize each teacher for just its domain, so each teacher achieves better results than a jack-of-all-trades model would. Again, the human analog holds up well here; as a student, you want to learn from the best, and the best teachers have deep specialty.
The other, more operational reason is it’s faster. Notice the structure of Stage 2 in the diagram: the teachers all train in parallel! Assuming the same amount of training data and compute, it’s faster in wall time to train many teachers at once and then do distillation at the end versus training one model on all that data in serial.
Xiaomi first unveiled the technique in the MiMo V2 technical report, and since then it has gradually become a standard post-training technique, as we first discussed in the Nemotron 3 Ultra technical report where MOPD plays a major role. It’s not an incredible breakthrough, more a combining of existing techniques into a new package, but they did present and name the approach. In this technical report, we will see its successor.
The Paper
As always for new models, let’s start out with the architecture.
There’s not a ton to report here actually, since V2.6 is a minor release and generally inherits from V2.5. In fact, pretty much the entire focus of the paper is on post-training, which is right in Scale’s wheelhouse. The lead researcher at Xiaomi has also said they are saving their major innovation for V3, which I expect will come out before the end of the year.
Anyway, let’s tick through what they’ve shown us in their architecture diagram. On the left we note vision and audio inputs, broken up into frames. Then all the inputs go to this “Hybrid-SWA Backbone”, which uses a bog-standard pattern of several local attention layers for every global attention layer. There are no fancy attention tricks, nothing with the KV cache, etc.
After you predict your token x_t, you also speculate three additional tokens with the MTP block over on the right. Again, MTP is standard and expected these days, although now the more common setup is speculative decoding with a separate draft model like DFlash or DSpark. And actually, after pretraining they compared their MTP setup with DFlash and found that DFlash worked better, so they dropped MTP. But if you want MTP you have to have it in pretraining, so you might as well do it and run the cheap comparison against DFlash later. Apparently pretraining with MTP also helps the model predict x_t better too, which I didn’t know.
Here are the architectural parameters all laid out. I want to call out three quick things here.
First is that V2.6 basically is a new mid-train and post-train on top of V2.5, although strangely it’s unclear if they actually did pretrain from scratch or just reused the V2.5 base as you’d expect from this being a minor version number bump. There is little detail about mid-training, just that they add a lot of agentic data in, so this technical report is mostly a post-training report. More on that later.
Second is the MoE setup. They are in a pretty normal range for sparseness, i.e. the ratio of total experts to activated experts. The only somewhat unconventional choice is to have no shared experts.
Third is the speculative decoder, which I mentioned on the last slide. The parameters in here aren’t important, I just want to cite the change from MTP to DFlash.
There is literally a single sentence about SFT, so we’ll jump straight to RL, which is really the focus of this model and this paper.
Of course before you can do any RL, you need the data to RL on. They spend a significant amount of the paper talking about their data synthesis methods, they have a whole other paper about deriving tasks from repos, and they even released 7k environments that I’ll touch on a bit later.
Much of the focus is on code, including the graphic here. On the left they enumerate sources of tasks, which are mostly self-explanatory but come with a few fun nuggets.
For one, “everyday workflows” includes data from employees, both vibe coders and professional developers. Harvesting employee data for training is increasingly common, with a mention in the most recent DeepSeek technical report and a mini rebellion for similar stuff at Meta. Given the increasing importance and general rarity of data from real, highly skilled workers, I can only see data harvesting becoming more common.
For “long-horizon engineering”, they take existing tasks and make them increasingly complicated with scope and dependencies. I strongly suspect the result is a high share of contrived tasks, but they don’t give any examples and none of the tasks they released come with the labels given here.
And for data vendors, sadly they don’t name names.
The supervision half on the right is worth describing just to demonstrate how much the dropping cost of inference unlocks in training:
“Each task is attempted four times by a coding agent. An auditing agent receives all four rollouts in a shared workspace containing the problem statement, tests, reference patch, and each rollout’s submitted patch, test output, and full conversation log. The auditor first articulates what a correct solution requires and assesses whether the tests capture those requirements. It then evaluates each submitted solution against the specification using its patch and conversation log, and compares this assessment with the observed reward. A passing solution judged incorrect is flagged as a potential false positive, suggesting incomplete verification. A failing solution judged correct is flagged as a potential false negative, suggesting overly restrictive tests or execution failures. These disagreements identify tasks requiring further review and provide concrete evidence of possible mismatches between the intended behavior and the implemented verification.”
That’s a lot of agentic coding and inspection per task! They also use an agent to examine the trajectory for evidence of reward hacking. Intelligence begets intelligence.
Now on the general agent side, the focus is more on diversity instead of complexity and difficulty, since “general” is a wider category. Also, because of that diversity of outputs, verification takes more thought.
Start with the environments. Here you have a diversity of interfaces - MCP, API, CLI, GUI - as well as a diversity of tools and applications. Add in the fact that many applications allow multiple interfaces, and you can see how managing the diversity becomes a primary challenge.
General agent environments also have a wider array of artifacts in them, like PDFs and spreadsheets and CAD files. Here they collected real-world artifacts, either for direct inclusion or as inputs for artifact synthesis.
And as before, they threw tons of intelligence at the problem, with multi-agent pipelines for making mocks of popular software and creating tasks based on the environment, and multiple rollouts to test tasks and make rubrics from.
They don’t have dedicated graphics for anything outside of code and general, but they specify visual and cybersecurity agents as additional areas of focus. The one neat thing I wanted to mention here is their split between “open-ended design” and “high-fidelity visual replication”. The latter is mostly objective and thus easier to verify, so of course they’re going to train on it, but they do find rubrics with a vision-equipped LLM judge work for the more aesthetic matters of open-ended design. We will see a bit of that effect at the end of the talk.
You can actually go look at all the environments they made. If you’ve never looked at an environment before and find the whole thing a bit abstract, find a few where you’re comfortable with the subject matter and poke around. I recommend filtering to General, where the tasks and environments are still complex but the subject matter is often straightforward.
It’s nice of Xiaomi to release this data. I think it’s part of their strategy to gain adoption via goodwill and engagement with the research community, which isn’t as measurable and straightforward as benchmark scores or pricing, but can be just as effective.
Now in terms of what and how they’re rewarding in these environments, they have two complementary tactics for their code agent tasks: Groupwise Reward Synthesis (GRS) and Groupwise Advantage Redistribution (GAR).
On the left we have GRS, which I think will be familiar to anyone who has worked on rubrics before. The idea is to review many different rollouts, many different responses, to then have an agent abstract away into a rubric. The interesting bit here is the two sets of criteria in the rubric: solution and behavior. The solution criteria are about the end state, whereas the behavior criteria are about the way you got there. The examples of behavior criteria in the paper are “gathering relevant evidence and checking the effects of code changes,” but surely there are many more - including some negative criteria for reward hacking.
I wish they showed or discussed an ablation with the behavior criteria, because my general understanding is that intermediate rewards can hurt output quality because they add constraints. Maybe they did the ablation and just didn’t include it. Either way, they definitely kept the behavior criteria in, and even made the reward formula multiplicative, meaning a zero on behavior means zero reward total. Pretty strong endorsement.
So GRS is for tasks with a high pass rate, where distinguishing by behavior adds a layer of refinement. The other method, GAR, covers the rest and can handle more disparate outcomes between rollouts.
GAR starts with a group of rollouts and their rewards. After a sweep for reward hacks, and a zero reward for any rollouts found guilty, the remaining rollouts get a look from an agentic grader that evaluates aspects of quality not represented in the rewards - maintainability, readability, consistency with codebase conventions, stuff like that that isn’t just a matter of running unit tests.
The resulting ranking of solutions then redistributes the advantages, which are basically the rewards but on a curve. The advantages then go into your loss function so you can update your model weights.
Here’s one fun result of all this agentic grading going on: your costs go up! Between the two models, they spent about $450k just on agentic grader inference.
Of course the lion’s share of the cost is on training and inference of the model you’re developing, and related models like the teachers. I was surprised to see rollout costs match training costs, but given the rollouts can be hundreds of thousands of tokens long, it’s believable.
Here are the classic RL graphs showing scores and rollout length increasing with training. The flatline on Flash for tokens on AutomationBench is surprising, but everything else looks as you’d expect.
One other, more recent graph I’ve been seeing a lot is how different harnesses perform over the course of training.
Here, the Xiaomi folks took a stripped-down approach to harnesses, creating four barebones variants for use in training rather than taking popular harnesses like OpenCode or Codex off the shelf. They also tested performance on existing third-party harnesses over the course of training, as shown on the right.
There’s actually not much to see here; performance generally improves over the course of training regardless of harness. Apparently Codex isn’t a great fit for MiMo, but otherwise it seems harness use is a pretty generalizable skill.
Okay, now remember how we talked about MOPD earlier? Well here is MOPD2, which stands for “multi-prefix multi-teacher on-policy distillation”.
Actually it’s a bit confusing to call MOPD2 a single technique, because really it contains three: standard MOPD, Teacher-Prefix OPD, and SFT-Prefix OPD.
Let’s start with what they share, which is part A at the top. For verifiable tasks, they use all the normal RL techniques to create teachers. For “open-domain” tasks, like game development and scientific research, they use SFT instead. Either way, they make many teachers.
For the RL teachers, we can use them in the expected way, which we discussed in the background slides. However, we can also use them in this new way, Teacher-Prefix OPD. Here, we have a series of turns in a trajectory from the teacher, and we split that trajectory into different prefixes, each with a different number of turns. The student then generates the next turn in the trajectory, and the teacher grades it as normal.
For the SFT teachers, we only use them in the new way. So for SFT-Prefix OPD, we instead have some existing trajectory, not from the teacher model. We do the same splitting and next-turn completion, and then the SFT teacher grades.
As we know from before, rollouts are getting quite long and very expensive, in both compute and in time. The big advantage of MOPD2 is you can train at the level of a single turn, which makes your updates much more precise.
In fact, the MiMo team recently patched V2.6 using just this method! After the initial release of V2.6, some users reported repeated tool calls - calls with similar or identical arguments that did not contribute to completing the task. Classic reward-hacking smell.
And actually, the team had anticipated this form of reward hacking and added in a tool repetition penalty. But apparently the penalty wasn’t harsh enough.
Now they could have gone back to an earlier checkpoint and done more RL with the harsher penalty added in. However, given the length of rollouts and the relative rarity of the bad behavior to begin with, the authors estimate it would have cost over $2M to fix with standard methods.
That’s where MOPD2 came in. Specifically, they could take trajectories where the bad behavior occurred, take the prefix up to the turn with the bad behavior, and then have a teacher with the correct behavior supervise the model at that turn. All it requires is training a teacher with the right behavior and distilling it onto the student.
So that’s what they did. And it fixed the problem, for $90k instead of $2M.
Now you might be thinking, if it’s such a small problem, why didn’t they just do a little extra RL on top of the final checkpoint to fix its repetition problem? Well the post doesn’t explain why, but the likely reason is that training in one domain can degrade performance in other domains. When they made the new teacher to address the repetition problem, they turned back a few steps in their existing MOPD2 training pipeline and inserted the new teacher in with the rest of the teachers, so the last few steps now included anti-repetition in addition to all the other skills.
Here’s what we get for all that work. 2.6 is a big jump from 2.5, and 2.6 Flash trounces 2.5 Pro. For folks who joined us for DeepSeek-V4.1 Flash, it’s that story on steroids.
What’s interesting to me is what competition they picked: leading American models. DeepSeek isn’t on the original table, and neither are Kimi and GLM. The authors don’t explain their choice of competition, but to me it implies they are shooting for the absolute frontier and not just the Chinese frontier or the open-weights frontier.
More cynically, maybe they are trying to hide their losses to non-frontier models. For example, DeepSeek beats them on DeepSWE (74.2) and AutomationBench (54.8), GLM beats them on Terminal Bench 4.0 (41.8) and ExploitBench (54.4), and Kimi beats them on OSWorld-Verified (84.8).
I don’t mean to diminish the accomplishment here. I just want to call your attention more generally to framing and choice of benchmarks in launches. For all the scientific rigor around benchmarks and other capabilities measurement, release blog posts and even technical reports have some element of marketing.
It’s not cherrypicked though. Putting aside the internal benchmarks, the ones they picked are mostly legit and appear in many other announcements by rival labs. I would only note CyberGym is saturated and therefore not so telling, and that the OSWorld folks have released a successor in OSWorld 2.0 that is probably a better indicator.

One final place we do have results to look at is here, in a distillation experiment they did.
To show the efficacy and transferability of their training data, the authors take Qwen3.5 9B and train it in two stages: SFT, using prompt-response pairs distilled from their final model; and RL, using all the assets they released and that we discussed earlier.
Of course the training worked and produced much improved benchmark scores, but the interesting bit is in the qualitative results they show here.
The prompt asks for a website about the heritage of Ethiopia and provides some high-level design guidance like the color palette. There are no mocks or wireframes; the model has wide latitude.
As we can see, the starting model - which has gotten its own SFT and RL from the Qwen team by the way - is perfectly fine but not especially enticing. The tones of orange are a little too close, and the emojis at the bottom are amateurish, but it has many fundamentals right.
After some SFT from MiMo, things get more interesting. For one, the color choices make more sense. The text is also significantly more compelling, although I know you can’t read most of it. It’s a more creative and thoughtful interpretation of the prompt.
Finally, after all the RL, the best version yet emerges. The hero image is striking, the orange is a sparing highlight rather than a frontal assault, and the text has once again improved. It’s not perfect - one button has overflowing text, and they use the historically outdated name “Abesinia” with a weird alternate spelling - but it’s good.
I appreciate the authors for throwing in an eval that is more subjective but can still show clear progress. It’s something I would like to see more of in these technical reports.
My Takeaways
It’s relatively easy to catch up
Especially if you can distill, which American open models that lag the Chinese open models likely haven’t been doing
But where are all the European and Korean/Japanese Xiaomis?
Faster/cheaper inference means faster/cheaper training
Rollouts are enormous now
“Given enough eyeballs, all bugs are shallow”
The very first words of this paper are “Recursive self-improvement”, which depends on fast/cheap inference
Expect more Jalapeños, and more chip-design + kernel-authoring RL

















