Originally presented as a live talk on November 12, 2025
🌍❓
I think we have to start with, what are world models?
And right at the top I want to say they are not necessarily models of The World, like the planet or the universe as we know it, although there are models that attempt something like this, where they can take a realistic scene and extrapolate forward, like what’s down the street in a city or what’s on the other side of a mountain. The Genie series from DeepMind is a world simulator like this.
That’s impressive and all but it’s kind of on a different axis than what our paper talks about. The world model concept here is more abstract, and the worlds they choose to build are very simple.
The world model concept is about whether a model can form a coherent and predictive view of a given environment based on limited information.
That may sound like a requirement for making good predictions, which we know models can do, but it’s actually not. Let me give you an example.
Let’s say you’re flipping a coin and you don’t already know that coin flips are inherently 50-50. One way you could discover they’re 50-50 is to flip the coin a bunch of times and extrapolate from that data to predict future data. You don’t need to understand the nature of coins or the laws of physics to do that extrapolation, you just need basic data analysis skills.
Of course the way a human would conclude coin flips are 50-50 is to look at the coin, maybe toss it a few times to make sure it’s not weighted, and then reason a bit to predict the coin toss will be 50-50. That does require a world model, because it requires underlying assumptions about how objects behave in the physical world.
The point of this world model line of inquiry is that LLMs do not have world models by design - they are just trained to predict the next token, much like the extrapolation from coin toss data we mentioned earlier.
So can they acquire world models emergently in this next-token training? Or maybe with certain data? Or does world model understanding require entirely new approaches to artificial intelligence?
One early but still quite relevant attempt at world model evaluation is the Abstract and Reasoning Corpus for Artificial General Intelligence, better known as ARC-AGI.
It’s from 2019, created by Francois Chollet, a longtime Google engineer who also created an influential ML tool called Keras.
It asks models to learn and then apply a pattern, as shown here.
Now for humans this is pretty easy, sometimes laughably so. Like the first example seems incredibly obvious, just complete the 2x2 square by adding a 1x1 square in the missing corner. But actually this can be pretty hard for models as we’ll see.
One note about the modality though: the default input is textual, specifically a 2D array of color names. If you don’t know what that means, just imagine each square here is written down as the name of its color and then put in a specially formatted list. Of course you can turn that into an image and then give that image to your model instead, or you can provide both, whatever works best. But of course not all models accept images, so those ones are going to be working with the textual representation only.
Regardless, in either case the output has to be textual.
This is what the ARC-AGI leaderboard looks like. They have a traditional ranked list version, but because throwing more thinking tokens at the problem does tend to improve outcomes, they have this chart of performance vs cost. So you can see all the way out on the right, it cost o3-preview hundreds of dollars per task on average. But it did also set a performance record for a model. The two orange dots at the top with papers rather than model names are custom systems built on top of models. Whereas the model names are just from feeding the textual and possibly image input straight to the model and not providing any special tools etc.
Even with all that token budget and all those tools though, the original ARC-AGI still hasn’t saturated! It’s perhaps the oldest ML benchmark fitting that bill.
But as the name at the top implies, there is a second one as we’ll see. And even a third one under development.
So here’s an example from ARC-AGI-2. Again, humans get these right virtually 100% of the time. In this example the rule is that blue-bordered squares move as far left as they can, and red-bordered squares move as far right as they can. Squares aren’t allowed to overlap so sometimes they stack. There are other differently themed puzzles but all of them still rely on visual reasoning and forming some hypothesis about the rules of the environment.
Here’s the leaderboard. The shape is similar but look at the y axis - the max score isn’t even at 30%. o3-preview spends hundreds of dollars per task just to fail on almost every one of them!
One model I do want to shout out here is Hierarchical Reasoning Model, and its successor Tiny Recursion Model - it’s the two orange triangles just past the $1 mark and hitting around 5%. These are small small models, like tens of millions of parameters rather than the tens or hundreds of billions of parameters on leading models. But they’re specifically designed for non-verbal reasoning and solving games like this. So they’re not the generalist workhorses that modern LLMs are, but they do show how much of a brute force approach LLMs are and how patchy their intelligence can be.
Anyway, Chollet left Google a year ago to found Ndea (pronounced “Endia”). Here’s how they describe program synthesis, their main research topic: “Instead of interpolating between data points in a continuous embedding space, program synthesis searches for discrete programs, or models, that perfectly explain observed data. This allows it to achieve much greater generalization power with extreme data-efficiency, requiring only a few examples to learn.”
So they want to combine the empirical, extrapolation-based intelligence of current LLMs with the principled, model-based intelligence of program synthesis.
The website claims all the big labs are working on this, though it’s not obvious there’s broad demand yet. There has however been a recent uptick on Google Trends for the term “world model”, and Meta recently produced a Code World Model. This isn’t a brand new line of research but it’s getting attention.
Speaking of Code World Model, let’s take a look at some of their data to understand how a coding model differs from a coding world model.
So normal code data for us would be a prompt and a response, if it’s SFT, or a prompt and a set of unit tests that verify whether the model’s answer works, if it’s RLVR. But in either case we’re training the model directly for writing code.
The coding world model has a different goal: to predict how a piece of code will work. If the model can do that, then it probably has a good world model for the world of code, its “laws of physics” so to speak.
The way they train that in is to provide a piece of code and an example of using that code, then ask the model to predict the state and action at each step of the program.
Here the state is in yellow and the action is in blue. As you go down the rows, or “frames” as the paper calls them, you can see what the model is keeping track of and how it changes after each action. Like for example in the second frame, we get a new variable n, with the starting value 0. Then in the third frame we see n being tracked in state, along with its current value, 0. The two dots in quotes is just a visual shorthand that means the value of the variable hasn’t changed, so like s is “strawberry” for the first frame and then just the two dots for the rest of the frames.
The paper we’re covering is a benchmark, so we won’t see any other training data, but hopefully this little snippet helps you compare against training data you’ve seen here at Scale that isn’t focused on world modeling.
Examples
The paper opens with four qualities a world model eval should have:
Interactive - the model should have the freedom or responsibility to gather information at its own direction
Behavior-based - the eval should only require looking at final predictions, not intermediate work or representations like a description of the AI’s world model
Goal-free - the model should not have an idea of what the test will be when it is forming its world model, so that it doesn’t reduce its model from the whole world to just the task-relevant stuff
Transferrable - the test environment should be related to, but different from, the learning environment; if the environment doesn’t change, you can’t test whether the AI developed a shared underlying world model versus just extrapolating from prior data in the learning environment
They basically take ARC-AGI and adapt it to fit their four qualities.
I’ll explain more in a minute, but it will make things much more concrete if I show you one of their examples first.
[Pick Masked Frame Prediction - P-7WWW9]
So here is one simple example. You have this little grid world, some shapes in it, and some basic inputs: the four arrow keys, plus clicking anywhere in the grid.
There are two stages to the task. In the first stage you just play around, get to know the environment. And you’ll see that when I provide input there’s a frame where the input happens, along with any resulting output, then frames continue to pass. So there’s a concept of time here, or at least sequence.
I also get to reset the environment if I want, so I can play around more with other choices. And then when I’m done playing I can take the test.
The type of test I picked here is Masked Frame Prediction, basically saying what a certain part of the world will look like at a certain time. So if I scroll through the frames I see eventually they apply a mask to part of it, and now I have to say what’s behind the mask.
Now let’s check how this matches up to our four qualities:
Interactive - yes, it’s interactive
Behavior-based - yes, all I had to do was pick the right option, I didn’t have to diagram or write out my understanding of the environment
Goal-free - yes, in the first part I was just playing around
Transferrable - yes, the rules I inferred transferred to the test environment, even though the test environment had the added factor of the mask
The big miss for ARC-AGI on that list is Interactive, although it also doesn’t have a goal-free setting and so it misses on 3 and 4 is kinda not applicable. Ultimately it all stems from not being interactive though, which apparently they are changing in ARC-AGI-3. So maybe this is a little preview of that.
Here’s a diagram of their eval, which they call WorldTest. Well technically WorldTest is a method or framework, and their particular instance is called AutumnBench, which I find a somewhat needless distinction. So I’m gonna just call it WorldTest.
Anyway, WorldTest has three different test tasks, one of which is the Masked Frame Prediction we saw in the example. The others are Planning, which sets a goal for the model to achieve, kind of like completing a maze; and Change Detection, where one rule of the environment changes and the model has to report the first time step where that rule change becomes visible.
The environments vary somewhat: size anywhere from 3x3 to 25x25, 1-12 different colors, 1-5 or so different object types. 43 different environments in total, each with the three task types, so 129 different tasks in total.
For human testing there’s a GUI, but for model testing they do only textual representations, not graphical.
So they tested human along with three different models: o3, Claude 4 Sonnet, and Gemini 2.5 Pro.
For the humans they recruited about 500 people from Prolific and had them pass an initial screener. Individual quality was pretty poor I suspect, but they compensated for it by taking the 80th percentile performance on each problem across 20 different human attempts. That 80th percentile was alway good if not perfect, so the contributor ceiling did not need to be that high. And that’s what we’d expect given the human results on ARC-AGI too.
As for the models, they did pretty poorly across the board. On the left is a breakdown by task type - CD for Change Detection, MFP for Masked Frame Prediction, PL for Planning - which shows relatively little variation. On the right is a plot of the distributions across all task types, where you can see the models sometimes doing well but mostly doing poorly; there’s a long tail of occasionally good or even great performance.
It is somewhat unfair that each model only got one attempt but the human score is based on an average, but I don’t think that explains most of the gap.
One interesting bit is the differences in human and model behavior. There are only four types of action you can take in these worlds: clicks, arrows (in one of four directions), resetting to the initial state, or nothing, also called “no-op” which is short for “no operation performed”.
There are two big differences here. One is how much more clicking the models did, which is weird but not that interesting. The other is the bias for action and pressing ahead, compared to the humans who reset or did nothing much much more.
That lack of epistemic humility, that overconfidence in knowledge is a big problem for modern AIs. It’s the same reason why models halllucinate, at least according to a recent OpenAI research paper; we train them to always provide an answer, to always know or at least think they know. If you ask a question the model doesn’t know the answer to, there’s a good chance it will answer anyway, because we don’t train on question where the right answer is “I don’t know” or “it might be this but it might be that”.
I think the same thing is happening here. That bias for action, that reluctance to explore more or start over again and implicitly admit you don’t know, that’s the same overconfidence showing. And I do think it’s down to how we do training.
The authors note a couple other model behaviors they saw in the chains of thought.
One is hypothesis formation and testing, but with somewhat of a blind eye, as some informative actions or results go unused.
The other is hesitancy to make updates. On the previous slide we saw Change Detection had the lowest score by task type; the models sometimes detected the change but stuck to the earlier version of the rule.
Again, all part of the same root issue in my view.
My Takeaways
Modern LLMs don’t have world models
Coding and tool use are prerequisites for LLM-guided program synthesis
As are more diffuse qualities like reasoning and creativity
World-model testing environments are an interesting artifact, though I’m not convinced you need a human touch to create them
Training data for LLM-guided program synthesis could be on the horizon right now for some research groups












