Grounded World Model:
Latent Planning with
Language Goals
What is a Grounded World Model?
A world model helps a robot predict what different actions might do before trying them. This lets the robot compare possible outcomes. Methods such as DINO-WM and LeWM specify that goal with a picture of the desired result (left). The robot predicts what each action would lead to and chooses actions that bring the scene closer to the example it was given. For a new task, getting that picture may mean setting up the desired result in advance. We want to describe the goal with an ordinary instruction instead (right), such as asking the robot to pick up an object used to serve food. With GWM, the planner compares predicted outcomes with what we asked for, letting us give new goals in words without a goal picture.
How to build a Grounded World Model
We start with a video-language embedding model that connects videos with language. Keeping it frozen, we train a predictor to imagine action outcomes in its latent representation. The robot can then compare possible futures with a language goal before acting.
A Video-Language Embedding Model as the Foundation
A video-language embedding model can find videos that match a description, or rank videos by how well they express the goal. It represents videos and text as vectors in a shared space, where similar directions indicate similar meaning. This allows us to describe an action, an object, or a desired outcome in words and compare it with what happens in a video.
Building GWM on the Frozen video-language embedding model
Learn from recorded futures
Given the current frame and rendered robot poses, the predictor learns to match the visual features of the recorded future using an MSE loss. The vision encoder stays frozen, and only the predictor learns. Training uses videos and actions, without task-language labels.
Planning with GWM in Latent Space
Keeping the current frame fixed, GWM predicts futures for many candidate action sequences. The same predicted futures can be compared with different language goals. Changing the task description changes each candidate’s cosine cost, reshaping the cost landscape and which action the planner prefers. No goal image or task-specific demonstrations are needed.
We use Qwen3-VL-Embedding as the frozen base model and train GWM on DROID and MolmoB0T data. No demonstrations are collected for any target task.
Watching language guide the search
We use pointing and pushing to characterize two properties of GWM-guided action generation: robust semantic grounding and fine-grained spatial precision. Pointing tests whether the search identifies the image named in an instruction. Pushing tests whether it selects trajectories that make precise contact with the named cube.
Point at the named image
Scoring camera view
Point at the dog
Point at the panda
Point at the banana
Point at the strawberry
The robot searches for a position above the photo named by the instruction. Each round samples possible endpoints, scores their predicted outcomes with GWM, and shifts the next round toward the best candidates. The selected endpoint lies inside the named photo’s cell in all four searches.
Push the named cube
Scoring camera view
Push the red cube
Push the green cube
Push the blue cube
Each endpoint defines a straight gripper slide. The search narrows toward the named cube, but reaching that region does not guarantee a push. Across 300 executed searches, 153 move the target by at least 3 cm, and none move an unnamed cube.
From Isaac Sim to the real world
For richer manipulation, we combine the GWM cost with language-blind trajectory proposals and geometric motion planning. Experiments in Isaac Sim and on a Franka robot test instructions that refer to objects, functions, and spatial relationships.
In Isaac Sim
The simulated tabletop setting includes 14 tasks with five trials each. GWM retains some ability to reason about spatial relations, visual attributes, and object affordances when deciding what the robot should pick up or where it should place an object.

14 tasks × 5 trials. System-level success.
* π0.5 uses a relaxed pick criterion: a lift above 1 cm at any point in the episode.
† V-JEPA 2-AC is given a goal image.
Task success
| Method | Success | Rate |
|---|---|---|
| GWM | 70/70 | 100% |
| GWM · opposite view | 65/70 | 93% |
| TiPToP | 66/70 | 94% |
| π0.5* | 37/70 | 53% |
| V-JEPA 2-AC† | 25/70 | 36% |
10 Picking Tasks
- pick up the fruit
- pick up the yellow object
- pick up the thing you could eat
- pick up the object that is neither a toy nor a container
- pick up the puzzle toy
- pick up the most colorful object
- pick up the object closest to the bowl that is not red
- pick up the object you would eat a meal from
- pick up the object between the cube and the banana
- pick up the round container
4 Placing Tasks
- put the block into the red box
- put the block into the green box
- put the block into the box that has the color of a tomato
- put the block into the box that has the color of grass
On a real robot
The same model guides a Franka in tabletop scenes with new objects, instructions, and camera viewpoints. It retains some ability to reason about spatial relations, visual attributes, and object affordances, using language to identify objects by their appearance, location, or potential use.
Successful GWM-system sub-tasks on hardware, compared with 52/60 for TiPToP. GWM inference runs on an RTX 5090 using up to 24 GB of CUDA memory, while TiPToP relies on a frontier Gemini model.