Accepted · ECCV 2026 Workshop X-Reason

Beyond Instance Slots:
Semantically Rich World Models for Physical Interaction Planning

A task-conditioned world model that turns visual entities into functional roles—then uses those roles to generate, evaluate, and repair robot actions.

Accepted at X-Reason: Visual Perception and Reasoning in the Interactable World, an ECCV Workshop.

Juntao Cheng1,2,* Jingkai Wang1 Yijun Shen1 Xiansheng Chen1 Zhiwei Yu1,†

1 Beijing Academy of Artificial Intelligence (BAAI)   ·   2 Shanghai Jiao Tong University

* Work done during the internship at BAAI. † Corresponding author.

Read on arXiv Watch demos Code Coming soon
Semantic state replay LIBERO Goal

Recorded simulator states · role contours visualize task binding

01 / Overview

World models should predict
what matters for the task.

Pixel or latent prediction can produce plausible futures without knowing whether the robot chose the right object, preserved a grasp, or completed the intended relation. SR-WM exposes those semantics directly.

SR-WM pipeline from observations through task-role binding and semantic rollout to decision making
Figure 1

One role-conditioned state connects observation history, action generation, semantic rollout, candidate selection, and suffix repair.

01

Bind entities to roles

Soft visual entities become task-specific gripper, target, and goal states instead of remaining exchangeable slots.

02

Predict semantic change

Action-conditioned dynamics forecast contact, grasp, task predicates, relation preservation, fixture state, and phase progression.

03

Select, execute, repair

A stage-aware selector ranks candidate chunks. When a violation appears, SR-WM keeps the valid prefix and resamples only the suffix.

02 / Simulation demos

From observation
to task completion.

Explore the paper's semantic replay alongside continuous human demonstrations from all four LIBERO benchmark suites.

Paper visualization

Inspect the semantic state.

Reconstructed from recorded simulator states for the paper figure bundle; these are concise state-sequence visualizations rather than continuous policy rollouts.

Task replay

LIBERO Goal · demo 0 · 92 states

Put the cream cheese in the bowl

Nine recorded states summarize the trajectory from initial observation through approach, grasp, transport, insertion, and task completion.

  1. 01Observe
  2. 02Approach
  3. 03Grasp
  4. 04Transport
  5. 05Inside
Role binding

Five semantic phases

Roles stay meaningful as geometry changes

Colored contours make the functional binding visible while the gripper, target, and goal move through the scene.

Gripper Target Goal
Side-by-side

Same state, two views

From pixels to a planning interface

The visualization isolates what changes when an observation is organized by task role. Contours are exact simulator instance masks used for this figure bundle; deployment does not require oracle masks.

Continuous benchmark demonstrations

Four suites, four task families.

Each card pairs an upstream LIBERO human demonstration with tracked gripper, target, and goal masks produced by a box-initialized SAM 2.1 video model.

SAM 2.1 role masks Gripper Target Goal Documented box/point prompts · video propagation
LIBERO-10

RGB + SAM 2.1 masks · Episode 290 · 15.0 s

Pick up the book and place it in the back compartment of the caddy

A compact single-object transfer with a clear grasp, transport, and placement into a specified storage compartment.

LIBERO Goal

RGB + SAM 2.1 masks · Episode 388 · 9.0 s

Turn on the stove

A goal-conditioned fixture task in which the gripper must localize and rotate the correct control.

LIBERO Object

RGB + SAM 2.1 masks · Episode 928 · 14.1 s

Pick up the milk and place it in the basket

An object-centric transfer task testing category grounding, stable grasping, transport, and containment.

LIBERO Spatial

RGB + SAM 2.1 masks · Episode 1397 · 9.5 s

Place the bowl between two objects onto the plate

A spatial-reference task that identifies the target through its relation to neighboring objects.

Interpretation. The RGB half is the upstream human demonstration; the mask half is a separate SAM 2.1 visualization. The masks are neither LIBERO ground truth nor SR-WM predictions and do not substantiate the paper's reported evaluation numbers.

LIBERO project LeRobot dataset SAM 2.1 model
Demo provenance and interpretation

Paper replay. LIBERO libero_goal, task “put the cream cheese in the bowl,” demonstration 0, agent-view camera, 640 × 640, 92 recorded states. Role contours use exact MuJoCo/robosuite instance rendering for traceability.

Benchmark gallery. Four complete episode ranges from lerobot/libero, pinned to revision a1aaacb. Each left panel preserves the upstream 10 fps agent-view demonstration.

Mask method. Right panels use facebook/sam2.1-hiera-small revision ee5bba1. All clips start from three verified boxes; selected clips additionally use documented correction boxes or points at contact and occlusion frames. They are SAM 2.1 outputs—not SAM 3, simulator ground truth, or SR-WM predictions.

03 / Results

Semantic structure closes
the planning loop.

Average rollout success across all four LIBERO suites. Candidate generation improves before selection; semantic ranking and targeted repair convert that coverage into executed success.

Average rollout success↑ better
83.0%

SR-WM semantic selector
with suffix repair

+37.5 points over LeWM Candidate 0
LeWMCandidate 0
45.5
SR-WMCandidate 0
62.8
SR-WMSemantic selector
75.3
SR-WM+ suffix repair
83.0
0255075100%
Rollout success (%) by LIBERO suite — reported in the paper’s main results table.
Decision ruleObjectSpatialGoalLIBERO-10Average
SR-WM Candidate 07467634762.8
Semantic selector8579756275.3
+ suffix repair9186837283.0
Oracle@4 analysis upper bound9692898189.5
62.8%Candidate 0 with role-conditioned generation, up from 45.5% with the monolithic LeWM latent.
78.2%Pairwise ranking accuracy, supporting semantic candidate scoring.
72.6%Patch-only rollout success without inference-time segmentation masks.

04 / Method

A shared semantic state
from perception to repair.

SR-WM is not “slots plus a language head.” Task-role binding defines the predictive interface used throughout generation, dynamics, selection, and repair.

Dynamic role-binding architecture using RGB, proprioception, language, and soft entities
1

Construct soft entities

A pretrained patch encoder produces exchangeable visual entity hypotheses. Segmentation proposals are optional attention priors rather than mandatory inference inputs.

2

Bind five functional roles

Language and robot state organize entities into gripper, target, goal, relation, and phase—capturing both task identity and progress.

3

Roll out candidate semantics

For each action chunk, the dynamics model predicts physical events and predicate changes. The selector prefers task-consistent futures.

4

Preserve valid work

When confidence drops, suffix repair retains the valid action prefix and regenerates from the first predicted violation.

05 / Abstract

World models for physical interaction are typically trained to predict future observations or latent features; however, a planning-oriented model must answer a fundamentally different question: whether a candidate action produces a task-consistent future while preserving essential relations.

We present the Semantically Rich World Model (SR-WM), a task-conditioned world model structured around five functional roles: gripper, target, goal, relation, and phase. A visual entity encoder extracts soft entity hypotheses; a role binder maps them to task-specific roles; and an action-conditioned dynamics model predicts task-critical semantic transitions.

This shared role state grounds multi-candidate action generation, stage-aware reranking, and violation-aware suffix resampling—turning object-centric prediction into a semantic interface for planning-oriented decisions.

06 / Citation

Build on this work.

Please cite the official arXiv preprint below.

@misc{cheng2026instanceslotssemanticallyrich,
  title         = {Beyond Instance Slots: Semantically Rich World Models
                   for Physical Interaction Planning},
  author        = {Juntao Cheng and Jingkai Wang and Yijun Shen and
                   Xiansheng Chen and Zhiwei Yu},
  year          = {2026},
  eprint        = {2608.22294},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2608.22294}
}