Bind entities to roles
Soft visual entities become task-specific gripper, target, and goal states instead of remaining exchangeable slots.
Accepted · ECCV 2026 Workshop X-Reason
A task-conditioned world model that turns visual entities into functional roles—then uses those roles to generate, evaluate, and repair robot actions.
Accepted at X-Reason: Visual Perception and Reasoning in the Interactable World, an ECCV Workshop.
Recorded simulator states · role contours visualize task binding
01 / Overview
Pixel or latent prediction can produce plausible futures without knowing whether the robot chose the right object, preserved a grasp, or completed the intended relation. SR-WM exposes those semantics directly.
One role-conditioned state connects observation history, action generation, semantic rollout, candidate selection, and suffix repair.
Soft visual entities become task-specific gripper, target, and goal states instead of remaining exchangeable slots.
Action-conditioned dynamics forecast contact, grasp, task predicates, relation preservation, fixture state, and phase progression.
A stage-aware selector ranks candidate chunks. When a violation appears, SR-WM keeps the valid prefix and resamples only the suffix.
02 / Simulation demos
Explore the paper's semantic replay alongside continuous human demonstrations from all four LIBERO benchmark suites.
Paper visualization
Reconstructed from recorded simulator states for the paper figure bundle; these are concise state-sequence visualizations rather than continuous policy rollouts.
LIBERO Goal · demo 0 · 92 states
Nine recorded states summarize the trajectory from initial observation through approach, grasp, transport, insertion, and task completion.
Five semantic phases
Colored contours make the functional binding visible while the gripper, target, and goal move through the scene.
Same state, two views
The visualization isolates what changes when an observation is organized by task role. Contours are exact simulator instance masks used for this figure bundle; deployment does not require oracle masks.
Continuous benchmark demonstrations
Each card pairs an upstream LIBERO human demonstration with tracked gripper, target, and goal masks produced by a box-initialized SAM 2.1 video model.
RGB + SAM 2.1 masks · Episode 290 · 15.0 s
A compact single-object transfer with a clear grasp, transport, and placement into a specified storage compartment.
RGB + SAM 2.1 masks · Episode 388 · 9.0 s
A goal-conditioned fixture task in which the gripper must localize and rotate the correct control.
RGB + SAM 2.1 masks · Episode 928 · 14.1 s
An object-centric transfer task testing category grounding, stable grasping, transport, and containment.
RGB + SAM 2.1 masks · Episode 1397 · 9.5 s
A spatial-reference task that identifies the target through its relation to neighboring objects.
Interpretation. The RGB half is the upstream human demonstration; the mask half is a separate SAM 2.1 visualization. The masks are neither LIBERO ground truth nor SR-WM predictions and do not substantiate the paper's reported evaluation numbers.
Paper replay. LIBERO libero_goal, task “put the cream cheese in the bowl,” demonstration 0, agent-view camera, 640 × 640, 92 recorded states. Role contours use exact MuJoCo/robosuite instance rendering for traceability.
Benchmark gallery. Four complete episode ranges from lerobot/libero, pinned to revision a1aaacb. Each left panel preserves the upstream 10 fps agent-view demonstration.
Mask method. Right panels use facebook/sam2.1-hiera-small revision ee5bba1. All clips start from three verified boxes; selected clips additionally use documented correction boxes or points at contact and occlusion frames. They are SAM 2.1 outputs—not SAM 3, simulator ground truth, or SR-WM predictions.
03 / Results
Average rollout success across all four LIBERO suites. Candidate generation improves before selection; semantic ranking and targeted repair convert that coverage into executed success.
SR-WM semantic selector
with suffix repair
| Decision rule | Object | Spatial | Goal | LIBERO-10 | Average |
|---|---|---|---|---|---|
| SR-WM Candidate 0 | 74 | 67 | 63 | 47 | 62.8 |
| Semantic selector | 85 | 79 | 75 | 62 | 75.3 |
| + suffix repair | 91 | 86 | 83 | 72 | 83.0 |
| Oracle@4 analysis upper bound | 96 | 92 | 89 | 81 | 89.5 |
04 / Method
SR-WM is not “slots plus a language head.” Task-role binding defines the predictive interface used throughout generation, dynamics, selection, and repair.
A pretrained patch encoder produces exchangeable visual entity hypotheses. Segmentation proposals are optional attention priors rather than mandatory inference inputs.
Language and robot state organize entities into gripper, target, goal, relation, and phase—capturing both task identity and progress.
For each action chunk, the dynamics model predicts physical events and predicate changes. The selector prefers task-consistent futures.
When confidence drops, suffix repair retains the valid action prefix and regenerates from the first predicted violation.
05 / Abstract
World models for physical interaction are typically trained to predict future observations or latent features; however, a planning-oriented model must answer a fundamentally different question: whether a candidate action produces a task-consistent future while preserving essential relations.
We present the Semantically Rich World Model (SR-WM), a task-conditioned world model structured around five functional roles: gripper, target, goal, relation, and phase. A visual entity encoder extracts soft entity hypotheses; a role binder maps them to task-specific roles; and an action-conditioned dynamics model predicts task-critical semantic transitions.
This shared role state grounds multi-candidate action generation, stage-aware reranking, and violation-aware suffix resampling—turning object-centric prediction into a semantic interface for planning-oriented decisions.
06 / Citation
Please cite the official arXiv preprint below.
@misc{cheng2026instanceslotssemanticallyrich,
title = {Beyond Instance Slots: Semantically Rich World Models
for Physical Interaction Planning},
author = {Juntao Cheng and Jingkai Wang and Yijun Shen and
Xiansheng Chen and Zhiwei Yu},
year = {2026},
eprint = {2608.22294},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2608.22294}
}