Learning Object-Centric Spatial Reasoning for Sequential Manipulation in Cluttered Environments

Chrisantus Eze1 Ryan C. Julian2,* Christopher Crick1
1Oklahoma State University 2Google DeepMind
*Work done while at Google DeepMind
Under review

Unveiler separates which object to remove from how to manipulate it: a Spatial Relationship Encoder picks the next object, and an independent Action Decoder executes the removal.

Four frames of a Dofbot-Pro removing a yellow block and a green block to reach a covered blue target.
Unveiler on a real robot (Dofbot-Pro), retrieving a covered target (the blue block). The SRE selects the yellow block for removal, then the green block beneath it, and the robot then grasps the exposed blue target. The red block covers nothing and is never selected, so it stays on the table.

Overview

Abstract

Sequential manipulation in clutter requires selecting which objects to remove and determining how to manipulate them. Learning both decisions jointly can complicate training and make failures difficult to diagnose. We present Unveiler, a modular framework that separates scene reasoning and object selection from action execution through a discrete object-index interface. Its Spatial Relationship Encoder (SRE) selects the best object to grasp, while a separate Action Decoder generates orientation-discretized push-grasp actions. The SRE is trained in two stages: imitation learning from heuristic demonstrations, followed by search-based policy improvement using simulated object removals. This training procedure requires no real-robot training data and allows the SRE to be paired with a different execution policy. We evaluate Unveiler in PyBullet simulation and on a physical Dofbot-Pro. In real-robot experiments, it retrieves two occluded targets in the prescribed order in 90% of scenes, compared with 20–30% for the heuristic, GPT-4o, and the imitation-only variant. In PyBullet, Unveiler achieves the highest task-completion rate across all six clutter-density and occlusion conditions, reaching 90.0% under full occlusion with 6–9 objects, compared with 67.5% for the heuristic and 60.0% for ThinkGrasp. These results demonstrate the potential of modular, object-centric spatial reasoning for sequential manipulation in clutter.

90%real-robot success retrieving two occluded targets in order (baselines: 20–30%)
6 / 6simulated clutter and occlusion conditions with the highest task completion
0real-robot training data: the SRE is trained entirely in simulation

Method

The Spatial Relationship Encoder (SRE) reasons over object-centric features of the scene and outputs the index of the next object to remove. An independent Action Decoder turns that index into an orientation-discretized push-grasp action. Because the two modules communicate only through a discrete object index, the selection policy can be trained in simulation and deployed with a different manipulation system.

Unveiler system architecture with the Spatial Relationship Encoder and the Action Decoder.
Unveiler system architecture. The SRE and the Action Decoder are trained independently and connected by a discrete object index.

The SRE is trained in two stages: imitation learning from heuristic demonstrations, followed by search-based policy improvement using simulated object removals. On the real robot, the SRE is paired with the robot's own pick-and-place primitive.

Real-World Execution

A Dofbot-Pro retrieving covered blocks. Magenta outline: block selected for removal. Green outline: target being grasped. p is the SRE's probability for its selection. Footage is sped up 4–5×.

UnveilerTask 2: retrieve red, then green

Each target lies under its own block. Unveiler uncovers and grasps red, then uncovers and grasps green.

Task 2: retrieve green, then red

One blue block lies across both targets. Unveiler removes it first; GPT-4o reaches for the green target while the blue block still covers it.

Task 1: retrieve blue under two blocks

Unveiler removes yellow, then green, then grasps the target. GPT-4o removes red, which covers nothing, then selects the still-covered target.

Task 1: target already free

Nothing needs to be removed, and both Unveiler and GPT-4o grasp the target directly.

Results

Real robot

Real-robot results on the Dofbot-Pro, 10 scenes per task. Task 1: one target. Task 2: two targets, in order. Sel.: correct selections (%).
Task 1Task 2
MethodSuccess (%)Sel. (%)StepsSuccess (%)Sel. (%)Steps
GPT-4o3045.02.02055.62.7
Heuristic4070.61.72071.42.1
Imitation only4069.22.63075.02.8
Unveiler (Ours)7085.72.1901003.4
Two rows of frames comparing Unveiler and GPT-4o on the same scene.
Top: Unveiler removes the blue block that lies across both targets, then grasps green and red in order. Bottom: in this run GPT-4o reaches for the green target while the blue block still covers it, and the target is not retrieved.

Simulation

In PyBullet, we vary the clutter density (number of objects) and whether the target is partially or fully occluded.

Task completion (%, higher is better) across clutter densities and occlusion levels. “—”: not reported.
2–6 objects6–9 objects9–12 objects
MethodPartialFullPartialFullPartialFull
GPT-4o86.066.780.060.066.726.7
CLIP-Grounding80.560.367.266.753.020.0
VILG75.050.080.040.046.733.3
ThinkGrasp73.353.366.760.046.733.3
PPG40.650.37.54.44.4—
Heuristic87.367.567.547.850.020.0
Unveiler (Ours)96.189.397.690.092.653.8
Average planning steps (lower is better).
2–6 objects6–9 objects9–12 objects
MethodPartialFullPartialFullPartialFull
GPT-4o1.502.501.814.223.103.00
CLIP-Grounding1.672.113.104.204.384.67
VILG3.002.863.173.332.863.40
ThinkGrasp3.002.253.293.783.423.56
PPG4.424.276.003.004.00—
Heuristic2.352.302.553.433.403.17
Unveiler (Ours)1.321.871.433.312.863.71

Evaluation Scenes

Real-robot evaluation scenes.
Evaluation scenes. Top: the layout rendered in the digital twin, from the arm camera. Bottom: the same layout built on the table. In Task 1 the target is covered by one block (first column), by two (second), or free (third); in Task 2 each of two targets is covered (fourth).
Selection probabilities of the SRE on real frames.
Selection probabilities of the SRE on real frames from Task 1, for every segmented object. Green: the target's outline from the reference frame taken before it was covered. Magenta: the selected object.

Citation

@article{eze2026unveiler,
  title   = {Learning Object-Centric Spatial Reasoning for Sequential
             Manipulation in Cluttered Environments},
  author  = {Eze, Chrisantus and Julian, Ryan C. and Crick, Christopher},
  journal = {arXiv preprint arXiv:2603.02511},
  year    = {2026}
}