Plan-Conditioned Imitation for Robust Object Retrieval
under Self-Occlusion in Dense Clutter
Teacher Rollouts for Adaptive Closed-loop Execution
A robot’s arm blocks its own view. TRACE uses one initial plan, memory,
and partial feedback to retrieve a target object from dense clutter.
Overview
Supplementary video · 2:59Target retrieval in dense clutter under self-occlusion.
INSIDE TRACE
Training & inference
View full resolution
View full resolution
On the real robot, we estimate object poses from instance masks produced by Mask R-CNN (He et al., ICCV 2017).
ON THE REAL ROBOT
Demonstrations
Choose a method and scene.
All recordings play at their original speed.
Scene 01 · Original playback speed
90.0% success67.3 s total2.9× faster than Online Teacher, with no sensing retractions during pushing.
Scene 01 · Original playback speed
95.0% success192.7 s totalHighest success, but 2.9× TRACE’s total time and repeated sensing retractions.
Teacher Replay
Coming soon
47.5% success49.5 s totalFaster than TRACE, but 42.5 percentage points less success without execution feedback.
Scene 01 · Original playback speed
85.0% success141.1 s total5 percentage points less success and 2.1× TRACE’s total time, with complete-scene sensing.
Scene 01 · Original playback speed
77.5% success44.4 s totalFastest, but 17.5% out-of-workspace (OOW) failures versus 0% for TRACE.
Selected videos; statistics cover the full hardware benchmark: 20 scenes × 2 trials per method. Total time includes initialization/planning, execution, and grasp/lift.