Abstract
Vision-language-action models have shown strong semantic understanding in tabletop settings, but household mobile manipulation is harder: a robot must reason jointly about scene layout, object geometry, its own configuration, and coordinated base, arm, torso, and gripper motion. Direct imitation learning supplies only a sparse action-prediction signal for this demanding 13-dimensional control problem.
SG-VLA addresses this gap by enriching both what the model observes and what it learns to reconstruct. The model combines head- and hand-camera RGB observations with depth, then co-trains auxiliary decoders for global robot position, grasp state, joint configuration, target-object pose, and target-object segmentation. These dense objectives encourage the shared visual-language representation to encode spatial and manipulation-relevant structure. A progressive training schedule first adapts randomly initialized decoders without letting their auxiliary gradients disrupt the backbone, then jointly refines the full model.