We first present demo-conditioned learning for adapting to out-of-distribution objects. Rather than fine-tuning, the policy is conditioned on a single demonstration provided at test time. We show that reasoning about the demonstration and the current observation jointly in 3D outperforms compressing the demonstration into a latent embedding, and that a single human hand demonstration can replace a teleoperated robot trajectory, improving real-world performance on challenging unseen objects.
We then present an uncertainty-aware hierarchical framework for tasks where sub-goals cannot be deterministically defined. Common heuristics, such as gripper open/close transitions or near-zero end-effector velocity, provide no signal for non-prehensile pushing, sliding, or manipulating levers and handles without a discrete grasp event. The framework derives candidate sub-goals through probabilistic changepoint segmentation, represents the high-level goal distribution as a mixture model over candidate sub-goals, and conditions the low-level policy on this distribution through goal-aware attention.
Finally, this thesis extends hierarchical manipulation policies to challenging settings: adapting to unseen objects, and modeling sub-goal uncertainty in trajectories that cannot be deterministically segmented.
David Held (advisor)
Zackory Erickson (advisor)
Shubham Tulsiani
