Student Talks
Self-supervised tactile perception for robot dexterity
Abstract: Humans are incredibly dexterous. We interact with and manipulate tools effortlessly, leveraging touch without a second thought. Yet, replicating this level of dexterity in robots is a major challenge. While the robotics community, recognizing the importance of touch in fine manipulation, has developed a wide variety of tactile sensors, how best to leverage these [...]
Efficient Visual Modeling with Adaptive Representations
Abstract: While image understanding, generation, and manipulation have matured rapidly in recent years, video remains challenging due to the significantly larger input size. As a result, tasks such as generating long videos or understanding extended video sequences remain out of reach for current models due to their computational cost. This talk presents a series of [...]
Toward Scalable Architectures for Multimodal LLM-based Cooperative Autonomous Driving
Abstract: Despite the tremendous progress made in autonomous driving over the years, the safety of autonomous vehicles still requires further improvement before they can operate worldwide with full human trust. One principal safety concern is that each individual vehicle may have a limited field of view due to finite detection ranges, potential sensor failures, or occlusions [...]
Towards Manipulation in the Blind: Motion Planning for Manipulation under Uncertainty using Contacts
Abstract: Humans routinely rely on the sense of touch to better perceive the world. In environments characterized by poor lighting, occlusions, limited fields of view, or sparse visual features, contact feedback often becomes a primary source of information for perceiving the environment and successfully completing manipulation tasks. Everyday examples include locating a light switch in the [...]
Rethinking Robot Safety: Adaptive and Scalable Methods for Real-World Autonomy
Abstract: Safe autonomy in the real world requires more than safety in structured, low-dimensional settings. Robots deployed in everyday environments must cope with non-stationarity—objectives and dynamics that change due to human preferences or evolving operating conditions—and must also scale safety reasoning to high-dimensional robots and environments, where perception, dynamics, and safety constraints can be complex [...]
Structured Policies for Efficient Knowledge-Guided Learning from Humans
Abstract: Imitation learning has achieved strong performance in sequential decision-making tasks, but typically requires large numbers of expert demonstrations, has limited generalization capability in unseen scenarios, and is challenging for laypeople without technical backgrounds. This thesis introduces structured policies, a framework that integrates human domain knowledge into imitation learning by using large language models (LLMs) to generate semantically meaningful policy structures from [...]
GRAPPA: Generalizing and Adapting Robot Policies via Online Agentic Guidance
Abstract: Robot learning approaches such as behavior cloning and reinforcement learning have shown great promise in synthesizing robot skills from human demonstrations in specific environments. However, these approaches often struggle to generalize to unseen real-world settings because they rely on task-specific demonstrations or complex simulators. While foundation models (e.g., LLMs, VLMs) offer rich semantic understanding [...]
Dynamic Route Guidance in Vehicle Networks by Simulating Future Traffic Patterns
Abstract: Roadway congestion leads to wasted time and money and environmental damage. One possible solution is adding more roadway capacity, but this can be impractical especially in urban environments and still may not make up for a poorly-calibrated traffic signal schedule. As such, it is becoming increasingly important to use existing road networks more efficiently. [...]
Correspondence-Preserving Transformers for Scalable 3D Lifting
Abstract: Takeo Kanade's famous quip - to infer geometry or motion from images, you must first know what in one image corresponds to what in another, has guided geometric vision for three decades. Deep learning seemed to bypass this: methods in 2017-2019 lifted 2D to 3D using only reprojection loss, exploiting an implicit bias toward smooth [...]
Empirically Grounded LLM-based Virtual Patients for Psychotherapy Training: Design, Modeling, and Evaluation
Abstract: The need for mental health care continues to outpace the supply of trained psychotherapists, while psychotherapy training remains constrained by limited supervision time and scarce opportunities for repeated, feedback-rich practice in realistic scenarios. Simulation-based training can mitigate these constraints, but actor-based standardized patients are costly and difficult to scale, and many clinically challenging moments [...]