BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260701T143000
DTEND;TZID=America/New_York:20260701T153000
DTSTAMP:20260911T204247
CREATED:20260625T193130Z
LAST-MODIFIED:20260625T193130Z
UID:151664-1782916200-1782919800@www.ri.cmu.edu
SUMMARY:Learning From History: Test-Time Verification and Adaptation for Robotics
DESCRIPTION:Abstract: The physical properties and dynamics that decide how an object or environment responds to a robot’s actions are often impossible to determine from visual observation alone. An object’s mass distribution and friction\, the kinematics of an articulated object: these latent factors dictate the correct action\, yet they leave little or no trace in a single image. This partial observability makes a purely feed-forward visual policy fundamentally limited\, and even a strong policy will inevitably make mistakes when deployed\, whether from this latent ambiguity or imperfect perception. \n\nThis thesis argues that both the missing information and the means to recover from error are supplied by the robot’s own history of interaction: by acting\, observing the outcome\, and reasoning about the mismatch between what was expected and what occurred\, an agent can infer the underlying dynamics online and adapt on the fly. \nWe develop this idea along two complementary axes. Through verification\, we introduce HAVE\, a History-Aware VErifier that scores action proposals from a generative policy by reasoning about past actions and their outcomes. We prove that any better-than-chance verifier improves expected reward over the generator alone\, and validate it across articulated objects\, multi-modal doors\, and uneven-mass objects in simulation and the real world. Through test-time training\, we introduce SCOUT\, a dynamics-aware meta-learning framework that couples a policy with a forward dynamics model through a shared belief latent; at deployment an inner loop revises this belief by minimizing the error between expected and observed outcomes\, while a meta-learned outer loop ensures the belief steers the policy correctly. SCOUT adapts substantially faster than history-conditioned baselines and directly sim-to-real transfers to real-world tasks. \nTogether\, these methods show that treating interaction history as a rich\, structured supervision signal\, rather than a sparse scalar reward\, yields manipulation policies that adapt quickly and robustly to the hidden physical properties and dynamics of objects and environments\, recovering from their own mistakes along the way. \nCommittee:\nProf. David Held (advisor)\nProf. Andrea Bajcsy\nMihir Prabhudesai
URL:https://www.ri.cmu.edu/event/learning-from-history-test-time-verification-and-adaptation-for-robotics/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260706T100000
DTEND;TZID=America/New_York:20260706T110000
DTSTAMP:20260911T204247
CREATED:20260630T134012Z
LAST-MODIFIED:20260630T134012Z
UID:151684-1783332000-1783335600@www.ri.cmu.edu
SUMMARY:View Generalizable Manipulation Policies via Sim-to-Real Transfer
DESCRIPTION:Abstract: Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. A key requirement for these manipulation policies is to exhibit robust generalization capabilities when deployed in the real world\, where the objects\, scenes\, and sensors a robot encounters differ from those seen during training. In practice\, learned policies often remain brittle to these changes\, which limits their usefulness beyond the narrow conditions in which they were trained.\nIn this thesis\, we study manipulation policies that remain performant under camera viewpoint shifts\, so that a single policy can be deployed across various camera poses in the real world. We approach this by grounding the policy in the robot frame in order to reason about the scene and the robot’s actions in a shared frame rather than relative to a particular camera. In the first part\, we present ArticuBot\, in which a single learned policy enables a robotics system to open diverse categories of unseen articulated objects in the real world. The policy operates on point clouds in the robot frame\, and we find it remains robust under camera viewpoint changes\, including camera poses not seen during training\, while generalizing across objects that vary widely in geometry\, size\, and articulation. By generating a large number of demonstrations in physics-based simulation and distilling the demonstrations into a hierarchical\, point cloud-based neural policy via imitation learning\, we demonstrate an effective policy learning approach that also achieves object-level generalization. \nIn the second part\, we bring this robot-frame reasoning to image-based policies\, which benefit from large-scale pretraining and scalability that point cloud policies do not. We present VGP\, an image-based policy that encodes the scene with geometry-aware visual features and grounds its visual\, proprioception\, and action tokens in a shared robot frame\, allowing it to remain robust across a wide range of camera poses\, outperforming 2D and 3D baselines\, while matching fixed-camera baselines. As a practical consequence\, our policy transfers zero-shot from simulation to the real world under random camera configurations. \nAcross these two parts\, we show how large-scale simulation and imitation learning\, together with grounding the policy in the robot frame\, can be used to train manipulation policies that remain robust as the camera viewpoint changes and transfer to the real world. \n\nCommittee: \nProf. David Held (co-chair)\nProf. Zackory Erickson (co-chair)\nProf. Shubham Tulsiani\nKrishna Suresh
URL:https://www.ri.cmu.edu/event/view-generalizable-manipulation-policies-via-sim-to-real-transfer/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260709T130000
DTEND;TZID=America/New_York:20260709T140000
DTSTAMP:20260911T204247
CREATED:20260706T145421Z
LAST-MODIFIED:20260706T145421Z
UID:151715-1783602000-1783605600@www.ri.cmu.edu
SUMMARY:Exploring High-Level Goal Prediction for Hierarchical Imitation Learning in Robotic Manipulation
DESCRIPTION:Abstract:\nHierarchical imitation learning has become an effective approach for robotic manipulation: a high-level policy predicts a sub-goal end-effector pose\, while a low-level policy executes the actions needed to reach it. This decomposition improves generalization and provides an interpretable interface\, but the design of the high-level goal predictor remains an open question. \n\nThis thesis studies high-level goal prediction through three investigations. First\, on the MimicGen benchmark\, we compare dense predicted goal point clouds with a sparse four-point end-effector representation\, finding that the sparse representation provides a more reliable conditioning signal across tasks. Second\, on RLBench\, we extend the four-point predictor with language\, RGB features\, gripper actions\, and collision-ignore decisions\, pairing it with motion planning to achieve competitive performance with state-of-the-art keyframe-action prediction methods. Third\, in a sim-to-real setting\, we make the high-level predictor steerable through prompting\, allowing users to select among multiple valid targets\, such as which drawer to open. \n\nTogether\, these studies clarify how representation\, capability\, and controllability shape high-level goal prediction for hierarchical imitation learning in robotic manipulation. \n\nThesis Committee:\nProf. David Held (co-advisor)\nProf. Zackory Erickson (co-advisor)\nProf. Shubham Tulsiani\nYilin Wu
URL:https://www.ri.cmu.edu/event/exploring-high-level-goal-prediction-for-hierarchical-imitation-learning-in-robotic-manipulation/
LOCATION:Gates 6115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260709T140000
DTEND;TZID=America/New_York:20260709T150000
DTSTAMP:20260911T204247
CREATED:20260706T191935Z
LAST-MODIFIED:20260706T191935Z
UID:151734-1783605600-1783609200@www.ri.cmu.edu
SUMMARY:Beyond Vision-Language-Action Models: Adapting\, Steering\, and Accelerating Generalist Robot Policies
DESCRIPTION:Abstract: Generalist robot policies\, vision-language-action models that combine a large pretrained vision-language model backbone with a diffusion or flow-matching action head\, are increasingly capable\, yet hard to deploy in the real world. Three gaps separate such a policy from a deployable one: a data gap (adapting to a new task still demands task-specific teleoperation data)\, an inference gap (the policy samples its action distribution with no control over how conservative or diverse the result is)\, and an architecture gap (a large policy is too slow to replan often\, so it must run its action chunks open-loop and cannot react mid-motion). This thesis argues that these gaps can be closed not by training larger models on more robot data\, but by changing how a pretrained policy generates its actions\, with little or no additional training. \nWe develop three methods. DemoDiffusion (data) imitates a single human demonstration instead of collecting teleoperation data: it retargets the human hand motion into a coarse robot trajectory\, then uses a frozen generalist diffusion policy to refine it into plausible robot actions. It needs no task-specific or paired human-robot data\, and succeeds even where the base policy fails outright. Temporal Score Rescaling (inference) rescales the learned score/flow at inference to draw from a sharper or broader distribution than the model was trained on. It is training-free\, works with any diffusion or flow model\, and improves image generation\, depth\, pose\, and protein design alongside robot policies. πR² (architecture\, inference) builds on diffusion forcing to split conditioning into a fast proprioceptive channel and a slow\, asynchronously updated vision-language channel\, so the policy reacts to fresh proprioception while tolerating stale vision. A latency-adaptive schedule lets a single model handle varying inference latency and emit actions in a single denoising step\, making a large policy reactive and real-time\, several times faster than the original. \nTogether\, these methods adapt\, steer\, and accelerate a pretrained policy\, taking vision-language-action models beyond what they can do as trained and toward real-world deployment. \nCommittee: \nProf. Shubham Tulsiani (advisor)\nProf. Katerina Fragkiadaki\nAndrew Wang
URL:https://www.ri.cmu.edu/event/beyond-vision-language-action-models-adapting-steering-and-accelerating-generalist-robot-policies/
LOCATION:GHC 9115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260710T110000
DTEND;TZID=America/New_York:20260710T120000
DTSTAMP:20260911T204247
CREATED:20260706T185409Z
LAST-MODIFIED:20260706T185409Z
UID:151732-1783681200-1783684800@www.ri.cmu.edu
SUMMARY:Towards Scalable Robot Learning: From Teleoperation to Web-scale Data
DESCRIPTION:Abstract:\nHumanoid robots operating in human environments must manipulate articulated objects under contact and kinematic constraints that human demonstrations do not satisfy. That mismatch makes the human–humanoid embodiment gap the central bottleneck for learning from human data: robot demonstrations are expensive and sparse\, while human demonstrations inhabit a different state-action space and often violate robot kinematic constraints. This thesis studies how to convert human behavior into supervision that remains executable for the target robot body. \nThe first part develops Humanoid Policy ~ Human Policy for cross-embodiment supervision in humanoid manipulation. It places humans and humanoids in a unified state-action representation\, enabling a transformer policy to co-train on human and robot demonstrations and retarget its predictions at deployment. To support this formulation\, we introduce PhD^2\, a task-oriented egocentric human demonstration dataset that expands data scale without discarding embodiment structure.\nThe second part presents EmbodyHOI\, which addresses a harder embodiment-gap setting in dexterous hand-object interaction. It starts from a flow-matching diffusion model trained in human hand-object space\, then applies a differentiable guidance function during sampling to steer trajectories toward a target humanoid embodiment\, jointly optimizing wrist reachability and base placement before downstream control. \nTogether\, these chapters show that scalable robot manipulation requires data transformations that preserve task structure while respecting the robot body. \n\n\nThesis Committee:\nProf. Guanya Shi (advisor)\nProf. Laszlo A. Jeni\nEliot Xing
URL:https://www.ri.cmu.edu/event/towards-scalable-robot-learning-from-teleoperation-to-web-scale-data/
LOCATION:Gates Hillman Center 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260713T133000
DTEND;TZID=America/New_York:20260713T143000
DTSTAMP:20260911T204247
CREATED:20260710T150347Z
LAST-MODIFIED:20260710T150347Z
UID:151756-1783949400-1783953000@www.ri.cmu.edu
SUMMARY:Toward Curiosity-Driven Embodied Learning Through World Models
DESCRIPTION:Abstract: Curiosity allows animals and humans to learn through interaction without explicit instruction. This thesis asks how principles of natural curiosity can be translated into embodied agents\, and what kinds of world models are needed as bodies\, action spaces\, and environments become more complex. We study this progression in simulation\, using animal behavior and neural dynamics as both inspiration and evaluation targets.\nWe first study futility-induced passivity in larval zebrafish. We introduce 3M-Progress\, a model-memory objective that compares an online world model with a frozen memory of normal action consequences\, and train the agent using intrinsic motivation alone. When swim commands no longer move the visual world\, the agent reproduces the active–passive cycling observed in zebrafish. Its latent dynamics also recapitulate the temporal structure of neural and glial activity in the biological circuit for passivity. \nWe then extend this perspective to walking Drosophila\, whose articulated body and richer action repertoire provide a more complex embodied setting. A simulated fly on a spherical treadmill allows us to compare behavior and internal activity with biological experiments showing neural asymmetry before spontaneous turns. We evaluate whether the agent produces fly-like walking and turning behavior and whether its recurrent activity contains information about future turn direction before movement begins. \nFinally\, we ask whether this approach can extend to larger\, more open-ended embodied environments that more closely resemble the settings in which human curiosity operates. Preliminary results reveal important limitations\, suggesting that the underlying world-model agents must first learn reliable behavior in these environments before curiosity can be meaningfully evaluated. \nTogether\, these projects treat world models as the substrate through which curiosity is expressed and assessed. They show that purpose-built models can reproduce biologically observed behavior and internal dynamics\, while extending curiosity to richer settings depends on whether the underlying world-model agent can first learn the environment reliably. \n\nThesis Committee:\nProf. Aran Nayebi (Chair)\nProf. Yonatan Bisk\nReece Keller
URL:https://www.ri.cmu.edu/event/toward-curiosity-driven-embodied-learning-through-world-models/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260713T140000
DTEND;TZID=America/New_York:20260713T153000
DTSTAMP:20260911T204247
CREATED:20260706T144344Z
LAST-MODIFIED:20260706T144344Z
UID:151711-1783951200-1783956600@www.ri.cmu.edu
SUMMARY:Consistent Modeling of 4D Scenes for Perception and Generation
DESCRIPTION:Abstract:\nA core challenge in vision is building representations that capture 3D scenes over time for both perception and generation. This thesis studies consistency across views\, time\, and modalities by moving from dense grid-based representations toward entity-centric scene representations that can be maintained across frames and used for interactive generation. \nThe first part of the thesis develops consistent 3D perception systems\, including methods for temporal multi-camera 3D detection\, depth completion\, semi-supervised detection\, and streaming semantic occupancy estimation. These works progressively move from dense spatial representations toward persistent\, query-based representations that model both foreground objects and background structure over time. \nThe second part of the thesis presents LatentWorld\, a generative model for entity-centric 4D scene generation. LatentWorld represents a scene as a sparse set of grounded 3D latents\, assigning persistent latents to foreground actors while using multiple latents for background regions. Generation is factorized into layout\, geometry\, and motion\, enabling temporally coherent semantic-occupancy rollouts with stable actor identity\, explicit ego and actor motion\, and direct scene-level control. \nTogether\, these works show how consistent\, entity-centric representations can serve as a common foundation for understanding dynamic 3D scenes and generating plausible\, controllable 4D worlds. \nThesis Committee Members:\nKris Kitani (Chair)\nDeva Ramanan\nShubham Tulsiani\nWei-Chiu Ma (Cornell University)
URL:https://www.ri.cmu.edu/event/consistent-modeling-of-4d-scenes-for-perception-and-generation/
LOCATION:GHC 4405
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260714T141500
DTEND;TZID=America/New_York:20260714T151500
DTSTAMP:20260911T204247
CREATED:20260708T191940Z
LAST-MODIFIED:20260708T191940Z
UID:151748-1784038500-1784042100@www.ri.cmu.edu
SUMMARY:Data Mining and Auto-Labeling for Promptable Driving Policies
DESCRIPTION:Abstract:\nAutonomous vehicles (AVs) are being deployed at scale today\, with companies like Waymo achieving upward of 500\,000 passenger rides per week. Two of the largest remaining problems in the field are 1) building a system that generalizes across the long-tail of edge cases that are represented few or no times within the training data and 2) validating the performance of the system in these rare scenarios prior to new deployments. \nThe first half of the thesis discusses RefAV\, a benchmark for retrieving text-specified scenarios of interest from uncurated driving logs. AVs collect terabytes of observational data during normal fleet testing\, with a large majority of it boring. Traditional scenario mining techniques are error-prone and prohibitively time-consuming\, often relying on hand-crafted structured queries. We revisit spatio-temporal scenario mining through the lens of recent vision-language models (VLMs) to detect whether a described scenario occurs in a driving log and\, if so\, precisely localize it in both time and space. We introduce a large-scale dataset of 10\,000 diverse natural language queries that describe complex multi-agent interactions relevant to motion planning. We evaluate several referential multi-object trackers and present an empirical analysis of our baselines. Notably\, we find that naively repurposing existing VLMs yields poor performance\, suggesting that scenario mining presents unique challenges. We discuss our recently held competition and share insights from the community. \nThe second half of the thesis explores training and evaluating vision-language-action (VLA) models for off-road driving. The text-centric VLM pretraining gives policies the potential to generalize to scenarios not observed during fleet testing. We recast driving as a text completion problem to fully leverage this pretraining. We aim to train a policy that is responsive to both long-horizon waypoints and free-form language commands such as “stay on the gravel” or “avoid the puddle”. We find that training naively on VLM-generated commands is not enough to elicit language following\, as the image observation is more predictive of the future trajectory than the language command. Instead\, we augment the observation with alternative trajectories collected from other times the robot visited a particular location. We show that adding these “counterfactual” trajectories decreases language following error by 0.28m ADE over a naively auto-labeled dataset. Finally\, we perform closed loop evaluation in 3D reconstructions of previously unseen environments. We show our dataset annotation method improves the ability of an agent to navigate autonomously to waypoints hundreds of meters away while enabling language-based interventions at important decision points. \n\nCommittee:\nDeva Ramanan (advisor)\nYonatan Bisk\nAyush Jain
URL:https://www.ri.cmu.edu/event/data-mining-and-auto-labeling-for-promptable-driving-policies/
LOCATION:GHC 6115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260715T110000
DTEND;TZID=America/New_York:20260715T120000
DTSTAMP:20260911T204247
CREATED:20260710T161451Z
LAST-MODIFIED:20260710T161451Z
UID:151758-1784113200-1784116800@www.ri.cmu.edu
SUMMARY:MapForest: A Modular Field Robotics System for Forest Mapping
DESCRIPTION:Abstract:\nForests present compounding challenges for mobile mapping systems. Dense canopy degrades GNSS\, uneven terrain demands deployment across diverse platforms\, and no single sensing platform can capture the full vertical structure of a forest — from the canopy above to the understory below. Yet precise\, georeferenced maps of individual trees are exactly what ecologists and forest managers need to monitor invasive species\, estimate biomass\, and track ecosystem change at scale. \nThis talk presents MapForest\, a modular field robotics system that turns multi-modal sensor data (LiDAR\, IMU\, GNSS\, and RGB imagery) into georeferenced 3D reconstructions and GIS-ready outputs. A single compact payload deploys across five carriers (handheld\, bicycle\, ATV\, and two UAVs) without hardware modification\, enabling seamless data collection across heterogeneous terrain. \nThe core mapping pipeline extends a LiDAR-inertial SLAM framework with covariance-aware GNSS priors and a Huber robust loss\, reducing trajectory error by 67% relative to a GNSS-blind baseline. To bridge the viewpoint gap between aerial and ground surveys\, we develop two complementary aerial-terrestrial alignment methods: Tensor-MI\, an analytical mutual-information approach operating on tree-likelihood fields\, and CRAF\, a learned registration model combining modality-specific encoders with a cross-attention transformer. As a concrete ecological application\, MapForest localizes invasive Tree-of-Heaven from onboard imagery using a fine-tuned detector\, projecting detections into the 3D map and exporting georeferenced GeoTIFF layers suitable for direct use in forestry workflows. \nMapForest is evaluated across six field sites spanning approximately 30 km of multi-modal traversal data\, demonstrating that robust\, actionable forest inventory is achievable at a granularity unavailable to satellite or conventional aerial methods. \nThesis Committee:\nProf. Abhisesh Silwal (chair)\nProf. Michael Kaess\nJohn Kim
URL:https://www.ri.cmu.edu/event/mapforest-a-modular-field-robotics-system-for-forest-mapping/
LOCATION:GHC 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260715T130000
DTEND;TZID=America/New_York:20260715T140000
DTSTAMP:20260911T204247
CREATED:20260716T132559Z
LAST-MODIFIED:20260716T132559Z
UID:152106-1784120400-1784124000@www.ri.cmu.edu
SUMMARY:Knowledge Graph-Augmented Reinforcement Learning: Injecting Structured Task Knowledge into Arbitrary Policy Architectures
DESCRIPTION:Abstract:\nReinforcement learning agents in complex tasks often require extensive exploration of large state spaces before useful structure emerges. Instead of pure exploration for learning tasks\, it is possible to leverage high-level semantic knowledge such as recipes\, instructions\, or labels\, and adapt to new tasks by grounding that prior knowledge in the environment. This talk presents Knowledge Graph-Augmented Reinforcement Learning (KG-RL)\, a method that augments RL policies with structured graph knowledge and injects it into a variety of policy networks across a variety of environments. \nThe method takes the form of a plug-in adapter. A task knowledge graph is merged each step with a scene graph\, encoded by a Graph Convolutional Recurrent Network\, and pooled through a small recommender into a fixed-width feature vector concatenated with the policy backbone’s observation features. The backbone itself is untouched\, so the adapter slots into learned policies as varied as CNN-MLP\, SoftMoE-LSTM\, GTrXL\, and PoliFormer without modification; bringing it to a new environment requires only enumerating a handful of entities and relation templates. The knowledge graphs are instantiated against simulator-exposed tables in symbolic environments\, or mined by a one-time pre-pass of the same perception pipeline that builds per-step scene graphs in open-world scenarios. \nWe evaluate the adapter on four environments (Overcooked-AI\, MiniGrid\, Craftax\, and AI-Habitat ObjectNav) and four backbones. It delivers two improvements independent of backbone choice: faster training to the same final policy\, reaching the same asymptotic reward in up to 3x fewer environment steps\, and a higher final policy under a fixed budget\, where on Craftax and AI-Habitat the method overtakes the available published baselines. Both gains scale with task complexity\, and the adapter is robust to substantial knowledge-graph corruption\, retaining its advantage with half of the graph’s nodes removed. These results argue that structured task knowledge belongs as a default\, low-cost input channel in modern reinforcement learning. \nCommittee:\nKatia Sycara (advisor)\nChangliu Liu\nRenos Zabounidis
URL:https://www.ri.cmu.edu/event/knowledge-graph-augmented-reinforcement-learning-injecting-structured-task-knowledge-into-arbitrary-policy-architectures/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260715T133000
DTEND;TZID=America/New_York:20260715T150000
DTSTAMP:20260911T204247
CREATED:20260709T134546Z
LAST-MODIFIED:20260709T143200Z
UID:151750-1784122200-1784127600@www.ri.cmu.edu
SUMMARY:Tracing Generated Content Back to Training Data
DESCRIPTION:Abstract: \nAI-generated content is inherently derived from training data\, yet it remains a mystery which specific data points large generative models rely on for a given generation. To address this\, my research focuses on training data attribution—identifying the training images that are most influential in synthesizing a specific output. The ideal objective is to find the exact subset of training data that\, if removed\, would prevent a retrained model from generating that content. However\, the required combinatorial search and iterative retraining are computationally intractable. \nTo make this search tractable\, we propose using model unlearning as a proxy for the counterfactual model. By forcing a model to unlearn a synthesized output\, we can trace which training data points it effectively “removes” by observing changes in the training loss. We show that this is a highly effective method for finding influential data in text-to-image models. \nNext\, we interpret these attribution results by identifying which features best predict them. We show that joint text-image features\, after light calibration\, serve as strong predictors of attribution. We then analyze whether textual or visual features are more salient in the attribution process. We find that this salience shifts drastically across different settings. \nTo enable more fine-grained analysis of the attribution results\, we introduce TPIPS\, a similarity metric conditioned on specific visual aspects. We repurpose a vision-language model to compute text-conditioned\, aspect-specific perceptual similarity grounded in human judgments. \nBased on the analysis\, I will conclude by discussing the next open challenges and potential next steps in the field. \nThesis Committee Members: \nJun-Yan Zhu\, Chair \nDeva Ramanan \nRuslan Salakhutdinov \nAlexei A. Efros\, UC Berkeley \nDavid Bau\, Northeastern
URL:https://www.ri.cmu.edu/event/tracing-generated-content-back-to-training-data/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260717T130000
DTEND;TZID=America/New_York:20260717T140000
DTSTAMP:20260911T204247
CREATED:20260715T135032Z
LAST-MODIFIED:20260715T135032Z
UID:151825-1784293200-1784296800@www.ri.cmu.edu
SUMMARY:Modeling Inter-Agent Interactions: A Spatiotemporal Attention Approach to Multi-Agent Action Anticipation
DESCRIPTION:Abstract:\nAnticipating the near-future actions of multiple people is central to embodied systems that plan and coordinate in shared environments\, yet most research targets a single agent and ignores the inter-agent dependencies that shape group behavior. This thesis presents InteractFormer\, a unified model that treats inter-agent interaction as a first-class signal: operating directly on fine-grained visual tokens\, it lets agents attend to one another spatially and temporally to jointly predict all their futures. On two benchmarks—LEMMA (household collaboration) and SportsHHI (team sports)—it consistently outperforms strong single- and multi-agent baselines\, with the largest gains in genuinely multi-agent scenarios. \nCommittee:\nKatia Sycara (advisor)\nJiaoyang Li\nRenos Zabounidis
URL:https://www.ri.cmu.edu/event/modeling-inter-agent-interactions-a-spatiotemporal-attention-approach-to-multi-agent-action-anticipation/
LOCATION:GHC 9115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260720T160000
DTEND;TZID=America/New_York:20260720T170000
DTSTAMP:20260911T204247
CREATED:20260716T132420Z
LAST-MODIFIED:20260716T132420Z
UID:152104-1784563200-1784566800@www.ri.cmu.edu
SUMMARY:Towards Smarter and Safer Self-Improving AI
DESCRIPTION:Abstract:\nAs AI systems become more capable\, further progress may depend not only on scaling models and training data\, but also on enabling systems to evaluate and improve their own behavior and development. This raises a dual challenge: how can we make self-improvement more effective while ensuring increasingly autonomous systems remain trustworthy?\n\nThis thesis investigates self-improvement across three complementary directions:\n\n\nImprove individual outputs. We develop iterative refinement methods for compositional visual generation\, enabling models to progressively refine their outputs using feedback from vision-language model critics. We study how different forms of test-time scaling — depth\, breadth\, and hybrid strategies — trade off accuracy\, quality\, and computational cost.\n\nImprove the research process. We extend the feedback loop from individual generations to automated experimentation. Using LLM agents for machine-learning and robotic policy-learning tasks\, we investigate whether ‘autoresearch’ agents can propose improvements\, run experiments\, learn from their outcomes\, and accumulate experience that transfers across tasks.\n\nImprove safety and oversight. As agents become increasingly autonomous within self-improvement loops\, they may learn to mislead evaluators in pursuit of their objectives. We investigate lying in LLMs\, identify internal mechanisms and representations associated with deception\, and evaluate interventions to mitigate it.\n\n\nTogether\, these directions frame self-improvement as a feedback loop involving generation\, evaluation\, revision\, and learning. The thesis presents work toward making such loops more capable and trustworthy as they scale toward increasingly autonomous scientific discovery and recursive AI development\, as envisioned in the AI 2027 scenario (https://ai-2027.com/). \n\nCommittee:\nDeepak Pathak (advisor)\nShubham Tulsiani\nMihir Prabhudesai
URL:https://www.ri.cmu.edu/event/towards-smarter-and-safer-self-improving-ai/
LOCATION:GHC 8115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260721T123000
DTEND;TZID=America/New_York:20260721T133000
DTSTAMP:20260911T204247
CREATED:20260716T132808Z
LAST-MODIFIED:20260716T132808Z
UID:152108-1784637000-1784640600@www.ri.cmu.edu
SUMMARY:Evaluating World Models in Embodied Question Answering through Computational Primitives and Difficulty Progressions
DESCRIPTION:Abstract:\nLanguage modeling progress is largely evidenced by steadily rising scores on benchmarks of increasing apparent difficulty. From this\, the field infers increasingly general capabilities\, many of which presuppose robust world modeling. Interpreting a score\, however\, requires understanding both the task’s computational requirements and how the test-taker generalizes from them. Unlike humans\, who demonstrably generalize well\, large language models (LLMs) often do not: they instead learn heuristics fit to the minimal sufficient computational requirements of a task\, which may be far simpler than the task appears to demand. This talk extends this analysis to embodiment\, where multimodal LLMs (MLLMs) serve as the perception and reasoning core of embodied agents\, whose reliable deployment depends on benchmarked capability including a robust world model of the agent’s environment. \nWe evaluate world model robustness by independently manipulating object count\, duplication\, trajectory length\, and viewpoint change in a controlled synthetic benchmark of multi-frame egocentric trajectories\, and find that frontier MLLMs degrade sharply as difficulty increases\, while humans remain near ceiling. We then characterize embodied question answering (EQA) task demands through three nested paradigms: selecting relevant observations\, clustering nearby observations into local spatial-semantic units\, and propagating semantic state across experience. Stratifying questions by the weakest sufficient paradigm yields a difficulty progression. Prominent EQA benchmarks predominantly test only the weakest paradigm\, single-frame selection\, and so we introduce Campus-Bench\, three multi-hour\, campus-scale episode histories with questions stratified into selection and propagation regimes\, directly comparing model performance on the two over the same episodes. We additionally develop a diagnostic method in which an MLLM incrementally constructs and traverses a hierarchical spatial-semantic memory\, executing propagation explicitly. A frontier long-context MLLM performs strongly on selection but collapses on propagation; the same model leveraging our method dramatically improves it while remaining competitive on selection\, suggesting models struggle to maintain state internally. \nFrom this\, we argue that frontier MLLMs do not yet maintain the robust world models their benchmark scores suggest\, and that the field needs to make progress in formalizing tasks’ computational requirements and evaluating along their progressions of difficulty. \nThesis Committee:\nYonatan Bisk (Advisor)\nWennie Tabib\nHaochen Zhang
URL:https://www.ri.cmu.edu/event/evaluating-world-models-in-embodied-question-answering-through-computational-primitives-and-difficulty-progressions/
LOCATION:NSH 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260723T120000
DTEND;TZID=America/New_York:20260723T130000
DTSTAMP:20260911T204247
CREATED:20260716T160912Z
LAST-MODIFIED:20260716T160912Z
UID:152251-1784808000-1784811600@www.ri.cmu.edu
SUMMARY:A Robotic System for Tree Nursery Automation
DESCRIPTION:  \nAbstract:\nThe United States Green Industry faces a persistent labor shortage that motivates the adoption of agricultural automation.  However\, existing systems are not designed for the unstructured\, densely planted environment of a tree nursery. \nThis thesis presents a robotic system intended to alleviate this shortage while remaining usable by non-technical farmers\, built around a map-based representation of the nursery environment. A custom robotic platform\, the mini-Amiga\, and an accompanying LiDAR-camera-IMU sensor rig were developed to satisfy the maneuverability and payload requirements of tight nursery inter-row spacing. Point cloud maps constructed with this platform\, using the GLIM LiDAR-inertial SLAM framework augmented with a custom GNSS georeferencing extension\, were processed with a new constrained Gaussian Mixture Model algorithm to segment individual trees without requiring trunk visibility or large annotated training datasets. \nThe resulting per-tree map was further augmented with photographic colorization and encoded as a hierarchically organized Universal Scene Description (USD) scene\, supporting non-destructive\, multi-mode visualization and per-tree metadata storage intended for intuitive interaction by non-technical operators\, and was used to derive a Nav2-compatible occupancy grid and row-traversal paths intended for autonomous task execution. These results demonstrate that individual nursery trees can be accurately and efficiently segmented from point cloud data\, and that the resulting map can be represented in a form suited to both non-technical human interaction and autonomous navigation\, providing a practical foundation for future work integrating localization and autonomous task execution to fully realize the labor-saving potential of this system. \n\nCommittee:\nGeorge Kantor (advisor)\nMichael Kaess\nEaston Potokar
URL:https://www.ri.cmu.edu/event/a-robotic-system-for-tree-nursery-automation/
LOCATION:NSH 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260727T093000
DTEND;TZID=America/New_York:20260727T103000
DTSTAMP:20260911T204247
CREATED:20260721T131907Z
LAST-MODIFIED:20260721T131907Z
UID:152651-1785144600-1785148200@www.ri.cmu.edu
SUMMARY:Understanding Image Intrinsics Through Light and Heat
DESCRIPTION:Abstract:\nThis talk examines the physical interplay between light and heat: what is absent in the visible image—light not reflected by the scene—is absorbed as heat and manifests in the thermal spectrum. This process intrinsically couples illumination\, material properties\, and heat transport. By exploiting this coupling\, a single visible–thermal image pair provides complementary information that enables the decomposition of photometric image intrinsics\, namely incident illumination (shading) and surface reflectance (albedo). More broadly\, modeling the flow of energy across light and heat transport opens new opportunities for vision\, imaging\, and inverse graphics. \nCommittee:\nProf. Srinivasa G. Narasimhan (advisor)\nProf. Aswin C. Sankaranarayanan (advisor)\nProf. Matthew P. O’Toole\nSriram N. Narayanan
URL:https://www.ri.cmu.edu/event/understanding-image-intrinsics-through-light-and-heat/
LOCATION:NSH 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260727T100000
DTEND;TZID=America/New_York:20260727T110000
DTSTAMP:20260911T204247
CREATED:20260722T172447Z
LAST-MODIFIED:20260722T172447Z
UID:152776-1785146400-1785150000@www.ri.cmu.edu
SUMMARY:Hierarchical Manipulation Policies: Adapting to Unseen Objects and Discovering Sub-goals
DESCRIPTION:Abstract:\nA robot that manipulates one object well may still fail on the next. Generalizing across diverse objects and tasks is hard because such objects vary widely in geometry\, articulation\, and interaction dynamics. Hierarchical policies offer a powerful approach: a high-level policy predicts sub-goal end-effector poses\, and a low-level policy generates the actions to reach them. However\, dominant approaches leave the high-level policy brittle to unseen objects and dependent on deterministic sub-goal heuristics that many tasks cannot provide. This thesis asks: How should sub-goals be defined\, represented\, and communicated from the high level to the low level policy so that a single hierarchical policy generalizes across diverse objects and tasks? We address this question through two complementary projects\, both building on a prior hierarchical policy that grounds 3D sub-goal prediction in the observed scene.  \nWe first present demo-conditioned learning for adapting to out-of-distribution objects. Rather than fine-tuning\, the policy is conditioned on a single demonstration provided at test time. We show that reasoning about the demonstration and the current observation jointly in 3D outperforms compressing the demonstration into a latent embedding\, and that a single human hand demonstration can replace a teleoperated robot trajectory\, improving real-world performance on challenging unseen objects. \n  \nWe then present an uncertainty-aware hierarchical framework for tasks where sub-goals cannot be deterministically defined. Common heuristics\, such as gripper open/close transitions or near-zero end-effector velocity\, provide no signal for non-prehensile pushing\, sliding\, or manipulating levers and handles without a discrete grasp event. The framework derives candidate sub-goals through probabilistic changepoint segmentation\, represents the high-level goal distribution as a mixture model over candidate sub-goals\, and conditions the low-level policy on this distribution through goal-aware attention. \n  \nFinally\, this thesis extends hierarchical manipulation policies to challenging settings: adapting to unseen objects\, and modeling sub-goal uncertainty in trajectories that cannot be deterministically segmented. \n\n\nCommittee:\nDavid Held (advisor)\nZackory Erickson (advisor)\nShubham Tulsiani\nJason Liu
URL:https://www.ri.cmu.edu/event/hierarchical-manipulation-policies-adapting-to-unseen-objects-and-discovering-sub-goals/
LOCATION:GHC 8115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260727T130000
DTEND;TZID=America/New_York:20260727T140000
DTSTAMP:20260911T204247
CREATED:20260721T190859Z
LAST-MODIFIED:20260721T190859Z
UID:152761-1785157200-1785160800@www.ri.cmu.edu
SUMMARY:From Following to Leading: Adaptive Collaboration and Influence for Multi-Agent Teaming
DESCRIPTION:Abstract: \nAutonomous agents and robots are taking on increasingly collaborative roles alongside people\, from self-driving vehicles that navigate roads alongside human drivers\, to language-model agents that understand user intentions and execute tasks independently. In each of these settings\, success is determined not only by an agent’s individual task competence\, but by its ability to work effectively with others whose preferences and strategies directly shape the outcome of the task. This talk explores collaborative intelligence\, focusing on the requisite components to move from static\, purely reactive agents to proactive collaborators that reason\, adapt to\, and shape the behaviors of their teammates. \nWe introduce TALENTS\, an ad hoc teamwork algorithm that analyzes teammate behavior and dynamically adapts its own policy to best suit them. To accomplish this\, we first learn a latent strategy space from offline trajectory data via a variational autoencoder\, cluster this space into discrete teammate types\, and use a regret-minimization algorithm to infer and track which strategy a partner is following\, allowing the cooperator to adapt online as the partner’s behavior changes over the course of an episode. In both agent-agent evaluations and a 119 participant human-agent study in a modified version of the Overcooked-ai benchmark\, we demonstrate that TALENTS outperforms existing baselines in both quantitative task reward as well as subjective measures of team fluency and trust. \nFinally\, we extend beyond adaptation to examine proactive collaboration through the lens of multi-agent influence. Rather than treating a partner’s strategy as fixed and simply best-responding to it\, we investigate how an agent equipped with knowledge of how its teammate will respond to its actions can deliberately shape that learning process\, motivating partners to shift toward more effective joint conventions. Together\, these contributions establish several important algorithmic foundations needed to build autonomous agents that not only intelligently adapt to humans and other artificial teammates\, but also actively help shape more effective collaboration. \n\nCommittee:\nKatia Sycara (advisor)\nJiaoyang Li\nRenos Zabounidis
URL:https://www.ri.cmu.edu/event/from-following-to-leading-adaptive-collaboration-and-influence-for-multi-agent-teaming/
LOCATION:GHC 8115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260728T110000
DTEND;TZID=America/New_York:20260728T120000
DTSTAMP:20260911T204247
CREATED:20260722T134629Z
LAST-MODIFIED:20260722T134629Z
UID:152769-1785236400-1785240000@www.ri.cmu.edu
SUMMARY:What needs to be learned in robot learning? A case study: learning battery insertion from a diagram
DESCRIPTION:Abstract:\nManual diagrams are a rich and common knowledge source humans use to learn new skills\, but their use for robot learning is still underexplored. A challenge with instruction diagrams is that they communicate task progression in a sparse\, qualitative visual format\, relying on the learner’s prior physical understanding to fill in unmentioned execution details. This thesis investigates the fundamental question of \emph{what needs to be learned in robot learning} by exploring the interplay between explicit information extracted from instructions and implicit physical knowledge discovered through practice or prior knowledge\, using the cylindrical battery insertion task as a case study. \nFirst\, we present a pipeline that compiles static 2D instruction diagrams into metrically accurate 3D simulation environments. We use Vision-Language Models (VLMs) to extract qualitative scene topology and contact modes\, followed by a geometric optimization solver that certifies and refines metric object dimensions and spatial subgoals. Second\, using the reconstructed task keyframes\, we manually designed a control strategy on physical hardware. This hardware deployment exposes the limitations of purely explicit instructions and reveals crucial implicit details such as “tricks” – open-loop primitives exploiting the task’s physical properties that yield robustness gains even over naive closed-loop methods. Finally\, we investigate whether these implicit physical behaviors can be discovered autonomously by reinforcement and imitation learning. \nIn summary\, this thesis demonstrates a paradigm for bridging human data such as is found on YouTube\, as well as instructional text\, diagrams\, and explicit demonstration videos\, showing a promising approach to make use of the vast knowledge base of humans. \nCommittee:\nChristopher Atkeson (advisor)\nShubham Tulsiani\nYuemin Mao
URL:https://www.ri.cmu.edu/event/what-needs-to-be-learned-in-robot-learning-a-case-study-learning-battery-insertion-from-a-diagram/
LOCATION:GHC 8115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260728T123000
DTEND;TZID=America/New_York:20260728T140000
DTSTAMP:20260911T204247
CREATED:20260721T190755Z
LAST-MODIFIED:20260721T190913Z
UID:152762-1785241800-1785247200@www.ri.cmu.edu
SUMMARY:Robust Visuomotor Policy Learning in Uncertain World Models
DESCRIPTION:Abstract:\nWorld models have shown promise in robotics by imagining future outcomes of robot actions directly from high-dimensional sensor observations. However\, learning visuomotor policies with world models still faces significant reliability challenges\, as imagined futures from world models often diverge from the outcomes that may actually occur. This unreliability arises from uncertainty in learned world models\, including epistemic uncertainty\, where the model lacks sufficient knowledge in out-of-distribution regions\, and aleatoric uncertainty\, where intrinsic randomness in the system allows multiple plausible outcomes to arise under the same robot action. This thesis develops methods for learning robust visuomotor policies using uncertain world models by explicitly reasoning about and mitigating these uncertainties. \nWe first introduce UNISafe\, an uncertainty-aware latent safety filter for mitigating epistemic uncertainty in learned world models. We propose a principled framework for detecting out-of-distribution world model predictions by quantifying epistemic uncertainty and calibrating an uncertainty threshold with conformal prediction. Moreover\, this out-of-distribution detection is incorporated into Hamilton-Jacobi reachability analysis\, synthesizing the safety filter to proactively avoid regions where world model predictions are unreliable and thereby achieve robust\, safe visuomotor control. We then introduce StressDream\, an inference-time steering method for mitigating aleatoric uncertainty in world models. Instead of relying on nominal samples from the world model\, StressDream actively steers the imaginations of video world models to expose plausible but critical outcomes of robot actions. This enables more robust policy evaluation by uncovering failure modes of robot actions\, as well as improved policy optimization by training policies against challenging but realistic imagined futures. Together\, these methods enable visuomotor policies relying on learned but uncertain world models to achieve robust control in complex\, uncertain environments with high-dimensional sensor observations by explicitly reasoning about the uncertainties of the learned world model. \nThesis Committee:\nAndrea Bajcsy (chair)\nMax Simchowitz\nJeff Schneider\nMichelle Zhao
URL:https://www.ri.cmu.edu/event/robust-visuomotor-policy-learning-in-uncertain-world-models/
LOCATION:NSH 3002
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260728T130000
DTEND;TZID=America/New_York:20260728T140000
DTSTAMP:20260911T204247
CREATED:20260721T191042Z
LAST-MODIFIED:20260721T191042Z
UID:152765-1785243600-1785247200@www.ri.cmu.edu
SUMMARY:Robust Visuomotor Policy Learning in Uncertain World Models
DESCRIPTION:Abstract:\nWorld models have shown promise in robotics by imagining future outcomes of robot actions directly from high-dimensional sensor observations. However\, learning visuomotor policies with world models still faces significant reliability challenges\, as imagined futures from world models often diverge from the outcomes that may actually occur. This unreliability arises from uncertainty in learned world models\, including epistemic uncertainty\, where the model lacks sufficient knowledge in out-of-distribution regions\, and aleatoric uncertainty\, where intrinsic randomness in the system allows multiple plausible outcomes to arise under the same robot action. This thesis develops methods for learning robust visuomotor policies using uncertain world models by explicitly reasoning about and mitigating these uncertainties. \nWe first introduce UNISafe\, an uncertainty-aware latent safety filter for mitigating epistemic uncertainty in learned world models. We propose a principled framework for detecting out-of-distribution world model predictions by quantifying epistemic uncertainty and calibrating an uncertainty threshold with conformal prediction. Moreover\, this out-of-distribution detection is incorporated into Hamilton-Jacobi reachability analysis\, synthesizing the safety filter to proactively avoid regions where world model predictions are unreliable and thereby achieve robust\, safe visuomotor control. We then introduce StressDream\, an inference-time steering method for mitigating aleatoric uncertainty in world models. Instead of relying on nominal samples from the world model\, StressDream actively steers the imaginations of video world models to expose plausible but critical outcomes of robot actions. This enables more robust policy evaluation by uncovering failure modes of robot actions\, as well as improved policy optimization by training policies against challenging but realistic imagined futures. Together\, these methods enable visuomotor policies relying on learned but uncertain world models to achieve robust control in complex\, uncertain environments with high-dimensional sensor observations by explicitly reasoning about the uncertainties of the learned world model. \nThesis Committee:\nAndrea Bajcsy (chair)\nMax Simchowitz\nJeff Schneider\nMichelle Zhao
URL:https://www.ri.cmu.edu/event/robust-visuomotor-policy-learning-in-uncertain-world-models-2/
LOCATION:NSH 3002
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260728T133000
DTEND;TZID=America/New_York:20260728T143000
DTSTAMP:20260911T204247
CREATED:20260721T132229Z
LAST-MODIFIED:20260721T132303Z
UID:152653-1785245400-1785249000@www.ri.cmu.edu
SUMMARY:Terrain-Aware Dynamics Models for High-Speed Off-Road Navigation
DESCRIPTION:Abstract:\nHigh-speed autonomy in the real world requires accurate control\, which often relies on dynamics models that capture the complex interaction between a robot and its environment. In off-road regimes\, this terrain interaction dominates the robot’s dynamics\, driven by formidable characteristics such as diverse surface properties\, complex geometries\, environment diversity\, and high-speed instability. This thesis investigates how terrain-aware perception can improve learned dynamics modeling and control at high speed and how such models can be rigorously evaluated before deployment. \nFirst\, we demonstrate how perception representations can make dynamics models terrain-aware\, capturing geometric and semantic details that physics-based models miss and that simplistic learned models overlook. We formulate a representation that extracts the most relevant terrain features given a robot’s motion\, yielding higher prediction accuracy. Second\, we introduce a rigorous evaluation method to mitigate real-world failures. We collect a challenging\, multi-season dataset at speeds up to 13 m/s and mine the most difficult evaluation samples using our benchmarking method. While models appear accurate on average\, our benchmark surfaces the long-tail failure cases where prior models fail catastrophically. \nTogether\, with our verified\, terrain-aware model\, we decrease the worst-case prediction error by 23.8%\, compared to physics-based and learned baselines. We further evaluate on a full-scale ATV platform across high-speed (>10m/s) and geometrically challenging courses with a 34.9% reduction in maximum cross-track error. These results demonstrate the importance of embedding environment context for locomotion-related tasks. \nThesis Committee:\nWenshan Wang (co-chair)\nSebastian Scherer (co-chair)\nAaron Johnson\nAnoushka Alavilli
URL:https://www.ri.cmu.edu/event/terrain-aware-dynamics-models-for-high-speed-off-road-navigation/
LOCATION:GHC 6115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260728T150000
DTEND;TZID=America/New_York:20260728T160000
DTSTAMP:20260911T204247
CREATED:20260721T141849Z
LAST-MODIFIED:20260721T141849Z
UID:152655-1785250800-1785254400@www.ri.cmu.edu
SUMMARY:Robotic Localization and Mapping of Disease in Apple Orchards
DESCRIPTION:Abstract:\nUnited States apple growers lose more than $100 million a year to fire blight\, a bacterial disease caused by Erwinia amylovora that is detrimental to pome fruit trees such as apple and pear.This bacterium infects blossoms\, shoots\, and branches during the bloom season causing tissue to die. Therefore\, the detection and removal of infected tissues over the dormant season is critical to prevent an outbreak in the following spring. However\, the symptoms of fire blight are subtle and difficult to detect\, and finding them still depends almost entirely on manual scouting\, a process that does not scale to large orchards and is prohibitively expensive and time-consuming for growers to perform at the scale required for effective disease management. With the decreasing availability of labor in the agricultural sector\, there is a pressing need for automated solutions that can perform this inspection at scale\, and with high accuracy. \nThis thesis aims to develop a robotic system capable of performing that inspection autonomously. Since the visual cues that distinguish infected tissue are subtle\, easily occluded\, vary with natural lighting\, and are reliably resolved only at close range\, the system relies on active perception: a manipulator positions a camera to acquire discriminative\, task-relevant views of the canopy. We have collected a multi-modal dataset of dormant apple trees\, the first to pair dense near-infrared imagery with flash-illuminated stereo RGB for this task\, and trained detectors to recognize disease symptoms across both modalities. We then introduce a confidence-aware semantic mapping method that fuses these per-view detections into a persistent 3D representation of disease confidence across the canopy\, and a next-best-view planner that actively selects viewpoints to refine the map’s least confident\, most contested predictions. Finally\, we integrated the full pipeline onto Erwin\, a mobile active perception robot built using an Amiga base and an xArm6 arm for manipulation of the camera rig. \n\nWe successfully validated the system both in simulation and in the field on 12 trees at the Penn State Fruit Research and Extension Center under a genuine train–test domain gap\, the semantic planner more than doubles the detection accuracy of a complete planar scan by the halfway point of the inspection\, concentrating its views on the map’s most contested disease evidence. This result demonstrates the central promise of active perception for orchard disease mapping: by deciding where to look next a robot can build maps of orchard disease that are accurate enough to support autonomous disease management\, while minimizing the time and energy spent on inspection. \nThesis Committee:\nProf. Abhisesh Silwal (chair)\nProf. Oliver Kroemer\nItamar Mishani
URL:https://www.ri.cmu.edu/event/robotic-localization-and-mapping-of-disease-in-apple-orchards/
LOCATION:NSH 1109
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260728T154500
DTEND;TZID=America/New_York:20260728T164500
DTSTAMP:20260911T204247
CREATED:20260723T192143Z
LAST-MODIFIED:20260723T192143Z
UID:152840-1785253500-1785257100@www.ri.cmu.edu
SUMMARY:The Tail and the Judge: Better Policy Gradients and World Models for Reinforcement Learning
DESCRIPTION:Abstract: \nA quiet shift has taken place in how reinforcement learning delivers results.\nIncreasingly\, performance comes not from a single attempt but from selection\nover many: a language model is sampled repeatedly and judged by its best\nverified answer\, and a planning agent imagines many action sequences and\nexecutes the one its learned world model scores highest. Selection can fail\nin exactly two ways. The candidate pool may contain nothing worth choosing\,\nor the judge that scores the candidates may be wrong precisely where it\nmatters. This thesis argues that standard training objectives invite both\nfailures\, and it redesigns them to prevent each. \nThe first part trains the policy for selection itself. Standard methods\nmaximize the average reward of a single attempt\, which spreads learning\neffort evenly across quality levels and neglects the rare\, excellent outcomes\nthat selection exists to find. Extending maximum-likelihood reinforcement\nlearning from binary to continuous rewards\, we introduce a tail-likelihood\nobjective that credits the policy for covering every level of quality and\nconcentrates its gradient on the levels it reaches only rarely. We prove that\nthis objective aligns training with best-of-many deployment at every sampling\nbudget simultaneously\, and that it is optimized by an unbiased\,\nhyperparameter-free estimator amounting to a one-line change in standard\npipelines. Empirically\, the method recovers the supervised gold-standard\ngradient where popular baselines remain permanently misaligned; it lifts\nmaze policies from a near-zero starting success rate to solving most mazes\noptimally; and it is evaluated further on object localization\,\nprogram-speed optimization\, and vision-language grounding. \nThe second part secures the judge. A planner that maximizes predicted return\nwill seek out exactly the regions where its world model is most optimistically\nwrong. We prove that the resulting loss in policy quality is governed by the\nsharpness of the model’s loss landscape\, and we show that sharpness-aware\ntraining of the world model alone\, leaving the planner and policy untouched\,\ndelivers large and consistent gains that transfer across high-dimensional\ncontinuous control and pixel-based discrete control\, and across two different\nmodel-based algorithms. \nGood selection needs something worth choosing and a judge worth trusting.\nProviding both\, and composing them\, is the program of this thesis. \nCommittee:\nProf. Jeff Schneider (advisor)\nProf. Andrea Zanette (advisor)\nProf. Ruslan Salakhutdinov\nWentse Chen
URL:https://www.ri.cmu.edu/event/the-tail-and-the-judge-better-policy-gradients-and-world-models-for-reinforcement-learning/
LOCATION:NSH 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260729T120000
DTEND;TZID=America/New_York:20260729T130000
DTSTAMP:20260911T204247
CREATED:20260722T134829Z
LAST-MODIFIED:20260722T134829Z
UID:152771-1785326400-1785330000@www.ri.cmu.edu
SUMMARY:Efficient cross-embodiment transfer via world models and policy steering
DESCRIPTION:Abstract:\nLearning robot manipulation policies typically requires large amounts of high-quality demonstration data\, which is expensive to collect on real robots. While large robot and human datasets exist\, embodiment gaps make transferring knowledge between platforms challenging. This thesis investigates cross-embodiment transfer—how robots can learn from experience collected on different robots and humans to reduce the need for new demonstrations.\n\nI present Latent Policy Steering (LPS)\, a framework that pretrains a world model across diverse embodiments using embodiment-agnostic visual dynamics represented by optical flow. The pretrained model can be adapted to an unseen robot using less than an hour of teleoperation data. At deployment\, the world model enables the policy to anticipate the consequences of its actions\, identify likely mistakes\, and steer itself back toward safer behaviors through latent-space search. Experiments on both real-world manipulation tasks and simulation benchmarks demonstrate that this approach consistently improves policy performance and outperforms methods relying on embodiment-specific representations.\nIn summary\, LPS demonstrates an effective paradigm for efficient cross-embodiment transfer\, leveraging diverse\, cost-effective robot and human data to adapt to an unseen robot with minimal demonstrations\, reducing data collection costs and enabling more scalable robot learning. \nCommittee:\nJeff Schneider (advisor)\nChristopher Atkeson\nTejus Gupta
URL:https://www.ri.cmu.edu/event/efficient-cross-embodiment-transfer-via-world-models-and-policy-steering/
LOCATION:NSH 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260729T143000
DTEND;TZID=America/New_York:20260729T153000
DTSTAMP:20260911T204247
CREATED:20260727T132903Z
LAST-MODIFIED:20260727T132903Z
UID:152892-1785335400-1785339000@www.ri.cmu.edu
SUMMARY:The Limits of Prompt-Based Self-Improvement in Language Model Agents: Meta-Optimization and Failure Attribution
DESCRIPTION:Abstract: \nModern language models’ in-context learning and instruction-following abilities make it possible to adapt agent behavior without updating model weights. This thesis studies online self-improvement\, where an agent encounters each task only once and must accumulate reusable experience across a sequence of tasks\, reflecting real-world deployments in which actions have persistent consequences and retries may be impossible. \nFirst\, we ask whether meta-optimization can be integrated with online self-improvement by allowing the prompt optimizer itself to learn across episodes. Building on Agentic Context Engineering (ACE)\, we evaluate three variants that equip its optimizer with persistent memory. On the AppWorld benchmark\, none reliably improves over ACE when using strong base language models. Analysis suggests that ACE-style adaptation is near saturation in this regime\, leaving little room for meta-optimization to provide further gains. \nTo probe the limits of ACE-style adaptation\, we evaluate it on Toolathlon\, a harder and more heterogeneous tool-use benchmark\, and observe no improvement over non-adapting agents. Auditing these apparent failures reveals ambiguous task specifications\, missing information\, brittle evaluators\, and unstable external services\, rather than cleanly attributable agent errors. Because agent trajectories can span hundreds of thousands of tokens\, we develop an automated\, evidence-grounded pipeline for analyzing trajectories and attributing failures at scale. Together\, these results show that effective online self-improvement depends not only on prompt optimization\, but also on transferable structure across tasks and trustworthy failure attribution. \n\n\nCommittee:\nAran Nayebi (advisor)\nDaniel Fried\nHaochen Zhang
URL:https://www.ri.cmu.edu/event/the-limits-of-prompt-based-self-improvement-in-language-model-agents-meta-optimization-and-failure-attribution/
LOCATION:NSH 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260729T160000
DTEND;TZID=America/New_York:20260729T170000
DTSTAMP:20260911T204247
CREATED:20260724T135742Z
LAST-MODIFIED:20260724T135742Z
UID:152867-1785340800-1785344400@www.ri.cmu.edu
SUMMARY:Exploiting Structure for Real-Time Robot Motion Planning and Control
DESCRIPTION:Abstract: \nReal-time robot motion in dynamic and uncertain environments requires algorithms that can make effective decisions under strict computational constraints. High-fidelity dynamics\, large search spaces\, uncertainty over future outcomes\, and the pursuit of optimal motion each impose substantial computational costs. These challenges become especially pronounced in partially observable and rapidly changing environments\, where a robot must continually update its predictions and react before a previously computed plan becomes obsolete. This thesis presents two model-based frameworks that exploit structure to bridge deliberative planning and reactive control\, enabling efficient closed-loop execution without solving a full planning problem at every step. \nFirst\, we study projectile interception\, where a robot must begin moving after only a few noisy measurements of a fast-moving object’s trajectory. Our framework couples uncertainty estimates from an RGB-D tracking system with a sparse kinodynamic graph of executable motion primitives. As the projectile estimate evolves\, the system rapidly reevaluates candidate motions and executes partial trajectories that remain robust across a distribution of possible future outcomes. This approach maintains millisecond-scale replanning and improves interception performance. \nThe second framework learns the parameters of convex optimization-based safety filters formulated as control barrier function quadratic programs (CBF-QPs). Imitation learning transfers goal-directed behavior from a computationally expensive global planner into a local reactive controller while retaining explicit model-based safety constraints. This combination shifts computation from online global search to offline learning\, enabling efficient closed-loop execution with reduced planning effort. We show that\, under matched online computational budgets\, the learned safety filter improves performance in planar navigation environments with moving obstacles and in constrained manipulator-planning tasks. \nTogether\, these contributions show how structured representations can preserve reasoning about dynamics\, uncertainty\, and safety while enabling efficient real-time robot motion. \n\nCommittee:\nMaxim Likhachev (co-advisor)\nHowie Choset (co-advisor)\nAndrea Bajscy\nItamar Mishani
URL:https://www.ri.cmu.edu/event/exploiting-structure-for-real-time-robot-motion-planning-and-control/
LOCATION:GHC 9115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260730T090000
DTEND;TZID=America/New_York:20260730T100000
DTSTAMP:20260911T204247
CREATED:20260723T152807Z
LAST-MODIFIED:20260723T152807Z
UID:152778-1785402000-1785405600@www.ri.cmu.edu
SUMMARY:Reference-Prompted Instance Segmentation for Growing Retail Catalogs
DESCRIPTION:Abstract:\n \nRetail and warehouse perception systems work against a catalog that never stops changing: a segmenter deployed today will eventually be asked to find products that did not exist when it was trained. A closed-set segmenter can only grow its vocabulary by widening its classification head and fine-tuning it\, which risks the classes it already handled\, or by retraining on the whole accumulated catalog\, which costs more with every product added. This thesis studies reference-prompted segmentation\, which specifies the target product with a canonical image rather than a class label or a text description\, so that a product’s identity enters the model as an input rather than as a slot in its output. We fine-tune a reference-conditioned segmenter one new product at a time — twenty-seven sequential steps growing a three-product catalog to thirty — using Learning without Forgetting\, and compare it against a closed-set segmenter given the identical recipe\, budget\, and data. Our method retains its earlier products substantially better than the closed-set baseline. Every dataset in this thesis comes from isaac_datagen\, a standalone system built on Isaac Sim as a synthetic-data engine for retail scenes. Under this protocol and at this scale\, reference prompting lets a perception system keep pace with a growing catalog without ever retraining on all of it.\n\nCommittee:\nJeff Ichnowski (advisor)\nShubham Tulsiani\nBardienus Duisterhof
URL:https://www.ri.cmu.edu/event/reference-prompted-instance-segmentation-for-growing-retail-catalogs/
LOCATION:NSH 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260730T110000
DTEND;TZID=America/New_York:20260730T120000
DTSTAMP:20260911T204247
CREATED:20260727T133112Z
LAST-MODIFIED:20260727T133112Z
UID:152894-1785409200-1785412800@www.ri.cmu.edu
SUMMARY:Risk-Aware Multi-Agent Navigation in Dynamic Smoke Environments
DESCRIPTION:Abstract:\nIn wildfire scenarios\, deploying autonomous drones requires safely anticipating the dynamic behavior of dense smoke to coordinate effectively. Unlike traditional rigid obstacles\, smoke represents a fast-moving\, complex fluid hazard that impairs visual navigation and onboard sensors. In this thesis\, we propose a novel risk-aware\, multi-agent path planning framework that treats dynamic smoke as a continuous physical hazard. By leveraging Probabilistic Fourier Neural Operators (PFNO)\, our architecture forecasts the spatiotemporal behavior of smoke density while simultaneously quantifying its inherent aleatoric uncertainty. These probabilistic predictions are mapped into a safety cost using a Conditional Value-at-Risk (CVaR) metric\, which directly informs a safe variation of a time-varying Model Predictive Path Integral (MPPI) controller to ensure collision-free coordination. We integrate inter-agent safety bounds directly within the sampling-based planner of each drone and shield its output using Control Barrier Functions (CBFs). Furthermore\, because individual agents might possess a limited field of view\, we address the critical challenge of partial observability. We explore a potential initial solution for an active search policy driven by epistemic uncertainty mapping\, demonstrating how the agents collaboratively construct partial global maps to navigate safely. Extensive results demonstrate that this unified multi-agent architecture successfully coordinates the agents\, conservatively overbounding active smoke fronts and minimizing both cumulative and high-density smoke exposure compared to standard reactive baselines. \nCommittee:\nProf. Katia Sycara (Advisor)\nProf. John Dolan\nAndrew Jong
URL:https://www.ri.cmu.edu/event/risk-aware-multi-agent-navigation-in-dynamic-smoke-environments/
LOCATION:Gates 6115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260730T153000
DTEND;TZID=America/New_York:20260730T163000
DTSTAMP:20260911T204247
CREATED:20260713T154004Z
LAST-MODIFIED:20260713T183008Z
UID:151761-1785425400-1785429000@www.ri.cmu.edu
SUMMARY:[MS Thesis Talk] Marble: An On-Manifold Approach to Solving Mathematical Programs with Complementarity Constraints
DESCRIPTION:Date: Thursday\, July 30\, 2026\nTime: 3:30 PM – 4:30 PM\nLocation / ZOOM Link: (GHC 6115 / https://cmu.zoom.us/j/96096959582 ) \nAbstract:\nMany problems in robotics require reasoning over a mix of continuous dynamics and discrete events\, such as making and breaking contact in manipulation and locomotion. These problems are locally well modeled by quadratic programs with complementarity constraints (QPCCs). While very expressive\, QPCCs are non-convex problems\, and few solvers exist for computing fast\, local solutions for use in planning pipelines. In this work\, we develop an open-source solver\, Marble\, designed to solve QPCCs using an on-manifold complementarity relaxation and homotopy technique. The resulting solver avoids many of the classical issues with complementarity constraints and exhibits competitive speed and robustness across QPCC benchmarks and robotics-specific examples compared to existing baselines. \nCommittee:\nZac Manchester (advisor)\nRajan Gill\nArun Bishop \n—\nMICAH REICH \nMSR Student\, Robotics Institute – CMU
URL:https://www.ri.cmu.edu/event/ms-thesis-talk-marble-an-on-manifold-approach-to-solving-mathematical-programs-with-complementarity-constraints/
LOCATION:Gates Hillman Center 6115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
END:VCALENDAR