BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260401T083000
DTEND;TZID=America/New_York:20260401T093000
DTSTAMP:20260922T134535
CREATED:20260323T185607Z
LAST-MODIFIED:20260323T185607Z
UID:150703-1775032200-1775035800@www.ri.cmu.edu
SUMMARY:Regression-based Multi-view Face Synthesis
DESCRIPTION:Abstract:\nSynthesizing photorealistic human faces from novel viewpoints using only a single frontal image remains a challenging problem in computer vision. Large viewpoint changes introduce geometric distortions\, self-occlusions\, and missing visual information\, making identity preservation and high-frequency detail reconstruction particularly difficult. While recent generative approaches such as diffusion models and 3D-aware neural representations produce visually compelling results\, they typically require expensive training and slow inference. In contrast\, lightweight feed-forward models enable efficient inference but often fail to capture fine-grained details and complex appearance variations. \nThis thesis presents a geometry-guided framework that integrates explicit 3D structure with learned image refinement to achieve both efficiency and realism. The method first builds a geometrically consistent prior by fitting a 3D Morphable Model\, estimating a texture map from the frontal image\, augmenting it with a lightweight hair representation\, and rendering the target viewpoint. A convolutional residual network then refines this prior by predicting residual corrections that restore fine details and enhance local realism while preserving geometric consistency. Adversarial supervision further improves perceptual quality\, encouraging sharper textures and more natural appearance without increasing inference cost. \nWe compare the proposed approach against Cap4D\, a state-of-the-art method\, in a single-image side-view synthesis setting. The results demonstrate substantially improved computational efficiency—achieving inference within seconds on a single GPU—while maintaining stronger identity preservation. These findings show that geometry-guided residual refinement offers a practical and scalable alternative to heavy 3D-aware generative pipelines for identity-consistent novel view synthesis.
URL:https://www.ri.cmu.edu/event/regression-based-multi-view-face-synthesis/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260403T100000
DTEND;TZID=America/New_York:20260403T113000
DTSTAMP:20260922T134535
CREATED:20260325T175307Z
LAST-MODIFIED:20260327T144650Z
UID:150733-1775210400-1775215800@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - Ananya Rao
DESCRIPTION:Who: Ananya Rao\nDate: Friday\, April 3rd\nTime: 10 AM ET\nLocation: NSH 3305\nZoom Link\nMeeting ID: 949 1271 6317\nPasscode: 809985\n\nTitle: Spectral-Based Coordination of Heterogeneous Multi-Agent Teams for Information Gathering\n\nAbstract:\n\nExtreme environments\, such as those encountered in planetary exploration or disaster response\, present complex\, time-sensitive tasks with significant uncertainty. In such settings\, heterogeneous teams of robots with diverse capabilities can offer more robust solutions. By efficiently coordinating robots with specialized skills\, these teams can adapt to unknowns\, enhance safety\, and successfully execute complex tasks. The challenge lies in coordinating these multi-agent teams and maximizing the unique strengths of each member to achieve a shared objective under harsh constraints and high costs of failure. Current strategies\, however\, are computationally intensive and often rely on expert oversight\, posing a significant burden on practitioners in real-world deployment. \nThis thesis focuses on developing efficient collaboration strategies for coordinating heterogeneous robotic teams for information gathering tasks. The work addresses multi-agent coordination through three main elements: goal decomposition\, robot-task allocation\, and deployment decisions. We not only present methods that improve the operations of a team when used as individual modules\, but also present an integrated pipeline that is a step towards autonomous heterogeneous multi-agent exploration. The foundation of this approach is a novel spectral-based framework. This framework utilizes spectral decomposition to represent the diverse capabilities of agents and the characteristics of the environment in a compact manner. This representation then serves as the basis for efficiently building and coordinating teams. The developed methods are rigorously evaluated using real-world datasets and expert feedback. \n\n\n\n\n\n\n\n\n  \nThesis Committee Members: \nDavid Wettergreen\, co-chair \nHowie Choset\, co-chair \nAndrea Bajscy \nRobin Murphy\, Texas A&M University \nGuillaume Sartoretti\, National University of Singapore \n  \nA draft of the thesis document is available at this link.
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-ananya-rao/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260407T100000
DTEND;TZID=America/New_York:20260407T120000
DTSTAMP:20260922T134535
CREATED:20260325T135207Z
LAST-MODIFIED:20260325T175439Z
UID:150727-1775556000-1775563200@www.ri.cmu.edu
SUMMARY:RI Ph.D. Thesis Defense - Mrinal Verghese
DESCRIPTION:Date: 7th April 2026\nTime: 10:00 a.m. (ET)\nLocation: NSH 3305\nZoom: Link\nType: Ph.D. Thesis Defense\nWho: Mrinal Verghese\nTitle: Strategies for Robot Learning from Human Data\n \nAbstract:\nRobot learning is fundamentally data-constrained. Internet-scale human data is a promising source of additional data about human environments\, tasks\, and common skills. This data comes in diverse representations\, such as human videos and Large Language Models\, and can provide supervision and priors at all levels of robot reasoning and control.  However\, leveraging this data for robot learning is not trivial. Regardless of the representation\, human data often lacks physical details and contains drastically different embodiments and environments from our target robot deployments.\n\nIn this thesis\, I argue that one of the keys to effectively leveraging human data for robot learning lies in identifying appropriate features and modalities in this human data for each level of robot reasoning. My work on robot learning from human data follows a three-step process: analyze deployed robot learning systems augmented with human data on common tasks\, identify appropriate features and failure modes\, and integrate models trained on these appropriate features into existing robot learning and reasoning paradigms. This thesis demonstrates these strategies across both task planning and skill learning domains. In task planning\, we identify high-level language representations of visual details as good features and design an approach that leverages Bayesian reasoning about information gain to better ground LLM-based planners to their environment. In skill learning\, we identify visual motion representations as good features and present an approach that uses dense reward signals learned from human video to rapidly improve robot performance in real-world experiments. Across our experiments\, we observe that for a robotics subproblem\, appropriate features and representations from human data for that subproblem are often ones that have a similar level of abstraction. For example\, we found high-level language details to be the most effective history representations for high-level task planning with LLMs and low-level visual motion features to most effectively capture salient information from human demonstrations when training robot visuomotor policies for skills. In addition to the specific methods and approaches presented here\, the work in this thesis offers a general strategy for robot learning from diverse human data.\n\n\n \nThesis Committee Members:\nChristopher Atkeson (Chair)\n\n\nOliver Kroemer\nDave Held\nRuta Desai (ex-FAIR)
URL:https://www.ri.cmu.edu/event/ri-ph-d-thesis-defense-mrinal-verghese/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260407T153000
DTEND;TZID=America/New_York:20260407T170000
DTSTAMP:20260922T134535
CREATED:20260325T191007Z
LAST-MODIFIED:20260331T183919Z
UID:150740-1775575800-1775581200@www.ri.cmu.edu
SUMMARY:RI PhD Defense - Neehar Peri
DESCRIPTION:Date: April 7th\, 2026\nTime: 3:30 PM (ET)\nLocation: NSH 3305 \nZoom Link \nType: Ph.D. Thesis Defense \n\n\nWho: Neehar Peri\nTitle: Towards Scalable Open-World 3D Perception \nAbstract:  \nState estimation is a fundamental component of embodied perception. For safe navigation\, we argue that robots (and autonomous vehicles (AVs) specifically) must detect\, track\, and forecast all object categories\, not just those seen during training. In this thesis\, we study open-world 3D perception along three complementary axes: (i) long-tailed recognition for offline data curation\, (ii) rapid model adaptation to new concepts via few-shot multi-modal examples\, and (iii) low-level 3D motion understanding for fast reactive control. \nContemporary autonomous vehicle (AV) benchmarks have advanced techniques for training 3D detectors on large-scale data. Notably\, although prior work has nearly solved 3D object detection for a few common classes (e.g.\, pedestrian and car)\, detecting many rare classes in-the-tail (e.g.\, debris and stroller) remains challenging. This limitation is especially critical for offline scenario mining\, where identifying rare but safety-critical events is essential. We show that fine-grained tail class accuracy is significantly improved via multi-modal fusion of RGB images with LiDAR; fine-grained classes are difficult to identify from sparse LiDAR geometry alone\, suggesting that multi-modal cues are crucial for long-tailed 3D detection. To this end\, we study a simple late-fusion framework that ensembles independently trained uni-modal LiDAR and RGB detectors. Importantly\, this formulation allows us to leverage large-scale uni-modal datasets (with more examples for rare classes) to train stronger RGB detectors\, unlike prevailing multimodal approaches that require paired multi-modal data. While such models improve the detection accuracy of rare categories\, open-world perception also requires adapting to new and evolving concepts from limited supervision. \nThe emergence of vision-language models (VLMs) trained on web-scale datasets challenges conventional formulations of open-world perception. We revisit few-shot object detection (FSOD) in the context of such foundation models. Zero-shot predictions from models such as GroundingDINO already outperform state-of-the-art few-shot detectors (48 vs. 33 AP) on COCO\, yet remain misaligned with out-of-distribution target domains. For instance\, trucks on the web (e.g. pickup trucks) may be defined differently from trucks in autonomous driving scenarios (e.g. semi-trucks). We therefore reformulate few-shot recognition as aligning foundation models to target concepts using a small number of examples. These examples can be naturally multi-modal\, combining text and visual cues\, analogous to how human annotators learn to annotate new categories. Concretely\, we propose Foundational FSOD\, a benchmark protocol that evaluates detectors that are pre-trained on arbitrary external data and are adapted using multi-modal K-shot examples per class. Together with long-tailed detection\, Foundational FSOD enables scalable discovery of rare and ambiguously defined object categories for scenario mining. \nFinally\, beyond semantic recognition and offline discovery\, on-robot open-world perception systems must support fast\, reactive decision-making. In safety critical scenarios\, we argue that accurate 3D motion estimation is more important for evasive maneuvering than explicit categorization. We therefore study LiDAR scene flow\, which formalizes the task of estimating per-point 3D motion between consecutive point clouds. Prior methods achieve centimeter-level accuracy but are typically trained on a single sensor\, limiting generalization. In contrast\, we learn motion priors that transfer across diverse and unseen LiDAR sensors. While prior work in LiDAR segmentation and detection suggests that naive multi-dataset training degrades performance\, we find that this conventional wisdom does not hold for motion estimation: scene flow models benefit substantially from cross-dataset training without architectural changes. Our analysis suggests that low-level motion cues are less sensitive to sensor configuration; indeed\, models trained on fast-moving objects (e.g.\, from highway datasets) perform well on fast-moving objects\, even across different datasets. Building on this insight\, we propose UniFlow\, a simple feedforward model trained jointly on multiple large-scale scene flow datasets with diverse sensor setups. UniFlow establishes a new state-of-the-art on Waymo and nuScenes\, improving over prior work by 5.1% and 35.2%\, respectively\, and generalizes to unseen datasets like TruckScenes and AEVAScenes. \nThesis Committee Members: \nDeva Ramanan\, Chair \nShubham Tulsiani\nKaterina Fragiadaki\nSanja Fidler\, University of Toronto\nGeorgia Gkioxari\, California Institute of Technology \n\nThesis Draft: https://drive.google.com/drive/folders/1qHhZgJnFcPG9ITdhICnwC9IpgPORUWcp?usp=sharing
URL:https://www.ri.cmu.edu/event/ri-phd-defense-neehar-peri/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260410T090000
DTEND;TZID=America/New_York:20260410T110000
DTSTAMP:20260922T134535
CREATED:20260330T170410Z
LAST-MODIFIED:20260330T170410Z
UID:150800-1775811600-1775818800@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - Abigail Breitfeld
DESCRIPTION:Date: April 10th\, 2026\nTime: 09:00 AM (ET)\nLocation: NSH 4305 \nZoom Link \nType: Ph.D. Thesis Defense \n\n\nWho: Abigail Breitfeld \nAbstract: Dynamic Multi-Objective Path Planning for Exploration \nAbstract:\n\nRobotic explorers play a crucial role in acquiring data from areas that are difficult or impossible for humans to reach. Whether for planetary exploration\, search and rescue missions\, agriculture\, or other scientific exploration tasks\, these robots can utilize pre-existing knowledge of the terrain to navigate effectively. In these scenarios\, robots must consider various factors including scientific data acquisition\, risk assessment\, and energy consumption. Importantly\, they must also adapt their strategies as these objectives evolve over time. This thesis addresses the unique challenge of autonomously planning paths for time- and information-evolving objectives.\n\nWe present planning methods that efficiently generate paths that honor a desired trade-off among objectives even in changing conditions. We introduce two approaches based on ergodic search\, one leveraging prior data to predict how to best balance objectives\, and another employing decomposition-based optimization to maintain a consistent trade-off even as objectives change. We further extend these ideas to graph search\, developing methods that adapt traditional algorithms like A* and Monte Carlo Tree Search for dynamic multi-objective planning.\n\nOur methods are validated using terrestrial and lunar terrain data\, demonstrating improved computational efficiency while maintaining solution quality compared to state-of-the-art methods. Additionally\, we present local trajectory planning strategies for real-time hazard avoidance\, which can be integrated with our global planning methods for robust end‑to‑end autonomy. In all\, these contributions advance the efficiency and adaptability of multi-objective planning for planetary and terrestrial exploration robots.\n\n\nCommittee:\n\nDavid Wettergreen\,chair\nMaxim Likhachev\nGeorge Kantor\nAlberto Candela\, NASA Jet Propulsion Laboratory\n\n\nThesis Draft:\nhttps://drive.google.com/drive/folders/1jKIfNL0VY0vySJjf2tdWncscOn7Zix4d?usp=drive_link
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-abigail-breitfeld/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260410T150000
DTEND;TZID=America/New_York:20260410T170000
DTSTAMP:20260922T134535
CREATED:20260407T141430Z
LAST-MODIFIED:20260407T141430Z
UID:150898-1775833200-1775840400@www.ri.cmu.edu
SUMMARY:Learning Dynamic Rope Manipulation with Task-Level Iterative Learning Control
DESCRIPTION:Abstract: Dynamic manipulation of deformable objects is challenging for humans and robots because they have infinite degrees of freedom and exhibit underactuated dynamics. This thesis introduces a Task-Level Iterative Learning Control method for dynamic manipulation of deformable objects and demonstrates this method on a non-planar rope manipulation task called the flying knot. Using a single human demonstration and a simplified rope model\, the method learns directly on hardware without reliance on large amounts of demonstration data or massive amounts of simulation. At each iteration\, the algorithm constructs a local inverse model of the robot and rope by solving a quadratic program to propagate task-space errors into action updates. We evaluate performance across 7 different kinds of ropes\, including chain\, latex surgical tubing\, and braided and twisted ropes\, ranging in thicknesses of 7-25mm and densities of 0.013-0.5 kg/m. Learning achieves a 100\% success rate within 10 trials on all ropes. Furthermore\, the method can successfully transfer between most rope types in approximately 2-5 trials. Project page: https://flying-knots.github.io\n \nCommittee:\nChris Atkeson (advisor)\nOliver Kroemer\nJeff Ichnowski\nArun Bishop
URL:https://www.ri.cmu.edu/event/learning-dynamic-rope-manipulation-with-task-level-iterative-learning-control/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260413T090000
DTEND;TZID=America/New_York:20260413T110000
DTSTAMP:20260922T134535
CREATED:20260407T193258Z
LAST-MODIFIED:20260407T193258Z
UID:150904-1776070800-1776078000@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Gaurav Parmar
DESCRIPTION:Who: Gaurav Parmar\nDate: 13 April 2026\nTime: 9:00 a.m. (ET)\nLocation:  NSH 3305\nZoom Link: Link\nType: Ph.D. Thesis Proposal\n\n\nTitle: Efficient and Controllable Diffusion Models\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\nAbstract:  \nGenerative models have made rapid advancements in recent years\, and text-conditioned diffusion models have become the standard paradigm for image and video synthesis. \nHowever\, controlling the output purely through text is not the ideal medium for many practical applications. In my thesis research\, I have focused on making diffusion models more controllable and efficient beyond text-only interaction. To this end\, I explore two complementary directions. \nPart A: \nI study methods for generating more expressive and diverse outputs. \nIn my first project\, I explore how users can specify desired output images through visual prompts rather than text. \nThis enables compositional image generation where object identities and appearances are controlled through reference images. \nIn my second project\, I address the issue of redundant generations when the task involves generating a group of images from the same text prompt. \nPart B: \nI study efficient\, structure preserving translations with diffusion models. \nIn my first project\, I explore zero-shot image editing method that shows how well-trained text-to-image models can be repurposed for editing real images. \nHowever\, such zero-shot methods are slow and struggle when the base text-to-image model is not trained for the target domain. \nIn my next project\, I explore how we can fine-tune text-to-image models to perform image translation in both paired and unpaired settings. \nIn my final project\, I will show how we can extend this to video-to-video translation\, and the unique challenges involved. \n \n \nThesis Committee: \nJun-Yan Zhu (Co-Chair) \nSrinivasa Narasimhan (Co-Chair) \nShubham Tulsiani \nDaniel Cohen-Or (Tel Aviv University) \n\n\n\n\n\n\n\n\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-gaurav-parmar/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260413T123000
DTEND;TZID=America/New_York:20260413T143000
DTSTAMP:20260922T134535
CREATED:20260325T191648Z
LAST-MODIFIED:20260403T192753Z
UID:150742-1776083400-1776090600@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - Tairan He
DESCRIPTION:Who: Tairan He\nDate: Monday\, April 13\, 2026\nTime: 12:30 PM ET\nRoom: GHC 8102\nZoom link\n\n\nTitle: Scalable Sim-to-Real Learning for General-Purpose Humanoid Skills \n\nAbstract:\n\nHumanoid robots are compelling because they share the physical interface of the human world\, including stairs\, doors\, tools\, shelves\, and workspaces designed around human embodiment. However\, this same embodiment also makes humanoids uniquely difficult to control\, since locomotion\, balance\, manipulation\, perception\, and hardware reliability are tightly coupled. This thesis studies how to build general-purpose humanoid skills through scalable sim-to-real learning.\n\nThe dissertation develops this theme in three parts. First\, it establishes the motor foundation for sim-to-real whole-body humanoid control\, including teleoperation\, dexterous loco-manipulation\, generalist whole-body control\, and explicit sim-to-real alignment for agile behaviors. Second\, it shows that robust mobility in unstructured environments requires perception-aware control rather than blind locomotion. Third\, it extends this framework to perceptive loco-manipulation\, demonstrating that large-scale visual sim-to-real learning can enable onboard RGB-based humanoid policies for navigation\, object transport\, placement\, and articulated interactions such as door opening.\n\nTaken together\, these contributions suggest a scalable recipe for general-purpose humanoid skills in which human motion provides priors\, simulation provides scale\, control abstractions provide reuse\, alignment preserves transfer\, and perception enables embodied interaction. More broadly\, this thesis argues that the path toward useful humanoids is cumulative: general-purpose capability becomes plausible when skill acquisition\, policy design\, and deployment are treated as one continuous sim-to-real systems problem.\n\n \n \nThesis Committee Members:\n\nGuanya Shi (co-chair)\nChangliu Liu (co-chair)\nKris Kitani\nMarco Hutter (ETH Zurich)\nPieter Abbeel (UC Berkeley)\n\n \n \nA draft of the thesis document is available at this link.
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-tairan-he/
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260413T151500
DTEND;TZID=America/New_York:20260413T170000
DTSTAMP:20260922T134535
CREATED:20260407T184256Z
LAST-MODIFIED:20260407T184256Z
UID:150900-1776093300-1776099600@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Xinyu (Rachel) Li
DESCRIPTION:Date: April 13\, 2026\nTime: 03:15 PM (ET)\nLocation: GHC 6121\nZoom Link\n\nType: Ph.D. Thesis Proposal \n\nWho: Xinyu (Rachel) Li\n\nTitle: Towards Accessible AI Agents\n \nAbstract:\nEmpowered by large language models (LLMs)\, AI agents have shown strong potential across tasks such as general-purpose assistance\, software coding\, and scientific research. However\, their practical utility in applications involving consequential decisions such as healthcare\, remains constrained by three major challenges. \nEvaluation. Existing agent evaluations often focus on well-structured tasks and final outcomes\, failing to fully capture the complexity of real-world workflows. We propose evaluation frameworks grounded in realistic machine learning engineering workflows\, providing skill-based\, multi-artifact\, and holistic assessments that systematically evaluate the practical utility of AI agents. \nLearning. Improving LLMs for agentic use typically relies on reinforcement learning with large amounts of high-quality labeled data\, which are costly and difficult to obtain in expert domains including healthcare. To address this limitation\, we aim to develop learning frameworks that require minimal external supervision\, improving the scalability and efficiency of agent learning. \nSpecialization. AI agents typically follow a one-size-fits-all paradigm at the time of deployment\, lacking mechanisms to account for task-specific or user-specific requirements. We propose methods that enable agent specialization for downstream tasks and users\, expanding their applicability across heterogeneous deployment settings. \nThis thesis aims to make AI agents more broadly accessible and impactful in important real-world applications by enhancing their practical utility\, making them more measurable\, more capable\, and better tailored to the needs of their users and applications. \n\nLink to thesis\n\nThesis committee members:\nArtur Dubrawski (Chair)\nAndrea Bajcsy\nBarnabás Póczos\nDaniel McDuff (Google)
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-xinyu-rachel-li/
LOCATION:GHC 6121
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260414T093000
DTEND;TZID=America/New_York:20260414T113000
DTSTAMP:20260922T134535
CREATED:20260325T192351Z
LAST-MODIFIED:20260407T230724Z
UID:150745-1776159000-1776166200@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Nikhil Keetha
DESCRIPTION:Date: April 14\, 2026\nTime: 9:30 AM (ET)\nLocation: NSH 3305\nZoom Link\n\nType: RI Ph.D. Thesis Proposal \n\nWho: Nikhil Keetha\n\nTitle: Scaling Representation Learning for Spatial Intelligence via Active Implicit Memory \n \nAutonomous systems and embodied agents that operate in the physical world require spatial intelligence: the ability to perceive\, understand\, and reason about the geometric and semantic structure of scenes from visual input.\nData-driven approaches have made progress in learning generalizable representations from large corpora\, following the success of large language models; however\, two fundamental challenges remain open.\nFirst\, on the learning paradigm: self-supervised objectives such as next-frame prediction are insufficient for visual streams\, and nascent multi-view completion approaches remain fragile and prone to model collapse.\nSecond\, on the architecture: current multi-view models maintain scene representations that are either fixed-size (limiting capacity)\, linearly growing (hitting compute walls)\, or managed by hand-crafted heuristics that reintroduce classical brittleness.\nThis thesis addresses both challenges through a unified framework in which bootstrapped representation learning\, active implicit memory\, and multi-task decoding are tightly coupled and mutually reinforcing.The first half of this thesis establishes the foundations through three completed contributions with increasing scope.\nAnyLoc demonstrates that foundation model features provide a universal substrate for visual place recognition across diverse environments without task-specific training\, establishing the power of data-driven representation learning.\nSplaTAM introduces covisibility-guided differentiable rendering for dense visual SLAM\, providing the memory management concepts and self-supervised rendering signal that the proposed work extends.\nMapAnything presents a unified feed-forward transformer that decodes twelve sub-tasks from a single factored representation at internet scale\, validating that task and modeling paradigm are orthogonal. \nThe second half proposes Memor\, the core contribution of this thesis.\nMemor introduces active implicit memory whose capacity adapts to the complexity of the input stream\, and unified multi-task decoding that recovers diverse spatial outputs from the resulting shared representation.\nMemory management is learned end-to-end rather than governed by hand-crafted rules\, and the diversity of decoded tasks jointly enriches the underlying representation.\nTwo application chapters extend Memor: one to holistic scene generation\, where supervised pre-training prevents the model collapse that limits purely self-supervised approaches; the other to open-world semantic exploration\, where the memory provides spatial context for language-guided navigation. \nTogether\, these contributions advance a unified approach to scaling representation learning for a wide breadth of spatial intelligence.\n\nThesis Committee Members:\nSebastian Scherer (Co-Chair)\nDeva Ramanan (Co-Chair)\nShubham Tulsiani\nPeter Kontschieder\, Meta Reality Labs\n \n\n\n\n \nLink to Thesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-nikhil-keetha/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260414T121500
DTEND;TZID=America/New_York:20260414T134500
DTSTAMP:20260922T134535
CREATED:20260402T155009Z
LAST-MODIFIED:20260402T155009Z
UID:150877-1776168900-1776174300@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - Rishi Veerapaneni
DESCRIPTION:Date: April 14th\, 2026\nTime: 12:15 PM (ET)\nLocation: NSH 3305\nZoom Link\nType: Ph.D. Thesis Defense\nWho: Rishi Veerapaneni\nTitle: Efficient Multi-Agent Motion Planning using Local Policies\n \nAbstract:\nMy thesis is motivated by a future full of multi-agent systems: robots running warehouses\, humanoids constructing buildings\, aerial and ground robots delivering packages in urban environments\, and autonomous rovers collaborating on the moon. One core problem in these multi-robot teams is effective multi-agent motion planning; robots in teams need to be able to efficiently find collision-free paths in order to complete their tasks. Multi-agent motion planning is challenging as agents may need to act non-greedily and as the number of possible solutions grows exponentially with the number of agents.\n\nMulti-Agent Path Finding (MAPF) is a subset of multi-agent motion planning that typically focuses on simple planar robots but can be generalized to other robot systems. The majority of modern MAPF methods compute complete start-goal (global) paths. However computing partial (local) paths offers several advantages including faster planning\, adaptability to changes\, and compatibility with decentralized systems. Additionally\, local planning makes it easier to incorporate machine learning as it makes the prediction problems more tractable and reduces data collection requirements. Despite these benefits\, local planning is challenging as it is likely to get stuck in livelock or deadlock. Thus\, the main objective of my thesis is to effectively solve global multi-agent motion planning leveraging local policies\, learned or planned\, while avoiding the typical pitfalls of local reasoning. \nThe first part of this talk will focus on demonstrating how augmenting local ML policies with heuristic search can dramatically improve scalability compared to just ML policies by themselves. The second part of the talk will describe different approaches for maintaining solution guarantees (i.e.\, not getting stuck in deadlock) with partial planning\, including using a learned ML policy. My last part of my talk will focus on generalizing MAPF techniques to robots running heterogeneous/independent motion planners (e.g.\, A*\, RRT\, optimization\, diffusion\, and RL) and can enable coordinating a team of quadruped robots.\n\nThesis Committee Members:\nMaxim Likhachev (co-chair)\nJiaoyang Li (co-chair)\nGuanya Shi\nMac Schwager (Stanford)\nVijay Kumar (University of Pennsylvania)\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-rishi-veerapaneni-2/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260415T090000
DTEND;TZID=America/New_York:20260415T103000
DTSTAMP:20260922T134535
CREATED:20260325T193218Z
LAST-MODIFIED:20260406T160421Z
UID:150749-1776243600-1776249000@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - JJ (Jeong Hun) Lee
DESCRIPTION:Date: April 15th\, 2026\nTime: 9:00 am (EST)\nRoom Location: NSH 3305\nZoom link\n\nType: Ph.D. Thesis Defense\nWho: Jeong Hun (JJ) Lee\nTitle: Just Keep Swimming: Harnessing Fluid-Robot Interaction for Bioinspired Aquatic Locomotion\n\n\nAbstract:\nMatching the swimming efficiency and agility of fish has remained an elusive goal in underwater robotics. Such locomotion capabilities require complex vortex interactions between the robot’s body and the surrounding fluid. Despite significant advances in biomimetic hardware\, realizing successful robot policies faces several additional challenges\, which include: the need for accurate and efficient simulation of fluid-robot multiphysics; utilizing the simulators effectively for downstream policy design; and successfully transferring policies from simulation to real hardware.\n\n\nThis thesis addresses these challenges by developing a differentiable fluid–robot interaction simulator and applying it to achieve aquatic locomotion. We first present a novel multiphysics framework that solves the strongly coupled fluid–robot dynamics as a single\, differentiable optimization problem. We then exploit the simulator’s differentiability to optimize control trajectories for a robotic eel\, achieving various swimming behaviors\, including steady undulatory locomotion and a highly dynamic C-start escape maneuver. To validate the simulation results\, we design and fabricate an accessible\, from-scratch hardware platform. This culminates in successful bioinspired robotic swimming in the real world.\n\nTogether\, these contributions advance the state of the art in simulation of underwater robots while providing a new sim-to-real platform for future research in bioinspired aquatic locomotion.\n\n\nCommittee Members:\n\nZachary Manchester\, chair\nCarmel Majidi\nKeenan Crane\nRobert Katzschmann\, ETH Zurich\n\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-jj-jeong-hun-lee/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260416T100000
DTEND;TZID=America/New_York:20260416T113000
DTSTAMP:20260922T134535
CREATED:20260409T191433Z
LAST-MODIFIED:20260409T191433Z
UID:150966-1776333600-1776339000@www.ri.cmu.edu
SUMMARY:Doppler Velocity Imaging Sonar
DESCRIPTION:Abstract:\nUnderwater robotics applications require accurate velocity sensing to enable long-term dead reckoning navigation in the absence of GPS or visual features. Velocity is typically measured with a Doppler Velocity Log\, which measures the Doppler frequency shift induced by the motion of the sensor along four beams\, from which the sensor velocity can be recovered. We expand upon this idea with the proposed Doppler Velocity Imaging Sonar (DVIS) device: a 3D multibeam imaging sonar that measures Doppler velocity in addition to range. It consists of a single omnidirectional transmitter that ensonifies a large portion of the scene\, and a 2D grid of receivers that collect echos. Digital beamforming techniques are applied to produce a 3D point cloud augmented with per-point Doppler velocity. We design and construct a proof of concept device and deploy it on an underwater vehicle for validation. The large number of Doppler velocity measurements should enable more accurate and robust vehicle velocity estimates in challenging underwater environments.\n\nCommittee:\nMichael Kaess (advisor)\nDavid Wettergreen\nEaston Potokar
URL:https://www.ri.cmu.edu/event/doppler-velocity-imaging-sonar/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260416T120000
DTEND;TZID=America/New_York:20260416T130000
DTSTAMP:20260922T134535
CREATED:20260409T191652Z
LAST-MODIFIED:20260409T191652Z
UID:150968-1776340800-1776344400@www.ri.cmu.edu
SUMMARY:Learned Metrics-Aware Covariance for Visual-Inertial Fusion
DESCRIPTION:Abstract:\nVisual-inertial state estimation integrates cameras and inertial measurement units (IMUs) to achieve accurate\, metric-scale state estimation for autonomous systems. The covariance matrices associated with visual and inertial measurements determine how the estimator weights each sensing modality\, making correct covariance modeling critical for fusion accuracy and consistency. However\, most existing VI state estimators rely on constant or heuristically tuned covariance parameters that fail to capture the observation-dependent nature of real sensor uncertainty. As a result\, they require tedious manual calibration for each new platform and environment and often yield suboptimal performance.\n\nThis thesis extends and improves learned metrics-aware covariance models (observation-dependent; uncertainty follows the error in the same physical scale) of both modalities for principled\, tuning-free visual-inertial fusion. We present two systems (MAC-I² and MAC-VIO) that leverage these covariance models.\n\nMAC-I² performs visual-inertial initialization and extrinsic calibration by fusing visual pose covariance with learned inertial covariance in a multi-frame optimization. MAC-VIO utilizes metrics-aware covariance to continuous visual-inertial odometry\, performing two-frame optimization with feature-level visual residuals and IMU preintegration constraints. Experiments show that our systems achieve robust and accurate state estimation\, even in challenging scenarios involving illumination changes\, dynamic objects\, and occlusions.\n\nCommittee Members:\nProf. Howie Choset (advisor)\nProf. Sebastian Scherer\nShibo Zhao
URL:https://www.ri.cmu.edu/event/learned-metrics-aware-covariance-for-visual-inertial-fusion/
LOCATION:GHC 6501
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260416T123000
DTEND;TZID=America/New_York:20260416T140000
DTSTAMP:20260922T134535
CREATED:20260410T144641Z
LAST-MODIFIED:20260410T144641Z
UID:150977-1776342600-1776348000@www.ri.cmu.edu
SUMMARY:Observational Study to Inform Wound Care Robotics Design
DESCRIPTION:Abstract:\nThe integration of robotics and assistive technology into wound care offers a promising solution to growing patient demand amid a global nursing shortage. While assistive technologies\, including robotics for dressing removal and AI for wound measurement\, have shown promise for isolated tasks\, research has yet to evaluate nurse wound care practices from a technology design perspective. Specifically\, where technology can most effectively support nurses\, what manipulation capabilities care tasks require\, and what design considerations emerge from clinical practice. To address this gap\, we conducted an observational field study in which we observed nurses perform 86 wound care sessions across a hospital and an assisted living facility. Using a mixed-methods approach\, we collected quantitative time-motion data alongside qualitative observations from each session. Through thematic analysis\, we define and characterize all observed tasks\, organizing them within a taxonomy that differentiates physical from social tasks and identifies associated interaction subjects (e.g.\, patients\, visitors\, aides). We report the time demands of each task category and the frequency with which nurses requested assistance across tasks. We then examine manipulation approaches across care tasks\, characterizing grasping strategies\, contact regions\, bimanual and mobile manipulation\, common motions\, and workspace use. For each area\, we identify dominant techniques\, describe the clinical context driving their use\, and highlight how approaches vary across tasks and interaction subjects. Building on these findings\, we present design considerations for assistive technology in wound care\, including tasks that serve as promising entry points for technological innovation. To illustrate how our findings translate into robotic design decisions\, we fabricate a novel end effector for wound dressing. Together\, these findings offer a foundation for developing robotic and assistive systems that align with the manipulation needs and workflow realities of wound care.\n\nCommittee:\nHenny Admoni\nZackory Erickson\nZeynep Temel\nZilin Si
URL:https://www.ri.cmu.edu/event/observational-study-to-inform-wound-care-robotics-design/
LOCATION:GHC 4405
CATEGORIES:MSR Thesis Presentation,PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260416T130000
DTEND;TZID=America/New_York:20260416T143000
DTSTAMP:20260922T134535
CREATED:20260409T191241Z
LAST-MODIFIED:20260409T191241Z
UID:150964-1776344400-1776349800@www.ri.cmu.edu
SUMMARY:Pairwise 3D Human Object Contact Estimation
DESCRIPTION:Abstract:\n \nUnderstanding real-world human-object interactions in images is an inherently many-to-many problem\, where disentangling fine-grained and concurrent physical contacts is particularly challenging. Existing semantic contact estimation methods are either limited to single-human settings or require object geometry (e.g.\, meshes) in addition to the input image. Current state-of-the-art method leverages a powerful VLM for category-level semantics\, but it still struggles in multi-human scenes and scales poorly at inference time.\n\nWe introduce Pi-HOC\, a single-pass\, instance-aware framework for dense 3D semantic contact prediction across all human-object pairs. Given an input image\, Pi-HOC detects human and object instances\, enumerates all human-object pairs\, and represents each pair with a dedicated human-object (HO) token. An InteractionFormer jointly refines HO tokens and image patch features to produce interaction-aware pair representations. A SAM-based contact decoder then predicts dense contact on SMPL human meshes for each pair. On the MMHOI and DAMON datasets\, Pi-HOC significantly improves accuracy and localization over state-of-the-art methods while achieving 20x higher throughput. These results establish Pi-HOC as an efficient and scalable solution for dense semantic contact reasoning in complex scenes.\n\nWe further show that the predicted contacts improve SAM-3D image-to-mesh reconstruction through a test-time optimization procedure and enable referential contact prediction from language queries without additional training.\n\nCommittee Members:\nDong Huang (advisor)\nFernando De La Torre\nJi Zhang\nAyush Jain
URL:https://www.ri.cmu.edu/event/pairwise-3d-human-object-contact-estimation/
LOCATION:Gates Hillman Center 6115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260416T140000
DTEND;TZID=America/New_York:20260416T153000
DTSTAMP:20260922T134535
CREATED:20260409T164517Z
LAST-MODIFIED:20260409T164517Z
UID:150945-1776348000-1776353400@www.ri.cmu.edu
SUMMARY:PhD Speaking Qualifier - Ce Zhang
DESCRIPTION:TBD – This information will be provided soon.
URL:https://www.ri.cmu.edu/event/phd-speaking-qualifier-ce-zhang/
LOCATION:GHC 6121
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260416T150000
DTEND;TZID=America/New_York:20260416T160000
DTSTAMP:20260922T134535
CREATED:20260409T191950Z
LAST-MODIFIED:20260409T191950Z
UID:150970-1776351600-1776355200@www.ri.cmu.edu
SUMMARY:Learning Generalizable Robot Skills from Diverse Data Sources and Modalities
DESCRIPTION:Abstract:\nRobust robot behavior in real-world environments requires generalization across diverse objects\, scenes\, and embodiments despite limited training data. This thesis studies how different sources and modalities of data can improve different forms of robot generalization. It explores three complementary directions: force information for object-level generalization in contact-rich manipulation\, human demonstration data for environment- and embodiment-level generalization\, and large-scale simulation for environment-level generalization in reactive motion generation. \nFirst\, this thesis presents FACTR\, a force-aware imitation learning framework that combines a low-cost bilateral teleoperation system with a force-attending curriculum\, improving generalization to unseen objects in contact-rich tasks. Second\, it presents DexWild\, a scalable framework that uses co-training on human and robot demonstrations to improve generalization to unseen environments while reducing robot-specific data requirements and supporting transfer across embodiments. Third\, it presents Deep Reactive Policy (DRP)\, a simulation-trained framework for reactive motion generation that transfers zero-shot to the real world and achieves robust performance in complex dynamic environments. Together\, these results show leveraging various sources and modalities of data leads to policies exhibiting different modes of generalization. \n \nCommittee:\nProf. Deepak Pathak (co-chair)\nProf. Ruslan Salakhutdinov (co-chair)\nProf. Katerina Fragkiadaki\nKenneth Shaw
URL:https://www.ri.cmu.edu/event/learning-generalizable-robot-skills-from-diverse-data-sources-and-modalities/
LOCATION:Gates Hillman Center 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260417T113000
DTEND;TZID=America/New_York:20260417T130000
DTSTAMP:20260922T134535
CREATED:20260413T154157Z
LAST-MODIFIED:20260413T154157Z
UID:150991-1776425400-1776430800@www.ri.cmu.edu
SUMMARY:Robust Zero-Shot Recognition via Spectral Calibration
DESCRIPTION:Abstract:\nZero-shot vision-language models such as CLIP have demonstrated remarkable recognition capabilities without task-specific training. However\, their text embeddings are derived from large-scale pretraining over vast domains\, which can introduce spurious correlations that hurt performance on minority groups when sensitive attributes are entangled with target classes. Correcting this bias is challenging in realistic settings\, where biases are numerous and overlapping\, and explicit annotations of spurious attributes are unavailable. \nThis thesis presents SPOT (Spectral Preconditioning of Text Embeddings)\, a post-training calibration method that adapts CLIP’s text axes to a target domain through a single closed-form spectral transformation. SPOT estimates the covariance of unlabeled target-domain image embeddings\, decomposes the text axis in the resulting eigenbasis\, and applies an anisotropic shrinkage that preserves energy where class semantics concentrate while attenuating directions aligned with spurious correlations. I also present DP-SPOT\, a differentially private extension that derives sensitivity bounds for the SPOT map and releases calibrated axes under formal guarantees\, enabling deployment in privacy-constrained settings. \nExperiments across multiple group-robustness benchmarks show that SPOT matches or approaches state-of-the-art debiasing methods on binary classification tasks and substantially outperforms them in multi-class settings. DP-SPOT achieves near-parity with its non-private counterpart under strict privacy budgets\, while all competing baselines lack any privacy guarantee. \nCommittee:\nFernando De la Torre (advisor)\nArtur Dubrawski\nYinong Wang
URL:https://www.ri.cmu.edu/event/robust-zero-shot-recognition-via-spectral-calibration/
LOCATION:GHC 6501
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260417T130000
DTEND;TZID=America/New_York:20260417T143000
DTSTAMP:20260922T134535
CREATED:20260410T192550Z
LAST-MODIFIED:20260410T192550Z
UID:150988-1776430800-1776436200@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Hanzhe Hu
DESCRIPTION:Date:  April 17\, 2026\nTime:  1 PM-2:30 PM\nLocation: NSH 3305\nZoom link\nType: RI PhD Thesis Proposal\nWho: Hanzhe Hu\n\nTitle: Learning to Create 3D Worlds via Multi-View Generation\n\n\nAbstract:\nA 3D world is a visual representation that can be rendered from any viewpoint at any moment in time. Creating such representations from minimal input — a single image\, a text prompt\, or a monocular video is a fundamental goal in computer vision and graphics. An emerging and promising alternative is multi-view generation. However\, multi-view generation introduces its own challenges: maintaining geometric consistency across views\, achieving practical inference speed\, and extending to dynamic scenes. This thesis addresses these three challenges and presents a path toward creating 3D worlds via multi-view generation. \nWe first present MVD-Fusion\, which tackles consistency by introducing depth-guided cross-view attention for multi-view RGB-D generation from a single image. Intermediate depth estimates enable reprojection-based feature aggregation\, enforcing geometric consistency and yielding direct 3D reconstruction without costly optimization. We then address efficiency with Turbo3D\, which generates 3D Gaussian Splatting assets from text in under one second. A dual-teacher distillation framework compresses a multi-step multi-view diffusion model into a 4-step generator\, while a latent-space reconstructor eliminates image decoding overhead. Finally\, we tackle dynamics with GeoVideo4D\, a framework for camera-controllable multi-view video generation that simultaneously produces synchronized RGB videos and aligned depth maps through a joint video diffusion process\, with a hybrid training strategy unifying static 3D\, monocular video\, and multi-view video data. \nLooking ahead\, we outline two directions. First\, unifying 3D reconstruction and generation by jointly training both tasks in a single model where cameras are learned in a self-supervised manner\, enabling training on large-scale unannotated data. Second\, extending video generation to long temporal horizons to support sustained\, coherent 3D world generation beyond the short clips produced by current methods. \n\nThesis Committee:\nShubham Tulsiani (chair)\,\nDeva Ramanan\,\nJun-Yan Zhu\,\nJiajun Wu (Stanford University)\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-hanzhe-hu/
LOCATION:NSH 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260420T093000
DTEND;TZID=America/New_York:20260420T110000
DTSTAMP:20260922T134535
CREATED:20260410T185351Z
LAST-MODIFIED:20260410T185455Z
UID:150981-1776677400-1776682800@www.ri.cmu.edu
SUMMARY:RI PhD Thesis - Gokul Swamy
DESCRIPTION:Date: April 20\, 2026\nTime: 9:30-11 AM\nLocation: NSH 3305\nZoom Link\n\n\nType: Ph.D. Thesis Defense \n\n\nWho: Gokul Swamy \n\n\nTitle: Efficient Interactive Learning: Learning More From Less\n\nAbstract: Even as we near the limits of what human-generated data we can scrape from the Internet\, today’s decision-making agents — from robots to large language models (LLMs) — are still far from perfect. Thus\, as we start to move past the era of simply scaling up training datasets\, I believe the most pressing question in decision-making is how we can learn more from less data.\n\nIn theory\, agents collecting their own data and learning from this experience might allow us to transcend the limits of static datasets. However\, there are two core challenges that make delivering on this promise of reinforcement learning (RL) practically challenging. The first is exploration: experiencing the right things. The second is specification: knowing if what you experienced was good or bad. My thesis focuses on algorithms that address both of these challenges in tandem.\nThis talk will cover both theoretical advancements and their practical implications. First\, I will discuss how we can teach robots to provably recover from their own mistakes without extensive trial-and-error exploration. Second\, I will describe provably robust algorithms for training language models from conflicting preferences. Third\, I will explain how RL learns more from less without having to create data ex nihilo. To conclude\, I will outline what I believe are the most promising directions to enable the next generation of agents to learn even more from less.\n\n\nCommittee:\n\nJ. Andrew Bagnell (co-chair)\, Carnegie Mellon University\nZhiwei Steven Wu (co-chair)\, Carnegie Mellon University\nGeoffrey J. Gordon\, Carnegie Mellon University\nArthur Gretton\, University College London\n\n\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-gokul-swamy/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260420T130000
DTEND;TZID=America/New_York:20260420T143000
DTSTAMP:20260922T134535
CREATED:20260409T150820Z
LAST-MODIFIED:20260410T193548Z
UID:150943-1776690000-1776695400@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - Cornelia Bauer
DESCRIPTION:Who: Cornelia Bauer\nDate: April 20\, 2026\nTime: 1:00 PM\nLocation: NSH 4305\nZoom Link: here\nType: Ph.D. Thesis Defense \nTitle: Embracing Contact: Leveraging Physical Interactions for Enhanced Robotic Control\n\n \nAbstract:\nAs robotic systems become increasingly capable and commercially available\, particularly in the form of humanoids and dexterous hands\, enabling them to physically interact with their environment remains a fundamental challenge. Humans effortlessly use contact with their surroundings to perform agile movements and manipulate both delicate and heavy objects.\nFor robots\, in contrast\, effectively leveraging physical contact\, whether through full-body interactions for enhanced agility\, or towards contact-rich manipulation with robot hands\, is still a complex and largely unsolved problem. \nThis thesis addresses key challenges in learning and control for physical interaction for agile robots and dexterous manipulation. Specifically\, it investigates how to effectively capture human demonstrations of contact-rich and dynamic tasks and how to translate them into robot capabilities. Analysis reveals that relatively simple models can sufficiently represent human dynamic interactions at an abstract level (e.g.\, hand motion and contact forces relative to the center of mass). Combined with reflex-like controllers\, these simple models can be used to recreate dynamic physical interaction behaviors in robots. \nBuilding on these insights\, this thesis extends these principles to dexterous manipulation with soft robotic hands. Just as the human body leverages compliance to safely and effectively interact with the environment\, soft hands offer compliance and robustness through their inherent material flexibility. However\, unlike the human hand\, which relies on rich multimodal sensing\, soft robot hands are usually limited by a lack of reliable sensing. To address this\, we first develop a multimodal sensing and learning approach for tendon-driven soft fingers\, combining actuation-side and embedded sensing to predict joint angles\, contact forces\, and contact locations. Through systematic ablations\, we show that accurate state estimation can be achieved with minimal sensing\, and that tendon-force measurements alone provide a strong signal for interaction understanding. \nWe then move toward a fully sensorless paradigm\, introducing a learning-based framework that infers hand state and contact events directly from motor-side signals only. By combining a learned underactuation model with a multi-task temporal network\, the approach predicts joint configurations\, contact forces\, and contact locations without any sensors on the hand\, and generalizes across fingers via zero-shot transfer. \nFinally\, these models are integrated into a teleoperation system with haptic feedback\, enabling users to perceive contact and grasp events without visual input. More broadly\, this work demonstrates that accurate perception of contact-rich interactions can emerge from actuation signals alone\, reducing the need for complex and fragile sensing hardware. This opens a path toward simpler\, more robust\, and scalable robotic systems capable of operating in unstructured\, real-world environments\, where reliable sensing remains a key bottleneck in order to embrace contact.\n \n\nThesis Committee Members:\nNancy Pollard (Chair)\nOliver Kroemer\nZackory Erickson\nJoohyung Kim (University of Illinois Urbana-Champaign)\n \nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-cornelia-bauer/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260420T150000
DTEND;TZID=America/New_York:20260420T163000
DTSTAMP:20260922T134535
CREATED:20260413T192430Z
LAST-MODIFIED:20260413T192430Z
UID:150995-1776697200-1776702600@www.ri.cmu.edu
SUMMARY:Geometry-Guided Perception for Rapid Prototyping Systems
DESCRIPTION:Abstract: \nRapid prototyping systems demand both adaptability to new components and precision in physical interaction. In this thesis\, we explore how geometry can be used in two complementary ways: as a prompt for detecting novel industrial objects\, and to support fine-grained perception in precision tasks. \nCAD-Prompted SAM3 enables instance segmentation directly from geometric specification\, supporting detection of unseen objects without object-specific training. A geometry-driven manipulation pipeline extends this representation to action\, integrating segmentation\, pose estimation\, and grasp generation for zero-shot robotic manipulation. Eye-in-Finger further leverages task-specific geometry with tool-integrated sensing\, enabling sub-millimeter perception accuracy in cluttered assembly tasks. \nTogether\, these works demonstrate how geometry-guided perception provides a scalable approach for rapid prototyping\, enabling systems to both adapt to newly introduced components and achieve the precision required for reliable assembly. \nCommittee:\nChangliu Liu (advisor)\nOliver Kroemer\nRuixuan Liu
URL:https://www.ri.cmu.edu/event/geometry-guided-perception-for-rapid-prototyping-systems/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260421T090000
DTEND;TZID=America/New_York:20260421T100000
DTSTAMP:20260922T134535
CREATED:20260415T141427Z
LAST-MODIFIED:20260415T141427Z
UID:151028-1776762000-1776765600@www.ri.cmu.edu
SUMMARY:Design\, Implementation\, and Validation of a State Estimator for the MoonRanger Lunar Rover
DESCRIPTION:Abstract:\nThe MoonRanger Lunar rover will demonstrate continuous\, on-board\, long-range autonomous navigation at the Moon’s South Pole. In this thesis\, I describe how MoonRanger estimates its position and orientation using input from several sensors: Inertial Measurement Unit (IMU)\, Sun Sensor\, wheel encoders\, and cameras. I thoroughly review the state estimation approaches for prior space rovers and rotorcraft. MoonRanger’s state estimator performs visual-wheel-inertial odometry\, the core of which is an attitude estimator. I derive the state estimator using detailed sensor models\, accounting for and modeling the many error sources corrupting the sensors’ measurements. Next\, I describe my custom simulation infrastructure\, which enables rapid iteration\, testing\, and Monte Carlo analyses of the algorithms. I use this simulator to validate my design choices through ablation studies. Lastly\, I prove the estimator’s effectiveness through physical tests. \nI cover several implementation details important for making the software computationally efficient\, memory-safe\, debug-able\, and robust to uncertainties. I provide a thorough background of 1) the mathematical fundamentals for state estimation\, 2) IMU noise characterization and 3) an in-field method for calibrating an IMU’s pitch to reduce elevation drift. These algorithms\, analyses\, and implementation details can serve as a guide to developing state estimators for future flight rovers on the Moon and beyond. \n\n\nCommittee:\nDavid Wettergreen\nRed Whittaker\nMichael Kaess\nEaston Potokar
URL:https://www.ri.cmu.edu/event/design-implementation-and-validation-of-a-state-estimator-for-the-moonranger-lunar-rover/
LOCATION:Gates Hillman Center 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260421T100000
DTEND;TZID=America/New_York:20260421T110000
DTSTAMP:20260922T134535
CREATED:20260414T191906Z
LAST-MODIFIED:20260414T191918Z
UID:151021-1776765600-1776769200@www.ri.cmu.edu
SUMMARY:Towards Scalable Real2Sim and Evaluations for VLAs
DESCRIPTION:Abstract:\nThe evaluation of generalist robot policies\, particularly Vision-Language-Action (VLA) models\, is limited by the cost and scalability of real-world testing. This thesis introduces RobotArena ∞\, a scalable benchmarking framework for large-scale simulation-based evaluation. Central to RobotArena ∞ is a fully automated reality-to-simulation (Real2Sim) pipeline that converts monocular video demonstrations from datasets such as Bridge\, DROID\, and RH20T into high-fidelity simulated environments. The pipeline integrates automated robot-camera calibration\, 3D asset reconstruction\, and system identification to align simulated dynamics with real-world behavior. To assess policy robustness\, RobotArena ∞ applies controlled domain perturbations\, including variations in background textures and object configurations. Policy performance is evaluated through two complementary mechanisms. First\, automated VLM-guided scoring leverages vision-language models to reason jointly over visual observations and simulator states\, producing structured estimates of task progress and success. Second\, scalable human preference evaluation employs crowdsourced pairwise comparisons between policy rollouts\, enabling the aggregation of human judgments into consistent global rankings of policy performance.\nWe evaluate six state-of-the-art VLAs across hundreds of environments and 8\,500+ human comparisons. Results show that while models perform well within training distributions\, they struggle under distribution shifts. Automated VLM rankings closely match human preferences\, supporting their use as scalable proxies. Existing benchmarks with limited diversity may overestimate policy capabilities. RobotArena ∞ provides a reproducible\, extensible platform for rigorous evaluation of robotic foundation models\, promoting standardized and trustworthy benchmarking for future generalist robots. Project page is available here. \nCommittee:\nDr. Katerina Fragkiadaki (co-chair)\nDr. Yonatan Bisk (co-chair)\nDr. Shubham Tulsiani\nMr. Ayush Jain
URL:https://www.ri.cmu.edu/event/towards-scalable-real2sim-and-evaluations-for-vlas/
LOCATION:GHC 6115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260421T140000
DTEND;TZID=America/New_York:20260421T153000
DTSTAMP:20260922T134535
CREATED:20260410T185801Z
LAST-MODIFIED:20260410T185801Z
UID:150984-1776780000-1776785400@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - Nupur Kumari
DESCRIPTION:Who: Nupur Kumari\nDate: April 21\, 2026\nTime: 2:00 PM\nLocation: NSH 3305 (3rd floor)\nZoom link\n\n\nType: Ph.D. Thesis Defense \n\n\n\n\nTitle: Customizing Text-to-Image Diffusion Models\n\n\n\nAbstract\nRecent advances in Generative AI highlight its growing potential to reshape content creation workflows\, with the potential to support highly personalized use cases for everyday users and independent creators. However\, realizing these applications\, requires going beyond text conditioning for which current large-scale models are typically pre-trained for\, owing to the abundance of text-annotated data on the web. In contrast\, practical applications often involve modifying existing content using multimodal cues\, where a core challenge is the scarcity of paired input–output data.\nIn this talk\, I will briefly overview my research on post-training methods that address this challenge along three directions:\n(1) Learning from few samples\, via parameter-efficient fine-tuning of the pre-trained model on a few user-provided data. This is computationally efficient but requires fine-tuning for each new task instance. \n(2) Learning from large-scale synthetic datasets\, where I propose a pipeline to create large-scale paired datasets using the capabilities of pre-trained generative models themselves to enable end-to-end training. However\, constructing such datasets requires careful curation\, filtering\, and risk becoming outdated as base pre-trained models evolve. \n(3) Learning from discriminative models\, which leverages vision–language models to evaluate task success and provide direct gradient-based feedback to the generative model. We show its potential as a scalable and robust framework for efficient customization of generative models for downstream tasks without relying on paired synthetic datasets. \n\n\nThesis Committee:\nJun-Yan Zhu (Chair)\nDeva Ramanan\nShubham Tulsiani\nPhillip Isola (MIT)\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-nupur-kumari/
LOCATION:NSH 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260422T100000
DTEND;TZID=America/New_York:20260422T230000
DTSTAMP:20260922T134535
CREATED:20260420T142325Z
LAST-MODIFIED:20260420T142325Z
UID:151077-1776852000-1776898800@www.ri.cmu.edu
SUMMARY:Creative and Reliable Exploration in Language Model Reasoning
DESCRIPTION:Abstract:\nLanguage models increasingly solve hard problems by exploring multiple reasoning paths at training and test time. But making this exploration effective requires balancing two goals that are often in tension: creativity and reliability. On the one hand\, models need to generate diverse\, non-myopic reasoning strategies rather than repeatedly sampling the same mistakes. On the other hand\, they need mechanisms to verify and refine these attempts in a trustworthy way. \nIn this talk\, I will argue that both goals require structured exploration. First\, I study creative exploration in language models\, showing that next-token prediction can be fundamentally limiting for open-ended reasoning and generation. I then show how explicitly structuring exploration across diverse reasoning modes can substantially improve test-time scaling. Finally\, I turn to self-verification for reliable exploration\, arguing that iterative improvement is most effective when verification is itself structured and informative\, rather than simply making the model think longer. I also discuss evidence from adversarial settings suggesting that reliability does not come automatically from adding more verification components. Together\, these results suggest that progress in reasoning will depend not just on more computation\, but on learning to use that computation for structured exploration that supports both creativity and reliability. \nCommittee:\nAditi Raghunathan\nShubham Tulsiani\nDaniel Fried\nZhiqiu Lin
URL:https://www.ri.cmu.edu/event/creative-and-reliable-exploration-in-language-model-reasoning/
LOCATION:Newell Simon Hall 4201
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260423T100000
DTEND;TZID=America/New_York:20260423T110000
DTSTAMP:20260922T134535
CREATED:20260420T155250Z
LAST-MODIFIED:20260423T140940Z
UID:151081-1776938400-1776942000@www.ri.cmu.edu
SUMMARY:Unified Spherical Frontend: Towards Universal Distortion-Free Lens-Agnostic Rotation-Equivariant Perception
DESCRIPTION:Abstract:\nModern perception increasingly relies on wide field-of-view cameras\, yet standard convolutional networks still operate on planar pixel grids designed for pinhole imagery. By Gauss’s Theorema Egregium\, no projection from the sphere to the plane preserves curvature\, so every planar map of a spherical signal introduces spatially-varying distortion. Models trained on one lens therefore overfit to its specific distortion and degrade sharply under camera changes or in-plane rotations. Prior remedies fall short in opposite ways: spherical harmonic CNNs recover rotation-equivariance but remain computationally infeasible at image-scale resolution\, while projection-based variants stay efficient but sacrifice equivariance. In this thesis\, we investigate a single spatial-domain primitive: geodesic-distance convolution on the unit sphere\, and its two complementary roles: unifying efficiency with equivariance for lens-agnostic 2D perception\, and extending the same geometric principle from the sphere to 3D Euclidean space.\n\nFirst\, we present the Unified Spherical Frontend (USF)\, a modular pipeline that replaces the planar frontend of any CNN with pixel-to-ray lifting\, near-uniform spherical resampling\, geodesic-distance convolution\, and spherical pooling\, while leaving the backbone untouched. Constraining kernels to depend only on geodesic distance makes the operation SO(3)-equivariant by construction\, and because all geometric quantities are camera-specific constants\, they are computed once and cached for near-zero runtime overhead. On Spherical MNIST\, PANDORA panoramic detection with YOLOv11\, and Stanford 2D-3D-S panoramic segmentation with DeepLab v3 and UNet\, USF matches or exceeds planar performance\, suffers less than 1% degradation under arbitrary SO(3) rotations without rotation augmentation\, and generalizes zero-shot to lens types unseen during training.\n\nSecond\, we extend the distance-only kernel design from the sphere S² to Euclidean ℝ³\, yielding an SE(3)-equivariant volumetric convolution for 3D point cloud processing. The same principle transfers directly\, suggesting a broader geometric primitive for equivariant learning on curved and flat manifolds alike.\n\nTogether\, these contributions position geodesic-distance aggregation as a unifying foundation for distortion-free\, lens-agnostic\, rotation-equivariant perception across 2D and 3D\, offering a principled alternative to the augmentation-heavy and spectral-transform-based approaches that dominate spherical deep learning today.\n\nCommittee:\nProf. László A. Jeni (co-chair)\nProf. Sebastian Scherer (co-chair)\nProf. Shubham Tulsiani\nMosam Dabhi
URL:https://www.ri.cmu.edu/event/unified-spherical-frontend-towards-universal-distortion-free-lens-agnostic-rotation-equivariant-perceptionunified-spherical-frontend-towards-universal-distortion-free-lens-agnostic-rotation-equivari/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260423T150000
DTEND;TZID=America/New_York:20260423T160000
DTSTAMP:20260922T134535
CREATED:20260420T155111Z
LAST-MODIFIED:20260420T155111Z
UID:151079-1776956400-1776960000@www.ri.cmu.edu
SUMMARY:Synthetic Data for Object Detection: Improving Robustness and Revealing Vulnerabilities
DESCRIPTION:Abstract:\nSynthetic data has emerged as a promising solution to the growing challenges of data acquisition in object detection. Modern detectors rely heavily on large-scale annotated datasets\, yet collecting real-world data with high-quality labels is often costly\, labor-intensive\, and impractical in diverse or rare scenarios. By enabling controllable generation with automatic annotations\, synthetic data provides a scalable alternative. In this thesis\, we investigate its two complementary roles: improving robustness under domain shifts and revealing fundamental vulnerabilities of object detection models. \n\nFirst\, we propose a synthetic data generation framework based on diffusion models to bridge distribution gaps between source and target domains in aerial imagery. By synthesizing high-quality images and corresponding annotations through cross-attention-guided labeling and multi-stage knowledge transfer\, our approach significantly improves detection robustness in unseen environments\, outperforming supervised learning on source domain data\, weakly supervised and unsupervised domain adaptation methods\, open-set object detectors\, and vision large language models.\n\nSecond\, we explore the adversarial potential of synthetic data via a controllable image-editing framework for realistic camouflage attacks. By formulating camouflaged adversarial example generation as a conditional image-editing problem\, we design image-level and scene-level strategies that produce stealthy\, physically plausible camouflages while effectively degrading detector performance. Extensive experiments demonstrate strong attack effectiveness\, improved human-perceived stealthiness\, and transferability to black-box models and to the physical world.\n\nTogether\, these complementary perspectives highlight the dual utility of synthetic data for object detection: as a powerful tool for improving robustness under domain shifts and as a principled lens for uncovering model vulnerabilities.\n\nCommittee:\nProf. Fernando De la Torre (Advisor)\nProf. Deva Kannan Ramanan\nJianjin Xu
URL:https://www.ri.cmu.edu/event/synthetic-data-for-object-detection-improving-robustness-and-revealing-vulnerabilities/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260424T110000
DTEND;TZID=America/New_York:20260424T120000
DTSTAMP:20260922T134535
CREATED:20260421T151322Z
LAST-MODIFIED:20260421T151322Z
UID:151098-1777028400-1777032000@www.ri.cmu.edu
SUMMARY:Online Policy Improvement via Reliable Critics and Deployment-Aligned Data
DESCRIPTION:Abstract:\n\nRecent progress in robotic foundation models has made broad\, reusable robot competence increasingly plausible\, but it has also brought a central challenge: how can robots continue to improve upon online deployment? This is especially acute in high-dimensional or contact-rich settings\, where reactive\, highly dexterous skills are critical but not trivial for BC priors. My thesis studies how to achieve reliable policy improvement from online experience via reinforcement learning (RL)\, grounded in the premise that scalable robotic adaptation requires reliable learning signals\, including accurate critics and distribution-aligned data curation.\nIn the first part of the talk\, I will present TD-M(PC)^2\, a model-based RL framework for high-dimensional continuous control. This work identifies a structural mismatch between the model-based exploration and the model-free critic learning as a major source of value overestimation. To address this issue\, I introduce constrained policy iteration that reduces out-of-distribution bootstrapping while preserving the advantages of test-time policy improvement via planning. The resulting method yields stable and sample-efficient learning. Second\, I present PLD (Probe\, Learn\, Distill)\, a post-training framework for vision-language-action models. PLD uses offline warm-start to initiate and stabilize online RL\, learning task-specific residual specialists\, collecting high-quality\, deployment-aligned recovery behaviors\, and distilling them back into a pretrained generalist without requiring massive human effort. This enables self-improvement while maintaining generalization across simulation and real-world dexterous manipulation. Finally\, I will discuss ongoing work on critic and reward modeling for long-horizon tasks that require precision and dexterity\, where reliable evaluation and credit assignment remain significant bottlenecks. Taken together\, the thesis suggests that progress in robotic autonomy depends not only on strong foundation priors but also on mechanisms for reliably assessing and bootstrapping experience.\n\n\nCommittee:\nProf. Guanya Shi (Co-chair)\nProf. Jeff Schneider (Co-chair)\nProf. Changliu Liu\nWenli Xiao
URL:https://www.ri.cmu.edu/event/online-policy-improvement-via-reliable-critics-and-deployment-aligned-data/
LOCATION:Gates Hillman Center 6115
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
END:VCALENDAR