BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251003T143000
DTEND;TZID=America/New_York:20251003T153000
DTSTAMP:20260922T113641
CREATED:20250902T162718Z
LAST-MODIFIED:20251006T153641Z
UID:148662-1759501800-1759505400@www.ri.cmu.edu
SUMMARY:Neural Certificates for Safe Robotic System Planning and Control
DESCRIPTION:Abstract:\nAchieving safety\, scalability\, and high performance in complex systems\, such as multi-agent systems (MAS) control\, is a central challenge in many real-world robotic deployments due to its computational complexity as a large-scale constrained optimal control problem. To address this\, we introduce a novel graph control barrier function (GCBF) as a core tool for large-scale distributed safe control\, which guarantees safety for arbitrarily large MAS with only local observations. For MAS with known dynamic models\, we present a self-supervised learning framework that can jointly learn GCBF and distributed control policies that consider actuation limits. For MAS with unknown dynamics\, we discuss how to blend GCBF in multi-agent reinforcement learning (MARL) to achieve high-performance and safe distributed policies.\n\nBio:\nChuchu Fan is an Associate Professor (pre-tenure) in the Department of Aeronautics and Astronautics (AeroAstro) and Laboratory for Information and Decision Systems (LIDS) at MIT. Before that\, she was a postdoc researcher at Caltech and got her Ph.D. at the University of Illinois at Urbana-Champaign. She earned her bachelor’s degree from Tsinghua University. Her research group\, the Realm at MIT\, works on developing computational tools that integrate rigorous mathematics into machine learning and AI for the design\, analysis\, and verification of safe\, large-scale\, and complex systems. Chuchu is the recipient of an NSF CAREER Award\, an AFOSR Young Investigator Program (YIP) Award\, an ONR YIP Award\, and the 2020 ACM Doctoral Dissertation Award.
URL:https://www.ri.cmu.edu/event/neural-certificates-for-safe-robotic-system-planning-and-controlri-seminar-w-chuchu-fan/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/ChuchuFan-042021.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251009T093000
DTEND;TZID=America/New_York:20251009T110000
DTSTAMP:20260922T113641
CREATED:20250903T175132Z
LAST-MODIFIED:20250930T131653Z
UID:148746-1760002200-1760007600@www.ri.cmu.edu
SUMMARY:Customizing Text-to-Image Diffusion Models
DESCRIPTION:Abstract: With the rapid advancement of generative models\, their potential to transform creative content creation is increasingly evident. However\, most large-scale generative models are primarily text-conditioned\, given the availability of large-scale paired text–image datasets. In contrast\, for most practical applications\, creators often begin from an existing asset and wish to generate variations or modify it in specific ways. For images\, this may involve placing an object in a new context\, adjusting local attributes\, or altering visual style. My research focuses on customizing pre-trained generative models\, primarily text-to-image diffusion models\, to facilitate such downstream tasks. A central challenge here is the lack of paired input–output data for these tasks. \nTo address this\, I explore three complementary directions: \nPart I: I study few-shot learning methods\, which are computationally efficient but require fine-tuning for each new task instance. This limitation motivates the second direction. \nPart II: Constructing synthetic paired datasets using the capabilities of pre-trained generative models themselves to train feed-forward models in a supervised manner. However\, constructing such datasets requires careful curation\, filtering\, and risk of becoming outdated as base pre-trained models evolve. Building on these insights\, my thesis proposes a third paradigm. \nPart III: Customizing generative models without paired supervision. Instead\, we plan to leverage vision–language models to evaluate task success and provide direct gradient-based feedback to the generative model. This approach has the potential to create a scalable and robust framework for efficient customization of generative models for downstream tasks without relying on synthetic datasets. \n\n\n \nThesis Committee:\nJun-Yan Zhu (Chair)\nDeva Ramanan\nShubham Tulsiani\nPhillip Isola (MIT)\n\nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/customizing-text-to-image-diffusion-models/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251010T110000
DTEND;TZID=America/New_York:20251010T130000
DTSTAMP:20260922T113641
CREATED:20250909T130917Z
LAST-MODIFIED:20250930T145809Z
UID:148785-1760094000-1760101200@www.ri.cmu.edu
SUMMARY:Pushing the Frontier of Robotic Tool Manipulation by Treating the Hand and the Tool Together as a Machine
DESCRIPTION:Abstract: Tool manipulation is an essential human skill. It expands our manipulation capability beyond the capability of the biological hand\, and is a defining feature of many tasks centered on physical interaction with the real world. For humanoid robots to become general-purpose\, they must master tool manipulation as well. However\, the state-of-the-art humanoid robots equipped with multi-finger hands still fall behind their human counterparts in tool manipulation performance. This thesis aims to narrow this gap by treating the hand and the tool together as a machine.\nSpecifically\, inspired by the analogy between multi-finger hands and CNC machines\, this thesis interprets a tool-manipulating hand as configuring itself and the tool into different tool-hand mechanisms in real time. To concretely represent each tool-hand mechanism—which consists of the tool\, the hand\, and the contacts—this thesis introduces two concepts: 1) foundational pose\, a pose and precondition that the tool and the hand must reach for the tool-hand mechanism to be successfully constructed and to run\, and a concise representation of tool-hand mechanism. 2) sub-assembly\, a set of contacts that independently fulfills part of the tool-hand mechanism’s function\, and a detailed\, modular representation of tool-hand mechanism. \nThis thesis first tests the validity of the concept of foundational pose via the question: “if a tool and a hand have reached a foundational pose\, can they act as the corresponding tool-hand mechanism and perform the tool manipulation motion?” To answer this question\, the thesis conducts a hand design experiment\, which uses foundational poses as constraints to sample many different hands and evaluates their tool manipulation motions. The results lead to a positive answer to the question\, verifying the concept of foundational pose. \n\nThen\, this thesis expands roll-slide contact-based tool manipulation motion planning—which previously was only possible for primitive shapes with global parametrizations—to manifold meshes\, which allows motion planning from foundational poses for arbitrarily shaped tools and hands. \nFinally\, for the proposed work\, this thesis aims to test the validity of the concept of sub-assembly via the question: “how many sub-assemblies are enough?” Based on the answer to this question\, this thesis aims to develop a sub-assembly-based control framework\, and test the framework on a real robotic hand for an entire tool manipulation sequence. \n\n \n\nThesis Committee Members: \nProf. Nancy Pollard (co-chair)\nProf. Jean Oh (co-chair)\nProf. Matthew Mason\nDr. Lael Odhner (The Robotics and AI Institute)\n\n\nDraft of the Thesis Proposal Document Link: https://drive.google.com/file/d/1L5ri7r0295poQOyTtqEOI3-LX3o7Gqlf/view?usp=drive_link
URL:https://www.ri.cmu.edu/event/pushing-the-frontier-of-robotic-tool-manipulation-by-treating-the-hand-and-the-tool-together-as-a-machine/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251010T143000
DTEND;TZID=America/New_York:20251010T153000
DTSTAMP:20260922T113641
CREATED:20250902T163332Z
LAST-MODIFIED:20251010T231714Z
UID:148665-1760106600-1760110200@www.ri.cmu.edu
SUMMARY:A Manipulation Journey
DESCRIPTION:Abstract:\nThe talk will revisit my career in manipulation research\, focusing on projects that might offer some useful lessons for others. We will start with my beginnings at the MIT AI Lab and my MS thesis\, which is still my most cited work\, then continue with my arrival at CMU\, a discussion with Allen Newell\, an exercise to envision a coherent research program\, and how that led to a second and third childhood. The talk will conclude with some discussion of lessons learned.\n\nBio:\nMatt has spent 50 years conducting research in Artificial Intelligence and Robotics\, starting as a student in the MIT AI Lab where he earned the BS\, MS\, and PhD degrees. He spent much of his career at CMU’s Robotics Institute\, where he was the founder and co-director of the Manipulation Laboratory\, and for ten years served as the Director of the Robotics Institute. Matt’s group studied the basic physics governing grasping and manipulation\, and demonstrated that sophisticated grasping and manipulation can be produced by the simple and robust grippers used in industrial automation.\n\nMason is a Fellow of the AAAI\, AAAS\, ACM\, and IEEE\, and a winner of the IEEE R&A Society’s Pioneer Award\, and the IEEE Technical Field Award in Robotics and Automation (the R&A Prize).\n\nMatt now focuses his attention on logistics and warehouse robotics\, helping to produce industry-leading solutions at Berkshire Grey.
URL:https://www.ri.cmu.edu/event/a-manipulation-journey/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/unnamed.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251010T143000
DTEND;TZID=America/New_York:20251010T163000
DTSTAMP:20260922T113641
CREATED:20250902T201509Z
LAST-MODIFIED:20251006T150600Z
UID:148687-1760106600-1760113800@www.ri.cmu.edu
SUMMARY:Generative Robotics: Self-Supervised Learning for Human-Robot Collaborative Creation
DESCRIPTION:Abstract:\nRobotic automation is generally welcomed for tasks that are dirty\, dull\, or dangerous\, but with expanding robotic capabilities\, robots are entering domains that are safe and enjoyable\, such as creative industries. Although there is a widespread rejection of automation in creative fields\, many people\, from amateurs to professionals\, would welcome supportive or collaborative creative tools. Supporting creative tasks is challenging with real-world robotics because there are limited relevant datasets\, creative tasks are abstract and high-level\, and real-world tools and materials are difficult to model and predict. Learning-based robotic intelligence is a promising method for creative support tools\, but since the task is so complex\, common approaches such as learning from demonstration would require too many samples and reinforcement learning may never converge. In this thesis\, we show that robots can learn to support acts of creativity purely through a few\, proposed self-supervised learning techniques. \nWe formalize robots that support people in the making of things from high-level goals in the real world as a new field\, Generative Robotics. We introduce an approach for supporting 2D visual art-making with paintings and drawings along with 3D clay sculpture from a fixed perspective. Because there are no robotic datasets for collaborative painting and sculpting\, we designed our approach to learn from small\, self-generated datasets to learn real-world constraints and support collaborative interactions. Our approach uses (1) Real2Sim2Real to enable a robot to teach itself about physical constraints (e.g.\, type of paint and brush)\, (2) semantic planning to plan from high-level\, abstract goals under severe real-world constraints (e.g.\, making a painting from a detailed photograph with only 4 colors and 64 brush strokes)\, and (3) self-supervised learning to generate data to train the robot to support creation rather than automate it. Our approach collaboratively creates paintings in heavily constrained settings. Lastly\, we generalize our approach to new materials\, tools\, action representations\, and state representations to perform long-horizon clay sculpting.\n\nDocument:\nhttps://drive.google.com/drive/folders/1H4-gyLhccQTtlipmBRGAHbrV3bHI7LtU?usp=sharing \n\n\nThesis Committee Members: \nJean Oh\, Chair \nJames McCann \nManuela Veloso \nKen Goldberg\, UC Berkeley
URL:https://www.ri.cmu.edu/event/generative-robotics-self-supervised-learning-for-human-robot-collaborative-creation-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251014T100000
DTEND;TZID=America/New_York:20251014T113000
DTSTAMP:20260922T113641
CREATED:20251001T150325Z
LAST-MODIFIED:20251024T233624Z
UID:148981-1760436000-1760441400@www.ri.cmu.edu
SUMMARY:Consistent Modeling of 4D Scenes for Perception and Generation
DESCRIPTION:Abstract:\n\nA core challenge in vision is building representations that capture 3D scenes over time for perception and interactive generation. For accurate perception and plausible generation\, we want consistency across views\, time\, and modalities. In this talk we explore consistency through the choice of representation\, moving from dense grid formulations to entity-centric scenes that are easier to maintain across frames\, and we extend that representation from perception to generation.\n\nOur past work follows this shift within perception tasks. SOLOFusion uses a grid representation with long- and short-baseline temporal stereo for multi-camera 3D detection\, improving foreground depth\, but it does not perform entity grouping and it does not model background. ASCFormer performs depth estimation and completion via pixel–point affinity\, grouping geometry coherently\, but the grouping is geometric rather than semantic and remains static. DetMatch\, together with our temporal follow-up\, addresses semi-supervised 2D and 3D detection\, aligning detections across modalities and video to produce consistent pseudo-labels and more stable tracklets\, but it focuses on foreground entities and does not model background. S2GO proposes a streaming query-based representation for semantic occupancy estimation that is entity-centric\, temporal\, and models both foreground and background: each persistent query decodes to semantic Gaussians\, and the state is carried across frames and supports short-horizon future prediction. This gives us a single\, stable representation suitable for both perception and sampling. \nWe propose two projects that make this representation generative. First\, we propose a static scene generation method: a diffusion model over grounded queries that represent both foreground and background and are decoded into Gaussians\, generating a complete semantic occupancy scene. This grounded latent representation enables intuitive\, consistent control. Then\, we propose motion generation: a model that generates trajectories for ego and foreground entities conditioned on the generated static scene\, producing coherent 4D rollouts and enabling interactive edits.\n \nThesis Committee Members:\nKris Kitani (Chair)\nDeva Ramanan\nShubham Tulsiani\nWei-Chiu Ma (Cornell University)\n \nLink to Proposal Draft: Link
URL:https://www.ri.cmu.edu/event/jinhyung-park-phd-thesis-proposal/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251014T140000
DTEND;TZID=America/New_York:20251014T153000
DTSTAMP:20260922T113641
CREATED:20250929T144547Z
LAST-MODIFIED:20251030T161352Z
UID:148937-1760450400-1760455800@www.ri.cmu.edu
SUMMARY:Embodied Artificial Intelligence for Emergency Care in Unstructured Environments
DESCRIPTION:Abstract: \nIn mass casualty events and resource-constrained scenarios\, limited responder capacity leads to preventable deaths. Time is of the essence particularly in severe trauma: the sooner individuals receive care\, the higher their chances of survival. Yet a single responder can only manage a few patients simultaneously\, leaving others unattended. This thesis addresses this capacity constraint by developing intelligent robotic systems that serve as medical force multipliers\, enabling effective emergency response when casualties outnumber available help. \nThis work presents two embodied Artificial Intelligence (AI) platforms for emergency medical response in unstructured field environments. The first performs multipatient automated assessment\, which uses contactless multimodal sensors to identify qualitative (e.g.\, wounds\, amputations\, hemorrhage\, respiratory distress) and quantitative vital signs (e.g.\, heart rate) to rapidly assess and prioritize the most critically injured. The second automates fluid resuscitation\, targeting hemorrhage\, the leading cause of preventable death in trauma. The pipeline comprises multiple stages: vessel localization and segmentation\, visualization and uncertainty quantification for safe decision-making\, bifurcation detection for anatomically-informed needle placement\, and real-time needle tracking. \nTo address the scarcity of training data in emergency medicine robotics\, this research embeds expert clinical knowledge and applies weak supervision techniques\, enabling robust performance with limited labeled examples. All algorithms execute in real time on resource-constrained platforms\, with key components designed to adapt to changing environmental conditions. \nThis thesis contributes to autonomous medical systems and offers new methodologies for developing AI solutions for life-critical applications in unstructured environments where traditional data-driven approaches may fail. By augmenting human responders\, we show how robotic systems can expand treatment capacity when it matters most\, potentially saving lives that would otherwise be lost. \nThesis Committee: \nArtur Dubrawski\, Chair \nJean Oh \nFernando de la Torre Frade \nDaniel McDuff\, Google \nLaura Brattain\, University of Central Florida \nDocument Link
URL:https://www.ri.cmu.edu/event/embodied-artificial-intelligence-for-emergency-care-in-unstructured-environments/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251015T130000
DTEND;TZID=America/New_York:20251015T140000
DTSTAMP:20260922T113641
CREATED:20251008T135546Z
LAST-MODIFIED:20251008T140726Z
UID:149052-1760533200-1760536800@www.ri.cmu.edu
SUMMARY:Seeing Deep Inside Scattering Tissue Using Efficient\, Noise-Robust Wavefront Shaping
DESCRIPTION:Abstract:\nScattering limits our ability to see inside biological tissue\, as light penetration is severely distorted by tissue components with varying refractive indices. One promising method to overcome scattering aberration is wavefront shaping. This technique involves placing a spatial light modulator (SLM) in the microscope’s optical path to correct the wavefront emitted from a point deep within the tissue. The goal is to bring light photons from a single target point to a single sensor point\, despite tissue aberrations. This technique has the potential to revolutionize tissue imaging by enabling high-SNR imaging deep within scattering biological targets. However\, estimating wavefront-shaping modulations in practice is challenging\, since the modulations must be estimated in real time\, using non-invasive feedback\, and under a low photon budget. \nIn the first part of this talk\, I will discuss efforts to derive noise-robust score functions that can identify effective modulation corrections using non-invasive feedback. I will review previous approaches and introduce a new\, simple\, noise-robust method that uses confocal correction of both incoming and outgoing light with linear single-photon fluorescent excitation. We show that despite the fact that we are only measuring light outside the tissue and have no direct way to measure how well light has focused inside the tissue\, maximizing the single-photon confocal intensity guarantees that we also focus all light into a spot inside the tissue. \nGiven a score function\, estimating the desired modulation becomes an optimization problem. However\, since the desired modulation depends on the unknown tissue structure\, typical optimization strategies involve slow sequential scanning\, where each modulation parameter is queried independently. In the second part of this talk\, I will present a novel approach for rapid modulation optimization. This method leverages optical computing ideas and uses the optical system to directly measure the gradient of the score function\, allowing simultaneous updates of all modulation parameters from a single measurement. \nBio: Anat Levin is a Professor at the department of Electrical and Computer Engineering\, Technion\, Israel\, doing research in the field of computational imaging. She received a Ph.D. in computer science from the Hebrew University in 2006. During the years 2007- 2009 she was a postdoc at MIT CSAIL\, and during 2009-2016 she was an Assistant and Associate Prof. at the department of Computer Science and Applied Math\, the Weizmann Inst. of Science.\nProf. Levin has received numerous awards for her research\, including the CVRP PAMI young researcher award in 2013; the eurographics young researcher award in 2010; the eurographics outstanding technical contributions award in 2024; the Blavatnik award in 2018; and 3 ERC grants. \nHomepage:  https://webee.technion.ac.il/people/anat.levin/ \n  \nSponsor \nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one AI-powered video generation platform that leverages advances in computer vision\, generative modeling\, and multimodal learning to make high-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/seeing-deep-inside-scattering-tissue-using-efficient-noise-robust-wavefront-shaping/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/10/10-15-25.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251020T150000
DTEND;TZID=America/New_York:20251020T163000
DTSTAMP:20260922T113641
CREATED:20251014T201636Z
LAST-MODIFIED:20251014T201948Z
UID:149083-1760972400-1760977800@www.ri.cmu.edu
SUMMARY:Unconstrained Perception for Scalable Robot Manipulation
DESCRIPTION:Abstract: Advances in visual imitation learning driven by large-scale data and expressive policy architectures have yielded impressive progress on long-horizon\, dexterous tasks. However\, current success rates remain insufficient for industrial deployment\, which demands near-perfect reliability on novel tasks. Compared to other fields such as NLP and CV\, the available data in robotics is several orders of magnitude smaller. This raises the question: how can we most effectively leverage priors from large-scale offline data? In this thesis\, I contribute methods to infer strong geometric and dynamic priors for robot manipulation. \nFirst\, geometric camera calibration is a critical prerequisite for real-world vision systems. I will discuss our work on MASt3R-SfM for unconstrained SfM from any image collection in linear complexity. Next\, I discuss how we use a large set of calibrated cameras in DeformGS for photorealistic digital twins with millimeter-accurate tracking of deformable cloth. Removing the need for costly multi-camera systems\, I introduce RaySt3R\, a method to generate complete object geometry from a single RGB-D image. \nBuilding on these works\, I will introduce our ongoing work on Flow2Flow – a flexible end-to-end feedforward architecture for zero-shot dynamics prediction. Many challenging tasks involve manipulating unseen articulated\, deformable\, and cluttered rigid objects; prior approaches rely on pre-trained VLMs\, large-scale 2D point tracking\, or previous interactions with the scene to inject priors. We cast dynamics prediction as a scene flow completion problem from a single RGB-D image\, and propose an optional two-stage adaptation procedure for unseen dynamics. We further study scene flow completion as a 3D pretraining objective for multi-task learning and propose scaling up training on real-world data for the first benchmark in dynamics prediction from a single image. \nThesis Committee Members:\nJeffrey Ichnowski (Chair)\nDeva Ramanan\nShubham Tulsiani\nAbhishek Gupta (University of Washington) \nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/unconstrained-perception-for-scalable-robot-manipulation/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251021T100000
DTEND;TZID=America/New_York:20251021T110000
DTSTAMP:20260922T113641
CREATED:20251007T172737Z
LAST-MODIFIED:20251007T172737Z
UID:149047-1761040800-1761044400@www.ri.cmu.edu
SUMMARY:Adaptive Robot Design for multimodal locomotion across diverse terrains
DESCRIPTION:Abstract: \nLocomotion across natural environments such as sand\, mud\, and water presents a fundamental challenge for robots due to the heterogeneous\, deformable\, and often unpredictable properties of these substrates. In this talk\, I will share how mechanical and structural adaptation can enable robust mobility in such complex settings through the development and characterization of two centimeter-scale robotic systems that leverage distinct modes of morphological adaptation. \nFirst\, TerraSkipper\, a mudskipper-inspired robot that integrates a spring-steel tail and magnetically encoded fins to achieve impulsive skipping and controlled crawling across a wide range of granular and muddy substrates. Through systematic experiments that vary substrate composition and moisture content\, we show that impulsive tail-driven locomotion achieves higher velocities and improved mobility where conventional frictional gaits fail\, providing new insights into substrate robot interactions at the centimeter scale. \nSecond\, PuffyBot\, an amphibious shape morphing robot that employs a scissor-lift mechanism\, coupled fins\, and a waterproof skin to actively modulate its volume and buoyancy. Our experimental results demonstrate multimodal locomotion\, including crawling on the land\, crawling on the underwater floor\, swimming on the water surface\, and bimodal buoyancy adjustment to submerge underwater or resurface. \nTogether\, these systems illustrate how embodied mechanical intelligence through energy-based and geometry-based adaptation extends the operational range of small robots across a wide range of terrestrial and aquatic domains. This work advances the understanding of morphology as a design variable in robot locomotion and lays the foundation for autonomous\, terrain-adaptive robotic platforms capable of robust operation in unstructured natural environments. \nCommittee: \nProf. Zeynep Temel \nProf. Sarah Bergbreiter \nProf. Guanya Shi \nRishi Veerapaneni
URL:https://www.ri.cmu.edu/event/adaptive-robot-design-for-multimodal-locomotion-across-diverse-terrains/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251024T090000
DTEND;TZID=America/New_York:20251024T110000
DTSTAMP:20260922T113641
CREATED:20251001T150706Z
LAST-MODIFIED:20251016T170428Z
UID:148983-1761296400-1761303600@www.ri.cmu.edu
SUMMARY:Title: Leveraging Geometric Priors for Robust Robotic Manipulation
DESCRIPTION:Abstract: \nThis thesis explores how explicit 3D geometric representations\, trained at scale on synthetic data\, can serve as priors to enhance robotic manipulation. Even with recent progress in geometric understanding\, generalization to unseen objects and environments remains constrained by the scale and diversity of existing 3D training data. Although more large-scale 3D datasets have been released\, their sizes are still considerably smaller than their image and language counterparts. In addition\, collecting diverse real-world 3D data is time-consuming and labor-intensive\, limiting the coverage of objects and scenes. To tackle this challenge\, this thesis explores how geometric understanding learned from large-scale synthetic 3D model datasets can improve generalization in robotic manipulation without further expanding real-world 3D training data. \nAs a first step\, Chapter 2 introduces RePOSE\, a fast and accurate 6D object pose refinement method that establishes a foundation for scalable geometric perception. Chapter 3 and Chapter 4 frame the acquisition of a universal geometric prior as a supervised learning problem on 3D geometry tasks. We propose two frameworks\, OctMAE and ZeroGrasp\, which learn a geometric prior through shape reconstruction and grasp pose prediction. We also introduce ZeroGrasp-11B\, a large-scale synthetic dataset containing 1M RGB-D images\, 12K 3D models\, and 11B grasps\, specifically designed for training such models. These methods achieve state-of-the-art performance on both shape reconstruction and grasp pose prediction of unseen objects on public benchmarks\, demonstrating the strength of the learned geometric prior. Real-world pick-and-place experiments further validate its generalization to practical robotic scenarios. \nWhile the learned geometric prior shows strong performance in pick-and-place tasks\, robotic manipulation involves a broader range of behaviors and longer temporal horizons. In Chapter 5\, we focus on integrating this prior into imitation learning to address more complex\, long-horizon tasks. To this end\, we propose GeoFlow\, a framework for flow-based 3D visuomotor policy learning that leverages geometry-aware pre-trained models as strong priors. GeoFlow achieves state-of-the-art performance across diverse benchmarks and demonstrates improved data efficiency and robustness under clutter and distractors\, highlighting that large-scale geometry pre-training and sparse voxel representations are key to scalable and generalizable robotic learning. \n\n\nThesis Committee: \nKris Kitani\, Chair \nDavid Held \nShubham Tulsiani \nSergey Zakharov\, Toyota Research Institute \nDocument Link
URL:https://www.ri.cmu.edu/event/shun-iwase-phd-thesis-defense/
LOCATION:Gates Hillman Center 6115
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251024T120000
DTEND;TZID=America/New_York:20251024T130000
DTSTAMP:20260922T113641
CREATED:20251015T183546Z
LAST-MODIFIED:20251015T183546Z
UID:149109-1761307200-1761310800@www.ri.cmu.edu
SUMMARY:KOROL: Learning Visualizable Object Feature with Koopman Operator Rollout for Manipulation
DESCRIPTION:Abstract: \nHumans possess an extraordinary ability to manipulate objects\, discerning position\, shape\, and other properties with just a glance. How can robots be endowed with similar perceptual and dexterous manipulation capabilities? In this talk\, I will present a method that combines the sample efficiency of traditional model-based approaches with the high generalizability of deep learning methods to tackle dexterous manipulation tasks. I will begin with a brief introduction to the model-based framework grounded in Koopman Operator Theory\, highlighting its reliance on ground-truth (GT) object states in real-world applications. \nTo address this limitation\, we propose Koopman Operator Rollout for Object Feature Learning (KOROL)—an approach that removes the dependency on GT states in model-based manipulation learning. KOROL learns visual features that predict robot states throughout dynamics model rollouts. Unlike prior approaches that learn implicit visual features for direct image-to-action policies\, KOROL explicitly trains on object-centric visual representations\, encoding essential scene information to improve robot state predictions during autoregressive rollouts. This establishes a synergistic relationship between the learned object features and the Koopman operator. \nOur experiments demonstrate that KOROL: (i) improves performance across various simulated manipulation tasks compared to Koopman operators using GT object states and other baselines\, (ii) extends Koopman-based methods to vision-based real-world tasks\, and (iii) enables multitasking through dimensionally aligned object features. \nCommittee: \nProf. Jeffrey Ichnowski \nProf. Guanya Shi \nProf. Oliver Kroemer \nBardienus Duisterhof
URL:https://www.ri.cmu.edu/event/adaptive-robot-design-for-multimodal-locomotion-across-diverse-terrains-2/
LOCATION:GHC 6115
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251024T143000
DTEND;TZID=America/New_York:20251024T153000
DTSTAMP:20260922T113641
CREATED:20250902T163833Z
LAST-MODIFIED:20251024T220347Z
UID:148668-1761316200-1761319800@www.ri.cmu.edu
SUMMARY:Bringing Dexterity to Robot Hands in the Real World
DESCRIPTION:Abstract:  Dexterous manipulation is a grand challenge of robotics\, and fine manipulation skills are required for many robotics applications that we envision.   In this overview talk\, I will discuss my view of some major factors that contribute to dexterity and discuss how we can incorporate them into our robots and systems.\n\nBio:  Nancy Pollard is a Professor in the Robotics Institute and the Computer Science Department at Carnegie Mellon University. She received her PhD in Electrical Engineering and Computer Science from the MIT Artificial Intelligence Laboratory\, where she developed grasp and manipulation planning algorithms for the Stanford/JPL and Utah/MIT dexterous hands. Prof. Pollard spent the next few decades studying human and robot dexterity\, with emphasis on bringing human manipulation strategies with performance guarantees to humanoid robots with dexterous hands.  She received the NSF CAREER award for research on “Quantifying Humanlike Enveloping Grasps”\,  the Okawa Research Grant for “Studies of Dexterity for Computer Graphics and Robotics\,” and was a recent recipient of an NSF Convergence Accelerator award for “Bio-Inspired Design of Robot Hands for Use-Driven Dexterity.”   She has led the development of several generations of dexterous soft robotic hands\, is a founder of FuturHand Robotics and leads the CMU Foam Hands Laboratory.
URL:https://www.ri.cmu.edu/event/bringing-dexterity-to-robot-hands-in-the-real-world/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/nsp-crop.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251030T113000
DTEND;TZID=America/New_York:20251030T123000
DTSTAMP:20260922T113641
CREATED:20251024T134300Z
LAST-MODIFIED:20251024T134300Z
UID:149169-1761823800-1761827400@www.ri.cmu.edu
SUMMARY:3D Thermal Perception for Autonomous Navigation in Visually Degraded & Unstructured Environments
DESCRIPTION:Abstract:\nAutonomous navigation in visually degraded and unstructured environments\, such as darkness\, smoke\, and rough off-road terrain remains a significant challenge for current robotic systems. RGB cameras fail without illumination\, and active sensors such as LiDAR degrade under aerosols and emit signals that are undesirable in sensitive or adversarial scenarios. In contrast\, long-wave infrared (thermal) sensing captures naturally emitted radiation and maintains visibility through many atmospheric obscurants. This thesis develops a framework for passive thermal autonomy\, enabling 3D stereo perception and autonomous navigation in challenging off-road terrain without active illumination. \nThe absence of large-scale thermal datasets\, sensor-to-perception pipeline\, and field deployments has long prevented reliable passive autonomy in low-visibility\, off-road environments. This work addresses the complete pipeline\, from rigorous sensor integration and custom cross-modality calibration\, through large-scale data collection\, to stereo thermal mapping and odometry\, and autonomy deployment. We introduce MACThermal\, which adapts metric- and uncertainty-aware covariance visual odometry to the thermal domain. It employs dynamic-range normalization and geometry-consistent flow augmentation to improve correspondence and uncertainty estimation. To address the critical gap in thermal 3D vision benchmarks\, we contributed two multi-modal datasets: FIReStereo for aerial platforms and TartanDrive 2.5T for ground vehicles\, covering diverse off-road terrains and visibility conditions. We integrate these perception modules into a complete autonomy stack\, utilizing visual foundation models for semantic understanding and self-supervised traversability estimation for adaptive decision-making. The full system is validated on a full-scale ATV platform deployed in previously unseen off-road terrain under complete darkness. Together\, this work demonstrates a shift from active LiDAR-based navigation to passive thermal autonomy\, enabling over 10x longer nighttime traversals in previously inaccessible terrain. \nCommittee:\nSebastian Scherer (advisor)\nSrinivasa Narasimhan\nNikhil Keetha
URL:https://www.ri.cmu.edu/event/3d-thermal-perception-for-autonomous-navigation-in-visually-degraded-unstructured-environments/
LOCATION:Gates Hillman Center 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251031T143000
DTEND;TZID=America/New_York:20251031T153000
DTSTAMP:20260922T113641
CREATED:20250902T164414Z
LAST-MODIFIED:20251030T145800Z
UID:148671-1761921000-1761924600@www.ri.cmu.edu
SUMMARY:Toward Generalist Humanoid Robots: Recent Advances\, Opportunities\, and Challenges
DESCRIPTION:Abstract: In an era of rapid AI progress\, leveraging accelerated computing and big data has unlocked new possibilities to develop generalist AI models. As AI systems like ChatGPT showcase remarkable performance in the digital realm\, we are compelled to ask: Can we achieve similar breakthroughs in the physical world — to create generalist humanoid robots capable of performing everyday tasks? In this talk\, I will outline our data-centric research principles and approaches for building general-purpose robot autonomy in the open world. I will present our recent work leveraging real-world\, synthetic\, and web data to train foundation models for humanoid robots. Furthermore\, I will discuss the opportunities and challenges of building the next generation of intelligent robots.\n\nBio: Yuke Zhu is an Associate Professor in the Computer Science Department of UT-Austin\, where he directs the Robot Perception and Learning (RPL) Lab. He is also a Director and Distinguished Research Scientist at NVIDIA Research\, where he co-leads the Generalist Embodied Agent Research (GEAR) lab. He focuses on developing intelligent algorithms for generalist robots and embodied agents to reason about and interact with the real world. He obtained his Ph.D. degree from Stanford University. He received the NSF CAREER Award\, the IEEE RAS Early Academic Career Award\, and various faculty fellowships and research awards from Amazon\, JP Morgan\, and Sony Research.
URL:https://www.ri.cmu.edu/event/toward-generalist-humanoid-robots-recent-advances-opportunities-and-challenges/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/yukezhu.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251103T153000
DTEND;TZID=America/New_York:20251103T163000
DTSTAMP:20260922T113641
CREATED:20251027T174614Z
LAST-MODIFIED:20251027T174614Z
UID:149194-1762183800-1762187400@www.ri.cmu.edu
SUMMARY:From Video Generation to Video World Models
DESCRIPTION:Abstract:\nVideo diffusion models have achieved remarkable success in content creation\, yet they still fall short of simulating interactive worlds that respond to users in real time. This talk examines the fundamental challenges preventing these models from evolving into true world simulators. I will present a series of works — CausVid\, Self-Forcing\, MotionStream\, and State-Space World Model — that collectively mark a paradigm shift from non-causal diffusion models to autoregressive–diffusion hybrids capable of streaming long-duration videos with real-time interactivity. These advances move beyond passive video generation toward dynamic\, immersive experiences\, unlocking new possibilities across gaming\, robotics\, live video editing\, and augmented/virtual reality. \nBio: Xun Huang was a Research Scientist at Adobe\, NVIDIA\, as well as an Adjunct Professor at Carnegie Mellon University. He is currently the Founder and CEO of a stealth startup. He obtained his Ph.D. from Cornell University in 2020 under the advisement of Professor Serge Belongie. His doctoral research was recognized with the Fellowship from NVIDIA\, Adobe\, and Snap. His research interests lie broadly in deep generative models\, with a recent focus on video world models. \nHomepage:  xunhuang.me \nSponsor \nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one AI-powered video generation platform that leverages advances in computer vision\, generative modeling\, and multimodal learning to make high-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/from-video-generation-to-video-world-models/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/10/11-3-25.jpeg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251103T160000
DTEND;TZID=America/New_York:20251103T173000
DTSTAMP:20260922T113641
CREATED:20251027T134212Z
LAST-MODIFIED:20251027T134212Z
UID:149192-1762185600-1762191000@www.ri.cmu.edu
SUMMARY:Unifying Perception and Creation with Generative Models
DESCRIPTION:Abstract: \nRecent advances in large-scale generative modeling have reshaped our understanding of visual intelligence. While models such as diffusion and autoregressive transformers have achieved remarkable success in image and video synthesis\, their potential for visual perception and understanding remains underexplored. This thesis investigates how generative models can serve as powerful visual learners—bridging the long-standing divide between generative and discriminative paradigms. \nWe begin with REM (Refer Everything Models)\, a framework for referring video segmentation built upon text-to-video diffusion models. By preserving generative representations and fine-tuning on narrow-domain datasets\, REM achieves state-of-the-art results on standard benchmarks and demonstrates strong generalization to unseen domains. \nBuilding on this foundation\, we introduce a unified perceptual–generative framework that repurposes a single diffusion model across a broad spectrum of computer vision and image restoration tasks. Through joint training and systematic evaluation over 15 tasks spanning perception and synthesis\, we show that diffusion-based models deliver superior or comparable performance to discriminative counterparts\, revealing their intrinsic ability to encode rich\, multi-modal world representations. \nFinally\, we extend our exploration to visual autoregressive (VAR) models\, presenting the first unified architecture capable of efficiently solving the same 15 tasks within a single framework. We show that latent-variable designs\, particularly those leveraging variational autoencoders\, are key to achieving coherent multi-modal understanding and consistent generation. Compared with diffusion counterparts\, VAR-based models offer substantial gains in latency\, scalability\, and output consistency. \nCollectively\, these studies offer a cohesive perspective on unifying perception and synthesis through generative modeling\, charting a path toward general-purpose visual foundation models that seamlessly integrate understanding\, reasoning\, and creation. \n\nThesis Committee: \nMartial Hebert\, Chair \nDeva Ramanan \nJun-Yan Zhu \nAlexei Efros\, University of California\, Berkeley \nYu-Xiong Wang\, University of Illinois Urbana-Champaign \nPavel Tokmakov\, Toyota Research Institute \nDraft of Document
URL:https://www.ri.cmu.edu/event/unifying-perception-and-creation-with-generative-models/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251107T150000
DTEND;TZID=America/New_York:20251107T163000
DTSTAMP:20260922T113641
CREATED:20251001T150038Z
LAST-MODIFIED:20251104T153929Z
UID:148979-1762527600-1762533000@www.ri.cmu.edu
SUMMARY:Visual-Tactile Synthesis for Texture Generation
DESCRIPTION:Abstract: Recent advances in generative models have enabled the creation of highly realistic visual content\, yet they remain limited to visual perception alone. In contrast\, human interaction with the physical world is inherently multimodal — we not only see textures but also feel them. This gap motivates the goal of my thesis: to build generative models that jointly synthesize visual and tactile modalities for material and texture generation. By unifying what we see and what we touch\, such models can drive new forms of physically grounded content creation\, from robotics and virtual reality to material design.\nHowever\, extending generative modeling to touch presents unique challenges: tactile data is scarce\, noisy\, and expensive to collect\, and there is no large-scale paired dataset linking visual appearance with tactile response. To address these challenges\, my research explores three synergistic directions. \nPart I: I introduce controllable visual-tactile synthesis models that jointly generate aligned visual and tactile textures from shared latent representations\, enabling explicit control over appearance and feel. \nPart II: I propose tactile-aware 3D generation frameworks that integrate tactile sensing into 3D diffusion pipelines\, allowing models to infer physically grounded material properties from visual cues and geometry. \nPart III: Building on these insights\, I aim to develop scalable multimodal generation systems that leverage large vision and language foundation models and physics priors to synthesize novel materials directly from text or image input\, without relying on extensive paired tactile data. \n\n \nThesis Committee:\nJun-Yan Zhu (Co-chair)\nWenzhen Yuan (Co-chair)\nShubham Tulsiani\nAndrew Owens (Cornell Tech)\n\nThesis Document
URL:https://www.ri.cmu.edu/event/ruihan-gao-phd-thesis-proposal/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251114T143000
DTEND;TZID=America/New_York:20251114T153000
DTSTAMP:20260922T113641
CREATED:20250902T170114Z
LAST-MODIFIED:20251125T172342Z
UID:148675-1763130600-1763134200@www.ri.cmu.edu
SUMMARY:Just Asking Questions
DESCRIPTION:Abstract: In the age of deep networks\, “learning” almost invariably means “learning from examples”. We train language models with human-generated text and labeled preference pairs\, image classifiers with large datasets of images\, and robot policies with rollouts or demonstrations. When human learners acquire new concepts and skills\, we often do so with richer supervision\, especially in the form of language—we learn new concepts from examples accompanied by descriptions or definitions\, and new skills from demonstrations accompanied by instructions. Crucially\, language-based supervision involves not only instructions but *questions*—students ask questions to elicit the most useful pieces of supervision\, and teachers ask questions to probe student knowledge and encourage them to acquire new skills or aspects of understanding on their own. This talk will focus on a few recent projects focused on building computational models that can ask good questions for both learning and teaching\, with applications spanning LM alignment\, policy learning\, and education. This is joint work with Belinda Li\, Alex Tamkin\, Noah Goodman\, Andi Peng\, Ilia Sucholutsky\, Nishanth Kumar\, Julie A Shah\, Andreea Bobu\, Alexis Ross\, Gabe Grand\, Valerio Pepe and Josh Tenenbaum. \nBio: Jacob Andreas is an associate professor at MIT in the Department of Electrical Engineering and Computer Science as well as the Computer Science and Artificial Intelligence Laboratory. His research aims to understand the computational foundations of language learning\, and to build intelligent systems that can learn from human guidance. Jacob earned his Ph.D. from UC Berkeley\, his M.Phil. from Cambridge (where he studied as a Churchill scholar) and his B.S. from Columbia. He has received a Sloan fellowship\, an NSF CAREER award\, MIT’s Junior Bose and Kolokotrones teaching awards\, and paper awards at ACL\, ICML and NAACL.
URL:https://www.ri.cmu.edu/event/just-asking-questions/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/head_small.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251117T103000
DTEND;TZID=America/New_York:20251117T113000
DTSTAMP:20260922T113641
CREATED:20251113T150929Z
LAST-MODIFIED:20251113T150929Z
UID:149416-1763375400-1763379000@www.ri.cmu.edu
SUMMARY:Towards Modernization of Long-Range Image-Space Planning for Off-Road Navigation
DESCRIPTION:Abstract: \nThis thesis revisits long-range\, image-space planning for off-road navigation and modernizes the classical first-person view (FPV) paradigm by building upon recent advances in perception. It introduces a lightweight depth calibration scheme\, analytic configuration-space (C-space) transforms\, interpretable frontier selection\, and a pixel-space A* planner with validated heuristic soundness. Concretely\, we (i) make monocular depth metrically usable at test time via an affine\, log-domain calibration with sparse LiDAR; (ii) derive and implement a depth-aware FPV C-space inflation that projects vehicle width/length analytically and realizes it with separable row/column sliding-maximum filters augmented by per-pixel depth consistency checks; (iii) propose transparent\, angular-sector frontiering that reasons jointly about traversability cost and minimal lethal depth\, alongside goal-aware revalidation; and (iv) preserve A* admissibility/consistency in image space through a simple cost renormalization that avoids silent suboptimality in low-cost free space. \nWe evaluate the resulting\, modular sub-stack in the high-fidelity Falcon simulator [2] under a shared ROS graph. Using a common perception front-end and planning back-end\, we compare three frontiering strategies — (1) a LAGR-style row-wise baseline\, (2) an LRN-inspired openness heuristic adapted to operate with explicit depth and cost\, and (3) an Angular Cost & Depth (ACD) variant that couples average cost with a minimum lethal-depth constraint. Across diverse courses (e.g.\, farm\, desert\, mixed terrain)\, the calibrated monocular depth reduces error versus raw monocular predictions\, and both LRN-inspired and ACD frontiering tend to outperform the purely row-wise baseline on longer traverses. We view these as encouraging indications rather than definitive claims: all results are in simulation\, with performance subject to perception quality\, calibration coverage\, and environment diversity. \n\nScope and limitations are explicit. The work was conducted over a short project window (May 2025–present) and prioritized stabilizing the proposed sub-stack in conjunction with a core FieldAI [1] stack through the high-fidelity Falcon simulator [2] and ROS integration into a reliable\, end-to-end testing framework. No real-world deployments were performed\, and the current implementation targets ROS1. Nevertheless\, the design is intentionally modular and auditable to make it relatively straightforward to integrate the sub-stack into existing off-road autonomy stacks lacking explicit long-range planning. Such integration will enable better autonomy by offering a practical bridge between classical image-space efficiency and metric-world robustness. \n[1] FieldAI: https://www.fieldai.com \n[2] Falcon: https://www.duality.ai \nCommittee: \nWenshan Wang\, Chair\nSebastian Scherer\, Co-Chair\nMaxim Likhachev\nSamuel Triest\, RI PhD
URL:https://www.ri.cmu.edu/event/towards-modernization-of-long-range-image-space-planning-for-off-road-navigation/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251119T130000
DTEND;TZID=America/New_York:20251119T140000
DTSTAMP:20260922T113641
CREATED:20251114T143444Z
LAST-MODIFIED:20251114T145002Z
UID:149445-1763557200-1763560800@www.ri.cmu.edu
SUMMARY:Building Robot Hands and Teaching Dexterity
DESCRIPTION:Abstract:  \nOur human hands are masterpieces of power and precision\, capable of typing\, hammering\, or delicately using chopsticks. Yet most robots today still rely on simple two-finger grippers in controlled settings because dexterous hands are costly and difficult to deploy. To close this gap\, I will introduce my LEAP Hands\, high-performance\, low-cost\, and easy-to-assemble robotic hands that have become the most widely used platform for dexterous manipulation research. LEAP Hand V1 employs motor-in-joint actuation for simplicity\, while V2 introduces a novel hybrid rigid–soft structure that delivers exceptional strength and durability.  I will then show how large-scale human video/motion data and simulation techniques can teach human-like manipulation skills across diverse environments.  By tightly integrating mechanical design and machine learning\, my open-source robot hands achieve unprecedented levels of dexterity for a variety of everyday tasks. \nCommittee:\nProf. Deepak Pathak (advisor)\nProf. Nancy Pollard \nProf. Abhinav Gupta \nProf. Jitendra Malik \nDr. Ankur Handa \n  \nA draft of my thesis proposal is available here: \nhttps://kennyshaw.net/phd_thesis_proposal.pdf
URL:https://www.ri.cmu.edu/event/building-robot-hands-and-teaching-dexterity-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251120T140000
DTEND;TZID=America/New_York:20251120T150000
DTSTAMP:20260922T113641
CREATED:20251114T205305Z
LAST-MODIFIED:20251114T205305Z
UID:149469-1763647200-1763650800@www.ri.cmu.edu
SUMMARY:Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models
DESCRIPTION:Abstract:\nTransferring skills between different objects remains one of the core challenges of open-world robot manipulation. Generalization needs to take into account the high-level structural differences between distinct objects while still maintaining similar low-level interaction control. In this paper\, we propose an example-based zero-shot approach to skill transfer. Rather than treating skills as atomic\, we decompose skills into a prioritized list of grounded task-axis (GTA) controllers. Each GTAC defines an adaptable controller\, such as a position or force controller\, along an axis. Importantly\, the GTACs are grounded in object key points and axes\, e.g.\, the relative position of a screw head or the axis of its shaft. Zero-shot transfer is thus achieved by finding semantically similar grounding features on novel target objects. We achieve this example-based grounding of the skills through the use of foundation models\, such as SD-DINO\, that can detect semantically similar keypoints of objects. We evaluate our framework on real-robot experiments\, including screwing\, pouring\, and spatula scraping tasks\, and demonstrate robust and versatile controller transfer for each.\n\nCommittee:\nProf. Oliver Kroemer\nProf. Katerina Fragkiadaki\nProf. Zeynep Temel\nMark Lee
URL:https://www.ri.cmu.edu/event/grounded-task-axes-zero-shot-semantic-skill-generalization-via-task-axis-controllers-and-visual-foundation-models/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251120T163000
DTEND;TZID=America/New_York:20251120T173000
DTSTAMP:20260922T113641
CREATED:20251112T161204Z
LAST-MODIFIED:20251112T161723Z
UID:149408-1763656200-1763659800@www.ri.cmu.edu
SUMMARY:OpenVDB
DESCRIPTION:Abstract: As the inventor of VDB and founder of OpenVDB\, I am excited to talk about its history\, motivation\, and diverse adoption. Specifically\, this lecture will cover the underlying VDB data structure\, and its adoption to computer graphics\, physics simulations and more recently machine learning. Since its open-source release in 2012\, OpenVDB has become an industry standard and has been used in numerous VFX franchises like “Avatar”\, “Avengers”\, “The Mummy”\, “Pirates of the Caribbean”\, “Kung Fu Panda”\, and “How to Train Your Dragon”. It is adopted by numerous commercial software packages used by the entertainment industry\, including Houdini\, RenderMan\, Arnold\, Blender\, and Unreal Engine\, just to mention a few. OpenVDB has also found use in many areas outside of media and entertainment\, including SLAM\, autonomous driving\, topology optimization\, semiconductor designs\, 3D printing\, medical imaging\, rocket design\, aerial surveillance\, robotics\, and many machine learning applications. Finally\, OpenVDB was the first open-source project to be adopted by the Academy Software Foundation (ASWF) and the Linux Foundation (in 2018). \nBio: Ken Museth is Sr Director of the High-Fidelity Physics Research team at Nvidia and chair of the Technical Steering Committee for OpenVDB under the Academy Software Foundation. He has a PhD in quantum physics from Copenhagen University and did his postgraduate studies in computer science at Caltech. Previously he was a Sr. Computational Scientist at SpaceX for six years\, working on CFD simulations of the Raptor engine\, Head of Simulation R&D at Weta FX for three years\, working on James Cameron’s Avatar 2\, director for FX and CFX simulation teams at DreamWorks Animation for eight years\, Sr Software Engineer at Digital Domain for three years\, full tenured professor in computer graphics at Linköping University for four years\, where he supervised five PhD and 15 MSc students\, and research scientist at NASA’s Jet Propulsion Laboratory for three years\, working on space-mission design and visualization. Ken invented and founded OpenVDB for which he won two Academy Awards from the Academy of Motion Picture Arts and Sciences; a Technical Achievement Award (Academy Certificate) in 2015 and a Scientific & Engineering Award (Academy Plaque) in 2024. Additionally\, in 2023 he was awarded the ACM SIGGRAPH Practitioner Award and was accepted into the ACM SIGGRAPH Academy. Ken has served on the Technical Papers Committee for ACM SIGGRAPH multiple times and has 29 movie credits\, including on franchises like “Avatar”\, “Avengers”\, “The Mummy”\, “Pirates of the Caribbean”\, “Kung Fu Panda”\, and “How to Train Your Dragon”. \nWebsite: https://research.nvidia.com/labs/prl/author/ken-museth/ \nPoster:  https://graphics.cs.cmu.edu/cmgc/ken-museth.pdf \nThe Carnegie Mellon Graphics Colloquium is hosted by the Carnegie Mellon Graphics Lab and supported by Meta and Adobe.
URL:https://www.ri.cmu.edu/event/openvdb/
LOCATION:Gates-Hillman Center 4401
CATEGORIES:Special Events
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/11/11-20-25-scaled.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251121T150000
DTEND;TZID=America/New_York:20251121T160000
DTSTAMP:20260922T113641
CREATED:20250915T203414Z
LAST-MODIFIED:20251125T191630Z
UID:148862-1763737200-1763740800@www.ri.cmu.edu
SUMMARY:How to Coordinate Thousands of Robots Efficiently and Robustly
DESCRIPTION:Abstract: \nLarge-scale robot fleets are increasingly deployed in warehouses\, factories\, transportation systems\, and emerging robotics applications. Coordinating hundreds or thousands of robots in shared\, cluttered spaces creates fundamental challenges in maintaining safety\, preventing deadlocks\, and minimizing congestion. In this talk\, I will present our recent work on scalable imitation learning methods for coordinating 10k robots\, automatic environment optimization techniques for alleviating traffic congestion\, and asynchronous execution frameworks that guarantee safe\, robust\, and deadlock-free multi-robot operations.\n\nBio: \nJiaoyang Li is an Assistant Professor in the Robotics Institute at Carnegie Mellon University. Her research lies at the intersection of artificial intelligence\, robotics\, and optimization\, with a particular focus on scalable multi-robot planning and coordination. She received her Ph.D. in Computer Science from the University of Southern California in 2022 and her B.Eng. in Automation from Tsinghua University in 2017. She is a recipient of the NSF CAREER Award\, best dissertation awards from ICAPS\, AAMAS\, and USC\, and multiple best paper awards from ICRA\, ICAPS\, and MRS.
URL:https://www.ri.cmu.edu/event/how-to-coordinate-thousands-of-robots-efficiently-and-robustly/
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/profile-scaled.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251124T130000
DTEND;TZID=America/New_York:20251124T143000
DTSTAMP:20260922T113641
CREATED:20251114T144910Z
LAST-MODIFIED:20251114T145326Z
UID:149447-1763989200-1763994600@www.ri.cmu.edu
SUMMARY:Sensorimotor-Aligned Design for Pareto-Efficient Haptic Immersion in Extended Reality
DESCRIPTION:Abstract: \nA new category of computing devices has emerged: augmented and virtual reality headsets\, collectively referred to as extended reality (XR). These devices can alter\, augment\, or even replace our reality. While these headsets have made impressive strides in audio-visual immersion over the past half-century\, XR interactions remain almost completely absent of appropriately expressive tactile sensations. At present\, even the most advanced mainstream consumer systems rely on vibrotactile haptic actuators in the controllers\, inherently limited to clicks and buzzes — an exceedingly limited range of expressivity with which to represent the rich tactile world. \nRealizing the holistic promise of XR requires full-body haptic immersion\, just as much as it requires full audio-visual immersion. To frame critical design considerations in haptics research\, I propose an immersion-practicality tradeoff model. These competing objectives underscore the inherent tension between providing rich sensory feedback (often e.g.. costly\, bulky)\, while maintaining consumer feasibility and usability (e.g.\, low cost\, easy to use). Under this framework\, I sought to identify and build Pareto-efficient haptic systems where the haptic approach is aligned with humans’ sensorimotor system. I present five finished projects that embody my design approach through tactile haptics to different regions of the body. This framework has also been extended to other major dimensions of haptics — kinesthetics (force feedback) and proprioception — to explore how these principles can generalize. These systems\, when combined together\, advance the vision of practical\, immersive\, full-body haptics. \nThesis Committee: \nChris Harrison\, Chair\, HCII CMU \nZeynep Temel\, RI CMU \nJim McCann\, RI CMU \nMar Gonzalez-Franco\, Google \nLink to the Thesis folder
URL:https://www.ri.cmu.edu/event/sensorimotor-aligned-design-for-pareto-efficient-haptic-immersion-in-extended-reality-2/
LOCATION:Newell-Simon Hall 3305
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251125T110000
DTEND;TZID=America/New_York:20251125T120000
DTSTAMP:20260922T113641
CREATED:20251118T142159Z
LAST-MODIFIED:20251118T142159Z
UID:149484-1764068400-1764072000@www.ri.cmu.edu
SUMMARY:Multi-View 4D Human Reconstruction under Interaction Scenarios
DESCRIPTION:Abstract:\nBuilding large-scale human datasets from multi-view videos is essential for advancing research in human behavior understanding\, virtual reality\, animation\, and robotics. Compared to traditional motion capture systems that rely on physical markers to track motion\, vision-based reconstruction not only enables the capture of human motion in unconstrained environments but also avoids altering human appearance with markers. However\, existing multi-view reconstruction methods often fail when humans interact closely with others or with objects due to severe occlusions and truncations introduced by complex activities. In this thesis\, we develop a markerless capture system capable of handling close human interactions and dexterous hand-object manipulations. Using this system\, we construct two large-scale human datasets\, Harmony4D and Contact4D\, which serve as landmarks for advancing fundamental research in human-centric AI\, such as human pose estimation\, contact estimation\, and motion generation. \nCommittee:\nKris Kitani\, Chair\nFernando De La Torre\nShubham Tulsiani\nErica Weng
URL:https://www.ri.cmu.edu/event/multi-view-4d-human-reconstruction-under-interaction-scenarios/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251125T140000
DTEND;TZID=America/New_York:20251125T150000
DTSTAMP:20260922T113641
CREATED:20251119T143738Z
LAST-MODIFIED:20251119T143738Z
UID:149496-1764079200-1764082800@www.ri.cmu.edu
SUMMARY:Vision-Based Multi-Wire Detection and Tracking for UAV Wire Approach
DESCRIPTION:Abstract: \nReliable detection and tracking of power lines is critical for enabling under-wire UAV approach and inductive power-line charging to extend UAV range. However\, wires are thin\, featureless\, and visually ambiguous structures that challenge traditional computer vision methods and degrade depth estimation accuracy. To address these challenges\, this thesis presents a fully passive\, camera-only multi-wire detection and tracking algorithm that operates in real time on lightweight onboard compute using only stereo RGB imagery. \nThe proposed framework integrates a lightweight classical vision pipeline with three-dimensional geometric reasoning and a multi-stage filtering process to produce robust wire instance detections. A complementary oriented object detection model is trained using labels generated by the classical pipeline\, leveraging its fine-tuned geometric outputs to improve resilience in challenging visual conditions. To track individual wire instances across frames\, we introduce a Kalman-filter-based tracking architecture that estimates both wire orientation and per-wire positional state while remaining robust to wire detection outliers and vehicle pose drift. The system is further expanded by testing a wire positional servoing approach using the tracked wire instances in simulation and is validated across a diverse range of data sources\, including simulation\, indoor testing\, handheld experiments\, and outdoor flight evaluations. \n\nCommittee:\nSebastian Scherer (advisor)\nWennie Tabib\nMohammad Mousaei
URL:https://www.ri.cmu.edu/event/vision-based-multi-wire-detection-and-tracking-for-uav-wire-approach/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251202T110000
DTEND;TZID=America/New_York:20251202T120000
DTSTAMP:20260922T113641
CREATED:20251125T145824Z
LAST-MODIFIED:20251125T145824Z
UID:149623-1764673200-1764676800@www.ri.cmu.edu
SUMMARY:Towards Scaling Embodied Data for Robot Learning
DESCRIPTION:Abstract:\nAs artificial intelligence advances quickly in the digital domain\, the next\nfrontier lies in physical intelligence: systems that learn through acting and\nsensing in the real world. In this thesis\, we explore practical ways of scaling\nsuch embodied data across three directions. AnyCar scales synthetic data\nthrough large-scale simulation\, training a universal dynamics transformer\nthat generalizes across vehicles and environments. FACTR improves\nthe efficiency of real robot data with a low-cost bilateral teleoperation\nsystem and a curriculum that teaches policies to integrate force and\nvision. DexWild scales human data through in-the-wild data collection\nand co-training with robot demonstrations\, enabling generalization to\nunseen objects and environments. Together\, these projects explore how a\ndata-centric approach can enable more adaptive and capable robots. \nCommittee:\nDeepak Pathak (chair)\nGuanya Shi\nKenneth Shaw
URL:https://www.ri.cmu.edu/event/towards-scaling-embodied-data-for-robot-learning/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251203T093000
DTEND;TZID=America/New_York:20251203T103000
DTSTAMP:20260922T113641
CREATED:20251201T142655Z
LAST-MODIFIED:20251201T142655Z
UID:149640-1764754200-1764757800@www.ri.cmu.edu
SUMMARY:Attractors and Their Applications in Heuristic Search
DESCRIPTION:Abstract:\nHeuristic search provides a principled way to guide exploration in large state spaces\, enabling efficient solution finding. As a result\, it is widely used across domains such as robotics\, games\, and planning. However\, its performance is often limited by memory consumption and computational overhead\, which have motivated extensive research on improving both. This thesis introduces a sparse representation called attractors and explores two of its applications in heuristic search. First\, we present Attractor-based Closed List Search (ACLS)\, a framework that uses attractors to sparsely represent the Closed list. ACLS intelligently identifies attractor states in a way that enables efficient solution reconstruction while preserving theoretical guarantees on the quality of the solution. We demonstrate that ACLS significantly reduces memory usage\, while achieving comparable planning times and outperforming state-of-the-art approaches. Second\, we introduce front-to-attractors (F2A) heuristics\, a family of heuristics that leverage attractors in bidirectional heuristic search (Bi-HS). We demonstrate that F2A heuristics substantially reduce the number of heuristic evaluations compared to front-to-front (F2F) heuristics\, while maintaining strong informativeness and reducing expansions relative to front-to-end (F2E) heuristics\, resulting in improved runtime performance. Together\, these projects demonstrate the broad potential of attractors in heuristic search.\n\nCommittee:\nMaxim Likhachev (chair)\nJiaoyang Li\nYorai Shaoul
URL:https://www.ri.cmu.edu/event/attractors-and-their-applications-in-heuristic-search/
LOCATION:GHC 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251203T150000
DTEND;TZID=America/New_York:20251203T163000
DTSTAMP:20260922T113641
CREATED:20251125T221309Z
LAST-MODIFIED:20251125T221309Z
UID:149634-1764774000-1764779400@www.ri.cmu.edu
SUMMARY:Robotic System Design Principles for Human-Human Collaboration
DESCRIPTION:Abstract: \nRobots possess unique affordances granted by combining software and hardware. Most existing research focuses on the impact of these affordances on human-robot collaboration\, but the theory of how robots can facilitate human-human collaboration is underdeveloped. Such a theory would be beneficial in education. An educational device can afford collaboration in both assembly and use. This thesis will enumerate and validate the design principles of educational devices that facilitate collaborative assembly and collaborative learning.\nThis research draws upon cognitive theories used in the disciplines of Computer-Supported Collaborative Work (CSCW)\, Computer-Supported Collaborative Learning (CSCL)\, Educational Robotics\, and Human-Robot Interaction (HRI). Each discipline uses theories that align with its respective goals to model different pieces of cognition. However\, they do not consider other factors outside their respective goals. Diverse analytical lenses are needed to understand the multiple dimensions of influence an educational device can have on human-human interaction to support collaborative assembly and collaborative learning.\nWe explore these dimensions first through the development and assessment of\nRoboLoom\, a robotic Jacquard loom kit designed for interdisciplinary\, collaborative education. Through the study of RoboLoom’s use and assembly in an undergraduate course\, we extract design features that facilitate student-student collaboration during classroom activities. These features encompass task complexity\, task parallelization\, physicality\, repetition of tasks\, specificity of hardware\, and familiarity with hardware.\nWe then explore these design principles through three studies: a comparison\nbetween two different looms\, a study of devices designed for and against the principles\, and a comparison of two versions of RoboLoom. We find five design principles that influence collaborative behavior: repetitiveness\, specificity\, parallelizability\, physicality\, and difficulty. These design principles were shown to causally change collaborative behaviors in controlled lab settings and in situ engineering education tasks. By evaluating these systems through multiple cognitive lenses\, we determine that these design principles are effective in facilitating collaborative assembly and promising for collaborative learning.\n\n\n\nCommittee Members: \n    Illah Nourbakhsh\, Co-Chair\n    Melisa Orta Martinez\, Co-Chair\nJames McCann\nKylie Peppler\, University of California\, Irvine \n\nLink to Thesis
URL:https://www.ri.cmu.edu/event/robotic-system-design-principles-for-human-human-collaboration/
LOCATION:GHC 8102
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
END:VCALENDAR