BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251003T143000
DTEND;TZID=America/New_York:20251003T153000
DTSTAMP:20260922T105618
CREATED:20250902T162718Z
LAST-MODIFIED:20251006T153641Z
UID:148662-1759501800-1759505400@www.ri.cmu.edu
SUMMARY:Neural Certificates for Safe Robotic System Planning and Control
DESCRIPTION:Abstract:\nAchieving safety\, scalability\, and high performance in complex systems\, such as multi-agent systems (MAS) control\, is a central challenge in many real-world robotic deployments due to its computational complexity as a large-scale constrained optimal control problem. To address this\, we introduce a novel graph control barrier function (GCBF) as a core tool for large-scale distributed safe control\, which guarantees safety for arbitrarily large MAS with only local observations. For MAS with known dynamic models\, we present a self-supervised learning framework that can jointly learn GCBF and distributed control policies that consider actuation limits. For MAS with unknown dynamics\, we discuss how to blend GCBF in multi-agent reinforcement learning (MARL) to achieve high-performance and safe distributed policies.\n\nBio:\nChuchu Fan is an Associate Professor (pre-tenure) in the Department of Aeronautics and Astronautics (AeroAstro) and Laboratory for Information and Decision Systems (LIDS) at MIT. Before that\, she was a postdoc researcher at Caltech and got her Ph.D. at the University of Illinois at Urbana-Champaign. She earned her bachelor’s degree from Tsinghua University. Her research group\, the Realm at MIT\, works on developing computational tools that integrate rigorous mathematics into machine learning and AI for the design\, analysis\, and verification of safe\, large-scale\, and complex systems. Chuchu is the recipient of an NSF CAREER Award\, an AFOSR Young Investigator Program (YIP) Award\, an ONR YIP Award\, and the 2020 ACM Doctoral Dissertation Award.
URL:https://www.ri.cmu.edu/event/neural-certificates-for-safe-robotic-system-planning-and-controlri-seminar-w-chuchu-fan/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/ChuchuFan-042021.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251009T093000
DTEND;TZID=America/New_York:20251009T110000
DTSTAMP:20260922T105618
CREATED:20250903T175132Z
LAST-MODIFIED:20250930T131653Z
UID:148746-1760002200-1760007600@www.ri.cmu.edu
SUMMARY:Customizing Text-to-Image Diffusion Models
DESCRIPTION:Abstract: With the rapid advancement of generative models\, their potential to transform creative content creation is increasingly evident. However\, most large-scale generative models are primarily text-conditioned\, given the availability of large-scale paired text–image datasets. In contrast\, for most practical applications\, creators often begin from an existing asset and wish to generate variations or modify it in specific ways. For images\, this may involve placing an object in a new context\, adjusting local attributes\, or altering visual style. My research focuses on customizing pre-trained generative models\, primarily text-to-image diffusion models\, to facilitate such downstream tasks. A central challenge here is the lack of paired input–output data for these tasks. \nTo address this\, I explore three complementary directions: \nPart I: I study few-shot learning methods\, which are computationally efficient but require fine-tuning for each new task instance. This limitation motivates the second direction. \nPart II: Constructing synthetic paired datasets using the capabilities of pre-trained generative models themselves to train feed-forward models in a supervised manner. However\, constructing such datasets requires careful curation\, filtering\, and risk of becoming outdated as base pre-trained models evolve. Building on these insights\, my thesis proposes a third paradigm. \nPart III: Customizing generative models without paired supervision. Instead\, we plan to leverage vision–language models to evaluate task success and provide direct gradient-based feedback to the generative model. This approach has the potential to create a scalable and robust framework for efficient customization of generative models for downstream tasks without relying on synthetic datasets. \n\n\n \nThesis Committee:\nJun-Yan Zhu (Chair)\nDeva Ramanan\nShubham Tulsiani\nPhillip Isola (MIT)\n\nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/customizing-text-to-image-diffusion-models/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251010T110000
DTEND;TZID=America/New_York:20251010T130000
DTSTAMP:20260922T105618
CREATED:20250909T130917Z
LAST-MODIFIED:20250930T145809Z
UID:148785-1760094000-1760101200@www.ri.cmu.edu
SUMMARY:Pushing the Frontier of Robotic Tool Manipulation by Treating the Hand and the Tool Together as a Machine
DESCRIPTION:Abstract: Tool manipulation is an essential human skill. It expands our manipulation capability beyond the capability of the biological hand\, and is a defining feature of many tasks centered on physical interaction with the real world. For humanoid robots to become general-purpose\, they must master tool manipulation as well. However\, the state-of-the-art humanoid robots equipped with multi-finger hands still fall behind their human counterparts in tool manipulation performance. This thesis aims to narrow this gap by treating the hand and the tool together as a machine.\nSpecifically\, inspired by the analogy between multi-finger hands and CNC machines\, this thesis interprets a tool-manipulating hand as configuring itself and the tool into different tool-hand mechanisms in real time. To concretely represent each tool-hand mechanism—which consists of the tool\, the hand\, and the contacts—this thesis introduces two concepts: 1) foundational pose\, a pose and precondition that the tool and the hand must reach for the tool-hand mechanism to be successfully constructed and to run\, and a concise representation of tool-hand mechanism. 2) sub-assembly\, a set of contacts that independently fulfills part of the tool-hand mechanism’s function\, and a detailed\, modular representation of tool-hand mechanism. \nThis thesis first tests the validity of the concept of foundational pose via the question: “if a tool and a hand have reached a foundational pose\, can they act as the corresponding tool-hand mechanism and perform the tool manipulation motion?” To answer this question\, the thesis conducts a hand design experiment\, which uses foundational poses as constraints to sample many different hands and evaluates their tool manipulation motions. The results lead to a positive answer to the question\, verifying the concept of foundational pose. \n\nThen\, this thesis expands roll-slide contact-based tool manipulation motion planning—which previously was only possible for primitive shapes with global parametrizations—to manifold meshes\, which allows motion planning from foundational poses for arbitrarily shaped tools and hands. \nFinally\, for the proposed work\, this thesis aims to test the validity of the concept of sub-assembly via the question: “how many sub-assemblies are enough?” Based on the answer to this question\, this thesis aims to develop a sub-assembly-based control framework\, and test the framework on a real robotic hand for an entire tool manipulation sequence. \n\n \n\nThesis Committee Members: \nProf. Nancy Pollard (co-chair)\nProf. Jean Oh (co-chair)\nProf. Matthew Mason\nDr. Lael Odhner (The Robotics and AI Institute)\n\n\nDraft of the Thesis Proposal Document Link: https://drive.google.com/file/d/1L5ri7r0295poQOyTtqEOI3-LX3o7Gqlf/view?usp=drive_link
URL:https://www.ri.cmu.edu/event/pushing-the-frontier-of-robotic-tool-manipulation-by-treating-the-hand-and-the-tool-together-as-a-machine/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251010T143000
DTEND;TZID=America/New_York:20251010T153000
DTSTAMP:20260922T105618
CREATED:20250902T163332Z
LAST-MODIFIED:20251010T231714Z
UID:148665-1760106600-1760110200@www.ri.cmu.edu
SUMMARY:A Manipulation Journey
DESCRIPTION:Abstract:\nThe talk will revisit my career in manipulation research\, focusing on projects that might offer some useful lessons for others. We will start with my beginnings at the MIT AI Lab and my MS thesis\, which is still my most cited work\, then continue with my arrival at CMU\, a discussion with Allen Newell\, an exercise to envision a coherent research program\, and how that led to a second and third childhood. The talk will conclude with some discussion of lessons learned.\n\nBio:\nMatt has spent 50 years conducting research in Artificial Intelligence and Robotics\, starting as a student in the MIT AI Lab where he earned the BS\, MS\, and PhD degrees. He spent much of his career at CMU’s Robotics Institute\, where he was the founder and co-director of the Manipulation Laboratory\, and for ten years served as the Director of the Robotics Institute. Matt’s group studied the basic physics governing grasping and manipulation\, and demonstrated that sophisticated grasping and manipulation can be produced by the simple and robust grippers used in industrial automation.\n\nMason is a Fellow of the AAAI\, AAAS\, ACM\, and IEEE\, and a winner of the IEEE R&A Society’s Pioneer Award\, and the IEEE Technical Field Award in Robotics and Automation (the R&A Prize).\n\nMatt now focuses his attention on logistics and warehouse robotics\, helping to produce industry-leading solutions at Berkshire Grey.
URL:https://www.ri.cmu.edu/event/a-manipulation-journey/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/unnamed.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251010T143000
DTEND;TZID=America/New_York:20251010T163000
DTSTAMP:20260922T105618
CREATED:20250902T201509Z
LAST-MODIFIED:20251006T150600Z
UID:148687-1760106600-1760113800@www.ri.cmu.edu
SUMMARY:Generative Robotics: Self-Supervised Learning for Human-Robot Collaborative Creation
DESCRIPTION:Abstract:\nRobotic automation is generally welcomed for tasks that are dirty\, dull\, or dangerous\, but with expanding robotic capabilities\, robots are entering domains that are safe and enjoyable\, such as creative industries. Although there is a widespread rejection of automation in creative fields\, many people\, from amateurs to professionals\, would welcome supportive or collaborative creative tools. Supporting creative tasks is challenging with real-world robotics because there are limited relevant datasets\, creative tasks are abstract and high-level\, and real-world tools and materials are difficult to model and predict. Learning-based robotic intelligence is a promising method for creative support tools\, but since the task is so complex\, common approaches such as learning from demonstration would require too many samples and reinforcement learning may never converge. In this thesis\, we show that robots can learn to support acts of creativity purely through a few\, proposed self-supervised learning techniques. \nWe formalize robots that support people in the making of things from high-level goals in the real world as a new field\, Generative Robotics. We introduce an approach for supporting 2D visual art-making with paintings and drawings along with 3D clay sculpture from a fixed perspective. Because there are no robotic datasets for collaborative painting and sculpting\, we designed our approach to learn from small\, self-generated datasets to learn real-world constraints and support collaborative interactions. Our approach uses (1) Real2Sim2Real to enable a robot to teach itself about physical constraints (e.g.\, type of paint and brush)\, (2) semantic planning to plan from high-level\, abstract goals under severe real-world constraints (e.g.\, making a painting from a detailed photograph with only 4 colors and 64 brush strokes)\, and (3) self-supervised learning to generate data to train the robot to support creation rather than automate it. Our approach collaboratively creates paintings in heavily constrained settings. Lastly\, we generalize our approach to new materials\, tools\, action representations\, and state representations to perform long-horizon clay sculpting.\n\nDocument:\nhttps://drive.google.com/drive/folders/1H4-gyLhccQTtlipmBRGAHbrV3bHI7LtU?usp=sharing \n\n\nThesis Committee Members: \nJean Oh\, Chair \nJames McCann \nManuela Veloso \nKen Goldberg\, UC Berkeley
URL:https://www.ri.cmu.edu/event/generative-robotics-self-supervised-learning-for-human-robot-collaborative-creation-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251014T100000
DTEND;TZID=America/New_York:20251014T113000
DTSTAMP:20260922T105618
CREATED:20251001T150325Z
LAST-MODIFIED:20251024T233624Z
UID:148981-1760436000-1760441400@www.ri.cmu.edu
SUMMARY:Consistent Modeling of 4D Scenes for Perception and Generation
DESCRIPTION:Abstract:\n\nA core challenge in vision is building representations that capture 3D scenes over time for perception and interactive generation. For accurate perception and plausible generation\, we want consistency across views\, time\, and modalities. In this talk we explore consistency through the choice of representation\, moving from dense grid formulations to entity-centric scenes that are easier to maintain across frames\, and we extend that representation from perception to generation.\n\nOur past work follows this shift within perception tasks. SOLOFusion uses a grid representation with long- and short-baseline temporal stereo for multi-camera 3D detection\, improving foreground depth\, but it does not perform entity grouping and it does not model background. ASCFormer performs depth estimation and completion via pixel–point affinity\, grouping geometry coherently\, but the grouping is geometric rather than semantic and remains static. DetMatch\, together with our temporal follow-up\, addresses semi-supervised 2D and 3D detection\, aligning detections across modalities and video to produce consistent pseudo-labels and more stable tracklets\, but it focuses on foreground entities and does not model background. S2GO proposes a streaming query-based representation for semantic occupancy estimation that is entity-centric\, temporal\, and models both foreground and background: each persistent query decodes to semantic Gaussians\, and the state is carried across frames and supports short-horizon future prediction. This gives us a single\, stable representation suitable for both perception and sampling. \nWe propose two projects that make this representation generative. First\, we propose a static scene generation method: a diffusion model over grounded queries that represent both foreground and background and are decoded into Gaussians\, generating a complete semantic occupancy scene. This grounded latent representation enables intuitive\, consistent control. Then\, we propose motion generation: a model that generates trajectories for ego and foreground entities conditioned on the generated static scene\, producing coherent 4D rollouts and enabling interactive edits.\n \nThesis Committee Members:\nKris Kitani (Chair)\nDeva Ramanan\nShubham Tulsiani\nWei-Chiu Ma (Cornell University)\n \nLink to Proposal Draft: Link
URL:https://www.ri.cmu.edu/event/jinhyung-park-phd-thesis-proposal/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251014T140000
DTEND;TZID=America/New_York:20251014T153000
DTSTAMP:20260922T105618
CREATED:20250929T144547Z
LAST-MODIFIED:20251030T161352Z
UID:148937-1760450400-1760455800@www.ri.cmu.edu
SUMMARY:Embodied Artificial Intelligence for Emergency Care in Unstructured Environments
DESCRIPTION:Abstract: \nIn mass casualty events and resource-constrained scenarios\, limited responder capacity leads to preventable deaths. Time is of the essence particularly in severe trauma: the sooner individuals receive care\, the higher their chances of survival. Yet a single responder can only manage a few patients simultaneously\, leaving others unattended. This thesis addresses this capacity constraint by developing intelligent robotic systems that serve as medical force multipliers\, enabling effective emergency response when casualties outnumber available help. \nThis work presents two embodied Artificial Intelligence (AI) platforms for emergency medical response in unstructured field environments. The first performs multipatient automated assessment\, which uses contactless multimodal sensors to identify qualitative (e.g.\, wounds\, amputations\, hemorrhage\, respiratory distress) and quantitative vital signs (e.g.\, heart rate) to rapidly assess and prioritize the most critically injured. The second automates fluid resuscitation\, targeting hemorrhage\, the leading cause of preventable death in trauma. The pipeline comprises multiple stages: vessel localization and segmentation\, visualization and uncertainty quantification for safe decision-making\, bifurcation detection for anatomically-informed needle placement\, and real-time needle tracking. \nTo address the scarcity of training data in emergency medicine robotics\, this research embeds expert clinical knowledge and applies weak supervision techniques\, enabling robust performance with limited labeled examples. All algorithms execute in real time on resource-constrained platforms\, with key components designed to adapt to changing environmental conditions. \nThis thesis contributes to autonomous medical systems and offers new methodologies for developing AI solutions for life-critical applications in unstructured environments where traditional data-driven approaches may fail. By augmenting human responders\, we show how robotic systems can expand treatment capacity when it matters most\, potentially saving lives that would otherwise be lost. \nThesis Committee: \nArtur Dubrawski\, Chair \nJean Oh \nFernando de la Torre Frade \nDaniel McDuff\, Google \nLaura Brattain\, University of Central Florida \nDocument Link
URL:https://www.ri.cmu.edu/event/embodied-artificial-intelligence-for-emergency-care-in-unstructured-environments/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251015T130000
DTEND;TZID=America/New_York:20251015T140000
DTSTAMP:20260922T105618
CREATED:20251008T135546Z
LAST-MODIFIED:20251008T140726Z
UID:149052-1760533200-1760536800@www.ri.cmu.edu
SUMMARY:Seeing Deep Inside Scattering Tissue Using Efficient\, Noise-Robust Wavefront Shaping
DESCRIPTION:Abstract:\nScattering limits our ability to see inside biological tissue\, as light penetration is severely distorted by tissue components with varying refractive indices. One promising method to overcome scattering aberration is wavefront shaping. This technique involves placing a spatial light modulator (SLM) in the microscope’s optical path to correct the wavefront emitted from a point deep within the tissue. The goal is to bring light photons from a single target point to a single sensor point\, despite tissue aberrations. This technique has the potential to revolutionize tissue imaging by enabling high-SNR imaging deep within scattering biological targets. However\, estimating wavefront-shaping modulations in practice is challenging\, since the modulations must be estimated in real time\, using non-invasive feedback\, and under a low photon budget. \nIn the first part of this talk\, I will discuss efforts to derive noise-robust score functions that can identify effective modulation corrections using non-invasive feedback. I will review previous approaches and introduce a new\, simple\, noise-robust method that uses confocal correction of both incoming and outgoing light with linear single-photon fluorescent excitation. We show that despite the fact that we are only measuring light outside the tissue and have no direct way to measure how well light has focused inside the tissue\, maximizing the single-photon confocal intensity guarantees that we also focus all light into a spot inside the tissue. \nGiven a score function\, estimating the desired modulation becomes an optimization problem. However\, since the desired modulation depends on the unknown tissue structure\, typical optimization strategies involve slow sequential scanning\, where each modulation parameter is queried independently. In the second part of this talk\, I will present a novel approach for rapid modulation optimization. This method leverages optical computing ideas and uses the optical system to directly measure the gradient of the score function\, allowing simultaneous updates of all modulation parameters from a single measurement. \nBio: Anat Levin is a Professor at the department of Electrical and Computer Engineering\, Technion\, Israel\, doing research in the field of computational imaging. She received a Ph.D. in computer science from the Hebrew University in 2006. During the years 2007- 2009 she was a postdoc at MIT CSAIL\, and during 2009-2016 she was an Assistant and Associate Prof. at the department of Computer Science and Applied Math\, the Weizmann Inst. of Science.\nProf. Levin has received numerous awards for her research\, including the CVRP PAMI young researcher award in 2013; the eurographics young researcher award in 2010; the eurographics outstanding technical contributions award in 2024; the Blavatnik award in 2018; and 3 ERC grants. \nHomepage:  https://webee.technion.ac.il/people/anat.levin/ \n  \nSponsor \nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one AI-powered video generation platform that leverages advances in computer vision\, generative modeling\, and multimodal learning to make high-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/seeing-deep-inside-scattering-tissue-using-efficient-noise-robust-wavefront-shaping/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/10/10-15-25.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251020T150000
DTEND;TZID=America/New_York:20251020T163000
DTSTAMP:20260922T105618
CREATED:20251014T201636Z
LAST-MODIFIED:20251014T201948Z
UID:149083-1760972400-1760977800@www.ri.cmu.edu
SUMMARY:Unconstrained Perception for Scalable Robot Manipulation
DESCRIPTION:Abstract: Advances in visual imitation learning driven by large-scale data and expressive policy architectures have yielded impressive progress on long-horizon\, dexterous tasks. However\, current success rates remain insufficient for industrial deployment\, which demands near-perfect reliability on novel tasks. Compared to other fields such as NLP and CV\, the available data in robotics is several orders of magnitude smaller. This raises the question: how can we most effectively leverage priors from large-scale offline data? In this thesis\, I contribute methods to infer strong geometric and dynamic priors for robot manipulation. \nFirst\, geometric camera calibration is a critical prerequisite for real-world vision systems. I will discuss our work on MASt3R-SfM for unconstrained SfM from any image collection in linear complexity. Next\, I discuss how we use a large set of calibrated cameras in DeformGS for photorealistic digital twins with millimeter-accurate tracking of deformable cloth. Removing the need for costly multi-camera systems\, I introduce RaySt3R\, a method to generate complete object geometry from a single RGB-D image. \nBuilding on these works\, I will introduce our ongoing work on Flow2Flow – a flexible end-to-end feedforward architecture for zero-shot dynamics prediction. Many challenging tasks involve manipulating unseen articulated\, deformable\, and cluttered rigid objects; prior approaches rely on pre-trained VLMs\, large-scale 2D point tracking\, or previous interactions with the scene to inject priors. We cast dynamics prediction as a scene flow completion problem from a single RGB-D image\, and propose an optional two-stage adaptation procedure for unseen dynamics. We further study scene flow completion as a 3D pretraining objective for multi-task learning and propose scaling up training on real-world data for the first benchmark in dynamics prediction from a single image. \nThesis Committee Members:\nJeffrey Ichnowski (Chair)\nDeva Ramanan\nShubham Tulsiani\nAbhishek Gupta (University of Washington) \nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/unconstrained-perception-for-scalable-robot-manipulation/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251021T100000
DTEND;TZID=America/New_York:20251021T110000
DTSTAMP:20260922T105618
CREATED:20251007T172737Z
LAST-MODIFIED:20251007T172737Z
UID:149047-1761040800-1761044400@www.ri.cmu.edu
SUMMARY:Adaptive Robot Design for multimodal locomotion across diverse terrains
DESCRIPTION:Abstract: \nLocomotion across natural environments such as sand\, mud\, and water presents a fundamental challenge for robots due to the heterogeneous\, deformable\, and often unpredictable properties of these substrates. In this talk\, I will share how mechanical and structural adaptation can enable robust mobility in such complex settings through the development and characterization of two centimeter-scale robotic systems that leverage distinct modes of morphological adaptation. \nFirst\, TerraSkipper\, a mudskipper-inspired robot that integrates a spring-steel tail and magnetically encoded fins to achieve impulsive skipping and controlled crawling across a wide range of granular and muddy substrates. Through systematic experiments that vary substrate composition and moisture content\, we show that impulsive tail-driven locomotion achieves higher velocities and improved mobility where conventional frictional gaits fail\, providing new insights into substrate robot interactions at the centimeter scale. \nSecond\, PuffyBot\, an amphibious shape morphing robot that employs a scissor-lift mechanism\, coupled fins\, and a waterproof skin to actively modulate its volume and buoyancy. Our experimental results demonstrate multimodal locomotion\, including crawling on the land\, crawling on the underwater floor\, swimming on the water surface\, and bimodal buoyancy adjustment to submerge underwater or resurface. \nTogether\, these systems illustrate how embodied mechanical intelligence through energy-based and geometry-based adaptation extends the operational range of small robots across a wide range of terrestrial and aquatic domains. This work advances the understanding of morphology as a design variable in robot locomotion and lays the foundation for autonomous\, terrain-adaptive robotic platforms capable of robust operation in unstructured natural environments. \nCommittee: \nProf. Zeynep Temel \nProf. Sarah Bergbreiter \nProf. Guanya Shi \nRishi Veerapaneni
URL:https://www.ri.cmu.edu/event/adaptive-robot-design-for-multimodal-locomotion-across-diverse-terrains/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251024T090000
DTEND;TZID=America/New_York:20251024T110000
DTSTAMP:20260922T105618
CREATED:20251001T150706Z
LAST-MODIFIED:20251016T170428Z
UID:148983-1761296400-1761303600@www.ri.cmu.edu
SUMMARY:Title: Leveraging Geometric Priors for Robust Robotic Manipulation
DESCRIPTION:Abstract: \nThis thesis explores how explicit 3D geometric representations\, trained at scale on synthetic data\, can serve as priors to enhance robotic manipulation. Even with recent progress in geometric understanding\, generalization to unseen objects and environments remains constrained by the scale and diversity of existing 3D training data. Although more large-scale 3D datasets have been released\, their sizes are still considerably smaller than their image and language counterparts. In addition\, collecting diverse real-world 3D data is time-consuming and labor-intensive\, limiting the coverage of objects and scenes. To tackle this challenge\, this thesis explores how geometric understanding learned from large-scale synthetic 3D model datasets can improve generalization in robotic manipulation without further expanding real-world 3D training data. \nAs a first step\, Chapter 2 introduces RePOSE\, a fast and accurate 6D object pose refinement method that establishes a foundation for scalable geometric perception. Chapter 3 and Chapter 4 frame the acquisition of a universal geometric prior as a supervised learning problem on 3D geometry tasks. We propose two frameworks\, OctMAE and ZeroGrasp\, which learn a geometric prior through shape reconstruction and grasp pose prediction. We also introduce ZeroGrasp-11B\, a large-scale synthetic dataset containing 1M RGB-D images\, 12K 3D models\, and 11B grasps\, specifically designed for training such models. These methods achieve state-of-the-art performance on both shape reconstruction and grasp pose prediction of unseen objects on public benchmarks\, demonstrating the strength of the learned geometric prior. Real-world pick-and-place experiments further validate its generalization to practical robotic scenarios. \nWhile the learned geometric prior shows strong performance in pick-and-place tasks\, robotic manipulation involves a broader range of behaviors and longer temporal horizons. In Chapter 5\, we focus on integrating this prior into imitation learning to address more complex\, long-horizon tasks. To this end\, we propose GeoFlow\, a framework for flow-based 3D visuomotor policy learning that leverages geometry-aware pre-trained models as strong priors. GeoFlow achieves state-of-the-art performance across diverse benchmarks and demonstrates improved data efficiency and robustness under clutter and distractors\, highlighting that large-scale geometry pre-training and sparse voxel representations are key to scalable and generalizable robotic learning. \n\n\nThesis Committee: \nKris Kitani\, Chair \nDavid Held \nShubham Tulsiani \nSergey Zakharov\, Toyota Research Institute \nDocument Link
URL:https://www.ri.cmu.edu/event/shun-iwase-phd-thesis-defense/
LOCATION:Gates Hillman Center 6115
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251024T120000
DTEND;TZID=America/New_York:20251024T130000
DTSTAMP:20260922T105618
CREATED:20251015T183546Z
LAST-MODIFIED:20251015T183546Z
UID:149109-1761307200-1761310800@www.ri.cmu.edu
SUMMARY:KOROL: Learning Visualizable Object Feature with Koopman Operator Rollout for Manipulation
DESCRIPTION:Abstract: \nHumans possess an extraordinary ability to manipulate objects\, discerning position\, shape\, and other properties with just a glance. How can robots be endowed with similar perceptual and dexterous manipulation capabilities? In this talk\, I will present a method that combines the sample efficiency of traditional model-based approaches with the high generalizability of deep learning methods to tackle dexterous manipulation tasks. I will begin with a brief introduction to the model-based framework grounded in Koopman Operator Theory\, highlighting its reliance on ground-truth (GT) object states in real-world applications. \nTo address this limitation\, we propose Koopman Operator Rollout for Object Feature Learning (KOROL)—an approach that removes the dependency on GT states in model-based manipulation learning. KOROL learns visual features that predict robot states throughout dynamics model rollouts. Unlike prior approaches that learn implicit visual features for direct image-to-action policies\, KOROL explicitly trains on object-centric visual representations\, encoding essential scene information to improve robot state predictions during autoregressive rollouts. This establishes a synergistic relationship between the learned object features and the Koopman operator. \nOur experiments demonstrate that KOROL: (i) improves performance across various simulated manipulation tasks compared to Koopman operators using GT object states and other baselines\, (ii) extends Koopman-based methods to vision-based real-world tasks\, and (iii) enables multitasking through dimensionally aligned object features. \nCommittee: \nProf. Jeffrey Ichnowski \nProf. Guanya Shi \nProf. Oliver Kroemer \nBardienus Duisterhof
URL:https://www.ri.cmu.edu/event/adaptive-robot-design-for-multimodal-locomotion-across-diverse-terrains-2/
LOCATION:GHC 6115
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251024T143000
DTEND;TZID=America/New_York:20251024T153000
DTSTAMP:20260922T105618
CREATED:20250902T163833Z
LAST-MODIFIED:20251024T220347Z
UID:148668-1761316200-1761319800@www.ri.cmu.edu
SUMMARY:Bringing Dexterity to Robot Hands in the Real World
DESCRIPTION:Abstract:  Dexterous manipulation is a grand challenge of robotics\, and fine manipulation skills are required for many robotics applications that we envision.   In this overview talk\, I will discuss my view of some major factors that contribute to dexterity and discuss how we can incorporate them into our robots and systems.\n\nBio:  Nancy Pollard is a Professor in the Robotics Institute and the Computer Science Department at Carnegie Mellon University. She received her PhD in Electrical Engineering and Computer Science from the MIT Artificial Intelligence Laboratory\, where she developed grasp and manipulation planning algorithms for the Stanford/JPL and Utah/MIT dexterous hands. Prof. Pollard spent the next few decades studying human and robot dexterity\, with emphasis on bringing human manipulation strategies with performance guarantees to humanoid robots with dexterous hands.  She received the NSF CAREER award for research on “Quantifying Humanlike Enveloping Grasps”\,  the Okawa Research Grant for “Studies of Dexterity for Computer Graphics and Robotics\,” and was a recent recipient of an NSF Convergence Accelerator award for “Bio-Inspired Design of Robot Hands for Use-Driven Dexterity.”   She has led the development of several generations of dexterous soft robotic hands\, is a founder of FuturHand Robotics and leads the CMU Foam Hands Laboratory.
URL:https://www.ri.cmu.edu/event/bringing-dexterity-to-robot-hands-in-the-real-world/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/nsp-crop.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251030T113000
DTEND;TZID=America/New_York:20251030T123000
DTSTAMP:20260922T105618
CREATED:20251024T134300Z
LAST-MODIFIED:20251024T134300Z
UID:149169-1761823800-1761827400@www.ri.cmu.edu
SUMMARY:3D Thermal Perception for Autonomous Navigation in Visually Degraded & Unstructured Environments
DESCRIPTION:Abstract:\nAutonomous navigation in visually degraded and unstructured environments\, such as darkness\, smoke\, and rough off-road terrain remains a significant challenge for current robotic systems. RGB cameras fail without illumination\, and active sensors such as LiDAR degrade under aerosols and emit signals that are undesirable in sensitive or adversarial scenarios. In contrast\, long-wave infrared (thermal) sensing captures naturally emitted radiation and maintains visibility through many atmospheric obscurants. This thesis develops a framework for passive thermal autonomy\, enabling 3D stereo perception and autonomous navigation in challenging off-road terrain without active illumination. \nThe absence of large-scale thermal datasets\, sensor-to-perception pipeline\, and field deployments has long prevented reliable passive autonomy in low-visibility\, off-road environments. This work addresses the complete pipeline\, from rigorous sensor integration and custom cross-modality calibration\, through large-scale data collection\, to stereo thermal mapping and odometry\, and autonomy deployment. We introduce MACThermal\, which adapts metric- and uncertainty-aware covariance visual odometry to the thermal domain. It employs dynamic-range normalization and geometry-consistent flow augmentation to improve correspondence and uncertainty estimation. To address the critical gap in thermal 3D vision benchmarks\, we contributed two multi-modal datasets: FIReStereo for aerial platforms and TartanDrive 2.5T for ground vehicles\, covering diverse off-road terrains and visibility conditions. We integrate these perception modules into a complete autonomy stack\, utilizing visual foundation models for semantic understanding and self-supervised traversability estimation for adaptive decision-making. The full system is validated on a full-scale ATV platform deployed in previously unseen off-road terrain under complete darkness. Together\, this work demonstrates a shift from active LiDAR-based navigation to passive thermal autonomy\, enabling over 10x longer nighttime traversals in previously inaccessible terrain. \nCommittee:\nSebastian Scherer (advisor)\nSrinivasa Narasimhan\nNikhil Keetha
URL:https://www.ri.cmu.edu/event/3d-thermal-perception-for-autonomous-navigation-in-visually-degraded-unstructured-environments/
LOCATION:Gates Hillman Center 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251031T143000
DTEND;TZID=America/New_York:20251031T153000
DTSTAMP:20260922T105618
CREATED:20250902T164414Z
LAST-MODIFIED:20251030T145800Z
UID:148671-1761921000-1761924600@www.ri.cmu.edu
SUMMARY:Toward Generalist Humanoid Robots: Recent Advances\, Opportunities\, and Challenges
DESCRIPTION:Abstract: In an era of rapid AI progress\, leveraging accelerated computing and big data has unlocked new possibilities to develop generalist AI models. As AI systems like ChatGPT showcase remarkable performance in the digital realm\, we are compelled to ask: Can we achieve similar breakthroughs in the physical world — to create generalist humanoid robots capable of performing everyday tasks? In this talk\, I will outline our data-centric research principles and approaches for building general-purpose robot autonomy in the open world. I will present our recent work leveraging real-world\, synthetic\, and web data to train foundation models for humanoid robots. Furthermore\, I will discuss the opportunities and challenges of building the next generation of intelligent robots.\n\nBio: Yuke Zhu is an Associate Professor in the Computer Science Department of UT-Austin\, where he directs the Robot Perception and Learning (RPL) Lab. He is also a Director and Distinguished Research Scientist at NVIDIA Research\, where he co-leads the Generalist Embodied Agent Research (GEAR) lab. He focuses on developing intelligent algorithms for generalist robots and embodied agents to reason about and interact with the real world. He obtained his Ph.D. degree from Stanford University. He received the NSF CAREER Award\, the IEEE RAS Early Academic Career Award\, and various faculty fellowships and research awards from Amazon\, JP Morgan\, and Sony Research.
URL:https://www.ri.cmu.edu/event/toward-generalist-humanoid-robots-recent-advances-opportunities-and-challenges/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/yukezhu.jpg
END:VEVENT
END:VCALENDAR