BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251103T153000
DTEND;TZID=America/New_York:20251103T163000
DTSTAMP:20260922T134536
CREATED:20251027T174614Z
LAST-MODIFIED:20251027T174614Z
UID:149194-1762183800-1762187400@www.ri.cmu.edu
SUMMARY:From Video Generation to Video World Models
DESCRIPTION:Abstract:\nVideo diffusion models have achieved remarkable success in content creation\, yet they still fall short of simulating interactive worlds that respond to users in real time. This talk examines the fundamental challenges preventing these models from evolving into true world simulators. I will present a series of works — CausVid\, Self-Forcing\, MotionStream\, and State-Space World Model — that collectively mark a paradigm shift from non-causal diffusion models to autoregressive–diffusion hybrids capable of streaming long-duration videos with real-time interactivity. These advances move beyond passive video generation toward dynamic\, immersive experiences\, unlocking new possibilities across gaming\, robotics\, live video editing\, and augmented/virtual reality. \nBio: Xun Huang was a Research Scientist at Adobe\, NVIDIA\, as well as an Adjunct Professor at Carnegie Mellon University. He is currently the Founder and CEO of a stealth startup. He obtained his Ph.D. from Cornell University in 2020 under the advisement of Professor Serge Belongie. His doctoral research was recognized with the Fellowship from NVIDIA\, Adobe\, and Snap. His research interests lie broadly in deep generative models\, with a recent focus on video world models. \nHomepage:  xunhuang.me \nSponsor \nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one AI-powered video generation platform that leverages advances in computer vision\, generative modeling\, and multimodal learning to make high-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/from-video-generation-to-video-world-models/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/10/11-3-25.jpeg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251103T160000
DTEND;TZID=America/New_York:20251103T173000
DTSTAMP:20260922T134536
CREATED:20251027T134212Z
LAST-MODIFIED:20251027T134212Z
UID:149192-1762185600-1762191000@www.ri.cmu.edu
SUMMARY:Unifying Perception and Creation with Generative Models
DESCRIPTION:Abstract: \nRecent advances in large-scale generative modeling have reshaped our understanding of visual intelligence. While models such as diffusion and autoregressive transformers have achieved remarkable success in image and video synthesis\, their potential for visual perception and understanding remains underexplored. This thesis investigates how generative models can serve as powerful visual learners—bridging the long-standing divide between generative and discriminative paradigms. \nWe begin with REM (Refer Everything Models)\, a framework for referring video segmentation built upon text-to-video diffusion models. By preserving generative representations and fine-tuning on narrow-domain datasets\, REM achieves state-of-the-art results on standard benchmarks and demonstrates strong generalization to unseen domains. \nBuilding on this foundation\, we introduce a unified perceptual–generative framework that repurposes a single diffusion model across a broad spectrum of computer vision and image restoration tasks. Through joint training and systematic evaluation over 15 tasks spanning perception and synthesis\, we show that diffusion-based models deliver superior or comparable performance to discriminative counterparts\, revealing their intrinsic ability to encode rich\, multi-modal world representations. \nFinally\, we extend our exploration to visual autoregressive (VAR) models\, presenting the first unified architecture capable of efficiently solving the same 15 tasks within a single framework. We show that latent-variable designs\, particularly those leveraging variational autoencoders\, are key to achieving coherent multi-modal understanding and consistent generation. Compared with diffusion counterparts\, VAR-based models offer substantial gains in latency\, scalability\, and output consistency. \nCollectively\, these studies offer a cohesive perspective on unifying perception and synthesis through generative modeling\, charting a path toward general-purpose visual foundation models that seamlessly integrate understanding\, reasoning\, and creation. \n\nThesis Committee: \nMartial Hebert\, Chair \nDeva Ramanan \nJun-Yan Zhu \nAlexei Efros\, University of California\, Berkeley \nYu-Xiong Wang\, University of Illinois Urbana-Champaign \nPavel Tokmakov\, Toyota Research Institute \nDraft of Document
URL:https://www.ri.cmu.edu/event/unifying-perception-and-creation-with-generative-models/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251107T150000
DTEND;TZID=America/New_York:20251107T163000
DTSTAMP:20260922T134536
CREATED:20251001T150038Z
LAST-MODIFIED:20251104T153929Z
UID:148979-1762527600-1762533000@www.ri.cmu.edu
SUMMARY:Visual-Tactile Synthesis for Texture Generation
DESCRIPTION:Abstract: Recent advances in generative models have enabled the creation of highly realistic visual content\, yet they remain limited to visual perception alone. In contrast\, human interaction with the physical world is inherently multimodal — we not only see textures but also feel them. This gap motivates the goal of my thesis: to build generative models that jointly synthesize visual and tactile modalities for material and texture generation. By unifying what we see and what we touch\, such models can drive new forms of physically grounded content creation\, from robotics and virtual reality to material design.\nHowever\, extending generative modeling to touch presents unique challenges: tactile data is scarce\, noisy\, and expensive to collect\, and there is no large-scale paired dataset linking visual appearance with tactile response. To address these challenges\, my research explores three synergistic directions. \nPart I: I introduce controllable visual-tactile synthesis models that jointly generate aligned visual and tactile textures from shared latent representations\, enabling explicit control over appearance and feel. \nPart II: I propose tactile-aware 3D generation frameworks that integrate tactile sensing into 3D diffusion pipelines\, allowing models to infer physically grounded material properties from visual cues and geometry. \nPart III: Building on these insights\, I aim to develop scalable multimodal generation systems that leverage large vision and language foundation models and physics priors to synthesize novel materials directly from text or image input\, without relying on extensive paired tactile data. \n\n \nThesis Committee:\nJun-Yan Zhu (Co-chair)\nWenzhen Yuan (Co-chair)\nShubham Tulsiani\nAndrew Owens (Cornell Tech)\n\nThesis Document
URL:https://www.ri.cmu.edu/event/ruihan-gao-phd-thesis-proposal/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251114T143000
DTEND;TZID=America/New_York:20251114T153000
DTSTAMP:20260922T134536
CREATED:20250902T170114Z
LAST-MODIFIED:20251125T172342Z
UID:148675-1763130600-1763134200@www.ri.cmu.edu
SUMMARY:Just Asking Questions
DESCRIPTION:Abstract: In the age of deep networks\, “learning” almost invariably means “learning from examples”. We train language models with human-generated text and labeled preference pairs\, image classifiers with large datasets of images\, and robot policies with rollouts or demonstrations. When human learners acquire new concepts and skills\, we often do so with richer supervision\, especially in the form of language—we learn new concepts from examples accompanied by descriptions or definitions\, and new skills from demonstrations accompanied by instructions. Crucially\, language-based supervision involves not only instructions but *questions*—students ask questions to elicit the most useful pieces of supervision\, and teachers ask questions to probe student knowledge and encourage them to acquire new skills or aspects of understanding on their own. This talk will focus on a few recent projects focused on building computational models that can ask good questions for both learning and teaching\, with applications spanning LM alignment\, policy learning\, and education. This is joint work with Belinda Li\, Alex Tamkin\, Noah Goodman\, Andi Peng\, Ilia Sucholutsky\, Nishanth Kumar\, Julie A Shah\, Andreea Bobu\, Alexis Ross\, Gabe Grand\, Valerio Pepe and Josh Tenenbaum. \nBio: Jacob Andreas is an associate professor at MIT in the Department of Electrical Engineering and Computer Science as well as the Computer Science and Artificial Intelligence Laboratory. His research aims to understand the computational foundations of language learning\, and to build intelligent systems that can learn from human guidance. Jacob earned his Ph.D. from UC Berkeley\, his M.Phil. from Cambridge (where he studied as a Churchill scholar) and his B.S. from Columbia. He has received a Sloan fellowship\, an NSF CAREER award\, MIT’s Junior Bose and Kolokotrones teaching awards\, and paper awards at ACL\, ICML and NAACL.
URL:https://www.ri.cmu.edu/event/just-asking-questions/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/head_small.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251117T103000
DTEND;TZID=America/New_York:20251117T113000
DTSTAMP:20260922T134536
CREATED:20251113T150929Z
LAST-MODIFIED:20251113T150929Z
UID:149416-1763375400-1763379000@www.ri.cmu.edu
SUMMARY:Towards Modernization of Long-Range Image-Space Planning for Off-Road Navigation
DESCRIPTION:Abstract: \nThis thesis revisits long-range\, image-space planning for off-road navigation and modernizes the classical first-person view (FPV) paradigm by building upon recent advances in perception. It introduces a lightweight depth calibration scheme\, analytic configuration-space (C-space) transforms\, interpretable frontier selection\, and a pixel-space A* planner with validated heuristic soundness. Concretely\, we (i) make monocular depth metrically usable at test time via an affine\, log-domain calibration with sparse LiDAR; (ii) derive and implement a depth-aware FPV C-space inflation that projects vehicle width/length analytically and realizes it with separable row/column sliding-maximum filters augmented by per-pixel depth consistency checks; (iii) propose transparent\, angular-sector frontiering that reasons jointly about traversability cost and minimal lethal depth\, alongside goal-aware revalidation; and (iv) preserve A* admissibility/consistency in image space through a simple cost renormalization that avoids silent suboptimality in low-cost free space. \nWe evaluate the resulting\, modular sub-stack in the high-fidelity Falcon simulator [2] under a shared ROS graph. Using a common perception front-end and planning back-end\, we compare three frontiering strategies — (1) a LAGR-style row-wise baseline\, (2) an LRN-inspired openness heuristic adapted to operate with explicit depth and cost\, and (3) an Angular Cost & Depth (ACD) variant that couples average cost with a minimum lethal-depth constraint. Across diverse courses (e.g.\, farm\, desert\, mixed terrain)\, the calibrated monocular depth reduces error versus raw monocular predictions\, and both LRN-inspired and ACD frontiering tend to outperform the purely row-wise baseline on longer traverses. We view these as encouraging indications rather than definitive claims: all results are in simulation\, with performance subject to perception quality\, calibration coverage\, and environment diversity. \n\nScope and limitations are explicit. The work was conducted over a short project window (May 2025–present) and prioritized stabilizing the proposed sub-stack in conjunction with a core FieldAI [1] stack through the high-fidelity Falcon simulator [2] and ROS integration into a reliable\, end-to-end testing framework. No real-world deployments were performed\, and the current implementation targets ROS1. Nevertheless\, the design is intentionally modular and auditable to make it relatively straightforward to integrate the sub-stack into existing off-road autonomy stacks lacking explicit long-range planning. Such integration will enable better autonomy by offering a practical bridge between classical image-space efficiency and metric-world robustness. \n[1] FieldAI: https://www.fieldai.com \n[2] Falcon: https://www.duality.ai \nCommittee: \nWenshan Wang\, Chair\nSebastian Scherer\, Co-Chair\nMaxim Likhachev\nSamuel Triest\, RI PhD
URL:https://www.ri.cmu.edu/event/towards-modernization-of-long-range-image-space-planning-for-off-road-navigation/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251119T130000
DTEND;TZID=America/New_York:20251119T140000
DTSTAMP:20260922T134536
CREATED:20251114T143444Z
LAST-MODIFIED:20251114T145002Z
UID:149445-1763557200-1763560800@www.ri.cmu.edu
SUMMARY:Building Robot Hands and Teaching Dexterity
DESCRIPTION:Abstract:  \nOur human hands are masterpieces of power and precision\, capable of typing\, hammering\, or delicately using chopsticks. Yet most robots today still rely on simple two-finger grippers in controlled settings because dexterous hands are costly and difficult to deploy. To close this gap\, I will introduce my LEAP Hands\, high-performance\, low-cost\, and easy-to-assemble robotic hands that have become the most widely used platform for dexterous manipulation research. LEAP Hand V1 employs motor-in-joint actuation for simplicity\, while V2 introduces a novel hybrid rigid–soft structure that delivers exceptional strength and durability.  I will then show how large-scale human video/motion data and simulation techniques can teach human-like manipulation skills across diverse environments.  By tightly integrating mechanical design and machine learning\, my open-source robot hands achieve unprecedented levels of dexterity for a variety of everyday tasks. \nCommittee:\nProf. Deepak Pathak (advisor)\nProf. Nancy Pollard \nProf. Abhinav Gupta \nProf. Jitendra Malik \nDr. Ankur Handa \n  \nA draft of my thesis proposal is available here: \nhttps://kennyshaw.net/phd_thesis_proposal.pdf
URL:https://www.ri.cmu.edu/event/building-robot-hands-and-teaching-dexterity-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251120T140000
DTEND;TZID=America/New_York:20251120T150000
DTSTAMP:20260922T134536
CREATED:20251114T205305Z
LAST-MODIFIED:20251114T205305Z
UID:149469-1763647200-1763650800@www.ri.cmu.edu
SUMMARY:Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models
DESCRIPTION:Abstract:\nTransferring skills between different objects remains one of the core challenges of open-world robot manipulation. Generalization needs to take into account the high-level structural differences between distinct objects while still maintaining similar low-level interaction control. In this paper\, we propose an example-based zero-shot approach to skill transfer. Rather than treating skills as atomic\, we decompose skills into a prioritized list of grounded task-axis (GTA) controllers. Each GTAC defines an adaptable controller\, such as a position or force controller\, along an axis. Importantly\, the GTACs are grounded in object key points and axes\, e.g.\, the relative position of a screw head or the axis of its shaft. Zero-shot transfer is thus achieved by finding semantically similar grounding features on novel target objects. We achieve this example-based grounding of the skills through the use of foundation models\, such as SD-DINO\, that can detect semantically similar keypoints of objects. We evaluate our framework on real-robot experiments\, including screwing\, pouring\, and spatula scraping tasks\, and demonstrate robust and versatile controller transfer for each.\n\nCommittee:\nProf. Oliver Kroemer\nProf. Katerina Fragkiadaki\nProf. Zeynep Temel\nMark Lee
URL:https://www.ri.cmu.edu/event/grounded-task-axes-zero-shot-semantic-skill-generalization-via-task-axis-controllers-and-visual-foundation-models/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251120T163000
DTEND;TZID=America/New_York:20251120T173000
DTSTAMP:20260922T134536
CREATED:20251112T161204Z
LAST-MODIFIED:20251112T161723Z
UID:149408-1763656200-1763659800@www.ri.cmu.edu
SUMMARY:OpenVDB
DESCRIPTION:Abstract: As the inventor of VDB and founder of OpenVDB\, I am excited to talk about its history\, motivation\, and diverse adoption. Specifically\, this lecture will cover the underlying VDB data structure\, and its adoption to computer graphics\, physics simulations and more recently machine learning. Since its open-source release in 2012\, OpenVDB has become an industry standard and has been used in numerous VFX franchises like “Avatar”\, “Avengers”\, “The Mummy”\, “Pirates of the Caribbean”\, “Kung Fu Panda”\, and “How to Train Your Dragon”. It is adopted by numerous commercial software packages used by the entertainment industry\, including Houdini\, RenderMan\, Arnold\, Blender\, and Unreal Engine\, just to mention a few. OpenVDB has also found use in many areas outside of media and entertainment\, including SLAM\, autonomous driving\, topology optimization\, semiconductor designs\, 3D printing\, medical imaging\, rocket design\, aerial surveillance\, robotics\, and many machine learning applications. Finally\, OpenVDB was the first open-source project to be adopted by the Academy Software Foundation (ASWF) and the Linux Foundation (in 2018). \nBio: Ken Museth is Sr Director of the High-Fidelity Physics Research team at Nvidia and chair of the Technical Steering Committee for OpenVDB under the Academy Software Foundation. He has a PhD in quantum physics from Copenhagen University and did his postgraduate studies in computer science at Caltech. Previously he was a Sr. Computational Scientist at SpaceX for six years\, working on CFD simulations of the Raptor engine\, Head of Simulation R&D at Weta FX for three years\, working on James Cameron’s Avatar 2\, director for FX and CFX simulation teams at DreamWorks Animation for eight years\, Sr Software Engineer at Digital Domain for three years\, full tenured professor in computer graphics at Linköping University for four years\, where he supervised five PhD and 15 MSc students\, and research scientist at NASA’s Jet Propulsion Laboratory for three years\, working on space-mission design and visualization. Ken invented and founded OpenVDB for which he won two Academy Awards from the Academy of Motion Picture Arts and Sciences; a Technical Achievement Award (Academy Certificate) in 2015 and a Scientific & Engineering Award (Academy Plaque) in 2024. Additionally\, in 2023 he was awarded the ACM SIGGRAPH Practitioner Award and was accepted into the ACM SIGGRAPH Academy. Ken has served on the Technical Papers Committee for ACM SIGGRAPH multiple times and has 29 movie credits\, including on franchises like “Avatar”\, “Avengers”\, “The Mummy”\, “Pirates of the Caribbean”\, “Kung Fu Panda”\, and “How to Train Your Dragon”. \nWebsite: https://research.nvidia.com/labs/prl/author/ken-museth/ \nPoster:  https://graphics.cs.cmu.edu/cmgc/ken-museth.pdf \nThe Carnegie Mellon Graphics Colloquium is hosted by the Carnegie Mellon Graphics Lab and supported by Meta and Adobe.
URL:https://www.ri.cmu.edu/event/openvdb/
LOCATION:Gates-Hillman Center 4401
CATEGORIES:Special Events
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/11/11-20-25-scaled.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251121T150000
DTEND;TZID=America/New_York:20251121T160000
DTSTAMP:20260922T134536
CREATED:20250915T203414Z
LAST-MODIFIED:20251125T191630Z
UID:148862-1763737200-1763740800@www.ri.cmu.edu
SUMMARY:How to Coordinate Thousands of Robots Efficiently and Robustly
DESCRIPTION:Abstract: \nLarge-scale robot fleets are increasingly deployed in warehouses\, factories\, transportation systems\, and emerging robotics applications. Coordinating hundreds or thousands of robots in shared\, cluttered spaces creates fundamental challenges in maintaining safety\, preventing deadlocks\, and minimizing congestion. In this talk\, I will present our recent work on scalable imitation learning methods for coordinating 10k robots\, automatic environment optimization techniques for alleviating traffic congestion\, and asynchronous execution frameworks that guarantee safe\, robust\, and deadlock-free multi-robot operations.\n\nBio: \nJiaoyang Li is an Assistant Professor in the Robotics Institute at Carnegie Mellon University. Her research lies at the intersection of artificial intelligence\, robotics\, and optimization\, with a particular focus on scalable multi-robot planning and coordination. She received her Ph.D. in Computer Science from the University of Southern California in 2022 and her B.Eng. in Automation from Tsinghua University in 2017. She is a recipient of the NSF CAREER Award\, best dissertation awards from ICAPS\, AAMAS\, and USC\, and multiple best paper awards from ICRA\, ICAPS\, and MRS.
URL:https://www.ri.cmu.edu/event/how-to-coordinate-thousands-of-robots-efficiently-and-robustly/
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/09/profile-scaled.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251124T130000
DTEND;TZID=America/New_York:20251124T143000
DTSTAMP:20260922T134536
CREATED:20251114T144910Z
LAST-MODIFIED:20251114T145326Z
UID:149447-1763989200-1763994600@www.ri.cmu.edu
SUMMARY:Sensorimotor-Aligned Design for Pareto-Efficient Haptic Immersion in Extended Reality
DESCRIPTION:Abstract: \nA new category of computing devices has emerged: augmented and virtual reality headsets\, collectively referred to as extended reality (XR). These devices can alter\, augment\, or even replace our reality. While these headsets have made impressive strides in audio-visual immersion over the past half-century\, XR interactions remain almost completely absent of appropriately expressive tactile sensations. At present\, even the most advanced mainstream consumer systems rely on vibrotactile haptic actuators in the controllers\, inherently limited to clicks and buzzes — an exceedingly limited range of expressivity with which to represent the rich tactile world. \nRealizing the holistic promise of XR requires full-body haptic immersion\, just as much as it requires full audio-visual immersion. To frame critical design considerations in haptics research\, I propose an immersion-practicality tradeoff model. These competing objectives underscore the inherent tension between providing rich sensory feedback (often e.g.. costly\, bulky)\, while maintaining consumer feasibility and usability (e.g.\, low cost\, easy to use). Under this framework\, I sought to identify and build Pareto-efficient haptic systems where the haptic approach is aligned with humans’ sensorimotor system. I present five finished projects that embody my design approach through tactile haptics to different regions of the body. This framework has also been extended to other major dimensions of haptics — kinesthetics (force feedback) and proprioception — to explore how these principles can generalize. These systems\, when combined together\, advance the vision of practical\, immersive\, full-body haptics. \nThesis Committee: \nChris Harrison\, Chair\, HCII CMU \nZeynep Temel\, RI CMU \nJim McCann\, RI CMU \nMar Gonzalez-Franco\, Google \nLink to the Thesis folder
URL:https://www.ri.cmu.edu/event/sensorimotor-aligned-design-for-pareto-efficient-haptic-immersion-in-extended-reality-2/
LOCATION:Newell-Simon Hall 3305
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251125T110000
DTEND;TZID=America/New_York:20251125T120000
DTSTAMP:20260922T134536
CREATED:20251118T142159Z
LAST-MODIFIED:20251118T142159Z
UID:149484-1764068400-1764072000@www.ri.cmu.edu
SUMMARY:Multi-View 4D Human Reconstruction under Interaction Scenarios
DESCRIPTION:Abstract:\nBuilding large-scale human datasets from multi-view videos is essential for advancing research in human behavior understanding\, virtual reality\, animation\, and robotics. Compared to traditional motion capture systems that rely on physical markers to track motion\, vision-based reconstruction not only enables the capture of human motion in unconstrained environments but also avoids altering human appearance with markers. However\, existing multi-view reconstruction methods often fail when humans interact closely with others or with objects due to severe occlusions and truncations introduced by complex activities. In this thesis\, we develop a markerless capture system capable of handling close human interactions and dexterous hand-object manipulations. Using this system\, we construct two large-scale human datasets\, Harmony4D and Contact4D\, which serve as landmarks for advancing fundamental research in human-centric AI\, such as human pose estimation\, contact estimation\, and motion generation. \nCommittee:\nKris Kitani\, Chair\nFernando De La Torre\nShubham Tulsiani\nErica Weng
URL:https://www.ri.cmu.edu/event/multi-view-4d-human-reconstruction-under-interaction-scenarios/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251125T140000
DTEND;TZID=America/New_York:20251125T150000
DTSTAMP:20260922T134536
CREATED:20251119T143738Z
LAST-MODIFIED:20251119T143738Z
UID:149496-1764079200-1764082800@www.ri.cmu.edu
SUMMARY:Vision-Based Multi-Wire Detection and Tracking for UAV Wire Approach
DESCRIPTION:Abstract: \nReliable detection and tracking of power lines is critical for enabling under-wire UAV approach and inductive power-line charging to extend UAV range. However\, wires are thin\, featureless\, and visually ambiguous structures that challenge traditional computer vision methods and degrade depth estimation accuracy. To address these challenges\, this thesis presents a fully passive\, camera-only multi-wire detection and tracking algorithm that operates in real time on lightweight onboard compute using only stereo RGB imagery. \nThe proposed framework integrates a lightweight classical vision pipeline with three-dimensional geometric reasoning and a multi-stage filtering process to produce robust wire instance detections. A complementary oriented object detection model is trained using labels generated by the classical pipeline\, leveraging its fine-tuned geometric outputs to improve resilience in challenging visual conditions. To track individual wire instances across frames\, we introduce a Kalman-filter-based tracking architecture that estimates both wire orientation and per-wire positional state while remaining robust to wire detection outliers and vehicle pose drift. The system is further expanded by testing a wire positional servoing approach using the tracked wire instances in simulation and is validated across a diverse range of data sources\, including simulation\, indoor testing\, handheld experiments\, and outdoor flight evaluations. \n\nCommittee:\nSebastian Scherer (advisor)\nWennie Tabib\nMohammad Mousaei
URL:https://www.ri.cmu.edu/event/vision-based-multi-wire-detection-and-tracking-for-uav-wire-approach/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251202T110000
DTEND;TZID=America/New_York:20251202T120000
DTSTAMP:20260922T134536
CREATED:20251125T145824Z
LAST-MODIFIED:20251125T145824Z
UID:149623-1764673200-1764676800@www.ri.cmu.edu
SUMMARY:Towards Scaling Embodied Data for Robot Learning
DESCRIPTION:Abstract:\nAs artificial intelligence advances quickly in the digital domain\, the next\nfrontier lies in physical intelligence: systems that learn through acting and\nsensing in the real world. In this thesis\, we explore practical ways of scaling\nsuch embodied data across three directions. AnyCar scales synthetic data\nthrough large-scale simulation\, training a universal dynamics transformer\nthat generalizes across vehicles and environments. FACTR improves\nthe efficiency of real robot data with a low-cost bilateral teleoperation\nsystem and a curriculum that teaches policies to integrate force and\nvision. DexWild scales human data through in-the-wild data collection\nand co-training with robot demonstrations\, enabling generalization to\nunseen objects and environments. Together\, these projects explore how a\ndata-centric approach can enable more adaptive and capable robots. \nCommittee:\nDeepak Pathak (chair)\nGuanya Shi\nKenneth Shaw
URL:https://www.ri.cmu.edu/event/towards-scaling-embodied-data-for-robot-learning/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251203T093000
DTEND;TZID=America/New_York:20251203T103000
DTSTAMP:20260922T134536
CREATED:20251201T142655Z
LAST-MODIFIED:20251201T142655Z
UID:149640-1764754200-1764757800@www.ri.cmu.edu
SUMMARY:Attractors and Their Applications in Heuristic Search
DESCRIPTION:Abstract:\nHeuristic search provides a principled way to guide exploration in large state spaces\, enabling efficient solution finding. As a result\, it is widely used across domains such as robotics\, games\, and planning. However\, its performance is often limited by memory consumption and computational overhead\, which have motivated extensive research on improving both. This thesis introduces a sparse representation called attractors and explores two of its applications in heuristic search. First\, we present Attractor-based Closed List Search (ACLS)\, a framework that uses attractors to sparsely represent the Closed list. ACLS intelligently identifies attractor states in a way that enables efficient solution reconstruction while preserving theoretical guarantees on the quality of the solution. We demonstrate that ACLS significantly reduces memory usage\, while achieving comparable planning times and outperforming state-of-the-art approaches. Second\, we introduce front-to-attractors (F2A) heuristics\, a family of heuristics that leverage attractors in bidirectional heuristic search (Bi-HS). We demonstrate that F2A heuristics substantially reduce the number of heuristic evaluations compared to front-to-front (F2F) heuristics\, while maintaining strong informativeness and reducing expansions relative to front-to-end (F2E) heuristics\, resulting in improved runtime performance. Together\, these projects demonstrate the broad potential of attractors in heuristic search.\n\nCommittee:\nMaxim Likhachev (chair)\nJiaoyang Li\nYorai Shaoul
URL:https://www.ri.cmu.edu/event/attractors-and-their-applications-in-heuristic-search/
LOCATION:GHC 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251203T150000
DTEND;TZID=America/New_York:20251203T163000
DTSTAMP:20260922T134536
CREATED:20251125T221309Z
LAST-MODIFIED:20251125T221309Z
UID:149634-1764774000-1764779400@www.ri.cmu.edu
SUMMARY:Robotic System Design Principles for Human-Human Collaboration
DESCRIPTION:Abstract: \nRobots possess unique affordances granted by combining software and hardware. Most existing research focuses on the impact of these affordances on human-robot collaboration\, but the theory of how robots can facilitate human-human collaboration is underdeveloped. Such a theory would be beneficial in education. An educational device can afford collaboration in both assembly and use. This thesis will enumerate and validate the design principles of educational devices that facilitate collaborative assembly and collaborative learning.\nThis research draws upon cognitive theories used in the disciplines of Computer-Supported Collaborative Work (CSCW)\, Computer-Supported Collaborative Learning (CSCL)\, Educational Robotics\, and Human-Robot Interaction (HRI). Each discipline uses theories that align with its respective goals to model different pieces of cognition. However\, they do not consider other factors outside their respective goals. Diverse analytical lenses are needed to understand the multiple dimensions of influence an educational device can have on human-human interaction to support collaborative assembly and collaborative learning.\nWe explore these dimensions first through the development and assessment of\nRoboLoom\, a robotic Jacquard loom kit designed for interdisciplinary\, collaborative education. Through the study of RoboLoom’s use and assembly in an undergraduate course\, we extract design features that facilitate student-student collaboration during classroom activities. These features encompass task complexity\, task parallelization\, physicality\, repetition of tasks\, specificity of hardware\, and familiarity with hardware.\nWe then explore these design principles through three studies: a comparison\nbetween two different looms\, a study of devices designed for and against the principles\, and a comparison of two versions of RoboLoom. We find five design principles that influence collaborative behavior: repetitiveness\, specificity\, parallelizability\, physicality\, and difficulty. These design principles were shown to causally change collaborative behaviors in controlled lab settings and in situ engineering education tasks. By evaluating these systems through multiple cognitive lenses\, we determine that these design principles are effective in facilitating collaborative assembly and promising for collaborative learning.\n\n\n\nCommittee Members: \n    Illah Nourbakhsh\, Co-Chair\n    Melisa Orta Martinez\, Co-Chair\nJames McCann\nKylie Peppler\, University of California\, Irvine \n\nLink to Thesis
URL:https://www.ri.cmu.edu/event/robotic-system-design-principles-for-human-human-collaboration/
LOCATION:GHC 8102
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251203T160000
DTEND;TZID=America/New_York:20251203T173000
DTSTAMP:20260922T134536
CREATED:20251202T153007Z
LAST-MODIFIED:20251202T153007Z
UID:149649-1764777600-1764783000@www.ri.cmu.edu
SUMMARY:F25 MRSD Poster Session
DESCRIPTION:Student teams from the Robotic Systems Development (MRSD) program will present posters\, videos\, and hardware related to their projects. Please come and see their efforts!\nThis year’s MRSD projects are robots for an autonomous wheelchair\, drone-based battlefield triage\, sandwich making\, knee surgery\, pepper picking\, manufacturing tote transporting\, fire search-and-rescue\, dual-arm QC inspection for manufacturing\, and sun-synchronous lunar circumnavigation. \nhttps://mrsd.ri.cmu.edu/project-examples/student-project-websites/spring-2025-fall-2025/
URL:https://www.ri.cmu.edu/event/f25-mrsd-poster-session/
LOCATION:NSH Atrium
CATEGORIES:Special Events
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251208T153000
DTEND;TZID=America/New_York:20251208T163000
DTSTAMP:20260922T134536
CREATED:20251202T190520Z
LAST-MODIFIED:20251206T151245Z
UID:149651-1765207800-1765211400@www.ri.cmu.edu
SUMMARY:What Can We Learn from a Million Models?
DESCRIPTION:Abstract: Machine learning has transformed many fields by learning from large collections of data. Yet\, it is rarely applied to its own outputs: the models themselves. Today\, with millions of publicly available models\, a natural question arises: what can we do with so many models? In this talk\, I will motivate two core applications that leverage this untapped potential\, demonstrating their utility in the context of computer vision: (i) identifying emerging trends in model design\, and (ii) reducing the need to train models from scratch through model recycling. To support these goals\, I introduce the Model Atlas: a structured graph that represents models\, their attributes\, and the weight-space transformations that interconnect them. My research into weight-space learning enables the construction of this atlas by treating models themselves as data and inferring properties such as functionality\, performance\, and lineage directly from their weights. I will present key observations and methodologies that make weight-space learning possible at scale. As a visual prelude\, you can explore the repository under study at: https://horwitz.ai/model-atlas . \nBio: Eliahu Horwitz is a Google PhD Fellow in Machine Learning and ML Foundations and a final-year PhD candidate in Computer Science at The Hebrew University of Jerusalem\, advised by Prof. Yedid Hoshen. His research centers on learning representations of neural network weights and understanding model populations directly in weight space. He is particularly interested in how weight-space learning can enable new downstream capabilities\, such as model forensics\, model discovery\, and interpretability\, and in how treating models as data points can advance broader areas of machine learning. Eliahu is also a recipient of the Israeli Council for Higher Education Scholarship and has previously interned at Google Research. \nHomepage:  https://horwitz.ai \nSponsor \nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one AI-powered video generation platform that leverages advances in computer vision\, generative modeling\, and multimodal learning to make high-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/what-can-we-learn-from-a-million-models/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/12/12-8-25.jpeg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251208T170000
DTEND;TZID=America/New_York:20251208T183000
DTSTAMP:20260922T134536
CREATED:20250929T144749Z
LAST-MODIFIED:20250929T144749Z
UID:148939-1765213200-1765218600@www.ri.cmu.edu
SUMMARY:Erica Weng - PhD Defense Info TBA
DESCRIPTION:More info coming soon
URL:https://www.ri.cmu.edu/event/erica-weng-phd-defense-info-tba/
LOCATION:NSH 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251209T120000
DTEND;TZID=America/New_York:20251209T133000
DTSTAMP:20260922T134536
CREATED:20251125T211907Z
LAST-MODIFIED:20251125T211907Z
UID:149629-1765281600-1765287000@www.ri.cmu.edu
SUMMARY:Design Optimization of Modular Manipulators for Manipulation in Cluttered Agricultural Environments
DESCRIPTION:Abstract:\nAlthough agriculture is a highly mechanized industry\, essential and high-value subsectors such as horticulture and floriculture remain heavily reliant on manual labor because they require complex\, contact-rich\, and highly selective handling of both plants and produce. The variability and density of tree-canopy clutter further complicate the automation process\, making robot performance difficult to quantify consistently and preventing the development of a single\, universally effective automation solution. Modular and reconfigurable robots (MRRs) can help address this challenge by reducing the cost of creating custom robots tailored to specific task requirements. However\, determining the optimal robot design configuration for an MRR system remains a complex and unintuitive process\, even for experts. This thesis addresses the problem of automating the robot design process by introducing a systematic design framework that unifies deterministic and consistent task-performance metrics with global optimization methods primarily targeting agricultural manipulation tasks. \nThe first contribution targets the challenge of computing self-motion manifolds (SMMs)\, which are global inverse-kinematics solutions for redundant manipulators. We solve this problem using Runge-Kutta solvers after posing the underlying ordinary differential equation problem in a form we call the SMM Initial Value Problem (SMM-IVP). The SMM-IVP is able to trace the manipulator’s self-motion configuration space reliably. Compared to existing predictor-corrector and linear step-corrector approaches\, the SMM-IVP exhibits improved convergence behavior and numerical stability. For design applications\, the SMM-IVP acts as a global inverse-kinematics procedure that provides consistent and initialization-independent performance characterization\, which is essential for design optimization in the cluttered conditions typical of agricultural manipulation. \nBuilding on the first contribution\, the second contribution develops a general framework for formulating and solving robot-design optimization problems. In parallel\, we introduce new SMM-based performance metrics that more accurately characterize dexterity metrics for redundant manipulators. We apply this framework and the new metrics to a manipulator placement optimization problem for a dual-arm pepper-harvesting system\, and we show that it produces highly performant\, non-intuitive configurations that outperform both human-expert designs and conventional dexterity-based baselines. \nThe third contribution grounds our design methods in real-world tree geometry data and directly addresses the inherent heterogeneity of robot performance. To this end\, we introduce a lexicographic design optimization framework for tuple-valued task metrics\, allowing robot performance to be represented as a set of multiple criteria ordered according to designer-specified priorities. This representation preserves the semantic meaning of each criterion\, enables explicit hierarchical prioritization\, and provides a principled alternative to ad-hoc scalarization methods. \nTogether\, these advances establish a reproducible foundation for task-driven robot design optimization. The methods integrate kinematic modeling\, performance evaluation\, and global optimization into a single\, coherent pipeline that extends beyond agricultural manipulation. More broadly\, this work supports the practical deployment of modular\, reconfigurable manipulators by lowering the barriers to designing task-specific robot designs for the highly cluttered conditions like those found in agricultural robot manipulation. \nThesis Committee Members:\nGeorge Kantor (Chair)\, CMU\nOliver Kroemer\, CMU\nZeynep Temel\, CMU\nChangying (Charlie) Li\, University of Florida \nThesis Draft
URL:https://www.ri.cmu.edu/event/design-optimization-of-modular-manipulators-for-manipulation-in-cluttered-agricultural-environments/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251212T110000
DTEND;TZID=America/New_York:20251212T120000
DTSTAMP:20260922T134536
CREATED:20251208T224116Z
LAST-MODIFIED:20251208T224116Z
UID:149685-1765537200-1765540800@www.ri.cmu.edu
SUMMARY:Examining Engagement and Motivation in a Conversational Robotic Exercise Coach for Older Adults
DESCRIPTION:Abstract: Exercise is essential for healthy aging\, but motivation and adherence to exercise often decline with age\, leading to a more sedentary lifestyle. At the same time\, the growing aging population continues to strain the availability of physical therapists and exercise coaches. In this thesis\, we introduce a conversational robotic exercise coach system designed to support older adults during exercise. To evaluate this system\, we conducted a user study with 10 participants aged 59 and above. We analyzed both survey responses and verbal interactions to understand how participants engaged with the robot and how motivation was expressed during exercise. Based on these findings\, this work presents design recommendations for future autonomous conversational exercise robots for older adults. \nCommittee: \nProf. Aaron Steinfeld (advisor) \nProf. Reid Simmons \nProf. Jean Oh \nMichelle Zhao
URL:https://www.ri.cmu.edu/event/examining-engagement-and-motivation-in-a-conversational-robotic-exercise-coach-for-older-adults/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251212T110000
DTEND;TZID=America/New_York:20251212T120000
DTSTAMP:20260922T134536
CREATED:20251209T143621Z
LAST-MODIFIED:20251209T143621Z
UID:149688-1765537200-1765540800@www.ri.cmu.edu
SUMMARY:Examining Engagement and Motivation in a Conversational Robotic Exercise Coach for Older Adults
DESCRIPTION:Abstract: Exercise is essential for healthy aging\, but motivation and adherence to exercise often decline with age\, leading to a more sedentary lifestyle. At the same time\, the growing aging population continues to strain the availability of physical therapists and exercise coaches. In this thesis\, we introduce a conversational robotic exercise coach system designed to support older adults during exercise. To evaluate this system\, we conducted a user study with 10 participants aged 59 and above. We analyzed both survey responses and verbal interactions to understand how participants engaged with the robot and how motivation was expressed during exercise. Based on these findings\, this work presents design recommendations for future autonomous conversational exercise robots for older adults. \nCommittee: \nProf. Aaron Steinfeld (advisor) \nProf. Reid Simmons \nProf. Jean Oh \nMichelle Zhao
URL:https://www.ri.cmu.edu/event/examining-engagement-and-motivation-in-a-conversational-robotic-exercise-coach-for-older-adults-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251212T150000
DTEND;TZID=America/New_York:20251212T163000
DTSTAMP:20260922T134536
CREATED:20251202T210118Z
LAST-MODIFIED:20251202T231025Z
UID:149656-1765551600-1765557000@www.ri.cmu.edu
SUMMARY:Modeling what Matters: Emergent Abstraction In Reinforcement Learning
DESCRIPTION:Abstract: Real-world decision-making is rife with partial observability\, long horizons\, and complex multi-agent interactions. This thesis argues that abstraction—forming simplified representations of the task that retain relevant information—offers a unifying principle for tackling these challenges across model-free and model-based reinforcement learning (RL). We develop methods in which abstractions are not hand-designed but emerge from learning objectives\, yielding representations that improve an agent’s ability to cope with high-dimensional observations\, extended temporal dependencies\, and inter-agent coupling.\n\nOn the model-free\, multi-agent side\, we introduce Partial Reward Decoupling (PRD)\, a game-abstraction mechanism that dynamically decomposes teams into subgroups\, simplifying cross-agent credit assignment and accelerating cooperative learning. We also study discrete communication learning under bandwidth constraints\, where agents learn what information to transmit\, to whom\, and how to encode it—linking communication learning to representation learning and generative modeling. \nWe also show how abstraction mitigates the misalignment between model-learning and task objectives typically found in model-based RL methods. By focusing limited model capacity on task-relevant factors and operating at an appropriate temporal scale\, abstraction improves the utility of world models for decision-making. Toward this end\, we explore the use of variational inference (VI) to learn both state and temporal abstractions. We demonstrate a state-abstraction method that ignores distracting details while retaining task-relevant features\, attaining strong results on distraction-rich control benchmarks without relying on data-augmentation heuristics. We also propose a latent-variable approach to temporal abstraction that extracts skills and learns a temporally abstract dynamics model from offline data\, enabling effective long-horizon prediction and planning for downstream tasks. \nFinally\, we present Unified RL\, which blends model-based and model-free updates by detecting when a learned model ceases to be useful for policy improvement and falling back to model-free learning updates. Empirically\, Unified RL retains the data efficiency of model-based methods while achieving asymptotic performance comparable to model-free RL. \nCommittee Members:  \nHowie Choset\, chair\nJeff Schneider\, co-chair\nRuslan Salakhutdinov\nRoberto Calandra\, TU Dresden\n \nLink to thesis
URL:https://www.ri.cmu.edu/event/modeling-what-matters-emergent-abstraction-in-reinforcement-learning/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251215T153000
DTEND;TZID=America/New_York:20251215T163000
DTSTAMP:20260922T134536
CREATED:20251203T162327Z
LAST-MODIFIED:20251210T163815Z
UID:149661-1765812600-1765816200@www.ri.cmu.edu
SUMMARY:Should we skip attention?
DESCRIPTION:Abstract: Transformers are ubiquitous. They influence nearly every aspect of modern AI. However\, the mechanics of their training remain poorly understood. This poses a problem for the field due to the immense amounts of data\, computational power\, and energy being invested in the training of these networks. I highlight a recent intriguing empirical result from our group. Specifically\, selfattention catastrophically fails to train unless it is paired with a skip connection. This contrasts with other components of a transformer that continue to demonstrate good performance (albeit suboptimal) when skip connections are removed. In this talk\, I explore why this is the case and what could be done to enhance the fundamental training efficiency of modern transformers. We even showcase some practical cases in which removing self-attention completely can lead to significantly improved performance. \nBio: Simon Lucey Ph.D. is the Director of the Australian Institute for Machine Learning (AIML) and a professor in the School of Computer and Mathematical Sciences\, at the University of Adelaide. He is also Director of the CommBank Foundational AI Research Centre. Prior to this he was an associate research professor at Carnegie Mellon University’s Robotics Institute (RI) in Pittsburgh USA; where he spent over 10 years as an academic. He was also Principal Research Scientist at the autonomous vehicle company Argo AI from 2017-2022. He has received various career awards\, notably the AmCham AI Scientist of the year in 2024. He is also currently a member of the Australian Government’s AI Expert Group\, and their National Robotics Strategy committee. Simon’s research interests span AI\, machine learning\, computer vision and robotics. \n  \nSponsor \nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one AI-powered video generation platform that leverages advances in computer vision\, generative modeling\, and multimodal learning to make high-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/should-we-skip-attention/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2025/12/12-12-25-Lucy.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251219T120000
DTEND;TZID=America/New_York:20251219T130000
DTSTAMP:20260922T134536
CREATED:20251211T153217Z
LAST-MODIFIED:20251211T153217Z
UID:149714-1766145600-1766149200@www.ri.cmu.edu
SUMMARY:Resilient Aerial Autonomy for Science\, Search\, and Survey
DESCRIPTION:
URL:https://www.ri.cmu.edu/event/resilient-aerial-autonomy-for-science-search-and-survey/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:Faculty Events
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260113T120000
DTEND;TZID=America/New_York:20260113T133000
DTSTAMP:20260922T134536
CREATED:20260106T153223Z
LAST-MODIFIED:20260106T153223Z
UID:149856-1768305600-1768311000@www.ri.cmu.edu
SUMMARY:Self-supervised tactile perception for robot dexterity
DESCRIPTION:Abstract: \nHumans are incredibly dexterous. We interact with and manipulate tools effortlessly\, leveraging touch without a second thought. Yet\, replicating this level of dexterity in robots is a major challenge. While the robotics community\, recognizing the importance of touch in fine manipulation\, has developed a wide variety of tactile sensors\, how best to leverage these sensors for both perception and manipulation is unclear. In this thesis\, we address how to efficiently integrate tactile sensing for robot perception and dexterous manipulation. \nSpecifically\, we turn to self-supervised learning (SSL) to train tactile representations that can generalize across sensors\, standardize usage across downstream tactile tasks\, and further alleviate the need to collect labeled task data which is often impractical to collect for tasks such as uncalibrated force field estimation. To this end\, we discuss Sparsh and Sparsh-skin\, a family of SSL models for vision and magnetic-skin based tactile sensors respectively. Sparsh and Sparsh-skin are trained via self-distillation for full-hand tactile sensors in downstream tasks. We find that both Sparsh and Sparsh-skin not only outperform task and sensor-specific end-to-end models by a large margin\, but also that they are data efficient for downstream task training. \nSecond\, we note that existing work often overlooks the multimodal aspects of human touch\, such as vibration and heat sensing. We discuss Sparsh-X\, a compact tactile representation fusing image\, pressure\, audio and inertial measurements from the DIGIT360 sensor. With Sparsh-X we demonstrate that multimodal sensing improves both passive perception tasks as well as dexterous manipulation tasks such as in-hand rotation. \nFinally\, we present privileged tactile latent distillation (PTLD)\, a novel method to imbue tactile sensing in dexterous manipulation policies trained via reinforcement learning. PTLD avoids simulating tactile sensors and uses privileged sensors to bridge the sim-to-real gap. With PTLD\, we first show that one can improve existing RL trained policies such as in-hand rotation and then that it can enable learning more challenging tasks such as in-hand reorientation. \nJointly these contributions provide a path to leverage tactile sensing in both imitation and reinforcement learning based robot manipulation. \nThesis Committee Members: \nMichael Kaess\, chair\nShubham Tulsiani\nGuanya Shi\nMustafa Mukadam\, Amazon Robotics\nJitendra Malik\, UC Berkeley & Amazon FAR \nThesis Draft
URL:https://www.ri.cmu.edu/event/self-supervised-tactile-perception-for-robot-dexterity/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260113T143000
DTEND;TZID=America/New_York:20260113T160000
DTSTAMP:20260922T134536
CREATED:20260106T180847Z
LAST-MODIFIED:20260106T180847Z
UID:149866-1768314600-1768320000@www.ri.cmu.edu
SUMMARY:Efficient Visual Modeling with Adaptive Representations
DESCRIPTION:Abstract: \n\nWhile image understanding\, generation\, and manipulation have matured rapidly in recent years\, video remains challenging due to the significantly larger input size. As a result\, tasks such as generating long videos or understanding extended video sequences remain out of reach for current models due to their computational cost. This talk presents a series of works that address this issue by adapting ideas from video compression to accelerate visual model training and inference. I will first introduce Run-Length Tokenization (RLT)\, which modifies the vision transformer architecture to exploit temporal redundancy\, enabling substantial speedups without compromising accuracy. Next\, I will present FlowTok\, which incorporates motion vectors to extend RLT to dynamic scenes\, maintaining efficiency even under camera and object motion. I will then discuss Adaptive Patch Transformers (APT)\, which apply these principles to images by dynamically assigning larger patch sizes in low-complexity regions to reduce computation while preserving performance. We next apply these principles to video generation\, and propose SkipSR\, a cascaded generation framework that combines fast video super-resolution with cascaded diffusion models. Finally\, we introduce FPS-Bench\, a benchmark to systematically evaluate the impact of frame rate and resolution on downstream video understanding tasks\, offering insights into which aspects of fidelity truly matter for model performance. By unifying efficient video tokenization with scalable video synthesis and principled evaluation\, this thesis enables significantly faster visual models in both understanding and generation tasks\, unlocking further scaling. \n\n\n\nThesis Committee Members:\n \n\n\nLászló A. Jeni(co-chair)\nKris M. Kitani (co-chair)\nJun-Yan Zhu\n\nRohit Girdhar (Meta GenAI)\n\nLu Jiang (ByteDance)\n\n\n\n\n\n\n\n\n\n\n\n\n\nA draft of the thesis is available here: Thesis Draft
URL:https://www.ri.cmu.edu/event/efficient-visual-modeling-with-adaptive-representations/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260115T083000
DTEND;TZID=America/New_York:20260115T100000
DTSTAMP:20260922T134536
CREATED:20260108T145112Z
LAST-MODIFIED:20260108T145112Z
UID:149897-1768465800-1768471200@www.ri.cmu.edu
SUMMARY:Toward Scalable Architectures for Multimodal LLM-based Cooperative Autonomous Driving
DESCRIPTION:Abstract: Despite the tremendous progress made in autonomous driving over the years\, the safety of autonomous vehicles still requires further improvement before they can operate worldwide with full human trust. One principal safety concern is that each individual vehicle may have a limited field of view due to finite detection ranges\, potential sensor failures\, or occlusions caused by nearby large objects such as buses or trucks. This limitation in perception introduces additional challenges for the downstream planning and control modules\, making it more difficult for autonomous vehicles to generate safe driving decisions and actions.\nTo address this issue\, recent research has proposed vehicle-to-vehicle (V2V) and vehicle-to-everything (V2X) cooperative perception for autonomous driving. In such systems\, connected autonomous vehicles (CAVs) share their individual perception features with one another to improve overall cooperative detection accuracy. However\, most existing work focuses solely on the cooperative detection task\, without leveraging temporal information about the dynamic environment or considering other critical components of autonomous driving\, such as prediction and planning. \nTo broaden the scope of cooperative driving research\, my proposed doctoral research aims to explore multimodal large language model (LLM)–based cooperative autonomous driving\, motivated by several potential advantages of LLMs. First\, a single LLM-based model offers the flexibility to perform multiple tasks\, including perception\, prediction\, and planning\, within a unified framework. Second\, LLMs exhibit strong generalizability due to large-scale pretraining on diverse data. Third\, LLM-based driving models possess reasoning capabilities that enable them to handle long-tail driving scenarios that may not appear in the training data. Fourth\, natural language can serve as an effective and efficient communication interface for V2V\, V2X\, and human–vehicle interactions. \nWe have developed multimodal LLM-based cooperative autonomous driving architectures that enable end-to-end cooperative driving and generate suggested future trajectories for all CAVs through V2V communication. In addition\, we have designed a graph-of-thoughts reasoning framework to further enhance the reliability and interpretability of our multimodal LLM-based architecture. Finally\, we propose to develop a decentralized V2V framework using multimodal LLMs to improve the scalability and feasibility of future large-scale deployment. \n\n \n \nThesis Committee:\n\nStephen F. Smith (Chair)\nJohn Dolan\nDeva Ramanan\nMin-Hung Chen (NVIDIA)\n\n\n\n\n\nThesis proposal link
URL:https://www.ri.cmu.edu/event/toward-scalable-architectures-for-multimodal-llm-based-cooperative-autonomous-driving/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260116T120000
DTEND;TZID=America/New_York:20260116T130000
DTSTAMP:20260922T134536
CREATED:20260112T160839Z
LAST-MODIFIED:20260112T160839Z
UID:149992-1768564800-1768568400@www.ri.cmu.edu
SUMMARY:Creative Physical AI
DESCRIPTION:Abstract:\nDo robots need creativity? I will share my stance that they do need creativity to solve general problems and support human values. Physical AI is a type of AI that enables robots to perceive and interact with a physical world. Trendy approaches in physical AI such as Vision-Language-Action (VLA) models directly map the observations to actions where robots make decisions dominantly based on sensed information. While sensing is crucial for understanding the current physical environments\, this paradigm of physical AI is fundamentally limited to support general tasks where humans see around corners and solve problems creatively based on not only what they can observe now but also various predictions of the latent spatiotemporal and social contexts. I will illustrate the examples where robots without creativity can fail to fulfill even simple goals and how we can develop physical AI for creative problem solving.\nIf equipped with creative physical AI\, can such robots promote human creativity as in creating arts? Generative AI has brought us numerous types of convenience in the digital art world. To create artifacts in the real world\, creative physical AI is needed\, for instance\, to preserve traditional craftsmanship such as wood carving or claymation\, which faces declining participation due to its labor-intensive nature. More broadly\, our innovations in creative physical AI aim to encourage people to participate in more creative activities such as educational and therapeutic art sessions. I would like to invite the audience to think about how we can use technologies to promote human creativity for the next generation.
URL:https://www.ri.cmu.edu/event/creative-physical-ai/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:Faculty Events
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260120T120000
DTEND;TZID=America/New_York:20260120T133000
DTSTAMP:20260922T134536
CREATED:20260113T200035Z
LAST-MODIFIED:20260113T200035Z
UID:150081-1768910400-1768915800@www.ri.cmu.edu
SUMMARY:Towards Manipulation in the Blind: Motion Planning for Manipulation under Uncertainty using Contacts
DESCRIPTION:Abstract:\n \nHumans routinely rely on the sense of touch to better perceive the world. In environments characterized by poor lighting\, occlusions\, limited fields of view\, or sparse visual features\, contact feedback often becomes a primary source of information for perceiving the environment and successfully completing manipulation tasks. Everyday examples include locating a light switch in the dark\, retrieving an item from a high shelf\, or reaching for a valve at the back of a kitchen sink cabinet where visual access is severely restricted. In such settings\, humans actively reason about contacts to infer both the poses of objects of interest and the environmental obstacles in the workspace. Enabling robots to exhibit similar capabilities remains a fundamental challenge in autonomous manipulation. This thesis investigates search-based planning techniques that allow robots to effectively leverage contact feedback as a sensing modality\, enabling robust manipulation under uncertainty. \nThis work focuses on two broad classes of manipulation problems in which contact plays a critical role. The first class concerns object pose uncertainty\, where precise estimation of target object pose is required to complete high-precision manipulation tasks such as charger plug insertion\, pipe assembly\, or other tight-tolerance manipulation problems. In these settings\, even small pose errors on the order of a few millimeters can lead to failure. This problem class studies how robots can actively use contacts during execution to reduce object pose uncertainty to a level sufficient for successful task completion. The second class of problems addresses manipulation under environmental uncertainty\, where the locations and geometries of environmental obstacles are unknown or only partially observable. For example\, when reaching into a cluttered kitchen sink cabinet with unknown obstacles and pipes\, a robot must detect contacts\, infer obstacle locations\, and adapt its motion accordingly in order to safely reach the target (valve). Together\, these two problem classes capture a wide range of real-world scenarios in which contact-driven reasoning is essential. \nPlanning under object pose uncertainty naturally falls within the framework of Partially Observable Markov Decision Processes (POMDPs)\, which are computationally expensive to solve\, particularly in continuous and high-dimensional robotic domains. This thesis presents three complementary frameworks to address this challenge. The first is an experience-based preprocessing approach designed for semi-structured environments that require strong online performance. This framework leverages solutions to previously solved\, similar POMDPs to accelerate future planning queries while maintaining theoretical guarantees on solution quality. An offline database of policies is constructed and queried at execution time based on the current problem instance\, enabling fast online decision-making. The second framework targets less structured domains where preprocessing is impractical. It introduces an online closed-loop planning and execution approach that employs a hierarchical representation of uncertainty. By adaptively representing and reasoning about uncertainty\, this method significantly reduces planning time\, making online planning and execution feasible in more complex settings. The third contribution addresses a key computational bottleneck in these settings\, namely the high cost of belief space transition computations. To mitigate this issue\, the thesis proposes lazy heuristic search algorithms for POMDPs that defer expensive belief updates until they are necessary\, using approximate Q-value estimators to guide search. These lazy solvers substantially reduce planning time while preserving solution quality. \nFor the problem class of manipulation under environmental uncertainty\, this thesis develops an iterative planning and execution framework that tightly couples contact sensing\, environment prediction\, and motion planning. The system employs a torque-based contact detection and localization module capable of detecting contacts occurring anywhere along the robot manipulator. The history of detected contacts is used to construct a partial occupancy map of the workspace\, which is then extrapolated using learned occupancy estimators. A motion planning module reasons over this estimated occupancy representation to compute actions that are likely to safely and efficiently move the robot toward the goal. The framework is evaluated in simulation and on a real UR10e manipulator across two challenging domestic tasks: manipulating a valve under a kitchen sink surrounded by pipes and retrieving a target object from a cluttered shelf. \nOverall\, this thesis takes a step toward manipulation in the blind\, where robots explicitly leverage contact interactions as a primary sensing modality to plan and execute manipulation tasks under uncertainty. Rather than treating contact as a failure mode to be avoided\, we treat it as an informative observation that can be actively exploited to reduce uncertainty and guide motion. \nThesis Committee: \nProf. Maxim Likhachev\, Chair \nProf. Jeffrey Ichnowski \nProf. Oliver Kroemer \nProf. Mehmet Dogar\, University of Leeds \nA draft of the thesis document is available at: https://drive.google.com/drive/folders/1TCSil_tJksTWC_RweWGS2FHoZBjyWM8w?usp=drive_link
URL:https://www.ri.cmu.edu/event/towards-manipulation-in-the-blind-motion-planning-for-manipulation-under-uncertainty-using-contacts/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260123T133000
DTEND;TZID=America/New_York:20260123T143000
DTSTAMP:20260922T134536
CREATED:20251208T163242Z
LAST-MODIFIED:20251208T163353Z
UID:149678-1769175000-1769178600@www.ri.cmu.edu
SUMMARY:RI Faculty Business Meeting
DESCRIPTION:Meeting for RI Faculty.
URL:https://www.ri.cmu.edu/event/ri-faculty-business-meeting-33-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:Faculty Events
ATTACH;FMTTYPE=image/png:https://www.ri.cmu.edu/app/uploads/2023/11/ri-new-mark-512-512-transparent.png
ORGANIZER;CN="RI Director's Office":MAILTO:lynnetta@cs.cmu.edu
END:VEVENT
END:VCALENDAR