BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260901T113000
DTEND;TZID=America/New_York:20260901T130000
DTSTAMP:20261003T070916
CREATED:20260820T203402Z
LAST-MODIFIED:20260820T203402Z
UID:153235-1788262200-1788267600@www.ri.cmu.edu
SUMMARY:Scalable Vision-Language Models through Unified 2D and 3D Representations
DESCRIPTION:Abstract:\nVision-language models have become remarkably capable on images and short videos\, yet they still struggle with two abilities central to embodied intelligence: spatial understanding and long-range temporal reasoning. A major reason is representational: today’s models process videos as long sequences of 2D patches\, so computation grows with observation length even when the underlying scene changes little. This thesis argues that organizing perception around persistent 3D structure rather than individual frames allows model complexity to scale with scene content rather than observation length\, leading to efficient inference on long videos while providing a stronger foundation for spatial reasoning.\n\nWe introduce a unified representation based on 3D feature clouds that handles both 2D and 3D inputs within a single architecture. Because the same model trains on abundant 2D image-text data alongside available 3D data\, it acquires better spatial understanding without sacrificing 2D performance. We demonstrate this on 3D instance segmentation (ODIN)\, extend it to broader vision-language tasks (UniVLG)\, and scale it to billion-parameter VLMs (Qwen-3D)\, showing consistent improvements in spatial understanding and inference efficiency. \nWe then move beyond static scenes to dynamic videos. We develop 3D scene representations (TrackEverything) that disentangle static and dynamic content\, deduplicate the scene across time\, and track all points in 3D throughout long videos. These representations convert videos into concise spatiotemporal structures that grow with scene complexity rather than video duration\, enabling new capabilities such as being able to track all points across all frames in long videos (1000+ frams). \nFinally\, we outline proposed work on two fronts: leveraging these dynamic 3D representations for spatial and motion reasoning in vision-language models\, and scaling dynamic 3D tracking to real-world video data using heterogeneous supervision beyond synthetic data. \nTogether\, this work makes the case for moving beyond 2D patch representations toward 3D-native models for better and more efficient video understanding. \nThesis Committee:\nKaterina Fragkiadaki (Chair)\nDeva Ramanan\nShubham Tulsiani\nLeonidas Guibas\, Stanford University\n\nThesis Proposal Link
URL:https://www.ri.cmu.edu/event/scalable-vision-language-models-through-unified-2d-and-3d-representations/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260904T133000
DTEND;TZID=America/New_York:20260904T143000
DTSTAMP:20261003T070916
CREATED:20260904T131506Z
LAST-MODIFIED:20260904T131506Z
UID:153464-1788528600-1788532200@www.ri.cmu.edu
SUMMARY:RI Faculty Business Meeting
DESCRIPTION:Meeting for RI Faculty.\nIn person location – NSH 4305.\nZoom link available via calendar invite.
URL:https://www.ri.cmu.edu/event/ri-faculty-business-meeting-33-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:Faculty Events
ATTACH;FMTTYPE=image/png:https://www.ri.cmu.edu/app/uploads/2023/11/ri-new-mark-512-512-transparent.png
ORGANIZER;CN="RI Director's Office":MAILTO:lynnetta@cs.cmu.edu
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260904T143000
DTEND;TZID=America/New_York:20260904T153000
DTSTAMP:20261003T070916
CREATED:20260828T164202Z
LAST-MODIFIED:20260910T180458Z
UID:153332-1788532200-1788535800@www.ri.cmu.edu
SUMMARY:CANCELED - Trust\, Sensing\, and Learning for Provable Multi-Robot Performance
DESCRIPTION:Seminar Canceled \nThis seminar has been canceled and may be rescheduled for a future date. Please check back for updates. ×Abstract: Multi-robot systems are physically embodied networks — they sense\, move\, and communicate through the physical world. The bar for safe decision-making rises as these systems enter safety-critical\, real-world settings where they must perform well under uncertainty. Our work shows that physicality is a resource against the two kinds of uncertainty they face: intentional (or adversarial)\, where data is manipulated by malicious agents\, and natural\, where aspects of the environment are simply unknown. Most of this talk concerns intentional uncertainty. Here\, one way to exploit physicality is by using communication as a sensor. Because the signals robots exchange are difficult to forge\, they carry evidence that can be cross-validated to yield a quantifiable likelihood that an agent’s data is trustworthy. This is the foundation of cy-trust\, in which stochastic observations of trust model an agent’s trustworthiness probabilistically from physical rather than cryptographic evidence. Each neighbor’s contribution is then weighted by its trust value. Under this framework\, we show that consensus\, distributed optimization\, and other core coordination tasks admit almost-sure convergence with bounded deviation from their nominal performance\, even when malicious agents exceed half of a node’s connectivity\, past the classical Byzantine bound. We support this finding with both theory and hardware experiments under adversarial attack. Against natural uncertainty\, we show that real-time sensing can be folded into rollout-based reinforcement learning\, where the same machinery reweights futures rather than neighbors. We apply this idea to routing a fleet of robots to stochastically appearing demand and\, with Project CETI\, to the first autonomous robotic rendezvous with sperm whales at sea. Finally\, we preview some of our future work combining trust with long-horizon sequential decision-making\, targeting planning that stays provably resilient when the data informing the plan may itself be corrupted. \nBio: Stephanie Gil is the John L. Loeb Associate Professor of Engineering and Applied Sciences at Harvard University and an Associate Faculty member of the Kempner Institute. Her research focuses on trust and coordination in multi-robot systems\, at the intersection of robotics\, communication\, and learning. Her contributions have been recognized through the DARPA Young Faculty Award (2024)\, the Office of Naval Research Young Investigator Award (2021)\, and the National Science Foundation CAREER Award (2019). She was named a 2020 Sloan Research Fellow for her work at the intersection of robotics and communication. She earned her Ph.D. at CSAIL at MIT\, specializing in multi-robot coordination and control\, and her B.S. at Cornell University.
URL:https://www.ri.cmu.edu/event/trust-sensing-and-learning-for-provable-multi-robot-performance/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2026/08/Stephanie-Gil_SQUARW.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260911T100000
DTEND;TZID=America/New_York:20260911T113000
DTSTAMP:20261003T070916
CREATED:20260908T140601Z
LAST-MODIFIED:20260908T140601Z
UID:153484-1789120800-1789126200@www.ri.cmu.edu
SUMMARY:From Simulation to the Real World: A Multi-Level System for Object Navigation with Vision-Language Models
DESCRIPTION:Abstract: \nObject navigation (ObjectNav) asks a robot to find an instance of a target object category in an unknown environment\, which demands perception\, spatial reasoning\, and long-horizon decision making at once. Today’s autonomous robots excel at mapping and moving through space yet lack high-level semantic intelligence\, while vision-language models (VLMs) offer rich commonsense reasoning but limited 3D spatial grounding and long-term spatial consistency. Most existing VLM-based navigation methods treat the model as a black-box oracle\, querying it at every step on unstructured local observations\, which leads to redundant backtracking\, inefficient exploration\, and brittle behavior outside clean simulation. This thesis argues that ObjectNav is a system-level problem rather than a single-policy learning task: its sub-challenges of semantic understanding\, complex spatial structure\, and long-horizon planning should be explicitly decoupled and handled by cooperating modules\, with the VLM asked to reason only at the level where it is reliable. We build such a system and carry it step by step from simulation to floor-scale\, cross-embodiment deployment in the real world. \nWe first develop the core of this system in simulation\, where the robot incrementally organizes what it has seen into a structured scene representation and the VLM reasons only at a high level over it\, while efficient geometry-based exploration handles fine-grained navigation. This design achieves state-of-the-art success rate and navigation efficiency across four widely used benchmarks. We then bring the system into the real world and extend it into three cooperating levels that decouple semantic reasoning\, navigation planning\, and motion control. At the high level\, the structured scene representation summarizes the environment and the VLM provides semantically grounded navigation guidance over it. At the mid level\, a hierarchical room-based navigation strategy reserves VLM guidance for room-level decisions\, which makes effective use of its reasoning while keeping the system efficient. At the low level\, planned waypoints are executed by embodiment-specific motion control. Because only the lowest level depends on the robot\, the same system runs on a custom-built wheeled robot\, the Unitree Go2 quadruped\, and the Unitree G1 humanoid. Across 190 real-world experiments\, it substantially improves success rate and navigates 4-5x more efficiently than existing baselines. To our knowledge\, it is the first system to reliably and efficiently complete floor-scale\, long-range object navigation in complex real-world environments. Together\, these results show that real-world ObjectNav is solved not by a larger model or a single end-to-end policy\, but by a system that balances semantic intelligence with spatial reliability and isolates embodiment-specific control from embodiment-invariant reasoning. \nCommittee:\nJean Oh (advisor)\nJi Zhang\nZhixuan Liu
URL:https://www.ri.cmu.edu/event/from-simulation-to-the-real-world-a-multi-level-system-for-object-navigation-with-vision-language-models/
LOCATION:NSH 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260911T120000
DTEND;TZID=America/New_York:20260911T130000
DTSTAMP:20261003T070916
CREATED:20260901T200426Z
LAST-MODIFIED:20260901T200426Z
UID:153397-1789128000-1789131600@www.ri.cmu.edu
SUMMARY:Computational Lensing
DESCRIPTION:Abstract: From the cameras in our phones to the lenses in head-mounted displays\, optics shape both how we capture the world and how we experience virtual reality. Most conventional lenses are designed to bring a single plane into focus. In this talk\, we will discuss a new class of computational lens—referred to as a Split-Lohmann lens—that provides spatially varying control over focal length. This capability is achieved by combining a phase-only spatial light modulator with the cubic phase plates used in Lohmann/Alvarez focus-tunable lenses. The resulting computational lens enables new imaging and display capabilities\, including the ability to (i) make a flat display appear to have three-dimensional shape\, and (ii) capture all-in-focus images of highly non-planar scenes.
URL:https://www.ri.cmu.edu/event/computational-lensing/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:Faculty Events
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260911T143000
DTEND;TZID=America/New_York:20260911T153000
DTSTAMP:20261003T070916
CREATED:20260828T164708Z
LAST-MODIFIED:20260911T212228Z
UID:153335-1789137000-1789140600@www.ri.cmu.edu
SUMMARY:Augmenting Bee Colonies with Robotics and AI Technologies for Ecosystem Support
DESCRIPTION:Abstract: Earth’s ecosystems are facing a rapid decline in biodiversity\, with honeybees —keystone pollinators critical to ecosystem stability— being among the most affected. The EU-funded RoboRoyale project addresses this crisis by integrating advanced robotics and AI to augment the beehive\, enabling observation at unprecedented resolutions and scales. Featured on the cover of Science Robotics and receiving the 6th Edge of Government Award at the World Government Summit in 2024 our system tracks the Queen’s behaviors\, colony efficiency\, comb states\, and long-term foraging activities\, while advancing micro-robotic intervention capabilities to support hive health. In this talk\, I will discuss the challenges of developing this system\, share key findings regarding complex social interactions\, and explore the future potential of bio-hybrid research. \nBio: Erol Şahin is a Professor of Computer Engineering at Middle East Technical University (METU) and the founding Director of the Center for Robotics and AI (ROMER). Established with over 5 million Euros in funding\, ROMER spans 25\,000 square feet of state-of-the-art facilities\, including prototyping workshops\, specialized research arenas\, and advanced robotic platforms. Dr. Şahin earned his PhD in Cognitive and Neural Systems from Boston University\, following a BSc in Electrical and Electronics Engineering from Bilkent University and an MSc in Computer Engineering from METU. Before assuming his current position\, he worked as postdoctoral researcher at the Université Libre de Bruxelles.   Between 2013 and 2015\, Dr. Sahin spent two years at the Robotics Institute of Carnegie Mellon University during his sabbatical. His research interests include swarm robotics\, robotic learning\, and human-robot interaction—work that has secured more than 2.5 million Euros from the European Union\, TUBITAK\, and industrial partners. Notably\, his contributions to robotic learning were awarded a 53-DOF iCub humanoid platform through the RobotCub project in 2007. Dr. Şahin has edited several journal special issues and books\, currently serves as an Associate Editor for Adaptive Behavior\, and is a member of the Editorial Board for the Swarm Intelligence journal.
URL:https://www.ri.cmu.edu/event/augmenting-bee-colonies-with-robotics-and-ai-technologies-for-ecosystem-support/
LOCATION:1403 Tepper School Building
CATEGORIES:RI Seminar,Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2026/08/erol-sahin.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260915T153000
DTEND;TZID=America/New_York:20260915T170000
DTSTAMP:20261003T070916
CREATED:20260903T124903Z
LAST-MODIFIED:20260916T130809Z
UID:153443-1789486200-1789491600@www.ri.cmu.edu
SUMMARY:Making is Decision Making
DESCRIPTION:Abstract:  Human creation of high-quality content requires making decisions – from coarse\, high-level decisions about content and style\, to precise low-level decisions about the color of an individual pixel. Modern generative AI promises to serve a collaborative assistant capable of executing design decisions specified in simple text prompts into high-quality content. Yet today’s AI systems are poor collaborators. They frequently misinterpret user intent\, while users lack a predictive conceptual model of how an AI will interpret a prompt or why it produces a particular result. Without shared conceptual grounding\, collaboration devolves into trial and error\, with users repeatedly rewriting prompts\, using the AI to generate a result and then adjusting the prompt to try again\, in the hope of obtaining the desired outcome. In this talk I’ll argue that for generative AI to fulfill its promise we must develop techniques and interfaces that enable users and AI models to establish shared conceptual grounding. I’ll outline two complementary research challenges; First\, we must identify the concepts that human creators commonly use when reasoning about and communicating within a content creation domain.  Here\, I’ll show how we might extend methods from cognitive psychology to elicit\, represent\, and analyze domain-specific conceptual structures.  Second\, we must develop interfaces for teaching these concepts to AI models. While machine learning is rapidly advancing methods for teaching AI new concepts\, I’ll show how adapting these techniques to the diverse domains of human content creation requires new interaction techniques that support the way people think\, communicate\, and create. Finally\, I’ll demonstrate a few implementations of these ideas that we have developed in our group at Stanford. \nBio: Maneesh Agrawala is the Forest Baskett Professor of Computer Science and Director of the Brown Institute for Media Innovation at Stanford University. He is also a consulting AI Scientist at Roblox. He works on computer graphics\, human computer interaction and visualization. His focus is on investigating how cognitive design principles can be used to improve the effectiveness of audio/visual media. The goals of this work are to discover the design principles and then instantiate them in both interactive and automated design tools. Honors include an Okawa Foundation Research Grant (2006)\, an Alfred P. Sloan Foundation Fellowship (2007)\, an NSF CAREER Award (2007)\, a SIGGRAPH Significant New Researcher Award (2008)\, a MacArthur Foundation Fellowship (2009)\, an Allen Distinguished Investigator Award (2014) and induction into the SIGCHI Academy (2021). He was named an ACM Fellow in 2022.
URL:https://www.ri.cmu.edu/event/making-is-decision-making/
LOCATION:Gates-Hillman Center 4401
CATEGORIES:Special Events
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2026/09/9-15-26-agrawala-macarthur3-head-square.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260918T100000
DTEND;TZID=America/New_York:20260918T113000
DTSTAMP:20261003T070916
CREATED:20260908T155758Z
LAST-MODIFIED:20260908T155947Z
UID:153494-1789725600-1789731000@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - Ingrid Navarro Anaya
DESCRIPTION:Date: September 18th\, 2026 \nTime: 10:00 AM (ET) \nZoom: link \nLocation: NSH 4305 \nType: PhD Thesis Defense \nWho: Ingrid Navarro Anaya \n  \nTitle: Towards Generalizable Motion Prediction under Distribution Shifts \n  \nAbstract: \nAutonomous robots are increasingly expected to operate in dynamic\, human-centered environments. To do so safely and efficiently\, they must reason about how people move and interact. In domains like driving\, social navigation\, and aviation\, a common approach is to learn models of human motion directly from recorded data and use the resulting priors to inform downstream systems like simulators and planning stacks. \n  \nDespite the growing availability of datasets\, benchmarks\, and modeling techniques\, state-of-the-art methods remain unreliable for real-world deployment\, often generalizing poorly to novel environments and rare events. Much of this stems from recorded datasets covering few environments and few safety-relevant events relative to what a deployed system will ultimately encounter. Broadly\, the field has addressed this in four main ways: validating autonomy stacks on the road\, collecting more data\, synthesizing relevant scenarios\, and adapting at test time. These are all valuable and necessary strategies\, but each is bounded by risk\, cost\, the sim-to-real gap\, or the difficulty of reliably detecting and adapting to a shift\, respectively. \n  \nThese strategies share the premise that the data we hold is insufficient. This dissertation argues that such data is also underexploited and thus pursues a complementary direction\, asking how much actionable signal existing datasets already contain but current practice overlooks. We do so through a recurring paradigm we call scenario characterization: describing a scenario in terms of a property of interest and acting on that description downstream. We use this paradigm in three ways. The first is for guidance and abstraction\, shaping what a model trains or optimizes. The second is for mining and benchmarking\, determining what a model trains on and what is withheld to test it. The third is for analysis\, fixing the basis on which results are studied. We apply these uses across two settings: in-distribution generalization\, which draws mainly on guidance and abstraction\, and generalization under distribution shift\, which draws on all three and is where the main contributions concentrate. \n  \nWe further argue that evidence for generalization is typically gathered only within a single domain\, so a claim that holds there is rarely challenged elsewhere. This is largely because\, outside autonomous driving\, no motion domain offers comparable infrastructure and scale. This dissertation therefore contributes aviation as a testbed\, introducing a large-scale framework and dataset for airport surface movement forecasting. \n  \nThrough this framework and our experimental settings\, we expose generalization failures that would otherwise have remained hidden in aggregate metrics. We also enable reading aviation and driving on a common basis\, hinting at what transfers and what does not. Finally\, we also provide early evidence that the signal recovered through our framework may help systems gain the risk awareness needed to handle real-world critical events. \n  \nThesis Committee: \nJean Oh (co-chair) \nJonathan Francis (co-chair) \nSebastian Scherer \nAndrea Bajcsy \nAlexandre Alahi (EPFL)
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-ingrid-navarro-anaya/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260918T113000
DTEND;TZID=America/New_York:20260918T123000
DTSTAMP:20261003T070916
CREATED:20260915T180113Z
LAST-MODIFIED:20260915T201300Z
UID:153571-1789731000-1789734600@www.ri.cmu.edu
SUMMARY:First Light: Spaceborne multi-view computational tomography (CT)
DESCRIPTION:Abstract:  We devise multiview tomography from space. This talk will show the First Light of CloudCT\, the very first image taken since we launched our first satellite. This opens opportunities for scientific observations and technologies\,  can transform climate research and potentially even affect medical X-ray CT.  Tasks and solutions involve both novel imaging hardware and computational algorithms\, based on machine learning and differential rendering. The key idea is that advanced computing enable computed tomography of volumetric scenes\, based scattered radiation.  CloudCT\, funded by the ERC\, involves 10 nano-satellites to fly in an unprecedented formation\, to capture the same scene (cloud fields) from multiple views simultaneously. We encounter challenges of  polarimetric self-calibration in orbit\, and estimation of 3D volumetric distribution of microphysical properties. \nBio:  Prof. Yoav Schechner\, heads the Asher Space Research Inst. and is a faculty member at the Viterbi Faculty of Electrical and Computer  Engineering\, Technion. He is a Principal Investigator and Coordinator of the ERC CloudCT space project\, and is on the science team of two additional space-imaging projects\, C3IEL (by ISA/CNES) and MAIA (by NASA/ASI). He was a Research Scientist at Columbia U.\, and a Visiting Associate in Caltech\, the Jet Propulsion Laboratory (JPL) and MIT.  His main interests are diverse forms of computational imaging.  We won the Best Paper Award\, ICCP 2013+2018\, Best Student Paper Award\, CVPR 2017\, Distinguished Teaching Awards\, Technion\, Fumio Okano Best 3D Paper Award\, and others. He trained in physics (BA+MSc) and EE (PhD). He is the Mark and Diane Seiden Chair in Science and a Landau Fellow – supported by the Taub Foundation. \nSponsor:\nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one\nAI-powered video generation platform that leverages advances in\ncomputer vision\, generative modeling\, and multimodal learning to make\nhigh-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/first-light-spaceborne-multi-view-computational-tomography-ct/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:Seminar,VASC Seminar
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260918T150000
DTEND;TZID=America/New_York:20260918T163000
DTSTAMP:20261003T070916
CREATED:20260914T133740Z
LAST-MODIFIED:20260914T133740Z
UID:153537-1789743600-1789749000@www.ri.cmu.edu
SUMMARY:A Reactive Vibration Compensation System for Discrete Material Deposition Processes
DESCRIPTION:Abstract:  \n\n\n\n\nUltra Large Format Deposition (ULF-D) systems perform precision material deposition over large workspaces\, yet their extended and mechanically compliant structures can produce configuration-dependent vibration at the tool mounted to the robot’s end-effector. In inkjet printing\, vibration perturbs the deposition tool such that each discrete deposit (a single ink drop) lands displaced from its intended location on the target surface\, thus creating errors* in the deposited pattern. Conventional motion-control vibration compensation focuses on suppressing the vibration and may be limited by uncertainty in vibration dynamics models\, insufficient control bandwidth\, or restricted access to the underlying motion controller. \n\n\nAn alternative approach to reducing deposition error is process control\, which monitors and adjusts process variables to achieve a desired output. In inkjet printing\, each drop is released by a digital trigger signal\, so the instant the trigger fires determines where along the path the drop lands. This timing can be controlled independently of the robot’s trajectory. \n\n\nRather than applying corrective robot motion\, this thesis presents a process-control framework that adapts deposition-event timing to compensate for errors caused by bounded tool vibration. The framework combines a synchronized\, low-latency sensing and actuation architecture; a tool state estimation pipeline that fuses high-rate (1 kHz) inertial measurements with lower-rate (100 Hz) global pose measurements; and a reactive scheduler that triggers deposition events according to the estimated progression of the tool along its planned path. Together\, these components decrease deposition error by adjusting deposition timing with high-rate vibrating tool state feedback without modifying the robot’s motion-control stack. \n\n\nExperiments on a ULF-D emulation testbench evaluate the framework across multiple trajectories under both structured and unstructured vibration conditions. The proposed scheduler reduces mean deposition error by 67% to 91% relative to fixed-time firing\, keeping error between 0.1 mm and 1.4 mm across all tested conditions despite tool vibration amplitudes of up to 11 mm.\n\n\n\n*The mean absolute difference between the desired and measured spacing of adjacent line features in a deposition pattern. \n\n\n\n\n\nCommittee:  \n\n\nDr. Howie Choset (advisor) \n\n\nDr. Wennie Tabib \n\n\nDarwin Mick
URL:https://www.ri.cmu.edu/event/a-reactive-vibration-compensation-system-for-discrete-material-deposition-processes/
LOCATION:Gates Hillman Center 4405
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260921T153000
DTEND;TZID=America/New_York:20260921T163000
DTSTAMP:20261003T070916
CREATED:20260908T153454Z
LAST-MODIFIED:20260908T153454Z
UID:153488-1790004600-1790008200@www.ri.cmu.edu
SUMMARY:Decision-Making in a World of Latent Particles
DESCRIPTION:Abstract: Robots must often make decisions in scenes containing many objects: they need to identify what is present\, understand where objects are\, predict how they will interact\, and choose actions accordingly. Learning these capabilities directly from pixels is challenging\, especially when the number and arrangement of objects can change from one scene to another. \nIn this talk\, I will presentDeep Latent Particles (DLP)\, aself-supervisedobject-centric representation that describes a visual scene as a set of compact latent particles. Each particle captures the location and visual properties of a discovered object or object part\, providing an interpretable bridge between raw images and multi-object decision-making. \nI will show how DLP can serve as a representation for learning robotic policies from online reinforcement learning\, offline data\, and demonstrations. In particular\, policies built on these representations can generalize compositionally to scenes containing more objects than were present during training. \nI will then introduce Latent Particle World Models(ICLR 2026 Oral)\, which learn to predict how collections of latent particles evolve over time. These object-centric world models support multi-view observations and flexible conditioning\, enabling prediction and decision-making in rich visual environments. I will discuss how this perspective connects to diffusion-based policies and world action models (WAMs)\, and conclude with a look toward self-supervised 3D object-centric learning for robots that can perceive\, predict\, and act in three-dimensional worlds. and which will be quietly subsumed by the next scale-up. \n  \nBio: Tal Daniel is a Postdoctoral Fellow at Carnegie Mellon University’s Robotics Institute\, working with Prof. Deepak Pathak and Prof. David Held. He received his Ph.D. in Electrical and Computer Engineering from the Technion\, advised by Prof. Aviv Tamar. His research spans self-supervised and object-centric representation learning\, generative modeling\, reinforcement learning\, and robotics\, with a focus on learning representations and world models. \n  \nHomepage: https://taldatech.github.io \n  \nSponsor: \nThe VASC seminar is generously sponsored by HeyGen\, an all-in-oneAI-powered video generation platform that leverages advances incomputer vision\, generative modeling\, and multimodal learning to makehigh-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/decision-making-in-a-world-of-latent-particles/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/png:https://www.ri.cmu.edu/app/uploads/2026/09/9-21-26-tal_daniel.png
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260923T141500
DTEND;TZID=America/New_York:20260923T163000
DTSTAMP:20261003T070916
CREATED:20260921T143237Z
LAST-MODIFIED:20260923T173509Z
UID:153762-1790172900-1790181000@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Defense - Bardienus Duisterhof
DESCRIPTION:Date: 9/23/2026 \nTime: 2:00 PM – 4:00 PM \nZoom: https://cmu.zoom.us/j/98474543738?pwd=PVQo6JfRsWvuB9eWLqbatVivVv7CTj.1 \nLocation: NSH 4305 \nType: PhD Thesis Defense \nWho: Bardienus Pieter Duisterhof \nTitle: Spatiotemporal World Models for Robot Manipulation \nAbstract: \nRobot manipulation aims to automate tasks that are too dull\, dirty\, or dangerous for humans. This future requires remarkable resource efficiency: robots must adapt to new tasks with limited data and compute while meeting stringent performance requirements. Current systems can succeed at dexterous tasks but require substantial resources to meet industrial standards. One possible explanation for this inefficiency is that frontier models learn directly from RGB images\, which are typically dominated by content irrelevant to robots. Previous work has addressed this problem through task-specific feature engineering\, improving efficiency by focusing on task-relevant information. Can we construct similarly focused representations that remain scalable and broadly applicable? This thesis explores spatiotemporal representations—representations of 3D geometry and motion over time—focusing on their reconstruction\, generation\, and robot applications. \nThe first part of this thesis addresses unconstrained spatiotemporal reconstruction. Robots may benefit from representations that place past observations in a precise spatial and temporal context. We contribute methods that improve 3D reconstruction and calibration from arbitrary image sets and lens configurations. We show that calibrated multi-camera setups and neural rendering yield precise reconstruction in dynamic scenes\, including highly deformable objects such as cloth. \nIn the second part of this thesis\, we investigate learning spatiotemporal generative priors for robot manipulation. Humans can infer plausible geometry and dynamics from a single observation. Can we instill similar priors into robots? We contribute methods that can infer depth maps\, complete object geometry\, and predict object dynamics. With Modality Forcing\, we investigate text-to-image pre-training as a way to learn geometric priors. With PointZero\, we use 3D point track completion to learn spatiotemporal priors without any robot data. \nThe final part of this thesis considers spatiotemporal world models applied to robot manipulation. First\, we show that PointZero improves performance in robot manipulation tasks\, including imitation learning and action-conditioned dynamics prediction. Next\, we investigate how predicting future scene states can guide action generation in world-action models (WAMs). In 3PoinTr\, we show that 3D point tracks can serve as compact task plans that support transfer from human demonstrations to robot execution. In ModAR\, we autoregressively denoise multiple future modalities and robot actions\, achieving the best performance among the WAM formulations tested. We systematically study which modalities contribute most to manipulation performance and find that predicting RGB images provides no consistent additional benefit at the scale studied. \nTogether\, these works connect reconstruction\, learned geometric and dynamics priors\, and action generation to study how spatiotemporal representations can support efficient\, scalable robot learning. \nCommittee:  \nJeffrey Ichnowski (Chair)\nDeva Ramanan\nShubham Tulsiani\nAbhishek Gupta (University of Washington)
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-defense-bardienus-duisterhof-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260925T110000
DTEND;TZID=America/New_York:20260925T130000
DTSTAMP:20261003T070916
CREATED:20260911T143145Z
LAST-MODIFIED:20260911T164700Z
UID:153524-1790334000-1790341200@www.ri.cmu.edu
SUMMARY:Adaptive Cameras: Bridging Novel Sensors and Robot Perception
DESCRIPTION:Abstract:  Most cameras on robots today capture images without considering scene content. In contrast\, animal eyes have fast mechanical movements that control how the scene is imaged in detail by the fovea\, where visual acuity is highest. The prevalence of active vision during biological imaging\, and the wide variety of it\, makes it very clear that this is an effective visual design strategy for robot vision. In this talk\, I will cover our recent work on creating *both* new camera designs and novel robot perception algorithms to enable adaptive and selective active vision and imaging inside cameras and sensors. \nBio:  Sanjeev J. Koppal is an Associate Professor at the University of Florida’s Electrical and Computer Engineering Department and is a Kent and Linda Fuchs Faculty Fellow. He also held a UF Term Professorship for 2021-23. Sanjeev is the Director of the FOCUS Lab at UF. Since 2022\, Sanjeev has been an Amazon Scholar with Amazon Robotics. Prior to joining UF\, he was a researcher at the Texas Instruments Imaging R&D lab. Sanjeev obtained his Masters and Ph.D. degrees from the Robotics Institute at Carnegie Mellon University. After CMU\, he was a postdoctoral research associate in the School of Engineering and Applied Sciences at Harvard University. He received his B.S. degree from the University of Southern California in 2003 as a Trustee Scholar. He is a co-author on best student paper awards for ECCV 2016 and NEMS 2018\, and work from his FOCUS lab was a CVPR 2019 best-paper finalist. Sanjeev won an NSF CAREER award in 2020 and is an IEEE Senior Member and an Optica Senior Member. He won a UF ECE Department Teaching Award in 2024. His interests span computer vision\, computational photography and optics\, novel cameras and sensors\, 3D reconstruction\, physics-based vision\, and active illumination.\n\nHomepage: https://focus.ece.ufl.edu/ \nSponsor:\nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one\nAI-powered video generation platform that leverages advances in\ncomputer vision\, generative modeling\, and multimodal learning to make\nhigh-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/adaptive-cameras-bridging-novel-sensors-and-robot-perception/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2026/09/9-25-26-SanjeevKoppal.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260928T153000
DTEND;TZID=America/New_York:20260928T163000
DTSTAMP:20261003T070916
CREATED:20260914T232638Z
LAST-MODIFIED:20260914T232638Z
UID:153556-1790609400-1790613000@www.ri.cmu.edu
SUMMARY:From Capturing People to Teaching Robots
DESCRIPTION:Abstract:   Equipping AI and robotic systems with the ability to understand human behavior is essential for enabling them to assist people across a wide range of everyday applications. This need is more pressing than ever: the heaviest consumers of such knowledge are no longer perception systems alone\, but robot policies that must learn to act in the physical world. Yet the high-quality 3D human motion data required to learn this knowledge remains extremely scarce. \nIn this talk\, I will present our lab’s efforts to scale and enrich 3D human motion data by capturing everyday movements and natural human-object interactions\, with the ultimate goal of teaching robots to move like humans. \nI will first introduce our multi-year effort in building multi-camera capture systems\, from Panoptic Studio to ParaHome\, and most recently OmniRoboHome\, a new system designed to capture human-object interactions in natural home environments. Next\, I will present a complementary direction: learning everyday interactions and affordances from generative image and video models\, which offer a scalable source of human behavior priors that capture systems alone cannot reach. Finally\, I will discuss the missing pieces that vision and image models cannot provide\, including physics\, contact\, and the gap across diverse embodiments\, together with our recent efforts to fill them. \nBio:  Hanbyul Joo is an associate professor at Seoul National University (SNU) in the Department of Computer Science and Engineering. Before joining SNU\, Hanbyul was a Research Scientist at Facebook AI Research (FAIR)\, Menlo Park. Hanbyul received his PhD from the Robotics Institute at Carnegie Mellon University. He is a recipient of the Samsung Scholarship\, the Okawa Foundation Research Grant\, and the Best Student Paper Award at CVPR 2018. \nHomepage: https://jhugestar.github.io/ \nSponsor:\nThe VASC seminar is generously sponsored by HeyGen\, an all-in-one\nAI-powered video generation platform that leverages advances in\ncomputer vision\, generative modeling\, and multimodal learning to make\nhigh-quality video creation both scalable and accessible.
URL:https://www.ri.cmu.edu/event/from-capturing-people-to-teaching-robots/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:Seminar,VASC Seminar
ATTACH;FMTTYPE=image/jpeg:https://www.ri.cmu.edu/app/uploads/2026/09/9-28-26-han_dec_2017-4.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260930T160000
DTEND;TZID=America/New_York:20260930T170000
DTSTAMP:20261003T070916
CREATED:20261001T142902Z
LAST-MODIFIED:20261001T142902Z
UID:153867-1790784000-1790787600@www.ri.cmu.edu
SUMMARY:RI PhD Speaking Qual - Jianjin Xu
DESCRIPTION:Date: Wednesday\, Sep 30\, 2026Time: 16:00 – 17:00 PMLocation: NSH 4305Zoom: https://cmu.zoom.us/j/97595860992?pwd=voNXd0zhn8HdqFakFOF3At8y0NbydV.1Title: Efficient 3D Avatar Reconstruction with Geometric Guidance \nAbstract:3D animatable human avatars are widely used in film making\, game characters\, and telepresence. To create these avatars efficiently\, researchers propose to reconstruct the avatar from several input images with a feedforward network. These networks are mostly transformers trained with large proprietary data and thousands of H100 GPU hours. However\, do we really need such a scale of data and compute for this task? \nOur answer is no. In this talk\, we present ARG-Avatar\, a lightweight network with only 68M trainable parameters\, yet achieves SOTA performance on OOD testing data with 13x less training compute to the best baseline. We will go through the two core components of ARG-Avatar. The first is FACRoPE\, which inject geometric guidance into the attention with RoPE mechanism. We propose a novel coordinate formulation named Foreground Avatar Coordinates (FAC)\, to associate let the network tokens focus on its corresponding image regions. The second is Intermediate Token Rendering (ITR)\, which decodes a coarse avatar from intermediate network tokens during forward pass. We show that these components are all beneficial in the ablation study. In conclusion\, we show that by properly injecting geometric into attention\, we make an architecture that learns more efficiently and effectively for feedforward avatar reconstruction. \n\n—\n \n\nBest regards\, \nJianjin Xu.\nCarnegie Mellon University\, Ph.D. in Robotics\nhttps://atlantixjj.github.io/
URL:https://www.ri.cmu.edu/event/ri-phd-speaking-qual-jianjin-xu/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
END:VCALENDAR