BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260113T120000
DTEND;TZID=America/New_York:20260113T133000
DTSTAMP:20260922T083157
CREATED:20260106T153223Z
LAST-MODIFIED:20260106T153223Z
UID:149856-1768305600-1768311000@www.ri.cmu.edu
SUMMARY:Self-supervised tactile perception for robot dexterity
DESCRIPTION:Abstract: \nHumans are incredibly dexterous. We interact with and manipulate tools effortlessly\, leveraging touch without a second thought. Yet\, replicating this level of dexterity in robots is a major challenge. While the robotics community\, recognizing the importance of touch in fine manipulation\, has developed a wide variety of tactile sensors\, how best to leverage these sensors for both perception and manipulation is unclear. In this thesis\, we address how to efficiently integrate tactile sensing for robot perception and dexterous manipulation. \nSpecifically\, we turn to self-supervised learning (SSL) to train tactile representations that can generalize across sensors\, standardize usage across downstream tactile tasks\, and further alleviate the need to collect labeled task data which is often impractical to collect for tasks such as uncalibrated force field estimation. To this end\, we discuss Sparsh and Sparsh-skin\, a family of SSL models for vision and magnetic-skin based tactile sensors respectively. Sparsh and Sparsh-skin are trained via self-distillation for full-hand tactile sensors in downstream tasks. We find that both Sparsh and Sparsh-skin not only outperform task and sensor-specific end-to-end models by a large margin\, but also that they are data efficient for downstream task training. \nSecond\, we note that existing work often overlooks the multimodal aspects of human touch\, such as vibration and heat sensing. We discuss Sparsh-X\, a compact tactile representation fusing image\, pressure\, audio and inertial measurements from the DIGIT360 sensor. With Sparsh-X we demonstrate that multimodal sensing improves both passive perception tasks as well as dexterous manipulation tasks such as in-hand rotation. \nFinally\, we present privileged tactile latent distillation (PTLD)\, a novel method to imbue tactile sensing in dexterous manipulation policies trained via reinforcement learning. PTLD avoids simulating tactile sensors and uses privileged sensors to bridge the sim-to-real gap. With PTLD\, we first show that one can improve existing RL trained policies such as in-hand rotation and then that it can enable learning more challenging tasks such as in-hand reorientation. \nJointly these contributions provide a path to leverage tactile sensing in both imitation and reinforcement learning based robot manipulation. \nThesis Committee Members: \nMichael Kaess\, chair\nShubham Tulsiani\nGuanya Shi\nMustafa Mukadam\, Amazon Robotics\nJitendra Malik\, UC Berkeley & Amazon FAR \nThesis Draft
URL:https://www.ri.cmu.edu/event/self-supervised-tactile-perception-for-robot-dexterity/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260113T143000
DTEND;TZID=America/New_York:20260113T160000
DTSTAMP:20260922T083157
CREATED:20260106T180847Z
LAST-MODIFIED:20260106T180847Z
UID:149866-1768314600-1768320000@www.ri.cmu.edu
SUMMARY:Efficient Visual Modeling with Adaptive Representations
DESCRIPTION:Abstract: \n\nWhile image understanding\, generation\, and manipulation have matured rapidly in recent years\, video remains challenging due to the significantly larger input size. As a result\, tasks such as generating long videos or understanding extended video sequences remain out of reach for current models due to their computational cost. This talk presents a series of works that address this issue by adapting ideas from video compression to accelerate visual model training and inference. I will first introduce Run-Length Tokenization (RLT)\, which modifies the vision transformer architecture to exploit temporal redundancy\, enabling substantial speedups without compromising accuracy. Next\, I will present FlowTok\, which incorporates motion vectors to extend RLT to dynamic scenes\, maintaining efficiency even under camera and object motion. I will then discuss Adaptive Patch Transformers (APT)\, which apply these principles to images by dynamically assigning larger patch sizes in low-complexity regions to reduce computation while preserving performance. We next apply these principles to video generation\, and propose SkipSR\, a cascaded generation framework that combines fast video super-resolution with cascaded diffusion models. Finally\, we introduce FPS-Bench\, a benchmark to systematically evaluate the impact of frame rate and resolution on downstream video understanding tasks\, offering insights into which aspects of fidelity truly matter for model performance. By unifying efficient video tokenization with scalable video synthesis and principled evaluation\, this thesis enables significantly faster visual models in both understanding and generation tasks\, unlocking further scaling. \n\n\n\nThesis Committee Members:\n \n\n\nLászló A. Jeni(co-chair)\nKris M. Kitani (co-chair)\nJun-Yan Zhu\n\nRohit Girdhar (Meta GenAI)\n\nLu Jiang (ByteDance)\n\n\n\n\n\n\n\n\n\n\n\n\n\nA draft of the thesis is available here: Thesis Draft
URL:https://www.ri.cmu.edu/event/efficient-visual-modeling-with-adaptive-representations/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260115T083000
DTEND;TZID=America/New_York:20260115T100000
DTSTAMP:20260922T083157
CREATED:20260108T145112Z
LAST-MODIFIED:20260108T145112Z
UID:149897-1768465800-1768471200@www.ri.cmu.edu
SUMMARY:Toward Scalable Architectures for Multimodal LLM-based Cooperative Autonomous Driving
DESCRIPTION:Abstract: Despite the tremendous progress made in autonomous driving over the years\, the safety of autonomous vehicles still requires further improvement before they can operate worldwide with full human trust. One principal safety concern is that each individual vehicle may have a limited field of view due to finite detection ranges\, potential sensor failures\, or occlusions caused by nearby large objects such as buses or trucks. This limitation in perception introduces additional challenges for the downstream planning and control modules\, making it more difficult for autonomous vehicles to generate safe driving decisions and actions.\nTo address this issue\, recent research has proposed vehicle-to-vehicle (V2V) and vehicle-to-everything (V2X) cooperative perception for autonomous driving. In such systems\, connected autonomous vehicles (CAVs) share their individual perception features with one another to improve overall cooperative detection accuracy. However\, most existing work focuses solely on the cooperative detection task\, without leveraging temporal information about the dynamic environment or considering other critical components of autonomous driving\, such as prediction and planning. \nTo broaden the scope of cooperative driving research\, my proposed doctoral research aims to explore multimodal large language model (LLM)–based cooperative autonomous driving\, motivated by several potential advantages of LLMs. First\, a single LLM-based model offers the flexibility to perform multiple tasks\, including perception\, prediction\, and planning\, within a unified framework. Second\, LLMs exhibit strong generalizability due to large-scale pretraining on diverse data. Third\, LLM-based driving models possess reasoning capabilities that enable them to handle long-tail driving scenarios that may not appear in the training data. Fourth\, natural language can serve as an effective and efficient communication interface for V2V\, V2X\, and human–vehicle interactions. \nWe have developed multimodal LLM-based cooperative autonomous driving architectures that enable end-to-end cooperative driving and generate suggested future trajectories for all CAVs through V2V communication. In addition\, we have designed a graph-of-thoughts reasoning framework to further enhance the reliability and interpretability of our multimodal LLM-based architecture. Finally\, we propose to develop a decentralized V2V framework using multimodal LLMs to improve the scalability and feasibility of future large-scale deployment. \n\n \n \nThesis Committee:\n\nStephen F. Smith (Chair)\nJohn Dolan\nDeva Ramanan\nMin-Hung Chen (NVIDIA)\n\n\n\n\n\nThesis proposal link
URL:https://www.ri.cmu.edu/event/toward-scalable-architectures-for-multimodal-llm-based-cooperative-autonomous-driving/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260120T120000
DTEND;TZID=America/New_York:20260120T133000
DTSTAMP:20260922T083157
CREATED:20260113T200035Z
LAST-MODIFIED:20260113T200035Z
UID:150081-1768910400-1768915800@www.ri.cmu.edu
SUMMARY:Towards Manipulation in the Blind: Motion Planning for Manipulation under Uncertainty using Contacts
DESCRIPTION:Abstract:\n \nHumans routinely rely on the sense of touch to better perceive the world. In environments characterized by poor lighting\, occlusions\, limited fields of view\, or sparse visual features\, contact feedback often becomes a primary source of information for perceiving the environment and successfully completing manipulation tasks. Everyday examples include locating a light switch in the dark\, retrieving an item from a high shelf\, or reaching for a valve at the back of a kitchen sink cabinet where visual access is severely restricted. In such settings\, humans actively reason about contacts to infer both the poses of objects of interest and the environmental obstacles in the workspace. Enabling robots to exhibit similar capabilities remains a fundamental challenge in autonomous manipulation. This thesis investigates search-based planning techniques that allow robots to effectively leverage contact feedback as a sensing modality\, enabling robust manipulation under uncertainty. \nThis work focuses on two broad classes of manipulation problems in which contact plays a critical role. The first class concerns object pose uncertainty\, where precise estimation of target object pose is required to complete high-precision manipulation tasks such as charger plug insertion\, pipe assembly\, or other tight-tolerance manipulation problems. In these settings\, even small pose errors on the order of a few millimeters can lead to failure. This problem class studies how robots can actively use contacts during execution to reduce object pose uncertainty to a level sufficient for successful task completion. The second class of problems addresses manipulation under environmental uncertainty\, where the locations and geometries of environmental obstacles are unknown or only partially observable. For example\, when reaching into a cluttered kitchen sink cabinet with unknown obstacles and pipes\, a robot must detect contacts\, infer obstacle locations\, and adapt its motion accordingly in order to safely reach the target (valve). Together\, these two problem classes capture a wide range of real-world scenarios in which contact-driven reasoning is essential. \nPlanning under object pose uncertainty naturally falls within the framework of Partially Observable Markov Decision Processes (POMDPs)\, which are computationally expensive to solve\, particularly in continuous and high-dimensional robotic domains. This thesis presents three complementary frameworks to address this challenge. The first is an experience-based preprocessing approach designed for semi-structured environments that require strong online performance. This framework leverages solutions to previously solved\, similar POMDPs to accelerate future planning queries while maintaining theoretical guarantees on solution quality. An offline database of policies is constructed and queried at execution time based on the current problem instance\, enabling fast online decision-making. The second framework targets less structured domains where preprocessing is impractical. It introduces an online closed-loop planning and execution approach that employs a hierarchical representation of uncertainty. By adaptively representing and reasoning about uncertainty\, this method significantly reduces planning time\, making online planning and execution feasible in more complex settings. The third contribution addresses a key computational bottleneck in these settings\, namely the high cost of belief space transition computations. To mitigate this issue\, the thesis proposes lazy heuristic search algorithms for POMDPs that defer expensive belief updates until they are necessary\, using approximate Q-value estimators to guide search. These lazy solvers substantially reduce planning time while preserving solution quality. \nFor the problem class of manipulation under environmental uncertainty\, this thesis develops an iterative planning and execution framework that tightly couples contact sensing\, environment prediction\, and motion planning. The system employs a torque-based contact detection and localization module capable of detecting contacts occurring anywhere along the robot manipulator. The history of detected contacts is used to construct a partial occupancy map of the workspace\, which is then extrapolated using learned occupancy estimators. A motion planning module reasons over this estimated occupancy representation to compute actions that are likely to safely and efficiently move the robot toward the goal. The framework is evaluated in simulation and on a real UR10e manipulator across two challenging domestic tasks: manipulating a valve under a kitchen sink surrounded by pipes and retrieving a target object from a cluttered shelf. \nOverall\, this thesis takes a step toward manipulation in the blind\, where robots explicitly leverage contact interactions as a primary sensing modality to plan and execute manipulation tasks under uncertainty. Rather than treating contact as a failure mode to be avoided\, we treat it as an informative observation that can be actively exploited to reduce uncertainty and guide motion. \nThesis Committee: \nProf. Maxim Likhachev\, Chair \nProf. Jeffrey Ichnowski \nProf. Oliver Kroemer \nProf. Mehmet Dogar\, University of Leeds \nA draft of the thesis document is available at: https://drive.google.com/drive/folders/1TCSil_tJksTWC_RweWGS2FHoZBjyWM8w?usp=drive_link
URL:https://www.ri.cmu.edu/event/towards-manipulation-in-the-blind-motion-planning-for-manipulation-under-uncertainty-using-contacts/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260127T120000
DTEND;TZID=America/New_York:20260127T133000
DTSTAMP:20260922T083157
CREATED:20260114T144316Z
LAST-MODIFIED:20260114T144316Z
UID:150084-1769515200-1769520600@www.ri.cmu.edu
SUMMARY:Rethinking Robot Safety: Adaptive and Scalable Methods for Real-World Autonomy
DESCRIPTION:Abstract:\nSafe autonomy in the real world requires more than safety in structured\, low-dimensional settings. Robots deployed in everyday environments must cope with non-stationarity—objectives and dynamics that change due to human preferences or evolving operating conditions—and must also scale safety reasoning to high-dimensional robots and environments\, where perception\, dynamics\, and safety constraints can be complex and tightly coupled. \nThis thesis presents safety methods that are both adaptive and scalable. We start with an approach for adapting to varying control objectives by incorporating short context demonstrations\, enabling rapid objective inference and controller adjustment in a physical human–robot collaboration task. Next\, we present a method for adapting safety to parameter-varying dynamics with formal guarantees\, leveraging the structure of provably safe controller synthesis to update safety certificates in real time to preserve feasibility. \nWe then shift to safety challenges that emerge in high-dimensional systems. For safety analysis at scale\, we propose λ-Reachability\, a method for learning safety value functions that interpolates between local self-consistency and long-horizon safety targets\, improving both feasible-boundary classification and safety-margin estimation in high-dimensional humanoid settings. Finally\, to handle the fact that learned safety representations can be imperfect in practice\, we present the Projected Safe Set Algorithm (p-SSA)\, a feasibility-aware safe control method that mitigates infeasible constraint sets arising from dense\, multi-body collision constraints in dexterous humanoid operation\, achieving robust collision avoidance in simulation and on real hardware. \nBy unifying adaptive control\, scalable safety analysis\, and feasibility-aware approaches\, this thesis takes a step toward robots that can operate safely amid the uncertainty\, complexity\, and interactivity of the real world. \nThesis Committee Members:\nProfessor Changliu Liu (chair)\nProfessor Zac Manchester\nProfessor Guanya Shi\nProfessor Chuchu Fan (MIT) \nLink to draft thesis document: https://www.dropbox.com/scl/fo/5yrayulkjhfndug6oc9n1/AIRsMhLiTr9EIlOoOAtXlio?rlkey=6o6u6jypjdg50o5psh8kopviy&st=f65avsbo&dl=0
URL:https://www.ri.cmu.edu/event/rethinking-robot-safety-adaptive-and-scalable-methods-for-real-world-autonomy/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
END:VCALENDAR