BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251103T160000
DTEND;TZID=America/New_York:20251103T173000
DTSTAMP:20260922T061750
CREATED:20251027T134212Z
LAST-MODIFIED:20251027T134212Z
UID:149192-1762185600-1762191000@www.ri.cmu.edu
SUMMARY:Unifying Perception and Creation with Generative Models
DESCRIPTION:Abstract: \nRecent advances in large-scale generative modeling have reshaped our understanding of visual intelligence. While models such as diffusion and autoregressive transformers have achieved remarkable success in image and video synthesis\, their potential for visual perception and understanding remains underexplored. This thesis investigates how generative models can serve as powerful visual learners—bridging the long-standing divide between generative and discriminative paradigms. \nWe begin with REM (Refer Everything Models)\, a framework for referring video segmentation built upon text-to-video diffusion models. By preserving generative representations and fine-tuning on narrow-domain datasets\, REM achieves state-of-the-art results on standard benchmarks and demonstrates strong generalization to unseen domains. \nBuilding on this foundation\, we introduce a unified perceptual–generative framework that repurposes a single diffusion model across a broad spectrum of computer vision and image restoration tasks. Through joint training and systematic evaluation over 15 tasks spanning perception and synthesis\, we show that diffusion-based models deliver superior or comparable performance to discriminative counterparts\, revealing their intrinsic ability to encode rich\, multi-modal world representations. \nFinally\, we extend our exploration to visual autoregressive (VAR) models\, presenting the first unified architecture capable of efficiently solving the same 15 tasks within a single framework. We show that latent-variable designs\, particularly those leveraging variational autoencoders\, are key to achieving coherent multi-modal understanding and consistent generation. Compared with diffusion counterparts\, VAR-based models offer substantial gains in latency\, scalability\, and output consistency. \nCollectively\, these studies offer a cohesive perspective on unifying perception and synthesis through generative modeling\, charting a path toward general-purpose visual foundation models that seamlessly integrate understanding\, reasoning\, and creation. \n\nThesis Committee: \nMartial Hebert\, Chair \nDeva Ramanan \nJun-Yan Zhu \nAlexei Efros\, University of California\, Berkeley \nYu-Xiong Wang\, University of Illinois Urbana-Champaign \nPavel Tokmakov\, Toyota Research Institute \nDraft of Document
URL:https://www.ri.cmu.edu/event/unifying-perception-and-creation-with-generative-models/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Defense,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251107T150000
DTEND;TZID=America/New_York:20251107T163000
DTSTAMP:20260922T061750
CREATED:20251001T150038Z
LAST-MODIFIED:20251104T153929Z
UID:148979-1762527600-1762533000@www.ri.cmu.edu
SUMMARY:Visual-Tactile Synthesis for Texture Generation
DESCRIPTION:Abstract: Recent advances in generative models have enabled the creation of highly realistic visual content\, yet they remain limited to visual perception alone. In contrast\, human interaction with the physical world is inherently multimodal — we not only see textures but also feel them. This gap motivates the goal of my thesis: to build generative models that jointly synthesize visual and tactile modalities for material and texture generation. By unifying what we see and what we touch\, such models can drive new forms of physically grounded content creation\, from robotics and virtual reality to material design.\nHowever\, extending generative modeling to touch presents unique challenges: tactile data is scarce\, noisy\, and expensive to collect\, and there is no large-scale paired dataset linking visual appearance with tactile response. To address these challenges\, my research explores three synergistic directions. \nPart I: I introduce controllable visual-tactile synthesis models that jointly generate aligned visual and tactile textures from shared latent representations\, enabling explicit control over appearance and feel. \nPart II: I propose tactile-aware 3D generation frameworks that integrate tactile sensing into 3D diffusion pipelines\, allowing models to infer physically grounded material properties from visual cues and geometry. \nPart III: Building on these insights\, I aim to develop scalable multimodal generation systems that leverage large vision and language foundation models and physics priors to synthesize novel materials directly from text or image input\, without relying on extensive paired tactile data. \n\n \nThesis Committee:\nJun-Yan Zhu (Co-chair)\nWenzhen Yuan (Co-chair)\nShubham Tulsiani\nAndrew Owens (Cornell Tech)\n\nThesis Document
URL:https://www.ri.cmu.edu/event/ruihan-gao-phd-thesis-proposal/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251117T103000
DTEND;TZID=America/New_York:20251117T113000
DTSTAMP:20260922T061750
CREATED:20251113T150929Z
LAST-MODIFIED:20251113T150929Z
UID:149416-1763375400-1763379000@www.ri.cmu.edu
SUMMARY:Towards Modernization of Long-Range Image-Space Planning for Off-Road Navigation
DESCRIPTION:Abstract: \nThis thesis revisits long-range\, image-space planning for off-road navigation and modernizes the classical first-person view (FPV) paradigm by building upon recent advances in perception. It introduces a lightweight depth calibration scheme\, analytic configuration-space (C-space) transforms\, interpretable frontier selection\, and a pixel-space A* planner with validated heuristic soundness. Concretely\, we (i) make monocular depth metrically usable at test time via an affine\, log-domain calibration with sparse LiDAR; (ii) derive and implement a depth-aware FPV C-space inflation that projects vehicle width/length analytically and realizes it with separable row/column sliding-maximum filters augmented by per-pixel depth consistency checks; (iii) propose transparent\, angular-sector frontiering that reasons jointly about traversability cost and minimal lethal depth\, alongside goal-aware revalidation; and (iv) preserve A* admissibility/consistency in image space through a simple cost renormalization that avoids silent suboptimality in low-cost free space. \nWe evaluate the resulting\, modular sub-stack in the high-fidelity Falcon simulator [2] under a shared ROS graph. Using a common perception front-end and planning back-end\, we compare three frontiering strategies — (1) a LAGR-style row-wise baseline\, (2) an LRN-inspired openness heuristic adapted to operate with explicit depth and cost\, and (3) an Angular Cost & Depth (ACD) variant that couples average cost with a minimum lethal-depth constraint. Across diverse courses (e.g.\, farm\, desert\, mixed terrain)\, the calibrated monocular depth reduces error versus raw monocular predictions\, and both LRN-inspired and ACD frontiering tend to outperform the purely row-wise baseline on longer traverses. We view these as encouraging indications rather than definitive claims: all results are in simulation\, with performance subject to perception quality\, calibration coverage\, and environment diversity. \n\nScope and limitations are explicit. The work was conducted over a short project window (May 2025–present) and prioritized stabilizing the proposed sub-stack in conjunction with a core FieldAI [1] stack through the high-fidelity Falcon simulator [2] and ROS integration into a reliable\, end-to-end testing framework. No real-world deployments were performed\, and the current implementation targets ROS1. Nevertheless\, the design is intentionally modular and auditable to make it relatively straightforward to integrate the sub-stack into existing off-road autonomy stacks lacking explicit long-range planning. Such integration will enable better autonomy by offering a practical bridge between classical image-space efficiency and metric-world robustness. \n[1] FieldAI: https://www.fieldai.com \n[2] Falcon: https://www.duality.ai \nCommittee: \nWenshan Wang\, Chair\nSebastian Scherer\, Co-Chair\nMaxim Likhachev\nSamuel Triest\, RI PhD
URL:https://www.ri.cmu.edu/event/towards-modernization-of-long-range-image-space-planning-for-off-road-navigation/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251119T130000
DTEND;TZID=America/New_York:20251119T140000
DTSTAMP:20260922T061750
CREATED:20251114T143444Z
LAST-MODIFIED:20251114T145002Z
UID:149445-1763557200-1763560800@www.ri.cmu.edu
SUMMARY:Building Robot Hands and Teaching Dexterity
DESCRIPTION:Abstract:  \nOur human hands are masterpieces of power and precision\, capable of typing\, hammering\, or delicately using chopsticks. Yet most robots today still rely on simple two-finger grippers in controlled settings because dexterous hands are costly and difficult to deploy. To close this gap\, I will introduce my LEAP Hands\, high-performance\, low-cost\, and easy-to-assemble robotic hands that have become the most widely used platform for dexterous manipulation research. LEAP Hand V1 employs motor-in-joint actuation for simplicity\, while V2 introduces a novel hybrid rigid–soft structure that delivers exceptional strength and durability.  I will then show how large-scale human video/motion data and simulation techniques can teach human-like manipulation skills across diverse environments.  By tightly integrating mechanical design and machine learning\, my open-source robot hands achieve unprecedented levels of dexterity for a variety of everyday tasks. \nCommittee:\nProf. Deepak Pathak (advisor)\nProf. Nancy Pollard \nProf. Abhinav Gupta \nProf. Jitendra Malik \nDr. Ankur Handa \n  \nA draft of my thesis proposal is available here: \nhttps://kennyshaw.net/phd_thesis_proposal.pdf
URL:https://www.ri.cmu.edu/event/building-robot-hands-and-teaching-dexterity-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251120T140000
DTEND;TZID=America/New_York:20251120T150000
DTSTAMP:20260922T061750
CREATED:20251114T205305Z
LAST-MODIFIED:20251114T205305Z
UID:149469-1763647200-1763650800@www.ri.cmu.edu
SUMMARY:Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models
DESCRIPTION:Abstract:\nTransferring skills between different objects remains one of the core challenges of open-world robot manipulation. Generalization needs to take into account the high-level structural differences between distinct objects while still maintaining similar low-level interaction control. In this paper\, we propose an example-based zero-shot approach to skill transfer. Rather than treating skills as atomic\, we decompose skills into a prioritized list of grounded task-axis (GTA) controllers. Each GTAC defines an adaptable controller\, such as a position or force controller\, along an axis. Importantly\, the GTACs are grounded in object key points and axes\, e.g.\, the relative position of a screw head or the axis of its shaft. Zero-shot transfer is thus achieved by finding semantically similar grounding features on novel target objects. We achieve this example-based grounding of the skills through the use of foundation models\, such as SD-DINO\, that can detect semantically similar keypoints of objects. We evaluate our framework on real-robot experiments\, including screwing\, pouring\, and spatula scraping tasks\, and demonstrate robust and versatile controller transfer for each.\n\nCommittee:\nProf. Oliver Kroemer\nProf. Katerina Fragkiadaki\nProf. Zeynep Temel\nMark Lee
URL:https://www.ri.cmu.edu/event/grounded-task-axes-zero-shot-semantic-skill-generalization-via-task-axis-controllers-and-visual-foundation-models/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Speaking Qualifier,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251125T110000
DTEND;TZID=America/New_York:20251125T120000
DTSTAMP:20260922T061750
CREATED:20251118T142159Z
LAST-MODIFIED:20251118T142159Z
UID:149484-1764068400-1764072000@www.ri.cmu.edu
SUMMARY:Multi-View 4D Human Reconstruction under Interaction Scenarios
DESCRIPTION:Abstract:\nBuilding large-scale human datasets from multi-view videos is essential for advancing research in human behavior understanding\, virtual reality\, animation\, and robotics. Compared to traditional motion capture systems that rely on physical markers to track motion\, vision-based reconstruction not only enables the capture of human motion in unconstrained environments but also avoids altering human appearance with markers. However\, existing multi-view reconstruction methods often fail when humans interact closely with others or with objects due to severe occlusions and truncations introduced by complex activities. In this thesis\, we develop a markerless capture system capable of handling close human interactions and dexterous hand-object manipulations. Using this system\, we construct two large-scale human datasets\, Harmony4D and Contact4D\, which serve as landmarks for advancing fundamental research in human-centric AI\, such as human pose estimation\, contact estimation\, and motion generation. \nCommittee:\nKris Kitani\, Chair\nFernando De La Torre\nShubham Tulsiani\nErica Weng
URL:https://www.ri.cmu.edu/event/multi-view-4d-human-reconstruction-under-interaction-scenarios/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251125T140000
DTEND;TZID=America/New_York:20251125T150000
DTSTAMP:20260922T061750
CREATED:20251119T143738Z
LAST-MODIFIED:20251119T143738Z
UID:149496-1764079200-1764082800@www.ri.cmu.edu
SUMMARY:Vision-Based Multi-Wire Detection and Tracking for UAV Wire Approach
DESCRIPTION:Abstract: \nReliable detection and tracking of power lines is critical for enabling under-wire UAV approach and inductive power-line charging to extend UAV range. However\, wires are thin\, featureless\, and visually ambiguous structures that challenge traditional computer vision methods and degrade depth estimation accuracy. To address these challenges\, this thesis presents a fully passive\, camera-only multi-wire detection and tracking algorithm that operates in real time on lightweight onboard compute using only stereo RGB imagery. \nThe proposed framework integrates a lightweight classical vision pipeline with three-dimensional geometric reasoning and a multi-stage filtering process to produce robust wire instance detections. A complementary oriented object detection model is trained using labels generated by the classical pipeline\, leveraging its fine-tuned geometric outputs to improve resilience in challenging visual conditions. To track individual wire instances across frames\, we introduce a Kalman-filter-based tracking architecture that estimates both wire orientation and per-wire positional state while remaining robust to wire detection outliers and vehicle pose drift. The system is further expanded by testing a wire positional servoing approach using the tracked wire instances in simulation and is validated across a diverse range of data sources\, including simulation\, indoor testing\, handheld experiments\, and outdoor flight evaluations. \n\nCommittee:\nSebastian Scherer (advisor)\nWennie Tabib\nMohammad Mousaei
URL:https://www.ri.cmu.edu/event/vision-based-multi-wire-detection-and-tracking-for-uav-wire-approach/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:MSR Thesis Presentation,Student Talks
END:VEVENT
END:VCALENDAR