BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20261007T153000
DTEND;TZID=America/New_York:20261007T170000
DTSTAMP:20261009T020012
CREATED:20260929T175722Z
LAST-MODIFIED:20260929T180200Z
UID:153850-1791387000-1791392400@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Shashwat Singh
DESCRIPTION:RI CALENDAR EVENT \n\n\n\n\n\n\n\n\n\n\nWho: Shashwat Singh\nType: RI Thesis Proposal\nDate: Wednesday\, October 7\, 2026\nTime: 03:30PM – 05:00PM (ET)\nLocation: NSH 4305 \nZoom link: Here \nTitle: Extending Robot Mobility Through Mechanical Adaptation \nAbstract: \nUnstructured terrain presents diverse challenges to robot locomotion\, including obstacles\, deformable surfaces\, and confined passages. Terrain conditions also change with time\, as rainfall alters surface traction or shifting debris and rubble block previously traversable terrain. Addressing these challenges requires robots to adapt how they move and interact with their environments. In this thesis\, I investigate how mechanical adaptation can extend the mobility of centimeter-scale robots. \nFirst\, I study how switching between complementary locomotion modes extends mobility. A springtail-inspired microrobot (2.1 cm long; 0.98 g)\, combines crawling and jumping using a single actuator to overcome obstacles. TerraSkipper (5.8 cm long; 28 g) uses impulsive skipping to traverse sand and mud where fin-based crawling is ineffective\, while the RESCUE Jumper (8.2cm long; 125 g) reuses its wheel actuators to jump onto steps it cannot climb by rolling. \nI further examine how physical collaboration extends mobility rather than integrating every capability into a single robot platform. A centimeter-scale RESCUE roller (9 cm long; 100g) couples with other rollers and with a soft growing robot\, also known as Vine robot to share strength and power\, and improve mobility. In this work\, I used RESCUE Rollers to manipulate a Vine robot in a two-dimensional plane\, expanding its operational workspace. To extend this capability even further\, I propose developing Drone Roller (19.4cm long; 200g)\, a compact aerial–ground version of the RESCUE Roller that can fly over obstacles it cannot climb and guide the Vine robot in three-dimensional space. \nFinally\, I study how incorporating an adaptive spine into a robot’s body extends mobility. I designed BaSiL (31 cm long\, 240g)\, a wheeled robot whose tendon-driven beaded spine combines passive compliance with active bending. Actuating the spine in predefined sequences increases the height of steps the robot climbed more than sixfold compared with a rigidly constrained spine. Furthermore\, I propose to investigate continuous morphological adaptation using the All Wheel Morph (AWM) robot (48 cm long\, 4640g)\, whose rotating wheels vary continuously in shape between round wheels and leg-like configurations. I will evaluate AWM through controlled laboratory experiments on steps\, sand\, and mud\, along with real-world experiments. These experiments will test whether the optimal morphology varies continuously with terrain conditions or changes abruptly between distinct wheel or leg configurations. \nWith the proposed work\, I aim to understand how robot morphology can be matched to different terrain conditions\, explore the benefits and trade-offs of adaptation\, and understand design guidelines for improving terrain traversal. \n  \nThesis Committee Members:\nZeynep Temel (Chair)\nSarah Bergbreiter\nAaron Johnson\nPakpong Chirarattananon (University of Toronto) \nDraft of the Thesis Proposal Document
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-shashwat-singh/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260901T113000
DTEND;TZID=America/New_York:20260901T130000
DTSTAMP:20261009T020012
CREATED:20260820T203402Z
LAST-MODIFIED:20260820T203402Z
UID:153235-1788262200-1788267600@www.ri.cmu.edu
SUMMARY:Scalable Vision-Language Models through Unified 2D and 3D Representations
DESCRIPTION:Abstract:\nVision-language models have become remarkably capable on images and short videos\, yet they still struggle with two abilities central to embodied intelligence: spatial understanding and long-range temporal reasoning. A major reason is representational: today’s models process videos as long sequences of 2D patches\, so computation grows with observation length even when the underlying scene changes little. This thesis argues that organizing perception around persistent 3D structure rather than individual frames allows model complexity to scale with scene content rather than observation length\, leading to efficient inference on long videos while providing a stronger foundation for spatial reasoning.\n\nWe introduce a unified representation based on 3D feature clouds that handles both 2D and 3D inputs within a single architecture. Because the same model trains on abundant 2D image-text data alongside available 3D data\, it acquires better spatial understanding without sacrificing 2D performance. We demonstrate this on 3D instance segmentation (ODIN)\, extend it to broader vision-language tasks (UniVLG)\, and scale it to billion-parameter VLMs (Qwen-3D)\, showing consistent improvements in spatial understanding and inference efficiency. \nWe then move beyond static scenes to dynamic videos. We develop 3D scene representations (TrackEverything) that disentangle static and dynamic content\, deduplicate the scene across time\, and track all points in 3D throughout long videos. These representations convert videos into concise spatiotemporal structures that grow with scene complexity rather than video duration\, enabling new capabilities such as being able to track all points across all frames in long videos (1000+ frams). \nFinally\, we outline proposed work on two fronts: leveraging these dynamic 3D representations for spatial and motion reasoning in vision-language models\, and scaling dynamic 3D tracking to real-world video data using heterogeneous supervision beyond synthetic data. \nTogether\, this work makes the case for moving beyond 2D patch representations toward 3D-native models for better and more efficient video understanding. \nThesis Committee:\nKaterina Fragkiadaki (Chair)\nDeva Ramanan\nShubham Tulsiani\nLeonidas Guibas\, Stanford University\n\nThesis Proposal Link
URL:https://www.ri.cmu.edu/event/scalable-vision-language-models-through-unified-2d-and-3d-representations/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260803T153000
DTEND;TZID=America/New_York:20260803T170000
DTSTAMP:20261009T020012
CREATED:20260727T145526Z
LAST-MODIFIED:20260727T145526Z
UID:152896-1785771000-1785776400@www.ri.cmu.edu
SUMMARY:Simulate to Learn\, Learn to Simulate for Dexterous Robot Control
DESCRIPTION:Abstract:\nSimulation enables robots to learn and evaluate behaviors at scale before real-world deployment. Yet the mismatch between simulation and the physical world remains a fundamental obstacle. This is particularly challenging for dexterous manipulation\, where contact-rich interactions and dynamics variations across objects and robot embodiments are difficult to model. In my thesis research\, I explore how robot learning can scale through simulation and how learned models can make simulation more accurate to the physical world\, through two complementary directions. \nPart I: Differentiable simulation for scalable robot learning. \nFirst\, I present a GPU-parallel differentiable multiphysics simulation and a first-order reinforcement learning algorithm that pairs simulation gradients with entropy regularization\, for smoother policy optimization on locomotion and manipulation tasks. Next\, I introduce hybrid analytic differentiability\, combining implicit differentiation\, auto-differentiation\, and custom analytic Jacobians to compute gradients through contact without modifying forward dynamics. With it\, I develop a production-ready differentiable simulation and show how its gradients support initial value problems\, trajectory optimization\, and system identification. \nPart II: Aligning simulation with the real world across diverse embodiments and tasks. \nFirst\, I introduce an algorithm for iterative real-to-sim alignment. Alongside\, I present a hybrid neural dynamics model that combines learned dynamics correction with analytical inverse dynamics while retaining the simulator’s contact resolution\, to produce physically consistent simulation trajectories. Next\, I build flexible real-time robot I/O infrastructure for synchronized data collection and policy deployment across different robots\, sensors\, and interfaces. \nIn my proposed work\, I will explore how differentiable simulation and differentiable rendering can support real-to-sim reconstruction of simulation environments from multimodal real-world data. In my final project\, I will study how neural dynamics can scale real-to-sim-to-real learning to dexterous hands and humanoid robots requiring high-dimensional continuous control. \n\n\nTogether\, these directions aim to establish a feedback loop in which robot policies and simulators continually improve one another. By turning physical experience into better simulators and using better simulators to train more capable robots\, this loop could scale robot learning across tasks and embodiments in ways that neither simulation nor real-world data can achieve alone.\n\n\nThesis Committee:\nJean Oh (co-chair)\nGuanya Shi (co-chair)\nJeff Ichnowski\nMiles Macklin (NVIDIA)\n\nThesis Link
URL:https://www.ri.cmu.edu/event/simulate-to-learn-learn-to-simulate-for-dexterous-robot-control/
LOCATION:Gates Hillman Center 4405
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260616T100000
DTEND;TZID=America/New_York:20260616T113000
DTSTAMP:20261009T020012
CREATED:20260609T175151Z
LAST-MODIFIED:20260609T175151Z
UID:151545-1781604000-1781609400@www.ri.cmu.edu
SUMMARY:Aligning Observations Across Viewpoint\, Time\, and Embodiment for Agricultural Perception and Manipulation
DESCRIPTION:Abstract:\n\nAgricultural specialists are actively turning to robotic and computer vision-based systems to reduce the manual labor required to inspect and manipulate crops. These tasks require robots to perceive and interact with plants from partial\, localized observations\, often in dense and cluttered environments. For perception\, a central challenge is that crops are small\, are easily occluded\, and may change in appearance and position over time. For manipulation\, the ability to learn visuomotor policies is limited by the lack of available datasets and the difficulty of collecting robot demonstrations in the field. This thesis addresses these challenges by aligning and associating partial observations across viewpoint and time for agricultural perception\, and across viewpoint and embodiment for learning wrist-camera manipulation policies from human demonstrations. \nIn the first part of this thesis\, we develop perception-based methods for visually inspecting small crops in agriculture from limited observations. We present a 3D reconstruction pipeline for non-destructive seed counting of sorghum panicles\, a next-best-view planning approach for autonomously imaging and sizing apple fruitlets\, and a transformer-based method for spatio-temporally associating apple fruitlets across days and viewpoints. \nThe second part of this thesis shifts towards robot manipulation and learning from human demonstrations. We present a method that transforms monocular egocentric human demonstrations into wrist-camera observations and robot actions for training visuomotor policies\, without requiring depth sensors\, multi-view camera setups\, or custom data collection hardware. Building on this work\, we propose to align egocentric and wrist-camera observations and actions in latent space\, reducing reliance on explicit object tracking and image-space rendering. We further propose to incorporate visuo-tactile sensing for grape cluster inspection and harvesting. Together\, these efforts investigate how aligning observations can support agricultural robots that reason from limited visual information and learn manipulation policies when robot data is difficult to collect. \n\nThesis Committee Members:\nGeorge Kantor (Chair)\nDavid Held\nJeffrey Ichnowski\nSoumik Sarkar (Iowa State University)\n \nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/aligning-observations-across-viewpoint-time-and-embodiment-for-agricultural-perception-and-manipulation/
LOCATION:1305 Newell Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260608T123000
DTEND;TZID=America/New_York:20260608T140000
DTSTAMP:20261009T020012
CREATED:20260602T163009Z
LAST-MODIFIED:20260602T163009Z
UID:151481-1780921800-1780927200@www.ri.cmu.edu
SUMMARY:Design and Evaluation of Low-Cost\, Open-Source Haptic Interfaces for Diverse Learning Applications
DESCRIPTION:Abstract: Touch is a powerful yet underused channel for learning. Prior research shows that haptic interaction can support both sensorimotor skill acquisition and the understanding of abstract concepts by grounding learning in bodily experience. However\, most haptic devices remain expensive\, technically complex\, and difficult to reproduce\, which keeps them largely confined to specialized laboratories. This limits their use in education and rehabilitation and has slowed progress toward scalable\, low-cost\, open-source solutions\, as well as toward a systematic understanding of how affordable haptic devices should be designed to reliably produce learning benefits. As a result\, the broader learning potential of haptics remains underexplored\, especially across diverse domains and beyond measures of immediate task success.\nThis thesis examines the design and evaluation of haptic systems for learning across three distinct domains. The first system\, HaptiClay\, explores how haptics and gesture can support mathematics learning by helping students construct concrete representations of terms in polynomial functions. The thesis traces the iterative design of the device and reports interventions with students that use haptics to encourage gestural movements while molding polynomial functions and relate those gestures to specific terms in the polynomials. We then analyze learning outcomes to understand the effectiveness of the intervention. The second system\, DexKit\, enables students to experience dexterity concepts in dexterous teleoperation through touch\, including robotic manipulation control\, object interaction\, and stiffness variation. It introduces a dexterous manipulation platform that combines a soft robotic hand with a three-finger haptic interface\, including a novel two-degree-of-freedom mechanism for the index and middle fingers and a soft delta mechanism for the thumb. The third system\, VibroGait\, is a wearable haptic device for gait correction that helps users learn improved walking patterns through vibrotactile feedback. The thesis presents the design of a flexible skin-interfacing device\, the gait prediction algorithms and their implementation\, and studies comparing multiple haptic feedback patterns for gait correction. \nAcross these case studies\, the thesis investigates how effective\, low-cost learning tools can be designed\, which design principles generalize across domains\, how haptics influence learning beyond task success\, and how haptic systems for learning can be evaluated rigorously. By bringing together mathematics learning\, robotic teleoperation\, and gait correction\, this work expands the evidence base for accessible haptic learning technologies and contributes practical design knowledge for future low-cost\, open-source haptic systems. \n\nCommittee\nMelisa Orta Martinez (chair)\nJames McCann\nEni Halilaj\nKylie Peppler (University of California\, Irvine)\n\n\nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/design-and-evaluation-of-low-cost-open-source-haptic-interfaces-for-diverse-learning-applications/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260521T153000
DTEND;TZID=America/New_York:20260521T170000
DTSTAMP:20261009T020012
CREATED:20260511T201231Z
LAST-MODIFIED:20260511T201231Z
UID:151239-1779377400-1779382800@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Anurag Ghosh
DESCRIPTION:Date: May 21st\, 2026\nTime: 3:30 – 5:00 pm\nRoom: NSH Room 4305\nZoom:  https://cmu.zoom.us/j/98318417145 \nType: RI PhD Thesis Proposal\nWho: Anurag Ghosh\n\n\nTitle: Scaling Long-Tailed Driving Perception and Planning with In-the-Wild Videos\n\n\nAbstract: Closed-loop driving\, where methods produce actions a simulator reacts to\, remains largely tied to driving logs from instrumented fleets. Thus\, reliably driving in rare-but-critical scenarios is still elusive. Meanwhile\, foundation models are increasingly common in autonomous driving\, vision-language and video world models are tackling open-loop tasks like scene description and video generation. Although the internet has revolutionized language and image generation\, planning in autonomous driving has not seen its ImageNet moment yet. Therefore\, an opportunity exists to leverage internet-scale data and tackle long-tailed autonomous driving.\n\nWe focus on work zones as a representative long-tail scenario as they are a major source of disengagements for commercial systems today. Work zones uniquely combine rare objects (e.g.\, construction vehicles\, arrow boards)\, unusual layouts (e.g.\, temporary closures\, crossing yellow lines)\, and unpredictable behaviors (e.g.\, flaggers\, sudden merges). These safety-critical scenarios are considered difficult to simulate at scale. \nFirst\, we mine long-tailed driving videos from a massive corpus. We find that foundation models fail at work zone perception and fine-tuning on our data combined with simple priors makes them effective. Second\, we introduce a resource-efficient\, geometry-based prior that improves scene perception and long-tail object detection. Third\, we focus on long-tailed closed-loop planning and develop an anytime language-action planner capable of real-time trajectory generation and contextual textual reasoning. Furthermore\, by developing a novel rules-based planner that effectively handles current benchmark scenarios\, we show that existing closed-loop driving benchmarks are insufficient for evaluating long-tailed behaviors. \nFinally\, this thesis proposes a framework that\, by carefully composing geometry-aware methods\, street-view imagery\, and foundation models\, lifts monocular videos into metric\, geo-referenced 4D driving logs compatible with existing simulators. Using this framework\, we create a new long-tail planning benchmark and propose to uncover insights and study the geographic scaling behavior of state-of-the-art planning methods. \nUltimately\, to advance autonomous driving beyond fleets\, we argue scaling of training and evaluation is achievable by harnessing internet-scale data while grounding foundation models with geometric and physical priors.\n\nCommittee:\nSrinivasa Narasimhan\, Chair\nDeva Ramanan\nMaxim Likhachev\nChristoph Mertz\nManmohan Chandraker\, UC San Diego\n\nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-anurag-ghosh/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260520T150000
DTEND;TZID=America/New_York:20260520T163000
DTSTAMP:20261009T020012
CREATED:20260512T145730Z
LAST-MODIFIED:20260512T145730Z
UID:151245-1779289200-1779294600@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Renos Zabounidis
DESCRIPTION:Date: May 20th\, 2026\nTime: 3:00 – 4:30 pm\nLocation: NSH 4305\nZoom Link \nType: RI PhD Thesis Proposal\nWho: Renos Zabounidis \nTitle: Enforcing Neuro-Symbolic Structure in Deep Reinforcement Learning \nAbstract: Monolithic deep reinforcement learning trains a single network to learn vision\, physics\, planning\, and control from reward alone. The result is poor sample efficiency\, brittle generalization\, and uninterpretable decisions. This thesis shows how to build domain knowledge into policy architecture and enforce these architectural priors during training. \nWe develop this claim at three levels of abstractions. Concept-level abstractions route predictions through human-interpretable predicates such as `door_present’ and `key_in_inventory’\, enabling runtime inspection and intervention. Action-level constraints enforce state-dependent action validity\, preventing unmasked training from suppressing rarely valid behaviors through shared representations. Compositional abstractions represent long-horizon tasks as reusable symbolic skills that a planner can sequence while RL grounds each skill in low-level control. \nBuilding on these foundations\, this proposal focuses on two future directions. Concept-conditioned latent action models impose semantic structure on variational motor representations\, allowing high-level controllers to sample behaviors by name. A planning-guided option critic learns dynamic skill scheduling under precondition constraints\, replacing static plan traversal with on-policy option selection. \nTogether\, these contributions show that domain knowledge in the policy architecture reduces sample complexity\, enables cross-task skill composition\, and makes internal decisions available for inspection and override. \nThesis Committee:\nKatia Sycara\, CMU (Chair)\nSebastian Scherer\, CMU\nYonatan Bisk\, CMU\nKevin Ellis\, Cornell University \n\nThesis URL
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-renos-zabounidis/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260518T100000
DTEND;TZID=America/New_York:20260518T113000
DTSTAMP:20261009T020012
CREATED:20260508T181152Z
LAST-MODIFIED:20260508T181152Z
UID:151234-1779098400-1779103800@www.ri.cmu.edu
SUMMARY:Deep Abstraction Learning for Neuro-Symbolic World Modeling
DESCRIPTION:Abstract: Modern foundation models have achieved remarkable progress by learning broad physical and semantic common sense from large-scale data. However\, robots operating in open-ended environments require more than general knowledge alone: they must continually specialize in new tasks\, environments\, and experiences encountered during deployment. Given only limited deployment-time data\, how can robots learn to solve substantially more—and conceptually harder—problems than those seen during training? \nThis thesis addresses this challenge through deep abstraction learning\, where robots discover relational abstractions grounded by deep neural networks to construct neuro-symbolic world models from high-dimensional and noisy observations. By ignoring task-irrelevant details\, abstractions can be learned efficiently from limited experience while enabling abstract planning for long-horizon decision-making problems involving many interacting objects. The central hypothesis is that abstractions and world models should co-evolve: abstractions enable efficient reasoning and planning\, while planning structure and execution failures drive the discovery of richer abstractions. \nThis framework consists of three key components. Deep state abstractions\, represented as relational predicates\, map high-dimensional observations into symbolic concepts that support reasoning and planning under noisy sensory inputs. Deep action abstractions\, represented as relational option policies\, capture reusable behaviors that enable robots to recover from failures and progressively acquire new abstractions from interaction. Neuro-symbolic world models describe how action abstractions transform state abstractions\, enabling abstract planning that generalizes to unseen long-horizon tasks. To study these challenges\, this thesis introduces benchmark suites that reveal the limitations of purely neural approaches and motivate abstraction-based world modeling. \nBuilding on these foundations\, this dissertation proposes two future directions. The first studies how pre-trained coding agents can synthesize expressive programmatic world models that support planning with recomposable tools such as abstractions and perception models. The second\, RoboSymphony\, studies multi-agent decision-making in which robots learn abstractions over the intentions and behaviors of other agents and humans\, enabling coordination in collaborative tasks. \nTogether\, these contributions advance a unified neuro-symbolic approach for robots that learn and plan with deep abstractions\, enabling efficient specialization and decision-making in open-ended real-world environments. \nThesis Committee: \nSebastian Scherer (Chair)\nMaxim Likhachev\nKatia Sycara\nTom Silver\, Princeton University\nLeslie P. Kaelbling\, Massachusetts Institute of Technology\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/deep-abstraction-learning-for-neuro-symbolic-world-modeling/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260511T083000
DTEND;TZID=America/New_York:20260511T100000
DTSTAMP:20261009T020012
CREATED:20260409T174428Z
LAST-MODIFIED:20260503T203935Z
UID:150953-1778488200-1778493600@www.ri.cmu.edu
SUMMARY:Think Globally\, Solve Locally: Non-sequential Planning for Robotic Manipulation
DESCRIPTION:Abstract:\n\nRobotic manipulation requires reasoning that bridges local competence over fine-grained dynamics with the  construction of valid long-horizon plans. This mirrors human reasoning\, where fast\, automatic processes propose locally plausible actions from experience\, while slower deliberation integrates them into a whole. Furthermore\, evidence from cognitive science suggests that humans do not reason sequentially from start to finish; instead\, they anchor on intermediate landmarks\, plan from multiple directions\, and let local insights reshape global structure.\n\nYet\, current approaches in robotics typically formulate planning as a monolithic\, unidirectional search from an initial state toward a goal or in the opposite direction. This thesis argues for a departure from that paradigm. We propose Non-Sequential Deliberative Planning\, a framework that distributes deliberation across multiple local regions in the problem space\, rather than committing to a systematic\, directed reasoning. By exploring simultaneously from these regions\, local solvers (whether geometric\, heuristic\, or generative) operate where they are most effective\, while a global search composes them into a coherent plan with formal guarantees. \nWe instantiate this principle across three algorithmic regimes. For high-dimensional motion planning\, we introduce Multi-Graph Search (MGS)\, which identifies key states as intermediate landmarks and simultaneously grows search trees from each\, merging local subgraphs into a global solution with provable completeness and bounded suboptimality guarantees.\nSecond\, for contact-rich manipulation\, we present MOSAIC\, which treats physics-validated skills as local competences. It composes sequences of skills\, such as pushing or grasping\, through a non-sequential search that connects local regions of reliable execution. Third\, for scenarios where deliberation time is severely limited\, we develop methodologies that shift non-sequential reasoning to an offline phase. By integrating manipulation behaviors directly into preprocessing\, we generate motions whose manipulation outcomes are provably reliable\, and the deliberation time is guaranteed to be within a user-defined time bound.\nTo complete this thesis\, we consider three extensions. First\, MOSAIC relies on physics simulation to estimate the outcome of contact-rich interactions during online planning—a significant computational bottleneck. We propose to address this by utilizing an offline phase to learn proxies and construct data structures that enable efficient online planning over long horizons. Second\, manipulation skills are typically designed and learned for interactions between a robot and a single object. Real-world deployment\, however\, brings scenes with many movable objects\, and tracking them jointly causes a combinatorial explosion in the planning state space. To address this\, we propose to extend our prior framework with partial state planning\, in which the global search operates over decoupled\, object-centric representations and evaluates multi-object interactions only when necessary. Third\, drawing on our prior work in multi-robot coordination—Experience-Accelerated Multi-Robot Planning (xECBS) and Multi-Robot Multi-Model Diffusion (MMD)—we plan to extend the non-sequential planning architecture to multi-arm settings\, enabling concurrent execution across multiple manipulators. \nUltimately\, this thesis establishes a unified framework for non-sequential decision making\, composing fast local competences into globally sound manipulation plans across a range of real-world settings. \n\n\nThesis Committee:\n\n\nProf. Maxim Likhachev (Chair)\nProf. Changliu Liu\nProf. David Held\nProf. Oren Salzman (Technion University)\n\n\n\nThesis proposal document draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-itamar-mishani/
LOCATION:Newell-Simon Hall 1305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260504T163000
DTEND;TZID=America/New_York:20260504T180000
DTSTAMP:20261009T020012
CREATED:20260427T175100Z
LAST-MODIFIED:20260427T175100Z
UID:151125-1777912200-1777917600@www.ri.cmu.edu
SUMMARY:Forecasting at Scale with Efficient Deep Learning Architectures
DESCRIPTION:Abstract:\nTime Series Foundation Models (TSFMs) have scaled rapidly\, with publicly reported pretraining corpora growing from 1.23 billion to 1 trillion data points between 2024 and 2026\, an approximately 800× increase in two years. Recent work has further supplemented real-world data with synthetic data to expose models to broader time series patterns. Yet\, this data-centric paradigm raises a fundamental question: must intelligent forecasting rely solely on scale\, or can intentional architectural design unlock better generalization? This thesis proposes that more intelligently and efficiently leveraging existing data\, rather than scale alone\, is key to achieving better forecasting generalization. We pursue this through three parallel architectural themes: exploiting cross-channel structure beyond temporal patterns\, enabling zero-shot generalization through structured composition\, and reducing gradient and forecast variance by design. Each theme aims to enhance generalization with available data while treating computational efficiency as a core design principle. \nIn this thesis\, we demonstrate that scale is not the only path to generalization by: developing multivariate architectures that leverage cross-channel dependencies efficiently while reducing forecast error; showing that architectures can generalize beyond their training distribution in both patterns and concepts; and verifying variance-aware architectural designs that extract richer training signals from existing data\, provably reducing gradient variance while reducing forecast error and improving calibration. \nWithin the first theme\, we further propose pretraining strategies for multivariate TSFMs to investigate whether data balancing and curriculum learning can improve downstream generalization given the same pretraining corpora. Within the second theme\, we propose an additional dimension of generalization\, extending beyond pattern and concept generalization to horizon generalization\, an important consideration for TSFMs applied across diverse tasks and domains. Overall\, this work contributes new insights into advancing time series forecasting generalization through efficient architectural design. \n\n\nCommittee:\nArtur Dubrawski\, Chair\nJohn Dolan\nBarnabás Póczos\nMichael W. Mahoney (University of California\, Berkeley)\n\n\nThesis Link
URL:https://www.ri.cmu.edu/event/forecasting-at-scale-with-efficient-deep-learning-architectures/
LOCATION:GHC 4405
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260504T133000
DTEND;TZID=America/New_York:20260504T150000
DTSTAMP:20261009T020012
CREATED:20260427T201525Z
LAST-MODIFIED:20260427T201525Z
UID:151127-1777901400-1777906800@www.ri.cmu.edu
SUMMARY:Leveraging Local Models for Planning and Control with Contact
DESCRIPTION:Abstract: Many planning and control approaches in robotics have converged on optimization-based formulations\, with recent advances achieved by leveraging significant data and compute to attempt to tackle these nonlinear and non-convex problems. In this thesis\, we instead focus on local models and demonstrate their benefits and surprising effectiveness. In the case of smooth optimization\, the local model is a convex quadratic program. We show how this structure enables efficient parallelization and scaling by mapping it to a neural network\, and that a fixed linear model is still capable of basic locomotion tasks even with a large sim-to-real gap. We then look at the non-smooth case that arises in contact-implicit approaches which can be expressed as quadratic programs with complementarity constraints and develop a C++ solver\, Marble\, that leverages the structure of relaxed complementarity. Finally\, we propose future work that unifies existing hard and soft contact models under a common framework and examines them in the context of planning. We also propose applying the resulting Marble solver for local motion retargeting tasks\, exploring applications in both simulation and on hardware in an iterative learning control context. \nThesis Committee: \nZac Manchester (chair)\nAaron Johnson\nLorenz Biegler\nPat Wensing (University of Notre Dame)\n\nThesis URL
URL:https://www.ri.cmu.edu/event/leveraging-local-models-for-planning-and-control-with-contact/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260430T160000
DTEND;TZID=America/New_York:20260430T173000
DTSTAMP:20261009T020012
CREATED:20260421T193956Z
LAST-MODIFIED:20260422T143815Z
UID:151107-1777564800-1777570200@www.ri.cmu.edu
SUMMARY:Longitudinal Human–Robot Interaction: Adaptive Personalization Across Repeated Encounters
DESCRIPTION:Abstract: \nAs robots increasingly move into homes\, healthcare settings\, and public environments\, many are expected to support people not through single encounters\, but through repeated interaction over time. In these settings\, successful human–robot interaction depends not only on immediate task performance\, but also on how users adapt to robotic systems\, how expectations change with repeated exposure\, and how interaction preferences evolve across sessions and contexts. Despite growing interest in personalization\, many robotic systems still assume that user preferences can be estimated once and treated as stable\, with limited understanding of how preferences develop longitudinally.\n\nThis thesis investigates longitudinal human–robot interaction by examining how user preferences\, engagement\, and interaction strategies change through repeated encounters with socially interactive robots. In robotic exercise support for older adults\, an exploratory Wizard-of-Oz study revealed substantial variation in how participants naturally engaged with a conversational exercise robot\, ranging from brief task-focused exchanges to extended social interaction. Building on these findings\, a four-week longitudinal study compared two contrasting robot personalities during repeated exercise sessions: a Social Buddy Personality emphasizing companionship and conversation\, and an Exercise Coach Personality emphasizing structured feedback and task-focused guidance. Participants responded positively to both personalities but valued different aspects of each\, with preferences varying across individuals and shifting across sessions.\n\nTo examine whether similar temporal dynamics extend beyond exercise\, this thesis also investigates repeated interaction in accessibility robotics through a longitudinal navigation study with blind and low-vision users. Across multiple weeks of navigation in public environments\, participants demonstrated evolving preferences for delegation and autonomy\, with assistance strategies changing according to context\, familiarity\, and accumulated experience. Together\, these studies show that user preferences in sustained human–robot interaction are not static\, but develop through repeated exposure and situational adaptation.\n\nMotivated by these findings\, this thesis proposes a multi-timescale adaptive robotic exercise coaching framework that dynamically adjusts social interaction and coaching behavior using multimodal estimates of user state\, behavior\, and engagement. By integrating exercise performance\, conversational behavior\, and interaction history\, the proposed system models preference across long-term trends\, session-level variation\, and moment-to-moment interaction. Overall\, this work contributes new insight into how temporally aware personalization can support long-term human–robot interaction in health\, accessibility\, and everyday assistive contexts.\n\n\nThesis Committee:\nAaron Steinfeld\, chair \nReid Simmons\, CMU\nHenny Admoni\, CMU\nMaja Matarić\, USC\nTaskin Padir\, Amazon Robotics\, Northeastern University\n\n\nThe link to the document can be found here: thesis_proposal 
URL:https://www.ri.cmu.edu/event/longitudinal-human-robot-interaction-adaptive-personalization-across-repeated-encounters/
LOCATION:1305 Newell Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260428T120000
DTEND;TZID=America/New_York:20260428T133000
DTSTAMP:20261009T020012
CREATED:20260325T185125Z
LAST-MODIFIED:20260421T154725Z
UID:150738-1777377600-1777383000@www.ri.cmu.edu
SUMMARY:Evolutionary Environment Optimization for Large-Scale Multi-Robot Systems
DESCRIPTION:Abstract:\n\n\nRecent advances in robotics have enabled researchers and practitioners to deploy large-scale multi-robot systems\, where hundreds to thousands of robots operate simultaneously in a shared environment. These systems are increasingly used in applications that require high efficiency and reliability\, such as automated warehouses\, robotic sorting systems\, and autonomous transportation. A fundamental challenge in such systems is to enable many robots to move efficiently while sharing limited space\, avoiding congestion\, and continually finishing new tasks. \nWhile the community has spent a significant amount of attention on studying how to more efficiently and effectively coordinate robots in multi-robot systems\, very few works focus on the environment where the systems are deployed. By environment\, we refer to both physical elements\, such as spatial arrangement of storage shelves in automated warehouses\, and virtual elements\, such as traffic rules that influence how robots move and interact. A well-designed environment can significantly reduce congestion\, foster implicit coordination among robots\, and improve system throughput. In this thesis\, we study the problem of Environment Optimization for large-scale multi-robot systems. We identify three main components of the environment that we find optimizable. We formally define each of them as black-box optimization problems and apply evolutionary-based optimizers to solve them. \nWe first discuss Task Mapping Optimization in the context of robotic sorting systems (RSS)\, where robots transport packages from induct workstations to eject chutes according to shipping destinations. The destination-to-chute assignment directly shapes traffic demand: poor assignments overload certain regions\, create conflicts\, and reduce throughput. To fill this gap\, we formally define the problem of Task Mapping Optimization (TMO) and propose methods to search for high-quality task mappings that improve throughput by balancing robot traffic. \nWe then discuss Layout Optimization in the context of automated warehouses\, where robots transport packages or inventory pods between locations. In these systems\, layout strongly affects robot coordination: narrow passages create bottlenecks and poorly placed shelves induce congestion\, even when state-of-the-art planning algorithms are applied. Existing warehouse layouts often follow grid-like structures designed for human accessibility\, but such patterns are not necessarily suitable for robots\, which do not share human ergonomic constraints. This motivates the question of how a warehouse should be designed when robot coordination is the primary objective. Inspired by scenario generation techniques\, we formally define the problem of Layout Optimization (LO) and propose methods to optimize layouts for maximum throughput. We then discuss our first proposed work: Multi-Objective Warehouse Layout Optimization for Throughput and Storage Capacity\, which addresses the trade-off between denser storage and efficient robot movement. We search for Pareto-optimal layouts that balance storage capacity and throughput. \nNext\, we investigate Guidance Graph Optimization in the context of multi-robot coordination in general. When robots are continuously receiving new tasks\, requiring them to move indefinitely\, existing planning algorithms inevitably make short-horizon decisions. In particular\, individually efficient paths may still create long-term congestion. To mitigate this limitation\, we introduce global movement guidance that encourages long-term cooperation among robots. We formally define the guidance graph as a representation of such guidance and the problem of Guidance Graph Optimization (GGO)\, and discuss several methods to optimize guidance graphs for throughput. However\, these methods suffer from the curse of dimensionality\, making existing GGO methods scale poorly to large graphs. Therefore\, we then discuss our second proposed work: Guidance Graph Optimization for Ultra Large Graphs\, which scales GGO methods to large graphs. \nFinally\, we address a common limitation shared by all environment optimization methods: their reliance on simplified simulation. Because realistic simulation is computationally expensive while evolutionary optimization requires many evaluations\, directly scaling these methods to realistic settings is difficult. To improve sample efficiency\, we discuss our third proposed work: Deep Surrogate Assisted Realistic Environment Optimization\, where we investigate data-driven surrogate models that approximate simulator outcomes and enable environment optimization under more realistic conditions. \nTogether\, these contributions demonstrate that optimizing physical layouts\, task assignments\, and virtual guidance can substantially improve coordination in large-scale multi-robot systems\, offering a unified perspective on environment design for scalable robotic operation. \n\n\n \n\nThesis Committee Members:\n\nJiaoyang Li (Chair)\,\nStephen Smith\,\nPeter Zhang\nStefanos Nokilaidis (University of Southern California)\n\n\nThesis Proposal Draft link: https://drive.google.com/file/d/1BwxjdPA0qOyiaOKhcuOUV4hNCcMgaTmq/view?usp=sharing
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-yulun-zhang/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260417T130000
DTEND;TZID=America/New_York:20260417T143000
DTSTAMP:20261009T020012
CREATED:20260410T192550Z
LAST-MODIFIED:20260410T192550Z
UID:150988-1776430800-1776436200@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Hanzhe Hu
DESCRIPTION:Date:  April 17\, 2026\nTime:  1 PM-2:30 PM\nLocation: NSH 3305\nZoom link\nType: RI PhD Thesis Proposal\nWho: Hanzhe Hu\n\nTitle: Learning to Create 3D Worlds via Multi-View Generation\n\n\nAbstract:\nA 3D world is a visual representation that can be rendered from any viewpoint at any moment in time. Creating such representations from minimal input — a single image\, a text prompt\, or a monocular video is a fundamental goal in computer vision and graphics. An emerging and promising alternative is multi-view generation. However\, multi-view generation introduces its own challenges: maintaining geometric consistency across views\, achieving practical inference speed\, and extending to dynamic scenes. This thesis addresses these three challenges and presents a path toward creating 3D worlds via multi-view generation. \nWe first present MVD-Fusion\, which tackles consistency by introducing depth-guided cross-view attention for multi-view RGB-D generation from a single image. Intermediate depth estimates enable reprojection-based feature aggregation\, enforcing geometric consistency and yielding direct 3D reconstruction without costly optimization. We then address efficiency with Turbo3D\, which generates 3D Gaussian Splatting assets from text in under one second. A dual-teacher distillation framework compresses a multi-step multi-view diffusion model into a 4-step generator\, while a latent-space reconstructor eliminates image decoding overhead. Finally\, we tackle dynamics with GeoVideo4D\, a framework for camera-controllable multi-view video generation that simultaneously produces synchronized RGB videos and aligned depth maps through a joint video diffusion process\, with a hybrid training strategy unifying static 3D\, monocular video\, and multi-view video data. \nLooking ahead\, we outline two directions. First\, unifying 3D reconstruction and generation by jointly training both tasks in a single model where cameras are learned in a self-supervised manner\, enabling training on large-scale unannotated data. Second\, extending video generation to long temporal horizons to support sustained\, coherent 3D world generation beyond the short clips produced by current methods. \n\nThesis Committee:\nShubham Tulsiani (chair)\,\nDeva Ramanan\,\nJun-Yan Zhu\,\nJiajun Wu (Stanford University)\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-hanzhe-hu/
LOCATION:NSH 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260414T093000
DTEND;TZID=America/New_York:20260414T113000
DTSTAMP:20261009T020012
CREATED:20260325T192351Z
LAST-MODIFIED:20260407T230724Z
UID:150745-1776159000-1776166200@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Nikhil Keetha
DESCRIPTION:Date: April 14\, 2026\nTime: 9:30 AM (ET)\nLocation: NSH 3305\nZoom Link\n\nType: RI Ph.D. Thesis Proposal \n\nWho: Nikhil Keetha\n\nTitle: Scaling Representation Learning for Spatial Intelligence via Active Implicit Memory \n \nAutonomous systems and embodied agents that operate in the physical world require spatial intelligence: the ability to perceive\, understand\, and reason about the geometric and semantic structure of scenes from visual input.\nData-driven approaches have made progress in learning generalizable representations from large corpora\, following the success of large language models; however\, two fundamental challenges remain open.\nFirst\, on the learning paradigm: self-supervised objectives such as next-frame prediction are insufficient for visual streams\, and nascent multi-view completion approaches remain fragile and prone to model collapse.\nSecond\, on the architecture: current multi-view models maintain scene representations that are either fixed-size (limiting capacity)\, linearly growing (hitting compute walls)\, or managed by hand-crafted heuristics that reintroduce classical brittleness.\nThis thesis addresses both challenges through a unified framework in which bootstrapped representation learning\, active implicit memory\, and multi-task decoding are tightly coupled and mutually reinforcing.The first half of this thesis establishes the foundations through three completed contributions with increasing scope.\nAnyLoc demonstrates that foundation model features provide a universal substrate for visual place recognition across diverse environments without task-specific training\, establishing the power of data-driven representation learning.\nSplaTAM introduces covisibility-guided differentiable rendering for dense visual SLAM\, providing the memory management concepts and self-supervised rendering signal that the proposed work extends.\nMapAnything presents a unified feed-forward transformer that decodes twelve sub-tasks from a single factored representation at internet scale\, validating that task and modeling paradigm are orthogonal. \nThe second half proposes Memor\, the core contribution of this thesis.\nMemor introduces active implicit memory whose capacity adapts to the complexity of the input stream\, and unified multi-task decoding that recovers diverse spatial outputs from the resulting shared representation.\nMemory management is learned end-to-end rather than governed by hand-crafted rules\, and the diversity of decoded tasks jointly enriches the underlying representation.\nTwo application chapters extend Memor: one to holistic scene generation\, where supervised pre-training prevents the model collapse that limits purely self-supervised approaches; the other to open-world semantic exploration\, where the memory provides spatial context for language-guided navigation. \nTogether\, these contributions advance a unified approach to scaling representation learning for a wide breadth of spatial intelligence.\n\nThesis Committee Members:\nSebastian Scherer (Co-Chair)\nDeva Ramanan (Co-Chair)\nShubham Tulsiani\nPeter Kontschieder\, Meta Reality Labs\n \n\n\n\n \nLink to Thesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-nikhil-keetha/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260413T151500
DTEND;TZID=America/New_York:20260413T170000
DTSTAMP:20261009T020012
CREATED:20260407T184256Z
LAST-MODIFIED:20260407T184256Z
UID:150900-1776093300-1776099600@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Xinyu (Rachel) Li
DESCRIPTION:Date: April 13\, 2026\nTime: 03:15 PM (ET)\nLocation: GHC 6121\nZoom Link\n\nType: Ph.D. Thesis Proposal \n\nWho: Xinyu (Rachel) Li\n\nTitle: Towards Accessible AI Agents\n \nAbstract:\nEmpowered by large language models (LLMs)\, AI agents have shown strong potential across tasks such as general-purpose assistance\, software coding\, and scientific research. However\, their practical utility in applications involving consequential decisions such as healthcare\, remains constrained by three major challenges. \nEvaluation. Existing agent evaluations often focus on well-structured tasks and final outcomes\, failing to fully capture the complexity of real-world workflows. We propose evaluation frameworks grounded in realistic machine learning engineering workflows\, providing skill-based\, multi-artifact\, and holistic assessments that systematically evaluate the practical utility of AI agents. \nLearning. Improving LLMs for agentic use typically relies on reinforcement learning with large amounts of high-quality labeled data\, which are costly and difficult to obtain in expert domains including healthcare. To address this limitation\, we aim to develop learning frameworks that require minimal external supervision\, improving the scalability and efficiency of agent learning. \nSpecialization. AI agents typically follow a one-size-fits-all paradigm at the time of deployment\, lacking mechanisms to account for task-specific or user-specific requirements. We propose methods that enable agent specialization for downstream tasks and users\, expanding their applicability across heterogeneous deployment settings. \nThis thesis aims to make AI agents more broadly accessible and impactful in important real-world applications by enhancing their practical utility\, making them more measurable\, more capable\, and better tailored to the needs of their users and applications. \n\nLink to thesis\n\nThesis committee members:\nArtur Dubrawski (Chair)\nAndrea Bajcsy\nBarnabás Póczos\nDaniel McDuff (Google)
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-xinyu-rachel-li/
LOCATION:GHC 6121
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260413T090000
DTEND;TZID=America/New_York:20260413T110000
DTSTAMP:20261009T020012
CREATED:20260407T193258Z
LAST-MODIFIED:20260407T193258Z
UID:150904-1776070800-1776078000@www.ri.cmu.edu
SUMMARY:RI PhD Thesis Proposal - Gaurav Parmar
DESCRIPTION:Who: Gaurav Parmar\nDate: 13 April 2026\nTime: 9:00 a.m. (ET)\nLocation:  NSH 3305\nZoom Link: Link\nType: Ph.D. Thesis Proposal\n\n\nTitle: Efficient and Controllable Diffusion Models\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\nAbstract:  \nGenerative models have made rapid advancements in recent years\, and text-conditioned diffusion models have become the standard paradigm for image and video synthesis. \nHowever\, controlling the output purely through text is not the ideal medium for many practical applications. In my thesis research\, I have focused on making diffusion models more controllable and efficient beyond text-only interaction. To this end\, I explore two complementary directions. \nPart A: \nI study methods for generating more expressive and diverse outputs. \nIn my first project\, I explore how users can specify desired output images through visual prompts rather than text. \nThis enables compositional image generation where object identities and appearances are controlled through reference images. \nIn my second project\, I address the issue of redundant generations when the task involves generating a group of images from the same text prompt. \nPart B: \nI study efficient\, structure preserving translations with diffusion models. \nIn my first project\, I explore zero-shot image editing method that shows how well-trained text-to-image models can be repurposed for editing real images. \nHowever\, such zero-shot methods are slow and struggle when the base text-to-image model is not trained for the target domain. \nIn my next project\, I explore how we can fine-tune text-to-image models to perform image translation in both paired and unpaired settings. \nIn my final project\, I will show how we can extend this to video-to-video translation\, and the unique challenges involved. \n \n \nThesis Committee: \nJun-Yan Zhu (Co-Chair) \nSrinivasa Narasimhan (Co-Chair) \nShubham Tulsiani \nDaniel Cohen-Or (Tel Aviv University) \n\n\n\n\n\n\n\n\n\nThesis Draft
URL:https://www.ri.cmu.edu/event/ri-phd-thesis-proposal-gaurav-parmar/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260325T110000
DTEND;TZID=America/New_York:20260325T123000
DTSTAMP:20261009T020012
CREATED:20260313T151609Z
LAST-MODIFIED:20260313T151812Z
UID:150608-1774436400-1774441800@www.ri.cmu.edu
SUMMARY:Robust\, Reliable Robot Odometry and its Certification
DESCRIPTION:Abstract:\n\nRobot odometry is the backbone of nearly all modern autonomous systems including\, but not limited to\, unmanned aerial vehicles\, autonomous underwater vehicles\, and autonomous ground vehicles. Most downstream tasks such as path planning\, perception\, and control require accurate knowledge of the vehicle position and orientation at any given moment. While odometry is well-studied and has many potential solutions\, due to its high-importance to the rest of the autonomy stack\, any increase in robustness and accuracy will only further drive the reliability and stability of the given autonomous system.\n\nThere are a number of areas where current odometry methods can be improved in accuracy\, robustness\, or reliability. More specifically\, we find that poor sensor models or assumptions can often reduce reliability and accuracy. Another area of concern is sensor failure\, where odometry methods that are too tightly coupled often fail entirely with a single sensor outage. Finally\, another potential issue is when odometry failures do occur\, most downstream tasks are unaware\, which can lead to erroneous behavior.\n\n\nIn this work we present methods that overcome these challenges. Specifically\, they (1) seek to correct any modeling errors that may occur\, (2) are robust to sensor failure\, and (3) provide sub-optimality metrics for downstream tasks. We first present a method for fusion of wheel encoder measurements for off-road autonomous vehicles that provides piecewise-planar constraints for non-planar environments\, while estimating wheel slip\, wheel radii\, and wheel baseline all in real-time. Additionally\, we present a novel LiDAR odometry method\, with a frontend informed by our empirical evaluations\, and a backend that smooths over a window of prior states while providing map corrections in real-time. This results in more accurate estimates and heightened robustness.\n\nFinally\, we propose finding a fast “sufficient condition” certificate for these optimization-based odometry methods utilizing novel semidefinite programming techniques. While not a perfect catch-all for odometry failures\, it aims to detect when sub-optimality or degeneracies in state estimation may be occurring and pass this information to downstream tasks\, allowing for reactionary behavior.\n\n\nThesis Committee Members:\nMichael Kaess\, chair\nSebastian Scherer\nGeorge Kantor\nTim Barfoot\, University of Toronto\n \nCurrent Thesis Proposal Draft
URL:https://www.ri.cmu.edu/event/robust-reliable-robot-odometry-and-its-certification/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260219T183000
DTEND;TZID=America/New_York:20260219T200000
DTSTAMP:20261009T020012
CREATED:20260209T220741Z
LAST-MODIFIED:20260209T221052Z
UID:150335-1771525800-1771531200@www.ri.cmu.edu
SUMMARY:Learning Dynamic and Competitive Human Skills and Strategies for Animation and Robotics
DESCRIPTION:Abstract:\nHumanoid control in animation and robotics requires physically realistic motion as well as the ability to adapt\, coordinate actions over time\, and make decisions in response to changing environments and other agents. Human motion data provides a powerful source of prior knowledge for learning natural and stable movement\, but many existing approaches rely on reference motions in ways that are difficult to generalize beyond demonstrated scenarios. As a result\, it remains challenging for humanoid agents to reuse and compose skills over time\, limiting their ability to operate in interactive and competitive environments. \nThis thesis investigates how human motion references can be used more flexibly\, not as strict templates to be reproduced\, but as behavioral priors that shape how agents move while allowing adaptation to new tasks and conditions. Rather than focusing on the replication of specific motions\, the emphasis is placed on learning reusable structure from human behavior that supports robustness and generalization across diverse goals\, environments\, and interactions. Through this perspective\, humanoid agents can preserve natural movement while remaining responsive to novel objectives and disturbances. \nBuilding on this foundation\, the thesis extends beyond single-agent skill execution to study higher-level behavior in interactive and competitive settings. Domains such as sports highlight that strong motor skills alone are insufficient: successful performance also depends on selecting appropriate actions\, timing them effectively\, and coordinating behavior over longer time horizons in response to both the environment and other agents. Overall\, this work presents a unified view of humanoid control in which reference motion supports generalization rather than constraining behavior\, and enables adaptive and strategic interaction in both animation and robotics. \nThesis Committee:\nJessica Hodgins\, chair\nDeva Ramanan\nGuanya Shi\,\nXue Bin Peng\, Simon Fraser University\, NVIDIA\nTaku Komura\, The University of Hong Kong \nA draft of the thesis proposal document is available here.
URL:https://www.ri.cmu.edu/event/learning-dynamic-and-competitive-human-skills-and-strategies-for-animation-and-robotics/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260218T100000
DTEND;TZID=America/New_York:20260218T113000
DTSTAMP:20261009T020012
CREATED:20260212T181825Z
LAST-MODIFIED:20260212T181825Z
UID:150366-1771408800-1771414200@www.ri.cmu.edu
SUMMARY:Toward Aligned Vision Models
DESCRIPTION:Abstract:\nModern vision and vision–language models (VLMs) achieve remarkable perceptual performance\, yet their internal representations often misalign with human-understandable concepts\, clinical reasoning\, or the causal structure of data. Such misalignment limits trust\, generalization\, and safety – particularly in high-stakes domains such as medical imaging. This thesis proposes a comprehensive framework for model alignment\, developing methods that bring model representations closer to clinically meaningful\, causally grounded\, and task-adaptive concepts across diverse visual tasks. \nFirst\, I introduce biomarker-grounded alignment for lung ultrasound (LUS)\, where domain-informed interpretable biomarkers serve as anchors to structure deep model representations. I develop methods that disentangle anatomical\, morphological\, and artifact-level biomarkers and demonstrate that these aligned representations improve interpretability while matching or exceeding fully supervised baselines across diagnostic and severity scoring tasks. \nSecond\, I propose a causal feature selection framework based on Markov blanket discovery to identify minimal yet causally relevant feature subsets across medical and non-medical datasets. By uncovering natural experiments embedded in observational data\, the method reveals features that are inherently robust\, reduces spurious correlations\, and provides theoretical and empirical evidence for improved generalization and interpretability. \nThird\, I explore prompt-tuning–based alignment for object detection\, showing that positive and negative few-shot exemplars can be leveraged for iterative prompt optimization. This strategy steers models toward task-relevant visual concepts\, improves detector robustness under domain shifts\, and reveals interpretable activation patterns associated with object-level reasoning.\n\nCollectively\, these contributions establish a cohesive strategy for aligning vision models with human-understandable\, causally grounded\, and task-relevant representations\, advancing the development of reliable\, interpretable\, and generalizable systems suitable for real-world clinical and broader perceptual deployment. \nFinally\, I outline two future research directions to further strengthen alignment and its evaluation. The first explores gradient-based soft prompt tuning\, which learns continuous prompt embeddings through backpropagation to enable more stable and scalable prompt optimization. The second develops saliency mapping for VLMs by evaluating region-proposal masks (e.g.\, from SAM) based on their impact on downstream performance\, measured through VQA-style scoring\, enabling quantitative assessment of visual explanation quality and faithfulness. \n \nThesis Committee Members:\nJohn Galeotti\, Co-chair\nDeva Ramanan\, Co-chair\nZachary Lipton\nTrevor Darrell\, University of California\, Berkeley\n \nA draft of the thesis proposal document is available here.
URL:https://www.ri.cmu.edu/event/toward-aligned-vision-models/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260211T090000
DTEND;TZID=America/New_York:20260211T103000
DTSTAMP:20261009T020012
CREATED:20260202T194426Z
LAST-MODIFIED:20260202T194426Z
UID:150247-1770800400-1770805800@www.ri.cmu.edu
SUMMARY:Plan What You Can\, Learn What You Must: Interleaving Planning and Learning for Multi-Robot Manipulation
DESCRIPTION:Abstract:\nMulti-robot manipulation is becoming an inevitability of modern robotics. As hardware costs fall\, the barrier to deploying robot teams has shifted from economics to algorithmic capability. To fulfill their promise\, multi-robot systems must jointly reason about geometric coordination\, contact interactions\, task assignments\, and scene dynamics\, while adapting to variable team sizes and diverse robot embodiments. \nThe field has long pursued two parallel yet complementary approaches to robotic manipulation. Classical planning methods explicitly model system dynamics and interactions\, enabling scalability\, adaptability\, and strong formal guarantees—but confining applicability to well-understood systems. In contrast\, recent work abandons explicit modeling in favor of data-driven flexibility\, capturing interaction dynamics that are difficult to model but struggling to generalize when data is scarce or task distributions shift. Even though the strengths of one mirror the weaknesses of the other\, only few efforts have sought principled ways to combine them within a unified framework for multi-robot manipulation. \nThis thesis targets the boundary between explicit planning and learned policy synthesis\, developing algorithms that plan what can be modeled and learn what cannot. By exploiting structure and identifying where learning is tractable\, we develop methods for multi-robot manipulation that generalize across team sizes and robot embodiments while retaining the flexibility to learn hard-to-model components. \nWe begin at the structured end of the spectrum\, addressing scalable motion planning for multi-robot-arm systems. After introducing two algorithms for labeled settings–where each robot has an assigned goal–we turn to our first proposed work: an anonymous multi-arm motion planner that generates coordinated trajectories under interchangeable-goal formulations. This capability is essential when policies specify contact objectives without prescribing which robot should execute them. \nFrom this structured foundation\, we dive into the boundary between planning and learning.\nWe develop two methods that compose simple\, data-efficient policies through lightweight planning—one for coordination under implicit objectives and another for planar collaborative manipulation—and propose our second research body direction: extending these ideas to realistic 3D domains involving contact-rich manipulation of large rigid objects. This work determines when analytical models suffice\, when learning is necessary\, and seeks to invoke learned models where they perform well (i.e.\, are in-distribution)\, defining a systematic interface between explicit planning and learned interaction dynamics. \nTogether\, these contributions reveal how structure and learning can coexist: flexible and generalizable multi-robot manipulation emerges not from choosing between planning and learning\, but from combining them. \n\nThesis Committee:\n\n\nProf. Jiaoyang Li (co-chair)\nProf. Maxim Likhachev (co-chair)\nProf. Andrea Bajcsy\nProf. Yilun Du (Harvard University)\n\n\nThesis proposal document draft
URL:https://www.ri.cmu.edu/event/plan-what-you-can-learn-what-you-must-interleaving-planning-and-learning-for-multi-robot-manipulation/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260206T083000
DTEND;TZID=America/New_York:20260206T100000
DTSTAMP:20261009T020012
CREATED:20260121T162246Z
LAST-MODIFIED:20260129T191639Z
UID:150154-1770366600-1770372000@www.ri.cmu.edu
SUMMARY:Empirically Grounded LLM-based Virtual Patients for Psychotherapy Training: Design\, Modeling\, and Evaluation
DESCRIPTION:Abstract:\n\n\nThe need for mental health care continues to outpace the supply of trained psychotherapists\, while psychotherapy training remains constrained by limited supervision time and scarce opportunities for repeated\, feedback-rich practice in realistic scenarios. Simulation-based training can mitigate these constraints\, but actor-based standardized patients are costly and difficult to scale\, and many clinically challenging moments are ethically and logistically hard to practice with real clients. Large language models (LLMs) make interactive virtual patients increasingly feasible as a complement to conventional training; however\, early systems are often prompt-driven rather than empirically grounded\, exhibit limited psychologically meaningful state and longitudinal change\, and predominantly focus on one-on-one sessions—leaving multi-party modalities such as couples therapy under-supported. \nThis thesis advances an empirical and design-oriented framework for building more realistic and pedagogically effective virtual patients grounded in psychotherapy process theory and real clinical data. Chapter 2 develops scalable\, LLM-based measurement of therapist behaviors and client responses from large psychotherapy transcript corpora and uses structural equation modeling to estimate process-level dynamics linking therapist micro-skills and relational factors to subsequent client disclosure and emotional expression. Chapter 3 translates these empirically informed requirements into a multimodal\, multi-agent couples therapy simulator that represents stage-structured sessions and recurrent interaction cycles such as demand–withdraw\, enabling trainees to practice timing- and wording-sensitive interventions in high-conflict moments; an evaluation with licensed therapists examines realism and training relevance. Chapter 4 proposes an adaptive virtual patient architecture in which an LLM agent maintains latent psychological states that update in real time according to SEM-derived interpersonal dynamics conditioned on detected therapist behaviors\, together with a plan for evaluating psychological fidelity. \nBy integrating theory-grounded measurement\, empirical causal modeling\, and interactive system design\, this work lays a pathway for scalable psychotherapy training tools that make the consequences of therapist choices visible\, support deliberate practice\, and responsibly expand access to high-quality skills development. \n\n \nThesis Committee:\n\n\n\n\n\nHaiyi Zhu (Chair)\, HCII\, CMU\nSherry Wu\, HCII & LTI\, CMU\nAaron Steinfeld\, RI\, CMU\nHolly Swartz\, Psychiatry\, UPMC\n\n\n\n\n\n\n\n\n \nA draft of the thesis proposal is available here.
URL:https://www.ri.cmu.edu/event/empirically-grounded-llm-based-virtual-patients-for-psychotherapy-training-design-modeling-and-evaluation/
LOCATION:Gates Hillman Center 6115
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260115T083000
DTEND;TZID=America/New_York:20260115T100000
DTSTAMP:20261009T020012
CREATED:20260108T145112Z
LAST-MODIFIED:20260108T145112Z
UID:149897-1768465800-1768471200@www.ri.cmu.edu
SUMMARY:Toward Scalable Architectures for Multimodal LLM-based Cooperative Autonomous Driving
DESCRIPTION:Abstract: Despite the tremendous progress made in autonomous driving over the years\, the safety of autonomous vehicles still requires further improvement before they can operate worldwide with full human trust. One principal safety concern is that each individual vehicle may have a limited field of view due to finite detection ranges\, potential sensor failures\, or occlusions caused by nearby large objects such as buses or trucks. This limitation in perception introduces additional challenges for the downstream planning and control modules\, making it more difficult for autonomous vehicles to generate safe driving decisions and actions.\nTo address this issue\, recent research has proposed vehicle-to-vehicle (V2V) and vehicle-to-everything (V2X) cooperative perception for autonomous driving. In such systems\, connected autonomous vehicles (CAVs) share their individual perception features with one another to improve overall cooperative detection accuracy. However\, most existing work focuses solely on the cooperative detection task\, without leveraging temporal information about the dynamic environment or considering other critical components of autonomous driving\, such as prediction and planning. \nTo broaden the scope of cooperative driving research\, my proposed doctoral research aims to explore multimodal large language model (LLM)–based cooperative autonomous driving\, motivated by several potential advantages of LLMs. First\, a single LLM-based model offers the flexibility to perform multiple tasks\, including perception\, prediction\, and planning\, within a unified framework. Second\, LLMs exhibit strong generalizability due to large-scale pretraining on diverse data. Third\, LLM-based driving models possess reasoning capabilities that enable them to handle long-tail driving scenarios that may not appear in the training data. Fourth\, natural language can serve as an effective and efficient communication interface for V2V\, V2X\, and human–vehicle interactions. \nWe have developed multimodal LLM-based cooperative autonomous driving architectures that enable end-to-end cooperative driving and generate suggested future trajectories for all CAVs through V2V communication. In addition\, we have designed a graph-of-thoughts reasoning framework to further enhance the reliability and interpretability of our multimodal LLM-based architecture. Finally\, we propose to develop a decentralized V2V framework using multimodal LLMs to improve the scalability and feasibility of future large-scale deployment. \n\n \n \nThesis Committee:\n\nStephen F. Smith (Chair)\nJohn Dolan\nDeva Ramanan\nMin-Hung Chen (NVIDIA)\n\n\n\n\n\nThesis proposal link
URL:https://www.ri.cmu.edu/event/toward-scalable-architectures-for-multimodal-llm-based-cooperative-autonomous-driving/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251119T130000
DTEND;TZID=America/New_York:20251119T140000
DTSTAMP:20261009T020012
CREATED:20251114T143444Z
LAST-MODIFIED:20251114T145002Z
UID:149445-1763557200-1763560800@www.ri.cmu.edu
SUMMARY:Building Robot Hands and Teaching Dexterity
DESCRIPTION:Abstract:  \nOur human hands are masterpieces of power and precision\, capable of typing\, hammering\, or delicately using chopsticks. Yet most robots today still rely on simple two-finger grippers in controlled settings because dexterous hands are costly and difficult to deploy. To close this gap\, I will introduce my LEAP Hands\, high-performance\, low-cost\, and easy-to-assemble robotic hands that have become the most widely used platform for dexterous manipulation research. LEAP Hand V1 employs motor-in-joint actuation for simplicity\, while V2 introduces a novel hybrid rigid–soft structure that delivers exceptional strength and durability.  I will then show how large-scale human video/motion data and simulation techniques can teach human-like manipulation skills across diverse environments.  By tightly integrating mechanical design and machine learning\, my open-source robot hands achieve unprecedented levels of dexterity for a variety of everyday tasks. \nCommittee:\nProf. Deepak Pathak (advisor)\nProf. Nancy Pollard \nProf. Abhinav Gupta \nProf. Jitendra Malik \nDr. Ankur Handa \n  \nA draft of my thesis proposal is available here: \nhttps://kennyshaw.net/phd_thesis_proposal.pdf
URL:https://www.ri.cmu.edu/event/building-robot-hands-and-teaching-dexterity-2/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251107T150000
DTEND;TZID=America/New_York:20251107T163000
DTSTAMP:20261009T020012
CREATED:20251001T150038Z
LAST-MODIFIED:20251104T153929Z
UID:148979-1762527600-1762533000@www.ri.cmu.edu
SUMMARY:Visual-Tactile Synthesis for Texture Generation
DESCRIPTION:Abstract: Recent advances in generative models have enabled the creation of highly realistic visual content\, yet they remain limited to visual perception alone. In contrast\, human interaction with the physical world is inherently multimodal — we not only see textures but also feel them. This gap motivates the goal of my thesis: to build generative models that jointly synthesize visual and tactile modalities for material and texture generation. By unifying what we see and what we touch\, such models can drive new forms of physically grounded content creation\, from robotics and virtual reality to material design.\nHowever\, extending generative modeling to touch presents unique challenges: tactile data is scarce\, noisy\, and expensive to collect\, and there is no large-scale paired dataset linking visual appearance with tactile response. To address these challenges\, my research explores three synergistic directions. \nPart I: I introduce controllable visual-tactile synthesis models that jointly generate aligned visual and tactile textures from shared latent representations\, enabling explicit control over appearance and feel. \nPart II: I propose tactile-aware 3D generation frameworks that integrate tactile sensing into 3D diffusion pipelines\, allowing models to infer physically grounded material properties from visual cues and geometry. \nPart III: Building on these insights\, I aim to develop scalable multimodal generation systems that leverage large vision and language foundation models and physics priors to synthesize novel materials directly from text or image input\, without relying on extensive paired tactile data. \n\n \nThesis Committee:\nJun-Yan Zhu (Co-chair)\nWenzhen Yuan (Co-chair)\nShubham Tulsiani\nAndrew Owens (Cornell Tech)\n\nThesis Document
URL:https://www.ri.cmu.edu/event/ruihan-gao-phd-thesis-proposal/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251020T150000
DTEND;TZID=America/New_York:20251020T163000
DTSTAMP:20261009T020012
CREATED:20251014T201636Z
LAST-MODIFIED:20251014T201948Z
UID:149083-1760972400-1760977800@www.ri.cmu.edu
SUMMARY:Unconstrained Perception for Scalable Robot Manipulation
DESCRIPTION:Abstract: Advances in visual imitation learning driven by large-scale data and expressive policy architectures have yielded impressive progress on long-horizon\, dexterous tasks. However\, current success rates remain insufficient for industrial deployment\, which demands near-perfect reliability on novel tasks. Compared to other fields such as NLP and CV\, the available data in robotics is several orders of magnitude smaller. This raises the question: how can we most effectively leverage priors from large-scale offline data? In this thesis\, I contribute methods to infer strong geometric and dynamic priors for robot manipulation. \nFirst\, geometric camera calibration is a critical prerequisite for real-world vision systems. I will discuss our work on MASt3R-SfM for unconstrained SfM from any image collection in linear complexity. Next\, I discuss how we use a large set of calibrated cameras in DeformGS for photorealistic digital twins with millimeter-accurate tracking of deformable cloth. Removing the need for costly multi-camera systems\, I introduce RaySt3R\, a method to generate complete object geometry from a single RGB-D image. \nBuilding on these works\, I will introduce our ongoing work on Flow2Flow – a flexible end-to-end feedforward architecture for zero-shot dynamics prediction. Many challenging tasks involve manipulating unseen articulated\, deformable\, and cluttered rigid objects; prior approaches rely on pre-trained VLMs\, large-scale 2D point tracking\, or previous interactions with the scene to inject priors. We cast dynamics prediction as a scene flow completion problem from a single RGB-D image\, and propose an optional two-stage adaptation procedure for unseen dynamics. We further study scene flow completion as a 3D pretraining objective for multi-task learning and propose scaling up training on real-world data for the first benchmark in dynamics prediction from a single image. \nThesis Committee Members:\nJeffrey Ichnowski (Chair)\nDeva Ramanan\nShubham Tulsiani\nAbhishek Gupta (University of Washington) \nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/unconstrained-perception-for-scalable-robot-manipulation/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251014T100000
DTEND;TZID=America/New_York:20251014T113000
DTSTAMP:20261009T020012
CREATED:20251001T150325Z
LAST-MODIFIED:20251024T233624Z
UID:148981-1760436000-1760441400@www.ri.cmu.edu
SUMMARY:Consistent Modeling of 4D Scenes for Perception and Generation
DESCRIPTION:Abstract:\n\nA core challenge in vision is building representations that capture 3D scenes over time for perception and interactive generation. For accurate perception and plausible generation\, we want consistency across views\, time\, and modalities. In this talk we explore consistency through the choice of representation\, moving from dense grid formulations to entity-centric scenes that are easier to maintain across frames\, and we extend that representation from perception to generation.\n\nOur past work follows this shift within perception tasks. SOLOFusion uses a grid representation with long- and short-baseline temporal stereo for multi-camera 3D detection\, improving foreground depth\, but it does not perform entity grouping and it does not model background. ASCFormer performs depth estimation and completion via pixel–point affinity\, grouping geometry coherently\, but the grouping is geometric rather than semantic and remains static. DetMatch\, together with our temporal follow-up\, addresses semi-supervised 2D and 3D detection\, aligning detections across modalities and video to produce consistent pseudo-labels and more stable tracklets\, but it focuses on foreground entities and does not model background. S2GO proposes a streaming query-based representation for semantic occupancy estimation that is entity-centric\, temporal\, and models both foreground and background: each persistent query decodes to semantic Gaussians\, and the state is carried across frames and supports short-horizon future prediction. This gives us a single\, stable representation suitable for both perception and sampling. \nWe propose two projects that make this representation generative. First\, we propose a static scene generation method: a diffusion model over grounded queries that represent both foreground and background and are decoded into Gaussians\, generating a complete semantic occupancy scene. This grounded latent representation enables intuitive\, consistent control. Then\, we propose motion generation: a model that generates trajectories for ego and foreground entities conditioned on the generated static scene\, producing coherent 4D rollouts and enabling interactive edits.\n \nThesis Committee Members:\nKris Kitani (Chair)\nDeva Ramanan\nShubham Tulsiani\nWei-Chiu Ma (Cornell University)\n \nLink to Proposal Draft: Link
URL:https://www.ri.cmu.edu/event/jinhyung-park-phd-thesis-proposal/
LOCATION:Newell-Simon Hall 3305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251010T110000
DTEND;TZID=America/New_York:20251010T130000
DTSTAMP:20261009T020012
CREATED:20250909T130917Z
LAST-MODIFIED:20250930T145809Z
UID:148785-1760094000-1760101200@www.ri.cmu.edu
SUMMARY:Pushing the Frontier of Robotic Tool Manipulation by Treating the Hand and the Tool Together as a Machine
DESCRIPTION:Abstract: Tool manipulation is an essential human skill. It expands our manipulation capability beyond the capability of the biological hand\, and is a defining feature of many tasks centered on physical interaction with the real world. For humanoid robots to become general-purpose\, they must master tool manipulation as well. However\, the state-of-the-art humanoid robots equipped with multi-finger hands still fall behind their human counterparts in tool manipulation performance. This thesis aims to narrow this gap by treating the hand and the tool together as a machine.\nSpecifically\, inspired by the analogy between multi-finger hands and CNC machines\, this thesis interprets a tool-manipulating hand as configuring itself and the tool into different tool-hand mechanisms in real time. To concretely represent each tool-hand mechanism—which consists of the tool\, the hand\, and the contacts—this thesis introduces two concepts: 1) foundational pose\, a pose and precondition that the tool and the hand must reach for the tool-hand mechanism to be successfully constructed and to run\, and a concise representation of tool-hand mechanism. 2) sub-assembly\, a set of contacts that independently fulfills part of the tool-hand mechanism’s function\, and a detailed\, modular representation of tool-hand mechanism. \nThis thesis first tests the validity of the concept of foundational pose via the question: “if a tool and a hand have reached a foundational pose\, can they act as the corresponding tool-hand mechanism and perform the tool manipulation motion?” To answer this question\, the thesis conducts a hand design experiment\, which uses foundational poses as constraints to sample many different hands and evaluates their tool manipulation motions. The results lead to a positive answer to the question\, verifying the concept of foundational pose. \n\nThen\, this thesis expands roll-slide contact-based tool manipulation motion planning—which previously was only possible for primitive shapes with global parametrizations—to manifold meshes\, which allows motion planning from foundational poses for arbitrarily shaped tools and hands. \nFinally\, for the proposed work\, this thesis aims to test the validity of the concept of sub-assembly via the question: “how many sub-assemblies are enough?” Based on the answer to this question\, this thesis aims to develop a sub-assembly-based control framework\, and test the framework on a real robotic hand for an entire tool manipulation sequence. \n\n \n\nThesis Committee Members: \nProf. Nancy Pollard (co-chair)\nProf. Jean Oh (co-chair)\nProf. Matthew Mason\nDr. Lael Odhner (The Robotics and AI Institute)\n\n\nDraft of the Thesis Proposal Document Link: https://drive.google.com/file/d/1L5ri7r0295poQOyTtqEOI3-LX3o7Gqlf/view?usp=drive_link
URL:https://www.ri.cmu.edu/event/pushing-the-frontier-of-robotic-tool-manipulation-by-treating-the-hand-and-the-tool-together-as-a-machine/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20251009T093000
DTEND;TZID=America/New_York:20251009T110000
DTSTAMP:20261009T020012
CREATED:20250903T175132Z
LAST-MODIFIED:20250930T131653Z
UID:148746-1760002200-1760007600@www.ri.cmu.edu
SUMMARY:Customizing Text-to-Image Diffusion Models
DESCRIPTION:Abstract: With the rapid advancement of generative models\, their potential to transform creative content creation is increasingly evident. However\, most large-scale generative models are primarily text-conditioned\, given the availability of large-scale paired text–image datasets. In contrast\, for most practical applications\, creators often begin from an existing asset and wish to generate variations or modify it in specific ways. For images\, this may involve placing an object in a new context\, adjusting local attributes\, or altering visual style. My research focuses on customizing pre-trained generative models\, primarily text-to-image diffusion models\, to facilitate such downstream tasks. A central challenge here is the lack of paired input–output data for these tasks. \nTo address this\, I explore three complementary directions: \nPart I: I study few-shot learning methods\, which are computationally efficient but require fine-tuning for each new task instance. This limitation motivates the second direction. \nPart II: Constructing synthetic paired datasets using the capabilities of pre-trained generative models themselves to train feed-forward models in a supervised manner. However\, constructing such datasets requires careful curation\, filtering\, and risk of becoming outdated as base pre-trained models evolve. Building on these insights\, my thesis proposes a third paradigm. \nPart III: Customizing generative models without paired supervision. Instead\, we plan to leverage vision–language models to evaluate task success and provide direct gradient-based feedback to the generative model. This approach has the potential to create a scalable and robust framework for efficient customization of generative models for downstream tasks without relying on synthetic datasets. \n\n\n \nThesis Committee:\nJun-Yan Zhu (Chair)\nDeva Ramanan\nShubham Tulsiani\nPhillip Isola (MIT)\n\nThesis Proposal Draft
URL:https://www.ri.cmu.edu/event/customizing-text-to-image-diffusion-models/
LOCATION:Newell-Simon Hall 4305
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20250919T100000
DTEND;TZID=America/New_York:20250919T120000
DTSTAMP:20261009T020012
CREATED:20250911T175647Z
LAST-MODIFIED:20250911T175647Z
UID:148841-1758276000-1758283200@www.ri.cmu.edu
SUMMARY:Accessible Dexterous Manipulation with Soft Hands: Designs\, Methods\, Models\, and the DexKit Platform
DESCRIPTION:Abstract: \nRobot dexterity remains an open challenge in robotics that has the potential to transform\nmanufacturing\, healthcare\, and daily life. Robots that safely and robustly interact with unstructured\nenvironments must combine compliant hardware with models and planners that tolerate uncertainty.\nAdditionally\, if robust robot dexterity is to be realized outside of research labs\, it must be accessible\nto a broad audience\, with low-cost hardware and open-source software. \nThis thesis advances dexterous manipulation with soft\, tendon-driven hands by integrating new\nfabrication and control methods\, data-driven models of manipulation capabilities\, a fast algorithm\nfor robust grasp synthesis\, and a first-of-its-kind accessible experimental platform. \nI introduce fully soft foam hands actuated by tendons routed on textile skins. I detail a\nsimple molding-and-casting pipeline\, validate a soft-body simulation framework\, compare inverse-\nkinematics control strategies\, and optimize nontrivial tendon routings. I further report a user study\non human-designed routings\, demonstrations of power/precision grasps and in-hand manipulation\,\nsub-millimeter repeatability\, and year-long durability\, alongside an analysis of limitations (e.g.\,\nrouting through foam\, sensing\, and sim-to-real gaps). \nTo tackle the challenge of planning with soft hands\, I develop data-driven models of manipulation\ncapabilities that capture the inherent uncertainty and redundancy of soft hands. Additionally I\ndemonstrate a fast\, anytime method to compute globally optimal Independent Contact Regions\n(ICRs) by iteratively building an incremental Delaunay triangulation over grasp configuration space.\nI show that ICRs guide simple policies that remain robust to real-world uncertainties in object size\,\npose\, and geometry. \nFinally\, to promote accessibility\, I contribute the DexKit system\, a low-cost\, anthropomorphic system (12\nactuated DoF hand on a 4-DoF gantry) that can be built for under $2000. \nTheses Committee Members:\nNancy Pollard (chair)\nMatthew Mason\nChristopher Atkeson\nJames Bern (Williams College)\n\nDraft of the Thesis Proposal Document Link
URL:https://www.ri.cmu.edu/event/accessible-dexterous-manipulation-with-soft-hands-designs-methods-models-and-the-dexkit-platform/
LOCATION:GHC 4405
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
END:VCALENDAR