BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Robotics Institute Carnegie Mellon University - ECPv6.15.12.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Robotics Institute Carnegie Mellon University
X-ORIGINAL-URL:https://www.ri.cmu.edu
X-WR-CALDESC:Events for Robotics Institute Carnegie Mellon University
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260901T113000
DTEND;TZID=America/New_York:20260901T130000
DTSTAMP:20261007T163631
CREATED:20260820T203402Z
LAST-MODIFIED:20260820T203402Z
UID:153235-1788262200-1788267600@www.ri.cmu.edu
SUMMARY:Scalable Vision-Language Models through Unified 2D and 3D Representations
DESCRIPTION:Abstract:\nVision-language models have become remarkably capable on images and short videos\, yet they still struggle with two abilities central to embodied intelligence: spatial understanding and long-range temporal reasoning. A major reason is representational: today’s models process videos as long sequences of 2D patches\, so computation grows with observation length even when the underlying scene changes little. This thesis argues that organizing perception around persistent 3D structure rather than individual frames allows model complexity to scale with scene content rather than observation length\, leading to efficient inference on long videos while providing a stronger foundation for spatial reasoning.\n\nWe introduce a unified representation based on 3D feature clouds that handles both 2D and 3D inputs within a single architecture. Because the same model trains on abundant 2D image-text data alongside available 3D data\, it acquires better spatial understanding without sacrificing 2D performance. We demonstrate this on 3D instance segmentation (ODIN)\, extend it to broader vision-language tasks (UniVLG)\, and scale it to billion-parameter VLMs (Qwen-3D)\, showing consistent improvements in spatial understanding and inference efficiency. \nWe then move beyond static scenes to dynamic videos. We develop 3D scene representations (TrackEverything) that disentangle static and dynamic content\, deduplicate the scene across time\, and track all points in 3D throughout long videos. These representations convert videos into concise spatiotemporal structures that grow with scene complexity rather than video duration\, enabling new capabilities such as being able to track all points across all frames in long videos (1000+ frams). \nFinally\, we outline proposed work on two fronts: leveraging these dynamic 3D representations for spatial and motion reasoning in vision-language models\, and scaling dynamic 3D tracking to real-world video data using heterogeneous supervision beyond synthetic data. \nTogether\, this work makes the case for moving beyond 2D patch representations toward 3D-native models for better and more efficient video understanding. \nThesis Committee:\nKaterina Fragkiadaki (Chair)\nDeva Ramanan\nShubham Tulsiani\nLeonidas Guibas\, Stanford University\n\nThesis Proposal Link
URL:https://www.ri.cmu.edu/event/scalable-vision-language-models-through-unified-2d-and-3d-representations/
LOCATION:3305 Newell-Simon Hall
CATEGORIES:PhD Thesis Proposal,Student Talks
END:VEVENT
END:VCALENDAR