From Video Generation to Video World Models
Abstract: Video diffusion models have achieved remarkable success in content creation, yet they still fall short of simulating interactive worlds that respond to users in real time. This talk examines the fundamental challenges preventing these models from evolving into true world simulators. I will present a series of works — CausVid, Self-Forcing, MotionStream, and State-Space [...]
Unifying Perception and Creation with Generative Models
Abstract: Recent advances in large-scale generative modeling have reshaped our understanding of visual intelligence. While models such as diffusion and autoregressive transformers have achieved remarkable success in image and video synthesis, their potential for visual perception and understanding remains underexplored. This thesis investigates how generative models can serve as powerful visual learners—bridging the long-standing divide [...]
Visual-Tactile Synthesis for Texture Generation
Abstract: Recent advances in generative models have enabled the creation of highly realistic visual content, yet they remain limited to visual perception alone. In contrast, human interaction with the physical world is inherently multimodal — we not only see textures but also feel them. This gap motivates the goal of my thesis: to build generative models [...]
Just Asking Questions
Abstract: In the age of deep networks, "learning" almost invariably means "learning from examples". We train language models with human-generated text and labeled preference pairs, image classifiers with large datasets of images, and robot policies with rollouts or demonstrations. When human learners acquire new concepts and skills, we often do so with richer supervision, especially [...]
Towards Modernization of Long-Range Image-Space Planning for Off-Road Navigation
Abstract: This thesis revisits long-range, image-space planning for off-road navigation and modernizes the classical first-person view (FPV) paradigm by building upon recent advances in perception. It introduces a lightweight depth calibration scheme, analytic configuration-space (C-space) transforms, interpretable frontier selection, and a pixel-space A* planner with validated heuristic soundness. Concretely, we (i) make monocular depth metrically [...]
Building Robot Hands and Teaching Dexterity
Abstract: Our human hands are masterpieces of power and precision, capable of typing, hammering, or delicately using chopsticks. Yet most robots today still rely on simple two-finger grippers in controlled settings because dexterous hands are costly and difficult to deploy. To close this gap, I will introduce my LEAP Hands, high-performance, low-cost, and easy-to-assemble robotic [...]
Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models
Abstract: Transferring skills between different objects remains one of the core challenges of open-world robot manipulation. Generalization needs to take into account the high-level structural differences between distinct objects while still maintaining similar low-level interaction control. In this paper, we propose an example-based zero-shot approach to skill transfer. Rather than treating skills as atomic, we [...]
OpenVDB
Abstract: As the inventor of VDB and founder of OpenVDB, I am excited to talk about its history, motivation, and diverse adoption. Specifically, this lecture will cover the underlying VDB data structure, and its adoption to computer graphics, physics simulations and more recently machine learning. Since its open-source release in 2012, OpenVDB has become an industry [...]
How to Coordinate Thousands of Robots Efficiently and Robustly
Abstract: Large-scale robot fleets are increasingly deployed in warehouses, factories, transportation systems, and emerging robotics applications. Coordinating hundreds or thousands of robots in shared, cluttered spaces creates fundamental challenges in maintaining safety, preventing deadlocks, and minimizing congestion. In this talk, I will present our recent work on scalable imitation learning methods for coordinating 10k robots, automatic environment [...]
Sensorimotor-Aligned Design for Pareto-Efficient Haptic Immersion in Extended Reality
Abstract: A new category of computing devices has emerged: augmented and virtual reality headsets, collectively referred to as extended reality (XR). These devices can alter, augment, or even replace our reality. While these headsets have made impressive strides in audio-visual immersion over the past half-century, XR interactions remain almost completely absent of appropriately expressive tactile [...]