Decision-Making in a World of Latent Particles - Robotics Institute Carnegie Mellon University
Loading Events

VASC Seminar

September

21
Mon
Tal Daniel Postdoctoral Fellow Robotics Institute,
Carnegie Mellon University
Monday, September 21
3:30 pm to 4:30 pm
3305 Newell-Simon Hall
Decision-Making in a World of Latent Particles

Abstract: Robots must often make decisions in scenes containing many objects: they need to identify what is present, understand where objects are, predict how they will interact, and choose actions accordingly. Learning these capabilities directly from pixels is challenging, especially when the number and arrangement of objects can change from one scene to another.

In this talk, I will presentDeep Latent Particles (DLP), aself-supervisedobject-centric representation that describes a visual scene as a set of compact latent particles. Each particle captures the location and visual properties of a discovered object or object part, providing an interpretable bridge between raw images and multi-object decision-making.

I will show how DLP can serve as a representation for learning robotic policies from online reinforcement learning, offline data, and demonstrations. In particular, policies built on these representations can generalize compositionally to scenes containing more objects than were present during training.

I will then introduce Latent Particle World Models(ICLR 2026 Oral), which learn to predict how collections of latent particles evolve over time. These object-centric world models support multi-view observations and flexible conditioning, enabling prediction and decision-making in rich visual environments. I will discuss how this perspective connects to diffusion-based policies and world action models (WAMs), and conclude with a look toward self-supervised 3D object-centric learning for robots that can perceive, predict, and act in three-dimensional worlds. and which will be quietly subsumed by the next scale-up.

 

Bio: Tal Daniel is a Postdoctoral Fellow at Carnegie Mellon University’s Robotics Institute, working with Prof. Deepak Pathak and Prof. David Held. He received his Ph.D. in Electrical and Computer Engineering from the Technion, advised by Prof. Aviv Tamar. His research spans self-supervised and object-centric representation learning, generative modeling, reinforcement learning, and robotics, with a focus on learning representations and world models.

 

Homepage: https://taldatech.github.io

 

Sponsor:

The VASC seminar is generously sponsored by HeyGen, an all-in-oneAI-powered video generation platform that leverages advances incomputer vision, generative modeling, and multimodal learning to makehigh-quality video creation both scalable and accessible.