Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

World models and embodied agents

When: Monday October 26th 2026

Overview

This session opens the “reasoning about the world” day. A 45-minute lecture, followed by questions, introduces world models: AI systems that learn to predict how their environment evolves, and to use these predictions to perceive, plan and act. The lecture will focus on video world models such as V-JEPA, and on what interpretability methods reveal about how they represent physical structure. It is followed by a two-hour hands-on tutorial on vision-language-action (VLA) models, which connect perception and language to motor control in embodied agents.

Instructors

Sonia Joseph is an AI researcher working on interpretable video world models, and a PhD student at McGill University and Mila. She conducted research on the JEPA team at Meta, where she also led the company’s internal interpretability community, and published the first mechanistic interpretability work on physical reasoning in video world models such as V-JEPA 2. She is also a developer of Prisma, an open source toolkit for mechanistic interpretability in vision and video models.

Glen Berseth is an assistant professor in the Department of Computer Science and Operations Research (DIRO) at Université de Montréal, a core academic member of Mila, and a Canada CIFAR AI Chair. He co-directs the Robotics and Embodied AI Lab (REAL). His research uses deep reinforcement learning to build generalist robots that learn from real-world, sequential decision making. He will lead the hands-on tutorial.

Objectives

Materials

Coming soon.