When: Monday October 26th 2026, 13:00-14:45
Overview¶
Vision-language-action (VLA) models extend large vision-language models with the ability to produce actions. This session will present the building blocks of VLA models and how they can be used as models of perception and behaviour in neuroAI.
Instructors¶
The session will be run by members of the teams of Shahab Bakhtiari and Pouya Bashivan.
Objectives¶
Learn the architecture of vision-language-action models, and how they are trained.
Run a pretrained VLA model on a simulated task.
Discuss how VLA models can be compared with human perception and behaviour.
Materials¶
Coming soon.

