When: Tuesday October 27th 2026, 09:00-12:00
Overview¶
This three-hour tutorial opens the “benchmarking” day. Participants will learn how to evaluate the reasoning abilities of large language models, drawing on paradigms from cognitive science to design benchmarks that go beyond accuracy.
Instructors¶
To be announced.
Objectives¶
Learn how reasoning is elicited and measured in large language and reasoning models.
Build a small reasoning benchmark inspired by a cognitive psychology task.
Run the benchmark on open models, and analyse their errors.
Materials¶
Coming soon.