Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Evaluating reasoning in large language models

When: Tuesday October 27th 2026, 09:00-12:00

Overview

This three-hour tutorial opens the “benchmarking” day. Participants will learn how to evaluate the reasoning abilities of large language models, drawing on paradigms from cognitive science to design benchmarks that go beyond accuracy.

Instructors

To be announced.

Objectives

Materials

Coming soon.