Understanding Theory-of-Mind in LLMs with LLMs
First page

Paper summary
A framework for procedurally generating ToM evaluations using LLMs themselves.
Ask this paper
01
LLM-generated benchmarks: Uses LLMs to procedurally create diverse ToM scenarios, avoiding benchmark contamination and enabling unlimited test generation.
02
Social reasoning study: Evaluates whether LLMs can track beliefs, intentions, and false beliefs of multiple agents - classic ToM challenges.
03
Controlled difficulty: Procedural generation allows varying difficulty (number of agents, nesting depth) to map capability boundaries.
04
Evaluation pattern: Early example of using LLMs to generate evaluations for LLMs - a pattern that would become standard in 2024 synthetic evaluation work.