Understanding Theory-of-Mind in LLMs with LLMs
Free while signed in. Answers cite the passages they came from.

A framework for procedurally generating ToM evaluations using LLMs themselves.
LLM-generated benchmarks: Uses LLMs to procedurally create diverse ToM scenarios, avoiding benchmark contamination and enabling unlimited test generation.
Social reasoning study: Evaluates whether LLMs can track beliefs, intentions, and false beliefs of multiple agents - classic ToM challenges.
Controlled difficulty: Procedural generation allows varying difficulty (number of agents, nesting depth) to map capability boundaries.
Evaluation pattern: Early example of using LLMs to generate evaluations for LLMs - a pattern that would become standard in 2024 synthetic evaluation work.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack