🚀NEW LABGetting Started with Claude AgentsStart lab
Agents

LLMs Still Can’t Plan

First page
LLMs Still Can’t Plan
Paper summary

evaluates whether large reasoning models such as o1 can plan; finds that a domain-independent planner can solve all instances of Mystery Blocksworld but LLMs struggle, even on small instances; o1-preview is effective on the task but tend to degrade in performance as plan length increases, concludes that while o1 shows progress on more challenging planning problems, the accuracy gains cannot be considered general or robust.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack