🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning · Evaluation

GSM-Symbolic

First page
GSM-Symbolic
Paper summary

tests several SoTA models on a benchmark created with symbolic templates that enable diverse mathematical problems; they find that LLMs exhibit variance when responding to variations of the same questions; the performance of all the models declines by adjusting the numerical values in the question; as questions are made more challenging (e.g., increasing the number of clauses) the performance significantly deteriorates; the authors hypothesize that the observed decline in performance is due to a lack of logical reasoning in current LLMs.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack