🚀NEW LABGetting Started with Claude AgentsStart lab
Agents

Grok-2

Full-paper indexing in progress
Paper preview
Grok-2
Paper summary

a new frontier model with strong code, math, and reasoning capabilities which includes a large and small model; outperforms both Claude 3.5 Sonnet and GPT-4-Turbo on the LMSYS Chatbot Arena; claims to improve capabilities including instruction following, retrieval, tool use, and enhancing factuality; competes with Claude 3.5 Sonnet (June release) and GPT-4o (May release) on MMLU and HumanEval.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack