🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Claude 3.5 Sonnet

Paper preview
Claude 3.5 Sonnet
Paper summary

a new model that achieves state-of-the-art performance on several common benchmarks such as MMLU and HumanEval; it outperforms Claude 3 Opus and GPT-4o on several benchmarks with the exception of math word problem-solving tasks; achieves strong performance on vision tasks which also helps power several new features like image-text transcription and generation of artifacts.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack