🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Evaluation · Code

TheAgentCompany

First page
TheAgentCompany
Paper summary

a new benchmark for evaluating AI agents on real-world professional tasks in a simulated software company environment; tasks span multiple professional roles including software engineering, project management, finance, and HR; when tested with various LLMs, including both API-based models like Claude-3.5-Sonnet and open-source models like Llama 3.1, the results show the current limitations of AI agents. The best-performing model, Claude-3.5-Sonnet, achieved only a 24% success rate on completing tasks fully while scoring 34.4% when accounting for partial progress.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack