🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Reasoning

CellVerse

First page
CellVerse
Paper summary

Introduces a benchmark to evaluate LLMs on single-cell biology tasks by converting multi-omics data into natural language. While generalist LLMs like DeepSeek and GPT-4 families show some reasoning ability, none significantly outperform random guessing on key tasks like drug response prediction, exposing major gaps in biological understanding by current LLMs.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack