
Open Problems and Limitations of RLHF
A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

Med-Flamingo
Stanford's Med-Flamingo is a multimodal medical model supporting in-context learning for few-shot medical visual QA.

ToolLLM
Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

Skeleton-of-Thought (SoT)
Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

MetaGPT
MetaGPT is a multi-agent framework that encodes standardized operating procedures (SOPs) for complex problem solving.

OpenFlamingo
An open-source family of autoregressive vision-language models spanning 3B to 9B parameters.

The Hydra Effect
DeepMind shows that language models exhibit self-repairing behavior when attention heads are ablated.

Self-Check
Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

Dynalang (Agents Model the World with Language)
UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

AutoRobotics-Zero
Discovers zero-shot adaptable robot policies from scratch, including the automatic discovery of Python control code.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack