
Transformers as Support Vector Machines
A theoretical paper establishing a formal connection between self-attention optimization and hard-margin SVM problems.

RLAIF (Scaling RLHF with AI Feedback)
Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

GPT Solves Math Problems Without a Calculator
Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

OPRO (LLMs as Optimizers)
DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

ImageBind-LLM
Shanghai AI Lab's ImageBind-LLM brings six-modality understanding to LLMs via the ImageBind joint embedding space.

Explaining Grokking
DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

Overview of AI Deception
A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

FLM-101B
A 101B parameter open LLM trainable on a $100K budget through a growth-based training strategy.

Cognitive Architectures for Language Agents (CoALA)
Princeton proposes CoALA, a systematic framework for understanding and building language agents.

Q-Transformer
Google's Q-Transformer is a scalable RL method for training multi-task robotic policies from large offline datasets.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack