
Let's Verify Step by Step
OpenAI's landmark paper on process reward models for mathematical reasoning.

No Positional Encodings (NoPE)
Shows explicit position embeddings aren't essential for decoder-only Transformers.

BiomedGPT
A unified biomedical GPT for vision, language, and multimodal tasks.

Thought Cloning
Imitation learning framework that learns to think as well as act.

Fine-Tuning Language Models with Just Forward Passes (MeZO)
A memory-efficient zeroth-order optimizer for LLM fine-tuning.

MERT
An acoustic music understanding model with large-scale self-supervised training.

Bytes Are All You Need
Performs classification directly on file bytes without decoding.

Direct Preference Optimization (DPO)
Rafailov et al.'s simpler alternative to RLHF that rivals full RL-based alignment.

SQL-PaLM
An LLM-based Text-to-SQL system built on PaLM-2.

CodeTF
An open-source Transformer library for state-of-the-art code LLMs.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack