🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papersIssue 42 of 176

The week of Jan 15 – Jan 21, 2024

10 papers, hand-picked and summarised.

AlphaGeometry

AlphaGeometry

DeepMind's AlphaGeometry is a theorem prover that solves Olympiad-level geometry problems at near gold-medallist performance, and crucially, without needing any human demonstrations.

01Reasoning
AlphaCodium

AlphaCodium

AlphaCodium is a test-based, iterative "flow" that turns off-the-shelf LLMs into strong competitive-programming solvers without model training.

02Code
RAG vs. Finetuning

RAG vs. Finetuning

Microsoft researchers systematically compare RAG and fine-tuning (and their combination) on LLMs like Llama 2 and GPT-4 using an agricultural domain dataset.

03Retrieval
Self-Rewarding Language Models

Self-Rewarding Language Models

Meta shows that an LLM can act as both actor and judge in its own alignment loop, generating training data without any external reward model.

04Reinforcement Learning
Tuning Language Models by Proxy

Tuning Language Models by Proxy

Proxy-tuning steers a large frozen LLM by *decoding-time* logit arithmetic using a much smaller fine-tuned model as a "proxy".

05Training
ReFT (Reinforced Fine-Tuning)

ReFT (Reinforced Fine-Tuning)

ByteDance's ReFT enhances LLM reasoning by combining supervised fine-tuning with online RL that samples alternative reasoning paths, without a learned reward model.

06Reasoning
Overview of LLMs for Evaluation

Overview of LLMs for Evaluation

A thorough survey of LLM-as-a-Judge and LLM-based evaluation methodologies, mapping strengths, limitations, and open problems.

07Evaluation
Patchscopes

Patchscopes

Patchscopes is a general framework for inspecting and intervening on LLM internals by "patching" hidden representations into a second inference pass.

08Reasoning
Easy-to-Hard Generalization

Easy-to-Hard Generalization

UNC researchers show that LLMs often generalize well from easy training data to hard evaluation data, with implications for scalable oversight.

09Evaluation
MoE-Mamba

MoE-Mamba

MoE-Mamba combines state-space models (Mamba) with Mixture-of-Experts to scale LLMs more efficiently than either Mamba or Transformer-MoE alone.

10Architecture
Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack