AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

Long-range Language Modeling with Self-Retrieval
Jointly trains a retrieval-augmented LM from scratch for long-range modeling.

Textbooks Are All You Need (phi-1)
Introduces a 1.3B parameter code LLM trained on textbook-quality data.

FinGPT
An open-source LLM for the finance sector with a data-centric approach.

Crowd Workers Widely Use LLMs for Text Production
Empirical evidence that 33-46% of MTurk crowd workers used LLMs on text tasks.

Reliability of Watermarks for LLMs
Studies whether watermarks survive human rewriting and LLM paraphrasing.

Benchmarking NN Training Algorithms (AlgoPerf)
A new benchmark for rigorously evaluating optimizers using realistic workloads.

Augmenting LLMs with Long-term Memory (LongMem)
Enables LLMs to memorize long history via memory-augmented adaptation.

TAPIR
Tracks any queried point on any physical surface throughout a video sequence faster than real-time.

Mind2Web
A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

Tracking Everything Everywhere All at Once (OmniMotion)
Test-time optimization for dense, long-range motion estimation.

AlphaDev
DeepMind's deep RL agent discovering faster sorting algorithms from scratch, now in LLVM.

Sparse-Quantized Representation (SpQR)
Tim Dettmers' near-lossless LLM compression technique.

MusicGen
A simple and controllable model for music generation using a single-stage Transformer.

Humor in ChatGPT
Explores ChatGPT's capabilities to grasp and reproduce humor.

Let's Verify Step by Step
OpenAI's landmark paper on process reward models for mathematical reasoning.

BiomedGPT
A unified biomedical GPT for vision, language, and multimodal tasks.

Thought Cloning
Imitation learning framework that learns to think as well as act.

MERT
An acoustic music understanding model with large-scale self-supervised training.

SQL-PaLM
An LLM-based Text-to-SQL system built on PaLM-2.

CodeTF
An open-source Transformer library for state-of-the-art code LLMs.

Model Evaluation for Extreme Risks
DeepMind's framework for evaluating models for catastrophic-risk capabilities.

LLM Research Directions
A list of research directions for students entering LLM research.

Reinventing RNNs for the Transformer Era (RWKV)
Combines parallelizable training of Transformers with efficient RNN inference.

Towards Expert-Level Medical Question Answering (Med-PaLM 2)
Google's second-generation medical LLM.