AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

Min-K% Prob (Detecting Pretraining Data)
Proposes Min-K% Prob as an effective detection method for determining whether specific text was in an LLM's pretraining data.

CommonCanvas
Releases CommonCanvas, a text-to-image dataset composed entirely of Creative-Commons-licensed images.

Llemma
Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

MentaLLaMA
An open-source LLM family specialized for interpretable mental-health analysis on social media.

Struc-Bench (LLMs for Structured Data)
Studies how LLMs handle complex structured-data generation and proposes a structure-aware fine-tuning method.

LMSYS-Chat-1M
LMSYS releases a large-scale dataset of 1 million real-world LLM conversations collected from the Vicuna demo and Chatbot Arena.

Textbooks Are All You Need II (phi-1.5)
Microsoft's phi-1.5 demonstrates that a 1.3B model trained on "textbook-quality" synthetic data rivals much larger models on reasoning.

Q-Transformer
Google's Q-Transformer is a scalable RL method for training multi-task robotic policies from large offline datasets.

SAM-Med2D
Adapts the Segment Anything Model (SAM) to 2D medical imaging through large-scale medical fine-tuning.

AnomalyGPT
Applies large vision-language models to industrial anomaly detection with synthetic data augmentation.

Survey on Instruction Tuning for LLMs
A comprehensive survey of instruction tuning covering methodology, dataset construction, and applications.

Platypus
Platypus is a family of fine-tuned and merged LLMs that topped the Open LLM Leaderboard in August 2023.

GEARS
Stanford's GEARS predicts cellular responses to genetic perturbation using deep learning + a gene-relationship knowledge graph.

OctoPack
Hugging Face releases OctoPack, a 4TB dataset of Git commits across 350 programming languages for instruction-tuning code LLMs.

Synthetic Data Reduces Sycophancy
Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

PUG (Photorealistic Unreal Graphics)
Meta's PUG uses Unreal Engine to generate photorealistic, semantically controllable synthetic datasets for vision research.

Textbooks Are All You Need (phi-1)
Introduces a 1.3B parameter code LLM trained on textbook-quality data.

Crowd Workers Widely Use LLMs for Text Production
Empirical evidence that 33-46% of MTurk crowd workers used LLMs on text tasks.

Mind2Web
A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

Let's Verify Step by Step
OpenAI's landmark paper on process reward models for mathematical reasoning.

TinyStories
Explores how small LMs can be and still speak coherent English.

DoReMi
Optimizes data mixtures for faster language model pretraining.

DataComp
A multimodal dataset benchmark with 12.8B image-text pairs.

Segment Anything (SAM)
Meta's foundational model for image segmentation with massive training data release.