AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

CodeT5+
An open code LLM family for code understanding and generation.

PaLM 2
Google's second-generation PaLM powering Bard and Google products.

InstructBLIP
Visual-language instruction tuning built on BLIP-2.

Active Retrieval Augmented LLMs (FLARE)
Actively decides when and what to retrieve during generation.

Are Emergent Abilities of LLMs a Mirage?
Stanford's critical re-examination of emergent abilities.

Interpretable ML for Science with PySR
An open-source library for practical symbolic regression in the sciences.

PMC-LLaMA
A LLaMA model fine-tuned on 4.8 million medical papers.

Distilling Step-by-Step!
A mechanism to train smaller models that outperform larger LLMs using fewer examples.

Scaling Transformer to 1M tokens with RMT
Recurrent Memory Transformer extends BERT's effective context to 2M tokens.

Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond
A practical guide for practitioners working with LLMs.

DataComp
A multimodal dataset benchmark with 12.8B image-text pairs.

ChatGPT for Information Extraction
A deeper assessment of ChatGPT on information extraction tasks.

Comparing Physician vs ChatGPT (JAMA)
A JAMA Internal Medicine study comparing physician and ChatGPT responses.

Evaluating Verifiability in Generative Search Engines
Audits popular generative search engines for citation accuracy.

OpenAGI: When LLM Meets Domain Experts
An open-source research platform for LLM agents manipulating domain expert models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
A benchmark using real human standardized exams.

Segment Everything Everywhere All at Once (SEEM)
A promptable, interactive segmentation model.

Eight Things to Know about Large Language Models
Sam Bowman's influential primer on key LLM considerations.

A Survey of Large Language Models
A 50-page comprehensive survey on LLMs.

MACHIAVELLI Benchmark
A benchmark of 134 text-based Choose-Your-Own-Adventure games for measuring ethical trade-offs.