AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

Active Retrieval Augmented LLMs (FLARE)
Actively decides when and what to retrieve during generation.

FrugalGPT
Strategies to reduce LLM inference cost while improving performance.

StarCoder
An open-access 15.5B code LLM with 8K context and 80+ programming languages.

MultiModal-GPT
A vision-language model for multi-round dialogue fine-tuned from OpenFlamingo.

scGPT
A foundation model for single-cell multi-omics pretrained on 10 million cells.

GPTutor
A ChatGPT-powered VSCode extension for code explanation.

Shap-E
OpenAI's conditional generative model for 3D assets producing implicit functions.

Are Emergent Abilities of LLMs a Mirage?
Stanford's critical re-examination of emergent abilities.

Interpretable ML for Science with PySR
An open-source library for practical symbolic regression in the sciences.

PMC-LLaMA
A LLaMA model fine-tuned on 4.8 million medical papers.

Distilling Step-by-Step!
A mechanism to train smaller models that outperform larger LLMs using fewer examples.

Poisoning Language Models During Instruction Tuning
Shows adversaries can poison LLMs via instruction tuning data.

Unlimiformer
Long-range Transformers with unlimited length input via external datastores.

Learning to Reason and Memorize with Self-Notes
LLMs that deviate from input to explicitly "think" and memorize.

Learning Agile Soccer Skills for a Bipedal Robot with Deep RL
DeepMind's bipedal humanoid robot playing soccer.

Scaling Transformer to 1M tokens with RMT
Recurrent Memory Transformer extends BERT's effective context to 2M tokens.

Track Anything
An interactive tool for video object tracking and segmentation built on Segment Anything.

A Cookbook of Self-Supervised Learning
A comprehensive overview of SSL techniques and practical considerations.

Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond
A practical guide for practitioners working with LLMs.

AudioGPT
Connects ChatGPT with audio foundational models for speech, music, sound, and talking head tasks.

DataComp
A multimodal dataset benchmark with 12.8B image-text pairs.

ChatGPT for Information Extraction
A deeper assessment of ChatGPT on information extraction tasks.

Comparing Physician vs ChatGPT (JAMA)
A JAMA Internal Medicine study comparing physician and ChatGPT responses.

Stable and Low-Precision Training for Large-Scale Vision-Language Models
Methods for accelerating and stabilizing large VLM training.