AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

Synthetic Data Reduces Sycophancy
Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

Open Problems and Limitations of RLHF
A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

Med-Flamingo
Stanford's Med-Flamingo is a multimodal medical model supporting in-context learning for few-shot medical visual QA.

AutoRobotics-Zero
Discovers zero-shot adaptable robot policies from scratch, including the automatic discovery of Python control code.

RT-2
Google DeepMind's end-to-end vision-language-action model that learns from both web and robotics data to control robots.

Tracking Anything in High Quality
A framework for high-quality tracking-anything in videos combining segmentation and refinement.

Foundation Models in Vision
A comprehensive survey on foundational models for computer vision and their open research directions.

LoraHub
Enables efficient cross-task generalization via dynamic LoRA composition.

Survey of Aligned LLMs
A comprehensive overview of alignment approaches covering data, training, and evaluation.

Llama 2
Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

CM3Leon
Meta's retrieval-augmented multi-modal language model that generates both text and images.

Secrets of RLHF in LLMs
A deep investigation into RLHF with a focus on the inner workings of PPO, including open-source code.

AnimateDiff
Animates frozen text-to-image diffusion models via a plug-in motion modeling module.

Generative Pretraining in Multimodality (Emu)
A transformer-based multimodal foundation model for generating images and text.

Multimodal Generation with Frozen LLMs
Maps images to LLM token space enabling models like PaLM and GPT-4 to handle visual tasks without parameter updates.

CodeGen2.5
Salesforce's new 7B code LLM trained on 1.5T tokens and optimized for fast sampling.

Extending Context Window of LLMs (PI)
Position Interpolation extends LLaMA's context to 32K with minimal fine-tuning (within 1000 steps).

Visual Navigation Transformer (ViNT)
A foundation model for vision-based robotic navigation built on flexible Transformers.

Textbooks Are All You Need (phi-1)
Introduces a 1.3B parameter code LLM trained on textbook-quality data.

RoboCat
DeepMind's self-improving foundation agent that operates different robotic arms from as few as 100 demonstrations.

ClinicalGPT
A language model optimized through extensive and diverse medical data and multi-turn dialogue.

LOMO
A memory-efficient optimizer that combines gradient computation and parameter update in one step.

LMFlow
An extensible and lightweight toolkit for fine-tuning and inference of large foundation models.

FinGPT
An open-source LLM for the finance sector with a data-centric approach.