AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

KTO (Kahneman-Tversky Optimization)
Contextual AI introduces KTO, an alignment objective derived from prospect theory that works with binary "good/bad" signals instead of preference pairs.

Seamless
Meta's Seamless is a family of models for end-to-end expressive, streaming cross-lingual speech communication.

Safe Deployment of Generative AI (Nature)
A Nature correspondence arguing that medical professionals - not commercial interests - must drive the development and deployment of generative AI in medicine.

Translatotron 3
Google's Translatotron 3 performs speech-to-speech translation using only monolingual data - no parallel corpora required.

Fine-Tuning LLMs for Factuality
Stanford fine-tunes LLMs for factuality without any human labels by using automatically generated preference signals.

MART (Multi-round Automatic Red-Teaming)
Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

LLMs Can Deceive Users (Trading Agent)
Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

Hallucination in LLMs Survey
A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

Evaluating LLMs Survey
A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

Zephyr
Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

Managing AI Risks (Bengio, Hinton, et al.)
A high-profile position paper by leading AI researchers laying out risks from upcoming advanced AI systems.

LLM Self-Explanations
Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

LLMs for Healthcare Survey
A comprehensive overview of LLMs applied to the healthcare domain.

LLaVA-RLHF
Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

LLM Alignment Survey
A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

MentaLLaMA
An open-source LLM family specialized for interpretable mental-health analysis on social media.

Chain-of-Verification (CoVe)
Meta's Chain-of-Verification adds a "deliberation" step where the LLM fact-checks its own draft before finalizing.

The Rise and Potential of LLM-Based Agents
A comprehensive survey of LLM-based agents covering construction, capability, and societal implications.

Rewindable Auto-regressive INference (RAIN)
Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

Hallucination Survey (Early)
Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

RLAIF (Scaling RLHF with AI Feedback)
Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

Explaining Grokking
DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

Overview of AI Deception
A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

LLMs for Illicit Purposes
A survey cataloguing threats and vulnerabilities arising from LLM deployment.