
Llama 2
Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

How is ChatGPT's Behavior Changing Over Time?
Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

FlashAttention-2
Tri Dao's follow-up to FlashAttention, dramatically improving attention throughput on modern GPUs.

Measuring Faithfulness in Chain-of-Thought Reasoning
Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

Generative TV & Showrunner Agents
Fable Studio's approach to generate episodic TV content using LLMs and multi-agent simulation.

Challenges & Application of LLMs
A comprehensive enumeration of open challenges and application domains for LLMs.

Retentive Network (RetNet)
Microsoft's proposed foundation architecture aiming to replace Transformer attention for LLMs.

Meta-Transformer
A unified framework performing learning across 12 different modalities with a shared backbone.

Retrieve In-Context Examples for LLMs
A framework to iteratively train dense retrievers that identify high-quality in-context examples.

FLASK
Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack