
Llama 2
Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

How is ChatGPT's Behavior Changing Over Time?
Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

FlashAttention-2
Tri Dao's follow-up to FlashAttention, dramatically improving attention throughput on modern GPUs.

Measuring Faithfulness in Chain-of-Thought Reasoning
Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

Generative TV & Showrunner Agents
Fable Studio's approach to generate episodic TV content using LLMs and multi-agent simulation.

Challenges & Application of LLMs
A comprehensive enumeration of open challenges and application domains for LLMs.

Retentive Network (RetNet)
Microsoft's proposed foundation architecture aiming to replace Transformer attention for LLMs.

Meta-Transformer
A unified framework performing learning across 12 different modalities with a shared backbone.

Retrieve In-Context Examples for LLMs
A framework to iteratively train dense retrievers that identify high-quality in-context examples.

FLASK
Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.