🚀NEW LABGetting Started with Claude AgentsStart lab
← All papersIssue 33 of 182

The week of Nov 13 – Nov 19, 2023

10 papers, hand-picked and summarised.

Emu Video and Emu Edit

Emu Video and Emu Edit

Meta releases Emu Video and Emu Edit, a pair of diffusion models targeting controlled text-to-video generation and instruction-based image editing.

01Multimodal
Chain-of-Note (CoN)

Chain-of-Note (CoN)

Tencent's Chain-of-Note adds an explicit note-taking step to RAG so the model can evaluate retrieved evidence before answering.

02Retrieval
LLMs for Scientific Discovery

LLMs for Scientific Discovery

A broad evaluation of GPT-4 across scientific disciplines including drug discovery, biology, and computational chemistry.

03Evaluation
Fine-Tuning LLMs for Factuality

Fine-Tuning LLMs for Factuality

Stanford fine-tunes LLMs for factuality without any human labels by using automatically generated preference signals.

04Training
Contrastive Chain-of-Thought

Contrastive Chain-of-Thought

Proposes contrastive CoT prompting where models see both valid *and* invalid reasoning demonstrations to reduce reasoning errors.

05Reasoning
Survey on Language Models for Code

Survey on Language Models for Code

A comprehensive survey of LLMs for code covering 50+ models, 30+ evaluation tasks, and 500 related works.

06Code
JARVIS-1

JARVIS-1

An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

07Agents
Learning to Filter Context for RAG (FILCO)

Learning to Filter Context for RAG (FILCO)

CMU's FILCO improves RAG by training a dedicated model to filter retrieved contexts before they reach the generator.

08Retrieval
MART (Multi-round Automatic Red-Teaming)

MART (Multi-round Automatic Red-Teaming)

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

09Safety
LLMs Can Deceive Users (Trading Agent)

LLMs Can Deceive Users (Trading Agent)

Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

10Agents
Every Monday
Get next week’s papers.
Subscribe on Substack