🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papersIssue 50 of 176

The week of Mar 11 – Mar 17, 2024

10 papers, hand-picked and summarised.

SIMA

SIMA

DeepMind's Scalable Instructable Multiworld Agent (SIMA) is a generalist AI agent that follows natural-language instructions across nine commercial 3D video games like No Man's Sky, Teardown, Valheim, and Space Engineers.

01Agents
Retrieval Augmented Thoughts (RAT)

Retrieval Augmented Thoughts (RAT)

RAT augments chain-of-thought by iteratively rewriting each reasoning step using retrieved context, sharply reducing hallucination on long-horizon generation tasks.

02Retrieval
Quiet-STaR

Quiet-STaR

Quiet-STaR generalizes the Self-Taught Reasoner (STaR) so that a language model learns to generate internal rationales between every token, not just for explicit QA problems.

03Reasoning
Knowledge Conflicts for LLMs

Knowledge Conflicts for LLMs

A survey that maps the landscape of knowledge conflicts in LLMs, covering how they arise, how models behave under them, and how to mitigate them.

04Memory
Stealing Part of a Production Language Model

Stealing Part of a Production Language Model

The paper demonstrates the first practical attack that extracts the embedding-projection layer of production LLMs through their ordinary logit APIs.

05Safety
Branch-Train-MiX (BTX)

Branch-Train-MiX (BTX)

Meta's BTX produces a single Mixture-of-Experts LLM by first training specialized experts in parallel and then mixing them, sidestepping the high cost of training one big generalist.

06Training
LLMs Predict Neuroscience Results (BrainBench)

LLMs Predict Neuroscience Results (BrainBench)

BrainBench asks both LLMs and human experts to predict the outcomes of neuroscience experiments from their abstracts, and finds LLMs outperform experts.

07Evaluation
C4AI Command-R

C4AI Command-R

Cohere for AI releases Command-R, a 35B open-weight LLM tuned specifically for retrieval-augmented generation, tool use, and multilingual workflows.

08Retrieval
Is Cosine-Similarity Really About Similarity?

Is Cosine-Similarity Really About Similarity?

This paper argues that cosine similarity between learned embeddings does not always measure semantic similarity, and gives analytical examples where it produces arbitrary or non-unique values.

09Safety
MM1: Multimodal LLM Pre-training

MM1: Multimodal LLM Pre-training

Apple's MM1 paper runs extensive ablations on multimodal LLM pretraining choices and releases a family of models up to 30B parameters that set competitive MLLM pretraining benchmarks.

10Training
Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack