🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papersIssue 51 of 176

The week of Mar 18 – Mar 25, 2024

10 papers, hand-picked and summarised.

Grok-1

Grok-1

xAI open-sources Grok-1, a 314B-parameter Mixture-of-Experts base model, making it the largest openly released LLM at the time of publication.

01Training
Evolutionary Model Merge

Evolutionary Model Merge

Sakana AI proposes using evolutionary algorithms to automatically discover effective merges of open-source models, producing strong composite models without any additional training.

02Evaluation
TacticAI

TacticAI

Google DeepMind, in collaboration with Liverpool FC, releases TacticAI, a geometric deep-learning system that analyzes football corner kicks and suggests alternative tactics for coaches to explore.

03Retrieval
What Are Tools Anyway? A Survey of Tool Use in LLMs

What Are Tools Anyway? A Survey of Tool Use in LLMs

This survey establishes a formal definition of tools as "external programs used by LMs" and systematizes when, why, and how tool-use improves LLM performance.

04Agents
RankPrompt: Step-by-Step Comparisons Make LLMs Better Reasoners

RankPrompt: Step-by-Step Comparisons Make LLMs Better Reasoners

RankPrompt is a prompting method that lets an LLM self-rank its own candidate answers via chains of pairwise comparisons, without needing an external verifier or additional fine-tuning.

05Reasoning
LLM4Decompile

LLM4Decompile

LLM4Decompile is the first open-source family of LLMs specialized for decompiling machine code back into readable, re-executable C source.

06Evaluation
Agent-FLAN

Agent-FLAN

Agent-FLAN redesigns fine-tuning data so that open models can learn agentic skills without sacrificing general capability, hitting new open-source SoTA for Llama2-7B-based agents.

07Agents
Logits of API-Protected LLMs Leak Proprietary Information

Logits of API-Protected LLMs Leak Proprietary Information

The paper shows that the softmax bottleneck in modern LLMs means even logit-level APIs leak enough information to reconstruct hidden architectural details.

08Safety
DROID

DROID

DROID is an open-source robot manipulation dataset that dramatically expands the diversity of real-world robot demonstrations available for imitation-learning research.

09Robotics
RAFT: Retrieval-Augmented Fine-Tuning

RAFT: Retrieval-Augmented Fine-Tuning

RAFT is a fine-tuning recipe that teaches LLMs to handle distractor documents during RAG and to answer with CoT-style citations to retrieved passages.

10Retrieval
Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack