🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papersIssue 43 of 176

The week of Jan 22 – Jan 28, 2024

10 papers, hand-picked and summarised.

Depth Anything

Depth Anything

A robust monocular depth estimator designed to handle "any image under any circumstance" by scaling self-training on unlabeled data rather than hunting for bigger labeled sets.

01Training
Knowledge Fusion of LLMs (FuseLLM)

Knowledge Fusion of LLMs (FuseLLM)

FuseLLM proposes fusing the capabilities of multiple existing LLMs into a single target model by distilling their output distributions rather than retraining from scratch.

02Training
MambaByte

MambaByte

MambaByte adapts the Mamba state-space architecture to learn directly from raw bytes, bypassing tokenization and all its well-known failure modes.

03Architecture
Diffuse to Choose

Diffuse to Choose

Amazon's Diffuse to Choose is a diffusion-based image-conditioned inpainting model built for "virtual try-on" scenarios where product images must be placed naturally into user scenes.

04Training
WARM (Weighted Averaged Reward Models)

WARM (Weighted Averaged Reward Models)

WARM averages multiple fine-tuned reward models in weight space rather than ensembling their predictions, dramatically reducing RLHF inference cost.

05Reinforcement Learning
Resource-efficient LLMs & Multimodal Foundation Models

Resource-efficient LLMs & Multimodal Foundation Models

A wide-ranging survey of efficiency techniques for LLMs and multimodal foundation models, spanning architecture, algorithms, and system design.

06Multimodal
Red Teaming Visual Language Models

Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

07Evaluation
Lumiere

Lumiere

Google's Lumiere is a space-time diffusion model for text-to-video that generates the entire video duration in a single forward pass rather than cascading short clips.

08Multimodal
Medusa

Medusa

Medusa accelerates LLM inference by bolting on multiple decoding heads that predict several future tokens in parallel, dramatically reducing decoding steps.

09Efficiency
AgentBoard

AgentBoard

AgentBoard is a benchmark and open-source evaluation framework for analytically evaluating LLM agents beyond the usual pass/fail metrics.

10Evaluation
Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack