🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papersIssue 41 of 176

The week of Jan 8 – Jan 14, 2024

10 papers, hand-picked and summarised.

InseRF

InseRF

InseRF inserts brand-new 3D objects into Neural Radiance Field scenes from just a text prompt plus a 2D bounding box, without requiring any explicit 3D input.

01Training
Sleeper Agents

Sleeper Agents

Anthropic shows that LLMs can be trained to act deceptively under specific triggers and that current safety training techniques fail to remove this hidden behavior.

02Safety
Blending Is All You Need

Blending Is All You Need

Small chat models (6B/13B) blended together can rival ChatGPT-class systems, without any new training.

03Evaluation
MagicVideo-V2

MagicVideo-V2

ByteDance's MagicVideo-V2 is an end-to-end text-to-video pipeline that stitches together four specialized modules into a high-fidelity generation system.

04Multimodal
TrustLLM (Trustworthiness in LLMs)

TrustLLM (Trustworthiness in LLMs)

A 100+ page study that defines a principled framework for trustworthy LLMs and benchmarks 16 mainstream models across it.

05Evaluation
Chain-of-Table

Chain-of-Table

Google's Chain-of-Table prompts LLMs to iteratively transform a complex table step-by-step to answer questions reliably, extending CoT reasoning to tabular data.

06Reasoning
Persuasive Adversarial Prompts (PAP)

Persuasive Adversarial Prompts (PAP)

Turns 40 human-persuasion techniques into a taxonomy of jailbreaks that achieve 92% attack success on frontier models without any optimization.

07Safety
RAISE

RAISE

RAISE is an advanced agent architecture that adds a dual-memory system on top of a ReAct-style backbone to better support long-running conversational agents.

08Memory
Quantifying Prompt-Format Sensitivity

Quantifying Prompt-Format Sensitivity

CMU researchers show that LLM few-shot performance is shockingly sensitive to superficial prompt-formatting choices.

09Evaluation
Adversarial Machine Learning (NIST)

Adversarial Machine Learning (NIST)

NIST's official taxonomy of adversarial machine learning, intended to standardize terminology for policy and practice.

10Evaluation
Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack