🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papersIssue 39 of 176

The week of Dec 25 – Dec 31, 2023

10 papers, hand-picked and summarised.

CogAgent

CogAgent

Tsinghua's CogAgent is an 18B-parameter visual-language model purpose-built for GUI understanding and navigation, with unusually high input resolution.

01Evaluation
From Gemini to Q-Star

From Gemini to Q-Star

A 300+-paper survey mapping the state of Generative AI and the research frontiers that followed the Gemini + rumored Q* news cycle.

02Multimodal
PromptBench

PromptBench

A unified library for comprehensive evaluation and analysis of LLMs that consolidates multiple evaluation concerns under one roof.

03Evaluation
Exploiting Novel GPT-4 APIs

Exploiting Novel GPT-4 APIs

A red-team study of three newer GPT-4 API surfaces - fine-tuning, function calling, and knowledge retrieval - that reveals each introduces new attack vectors.

04Training
Fact Recalling in LLMs

Fact Recalling in LLMs

A mechanistic-interpretability study showing that early MLP layers function as a lookup table for factual recall.

05Safety
Generative AI for Math (OpenWebMath / MathPile)

Generative AI for Math (OpenWebMath / MathPile)

Releases a diverse, high-quality math-centric corpus of ~9.5B tokens designed for training math-capable foundation models.

06Reasoning
Principled Instructions Are All You Need

Principled Instructions Are All You Need

Distills effective LLM prompting into 26 guiding principles and validates them across multiple model families.

07Training
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

08Reasoning
LLaRA

LLaRA

LLaRA adapts a decoder-only LLM for dense retrieval via two tailored pretext tasks that leverage text embeddings from the LLM itself.

09Retrieval
Gemini vs GPT-4V

Gemini vs GPT-4V

A qualitative side-by-side comparison of Gemini and GPT-4V across vision-language tasks, documenting systematic behavioral differences.

10Multimodal
Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack