
A Survey on Evaluation of LLMs
A comprehensive overview of evaluation methods covering what, where, and how to evaluate LLMs.

How Language Models Use Long Contexts (Lost-in-the-Middle)
Shows LLM performance drops when relevant information is in the middle of a long context.

LLMs as Effective Text Rankers
A prompting technique that enables open-source LLMs to perform SOTA text ranking.

Multimodal Generation with Frozen LLMs
Maps images to LLM token space enabling models like PaLM and GPT-4 to handle visual tasks without parameter updates.

CodeGen2.5
Salesforce's new 7B code LLM trained on 1.5T tokens and optimized for fast sampling.

Elastic Decision Transformer
An advance over Decision Transformers that enables trajectory stitching at inference time.

Robots That Ask for Help
A framework for calibrating LLM-based robot planners so they ask for help when uncertain.

Physics-based Motion Retargeting in Real-Time
Uses RL to retarget motions from sparse human sensor data to characters of various morphologies.

Scaling Transformer to 1 Billion Tokens (LongNet)
Microsoft's Transformer variant scaling sequence length past 1B tokens.

InterCode
A framework treating interactive coding as a reinforcement learning environment.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack