🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents

Skill Lift

First page
Skill Lift
Paper summary

Enterprise teams reviewing shared skill libraries almost always gate on a scanner that checks structure, style, and security. NVIDIA measured whether that gate predicts anything about how a skill actually performs, and the answer is close to no.

Ask this paper

Key points
01

The review gate is nearly uncorrelated with quality: Across 145 real skills from internal and public catalogs, structural scan scores correlate with LLM-judge quality at a Spearman rho of 0.14. Passing the scanner tells you the skill is well formatted, nothing more.

02

Measure the delta, not the document: ACES proposes Skill Lift. Run the same task twice under the same model, sandbox, workspace, and scorer, once with the skill loaded and once without, then measure the difference in what the agent completed.

03

Results compare across harnesses: 947 paired cases from 58 production skills were scored across four harnesses, with trajectories normalized into a shared Agent Trajectory Interchange Format so a skill's lift in Claude Code can be read against its lift in Cursor.

04

Why it matters: The largest process-metric gains show up in skill execution, behavior check, and skill efficiency, which points at what skills are actually for. If you run a review process today, this gives you the paired-run design to replace it with something that measures outcomes.

Every Monday
Get next week’s papers.
Subscribe on Substack