🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Agents

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

First page
D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?
The curator’s take

Daifeng Li (HKUST and Alibaba) with Huiqiang Jiang, Dayiheng Liu and colleagues at Alibaba Group introduce D2K-Bench, which measures how well coding agents turn expert design guidance into efficient Triton GPU kernels.

Ask this paper

Key points
01

Benchmark. 26 tasks and 85 workloads, each paired with an expert reference and guidance at three levels: high-level algorithmic insight, dataflow design and low-level optimization tricks, with dependencies between levels recorded.

02

Paired protocol. Each model solves every task twice, with and without guidance, using the same task text, workloads, tools, B200 hardware and a 350-turn budget. An LLM judge scores which design properties the source code implements, with runtime hidden.

03

Results. Across five models including GPT-6-Astra, Claude-Opus-4.8 and GPT-5.6-Sol, guidance raises correctness over 130 model-task pairs from 93.1% to 98.5% and the Performance Score from 1.46 to 1.95, a 33.9% increase.

04

Remaining gap. The mean implementation score rises only from 57 to 70 out of 100, so agents given the expert design still leave assessed properties unimplemented. Guidance raises API cost by 29.9%, almost all from cached input tokens.

Abstract

GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2: dataflow design, and L3: low-level optimization tricks, including dependencies among these levels. Pairwise runs with and without guidance share task descriptions, workloads, tools, hardware, and a 350-turn budget. Complementary assessments examine independently proposed designs and the design properties implemented in generated code. Across five models on NVIDIA B200 GPUs, guidance raises correctness over 130 model-task pairs from 93.1% to 98.5% and increases the Performance Score over all 26 tasks from 1.46 to 1.95. For the three frontier models with correct submissions on all 26 tasks in both runs (GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol), geometric mean speedup increases from $1.69\times$ to $2.49\times$. Across all five models, the mean combined implementation score increases from 57 to 70 out of 100. These results show the value of expert design guidance while identifying design properties that remain unimplemented.

Every Monday
Get next week’s papers.
Subscribe on Substack