🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 1, 2026
Agents

AI as a Compiler: Compiling Triton kernels without the Triton compiler

First page
AI as a Compiler: Compiling Triton kernels without the Triton compiler
The curator’s take

Francois Costa, Azalia Mirhoseini and colleagues at Stanford (with EPFL) test AI lowering, in which an LLM agent translates Triton kernels directly into NVIDIA PTX instead of running the Triton compiler pipeline.

Ask this paper

Key points
01

Setup. An agentic harness translates Triton to PTX and an evaluation environment checks each candidate for correctness and speed.

02

Results. Across twelve common kernels on Ada, Hopper and Blackwell and ten kernels from recent ML papers, AI-lowered PTX runs at 0.83x to 3.34x the speed of autotuned Triton.

03

Where the speedups come from. Transformations Triton does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), giving each thread a full softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x).

04

Verification. The Volta PTX verifier is extended to Blackwell's tcgen05 Tensor Core interface, which requires modeling tensor memory, descriptor-based operand layouts, and asynchronous commits, waits, barriers and proxy fences.

Abstract

Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.

Every Monday
Get next week’s papers.
Subscribe on Substack