🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Aug 28 – Aug 28, 2026
Agents

Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI

First page
Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI
The curator’s take

Architect Labs report Redwood, a frontier inference accelerator whose performance model, RTL, UVM environments, formal proofs, firmware and kernels were generated end to end by an AI system in under two weeks from a specification written by two human architects.

Ask this paper

Key points
01

No human intervention below the spec: Two architects wrote a high-level specification; every artifact below it, including verification environments and formal proofs, was machine-generated. Every block reached 95% coverage via commercial EDA tools plus a proprietary formal engine and hardware-in-the-loop validation.

02

48-hour respin: Specification changes were reverified and redeployed to hardware in under 48 hours, which is the actual claim worth stress-testing: cycle time, not one-shot success.

03

Measured silicon numbers: Projected onto Samsung 8 nm, Redwood delivers 1.75x the throughput at 1.9x lower power against a measured Jetson Orin Nano baseline on the same models, a 3.4x performance-per-watt gain. Redwood Nano, the FPGA variant, runs multi-billion-parameter Llama and Qwen models.

04

A recursive-self-improvement footnote: Qwen running on Redwood helped design the next-generation Redwood, which the authors frame as an early RSI step rather than a completed loop.

05

Why it matters: Hardware design is the hardest possible test of long-horizon agentic verification, because the artifact is expensive, formally checkable, and physically deployed. Read the coverage methodology closely before accepting the headline.

Abstract

Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months. Design decisions are therefore committed under deep uncertainty and paid for twice, once in the generality added as a hedge, and again when new workloads map poorly onto frozen silicon. As Moore's Law stagnates, specialization is the main remaining source of performance-per-watt and demands a design cycle that runs at the cadence of the workloads. We present an end-to-end AI system that collapses the software-to-silicon stack into a single optimization loop, where hardware and software are co-designed and verified under one objective. Its first demonstration is Redwood, a frontier AI accelerator built for single-batch, low-power, ultra-low-latency inference for physical AI. From a high-level specification by two human architects, the system autonomously generated the performance model, RTL design, UVM environments, formal proofs, firmware, and kernels in under two weeks with no human intervention below the specification. Every block reached 95% coverage via commercial EDA tools, our proprietary formal engine, and hardware-in-the-loop validation. Specification changes were reverified and redeployed to hardware in under 48 hours. Redwood Nano, its ultra-low-power FPGA variant, runs multi-billion-parameter models like Llama and Qwen. Projected onto Samsung 8 nm, the Jetson Orin Nano's process class, Redwood delivers 1.75x the throughput at 1.9x lower power, a 3.4x performance-per-watt gain against a measured Jetson baseline on the same models. Qwen running on Redwood also helped design next-generation Redwood, an early step toward recursive self-improvement. To our knowledge, this is the first production-worthy AI accelerator designed end-to-end by an AI system and running a modern AI model.

Every Monday
Get next week’s papers.
Subscribe on Substack