🚀NEW LABGetting Started with Claude AgentsStart lab
Hands-on LabIntermediateFree

Jev-as-a-Judge: Evaluating Agents with a Decision Model

Evaluate agent tool choices, execution traces, and outcomes with Jev through OpenRouter.

4 labsabout 28 minUpdated Oct 2026

What you'll build

Inspect an agent's tool choices and execution trace
Score runs against a rubric with Jev as the judge
Check the judge before you trust its verdicts
Repair the agent's policy and measure the change

Curriculum

4 labs, about 28 min

  1. 1

    Inspect an agent run

    4 min

  2. 2

    Define a rubric and run Jev

    8 min

    Locked
  3. 3

    Check the judge

    8 min

    Locked
  4. 4

    Repair the agent and reevaluate

    8 min

    Locked

Your instructor

ES

Elvis Saravia

Founder, DAIR.AI

Elvis founded DAIR.AI and wrote the Prompt Engineering Guide, one of the most widely used references on working with language models. He has worked at companies like Meta AI and Elastic, and specializes in AI agents, RAG, and harness engineering.

How labs work

Live workspace

You run the real tools in a browser workspace. Nothing to install.

Checkpoints

Each lab checks your files and output, then unlocks the next one.

Your work stays

The workspace persists between labs so you finish with a project.

Questions