Jev-as-a-Judge: Evaluating Agents with a Decision Model
Evaluate agent tool choices, execution traces, and outcomes with Jev through OpenRouter.
$ open "Inspect an agent run"
checkpoints run as you work
✓ checkpoint passed, lab 2 unlocked
What you'll build
Curriculum
4 labs, about 28 min
- 1
Inspect an agent run
4 min
- 2
Define a rubric and run Jev
8 min
Locked - 3
Check the judge
8 min
Locked - 4
Repair the agent and reevaluate
8 min
Locked
Your instructor
Elvis Saravia
Founder, DAIR.AI
Elvis founded DAIR.AI and wrote the Prompt Engineering Guide, one of the most widely used references on working with language models. He has worked at companies like Meta AI and Elastic, and specializes in AI agents, RAG, and harness engineering.
How labs work
Live workspace
You run the real tools in a browser workspace. Nothing to install.
Checkpoints
Each lab checks your files and output, then unlocks the next one.
Your work stays
The workspace persists between labs so you finish with a project.