LLM-as-a-judge means using a large language model (LLM) to evaluate the work of another model. You give the judge an output, or a pair of outputs, along with the criteria that matter, and it scores, ranks, or compares them. Did the answer follow the policy? Is response A better than response B? Did the agent actually finish the task?
This approach is useful because manual review does not scale. A person can carefully read a few dozen outputs, but a judge model can apply the same criteria to thousands of runs, every time you change a prompt, a tool, or a model.
Jev is a new model built specifically for this kind of decision, which makes it a natural fit for the judge role. Using Jev as the judging LLM is what this tutorial means by Jev-as-a-Judge.
Judging agents adds a twist, because an agent can sound correct even when its work failed. Imagine a refund agent telling a customer, "Your refund has been processed," when the refund tool actually timed out.
Jev-as-a-Judge catches this by checking what the agent did, not just what it said. Jev reads the request, each tool call and its result, and the final reply, then returns a verdict with a probability attached.
This guide walks through the idea with one refund example and a live playground. The full lab builds the complete evaluation workflow.
