Articles

JEV-as-a-Judge for Agent Evals

JEV-as-a-judge is TypeSafe AI's typed evaluator for agent evals. See how it compares to LLM-as-a-judge, where it wins, where it breaks, and when to use each.

· Updated
· 12 min read
agent-evals llm-as-a-judge jev typed-judge semantic-verifier
Comparison of a typed judge, an LLM-as-a-judge, and a code-based check scoring an agent output for agent evals
Table of Contents

Every team building agents hits the same wall: you can generate a thousand agent runs a day, but you cannot read a thousand transcripts a day. Something has to grade them for you, and that grader becomes the most important part of your pipeline.

The default grader today is LLM-as-a-judge: you ask a strong model to score another model’s output against a rubric. It is flexible and it reads like a human, but it is also slow, expensive, and noisy, giving you a slightly different score each time you run it.

So when a typed alternative showed up, teams paid attention. JEV-as-a-judge is TypeSafe AI’s take on the problem: instead of writing a paragraph of reasoning, it returns a typed verdict with a probability attached.

This post is a practitioner’s read on where that helps: what JEV actually is, how it sits next to LLM-as-a-judge and plain code checks, the benchmarks worth trusting, the failure modes to watch, and a simple framework for picking the right method for each thing you grade.

TL;DR

  • JEV-as-a-judge is TypeSafe AI’s typed evaluator. It returns a verdict (a yes/no, a choice, or a score) plus a calibrated probability, with no written reasoning.
  • In the September 29, 2026 revision of the JEV-as-a-Judge paper, JEV recorded a 0.15-second median latency and cost 0.36% of the strongest comparator’s fee on the tested workloads.
  • It came within three percentage points of GPT-6 when a verdict could be read directly from supplied text, but fell further behind on math, code, logic, and other derivation-heavy tasks.
  • Typed and generative judges both process untrusted content. Treat their inputs as untrusted and escalate uncertain or security-sensitive decisions.
  • Choose the method by what you are grading: code checks for formats, semantic verifiers for grounding, LLM judges for nuance, typed judges for the high-volume gate. Most production agents need several at once.

What Is JEV-as-a-Judge?

JEV is an evaluator model from TypeSafe AI, released in September 2026 and framed by its makers as a “System One” model: fast, instinctive scoring rather than slow deliberation. The practical difference from a chat model is the output.

A normal LLM judge may return a rationale and a score. JEV returns a typed verdict, a yes/no, a multiple-choice pick, or a point on an ordered scale, along with a calibrated probability for that verdict. It does not return a written rationale.

That design choice is the whole pitch. Because the model is trained to emit a decision instead of prose, it runs fast and cheap and gives you close to the same answer every time.

The latest Carnegie Mellon JEV-as-a-Judge paper reports a 0.15-second median latency and a fee equal to 0.36% of its strongest comparator on the tested workloads. Those figures describe the paper’s benchmark setup, not a universal latency or cost guarantee. LangChain lists it in LangSmith Evals as jev-latest and frames typed judges as a category alongside code-based checks and LLM-as-a-judge.

The trade you are making is reasoning for repeatability. An LLM judge hands you a paragraph you can read and argue with. A typed judge hands you a number you can gate on. For a high-volume regression suite that runs on every commit, that repeatability is often worth more than the prose.

The probability attached to each verdict is not decoration. Because it is calibrated, you can act on it: accept the verdict when confidence is high, and route the low-confidence cases to a stronger judge or a human reviewer.

The Carnegie Mellon research on this confidence-cascade approach is titled around exactly that idea, accept when confident and escalate when unsure, which turns a single judge into a triage step rather than a final word.

It is also a different tool from agent-as-a-judge, which runs a full agent to grade another agent. JEV is a single decision-model call, not an agent.

JEV vs LLM-as-a-Judge vs Code-Based Evals

It helps to see all three side by side, because they are not competitors so much as points on a spectrum from rigid to flexible. A code-based check is a rule you write: a regex, a JSON-schema validation, an exact-match assertion. It is free, instant, and perfectly repeatable, and it only works when correctness is something you can spell out in advance.

At the other end, an LLM judge handles open-ended cases a rule cannot, at the cost of additional latency, money, and potential variance. Typed judges like JEV and embedding-based semantic verifiers sit in the middle: more constrained than a generative judge and more flexible than a hardcoded rule. This is the same deterministic vs LLM-judge axis, with two additional options between the extremes.

The axis that matters most in production is variance. A code check and a semantic verifier return the same result on the same input every time, so a failing case is a real regression rather than noise. An LLM judge can score the same answer 0.7 today and 0.6 tomorrow, which means a red build might just be the judge changing its mind.

A typed judge sits close to the deterministic end here, and that stability is the main reason teams reach for it on a gate that has to hold steady across thousands of runs.

Three categories of agent eval: a code-based check, an LLM-as-a-judge, and a typed judge scoring the same agent output

MethodOutputReasoning traceVarianceCost and latencyBest for
Code-based checkPass or fail from a ruleNone, fully deterministicZeroLowestExact formats, schemas, known strings
LLM-as-a-judgeScore plus a written reasonFull text rationaleHighHighestNuanced quality, tone, open-ended answers
Typed judge (JEV)Typed verdict plus a probabilityNone, typed output onlyLowLowHigh-volume, repeatable pass/fail gates
Semantic verifierSimilarity score plus pass or failNone, embedding-basedLowLowFactual grounding against a reference

The lesson from the table is that no single row wins. The right eval suite usually uses several rows at once, one per kind of check.

Where Typed Judges Win, and Where They Break

The September 29 revision of the Carnegie Mellon paper reports a clearer boundary than one headline benchmark can capture. JEV came within three percentage points of GPT-6 when a verdict could be read directly from the supplied text, at 0.36% of the comparator’s fee and a 0.15-second median latency in that setup.

The gap widened when the verdict had to be derived through math, code, logic, or other multi-step reasoning. The paper’s strongest deployment pattern is therefore a cascade: accept sufficiently confident JEV verdicts and escalate uncertain cases to a reasoning judge or human reviewer.

Typed judges fit well-defined, high-volume first-pass checks. Subtle or derivation-heavy judgments still need a reasoning judge or human, and teams should validate confidence thresholds on their own labeled data before using them as gates.

The prompt-injection problem

Typed and generative judges both read content they grade, so that content can carry adversarial instructions. A written rationale may help an investigation, but it does not reliably reveal prompt injection, and a typed verdict provides less diagnostic context when something goes wrong.

If a judge influences a security or moderation decision, treat its input as untrusted. Isolate the evaluation instructions, test adversarial examples, use deterministic controls where the policy can be stated exactly, and route uncertain cases to human review. That caution is the same reason LLM-as-a-judge works best when it is calibrated and monitored rather than trusted blindly.

What a Semantic Verifier Actually Checks

Before you reach for any judge model, there is a cheaper option that covers a surprising number of cases: a semantic verifier. Instead of asking a model to reason, you embed both the agent’s answer and a reference answer, then measure how close their meanings are. If the cosine similarity clears a threshold, it passes.

This catches the common situation where the agent is right but phrased differently from your reference, which an exact-match rule would wrongly fail.

# Code 1 - semantic verifier: does the answer match the reference meaning?
from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("all-MiniLM-L6-v2")

def semantic_match(answer: str, reference: str, threshold: float = 0.75) -> dict:
    emb = model.encode([answer, reference], convert_to_tensor=True)
    score = util.cos_sim(emb[0], emb[1]).item()
    return {"score": round(score, 3), "passed": score >= threshold}

def _demo():
    r = semantic_match("The capital of France is Paris.", "Paris is France's capital.")
    assert r["passed"] is True

if __name__ == "__main__":
    _demo()

The all-MiniLM-L6-v2 model is small and can run locally, making this check relatively inexpensive at high volume, although latency and infrastructure cost depend on the deployment. The _demo proves the narrower point: two sentences that say the same thing in different word order clear the threshold.

Two things decide whether this check earns its keep. The first is the threshold. Set it too low and paraphrases of the wrong answer slip through; set it too high and correct-but-reworded answers fail. Tune it against a handful of known-good and known-bad pairs instead of guessing.

The second is a blind spot worth knowing about: embedding similarity measures topical closeness, not truth, so “Paris is the capital of France” and “Paris is not the capital of France” score as very similar even though one is wrong.

For anything where a negation or a single flipped fact changes the verdict, pair the verifier with a rule or a judge that can catch it. A semantic verifier is best described as measuring agreement with a trusted reference, not factual correctness; similarity alone cannot establish that a claim is true.

Choosing an Eval Method for Your Agent Evals

Put the pieces together and method selection stops being a matter of taste. It comes down to one question: what are you actually grading? Match the check to the answer type and each method lands in its natural place.

Take a support agent as a worked example. Whether its reply is valid JSON with the required fields is a format question, so a code check owns it. Whether the answer matches the known resolution for that ticket is grounding, so a semantic verifier handles it cheaply.

Whether the tone stayed polite and on-brand is a judgment call, so an LLM judge earns its cost there. And the nightly run across ten thousand past tickets that has to stay stable is the high-volume gate, which is where a typed judge fits. One agent, four checks, four different methods, each picked by the answer type rather than by habit.

Decision framework mapping what you are grading to the recommended eval method for agent evals

What you are gradingRecommended methodWhyMain failure mode to watch
Exact format or schemaCode-based checkDeterministic, instant, freeFails answers that are right but formatted differently
Semantic agreement with a trusted referenceSemantic verifierCheap meaning match to a referenceHigh similarity does not prove factual correctness and may miss negation
Nuanced quality or toneLLM-as-a-judgeHandles open-ended, subjective judgmentVariance and cost; calibrate against human labels
High-volume first-pass regression checkTyped judge (JEV)Fast, low-cost typed decisions at scaleValidate confidence thresholds and escalate uncertain cases

Most production agents need three or four of these running together, and that is the point where an ad hoc script stops scaling and you want an agent eval harness to hold them. The harness runs every check against every case and rolls the results into one number you can gate a deploy on.

Running an LLM-Judge Eval in Python

For the nuanced cases a rule cannot cover, here is the LLM-judge pattern in its simplest honest form: a rubric, a model call constrained to JSON, and a threshold that turns the score into a pass or fail.

# Code 2 - LLM-as-a-judge: rubric scoring that returns a score + reason
import json
from openai import OpenAI

client = OpenAI()

RUBRIC = (
    "Score from 0 to 1 how fully the answer resolves the user's request. "
    'Return strict JSON: {"score": <float 0-1>, "reason": "<one sentence>"}.'
)

def llm_judge(question: str, answer: str) -> dict:
    resp = client.chat.completions.create(
        model="gpt-4o",
        response_format={"type": "json_object"},
        messages=[
            {"role": "system", "content": RUBRIC},
            {"role": "user", "content": f"Question: {question}\nAnswer: {answer}"},
        ],
    )
    verdict = json.loads(resp.choices[0].message.content)
    verdict["passed"] = verdict["score"] >= 0.7
    return verdict

The response_format set to a JSON object is what keeps this reliable: the model must return parseable JSON, so you get a score and a reason you can log rather than free text you have to scrape.

The written reason is the audit trail a typed judge does not give you, which is exactly why you accept the extra cost here and save the typed judge for the high-volume gate.

One structural choice shapes the rest: pointwise or pairwise. The code above is pointwise, scoring one answer on its own, which is simple, cheap, and what you want for a threshold gate.

Pairwise judging, asking the model which of two answers is better, tends to track human preference more closely, but it carries a known position bias where the model favors whichever answer it reads first, so you run both orders and average them.

Start pointwise, and move to pairwise only for the rankings where the extra agreement is worth the extra calls. Two cautions carry over from every LLM-judge study either way: the score has variance, so average across runs on the cases that matter, and calibrate the rubric against a small set of human labels before you trust the threshold.

Where Future AGI Fits for Agent Evals

The decision framework above, several methods running together and rolling into one gate, is exactly what Future AGI’s Evaluation product is built to run.

Rather than committing to a single judge, you write a custom eval in plain language, choose whether it grades with an LLM-as-judge or a deterministic code check, pick the dataset columns to grade, and set the pass or fail threshold that gates your pipeline.

Each run returns a 0 to 1 score with a reason, checked against that threshold, so the output is a decision you can act on rather than a transcript to read.

For agent work specifically, built-in evaluators cover the checks teams reach for most: Groundedness, Detect Hallucination, Context Adherence, Task Completion, Instruction Adherence, and Conversation Resolution. You can run evals on your agents across a dataset and wire the whole suite in as a CI gate, mixing code checks, semantic checks, and LLM judges in one place.

Future AGI ships an open-source SDK at its core, so you can start from the evaluators above and add your own without rebuilding the harness.

Picking the Right Judge for Your Agents

The bottleneck we opened with, too many agent runs and no way to read them all, does not have a single-tool answer. No judge wins every category.

A code check is the strongest default for exact formats, a semantic verifier cheaply measures agreement with a reference, an LLM judge handles nuance, and a typed judge like JEV is a fast first pass for high-volume checks when uncertain cases are escalated.

Reliable agents come from matching each check to the right method and running the whole mix automatically on every change. Pick the one output type your agent gets wrong most often, choose its method from the table above, and write that first custom eval today. That single gate is where trustworthy agent evals start.

Frequently Asked Questions

What is JEV-as-a-judge?

JEV-as-a-judge is TypeSafe AI's typed evaluator, a System One model released in September 2026. Instead of writing a rationale, it returns a typed verdict, a yes/no, a choice, or a point on an ordered scale, with a calibrated probability attached. That makes it useful as a fast first-pass judge for high-volume agent evals, with uncertain or reasoning-heavy cases escalated to a stronger judge or human reviewer.

What is LLM-as-a-judge?

LLM-as-a-judge uses a strong language model to score another model's output against a rubric, usually returning a number and a short written reason. It scales human-style judgment to volumes no person could read, which made it the default grader for agent evals. The costs are latency, money, and variance: the same answer can score 0.7 on one run and 0.6 on the next, so you average across runs and calibrate the rubric against human labels before trusting it.

What is an alternative to LLM-as-a-judge?

Three alternatives cover most cases. A deterministic code check (a regex, a schema validation, an exact match) is repeatable when you can state correctness in advance. An embedding-based semantic verifier measures agreement with a trusted reference, so it can pass paraphrases that exact matching rejects, but similarity alone does not prove truth. A typed judge like JEV returns a verdict and probability at high speed and low cost. Most suites combine several methods.

What is a semantic verifier?

A semantic verifier checks whether an answer matches a reference meaning by embedding both and measuring how close they are, rather than comparing exact text. It passes the common case where the agent is right but phrased differently from your reference, which an exact-match rule would wrongly fail. It runs locally in milliseconds, so it costs almost nothing per case. But similarity measures topical closeness, not truth, so pair it with a rule or a judge wherever one flipped fact changes the verdict.

How accurate is LLM-as-a-judge?

On clear rubrics, LLM-as-a-judge often reaches roughly human-level agreement, which is why it works well for tone, quality, and other open-ended checks. It drifts on nuanced, argue-both-sides cases and on adversarial inputs, and its scores carry run-to-run variance. Accuracy depends on the rubric wording, the judge model, and how well you calibrate against a small set of human labels. Treat it as a monitored instrument rather than a source of truth, and average across runs on the cases that matter.
Related Articles
View all