JEVAL
benchmark // jev vs llm judges // gold-labelled

Same rubric. Same cases. Every judge.

Every judge answers the identical evaluator questions, built from the same rubric text jeval sends to Jev, over hand-labelled cases. We measure how often each judge matches the gold label, how well its confidence is calibrated, how consistent it is across repeats, and what it costs in time and money. No estimates: every number below is a real call.

accuracymatch with gold label (score levels within ±1)
briercalibration of P(yes); lower is better
consistencyidentical answers across repeats
latencyp50 / p95 per case, one call per case
costfrom returned token usage × list price
the datasetloading…
    click a category to include only some · labels are hand-written; disagree → PR
    judges
    repeats
    4 judges · 0 cases · 1 repeat · 0 calls
    provider keys
    Stored only in this browser and sent with your run. This deployment has no server keys; bring your own.