benchmark // jev vs llm judges // gold-labelled
Same rubric. Same cases. Every judge.
Every judge answers the identical evaluator questions, built from the same rubric text jeval sends to Jev, over hand-labelled cases. We measure how often each judge matches the gold label, how well its confidence is calibrated, how consistent it is across repeats, and what it costs in time and money. No estimates: every number below is a real call.
accuracymatch with gold label (score levels within ±1)
briercalibration of P(yes); lower is better
consistencyidentical answers across repeats
latencyp50 / p95 per case, one call per case
costfrom returned token usage × list price
the datasetloading…
judges
repeats
4 judges · 0 cases · 1 repeat · 0 calls
provider keys
Stored only in this browser and sent with your run. This deployment has no server keys; bring your own.