0
The Benchmark Behind the Benchmark
https://browser-use.com/posts/benchmark-behind-the-benchmark(browser-use.com)The quality of an AI benchmark depends less on its tasks and more on its "judge," the often-overlooked system that grades performance. A weak judge introduces significant "noise," where the same completed task can receive wildly different scores, sometimes varying more than the performance gap between the models being tested. To create a more stable and reliable benchmark, the judge should be an agent that actively investigates evidence and provides a nuanced percentage score instead of a binary pass/fail. Further stability is achieved by using an ensemble of multiple judges and taking their median score, which effectively cancels out random errors and disagreements on borderline cases.
0 points•by chrisf•1 hour ago