0
Who Grades the Graders? Rethinking Verifier Design for Computer-Use Agents
https://labs.scale.com/blog/verifier-design-for-cua(labs.scale.com)Verifiers used to evaluate and train Computer Use Agents (CUAs) often have significant flaws, such as all-or-nothing binary grading and brittle requirements found in benchmarks like OSWorld 2.0. These weaknesses lead to inaccurate model assessment and can be exploited through "reward hacking," where agents achieve high scores without correctly completing tasks. Two primary evaluation methods, programmatic checks and VLM Agent Judges, each have distinct strengths and weaknesses regarding determinism, cost, and scope. A more robust approach combines the strengths of both, using human-written rubrics to create deterministic checks for objective criteria and VLM judges for subjective aspects, while actively testing for and mitigating potential gaming strategies.
0 points•by ogg•1 hour ago