0

Who Grades the Graders? Rethinking Verifier Design for Computer-Use Agents

https://labs.scale.com/blog/verifier-design-for-cua(labs.scale.com)
Verifiers used to evaluate and train Computer Use Agents (CUAs) often have significant flaws, such as all-or-nothing binary grading and brittle requirements found in benchmarks like OSWorld 2.0. These weaknesses lead to inaccurate model assessment and can be exploited through "reward hacking," where agents achieve high scores without correctly completing tasks. Two primary evaluation methods, programmatic checks and VLM Agent Judges, each have distinct strengths and weaknesses regarding determinism, cost, and scope. A more robust approach combines the strengths of both, using human-written rubrics to create deterministic checks for objective criteria and VLM judges for subjective aspects, while actively testing for and mitigating potential gaming strategies.
0 pointsby ogg1 hour ago

Comments (0)

No comments yet. Be the first to comment!

Want to join the discussion?