0
Measuring benchmark optimization in speech recognition
https://huggingface.co/blog/asr-benchmark-optimization(huggingface.co)High scores on public speech recognition benchmarks can be misleading, as models may be learning to game the test rather than truly improving at transcription. New research introduces three tests to measure this "benchmark optimization," revealing that some top-performing models reproduce known errors from popular datasets. In one striking example, several models transcribed an incorrect phrase to match a flawed benchmark transcript, even when the actual audio clearly said something different. These models often corrected their transcription only when the audio was re-recorded in a new voice, suggesting they rely on subtle acoustic cues to identify and overfit to specific benchmark data.
0 points•by will22•1 hour ago