FFASR Leaderboard: Open Benchmark for Far-Field ASR

Treble Technologies and Hugging Face launched the FFASR Leaderboard, the first open community-driven benchmark for far-field automatic speech recognition (ASR) evaluation.

It is designed to quantify the gap between near-field clean-speech performance and real-world far-field conditions caused by reverberation, background noise, and microphone distance.

The leaderboard evaluates models across 14 simulated rooms (20–470 m³) including bathrooms, offices, classrooms, and restaurants, using Treble’s hybrid simulation engine that combines wave-based and geometrical-acoustics modeling.

The primary ranking score uses four conditions: near-field dry, far-field high SNR (>14 dB), far-field mid SNR (8–12 dB), and far-field low SNR (<6 dB). Additional tracks include lab-measured vs.

simulated validation and moving-source splits in beta.

All models report word error rate (WER) alongside RTFx (audio seconds per inference second) on NVIDIA L4 GPUs, with a Pareto front visualization in the Analysis tab.

Early data shows far-field WER at low SNR is consistently several times higher than near-field WER on the same speech content.

The benchmark supports Whisper, IBM Granite Speech, Cohere Transcribe, Wav2Vec2, HuBERT CTC heads, and SpeechBrain via model ID submission on Hugging Face, with a custom evaluator option for complex pipelines.

The held-out test set uses 2,000 anechoic speech samples across conditions (about 8 hours per condition) with consistent Whisper-style text normalization, and audio is not exposed to submitters to avoid contamination.

Future roadmap includes multi-talker scenarios, microphone array support, and echo cancellation.

Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World

View Original