Word error rate
Published price per hour
List prices as of 19 August 2026. The leaderboard does not measure them.Batch
Streaming
Worth knowing: these are prices per audio hour and exclude enrichment flags.
speaker_diarization, time_stamps, and PII/PHI tagging each change cost and latency, as Transcription documents per flag.Reading WER against your own requirement
WER counts substitutions, insertions, and deletions equally, with no notion of which words carry the decision. A transcript that renders a sentence correctly except the account number scores better than one that drops two filler words. An application that reads names, amounts, or identifiers out of the transcript needs accuracy measured on those spans. Normalisation determines comparability. The leaderboard applies one normaliser to every entrant. Figures taken from two vendors’ own marketing were normalised differently, and this tab does not mix them.Sources
- Open ASR Leaderboard, Hugging Face. Live table and evaluation code.
- Earnings-21: A Practical Benchmark for ASR in the Wild, arXiv:2104.11348. Background on the earnings-call corpora.
- ESB: A Benchmark For Multi-Domain End-to-End Speech Recognition, arXiv:2210.13352. Dataset selection and normalisation rationale.