Speech DF Arena: 14 deepfake datasets
Speech DF Arena scores deepfake detection across 14 evaluation sets spanning synthesis families, languages, and channel conditions.
In The Wild is the only set drawn from deepfakes made to deceive rather than to populate a corpus. Codecfake is the only set whose audio starts as genuine speech, which defeats detectors keyed to synthesis artefacts.
Results: Deepfake Detection.
Open ASR Leaderboard: 8 transcription datasets
The Open ASR Leaderboard scores English transcription across eight corpora, applying one text normaliser to every entrant.
Earnings-22 and VoxPopuli are the two sets closest to contact-centre audio. Modulate reports that pair as a named subset alongside the full leaderboard result.
Results: Transcription.
What these datasets do not cover
- 8 kHz telephony throughout. ASVspoof 2021 LA applies telephony codecs and Earnings-22 is telephony-grade, but most sets in both benchmarks are wideband. A pipeline running narrowband PCM end to end is outside the measured distribution.
- Non-speech audio between speech. IVR prompts, hold music, ringback, and transfer tones are absent from every set listed here.
- Streaming. Both benchmarks score complete files. Neither measures partial-result latency or accuracy on a live socket.
- Languages beyond the sets listed. Speech DF Arena covers English and Mandarin. The Open ASR Leaderboard is English only. Neither speaks to the rest of the languages Multilingual Transcription accepts.
- Speaker diarization, emotion, accent, and PII/PHI. No public benchmark in this tab scores them. Modulate evaluates them internally.