Skip to main content
Reviewed as of 19 August 2026. Deepfake Detection is ranked 1st on Speech DF Arena, with an average EER of 1.104% across the arena’s 14 evaluation datasets.
Snapshot: 19 August 2026. Speech DF Arena accepts submissions continuously, so this standing is accurate on that date and not after. The live leaderboard is authoritative. Modulate’s row is listed under a legacy VELMA-2 identifier rather than the canonical model name.

Scores

Top four systems at the snapshot date. Lower is better in both EER columns. Detection accuracy is 100% - average EER. It restates the first column. The two EER columns order these systems differently. Resemble Detect 3B Omni posts the second-best pooled EER and the third-best average EER; Hiya is the reverse. Deepfake Detection leads on both, so one threshold holds across all 14 datasets. Pooled EER and average EER covers the distinction.

At an equal-error operating point

At the equal-error threshold the same rate applies in both error directions. Each figure below is simultaneously the missed synthetic calls per 1,000 synthetic calls and the false alarms per 1,000 genuine calls. This is arithmetic on the average EER column, not a separate measurement. The next-best system by average EER produces roughly 1.9 times as many errors at this operating point.
Worth knowing: production systems rarely run at the equal-error threshold. Moving the threshold to catch more synthetic calls raises false alarms on genuine ones, and the published EER does not describe that curve. Set the threshold against your own labelled recordings, as voice fraud screening describes.

Deployment characteristics

Minimum audio duration determines whether a model can gate a live authentication flow. Deepfake Detection returns a verdict per 192 ms frame once 2.5 seconds of speech has accumulated. Competitor prices are list prices at the snapshot date. The arena does not measure them.

What the arena does not measure

  • Streaming behaviour. The arena scores complete files. The per-frame latency of Deepfake Detection streaming is not part of this result.
  • Narrowband telephony end to end. ASVspoof 2021 LA applies telephony codecs. Most arena datasets are wideband.
  • Non-speech audio. Hold music, IVR prompts, and ringback are absent from every arena dataset. In production those frames return verdict: "no-content" and are excluded from scoring.
  • Languages beyond English and Mandarin.

Sources