> ## Documentation Index
> Fetch the complete documentation index at: https://docs.modulate.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Datasets

> Every public dataset behind Modulate's published benchmark results, the attack family or acoustic condition each one stresses, and the conditions none of them cover.

*Reviewed as of 19 August 2026.*

Modulate's published results are measured on the datasets its benchmarks define. Neither corpus set is chosen by Modulate.

## Speech DF Arena: 14 deepfake datasets

Speech DF Arena scores deepfake detection across 14 evaluation sets spanning synthesis families, languages, and channel conditions.

| Dataset            | Attack family                              | Condition it stresses                                                                                                                                  |
| ------------------ | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| ASVspoof 2019      | Text-to-speech and voice conversion        | Studio-clean synthesis. The long-standing baseline, and the set most detectors are tuned against.                                                      |
| ASVspoof 2021 LA   | Text-to-speech and voice conversion        | The 2019 attacks after transmission through telephony codecs. Separates codec robustness from detection ability.                                       |
| ASVspoof 2021 DF   | Compressed synthetic speech                | Varied lossy encoders and bitrates, as media re-encoding would apply.                                                                                  |
| ASVspoof 2024 Eval | Crowdsourced and adversarial synthesis     | Newer generation methods over non-studio source recordings.                                                                                            |
| Fake or Real       | Commercial text-to-speech                  | Synthesis from deployed commercial systems, across mixed recording conditions.                                                                         |
| Codecfake          | Neural audio codec resynthesis             | Speech reconstructed through a learned codec rather than a vocoder. Genuine speech enters the pipeline, so there are no synthesis artefacts to key on. |
| ADD 2022 Track 1   | Mandarin full-utterance fakes              | Low-quality and noisy Mandarin audio.                                                                                                                  |
| ADD 2022 Track 3   | Adversarial Mandarin fakes                 | Attacks constructed to evade detection rather than to sound natural.                                                                                   |
| ADD 2023 Round 1   | Mandarin deepfake detection                | Second-edition challenge audio, first evaluation round.                                                                                                |
| ADD 2023 Round 2   | Mandarin deepfake detection                | Second evaluation round of the same challenge.                                                                                                         |
| DFADD              | Diffusion and flow-matching text-to-speech | Generation families that postdate most detectors' training data.                                                                                       |
| LibriVoc           | Vocoder artefacts                          | Multiple vocoders over read speech, isolating the vocoder as the only synthetic signal.                                                                |
| SONAR              | Recent end-to-end text-to-speech           | Current-generation systems, including synthesis with no separate vocoder stage.                                                                        |
| In The Wild        | Real-world deepfakes                       | Deepfaked speech of public figures collected from social media. No controlled synthesis pipeline and no matching training distribution.                |

In The Wild is the only set drawn from deepfakes made to deceive rather than to populate a corpus. Codecfake is the only set whose audio starts as genuine speech, which defeats detectors keyed to synthesis artefacts.

Results: [Deepfake Detection](/benchmarks/deepfake-detection).

## Open ASR Leaderboard: 8 transcription datasets

The Open ASR Leaderboard scores English transcription across eight corpora, applying one text normaliser to every entrant.

| Dataset           | Speech type                         | Condition it stresses                                                                                                     |
| ----------------- | ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| AMI               | Multi-party meetings                | Far-field microphones, overlapping speech, spontaneous turn-taking. The hardest set on the leaderboard for every entrant. |
| Earnings-22       | Earnings calls                      | Telephony-grade audio, accented English, dense financial vocabulary and named entities.                                   |
| GigaSpeech        | Mixed podcast, audiobook, and video | Broad-domain spontaneous speech across recording qualities.                                                               |
| LibriSpeech clean | Read audiobooks                     | Well-recorded read speech. Near-saturated across entrants.                                                                |
| LibriSpeech other | Read audiobooks                     | The harder speaker split of the same corpus.                                                                              |
| SPGISpeech        | Financial calls                     | Long-form professional speech with heavy domain vocabulary.                                                               |
| TED-LIUM          | Conference talks                    | Prepared single-speaker delivery, varied accents.                                                                         |
| VoxPopuli         | European Parliament proceedings     | Non-native English accents and parliamentary register.                                                                    |

Earnings-22 and VoxPopuli are the two sets closest to contact-centre audio. Modulate reports that pair as a named subset alongside the full leaderboard result.

Results: [Transcription](/benchmarks/transcription).

## What these datasets do not cover

* **8 kHz telephony throughout.** ASVspoof 2021 LA applies telephony codecs and Earnings-22 is telephony-grade, but most sets in both benchmarks are wideband. A pipeline running narrowband PCM end to end is outside the measured distribution.
* **Non-speech audio between speech.** IVR prompts, hold music, ringback, and transfer tones are absent from every set listed here.
* **Streaming.** Both benchmarks score complete files. Neither measures partial-result latency or accuracy on a live socket.
* **Languages beyond the sets listed.** Speech DF Arena covers English and Mandarin. The Open ASR Leaderboard is English only. Neither speaks to the rest of the languages [Multilingual Transcription](/get-started/stt) accepts.
* **Speaker diarization, emotion, accent, and PII/PHI.** No public benchmark in this tab scores them. Modulate evaluates them internally.
