Skip to main content
Language Detection identifies the spoken language of an audio file and returns a confidence score alongside the result. It analyzes up to the first 30 seconds of audio and responds synchronously — one POST in, one JSON response out.
Only the first 30 seconds of audio are analyzed. Longer files are accepted but the additional audio is ignored. For best results, ensure at least 3–5 seconds of clear speech in the first 30 seconds.

Make a request

Working with confidence scores

The confidence field is a probability — values close to 1.0 mean the model is highly certain, values close to 0.0 mean it could not commit to any language. A common pattern is to set a threshold below which you fall back to a default behavior:
There is no universally correct threshold — tune it based on your application’s tolerance for misclassification. If an audio file contains a language outside the supported set, the model returns the closest supported match, typically with low confidence.

Supported languages

100 spoken languages are recognized: Afrikaans, Albanian, Amharic, Arabic, Armenian, Assamese, Azerbaijani, Bashkir, Basque, Belarusian, Bengali, Bosnian, Breton, Bulgarian, Cantonese, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Faroese, Finnish, French, Galician, Georgian, German, Greek, Gujarati, Haitian Creole, Hausa, Hawaiian, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Javanese, Kannada, Kazakh, Khmer, Korean, Lao, Latin, Latvian, Lingala, Lithuanian, Luxembourgish, Macedonian, Malagasy, Malay, Malayalam, Maltese, Maori, Marathi, Mongolian, Myanmar, Nepali, Norwegian, Nynorsk, Occitan, Pashto, Persian, Polish, Portuguese, Punjabi, Romanian, Russian, Sanskrit, Serbian, Shona, Sindhi, Sinhala, Slovak, Slovenian, Somali, Spanish, Sundanese, Swahili, Swedish, Tagalog, Tajik, Tamil, Tatar, Telugu, Thai, Tibetan, Turkish, Turkmen, Ukrainian, Urdu, Uzbek, Vietnamese, Welsh, Yiddish, Yoruba.

API reference