AI Music Detection Batch
Detect AI-generated music in an audio file. Returns a clip-level verdict plus a per-window breakdown of vocal and instrumental AI content.
Authorizations
API key used for authentication and usage tracking.
Body
Audio file to analyse. Must be non-empty. Supported formats:
.3g2, .3ga, .3gp, .3gpp, .8svx, .aa3, .aac,
.ac3, .act, .adts, .aif, .aifc, .aiff, .alac,
.amb, .amr, .ape, .asf, .at3, .au, .avi,
.avr, .awb, .bwf, .c2, .caf, .dss, .dts,
.dtshd, .eac3, .ec3, .f4a, .f4b, .flac, .flv,
.gsm, .iff, .m2a, .m2ts, .m4a, .m4b, .m4r,
.m4v, .mka, .mkv, .mlp, .mmf, .mov, .mp+,
.mp1, .mp2, .mp3, .mp4, .mpa, .mpc, .mpeg,
.mpg, .mpga, .mpp, .mts, .mxf, .nist, .oga,
.ogg, .ogx, .oma, .omg, .opus, .paf, .pvf,
.qcp, .ra, .rf64, .rka, .rm, .rmvb, .sf,
.shn, .snd, .sph, .svx, .tak, .thd,
.ts, .tta, .vob, .voc, .vqf, .w64, .wav,
.wave, .weba, .webm, .wma, .wmv,
.wv. Maximum file size: 100 MB.
Response
Detection completed successfully.
Name of the submitted audio file. Empty string if no filename was provided in the upload.
"my_audio.mp3"
Total duration of the analysed audio in seconds.
x >= 089.28
Clip-level classification:
ai-vocal-music- AI-generated music with a detected synthetic voice (covers AI songs and AI synthetic vocal tracks).ai-instrumental- AI-generated instrumental music with no detectable synthetic voice.not-ai-music- the clip does not appear to contain AI-generated music.
ai-vocal-music, ai-instrumental, not-ai-music "ai-vocal-music"
Clip-level average percentage of the audio that contains vocal content, averaged across all windows.
0 <= x <= 10087.5
Clip-level AI-vocal score, on a 0-100 scale: each scored window's probability-weighted duration as a percentage of the full clip duration. Windows without enough vocal content to score contribute zero to the numerator but their duration still counts toward the total, so a clip that is mostly instrumental with only a brief, confidently-synthetic vocal patch does not read as overwhelmingly AI-vocal.
0 <= x <= 10056.5
Average confidence across all windows that were scored for AI-generated vocals. Not diluted by windows without enough vocal content to score.
0 <= x <= 10.89
Clip-level average percentage of the audio that contains instrumental music content, averaged across all windows.
0 <= x <= 10064.3
Clip-level AI-instrumental score, on a 0-100 scale: the duration-weighted average of each scored window's AI-instrumental probability. Windows that are mostly silence do not contribute.
0 <= x <= 10010.5
Average confidence across all windows that were scored for AI-generated instrumental content. Zero if no window was scored.
0 <= x <= 10.95
Clip-level average percentage of the audio that contains neither vocal nor instrumental content, averaged across all windows.
0 <= x <= 1003.51
Per-window breakdown of detection results.
End-to-end inference time in milliseconds.
x >= 01333