Music Detection Batch
Classify music and speech in an audio file. Returns frame-level probabilities, a primary label, and percentage breakdowns of content type.
Authorizations
API key used for authentication and usage tracking.
Body
Audio file to analyse. Must be non-empty. Supported formats:
.3g2, .3ga, .3gp, .3gpp, .8svx, .aa3, .aac,
.ac3, .act, .adts, .aif, .aifc, .aiff, .alac,
.amb, .amr, .ape, .asf, .at3, .au, .avi,
.avr, .awb, .bwf, .c2, .caf, .dss, .dts,
.dtshd, .eac3, .ec3, .f4a, .f4b, .flac, .flv,
.gsm, .iff, .m2a, .m2ts, .m4a, .m4b, .m4r,
.m4v, .mka, .mkv, .mlp, .mmf, .mov, .mp+,
.mp1, .mp2, .mp3, .mp4, .mpa, .mpc, .mpeg,
.mpg, .mpga, .mpp, .mts, .mxf, .nist, .oga,
.ogg, .ogx, .oma, .omg, .opus, .paf, .pvf,
.qcp, .ra, .rf64, .rka, .rm, .rmvb, .sf,
.shn, .snd, .sph, .spx, .svx, .tak, .thd,
.ts, .tta, .vob, .voc, .vqf, .w64, .wav,
.wave, .weba, .webm, .wma, .wmv,
.wv. Maximum file size: 100 MB.
Response
Detection completed successfully.
Name of the submitted audio file. Empty string if no filename was provided in the upload.
"my_audio.wav"
Total duration of the analysed audio in seconds.
x >= 05.76
Overall classification of the clip:
music- music covers at least as much of the clip as speech, and more than zero.speech- speech covers more of the clip than music, and more than zero.neither- neither music nor speech reached the dominant threshold for any portion of the clip.unknown- no frames could be produced from the audio.
music, speech, neither, unknown "speech"
Percentage of the clip classified as containing music.
0 <= x <= 1000
Percentage of the clip classified as containing speech.
0 <= x <= 10086.7
End-to-end inference time in milliseconds.
x >= 01243.5
Ordered list of per-frame classification results covering the full duration of the clip.