Abstract:
Discriminating features between speech and music are analyzed, including perceptual features like pitch, brightness and harmonicity, etc, and Mel-Frequency Cepstral Coefficients (MFCC). Their performances are evaluated in a left-right discrete HMM-based audio classifier, which is used to classify audio into speech, music, their mixed sound and such-like three categories with maximum likelihood criterion. The experiment results show that the features selected are effective for speech/music classification, and the classification accuracy is excellent.