Abstract:
An audio-visual hierarchical model is used to detect explosion scenes from MPEG stream based on compressed features. First, a coarse SVM is applied to discriminate explosion and explosion-like audio from others, then several fine-grained SVMs are used to determine explosion audio from explosion-like one. From these coarse to fine-grained SVMs, the audio explosion candidates are selected out. Because most explosion scenes have obvious visual change, the corresponding video is checked to get the final result.