高级检索

基于视听分层模型的实时爆炸场景识别

Real-Time Recognition of Explosion Scenes Based on Audio-Visual Hierarchical Model

  • 摘要: 提出在实时环境下使用基于听觉和视觉的分层模型对MPEG多媒体数据流中的"爆炸"场景在压缩域进行识别的算法.首先用一个粗分支持向量机把爆炸和类似爆炸的音频从别的音频中识别出来,然后再分别用几个精细支持向量机把爆炸和类似爆炸的音频区分开,由此得到音频爆炸备选场景.由于大多数爆炸场景均伴随剧烈的视觉突变,因此对得到的音频爆炸备选场景再判断其对应的视觉特征是否发生了变化,得到最后的识别结果.

     

    Abstract: An audio-visual hierarchical model is used to detect explosion scenes from MPEG stream based on compressed features. First, a coarse SVM is applied to discriminate explosion and explosion-like audio from others, then several fine-grained SVMs are used to determine explosion audio from explosion-like one. From these coarse to fine-grained SVMs, the audio explosion candidates are selected out. Because most explosion scenes have obvious visual change, the corresponding video is checked to get the final result.

     

/

返回文章
返回