高级检索

用于文本区域提取的边缘像素聚类方法

Edge-Pixels Clustering for Text Area Extraction

  • 摘要: 根据边缘点的位置和颜色信息采取逐步松弛的聚类方法将图像分割成像素子集,应用文本区域边缘的分布特征提取初始文本区,并进行边界扩展得到完整的文本区域;同时给出了一种文本区域二值化方法,减少了在文本颜色极性未知时的二值图像个数,可提高字符分割等后续处理的计算效率.实验结果表明,该方法对文本区域提取是有效的,提取完整率达99%.

     

    Abstract: An approach based on edge-pixels clustering to extract Chinese and English text areas from an image is proposed.The image is segmented into pixel-subclasses based on the colors and positions of edgepixels.And then the initial text areas are extracted according to the features of edges in text area. The boundaries of the initial text areas are expanded for the entire text areas.Furthermore,an algorithm of text area binarization is presented to improve the efficiency of post-processing by reducing the number of binary images when the text color polarity is unknown.The experimental results show that the proposed approach is effective with integrality up to 99%.

     

/

返回文章
返回