高级检索

基于特征细节注意力计算的高效中文文本识别模型

Efficient Chinese Text Recognition Model Utilizing Feature Detail Attention

  • 摘要: 文本识别是计算机视觉领域的一个热门研究课题。当下针对中文文本识别的研究成果较少,中文字体结构复杂,字符类别多,识别难度较大。近年提出的多模态方法识别更加准确,但参数较多且推理速度慢。针对这一问题,该文提出了一种新的单一视觉中文识别模型FUD,利用注意力计算实现对图像特征的充分挖掘。首先提出位置编码矫正模块PECM对特征位置信息强化、矫正,显著提高注意力计算有效性。然后针对中文复杂结构,提出局部-全局交融的注意力计算方式,同时通过该文提出的CLA来获取归纳偏置优势,使模型仅需处理图像输入就可以兼顾特征细节与全局信息。模型FUD-Large在中文场景文本基准数据集上识别准确率达74.4%,而参数仅为44 M,达到了目前的领先水平,解决了已有的单一视觉中文文本识别模型准确率不够高的难题,利于实际应用。

     

    Abstract: Text recognition is a research topic in computer vision, with limited achievements in Chinese text recognition. The intricate font structures and extensive character categories inherent to Chinese characters pose significant challenges for recognition systems. To address these challenges, this paper proposes a new single visual Chinese recognition model called FUD. Initially, a Position Encoding Correction Module (PECM) is developed to rectify and augment the positional information within the features. Then, to capture the intricate Chinese structures, a local-global fusion attention computation approach is proposed. In addition, a CLA mechanism is introduced to estimate the inductive bias advantages. This model effectively concentrates on both fine-grained feature details and the global information in the input image. The proposed method outperforms the existing methods by an impressive accuracy of 74.4% on a benchmark dataset for Chinese scene text recognition, with only 44 million parameters.

     

/

返回文章
返回