高级检索

基于Mcross-wmBiFPN与混合注意力机制的藏文古籍图像版面分析

Tibetan Ancient Manuscript Layout Analysis Based on Mcross-wmBiFPN and Hybrid Attention Mechanism

  • 摘要: 藏文古籍是中华民族历史的珍贵遗产,对于研究藏族历史、语言和文化等领域具有极高的价值。然而,藏文古籍的版面通常十分复杂,给版面分析带来了极大挑战。图像版面分析在藏文古籍数字化流程中至关重要,直接决定着后续藏文文本识别的效果。为解决这一文字识别中的关键问题,该文以版面复杂的敦煌藏文文献为研究对象,提出了一种基于Mcross-wmBiFPN与混合注意力机制的藏文古籍图像版面分析方法。该方法通过引入混合注意力机制来评估特征的重要性,以增强重要特征并抑制无效特征,提升多尺度特征金字塔的表征能力;通过融合不同层级的跨级连接结构进一步改善BiFPN在多次级联后可能出现的信息流失问题,从而提高了模型的检测精度。在构建的影印版敦煌藏文古籍文献图像数据集上,所提方法取得了99.17%的mAP,相较于基准模型EfficientDet提高了6.14%,为类似古籍的版面分析提供了重要的参考和借鉴。为进一步验证该文方法的有效性和扩展应用范围,收集了中英文以及其他语种版面图像进行测试实验,实验结果证明,所提方法可有效应用于其他版面的版面分析,具有较好的应用前景。

     

    Abstract: Tibetan ancient texts are invaluable treasures of Chinese cultural heritage, offering significant insights into Tibetan history, language, and culture. In this sense, the image-based layout analysis for Tibetan manuscripts is crucial in that it directly impacts the effectiveness of subsequent Tibetan text recognition. This paper focuses on the complex layout of Dunhuang Tibetan manuscripts and proposes an image layout analysis method based on Mcross-wmBiFPN and a hybrid attention mechanism. This method incorporates a hybrid attention mechanism to evaluate the importance of features, which improves the representational capability of the multi-scale feature pyramid. Additionally, by integrating cross-level connections across different layers, the method mitigates the potential information loss that may occur after multiple cascades of BiFPN. On the constructed dataset of Dunhuang Tibetan manuscript images, the proposed method achieved a mean Average Precision (mAP) of 99.17%, i.e. a 6.14% improvement over the baseline model EfficientDet.

     

/

返回文章
返回