高级检索

基于LLM可解释的篇章级小句复合体边界自动识别

Discourse-level Clause Complex Boundary Identification Based on Explainable Large Language Models

  • 摘要: 小句复合体是汉语篇章的基本语义单元,对其边界的准确识别是实现复杂篇章理解与深层次语义解析的基础。现有方法多依赖端到端的深度学习模型,虽能取得较高性能,但普遍缺乏可解释性。为此,该文提出并验证了一种结合依存句法特征与LLM语义特征的可解释识别范式。该方法通过融合决策算法,将LTP提取的句法特征与LLM生成的语义标签结合,以加权方式共同参与边界判定。实验结果验证了该范式的有效性,LTP-LLM混合模型准确率达到88.24%,在保证结果可解释性的同时,取得了较高的准确率。

     

    Abstract: Accurate clause complex boundary identification is foundational for advanced discourse comprehension and deep semantic parsing. Existing end-to-end deep learning models achieve high performance but provide less interpretability. An interpretable recognition paradigm is proposed in this paper, combining syntactic features from the Language Technology Platform (LTP) with semantic features from a large language model (LLM). The approach employed a fusion decision algorithm to integrate LTP-extracted syntactic features and LLM-generated semantic labels in a weighted manner for joint boundary determination. Experimental results demonstrate the hybrid LTP-LLM model achieves an 88.24% accuracy, which is an effective balance between performance and interpretability.

     

/

返回文章
返回