Advanced Search
XU Nan, WANG Peiyan. Error Detection Model with Miswritten Terminology Feature Matrix for Process Specification TextsJ. Journal of Chinese Information Processing, 2026, 40(8): 146-155. DOI: 10.3969/j.issn.1003-0077.2026.08.014
Citation: XU Nan, WANG Peiyan. Error Detection Model with Miswritten Terminology Feature Matrix for Process Specification TextsJ. Journal of Chinese Information Processing, 2026, 40(8): 146-155. DOI: 10.3969/j.issn.1003-0077.2026.08.014

Error Detection Model with Miswritten Terminology Feature Matrix for Process Specification Texts

  • Text error detection is an effective approach to reduce the time and cost of reviewing process specification texts. As a specialized technical document, process specification texts contain a large amount of terminology, which makes general domain error detection methods ineffective for process specification texts. To tackle this challenge, this paper proposes an error detection model with miswritten terminology feature matrix, which utilizes the n-gram similarity matching method to pre-check potential miswritten terminology. Meanwhile, this paper constructs a process specification text error detection corpus (PS-TEDC), which includes four types of errors: redundant words, missing words, word selection errors, and word ordering errors, totaling 12157 sentences. Experimental results demonstrate the F1 value of all baseline models has improved after adding the miswritten terminology pre-check matrix, with the highest F1 value of 75.11%.
  • loading

Catalog

    Turn off MathJax
    Article Contents

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return