Error Detection Model with Miswritten Terminology Feature Matrix for Process Specification Texts
-
Abstract
Text error detection is an effective approach to reduce the time and cost of reviewing process specification texts. As a specialized technical document, process specification texts contain a large amount of terminology, which makes general domain error detection methods ineffective for process specification texts. To tackle this challenge, this paper proposes an error detection model with miswritten terminology feature matrix, which utilizes the n-gram similarity matching method to pre-check potential miswritten terminology. Meanwhile, this paper constructs a process specification text error detection corpus (PS-TEDC), which includes four types of errors: redundant words, missing words, word selection errors, and word ordering errors, totaling 12157 sentences. Experimental results demonstrate the F1 value of all baseline models has improved after adding the miswritten terminology pre-check matrix, with the highest F1 value of 75.11%.
-
-