Knowledge Distillation-based Optimization Strategy for Low-frequency Word Translation in Neural Machine Translation
-
Abstract
Low-frequency words remain difficult for neural machine translation This paper proposes a knowledge distillation method to improve translation quality for informative but relatively infrequent words. A teacher model with stronger low-frequency-word translation ability is first trained, and a single-teacher distillation framework is then used to guide student learning. A dual-teacher distillation model is further designed to improve training stability and preserve high-frequency-word performance. Experimental results showed that the single-teacher model increases BLEU by 0.64 in the English-German task over the state-of-the-art system, and the dual-teacher model improve BLEU by 1.24 in English-German, 0.47 in English-Czech, and 0.87 in English-French.
-
-