中文文本体裁的自动分类机制

方鸷飞,林鸿飞,杨志豪,赵晶

PDF(265 KB)
PDF(265 KB)
中文信息学报 ›› 2006, Vol. 20 ›› Issue (2) : 26-34.

中文文本体裁的自动分类机制

  • 方鸷飞,林鸿飞,杨志豪,赵晶
作者信息 +

Automatic Classification of Chinese Text Genre

  • FANG Zhi-fei,LIN Hong-fei,YANG Zhi-hao,ZHAO Jing
Author information +
History +

摘要

文本按体裁自动分类属于按文本的形式分类的范畴,所以它与按内容自动分类问题有许多的不同之处,本文提出了一种关于中文文本体裁自动分类的新机制。在体裁分类过程中首要的问题是分类特征的选取,体裁分类特征项分为两种方式加以描述,一是集合形式,如基于分类词典和语料统计的政论性词汇和情感词汇等,二是规则形式,如公文标识信息和条文句等。基于根据特征之间的关联性和差异性,采用样本分布决策的方法抽取相应的特征项。最后利用支撑向量机算法进行自动分类。该机制已经在五类体裁的语料上得到实现,并获得了较好的效果。

Abstract

Genre is defined as a category on the basis of external criteria , so its classification is different from the classification based on content. A new mechanism for automatic classification of Chinese text genre is presented , and its main idea is as follows. Features for genre classification , as an essential factor in the mechanism , are described in two ways : one is in word-set , such as affective words and political words derived from some related dictionaries and corpus statistics ; another one is in rule format , such as document identifiers and items. In terms of the correlativeness and variance of features , an approach of parametric distribution is applied to evaluate various features of the genres and extract the features for genre classification. Support Vector Machine is then used as the learning algorithm to build the classifier. The experiment on automatic classification of Chinese text genres , running on a text corpus consisting of five genres , shows that it can improve the precision of classification.

关键词

计算机应用 / 中文信息处理 / 体裁分类 / 特征项选取 / 样本分布决策 / 支撑向量机

Key words

computer application / Chinese information processing / text genre classification / feature selection / parametric distribution / support vector machine

引用本文

导出引用
方鸷飞,林鸿飞,杨志豪,赵晶. 中文文本体裁的自动分类机制. 中文信息学报. 2006, 20(2): 26-34
FANG Zhi-fei,LIN Hong-fei,YANG Zhi-hao,ZHAO Jing. Automatic Classification of Chinese Text Genre. Journal of Chinese Information Processing. 2006, 20(2): 26-34

参考文献

[1] Barbara H Kwasnik ,Kevin Crowston. Genres of Digital Documents[A] . In : Proceedings of the 37th Annual Hawaii International Conference on System Sciences (HICSS’04) [C] ,Big Island ,Hawaii ,2004.
[2] Douglas Biber. Using register-diversified corpora for general language study[J ] . Computational Linguistics ,1993 ,19 : 219 - 241.
[3] Brett Kessler ,Geoffrey Nunberg ,Hinrich Schutze. Automatic Detection of Text Genre [A] . In : Proceedings of 35th Annual Meeting of Association for Computational Linguistics and 8th Conference of European Chapter of Association for Computational Linguistics[C] ,Madrid ,Spain ,1997 ,32 - 38.
[4] E. Stamatatos ,N. Fakotakis & G. Kokkinakis ,Text genre detection using common word frequencies[A] . In : Proceedings of 18 International Conference on Computational Linguistics[C] ,Luxemburg ,2001 ,808 - 814.
[5] Marina Santini. A Shallow Approach To Syntactic Feature Extraction For Genre Classification[A]. In : Proceedings of the 7th Annual Colloquiumfor the UK Special Interest Group for Computational Linguistics[C] ,Birmingham,UK. 2004.
[6] Jussi Karlgren ,Douglass Cutting ,Recognizing text genres with simple metrics using discriminant analysis[A] ,In : Proceedings of the 15th conference on Computational linguistics [C], Kyoto ,Japan ,1994.
[7] Aidan Finn ,Nicholas Kushmerick. Learning to classify documents according to genre[A] . In : proceedings of IJCAI - 03 Workshop on Computational Approaches to Style Analysis and Synthesis[C] ,California : IJCAI ,2003 ,35 - 45.
[8] Yong-Bae Lee ,Sung Hyon Myaeng. Automatic Identification of Text Genres and Their Roles in Subject-Based Categorization[A] . In : Proceedings of the 37th Hawaii International Conference on System Sciences (HICSS 2004) [C] ,Big Island ,Hawaii ,USA. 2004.
[9] Andreas Rauber and Alexander Muller-Kogler. Integrating automatic genre analysis into digital libraries[A] , In : Proceedings of the First ACM IEEE Joint Conference on Digital Libraries (JCDL’01) [C] . 2001 ,1 - 10.
[10] 王慧玲,宋柔,戴伟长. 汉语文本按语体分类的研究[A] . 自然语言理解与机器翻译[C] ,北京: 清华大学出版社,2001 ,344 - 352.
[11] 金振邦. 文章体裁辞典[M] . 第1版,长春: 东北师范大学出版社. 1986.
[12] 寸镇东. 语境与修辞[M] . 贵阳:贵州人民出版社,1996 ,290 - 305.
[13] 董大年. 现代汉语分类词典[M] . 第1版,上海: 汉语大词典出版社. 1998.
[14] 何劲松,郑浩然,王煦法. 从熵均值决策到样本分布决策[J] . 软件学报,2003 ,14 (3) :480 - 483.
[15] Thorsten Joachims ,Text Categorization with Support Vector Machines : Learning with Many Relevant Features[A] . In : Proceedings of European Conference on Machine Learning (ECML) ,Claire Nédellec and Céline Rouveirol (ed.) [C] ,Chemnitz ,Germany ,1998.

基金

国家自然科学基金资助项目(60373095)
PDF(265 KB)

Accesses

Citation

Detail

段落导航
相关文章

/