A Multistage Feature Selection Model for Document Classification Using Information Gain and Rough Set

Mrs. Leena. H. Patil; Dr. Mohammed Atique

doi:10.14569/IJARAI.2014.031103

DOI: 10.14569/IJARAI.2014.031103

PDF

A Multistage Feature Selection Model for Document Classification Using Information Gain and Rough Set

Author 1: Mrs. Leena. H. Patil

Author 2: Dr. Mohammed Atique

International Journal of Advanced Research in Artificial Intelligence(IJARAI), Volume 3 Issue 11, 2014.

Abstract and Keywords
How to Cite this Article
{} BibTeX Source

Abstract: Huge number of documents are increasing rapidly, therefore, to organize it in digitized form text categorization becomes an challenging issue. A major issue for text categorization is its large number of features. Most of the features are noisy, irrelevant and redundant, which may mislead the classifier. Hence, it is most important to reduce dimensionality of data to get smaller subset and provide the most gain in information. Feature selection techniques reduce the dimensionality of feature space. It also improves the overall accuracy and performance. Hence, to overcome the issues of text categorization feature selection is considered as an efficient technique . Therefore, we, proposed a multistage feature selection model to improve the overall accuracy and performance of classification. In the first stage document preprocessing part is performed. Secondly, each term within the documents are ranked according to their importance for classification using the information gain. Thirdly rough set technique is applied to the terms which are ranked importantly and feature reduction is carried out. Finally a document classification is performed on the core features using Naive Bayes and KNN classifier. Experiments are carried out on three UCI datasets, Reuters 21578, Classic 04 and Newsgroup 20. Results show the better accuracy and performance of the proposed model.

Keywords: Introduction; Document Preprocessing; Information Gain; Rough Set; Classifiers

Mrs. Leena. H. Patil and Dr. Mohammed Atique, “A Multistage Feature Selection Model for Document Classification Using Information Gain and Rough Set” International Journal of Advanced Research in Artificial Intelligence(IJARAI), 3(11), 2014. http://dx.doi.org/10.14569/IJARAI.2014.031103

@article{Patil2014,
title = {A Multistage Feature Selection Model for Document Classification Using Information Gain and Rough Set},
journal = {International Journal of Advanced Research in Artificial Intelligence},
doi = {10.14569/IJARAI.2014.031103},
url = {http://dx.doi.org/10.14569/IJARAI.2014.031103},
year = {2014},
publisher = {The Science and Information Organization},
volume = {3},
number = {11},
author = {Mrs. Leena. H. Patil and Dr. Mohammed Atique}
}

Copyright Statement: This is an open access article licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, even commercially as long as the original work is properly cited.

A Multistage Feature Selection Model for Document Classification Using Information Gain and Rough Set

Upcoming Conferences