Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

A Strategy for Training Set Selection in Text Classification Problems

Author 1: Maria Luiza C. Passini Author 2: Katiusca B. Estébanez Author 3: Grazziela P. Figueredo Author 4: Nelson F. F. Ebecken
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 4, No. 6 · Published 2013 · Cited by 13

DOI: https://doi.org/10.14569/IJACSA.2013.040608

Abstract

An issue in text classification problems involves the choice of good samples on which to train the classifier. Training sets that properly represent the characteristics of each class have a better chance of establishing a successful predictor. Moreover, sometimes data are redundant or take large amounts of computing time for the learning process. To overcome this issue, data selection techniques have been proposed, including instance selection. Some data mining techniques are based on nearest neighbors, ordered removals, random sampling, particle swarms or evolutionary methods. The weaknesses of these methods usually involve a lack of accuracy, lack of robustness when the amount of data increases, over?tting and a high complexity. This work proposes a new immune-inspired suppressive mechanism that involves selection. As a result, data that are not relevant for a classifier’s ?nal model are eliminated from the training process. Experiments show the e?ectiveness of this method, and the results are compared to other techniques; these results show that the proposed method has the advantage of being accurate and robust for large data sets, with less complexity in the algorithm.

Keywords

How to Cite this Article

Passini, M. L. C., Estébanez, K. B., Figueredo, G. P., & Ebecken, N. F. F. (2013). A Strategy for Training Set Selection in Text Classification Problems. International Journal of Advanced Computer Science and Applications, 4(6). https://doi.org/10.14569/IJACSA.2013.040608

Passini, Maria Luiza C., et al.. "A Strategy for Training Set Selection in Text Classification Problems." International Journal of Advanced Computer Science and Applications, vol. 4, no. 6, 2013, https://doi.org/10.14569/IJACSA.2013.040608.

@article{Passini2013,
  title     = {A Strategy for Training Set Selection in Text Classification Problems},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {4},
  number    = {6},
  year      = {2013},
  publisher = {The Science and Information Organization},
  author    = {Maria Luiza C. Passini and Katiusca B. Estébanez and Grazziela P. Figueredo and Nelson F. F. Ebecken},
  doi       = {10.14569/IJACSA.2013.040608},
  url       = {https://doi.org/10.14569/IJACSA.2013.040608}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.