Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Towards Stopwords Identification in Tamil Text Clustering

Author 1: M. S. Faathima Fayaza Author 2: F. Fathima Farhath
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 12, No. 12 · Published 2021 · Cited by 13

DOI: https://doi.org/10.14569/IJACSA.2021.0121267

Abstract

Now-a-days, digital documents have become the primary source of information. Therefore, natural language processing is widely utilized in information retrieval, topic modeling, document classification, and document clustering. Preprocessing plays a significant role in all of these applications. One of the critical steps in preprocessing is removing stopwords. Many languages have defined their list of stopwords. However, a publicly available stopwords list isn't available for the Tamil language since it is under-resourced. This study identified 93 general and some domain-specific stopwords for sports, entertainment, local and foreign news by analyzing more than 1.7 million Tamil documents with more than 21 million words. Also, this study shows that removing stopwords improves the accuracy of a Tamil document clustering system. It showed an improvement of 2.4%, 0.95% in the F-score for TF-IDF with one pass algorithm and FastText with the one-pass algorithm, respectively.

Keywords

How to Cite this Article

Fayaza, M. S. F., & Farhath, F. F. (2021). Towards Stopwords Identification in Tamil Text Clustering. International Journal of Advanced Computer Science and Applications, 12(12). https://doi.org/10.14569/IJACSA.2021.0121267

Fayaza, M. S. Faathima, and F. Fathima Farhath. "Towards Stopwords Identification in Tamil Text Clustering." International Journal of Advanced Computer Science and Applications, vol. 12, no. 12, 2021, https://doi.org/10.14569/IJACSA.2021.0121267.

@article{Fayaza2021,
  title     = {Towards Stopwords Identification in Tamil Text Clustering},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {12},
  number    = {12},
  year      = {2021},
  publisher = {The Science and Information Organization},
  author    = {M. S. Faathima Fayaza and F. Fathima Farhath},
  doi       = {10.14569/IJACSA.2021.0121267},
  url       = {https://doi.org/10.14569/IJACSA.2021.0121267}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.