Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

NADA: New Arabic Dataset for Text Classification

Author 1: Nada Alalyani Author 2: Souad Larabi Marie-Sainte
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 9, No. 9 · Published 2018 · Cited by 26

DOI: https://doi.org/10.14569/IJACSA.2018.090928

Abstract

In the recent years, Arabic Natural Language Processing, including Text summarization, Text simplification, Text Categorization and other Natural Language-related disciplines, are attracting more researchers. Appropriate resources for Arabic Text Categorization are becoming a big necessity for the development of this research. The few existing corpora are not ready for use, they require preprocessing and filtering operations. In addition, most of them are not organized based on standard classification methods which makes unbalanced classes and thus reduced the classification accuracy. This paper proposes a New Arabic Dataset (NADA) for Text Categorization purpose. This corpus is composed of two existing corpora OSAC and DAA. The new corpus is preprocessed and filtered using the recent state of the art methods. It is also organized based on Dewey decimal classification scheme and Synthetic Minority Over-Sampling Technique. The experiment results show that NADA is an efficient dataset ready for use in Arabic Text Categorization.

Keywords

How to Cite this Article

Alalyani, N., & Marie-Sainte, S. L. (2018). NADA: New Arabic Dataset for Text Classification. International Journal of Advanced Computer Science and Applications, 9(9). https://doi.org/10.14569/IJACSA.2018.090928

Alalyani, Nada, and Souad Larabi Marie-Sainte. "NADA: New Arabic Dataset for Text Classification." International Journal of Advanced Computer Science and Applications, vol. 9, no. 9, 2018, https://doi.org/10.14569/IJACSA.2018.090928.

@article{Alalyani2018,
  title     = {NADA: New Arabic Dataset for Text Classification},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {9},
  number    = {9},
  year      = {2018},
  publisher = {The Science and Information Organization},
  author    = {Nada Alalyani and Souad Larabi Marie-Sainte},
  doi       = {10.14569/IJACSA.2018.090928},
  url       = {https://doi.org/10.14569/IJACSA.2018.090928}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.