Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

SDCT: Multi-Dialects Corpus Classification for Saudi Tweets

Author 1: Afnan Bayazed Author 2: Ola Torabah Author 3: Redha AlSulami Author 4: Dimah Alahmadi Author 5: Amal Babour Author 6: Kawther Saeedi
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 11, No. 11 · Published 2020 · Cited by 13

DOI: https://doi.org/10.14569/IJACSA.2020.0111128

Abstract

There is an increasing demand for analyzing the contents of social media. However, the process of sentiment analysis in Arabic language especially Arabic dialects can be very complex and challenging. This paper presents details of collecting and constructing a classified corpus of 4180 multi-dialectal Saudi tweets (SDCT). The tweets were annotated manually by five native speakers in two stages. The first stage annotated the tweets as Hijazi, Najdi, and Eastern based on some Saudi regions. The second stage annotated the sentiment as positive, negative, and natural. The annotation process was evaluated using Kappa Score. The validation process used cross validation technique through eight baseline experiments for training different classifier models. The results present that the 10-folds validation provides greater accuracy than 5-folds across the eight experiments and the classification of the Eastern dialects achieved the best accuracy compared to the other dialects with an accuracy of 91.48%.

Keywords

How to Cite this Article

Bayazed, A., Torabah, O., AlSulami, R., Alahmadi, D., Babour, A., & Saeedi, K. (2020). SDCT: Multi-Dialects Corpus Classification for Saudi Tweets. International Journal of Advanced Computer Science and Applications, 11(11). https://doi.org/10.14569/IJACSA.2020.0111128

Bayazed, Afnan, et al.. "SDCT: Multi-Dialects Corpus Classification for Saudi Tweets." International Journal of Advanced Computer Science and Applications, vol. 11, no. 11, 2020, https://doi.org/10.14569/IJACSA.2020.0111128.

@article{Bayazed2020,
  title     = {SDCT: Multi-Dialects Corpus Classification for Saudi Tweets},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {11},
  number    = {11},
  year      = {2020},
  publisher = {The Science and Information Organization},
  author    = {Afnan Bayazed and Ola Torabah and Redha AlSulami and Dimah Alahmadi and Amal Babour and Kawther Saeedi},
  doi       = {10.14569/IJACSA.2020.0111128},
  url       = {https://doi.org/10.14569/IJACSA.2020.0111128}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.