Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

An Enhanced Twitter Corpus for the Classification of Arabic Speech Acts

Author 1: Majdi Ahed Author 2: Bassam H. Hammo Author 3: Mohammad A. M. Abushariah
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 11, No. 3 · Published 2020

DOI: https://doi.org/10.14569/IJACSA.2020.0110325

Abstract

Twitter has gained wide attention as a major social media platform where many topics are discussed on daily basis through millions of tweets. A tweet can be viewed as a speech act (SA), which is an utterance for presenting information, hiding indirect meaning, or carrying out an action. According to SA theory, SA can represent an assertion, a question, a recommendation, or many other things. In this paper, we tackle the problem of constructing a reference corpus of Arabic tweets for the classification of Arabic speech acts. We refer to this corpus as the Arabic Tweets Speech Act Corpus (ArTSAC). It is an enhancement of a modern standard Arabic (MSA) tweet corpus of speech acts called ArSAS. ArTSAC is more advantageous than ArSAS in terms of its richness of annotated features. The goal of ArTSAC is twofold: Firstly, to understand the purpose and intention of tweets which act in accordance with the SA theory, and hence positively influencing the development of many natural language processing (NLP) applications. Secondly, as a future goal, to be used as a benchmark annotated dataset for testing and evaluating state-of-the-art Arabic SA classification algorithms and applications. ArTSAC has been put in practice to classify Arabic tweets containing speech acts using the Support Vector Machine (SVM) classification algorithm. The results of the experiments show that the enhanced ArTSAC corpus achieved an average precision of 90.6% and an F-score of 89.6%. Substantially it outperformed the results of its predecessor ArTSAC corpus.

Keywords

How to Cite this Article

Ahed, M., Hammo, B. H., & Abushariah, M. A. M. (2020). An Enhanced Twitter Corpus for the Classification of Arabic Speech Acts. International Journal of Advanced Computer Science and Applications, 11(3). https://doi.org/10.14569/IJACSA.2020.0110325

Ahed, Majdi, et al.. "An Enhanced Twitter Corpus for the Classification of Arabic Speech Acts." International Journal of Advanced Computer Science and Applications, vol. 11, no. 3, 2020, https://doi.org/10.14569/IJACSA.2020.0110325.

@article{Ahed2020,
  title     = {An Enhanced Twitter Corpus for the Classification of Arabic Speech Acts},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {11},
  number    = {3},
  year      = {2020},
  publisher = {The Science and Information Organization},
  author    = {Majdi Ahed and Bassam H. Hammo and Mohammad A. M. Abushariah},
  doi       = {10.14569/IJACSA.2020.0110325},
  url       = {https://doi.org/10.14569/IJACSA.2020.0110325}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.