Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Zero-resource Multi-dialectal Arabic Natural Language Understanding

Author 1: Muhammad Khalifa Author 2: Hesham Hassan Author 3: Aly Fahmy
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 12, No. 3 · Published 2021 · Cited by 6

DOI: https://doi.org/10.14569/IJACSA.2021.0120369

Abstract

A reasonable amount of annotated data is required for fine-tuning pre-trained language models (PLM) on down-stream tasks. However, obtaining labeled examples for different language varieties can be costly. In this paper, we investigate the zero-shot performance on Dialectal Arabic (DA) when fine-tuning a PLM on modern standard Arabic (MSA) data only— identifying a significant performance drop when evaluating such models on DA. To remedy such performance drop, we propose self-training with unlabeled DA data and apply it in the context of named entity recognition (NER), part-of-speech (POS) tagging, and sarcasm detection (SRD) on several DA varieties. Our results demonstrate the effectiveness of self-training with unlabeled DA data: improving zero-shot MSA-to-DA transfer by as large as ~10% F₁ (NER), 2% accuracy (POS tagging), and 4.5% F₁ (SRD). We conduct an ablation experiment and show that the performance boost observed directly results from the unlabeled DA examples used for self-training. Our work opens up opportunities for leveraging the relatively abundant labeled MSA datasets to develop DA models for zero and low-resource dialects. We also report new state-of-the-art performance on all three tasks and open-source our fine-tuned models for the research community.

Keywords

How to Cite this Article

Khalifa, M., Hassan, H., & Fahmy, A. (2021). Zero-resource Multi-dialectal Arabic Natural Language Understanding. International Journal of Advanced Computer Science and Applications, 12(3). https://doi.org/10.14569/IJACSA.2021.0120369

Khalifa, Muhammad, et al.. "Zero-resource Multi-dialectal Arabic Natural Language Understanding." International Journal of Advanced Computer Science and Applications, vol. 12, no. 3, 2021, https://doi.org/10.14569/IJACSA.2021.0120369.

@article{Khalifa2021,
  title     = {Zero-resource Multi-dialectal Arabic Natural Language Understanding},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {12},
  number    = {3},
  year      = {2021},
  publisher = {The Science and Information Organization},
  author    = {Muhammad Khalifa and Hesham Hassan and Aly Fahmy},
  doi       = {10.14569/IJACSA.2021.0120369},
  url       = {https://doi.org/10.14569/IJACSA.2021.0120369}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.