Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Normalisation of Indonesian-English Code-Mixed Text and its Effect on Emotion Classification

Author 1: Evi Yulianti Author 2: Ajmal Kurnia Author 3: Mirna Adriani Author 4: Yoppy Setyo Duto
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 12, No. 11 · Published 2021 · Cited by 23

DOI: https://doi.org/10.14569/IJACSA.2021.0121177

Abstract

Usage of code-mixed text has increased in re-cent years among Indonesian internet users, who often mix Indonesian-language with English-language text. Normalisation of this code-mixed text into Indonesian needs to be performed to capture the meaning of English parts of the text and process them effectively. We improve a state-of-the-art code-mixed Indonesian-English normalisation system by modifying its pipeline modules. We further analyse the effect of code-mixed normalisation on emotion classification tasks. Our approach significantly improved on a state-of-the-art Indonesian-English code-mixed text normal-isation system in both the individual pipeline modules and the overall system. The new feature set in the language identification module showed an improvement of 4.26% in terms of F1 score. The combination of machine translation and ruleset in the lexical normalisation module improved BLEU score by 25.22% and lowered WER by 62.49%. The use of context in the translation module improved BLEU score by 2.5% and lowered WER by 8.84%. The effectiveness of the overall pipeline normalisation system increased by 32.11% and 33.82%, in terms of BLEU score and WER, respectively. Code-mixed normalisation also improved the accuracy of emotion classification by up to 37.74% in terms of F1 score.

Keywords

How to Cite this Article

Yulianti, E., Kurnia, A., Adriani, M., & Duto, Y. S. (2021). Normalisation of Indonesian-English Code-Mixed Text and its Effect on Emotion Classification. International Journal of Advanced Computer Science and Applications, 12(11). https://doi.org/10.14569/IJACSA.2021.0121177

Yulianti, Evi, et al.. "Normalisation of Indonesian-English Code-Mixed Text and its Effect on Emotion Classification." International Journal of Advanced Computer Science and Applications, vol. 12, no. 11, 2021, https://doi.org/10.14569/IJACSA.2021.0121177.

@article{Yulianti2021,
  title     = {Normalisation of Indonesian-English Code-Mixed Text and its Effect on Emotion Classification},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {12},
  number    = {11},
  year      = {2021},
  publisher = {The Science and Information Organization},
  author    = {Evi Yulianti and Ajmal Kurnia and Mirna Adriani and Yoppy Setyo Duto},
  doi       = {10.14569/IJACSA.2021.0121177},
  url       = {https://doi.org/10.14569/IJACSA.2021.0121177}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.