Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Factored Phrase-based Statistical Machine Pre-training with Extended Transformers

Author 1: Vivien L. Beyala Author 2: Marcellin J. Nkenlifack Author 3: Perrin Li Litet
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 11, No. 9 · Published 2020

DOI: https://doi.org/10.14569/IJACSA.2020.0110907

Abstract

This paper presents the development of a cascaded hybrid multi- lingual automatic translation system, by allowing a tight coupling between the two underlying research approach in machine translation, namely, the neuronal (deterministic approach) and statistical (probabilistic approach), while fully taking advantage of each method in order to improve translation performance. This architecture addresses two major problems frequently occurring when dealing with morphologically richer languages in MT, that is, the significant number unknown tokens generated due to the presence of out of vocabulary (OOV) words, and size of the output vocabulary. Additionally, we incorporated factors (additional word-level linguistic information) in order to alleviate data sparseness problem or potentially reduce language ambiguity, the factors we considered are lemmatization and Part-of-Speech tags (taking into consideration its various compounds). We combined a fully-factored transformer and a factored PB-SMT, where, the training data is pre-translated using the trained fully-factored transformer, and afterwards employed to build an PB-SMT system, parallelly using the pre-translated development set to tune parameters. Finally, in order to produce the desired results, we operated the FPB-SMT system to re-decode the pre-translated test set in a post-processing step. Experiments performed on translations from Japanese to English and English to Japanese reveals that our proposed cascaded hybrid framework outperforms the strong HMT state-of-the-art by over 8.61% BLEU and 7.25% BLEU, respectively, for validation set, and over 8.70% BLEU and 7.70% BLEU, respectively, for test set.

Keywords

How to Cite this Article

Beyala, V. L., Nkenlifack, M. J., & Litet, P. L. (2020). Factored Phrase-based Statistical Machine Pre-training with Extended Transformers. International Journal of Advanced Computer Science and Applications, 11(9). https://doi.org/10.14569/IJACSA.2020.0110907

Beyala, Vivien L., et al.. "Factored Phrase-based Statistical Machine Pre-training with Extended Transformers." International Journal of Advanced Computer Science and Applications, vol. 11, no. 9, 2020, https://doi.org/10.14569/IJACSA.2020.0110907.

@article{Beyala2020,
  title     = {Factored Phrase-based Statistical Machine Pre-training with Extended Transformers},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {11},
  number    = {9},
  year      = {2020},
  publisher = {The Science and Information Organization},
  author    = {Vivien L. Beyala and Marcellin J. Nkenlifack and Perrin Li Litet},
  doi       = {10.14569/IJACSA.2020.0110907},
  url       = {https://doi.org/10.14569/IJACSA.2020.0110907}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.