The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |
First page preview

A Hybrid Semantic-Statistical Feature Fusion Framework for Bilingual Text Classification on Multilingual Big Data Corpora

Author 1: Kavitha M Author 2: Purohit Shrinivasacharya Author 3: Y S Nijagunarya
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 17, No. 6 · Published 2026

DOI: https://doi.org/10.14569/IJACSA.2026.0170684

Abstract

As the volume of scientific information is growing exponentially in several languages, there is a need for practical and scalable bilingual classification systems for large aligned scientific text corpora. To address this challenge, this study makes two key contributions - first, a large-scale bilingual English–Hindi aligned arXiv scientific text corpus, providing a structured resource for multilingual scientific text analytics and classification research is constructed. Second, the study proposes a bilingual scientific text classification framework and performs a rigorous experimental evaluation using three strong multilingual transformer models, namely Mini Language Model-12 Layers (MiniLM-L12), Multilingual Bidirectional Encoder Representations from Transformers (mBERT), and Cross-lingual Language Model-Robustly Optimized Bidirectional Encoder Representations from Transformers Pretraining Approach (XLM-RoBERTa), on the developed English–Hindi aligned arXiv big data corpus. English and Hindi summaries are categorized independently to investigate the performance trade-offs in each language. The proposed hybrid MiniLM-L12 + Multi-Layer Perceptron (MLP) architecture enhance the classification capability through the integration of statistical feature design with contextual sentence embeddings. Empirical analysis indicates that the proposed bilingual classification framework consistently outperforms the baseline transformer-only models, achieving a higher accuracy of 95.56% and weighted F1-score of 95.31%, while maintaining computational efficiency. The findings emphasize the effectiveness of hybrid representation learning for bilingual big data corpora and provide practical insights for scalable multilingual scholarly text analytics.

Keywords

How to Cite this Article

Kavitha M, Purohit Shrinivasacharya and Y S Nijagunarya. "A Hybrid Semantic-Statistical Feature Fusion Framework for Bilingual Text Classification on Multilingual Big Data Corpora". International Journal of Advanced Computer Science and Applications (IJACSA), Vol. 17, No. 6, 2026. https://doi.org/10.14569/IJACSA.2026.0170684

BibTeX

@article{M2026,
  title     = {A Hybrid Semantic-Statistical Feature Fusion Framework for Bilingual Text Classification on Multilingual Big Data Corpora},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {17},
  number    = {6},
  year      = {2026},
  publisher = {The Science and Information Organization},
  author    = {Kavitha M and Purohit Shrinivasacharya and Y S Nijagunarya},
  doi       = {10.14569/IJACSA.2026.0170684},
  url       = {https://doi.org/10.14569/IJACSA.2026.0170684}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.