As the volume of scientific information is growing exponentially in several languages, there is a need for practical and scalable bilingual classification systems for large aligned scientific text corpora. To address this challenge, this study makes two key contributions - first, a large-scale bilingual English–Hindi aligned arXiv scientific text corpus, providing a structured resource for multilingual scientific text analytics and classification research is constructed. Second, the study proposes a bilingual scientific text classification framework and performs a rigorous experimental evaluation using three strong multilingual transformer models, namely Mini Language Model-12 Layers (MiniLM-L12), Multilingual Bidirectional Encoder Representations from Transformers (mBERT), and Cross-lingual Language Model-Robustly Optimized Bidirectional Encoder Representations from Transformers Pretraining Approach (XLM-RoBERTa), on the developed English–Hindi aligned arXiv big data corpus. English and Hindi summaries are categorized independently to investigate the performance trade-offs in each language. The proposed hybrid MiniLM-L12 + Multi-Layer Perceptron (MLP) architecture enhance the classification capability through the integration of statistical feature design with contextual sentence embeddings. Empirical analysis indicates that the proposed bilingual classification framework consistently outperforms the baseline transformer-only models, achieving a higher accuracy of 95.56% and weighted F1-score of 95.31%, while maintaining computational efficiency. The findings emphasize the effectiveness of hybrid representation learning for bilingual big data corpora and provide practical insights for scalable multilingual scholarly text analytics.
Kavitha M, Purohit Shrinivasacharya and Y S Nijagunarya. "A Hybrid Semantic-Statistical Feature Fusion Framework for Bilingual Text Classification on Multilingual Big Data Corpora". International Journal of Advanced Computer Science and Applications (IJACSA), Vol. 17, No. 6, 2026. https://doi.org/10.14569/IJACSA.2026.0170684
BibTeX
@article{M2026,
title = {A Hybrid Semantic-Statistical Feature Fusion Framework for Bilingual Text Classification on Multilingual Big Data Corpora},
journal = {International Journal of Advanced Computer Science and Applications},
volume = {17},
number = {6},
year = {2026},
publisher = {The Science and Information Organization},
author = {Kavitha M and Purohit Shrinivasacharya and Y S Nijagunarya},
doi = {10.14569/IJACSA.2026.0170684},
url = {https://doi.org/10.14569/IJACSA.2026.0170684}
}
Open Access — licensed under a
Creative Commons Attribution 4.0 International License.
Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.