Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

ForenVoice-Secure: Robust and Privacy-Aware Audio Data Mining for Forensic Speaker Identification

Author 1: Mubarak Albathan
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 17, No. 1 · Published 2026

DOI: https://doi.org/10.14569/IJACSA.2026.0170115

Abstract

Speech is now routine evidence in criminal investigations, but forensic audio rarely matches the clean assumptions of standard speaker recognition. Clips are short, noisy, codec-compressed, and channel-mismatched, and they are increasingly exposed to replay and synthetic speech manipulation. Therefore, the cast criminal voice identification is forensic audio data mining, aiming to extract a stable identity structure from heterogeneous and potentially adversarial evidence, while respecting operational and privacy constraints. In this study, a novel ForenVoice-Secure system is proposed, a unified pipeline that combines robust representation learning, spoof-aware decisioning, and privacy-preserving training. Audio is mapped to log-Mel spectrograms and encoded with a CNN, while an LSTM aggregates temporal identity cues from irregular utterances. Robustness is improved through multi-task learning (identity + spoof), adversarial training, and spectro-temporal consistency checks for replay/deepfake artifacts. Privacy is addressed using federated learning, keeping raw recordings local and sharing only model updates. Experiments on VoxCeleb2, ASVspoof 2021, and a forensic-style speaker comparison corpus achieve statistically significant performance gains, 98.43% mean identification accuracy with strong class-balanced performance (macro F1 = 98.10%, precision = 98.22%, recall = 98.01%) and statistically significant gains over strong baselines across repeated folds (F1: p=8.0×〖10〗^(-4); precision: p=1.1×〖10〗^(-3); recall: p=9.0×〖10〗^(-4)). The model remains lightweight (≈4.3M parameters, ≈1.2 GFLOPs per 3 s), enabling near real-time inference with modest overhead from consistency checks (<6%). Overall, ForenVoice-Secure provides a compact and reproducible forensic audio data mining framework for scalable, spoof-resilient, privacy-aware law-enforcement identification.

Keywords

How to Cite this Article

Albathan, M. (2026). ForenVoice-Secure: Robust and Privacy-Aware Audio Data Mining for Forensic Speaker Identification. International Journal of Advanced Computer Science and Applications, 17(1). https://doi.org/10.14569/IJACSA.2026.0170115

Albathan, Mubarak. "ForenVoice-Secure: Robust and Privacy-Aware Audio Data Mining for Forensic Speaker Identification." International Journal of Advanced Computer Science and Applications, vol. 17, no. 1, 2026, https://doi.org/10.14569/IJACSA.2026.0170115.

@article{Albathan2026,
  title     = {ForenVoice-Secure: Robust and Privacy-Aware Audio Data Mining for Forensic Speaker Identification},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {17},
  number    = {1},
  year      = {2026},
  publisher = {The Science and Information Organization},
  author    = {Mubarak Albathan},
  doi       = {10.14569/IJACSA.2026.0170115},
  url       = {https://doi.org/10.14569/IJACSA.2026.0170115}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.