Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Convolutional Transformer based Local and Global Feature Learning for Speech Enhancement

Author 1: Chaitanya Jannu Author 2: Sunny Dayal Vanambathina
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 14, No. 1 · Published 2023 · Cited by 16

DOI: https://doi.org/10.14569/IJACSA.2023.0140181

Abstract

Speech enhancement (SE) is an important method for improving speech quality and intelligibility in noisy environments where received speech is severely distorted by noise. An efficient speech enhancement system relies on accurately modelling the long-term dependencies of noisy speech. Deep learning has greatly benefited by the use of transformers where long-term dependencies can be modelled more efficiently with multi-head attention (MHA) by using sequence similarity. Transformers frequently outperform recurrent neural network (RNN) and convolutional neural network (CNN) models in many tasks while utilizing parallel processing. In this paper we proposed a two-stage convolutional transformer for speech enhancement in time domain. The transformer considers global information as well as parallel computing, resulting in a reduction of long-term noise. In the proposed work unlike two-stage transformer neural network (TSTNN) different transformer structures for intra and inter transformers are used for extracting the local as well as global features of noisy speech. Moreover, a CNN module is added to the transformer so that short-term noise can be reduced more effectively, based on the ability of CNN to extract local information. The experimental findings demonstrate that the proposed model outperformed the other existing models in terms of STOI (short-time objective intelligibility), and PESQ (perceptual evaluation of the speech quality).

Keywords

How to Cite this Article

Jannu, C., & Vanambathina, S. D. (2023). Convolutional Transformer based Local and Global Feature Learning for Speech Enhancement. International Journal of Advanced Computer Science and Applications, 14(1). https://doi.org/10.14569/IJACSA.2023.0140181

Jannu, Chaitanya, and Sunny Dayal Vanambathina. "Convolutional Transformer based Local and Global Feature Learning for Speech Enhancement." International Journal of Advanced Computer Science and Applications, vol. 14, no. 1, 2023, https://doi.org/10.14569/IJACSA.2023.0140181.

@article{Jannu2023,
  title     = {Convolutional Transformer based Local and Global Feature Learning for Speech Enhancement},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {14},
  number    = {1},
  year      = {2023},
  publisher = {The Science and Information Organization},
  author    = {Chaitanya Jannu and Sunny Dayal Vanambathina},
  doi       = {10.14569/IJACSA.2023.0140181},
  url       = {https://doi.org/10.14569/IJACSA.2023.0140181}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.