Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Cross-Modal Video Retrieval Model Based on Video-Text Dual Alignment

Author 1: Zhanbin Che Author 2: Huaili Guo
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 15, No. 2 · Published 2024

DOI: https://doi.org/10.14569/IJACSA.2024.0150232

Abstract

Cross-modal video retrieval remains a major challenge in natural language processing due to the natural semantic divide between video and text. Most approaches use a single encoder to extract video and text features separately, and train video-text pairs by means of contrastive learning, but this global alignment of video and text is prone to neglecting more fine-grained features of both. In addition, some studies focus only on profiling the video description text, ignoring the correlation relationship with the video. Therefore, this paper proposes a video retrieval method based on video-text alignment, which realizes both global and fine-grained alignment between video and text. For global alignment, the video and text are aligned by a single encoder and after linear projection; for fine-grained alignment, the video encoder is trained to align the video and text by masking some semantic information in the text. By experimentally comparing with multiple existing methods on MSR-VTT and MSVD datasets, the model achieves R@1 (recall at 1) metrics of 51.5% and 52.4% on MSR-VTT and MSVD datasets, respectively, which indicates that the proposed model can improve the efficiency of cross-modal video retrieval.

Keywords

How to Cite this Article

Che, Z., & Guo, H. (2024). Cross-Modal Video Retrieval Model Based on Video-Text Dual Alignment. International Journal of Advanced Computer Science and Applications, 15(2). https://doi.org/10.14569/IJACSA.2024.0150232

Che, Zhanbin, and Huaili Guo. "Cross-Modal Video Retrieval Model Based on Video-Text Dual Alignment." International Journal of Advanced Computer Science and Applications, vol. 15, no. 2, 2024, https://doi.org/10.14569/IJACSA.2024.0150232.

@article{Che2024,
  title     = {Cross-Modal Video Retrieval Model Based on Video-Text Dual Alignment},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {15},
  number    = {2},
  year      = {2024},
  publisher = {The Science and Information Organization},
  author    = {Zhanbin Che and Huaili Guo},
  doi       = {10.14569/IJACSA.2024.0150232},
  url       = {https://doi.org/10.14569/IJACSA.2024.0150232}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.