Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Cross-Modal Sentiment Analysis Based on CLIP Image-Text Attention Interaction

Author 1: Xintao Lu Author 2: Yonglong Ni Author 3: Zuohua Ding
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 15, No. 2 · Published 2024 · Cited by 11

DOI: https://doi.org/10.14569/IJACSA.2024.0150290

Abstract

Multimodal sentiment analysis is a traditional text-based sentiment analysis technique. However, the field of multi-modal sentiment analysis still faces challenges such as inconsistent cross-modal feature information, poor interaction capabilities, and insufficient feature fusion. To address these issues, this paper proposes a cross-modal sentiment model based on CLIP image-text attention interaction. The model utilizes pre-trained ResNet50 and RoBERTa to extract primary image-text features. After contrastive learning with the CLIP model, it employs a multi-head attention mechanism for cross-modal feature interaction to enhance information exchange between different modalities. Subsequently, a cross-modal gating module is used to fuse feature networks, combining features at different levels while controlling feature weights. The final output is fed into a fully connected layer for sentiment recognition. Comparative experiments are conducted on the publicly available datasets MSVA-Single and MSVA-Multiple. The experimental results demonstrate that our model achieved accuracy rates of 75.38%and 73.95% , and F1-scores of 75.21% and 73.83% on the mentioned datasets, respectively. This indicates that the proposed approach exhibits higher generalization and robustness compared to existing sentiment analysis models.

Keywords

How to Cite this Article

Lu, X., Ni, Y., & Ding, Z. (2024). Cross-Modal Sentiment Analysis Based on CLIP Image-Text Attention Interaction. International Journal of Advanced Computer Science and Applications, 15(2). https://doi.org/10.14569/IJACSA.2024.0150290

Lu, Xintao, et al.. "Cross-Modal Sentiment Analysis Based on CLIP Image-Text Attention Interaction." International Journal of Advanced Computer Science and Applications, vol. 15, no. 2, 2024, https://doi.org/10.14569/IJACSA.2024.0150290.

@article{Lu2024,
  title     = {Cross-Modal Sentiment Analysis Based on CLIP Image-Text Attention Interaction},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {15},
  number    = {2},
  year      = {2024},
  publisher = {The Science and Information Organization},
  author    = {Xintao Lu and Yonglong Ni and Zuohua Ding},
  doi       = {10.14569/IJACSA.2024.0150290},
  url       = {https://doi.org/10.14569/IJACSA.2024.0150290}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.