高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

对抗优化三重软对比学习的多模态情感分析方法

康守强 张博浩 谢金宝

康守强, 张博浩, 谢金宝. 对抗优化三重软对比学习的多模态情感分析方法[J]. 电子与信息学报. doi: 10.11999/JEIT260255
引用本文: 康守强, 张博浩, 谢金宝. 对抗优化三重软对比学习的多模态情感分析方法[J]. 电子与信息学报. doi: 10.11999/JEIT260255
KANG Shouqiang, ZHANG Bohao, XIE Jinbao. Research on Multimodal Sentiment Analysis Method Based on Adversarial Optimization and Triplet Soft Contrastive Learning[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260255
Citation: KANG Shouqiang, ZHANG Bohao, XIE Jinbao. Research on Multimodal Sentiment Analysis Method Based on Adversarial Optimization and Triplet Soft Contrastive Learning[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260255

对抗优化三重软对比学习的多模态情感分析方法

doi: 10.11999/JEIT260255 cstr: 32379.14.JEIT260255
详细信息
    作者简介:

    康守强:男,教授,研究方向为非平稳信号处理,故障诊断、状态评估

    张博浩:男,硕士生,研究方向为自然语言处理

    谢金宝:男,副教授,研究方向为自然语言处理

    通讯作者:

    谢金宝 jbxpost@163.com

  • 中图分类号: TP391.4

Research on Multimodal Sentiment Analysis Method Based on Adversarial Optimization and Triplet Soft Contrastive Learning

  • 摘要: 多模态情感分析是人机交互领域的关键环节,但在实际应用中往往忽视了各模态之间的时间依赖关系,导致时序信息未能得到充分利用。同时,现有工作多通过简单的特征拼接来融合多模态信息,跨模态交互有限,容易引入冗余特征。针对上述问题,提出一种对抗优化三重软对比学习的多模态情感分析方法。本方法构建了三重软对比学习框架,即引入软对比学习策略,为样本赋予连续权重,以刻画样本间的细微差异。并在模态内、模态间以及时间序列三个层面开展对比任务,全面捕捉多模态数据在空间与时间维度中的关联信息。同时结合生成对抗网络,提升样本质量,缓解特征冗余问题,从而优化多模态情感特征的学习过程与分类效果。实验结果表明,在CMU-MOSI数据集上准确率和F1 值分别提升了1.3%和1%;而在CMU-MOSEI数据集上准确率和F1 值相较于最新模型均提升了0.7%,验证了所提模型在多模态情感分析任务中的有效性。
  • 图  1  多模态表达示例图

    图  2  对抗优化三重软对比学习模型结构图

    图  3  TCN网络结构图

    图  4  软硬对比学习差异示例图

    图  5  GAN网络结构图

    图  6  不同超参数配置结果对比图

    表  1  对照组参数配置表

    组别mτ
    10.00.1
    20.30.1
    30.30.2
    40.50.2
    下载: 导出CSV

    表  2  参数配置对比实验

    组别ACC2(%)F1(%)MAECorr
    185.70.870.490.90
    287.50.890.450.92
    386.80.880.470.91
    484.90.850.500.89
    下载: 导出CSV

    表  3  数据集数据量

    数据集训练集(个)验证集(个)测试集(个)总计(个)语言
    CMU-MOSI
    CMU-MOSEI
    1284
    16326
    229
    1871
    686
    4659
    2199
    22856
    英文
    英文
    下载: 导出CSV

    表  4  损失函数配置对比实验

    组别λintraλinterλtemλganACC2(%)
    A0.50.50.50.573.7
    B0.50.30.21.082.3
    C0.20.30.51.078.7
    D1.01.01.01.082.0
    E0.50.20.31.085.7
    F0.30.20.51.087.6
    下载: 导出CSV

    表  5  在CMU-MOSI数据集的实验结果

    模型MOSI
    Acc-2(%)F1(%)MAECorr
    TFN[28]80.874.50.9010.700
    LMF[29]82.582.50.9170.677
    Self-MM[30]85.985.90.7130.798
    CRIL[31]86.986.80.6950.812
    MAG-BERT[32]83.082.80.8710.559
    MISA[33]83.583.50.7900.593
    HyCon[22]86.486.40.6440.832
    ConFEDE[24]86.586.50.7080.796
    MSA-HCL[34]86.486.40.7260.789
    MLCL[35]86.486.30.7010.798
    Ours87.887.50.6050.804
    下载: 导出CSV

    表  6  在CMU-MOSEI数据集的实验结果

    模型MOSEI
    Acc-2(%)F1(%)MAECorr
    TFN[28]82.682.10.5930.700
    LMF[29]82.082.50.6230.677
    Self-MM[30]85.385.10.5300.765
    CRIL[31]86.286.10.5290.767
    MAG-BERT[32]85.685.00.6020.778
    MISA[33]85.585.30.5550.756
    HyCon[22]86.586.40.5900.792
    ConFEDE[24]86.886.90.5220.780
    MSA-HCL[34]86.485.90.5300.771
    MLCL[35]86.386.20.5510.756
    Ours87.5 87.60.506 0.740
    下载: 导出CSV

    表  7  MOSI数据集上模块消融实验结果

    模型MOSI
    Acc-2(%)F1(%)MAECorr
    w/o TemCL74.475.90.6420.801
    w/o InterCL79.978.60.6250.802
    w/o IntraCL60.781.30.6020.548
    w/o SoftCL86.287.40.5480.742
    w/o GAN80.881.60.5900.746
    Ours87.887.50.6050.804
    下载: 导出CSV

    表  8  MOSEI数据集上模块消融实验结果

    模型MOSEI
    Acc-2(%)F1(%)MAECorr
    w/o TemCL76.079.70.5970.845
    w/o InterCL81.383.40.5870.860
    w/o IntraCL62.085.10.5630.893
    w/o SoftCL85.583.90.4900.881
    w/o GAN81.786.80.5370.881
    Ours87.587.60.4500.920
    下载: 导出CSV

    表  9  MOSEI数据集测试鲁棒性实验结果

    模型TrainTestAcc-2(%)F1(%)MAECorr
    HyCon[22]MOSIMOSEI84.984.80.6760.823
    ConFEDE[24]MOSIMOSEI85.285.10.7350.781
    MSA-HCL[34]MOSIMOSEI85.084.90.7480.774
    MLCL[35]MOSIMOSEI85.185.00.7210.785
    OursMOSIMOSEI86.686.40.6420.807
    下载: 导出CSV

    表  10  MOSI数据集测试鲁棒性实验结果

    模型TrainTestAcc-2(%)F1(%)MAECorr
    HyCon[22]MOSEIMOSI87.187.00.5610.810
    ConFEDE[24]MOSEIMOSI87.487.50.4960.802
    MSA-HCL[34]MOSEIMOSI87.086.70.5030.795
    MLCL[35]MOSEIMOSI86.986.80.5200.782
    OursMOSEIMOSI88.388.40.4780.785
    下载: 导出CSV

    表  11  MOSI数据集和MOSEI数据集模态缺失鲁棒性实验

    模型MOSIMOSEI
    Acc-2(%)F1(%)MAECorrAcc-2(%)F1(%)MAECorr
    全模态87.887.50.6050.80487.587.60.5060.740
    缺失文本模态82.382.10.7200.69081.581.30.6420.635
    缺失语音模态86.986.80.6250.78586.286.30.5350.720
    缺失视觉模态85.785.60.6510.76084.884.90.5660.695
    下载: 导出CSV
  • [1] 黄辰, 刘会杰, 张龑, 等. 带全局噪声增强的多模态超图学习引导用于模态信息缺失情感分析[J]. 电子与信息学报, 2025, 47(12): 5192–5202. doi: 10.11999/JEIT250649.

    HUANG Chen, LIU Huijie, ZHANG Yan, et al. Multimodal hypergraph learning guidance with global noise enhancement for sentiment analysis under missing modality information[J]. Journal of Electronics & Information Technology, 2025, 47(12): 5192–5202. doi: 10.11999/JEIT250649.
    [2] 张乐, 陈岩松, 张雷瀚. 大模型特征增强与多层次交叉融合的多模态情感分析方法[J]. 数据分析与知识发现, 2025, 9(8): 47–58. doi: 10.11925/infotech.2096-3467.2024.0625.

    ZHANG Le, CHEN Yansong, and ZHANG Leihan. A multimodal sentiment analysis method based on LLM feature enhancement and multi-level cross-fusion[J]. Data Analysis and Knowledge Discovery, 2025, 9(8): 47–58. doi: 10.11925/infotech.2096-3467.2024.0625.
    [3] 陈杰, 马静, 李晓峰, 等. 基于DR-Transformer模型的多模态情感识别研究[J]. 情报科学, 2022, 40(3): 117–125. doi: 10.13833/j.issn.1007-7634.2022.03.015.

    CHEN Jie, MA Jing, LI Xiaofeng, et al. Multi-modal emotion recognition based on DR-Transformer model[J]. Information Science, 2022, 40(3): 117–125. doi: 10.13833/j.issn.1007-7634.2022.03.015.
    [4] 林宜山, 左景, 卢树华. 基于音视频特征优化与跨模态Transformer的多模态情感分析[J]. 北京航空航天大学学报, 2026, 52(6): 2219–2228. doi: 10.13700/j.bh.1001-5965.2024.0247.

    LIN Yishan, ZUO Jing, and LU Shuhua. A multimodal sentiment analysis based on audio and video features optimization and cross-modal Transformer[J]. Journal of Beijing University of Aeronautics and Astronautics, 2026, 52(6): 2219–2228. doi: 10.13700/j.bh.1001-5965.2024.0247.
    [5] WU Yujin, DAOUDI M, and AMAD A. Transformer-based self-supervised multimodal representation learning for wearable emotion recognition[J]. IEEE Transactions on Affective Computing, 2024, 15(1): 157–172. doi: 10.1109/TAFFC.2023.3263907.
    [6] CAI Yujian, LI Xingguang, ZHANG Yingyu, et al. Multimodal sentiment analysis based on multi-layer feature fusion and multi-task learning[J]. Scientific Reports, 2025, 15(1): 2126. doi: 10.1038/s41598-025-85859-6.
    [7] 冯广, 周垣桦, 钟婷, 等. 结合自适应特征加权与权值优化策略的多模态情感分析[J]. 计算机工程与应用, 2026, 62(6): 194–204. doi: 10.3778/j.issn.1002-8331.2501-0164.

    FENG Guang, ZHOU Yuanhua, ZHONG Ting, et al. Multimodal sentiment analysis combining adaptive feature weighting and weight optimization strategy[J]. Computer Engineering and Applications, 2026, 62(6): 194–204. doi: 10.3778/j.issn.1002-8331.2501-0164.
    [8] 刘佳, 宋泓, 陈大鹏, 等. 非语言信息增强和对比学习的多模态情感分析模型[J]. 电子与信息学报, 2024, 46(8): 3372–3381. doi: 10.11999/JEIT231274.

    LIU Jia, SONG Hong, CHEN Dapeng, et al. A multimodal sentiment analysis model enhanced with non-verbal information and contrastive learning[J]. Journal of Electronics & Information Technology, 2024, 46(8): 3372–3381. doi: 10.11999/JEIT231274.
    [9] SCHULLER B, RIGOLL G, and LANG M. Speech emotion recognition: Features and classification[J]. Speech Communication, 2009, 51(10): 975–982. doi: 10.1016/j.dsp.2012.05.007. (查阅网上资料,未找到本条文献信息,请确认).
    [10] YE Jiaxin, WEN Xincheng, WEI Yujie, et al. Temporal modeling matters: A novel temporal emotional modeling approach for speech emotion recognition[C]. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Rhodes Island, Greece, 2023: 1–5. doi: 10.1109/ICASSP49357.2023.10096370.
    [11] MIKOLOV T, CHEN Kai, CORRADO G, et al. Efficient estimation of word representations in vector space[C]. International Conference on Learning Representations, Scottsdale, USA, 2013: 1301–3781. (查阅网上资料, 未找到本条文页码, 请确认).
    [12] DEVLIN J, CHANG Mingwei, LEE K, et al. BERT: Pre-training of deep bidirectional Transformers for language understanding[C]. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, Minnesota, 2019: 4171–4186. doi: 10.18653/v1/N19-1423.
    [13] WANG Yabing, HUANG Guimin, LI Maolin, et al. Automatically constructing a fine-grained sentiment lexicon for sentiment analysis[J]. Cognitive Computation, 2023, 15(1): 254–271. doi: 10.1007/s12559-022-10043-1.
    [14] JASSIM M A, ABD D H, and OMRI M N. A survey of sentiment analysis from film critics based on machine learning, lexicon and hybridization[J]. Neural Computing and Applications, 2023, 35(13): 9437–9461. doi: 10.1007/s00521-023-08359-6.
    [15] 曹银妮, 韩虎, 黄明伟, 等. 基于多视角融合表示的多模态方面级情感分析模型[J]. 数据分析与知识发现, 2025, 9(10): 54–67. doi: 10.11925/infotech.2096-3467.2024.1114.

    CAO Yinni, HAN Hu, HUANG Mingwei, et al. Multi-modal aspect-level sentiment analysis model with multi-view fusion representation[J]. Data Analysis and Knowledge Discovery, 2025, 9(10): 54–67. doi: 10.11925/infotech.2096-3467.2024.1114.
    [16] 赵川斌, 许伟华, 林博, 等. 融合视觉的多模态通信感知一体化关键技术及原型验证[J]. 电子与信息学报, 2026, 48(2): 487–498. doi: 10.11999/JEIT250685.

    ZHAO Chuanbin, XU Weihua, LIN Bo, et al. Vision enabled multimodal integrated sensing and communications: Key technologies and prototype validation[J]. Journal of Electronics & Information Technology, 2026, 48(2): 487–498. doi: 10.11999/JEIT250685.
    [17] YOU Quanzeng, JIN Hailin, and LUO Jiebo. Visual sentiment analysis by attending on local image regions[C]. Proceedings of the 31st AAAI Conference on Artificial Intelligence, San Francisco, USA, 2017: 231–237.
    [18] LIU Yunze, FAN Qingnan, ZHANG Shanghang, et al. Contrastive multimodal fusion with TupleInfoNCE[C]. Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 734–743. doi: 10.1109/ICCV48922.2021.00079.
    [19] YANG Liu, WU Zhenjie, HONG Junkun, et al. MCL: A contrastive learning method for multimodal data fusion in violence detection[J]. IEEE Signal Processing Letters, 2023, 30: 408–412. doi: 10.1109/LSP.2022.3227818.
    [20] GRILL J B, STRUB F, ALTCHÉ F, et al. Bootstrap your own latent a new approach to self-supervised learning[C]. Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2020: 1786.
    [21] WANG Huiru, LI Xiuhong, REN Zenyu, et al. Multimodal sentiment analysis representations learning via contrastive learning with condense attention fusion[J]. Sensors, 2023, 23(5): 2679. doi: 10.3390/s23052679.
    [22] MAI Sijie, ZENG Ying, ZHENG Shuangjia, et al. Hybrid contrastive learning of tri-modal representation for multimodal sentiment analysis[J]. IEEE Transactions on Affective Computing, 2023, 14(3): 2276–2289. doi: 10.1109/TAFFC.2022.3172360.
    [23] QUAN Zhibang, SUN Tao, SU Mengli, et al. Multimodal sentiment analysis based on nonverbal representation optimization network and contrastive interaction learning[C]. Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics, Prague, Czech Republic, 2022: 3086–3091. doi: 10.1109/SMC53654.2022.9945514.
    [24] YANG Jiuding, YU Yakun, NIU Di, et al. ConFEDE: Contrastive feature decomposition for multimodal sentiment analysis[C]. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, Canada, 2023: 7617–7630. doi: 10.18653/v1/2023.acl-long.421.
    [25] WANG Senzhang, YAN Hao, DU Jinlong, et al. Adversarial hard negative generation for complementary graph contrastive learning[C]. SIAM International Conference on Data Mining, Austin, USA, 2023: 163–171. doi: 10.1137/1.9781611977653.ch19. (查阅网上资料,未找到本条文献出版地,请确认).
    [26] ZADEH A, ZELLERS R, PINCUS E, et al. MOSI: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos[EB/OL]. https://arxiv.org/abs/1606.06259, 2016.
    [27] ZADEH A, LIANG P P, PORIA S, et al. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph[C]. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia, 2018: 2236–2246. doi: 10.18653/v1/P18-1208.
    [28] ZADEH A, CHEN Minghai, PORIA S, et al. Tensor fusion network for multimodal sentiment analysis[C]. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark, 2017: 1103–1114. doi: 10.18653/v1/D17-1115.
    [29] LIU Zhun, SHEN Ying, LAKSHMINARASIMHAN V B, et al. Efficient low-rank multimodal fusion with modality-specific factors[C]. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia, 2018: 2247–2256. doi: 10.18653/v1/P18-1209.
    [30] YU Wenmeng, XU Hua, YUAN Ziqi, et al. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis[C]. Proceedings of the 35th AAAI Conference on Artificial Intelligence, 2021: 10790–10797. doi: 10.1609/aaai.v35i12.17289. (查阅网上资料,未找到本条文献出版地,请确认).
    [31] HUANG Jian, JI Yanli, YANG Yang, et al. Cross-modality representation interactive learning for multimodal sentiment analysis[C]. Proceedings of the 31st ACM International Conference on Multimedia, Ottawa, Canada, 2023: 426–434. doi: 10.1145/3581783.3612295.
    [32] RAHMAN W, HASAN K, LEE S, et al. Integrating multimodal information in large pretrained transformers[C]. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020: 2359–2369. doi: 10.18653/v1/2020.acl-main.214. (查阅网上资料,未找到本条文献出版地,请确认).
    [33] HAZARIKA D, ZIMMERMANN R, and PORIA S. MISA: Modality-invariant and -specific representations for multimodal sentiment analysis[C]. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, USA, 2020: 1122–1131.
    [34] ZHAO Wang, ZHANG Yong, HUA Qiang, et al. MSA-HCL: Multimodal sentiment analysis model with hybrid contrastive learning[J]. Mathematical Foundations of Computing, 2025, 8(3): 433–447. doi: 10.3934/mfc.2024017.
    [35] ZHUANG Yan, BAI Wei, ZHANG Yanru, et al. Multi-level contrastive learning for multimodal sentiment analysis[J]. IEEE Transactions on Multimedia, 2025, 27: 9044–9058. doi: 10.1109/TMM.2025.3613116.
  • 加载中
图(6) / 表(11)
计量
  • 文章访问数:  12
  • HTML全文浏览量:  1
  • PDF下载量:  0
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-03-06
  • 修回日期:  2026-09-13
  • 录用日期:  2026-09-13
  • 网络出版日期:  2026-09-18

目录

    /

    返回文章
    返回