高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

隐空间掩码语义预测驱动的群体行为表征学习

张雅琪 李承阳 朱丽萍 李瑞娜

张雅琪, 李承阳, 朱丽萍, 李瑞娜. 隐空间掩码语义预测驱动的群体行为表征学习[J]. 电子与信息学报. doi: 10.11999/JEIT260730
引用本文: 张雅琪, 李承阳, 朱丽萍, 李瑞娜. 隐空间掩码语义预测驱动的群体行为表征学习[J]. 电子与信息学报. doi: 10.11999/JEIT260730
ZHANG Yaqi, LI Chengyang, ZHU Liping, LI Ruina. Group Activity Representation Learning via Masked Semantic Prediction in Latent Space[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260730
Citation: ZHANG Yaqi, LI Chengyang, ZHU Liping, LI Ruina. Group Activity Representation Learning via Masked Semantic Prediction in Latent Space[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260730

隐空间掩码语义预测驱动的群体行为表征学习

doi: 10.11999/JEIT260730 cstr: 32379.14.JEIT260730
基金项目: 国家自然科学基金青年科学基金(批准号:62502538),中国石油大学(北京)科研基金资助(批准号:2462025YJRC006)
详细信息
    作者简介:

    张雅琪:女,硕士研究生,研究方向为群体行为识别、世界模型, 2026211286@student.cup.edu.cn

    李承阳:男,特任岗位副教授,研究方向为复杂环境智能感知, cyangli@cup.edu.cn

    朱丽萍:女,副教授,研究方向为计算机视觉

    李瑞娜:女,硕士研究生,研究方向为群体行为识别、多模态学习

    通讯作者:

    李承阳 cyangli@cup.edu.cn(通信和一作均需唯一)

  • 中图分类号: TP391.4

Group Activity Representation Learning via Masked Semantic Prediction in Latent Space

Funds: National Natural Science Foundation of China (No.62502538), Science Foundation of China University of Petroleum, Beijing (No.2462025YJRC006)
  • 摘要: 针对现有无群体活动标签下的群体行为表征方法难以充分建模多主体连续交互的问题,该文提出一种隐空间掩码语义预测框架。该框架利用人物轨迹掩码构造预测任务,通过双路径Transformer编码可见个体的时空交互,并恢复被掩码个体的高层语义特征;进一步结合余弦语义对齐与跨视频InfoNCE约束,改善特征的语义一致性和实例级分布。在Volleyball数据集上,该方法的Hit@1达到86.5%,较基线提高1.7个百分点,并在mAP、K-NN等指标上取得改善,验证了所提框架在无群体活动标签条件下群体行为表征学习的有效性。在CAD数据集上,本文方法的Hit@1达到96.5%,高于基线的94.9%,进一步验证了所提框架的跨场景适用性。
  • 图  1  两种掩码预测式表征范式对比

    图  2  掩码语义特征预测架构图

    图  3  掩码人数对模型性能的影响

    图  4  基于 t-SNE 的特征分布可视化结果

    图  5  混淆矩阵与掩码语义预测可视化

    表  1  Baseline与本文方法在三个随机种子下的检索性能(均值±标准差,%)

    方法Hit@1Hit@2Hit@3mAP1NN3NN
    Baseline78.83±10.0486.24±5.8389.15±4.2852.71±8.4778.83±10.0480.68±8.72
    本文方法85.49±0.9189.68±0.0790.92±0.3558.37±1.1385.49±0.9186.54±0.59
    下载: 导出CSV

    表  2  群体活动检索性能对比(%)

    方法 骨干网络 Hit@1 Hit@2 Hit@3
    HiGCIN ResNet-18 50.0 66.3 74.5
    DIN VGG-16 57.0 73.1 81.1
    Dual-AI Inception-v3 64.4 76.5 82.0
    Group-DINOmics DINOv3 (ViT-L) 82.7 90.0 93.0
    Baseline VGG-16 84.8 89.6 91.8
    本文方法 VGG-16 86.5 89.9 91.5
    注:Baseline与本文方法采用相同VGG-16骨干网络和预训练设置,构成主要受控对比。Group-DINOmics采用DINOv3(ViT-L),全监督方法的监督信号也与本文不同,相关结果仅作跨配置参考。
    下载: 导出CSV

    表  3  特征判别性与邻域一致性对比(%)

    方法mAPmAP rank1NN3NN5NN
    Baseline57.0257.9984.8285.7985.86
    本文方法59.5860.6686.4687.2187.21
    注:表中结果均来自随机种子0对应的固定单次主实验模型,非三个随机种子的均值;多随机种子统计结果见表1
    下载: 导出CSV

    表  4  各类群体活动检索性能(Hit@1, %)

    类别Baseline本文方法
    r-set79.1776.56
    r-spike89.6089.60
    r-pass85.2490.48
    r-winpoint75.8681.61
    l-set88.1083.93
    l-spike86.5992.74
    l-pass84.9688.05
    l-winpoint85.2985.29
    下载: 导出CSV

    表  5  不同规模类别均衡检索库上的Hit@1(均值±标准差,%)

    每类样本数 检索库总规模 Baseline 本文方法
    5 40 75.78±0.98 76.83±2.34
    10 80 75.77±1.92 78.88±1.45
    20 160 78.92±1.49 80.40±2.41
    50 400 80.94±1.95 83.07±0.99
    100 800 83.01±0.75 84.76±0.84
    200 1,600 83.53±0.52 85.46±0.59
    208(最大平衡规模) 1,664 84.62±0.72 85.71±0.30
    完整检索库 3,493 84.82 86.46
    下载: 导出CSV

    表  6  CAD数据集上的群体活动检索结果(%)

    方法Hit@1Hit@2Hit@3mAP
    Baseline94.9095.5696.3495.32
    本文方法96.4796.8696.9987.83
    下载: 导出CSV

    表  7  对齐-对比联合约束的消融分析(%)

    ID Semantic Cosine InfoNCE Hit@1 mAP
    A × × × 84.82 57.02
    B × × 84.59 57.63
    C × 85.71 55.66
    D 86.46 59.58
    下载: 导出CSV

    表  8  不同预测目标对表征学习效果的影响(%)

    预测目标损失函数Hit@1mAP
    无(Baseline)84.8257.02
    底层视觉特征MSE85.0459.89
    高层语义特征MSE84.5957.63
    高层语义特征(本文方法)Cosine + InfoNCE86.4659.58
    下载: 导出CSV

    表  9  位置先验与坐标扰动实验(%)

    实验设置Hit@1mAP
    完整模型86.4659.58
    仅位置布局(无外观输入)39.1216.35
    去除预测端位置引导73.2235.90
    中心坐标扰动10%73.7534.24
    中心坐标扰动20%73.6734.38
    下载: 导出CSV

    表  10  超参数与训练批量大小的敏感性分析(%)

    变化因素 参数设置 Hit@1 mAP
    默认配置 $ \lambda $= 0.10, $ \alpha $= 0.10, $ \tau $ = 0.10,B = 8 86.46 59.58
    余弦损失权重 $ \lambda $=0.05 86.38 60.94
    对比损失权重 $ \alpha $=0.05 86.09 59.81
    温度参数 $ \tau $=0.20 86.08 59.63
    训练批量大小 B=4 85.04 55.45
    下载: 导出CSV
  • [1] WU Dekun, ZHAO He, BAO Xingce, et al. Sports video analysis on large-scale data[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 19–36. doi: 10.1007/978-3-031-19836-6_2.
    [2] RANASINGHE K, NASEER M, KHAN S, et al. Self-supervised video transformer[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 2864–2874. doi: 10.1109/CVPR52688.2022.00289.
    [3] EHSANPOUR M, ABEDIN A, SALEH F, et al. Joint learning of social groups, individuals action and sub-group activities in videos[C]. 16th European Conference Computer Vision – ECCV 2020, Glasgow, UK, 2020: 177–195. doi: 10.1007/978-3-030-58545-7_11.
    [4] FU Jun, LIU Jing, TIAN Haijie, et al. Dual attention network for scene segmentation[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 3141–3149. doi: 10.1109/CVPR.2019.00326.
    [5] TAMURA M, VISHWAKARMA R, and VENNELAKANTI R. Hunting group clues with transformers for social group activity recognition[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 19–35. doi: 10.1007/978-3-031-19772-7_2.
    [6] ZHOU Honglu, KADAV A, SHAMSIAN A, et al. COMPOSER: Compositional reasoning of group activity in videos with keypoint-only modality[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 249–266. doi: 10.1007/978-3-031-19833-5_15.
    [7] HE Kaiming, CHEN Xinlei, XIE Saining, et al. Masked autoencoders are scalable vision learners[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 15979–15988. doi: 10.1109/CVPR52688.2022.01553.
    [8] ASSRAN M, DUVAL Q, MISRA I, et al. Self-supervised learning from images with a joint-embedding predictive architecture[C]. Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, Canada, 2023: 15619–15629. doi: 10.1109/CVPR52729.2023.01499.
    [9] AMER M R, LEI Peng, and TODOROVIC S. HiRF: Hierarchical random field for collective activity recognition in videos[C]. 13th European Conference Computer Vision -- ECCV 2014, Zurich, Switzerland, 2014: 572–585. doi: 10.1007/978-3-319-10599-4_37.
    [10] ZHENG Yihao, WANG Zhuming, GU Ke, et al. Multi-scale motion-based relational reasoning for group activity recognition[J]. Engineering Applications of Artificial Intelligence, 2025, 139: 109570. doi: 10.1016/j.engappai.2024.109570.
    [11] DU Zexing and WANG Qing. Exploring global context and position-aware representation for group activity recognition[J]. Image and Vision Computing, 2024, 149: 105181. doi: 10.1016/j.imavis.2024.105181.
    [12] WU Jianchao, WANG Limin, WANG Li, et al. Learning actor relation graphs for group activity recognition[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 9956–9966. doi: 10.1109/CVPR.2019.01020.
    [13] HAN Mingfei, ZHANG D J, WANG Yali, et al. Dual-AI: Dual-path actor interaction learning for group activity recognition[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 2980–2989. doi: 10.1109/CVPR52688.2022.00300.
    [14] 朱丽萍, 吴祀霖, 陈晓禾, 等. 多尺度子群体交互关系下的群体行为识别方法[J]. 电子与信息学报, 2024, 46(5): 2228–2236. doi: 10.11999/JEIT231304.

    ZHU Liping, WU Silin, CHEN Xiaohe, et al. Group activity recognition under multi-scale sub-group interaction relationships[J]. Journal of Electronics & Information Technology, 2024, 46(5): 2228–2236. doi: 10.11999/JEIT231304.
    [15] 韩宗旺, 杨涵, 吴世青, 等. 时空自适应图卷积与Transformer结合的动作识别网络[J]. 电子与信息学报, 2024, 46(6): 2587–2595. doi: 10.11999/JEIT230551.

    HAN Zongwang, YANG Han, WU Shiqing, et al. Action recognition network combining spatio-temporal adaptive graph convolution and Transformer[J]. Journal of Electronics & Information Technology, 2024, 46(6): 2587–2595. doi: 10.11999/JEIT230551.
    [16] RAVITEJA CHAPPA N V, NGUYEN P, NELSON A H, et al. SoGAR: Self-supervised spatiotemporal attention-based social group activity recognition[J]. IEEE Access, 2025, 13: 33631–33642. doi: 10.1109/ACCESS.2025.3541986.
    [17] GUO Jie and GE Yongxin. Temporal contrastive and spatial enhancement coarse grained network for weakly supervised group activity recognition[J]. Engineering Applications of Artificial Intelligence, 2024, 133: 108115. doi: 10.1016/j.engappai.2024.108115.
    [18] IBRAHIM M S and MORI G. Hierarchical relational networks for group activity recognition and retrieval[C]. 15th European Conference Computer Vision – ECCV 2018, Munich, Germany, 2018: 742–758. doi: 10.1007/978-3-030-01219-9_44.
    [19] NAKATANI C, KAWASHIMA H, and UKITA N. Learning group activity features through person attribute prediction[C]. Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 18233–18242. doi: 10.1109/CVPR52733.2024.01726.
    [20] CHEN Ting, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations[C]. Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, 119: 1597–1607. (查阅网上资料, 未找到本条文献出版地信息, 请确认).
    [21] 孙中华, 吴双, 贾克斌, 等. 基于对比学习的动作识别研究综述[J]. 电子与信息学报, 2025, 47(8): 2473–2485. doi: 10.11999/JEIT250131.

    SUN Zhonghua, WU Shuang, JIA Kebin, et al. A review on action recognition based on contrastive learning[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2473–2485. doi: 10.11999/JEIT250131.
    [22] 刁文辉, 龚铄, 辛林霖, 等. 针对多模态遥感数据的自监督策略模型预训练方法[J]. 电子与信息学报, 2025, 47(6): 1658–1668. doi: 10.11999/JEIT241016.

    DIAO Wenhui, GONG Shuo, XIN Linlin, et al. A model pre-training method with self-supervised strategies for multimodal remote sensing data[J]. Journal of Electronics & Information Technology, 2025, 47(6): 1658–1668. doi: 10.11999/JEIT241016.
    [23] ASSRAN M, BARDES A, FAN D, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning[EB/OL]. https://arxiv.org/abs/2506.09985, 2025.
    [24] CHEN Delong, SHUKOR M, MOUTAKANNI T, et al. VL-JEPA: Joint embedding predictive architecture for vision-language[EB/OL]. https://arxiv.org/abs/2512.10942v1, 2025.
    [25] NAM H, LE LIDEC Q, MAES L, et al. Causal-JEPA: Learning world models through object-level latent interventions[EB/OL]. https://arxiv.org/abs/2602.11389, 2026.
    [26] IBRAHIM M S, MURALIDHARAN S, DENG Zhiwei, et al. A hierarchical deep temporal model for group activity recognition[C]. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, USA, 2016: 1971–1980. doi: 10.1109/CVPR.2016.217.
    [27] CHOI W, SHAHID K, and SAVARESE S. What are they doing?: Collective activity classification using spatio-temporal relationship among people[C]. 2009 IEEE 12th International Conference on Computer Vision Workshops, ICCV Workshops, Kyoto, Japan, 2009: 1282–1289. doi: 10.1109/ICCVW.2009.5457461.
    [28] YAN Rui, XIE Lingxi, TANG Jinhui, et al. HiGCIN: Hierarchical graph-based cross inference network for group activity recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(6): 6955–6968. doi: 10.1109/TPAMI.2020.3034233.
    [29] YUAN Hangjie, NI Dong, and WANG Mang. Spatio-temporal dynamic inference network for group activity recognition[C]. Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, Canada, 2021: 7456–7465. doi: 10.1109/ICCV48922.2021.00738.
    [30] TEZUKA R, NAKATANI C, and UKITA N. Group-DINOmics: Incorporating people dynamics into DINO for self-supervised group activity feature learning[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026. (查阅网上资料, 未找到本条文献出版地信息, 请确认).
  • 加载中
图(5) / 表(10)
计量
  • 文章访问数:  9
  • HTML全文浏览量:  1
  • PDF下载量:  0
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-06-02
  • 录用日期:  2026-09-17
  • 网络出版日期:  2026-09-24

目录

    /

    返回文章
    返回