Group Activity Representation Learning via Masked Semantic Prediction in Latent Space
-
摘要: 针对现有无群体活动标签下的群体行为表征方法难以充分建模多主体连续交互的问题,该文提出一种隐空间掩码语义预测框架。该框架利用人物轨迹掩码构造预测任务,通过双路径Transformer编码可见个体的时空交互,并恢复被掩码个体的高层语义特征;进一步结合余弦语义对齐与跨视频InfoNCE约束,改善特征的语义一致性和实例级分布。在Volleyball数据集上,该方法的Hit@1达到86.5%,较基线提高1.7个百分点,并在mAP、K-NN等指标上取得改善,验证了所提框架在无群体活动标签条件下群体行为表征学习的有效性。在CAD数据集上,本文方法的Hit@1达到96.5%,高于基线的94.9%,进一步验证了所提框架的跨场景适用性。Abstract:
Objective Group activity recognition aims to understand collective behavior among interacting individuals. Unlike individual action recognition, it depends on both actor states and their spatial and temporal interactions. Most methods rely on group activity annotations and individual action labels, making annotation expensive. Recent representation-learning approaches reduce reliance on group labels, but their objectives are often transferred from generic visual learning. Pixel- or low-level feature reconstruction may emphasize appearance unrelated to collective behavior, whereas discrete classification compresses evolving interactions into a few categories. This study investigates group activity representation learning without group labels during representation training. Individual action labels from the standard training split are retained as auxiliary supervision; thus, the setting is not fully label-free or purely self-supervised. Inspired by joint-embedding predictive architectures, we propose a framework that predicts masked actors’ high-level latent semantics and encourages group representations to encode relationships between visible interactions and missing actor states. Methods The framework contains three components: actor feature extraction and masking, latent semantic context reasoning, and joint alignment–contrastive optimization. An ImageNet-pretrained VGG-16 extracts multi-scale feature maps, RoIAlign obtains actor features from trajectory boxes, and a fully connected layer with LayerNorm and ReLU maps them into a 1,024-dimensional semantic space. The backbone and linear projection are initialized by Stage-1 individual action prediction on Volleyball using nine action classes without group activity labels. In Stage 2, they are jointly optimized with the context encoder and predictor. For each 10-frame clip, three trajectories are masked consistently. The complete semantic tensor provides targets, while a binary mask excludes selected actors from spatial aggregation and group pooling. A dual-path Transformer models space-to-time and time-to-space interactions, producing a 2,048-dimensional group representation. Combined with target-position encoding, a three-layer predictor recovers masked semantics. Stop-gradient is applied only to targets, without a teacher network. Cosine loss aligns semantic directions, while cross-video InfoNCE treats visible actors from other clips as unfiltered instance-level negatives. Auxiliary action classification is retained. The loss weights are 1, 0.10, and 0.10, with a temperature of 0.10. For CAD, the framework uses an Inception-v3 backbone and the standard split. Results and Discussions Experiments on Volleyball used retrieval and nearest-neighbor evaluation. In a controlled single-run comparison with the same VGG-16 backbone, Stage-1 initialization, data split, and evaluation protocol, the proposed method achieved 86.46% Hit@1 and 59.58% mAP, versus 84.82% and 57.02% for the baseline. Hit@2 increased from 89.6% to 89.9%, whereas Hit@3 decreased from 91.8% to 91.5%, indicating nonuniform gains across retrieval depths. On CAD, after dataset-specific retraining, Hit@1 increased from 94.90% to 96.47%, whereas mAP decreased from 95.32% to 87.83%, indicating improved local discrimination but weaker global ranking. Across three random seeds, the method obtained 85.49%±0.91% Hit@1 and 58.37%±1.13% mAP, with lower variation than the baseline. With class-balanced galleries of 5–208 samples per class, it exceeded the baseline in Hit@1 by 1.05–3.11 percentage points. Ablations showed that semantic prediction and cosine alignment improve nearest-neighbor discrimination, while InfoNCE improves global ranking. Three masked trajectories best balanced visible context and prediction difficulty. A post hoc analysis found that cross-video pairs sharing the same video-level activity label comprised 13.68%±0.12% of candidates. This is a video-class collision rate, not the true actor-level semantic false-negative rate. Position-only input produced 39.12% Hit@1, showing that court layout alone is insufficient. Removing positional guidance retained 73.22% Hit@1, indicating that appearance and multi-actor context remain informative, although coordinate perturbation revealed sensitivity to localization errors. The Davies–Bouldin index decreased from 3.2339 to2.8287 , supporting improved compactness and separation. Group-DINOmics uses DINOv3 with a ViT-L backbone and different pretraining resources; therefore, its result is a cross-configuration reference rather than a controlled comparison.Conclusions Masked high-level semantic prediction provides an effective objective for learning interaction-aware group representations when group activity labels are excluded from representation training. The proposed framework uses visible actors to infer missing actor semantics, while cosine alignment and instance-level contrastive learning jointly improve local discrimination and global ranking. The results also show that individual action supervision supplies an important semantic anchor and that positional information is helpful but insufficient on its own. The CAD results further support applicability beyond sports scenes. The current method remains sensitive to localization errors and is less effective for several short-duration or weakly collaborative activities. Future work will investigate more robust spatial relation modeling, finer temporal dynamics, and stronger backbones under unified settings. -
表 1 Baseline与本文方法在三个随机种子下的检索性能(均值±标准差,%)
方法 Hit@1 Hit@2 Hit@3 mAP 1NN 3NN Baseline 78.83±10.04 86.24±5.83 89.15±4.28 52.71±8.47 78.83±10.04 80.68±8.72 本文方法 85.49±0.91 89.68±0.07 90.92±0.35 58.37±1.13 85.49±0.91 86.54±0.59 表 2 群体活动检索性能对比(%)
方法 骨干网络 Hit@1 Hit@2 Hit@3 HiGCIN ResNet-18 50.0 66.3 74.5 DIN VGG-16 57.0 73.1 81.1 Dual-AI Inception-v3 64.4 76.5 82.0 Group-DINOmics DINOv3 (ViT-L) 82.7 90.0 93.0 Baseline VGG-16 84.8 89.6 91.8 本文方法 VGG-16 86.5 89.9 91.5 注:Baseline与本文方法采用相同VGG-16骨干网络和预训练设置,构成主要受控对比。Group-DINOmics采用DINOv3(ViT-L),全监督方法的监督信号也与本文不同,相关结果仅作跨配置参考。 表 3 特征判别性与邻域一致性对比(%)
方法 mAP mAP rank 1NN 3NN 5NN Baseline 57.02 57.99 84.82 85.79 85.86 本文方法 59.58 60.66 86.46 87.21 87.21 注:表中结果均来自随机种子0对应的固定单次主实验模型,非三个随机种子的均值;多随机种子统计结果见表1。 表 4 各类群体活动检索性能(Hit@1, %)
类别 Baseline 本文方法 r-set 79.17 76.56 r-spike 89.60 89.60 r-pass 85.24 90.48 r-winpoint 75.86 81.61 l-set 88.10 83.93 l-spike 86.59 92.74 l-pass 84.96 88.05 l-winpoint 85.29 85.29 表 5 不同规模类别均衡检索库上的Hit@1(均值±标准差,%)
每类样本数 检索库总规模 Baseline 本文方法 5 40 75.78±0.98 76.83±2.34 10 80 75.77±1.92 78.88±1.45 20 160 78.92±1.49 80.40±2.41 50 400 80.94±1.95 83.07±0.99 100 800 83.01±0.75 84.76±0.84 200 1,600 83.53±0.52 85.46±0.59 208(最大平衡规模) 1,664 84.62±0.72 85.71±0.30 完整检索库 3,493 84.82 86.46 表 6 CAD数据集上的群体活动检索结果(%)
方法 Hit@1 Hit@2 Hit@3 mAP Baseline 94.90 95.56 96.34 95.32 本文方法 96.47 96.86 96.99 87.83 表 7 对齐-对比联合约束的消融分析(%)
ID Semantic Cosine InfoNCE Hit@1 mAP A × × × 84.82 57.02 B √ × × 84.59 57.63 C √ √ × 85.71 55.66 D √ √ √ 86.46 59.58 表 8 不同预测目标对表征学习效果的影响(%)
预测目标 损失函数 Hit@1 mAP 无(Baseline) — 84.82 57.02 底层视觉特征 MSE 85.04 59.89 高层语义特征 MSE 84.59 57.63 高层语义特征(本文方法) Cosine + InfoNCE 86.46 59.58 表 9 位置先验与坐标扰动实验(%)
实验设置 Hit@1 mAP 完整模型 86.46 59.58 仅位置布局(无外观输入) 39.12 16.35 去除预测端位置引导 73.22 35.90 中心坐标扰动10% 73.75 34.24 中心坐标扰动20% 73.67 34.38 表 10 超参数与训练批量大小的敏感性分析(%)
变化因素 参数设置 Hit@1 mAP 默认配置 $ \lambda $= 0.10, $ \alpha $= 0.10, $ \tau $ = 0.10,B = 8 86.46 59.58 余弦损失权重 $ \lambda $=0.05 86.38 60.94 对比损失权重 $ \alpha $=0.05 86.09 59.81 温度参数 $ \tau $=0.20 86.08 59.63 训练批量大小 B=4 85.04 55.45 -
[1] WU Dekun, ZHAO He, BAO Xingce, et al. Sports video analysis on large-scale data[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 19–36. doi: 10.1007/978-3-031-19836-6_2. [2] RANASINGHE K, NASEER M, KHAN S, et al. Self-supervised video transformer[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 2864–2874. doi: 10.1109/CVPR52688.2022.00289. [3] EHSANPOUR M, ABEDIN A, SALEH F, et al. Joint learning of social groups, individuals action and sub-group activities in videos[C]. 16th European Conference Computer Vision – ECCV 2020, Glasgow, UK, 2020: 177–195. doi: 10.1007/978-3-030-58545-7_11. [4] FU Jun, LIU Jing, TIAN Haijie, et al. Dual attention network for scene segmentation[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 3141–3149. doi: 10.1109/CVPR.2019.00326. [5] TAMURA M, VISHWAKARMA R, and VENNELAKANTI R. Hunting group clues with transformers for social group activity recognition[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 19–35. doi: 10.1007/978-3-031-19772-7_2. [6] ZHOU Honglu, KADAV A, SHAMSIAN A, et al. COMPOSER: Compositional reasoning of group activity in videos with keypoint-only modality[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 249–266. doi: 10.1007/978-3-031-19833-5_15. [7] HE Kaiming, CHEN Xinlei, XIE Saining, et al. Masked autoencoders are scalable vision learners[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 15979–15988. doi: 10.1109/CVPR52688.2022.01553. [8] ASSRAN M, DUVAL Q, MISRA I, et al. Self-supervised learning from images with a joint-embedding predictive architecture[C]. Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, Canada, 2023: 15619–15629. doi: 10.1109/CVPR52729.2023.01499. [9] AMER M R, LEI Peng, and TODOROVIC S. HiRF: Hierarchical random field for collective activity recognition in videos[C]. 13th European Conference Computer Vision -- ECCV 2014, Zurich, Switzerland, 2014: 572–585. doi: 10.1007/978-3-319-10599-4_37. [10] ZHENG Yihao, WANG Zhuming, GU Ke, et al. Multi-scale motion-based relational reasoning for group activity recognition[J]. Engineering Applications of Artificial Intelligence, 2025, 139: 109570. doi: 10.1016/j.engappai.2024.109570. [11] DU Zexing and WANG Qing. Exploring global context and position-aware representation for group activity recognition[J]. Image and Vision Computing, 2024, 149: 105181. doi: 10.1016/j.imavis.2024.105181. [12] WU Jianchao, WANG Limin, WANG Li, et al. Learning actor relation graphs for group activity recognition[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 9956–9966. doi: 10.1109/CVPR.2019.01020. [13] HAN Mingfei, ZHANG D J, WANG Yali, et al. Dual-AI: Dual-path actor interaction learning for group activity recognition[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 2980–2989. doi: 10.1109/CVPR52688.2022.00300. [14] 朱丽萍, 吴祀霖, 陈晓禾, 等. 多尺度子群体交互关系下的群体行为识别方法[J]. 电子与信息学报, 2024, 46(5): 2228–2236. doi: 10.11999/JEIT231304.ZHU Liping, WU Silin, CHEN Xiaohe, et al. Group activity recognition under multi-scale sub-group interaction relationships[J]. Journal of Electronics & Information Technology, 2024, 46(5): 2228–2236. doi: 10.11999/JEIT231304. [15] 韩宗旺, 杨涵, 吴世青, 等. 时空自适应图卷积与Transformer结合的动作识别网络[J]. 电子与信息学报, 2024, 46(6): 2587–2595. doi: 10.11999/JEIT230551.HAN Zongwang, YANG Han, WU Shiqing, et al. Action recognition network combining spatio-temporal adaptive graph convolution and Transformer[J]. Journal of Electronics & Information Technology, 2024, 46(6): 2587–2595. doi: 10.11999/JEIT230551. [16] RAVITEJA CHAPPA N V, NGUYEN P, NELSON A H, et al. SoGAR: Self-supervised spatiotemporal attention-based social group activity recognition[J]. IEEE Access, 2025, 13: 33631–33642. doi: 10.1109/ACCESS.2025.3541986. [17] GUO Jie and GE Yongxin. Temporal contrastive and spatial enhancement coarse grained network for weakly supervised group activity recognition[J]. Engineering Applications of Artificial Intelligence, 2024, 133: 108115. doi: 10.1016/j.engappai.2024.108115. [18] IBRAHIM M S and MORI G. Hierarchical relational networks for group activity recognition and retrieval[C]. 15th European Conference Computer Vision – ECCV 2018, Munich, Germany, 2018: 742–758. doi: 10.1007/978-3-030-01219-9_44. [19] NAKATANI C, KAWASHIMA H, and UKITA N. Learning group activity features through person attribute prediction[C]. Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 18233–18242. doi: 10.1109/CVPR52733.2024.01726. [20] CHEN Ting, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations[C]. Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, 119: 1597–1607. (查阅网上资料, 未找到本条文献出版地信息, 请确认). [21] 孙中华, 吴双, 贾克斌, 等. 基于对比学习的动作识别研究综述[J]. 电子与信息学报, 2025, 47(8): 2473–2485. doi: 10.11999/JEIT250131.SUN Zhonghua, WU Shuang, JIA Kebin, et al. A review on action recognition based on contrastive learning[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2473–2485. doi: 10.11999/JEIT250131. [22] 刁文辉, 龚铄, 辛林霖, 等. 针对多模态遥感数据的自监督策略模型预训练方法[J]. 电子与信息学报, 2025, 47(6): 1658–1668. doi: 10.11999/JEIT241016.DIAO Wenhui, GONG Shuo, XIN Linlin, et al. A model pre-training method with self-supervised strategies for multimodal remote sensing data[J]. Journal of Electronics & Information Technology, 2025, 47(6): 1658–1668. doi: 10.11999/JEIT241016. [23] ASSRAN M, BARDES A, FAN D, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning[EB/OL]. https://arxiv.org/abs/2506.09985, 2025. [24] CHEN Delong, SHUKOR M, MOUTAKANNI T, et al. VL-JEPA: Joint embedding predictive architecture for vision-language[EB/OL]. https://arxiv.org/abs/2512.10942v1, 2025. [25] NAM H, LE LIDEC Q, MAES L, et al. Causal-JEPA: Learning world models through object-level latent interventions[EB/OL]. https://arxiv.org/abs/2602.11389, 2026. [26] IBRAHIM M S, MURALIDHARAN S, DENG Zhiwei, et al. A hierarchical deep temporal model for group activity recognition[C]. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, USA, 2016: 1971–1980. doi: 10.1109/CVPR.2016.217. [27] CHOI W, SHAHID K, and SAVARESE S. What are they doing?: Collective activity classification using spatio-temporal relationship among people[C]. 2009 IEEE 12th International Conference on Computer Vision Workshops, ICCV Workshops, Kyoto, Japan, 2009: 1282–1289. doi: 10.1109/ICCVW.2009.5457461. [28] YAN Rui, XIE Lingxi, TANG Jinhui, et al. HiGCIN: Hierarchical graph-based cross inference network for group activity recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(6): 6955–6968. doi: 10.1109/TPAMI.2020.3034233. [29] YUAN Hangjie, NI Dong, and WANG Mang. Spatio-temporal dynamic inference network for group activity recognition[C]. Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, Canada, 2021: 7456–7465. doi: 10.1109/ICCV48922.2021.00738. [30] TEZUKA R, NAKATANI C, and UKITA N. Group-DINOmics: Incorporating people dynamics into DINO for self-supervised group activity feature learning[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026. (查阅网上资料, 未找到本条文献出版地信息, 请确认). -
下载: