Advanced Search
Turn off MathJax
Article Contents
ZHANG Yaqi, LI Chengyang, ZHU Liping, LI Ruina. Group Activity Representation Learning via Masked Semantic Prediction in Latent Space[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260730
Citation: ZHANG Yaqi, LI Chengyang, ZHU Liping, LI Ruina. Group Activity Representation Learning via Masked Semantic Prediction in Latent Space[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260730

Group Activity Representation Learning via Masked Semantic Prediction in Latent Space

doi: 10.11999/JEIT260730 cstr: 32379.14.JEIT260730
Funds:  National Natural Science Foundation of China (No.62502538), Science Foundation of China University of Petroleum, Beijing (No.2462025YJRC006)
  • Received Date: 2026-06-02
  • Accepted Date: 2026-09-17
  • Available Online: 2026-09-24
  •   Objective  Group activity recognition aims to understand collective behavior among interacting individuals. Unlike individual action recognition, it depends on both actor states and their spatial and temporal interactions. Most methods rely on group activity annotations and individual action labels, making annotation expensive. Recent representation-learning approaches reduce reliance on group labels, but their objectives are often transferred from generic visual learning. Pixel- or low-level feature reconstruction may emphasize appearance unrelated to collective behavior, whereas discrete classification compresses evolving interactions into a few categories. This study investigates group activity representation learning without group labels during representation training. Individual action labels from the standard training split are retained as auxiliary supervision; thus, the setting is not fully label-free or purely self-supervised. Inspired by joint-embedding predictive architectures, we propose a framework that predicts masked actors’ high-level latent semantics and encourages group representations to encode relationships between visible interactions and missing actor states.  Methods  The framework contains three components: actor feature extraction and masking, latent semantic context reasoning, and joint alignment–contrastive optimization. An ImageNet-pretrained VGG-16 extracts multi-scale feature maps, RoIAlign obtains actor features from trajectory boxes, and a fully connected layer with LayerNorm and ReLU maps them into a 1,024-dimensional semantic space. The backbone and linear projection are initialized by Stage-1 individual action prediction on Volleyball using nine action classes without group activity labels. In Stage 2, they are jointly optimized with the context encoder and predictor. For each 10-frame clip, three trajectories are masked consistently. The complete semantic tensor provides targets, while a binary mask excludes selected actors from spatial aggregation and group pooling. A dual-path Transformer models space-to-time and time-to-space interactions, producing a 2,048-dimensional group representation. Combined with target-position encoding, a three-layer predictor recovers masked semantics. Stop-gradient is applied only to targets, without a teacher network. Cosine loss aligns semantic directions, while cross-video InfoNCE treats visible actors from other clips as unfiltered instance-level negatives. Auxiliary action classification is retained. The loss weights are 1, 0.10, and 0.10, with a temperature of 0.10. For CAD, the framework uses an Inception-v3 backbone and the standard split.  Results and Discussions  Experiments on Volleyball used retrieval and nearest-neighbor evaluation. In a controlled single-run comparison with the same VGG-16 backbone, Stage-1 initialization, data split, and evaluation protocol, the proposed method achieved 86.46% Hit@1 and 59.58% mAP, versus 84.82% and 57.02% for the baseline. Hit@2 increased from 89.6% to 89.9%, whereas Hit@3 decreased from 91.8% to 91.5%, indicating nonuniform gains across retrieval depths. On CAD, after dataset-specific retraining, Hit@1 increased from 94.90% to 96.47%, whereas mAP decreased from 95.32% to 87.83%, indicating improved local discrimination but weaker global ranking. Across three random seeds, the method obtained 85.49%±0.91% Hit@1 and 58.37%±1.13% mAP, with lower variation than the baseline. With class-balanced galleries of 5–208 samples per class, it exceeded the baseline in Hit@1 by 1.05–3.11 percentage points. Ablations showed that semantic prediction and cosine alignment improve nearest-neighbor discrimination, while InfoNCE improves global ranking. Three masked trajectories best balanced visible context and prediction difficulty. A post hoc analysis found that cross-video pairs sharing the same video-level activity label comprised 13.68%±0.12% of candidates. This is a video-class collision rate, not the true actor-level semantic false-negative rate. Position-only input produced 39.12% Hit@1, showing that court layout alone is insufficient. Removing positional guidance retained 73.22% Hit@1, indicating that appearance and multi-actor context remain informative, although coordinate perturbation revealed sensitivity to localization errors. The Davies–Bouldin index decreased from 3.2339 to 2.8287, supporting improved compactness and separation. Group-DINOmics uses DINOv3 with a ViT-L backbone and different pretraining resources; therefore, its result is a cross-configuration reference rather than a controlled comparison.  Conclusions  Masked high-level semantic prediction provides an effective objective for learning interaction-aware group representations when group activity labels are excluded from representation training. The proposed framework uses visible actors to infer missing actor semantics, while cosine alignment and instance-level contrastive learning jointly improve local discrimination and global ranking. The results also show that individual action supervision supplies an important semantic anchor and that positional information is helpful but insufficient on its own. The CAD results further support applicability beyond sports scenes. The current method remains sensitive to localization errors and is less effective for several short-duration or weakly collaborative activities. Future work will investigate more robust spatial relation modeling, finer temporal dynamics, and stronger backbones under unified settings.
  • loading
  • [1]
    WU Dekun, ZHAO He, BAO Xingce, et al. Sports video analysis on large-scale data[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 19–36. doi: 10.1007/978-3-031-19836-6_2.
    [2]
    RANASINGHE K, NASEER M, KHAN S, et al. Self-supervised video transformer[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 2864–2874. doi: 10.1109/CVPR52688.2022.00289.
    [3]
    EHSANPOUR M, ABEDIN A, SALEH F, et al. Joint learning of social groups, individuals action and sub-group activities in videos[C]. 16th European Conference Computer Vision – ECCV 2020, Glasgow, UK, 2020: 177–195. doi: 10.1007/978-3-030-58545-7_11.
    [4]
    FU Jun, LIU Jing, TIAN Haijie, et al. Dual attention network for scene segmentation[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 3141–3149. doi: 10.1109/CVPR.2019.00326.
    [5]
    TAMURA M, VISHWAKARMA R, and VENNELAKANTI R. Hunting group clues with transformers for social group activity recognition[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 19–35. doi: 10.1007/978-3-031-19772-7_2.
    [6]
    ZHOU Honglu, KADAV A, SHAMSIAN A, et al. COMPOSER: Compositional reasoning of group activity in videos with keypoint-only modality[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 249–266. doi: 10.1007/978-3-031-19833-5_15.
    [7]
    HE Kaiming, CHEN Xinlei, XIE Saining, et al. Masked autoencoders are scalable vision learners[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 15979–15988. doi: 10.1109/CVPR52688.2022.01553.
    [8]
    ASSRAN M, DUVAL Q, MISRA I, et al. Self-supervised learning from images with a joint-embedding predictive architecture[C]. Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, Canada, 2023: 15619–15629. doi: 10.1109/CVPR52729.2023.01499.
    [9]
    AMER M R, LEI Peng, and TODOROVIC S. HiRF: Hierarchical random field for collective activity recognition in videos[C]. 13th European Conference Computer Vision -- ECCV 2014, Zurich, Switzerland, 2014: 572–585. doi: 10.1007/978-3-319-10599-4_37.
    [10]
    ZHENG Yihao, WANG Zhuming, GU Ke, et al. Multi-scale motion-based relational reasoning for group activity recognition[J]. Engineering Applications of Artificial Intelligence, 2025, 139: 109570. doi: 10.1016/j.engappai.2024.109570.
    [11]
    DU Zexing and WANG Qing. Exploring global context and position-aware representation for group activity recognition[J]. Image and Vision Computing, 2024, 149: 105181. doi: 10.1016/j.imavis.2024.105181.
    [12]
    WU Jianchao, WANG Limin, WANG Li, et al. Learning actor relation graphs for group activity recognition[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 9956–9966. doi: 10.1109/CVPR.2019.01020.
    [13]
    HAN Mingfei, ZHANG D J, WANG Yali, et al. Dual-AI: Dual-path actor interaction learning for group activity recognition[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 2980–2989. doi: 10.1109/CVPR52688.2022.00300.
    [14]
    朱丽萍, 吴祀霖, 陈晓禾, 等. 多尺度子群体交互关系下的群体行为识别方法[J]. 电子与信息学报, 2024, 46(5): 2228–2236. doi: 10.11999/JEIT231304.

    ZHU Liping, WU Silin, CHEN Xiaohe, et al. Group activity recognition under multi-scale sub-group interaction relationships[J]. Journal of Electronics & Information Technology, 2024, 46(5): 2228–2236. doi: 10.11999/JEIT231304.
    [15]
    韩宗旺, 杨涵, 吴世青, 等. 时空自适应图卷积与Transformer结合的动作识别网络[J]. 电子与信息学报, 2024, 46(6): 2587–2595. doi: 10.11999/JEIT230551.

    HAN Zongwang, YANG Han, WU Shiqing, et al. Action recognition network combining spatio-temporal adaptive graph convolution and Transformer[J]. Journal of Electronics & Information Technology, 2024, 46(6): 2587–2595. doi: 10.11999/JEIT230551.
    [16]
    RAVITEJA CHAPPA N V, NGUYEN P, NELSON A H, et al. SoGAR: Self-supervised spatiotemporal attention-based social group activity recognition[J]. IEEE Access, 2025, 13: 33631–33642. doi: 10.1109/ACCESS.2025.3541986.
    [17]
    GUO Jie and GE Yongxin. Temporal contrastive and spatial enhancement coarse grained network for weakly supervised group activity recognition[J]. Engineering Applications of Artificial Intelligence, 2024, 133: 108115. doi: 10.1016/j.engappai.2024.108115.
    [18]
    IBRAHIM M S and MORI G. Hierarchical relational networks for group activity recognition and retrieval[C]. 15th European Conference Computer Vision – ECCV 2018, Munich, Germany, 2018: 742–758. doi: 10.1007/978-3-030-01219-9_44.
    [19]
    NAKATANI C, KAWASHIMA H, and UKITA N. Learning group activity features through person attribute prediction[C]. Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 18233–18242. doi: 10.1109/CVPR52733.2024.01726.
    [20]
    CHEN Ting, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations[C]. Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, 119: 1597–1607. (查阅网上资料, 未找到本条文献出版地信息, 请确认).
    [21]
    孙中华, 吴双, 贾克斌, 等. 基于对比学习的动作识别研究综述[J]. 电子与信息学报, 2025, 47(8): 2473–2485. doi: 10.11999/JEIT250131.

    SUN Zhonghua, WU Shuang, JIA Kebin, et al. A review on action recognition based on contrastive learning[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2473–2485. doi: 10.11999/JEIT250131.
    [22]
    刁文辉, 龚铄, 辛林霖, 等. 针对多模态遥感数据的自监督策略模型预训练方法[J]. 电子与信息学报, 2025, 47(6): 1658–1668. doi: 10.11999/JEIT241016.

    DIAO Wenhui, GONG Shuo, XIN Linlin, et al. A model pre-training method with self-supervised strategies for multimodal remote sensing data[J]. Journal of Electronics & Information Technology, 2025, 47(6): 1658–1668. doi: 10.11999/JEIT241016.
    [23]
    ASSRAN M, BARDES A, FAN D, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning[EB/OL]. https://arxiv.org/abs/2506.09985, 2025.
    [24]
    CHEN Delong, SHUKOR M, MOUTAKANNI T, et al. VL-JEPA: Joint embedding predictive architecture for vision-language[EB/OL]. https://arxiv.org/abs/2512.10942v1, 2025.
    [25]
    NAM H, LE LIDEC Q, MAES L, et al. Causal-JEPA: Learning world models through object-level latent interventions[EB/OL]. https://arxiv.org/abs/2602.11389, 2026.
    [26]
    IBRAHIM M S, MURALIDHARAN S, DENG Zhiwei, et al. A hierarchical deep temporal model for group activity recognition[C]. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, USA, 2016: 1971–1980. doi: 10.1109/CVPR.2016.217.
    [27]
    CHOI W, SHAHID K, and SAVARESE S. What are they doing?: Collective activity classification using spatio-temporal relationship among people[C]. 2009 IEEE 12th International Conference on Computer Vision Workshops, ICCV Workshops, Kyoto, Japan, 2009: 1282–1289. doi: 10.1109/ICCVW.2009.5457461.
    [28]
    YAN Rui, XIE Lingxi, TANG Jinhui, et al. HiGCIN: Hierarchical graph-based cross inference network for group activity recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(6): 6955–6968. doi: 10.1109/TPAMI.2020.3034233.
    [29]
    YUAN Hangjie, NI Dong, and WANG Mang. Spatio-temporal dynamic inference network for group activity recognition[C]. Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, Canada, 2021: 7456–7465. doi: 10.1109/ICCV48922.2021.00738.
    [30]
    TEZUKA R, NAKATANI C, and UKITA N. Group-DINOmics: Incorporating people dynamics into DINO for self-supervised group activity feature learning[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026. (查阅网上资料, 未找到本条文献出版地信息, 请确认).
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(5)  / Tables(10)

    Article Metrics

    Article views (22) PDF downloads(0) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return