| Citation: | ZHANG Yaqi, LI Chengyang, ZHU Liping, LI Ruina. Group Activity Representation Learning via Masked Semantic Prediction in Latent Space[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260730 |
| [1] |
WU Dekun, ZHAO He, BAO Xingce, et al. Sports video analysis on large-scale data[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 19–36. doi: 10.1007/978-3-031-19836-6_2.
|
| [2] |
RANASINGHE K, NASEER M, KHAN S, et al. Self-supervised video transformer[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 2864–2874. doi: 10.1109/CVPR52688.2022.00289.
|
| [3] |
EHSANPOUR M, ABEDIN A, SALEH F, et al. Joint learning of social groups, individuals action and sub-group activities in videos[C]. 16th European Conference Computer Vision – ECCV 2020, Glasgow, UK, 2020: 177–195. doi: 10.1007/978-3-030-58545-7_11.
|
| [4] |
FU Jun, LIU Jing, TIAN Haijie, et al. Dual attention network for scene segmentation[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 3141–3149. doi: 10.1109/CVPR.2019.00326.
|
| [5] |
TAMURA M, VISHWAKARMA R, and VENNELAKANTI R. Hunting group clues with transformers for social group activity recognition[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 19–35. doi: 10.1007/978-3-031-19772-7_2.
|
| [6] |
ZHOU Honglu, KADAV A, SHAMSIAN A, et al. COMPOSER: Compositional reasoning of group activity in videos with keypoint-only modality[C]. 17th European Conference Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 249–266. doi: 10.1007/978-3-031-19833-5_15.
|
| [7] |
HE Kaiming, CHEN Xinlei, XIE Saining, et al. Masked autoencoders are scalable vision learners[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 15979–15988. doi: 10.1109/CVPR52688.2022.01553.
|
| [8] |
ASSRAN M, DUVAL Q, MISRA I, et al. Self-supervised learning from images with a joint-embedding predictive architecture[C]. Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, Canada, 2023: 15619–15629. doi: 10.1109/CVPR52729.2023.01499.
|
| [9] |
AMER M R, LEI Peng, and TODOROVIC S. HiRF: Hierarchical random field for collective activity recognition in videos[C]. 13th European Conference Computer Vision -- ECCV 2014, Zurich, Switzerland, 2014: 572–585. doi: 10.1007/978-3-319-10599-4_37.
|
| [10] |
ZHENG Yihao, WANG Zhuming, GU Ke, et al. Multi-scale motion-based relational reasoning for group activity recognition[J]. Engineering Applications of Artificial Intelligence, 2025, 139: 109570. doi: 10.1016/j.engappai.2024.109570.
|
| [11] |
DU Zexing and WANG Qing. Exploring global context and position-aware representation for group activity recognition[J]. Image and Vision Computing, 2024, 149: 105181. doi: 10.1016/j.imavis.2024.105181.
|
| [12] |
WU Jianchao, WANG Limin, WANG Li, et al. Learning actor relation graphs for group activity recognition[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 9956–9966. doi: 10.1109/CVPR.2019.01020.
|
| [13] |
HAN Mingfei, ZHANG D J, WANG Yali, et al. Dual-AI: Dual-path actor interaction learning for group activity recognition[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 2980–2989. doi: 10.1109/CVPR52688.2022.00300.
|
| [14] |
朱丽萍, 吴祀霖, 陈晓禾, 等. 多尺度子群体交互关系下的群体行为识别方法[J]. 电子与信息学报, 2024, 46(5): 2228–2236. doi: 10.11999/JEIT231304.
ZHU Liping, WU Silin, CHEN Xiaohe, et al. Group activity recognition under multi-scale sub-group interaction relationships[J]. Journal of Electronics & Information Technology, 2024, 46(5): 2228–2236. doi: 10.11999/JEIT231304.
|
| [15] |
韩宗旺, 杨涵, 吴世青, 等. 时空自适应图卷积与Transformer结合的动作识别网络[J]. 电子与信息学报, 2024, 46(6): 2587–2595. doi: 10.11999/JEIT230551.
HAN Zongwang, YANG Han, WU Shiqing, et al. Action recognition network combining spatio-temporal adaptive graph convolution and Transformer[J]. Journal of Electronics & Information Technology, 2024, 46(6): 2587–2595. doi: 10.11999/JEIT230551.
|
| [16] |
RAVITEJA CHAPPA N V, NGUYEN P, NELSON A H, et al. SoGAR: Self-supervised spatiotemporal attention-based social group activity recognition[J]. IEEE Access, 2025, 13: 33631–33642. doi: 10.1109/ACCESS.2025.3541986.
|
| [17] |
GUO Jie and GE Yongxin. Temporal contrastive and spatial enhancement coarse grained network for weakly supervised group activity recognition[J]. Engineering Applications of Artificial Intelligence, 2024, 133: 108115. doi: 10.1016/j.engappai.2024.108115.
|
| [18] |
IBRAHIM M S and MORI G. Hierarchical relational networks for group activity recognition and retrieval[C]. 15th European Conference Computer Vision – ECCV 2018, Munich, Germany, 2018: 742–758. doi: 10.1007/978-3-030-01219-9_44.
|
| [19] |
NAKATANI C, KAWASHIMA H, and UKITA N. Learning group activity features through person attribute prediction[C]. Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 18233–18242. doi: 10.1109/CVPR52733.2024.01726.
|
| [20] |
CHEN Ting, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations[C]. Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, 119: 1597–1607. (查阅网上资料, 未找到本条文献出版地信息, 请确认).
|
| [21] |
孙中华, 吴双, 贾克斌, 等. 基于对比学习的动作识别研究综述[J]. 电子与信息学报, 2025, 47(8): 2473–2485. doi: 10.11999/JEIT250131.
SUN Zhonghua, WU Shuang, JIA Kebin, et al. A review on action recognition based on contrastive learning[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2473–2485. doi: 10.11999/JEIT250131.
|
| [22] |
刁文辉, 龚铄, 辛林霖, 等. 针对多模态遥感数据的自监督策略模型预训练方法[J]. 电子与信息学报, 2025, 47(6): 1658–1668. doi: 10.11999/JEIT241016.
DIAO Wenhui, GONG Shuo, XIN Linlin, et al. A model pre-training method with self-supervised strategies for multimodal remote sensing data[J]. Journal of Electronics & Information Technology, 2025, 47(6): 1658–1668. doi: 10.11999/JEIT241016.
|
| [23] |
ASSRAN M, BARDES A, FAN D, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning[EB/OL]. https://arxiv.org/abs/2506.09985, 2025.
|
| [24] |
CHEN Delong, SHUKOR M, MOUTAKANNI T, et al. VL-JEPA: Joint embedding predictive architecture for vision-language[EB/OL]. https://arxiv.org/abs/2512.10942v1, 2025.
|
| [25] |
NAM H, LE LIDEC Q, MAES L, et al. Causal-JEPA: Learning world models through object-level latent interventions[EB/OL]. https://arxiv.org/abs/2602.11389, 2026.
|
| [26] |
IBRAHIM M S, MURALIDHARAN S, DENG Zhiwei, et al. A hierarchical deep temporal model for group activity recognition[C]. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, USA, 2016: 1971–1980. doi: 10.1109/CVPR.2016.217.
|
| [27] |
CHOI W, SHAHID K, and SAVARESE S. What are they doing?: Collective activity classification using spatio-temporal relationship among people[C]. 2009 IEEE 12th International Conference on Computer Vision Workshops, ICCV Workshops, Kyoto, Japan, 2009: 1282–1289. doi: 10.1109/ICCVW.2009.5457461.
|
| [28] |
YAN Rui, XIE Lingxi, TANG Jinhui, et al. HiGCIN: Hierarchical graph-based cross inference network for group activity recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(6): 6955–6968. doi: 10.1109/TPAMI.2020.3034233.
|
| [29] |
YUAN Hangjie, NI Dong, and WANG Mang. Spatio-temporal dynamic inference network for group activity recognition[C]. Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, Canada, 2021: 7456–7465. doi: 10.1109/ICCV48922.2021.00738.
|
| [30] |
TEZUKA R, NAKATANI C, and UKITA N. Group-DINOmics: Incorporating people dynamics into DINO for self-supervised group activity feature learning[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026. (查阅网上资料, 未找到本条文献出版地信息, 请确认).
|