Advanced Search
Turn off MathJax
Article Contents
CHEN Xiaoyu, ZHANG Fengzhuo, CHEN Yang, LIU Wenyuan, KONG Deming. A Spatial-Temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260512
Citation: CHEN Xiaoyu, ZHANG Fengzhuo, CHEN Yang, LIU Wenyuan, KONG Deming. A Spatial-Temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260512

A Spatial-Temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction

doi: 10.11999/JEIT260512 cstr: 32379.14.JEIT260512
Funds:  The Natural Science Foundation of Hebei Province (No.F2025203055)
  • Received Date: 2026-04-24
  • Accepted Date: 2026-07-01
  • Rev Recd Date: 2026-07-01
  • Available Online: 2026-07-23
  •   Objective  In port operation videos, the grab is a continuously moving target, and accurate trajectory extraction is of great importance for operation monitoring, equipment coordination, and collision warning. However, complex backgrounds, scale variations, partial occlusion, and boundary degradation often make target region extraction unstable, leading to centroid deviation, trajectory jitter, missed detections, and trajectory interruption. To address these issues, this paper proposes a spatial-temporal collaborative optimization method for stable and continuous grab trajectory extraction. While maintaining a relatively high inference speed, the proposed method effectively improves both the accuracy and continuity of trajectory extraction, providing a practical solution for stable perception of continuously moving targets in port industrial video scenarios.  Methods  Built on YOLOv8-seg, the proposed framework integrates spatial representation enhancement and temporal consistency constraint. First, CBAM, BiFPN-lite, and shallow feature aggregation are introduced to improve target-background separability, strengthen multi-scale representation, and preserve boundary details. Then, a temporal consistency constraint is imposed on prototype features through global average pooling, a cache-based pairing mechanism, and a weighted Charbonnier loss, thereby suppressing the temporal accumulation of local errors. In addition, a stage-wise training strategy with warmup epochs and a unified loss function is adopted to ensure stable convergence.  Results and Discussions  Experiments are conducted on DAVIS2016, SegTrackV2, and a real portal crane grab dataset to evaluate the proposed method in terms of segmentation performance, trajectory stability, and occlusion robustness. The results show that the proposed method achieves the best segmentation performance on the grab dataset, with J and F scores of 90.05% and 98.56%, respectively. It also shows clear improvement on DAVIS2016 and maintains comparable performance with slight gains on SegTrackV2 (Tables 1 and 2, Fig. 2). In terms of trajectory stability, compared with YOLOv8-seg, the proposed method reduces MAE and RMSE by approximately 55.3% and 52.6%, respectively, while lowering the miss rate to 0.56% (Table 5). It also produces a more concentrated trajectory error distribution and a smaller fluctuation range (Fig. 3). Occlusion robustness experiments further show that, under different occlusion ratios, the proposed method maintains good region integrity and continuous extraction capability, reducing the maximum number of consecutive missed frames from 52 to 47 (Table 6, Figs. 4 and 5). Ablation studies verify the complementarity of spatial representation enhancement and the temporal consistency constraint, while parameter analysis shows that setting the TCONS weight to 0.3 provides the best balance between segmentation quality and trajectory stability (Tables 7 and 8).  Conclusions  This paper proposes a spatial-temporal collaborative optimization method to address the challenge of stable grab trajectory extraction in port operation videos. Experimental results demonstrate that the proposed method achieves favorable segmentation accuracy and trajectory stability on DAVIS2016 and the real grab dataset, maintains comparable segmentation performance on SegTrackV2, shows strong continuous extraction capability under occlusion, and incurs no significant loss in inference speed. Since the current study is limited to fixed crane viewpoints, future work will focus on cross-scene generalization and long-term continuous perception under more complex operating conditions to further enhance the robustness and applicability of the proposed method in real-world environments.
  • loading
  • [1]
    HE Kaiming, GKIOXARI G, DOLLÁR P, et al. Mask R-CNN[C]. Proceedings of the 2017 IEEE International Conference on Computer Vision, Venice, Italy, 2017: 2980–2988. doi: 10.1109/ICCV.2017.322.
    [2]
    CAELLES S, MANINIS K K, PONT-TUSET J, et al. One-shot video object segmentation[C]. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 5320–5329. doi: 10.1109/CVPR.2017.565.
    [3]
    OH S W, LEE J Y, XU Ning, et al. Video object segmentation using space-time memory networks[C]. Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Korea (South), 2019: 9225–9234. doi: 10.1109/ICCV.2019.00932.
    [4]
    CHENG H K and SCHWING A G. XMem: Long-term video object segmentation with an Atkinson-Shiffrin memory model[C]. 17th European Conference on Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 640–658. doi: 10.1007/978-3-031-19815-1_37.
    [5]
    MONNIER Q, POULI T, and KPALMA K. Survey on fast dense video segmentation techniques[J]. Computer Vision and Image Understanding, 2024, 241: 103959. doi: 10.1016/j.cviu.2024.103959.
    [6]
    XU Guoping, UDUPA J K, YU Yajun, et al. Segment anything for video: A comprehensive review of video object segmentation and tracking from past to future[J]. Neurocomputing, 2026, 682: 133439. doi: 10.1016/j.neucom.2026.133439.
    [7]
    HOU Zhiqiang, LI Fucheng, DONG Jiale, et al. Video object segmentation based on dynamic perception update and feature fusion[J]. Image and Vision Computing, 2024, 150: 105156. doi: 10.1016/j.imavis.2024.105156.
    [8]
    WANG Jingxin, ZHANG Yunfeng, BAO Fangxun, et al. Video object segmentation by multi-scale attention using bidirectional strategy[J]. Image and Vision Computing, 2024, 148: 105136. doi: 10.1016/j.imavis.2024.105136.
    [9]
    KIM J, KIM J, and HONG S. G-TRACE: Grouped temporal recalibration for video object segmentation[J]. Image and Vision Computing, 2024, 147: 105050. doi: 10.1016/j.imavis.2024.105050.
    [10]
    HOU Zhiqiang, WANG Chenxu, MA Sugang, et al. Lightweight video object segmentation: Integrating online knowledge distillation for fast segmentation[J]. Knowledge-Based Systems, 2025, 308: 112759. doi: 10.1016/j.knosys.2024.112759.
    [11]
    LIU Yong, YU Ran, YIN Fei, et al. Learning high-quality dynamic memory for video object segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(5): 3452–3468. doi: 10.1109/TPAMI.2025.3532306.
    [12]
    SUN Maojin and SUN Minghui. STSim-Mamb: A spatiotemporal similarity learning framework for unsupervised video object segmentation[J]. Image and Vision Computing, 2026, 169: 105945. doi: 10.1016/j.imavis.2026.105945.
    [13]
    KAZEMI E S, TOUBAL I E, RAHMON G, et al. Domain generalization for multiple video object segmentation and tracking using transformers and smart memory[J]. International Journal of Computer Vision, 2026, 134(5): 206. doi: 10.1007/s11263-026-02742-1.
    [14]
    LU Hannan, TIAN Zhi, WEI Pengxu, et al. Integrating instance-level knowledge to see the unseen: A two-stream network for video object segmentation[J]. Neurocomputing, 2024, 602: 127878. doi: 10.1016/j.neucom.2024.127878.
    [15]
    侯志强, 董佳乐, 马素刚, 等. 基于多尺度特征增强与全局-局部特征聚合的视频目标分割算法[J]. 电子与信息学报, 2024, 46(11): 4198–4207. doi: 10.11999/JEIT231394.

    HOU Zhiqiang, DONG Jiale, MA Sugang, et al. Video object segmentation algorithm based on multi-scale feature enhancement and global-local feature aggregation[J]. Journal of Electronics & Information Technology, 2024, 46(11): 4198–4207. doi: 10.11999/JEIT231394.
    [16]
    陈雷, 杨吉斌, 曹铁勇, 等. 一种基于Transformer特征金字塔的自蒸馏目标分割方法[J]. 电子与信息学报, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.

    CHEN Lei, YANG Jibin, CAO Tieyong, et al. A self-distillation object segmentation method based on Transformer feature pyramid[J]. Journal of Electronics & Information Technology, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.
    [17]
    WOO S, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. 15th European Conference on Computer Vision – ECCV 2018, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1.
    [18]
    LIN T Y, DOLLÁR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 936–944. doi: 10.1109/CVPR.2017.106.
    [19]
    LIU Shu, QI Lu, QIN Haifang, et al. Path aggregation network for instance segmentation[C]. Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 8759–8768. doi: 10.1109/CVPR.2018.00913.
    [20]
    丁建睿, 张听, 刘家栋, 等. 融合邻域注意力和状态空间模型的医学视频分割算法[J]. 电子与信息学报, 2025, 47(5): 1582–1595. doi: 10.11999/JEIT240755.

    DING Jianrui, ZHANG Ting, LIU Jiadong, et al. Medical video segmentation algorithm integrating neighborhood attention and state space model[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1582–1595. doi: 10.11999/JEIT240755.
    [21]
    LI Jun, SUN Lijuan, REN Hengyi, et al. Learning effective feature representation for video object segmentation via memory[J]. Knowledge-Based Systems, 2024, 299: 112020. doi: 10.1016/j.knosys.2024.112020.
    [22]
    WANG Hui, ZHAO Yuqian, ZHANG Fan, et al. Multi-scale spatio-temporal memory network for semi-supervised video object segmentation[J]. Neurocomputing, 2025, 642: 130487. doi: 10.1016/j.neucom.2025.130487.
    [23]
    MIAO Bo, BENNAMOUN M, GAO Yongsheng, et al. Temporally consistent referring video object segmentation with hybrid memory[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(11): 11373–11385. doi: 10.1109/TCSVT.2024.3419119.
    [24]
    HOU Zhiqiang, CUI Hao, WANG Chenxu, et al. Frequency-aware fusion for improved video object segmentation[J]. Neurocomputing, 2025, 656: 131585. doi: 10.1016/j.neucom.2025.131585.
    [25]
    PERAZZI F, PONT-TUSET J, MCWILLIAMS B, et al. A benchmark dataset and evaluation methodology for video object segmentation[C]. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016: 724–732. doi: 10.1109/CVPR.2016.85.
    [26]
    LI Fuxin, KIM T, HUMAYUN A, et al. Video segmentation by tracking many figure-ground segments[C]. Proceedings of the 2013 IEEE International Conference on Computer Vision, Sydney, Australia, 2013: 2192–2199. doi: 10.1109/ICCV.2013.273.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(5)  / Tables(8)

    Article Metrics

    Article views (43) PDF downloads(4) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return