Advanced Search
Turn off MathJax
Article Contents
WU You, ZHANG Mingxuan, CHANG Ren, YANG Kang, TAO Shifei. STAVFT: Spatio-Temporal Contrastive Learning for AIS-Video Fusion Tracking[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260236
Citation: WU You, ZHANG Mingxuan, CHANG Ren, YANG Kang, TAO Shifei. STAVFT: Spatio-Temporal Contrastive Learning for AIS-Video Fusion Tracking[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260236

STAVFT: Spatio-Temporal Contrastive Learning for AIS-Video Fusion Tracking

doi: 10.11999/JEIT260236 cstr: 32379.14.JEIT260236
Funds:  State Key Laboratory Fund
  • Accepted Date: 2026-09-15
  • Rev Recd Date: 2026-09-15
  • Available Online: 2026-09-20
  •   Objective  As maritime ship density increases and the demand for intelligent maritime supervision rises, utilizing multi-source heterogeneous sensors to achieve continuous and robust target tracking has become a critical means of enhancing navigational perception. In nearshore ship monitoring, visual modalities are highly susceptible to environmental interference such as low visibility and occlusions. Furthermore, traditional tracking methods struggle with the severe temporal asynchrony between Automatic Identification System (AIS) data and video trajectories, as well as the instability of identity association in multi-target tracking. This research aims to address these challenges by proposing a ship tracking algorithm based on the deep fusion of AIS and video information combined with spatio-temporal contrastive learning, referred to as STAVFT. The primary goal is to leverage AIS spatial priors to guide visual detectors in locking onto targets under low-visibility conditions and to ensure stable identity consistency in complex multi-target scenarios.  Methods  The STAVFT algorithm integrates perception-level enhancement with trajectory-level association to create a unified tracking pipeline. At the detection stage, the algorithm employs a Soft Mask feature enhancement strategy and an ROI re-detection compensation strategy. The Soft Mask strategy maps AIS spatial priors into Gaussian-weighted guidance maps that are injected into the feature extraction and fusion stages of the YOLOv11 detector, forcing the network to focus its representational power on high-probability target regions. For targets missed by the initial detector, the ROI re-detection strategy generates high-confidence candidate windows based on AIS priors to perform secondary detection and recover lost tracks. In the association stage, the algorithm constructs a Spatio-Temporal Aware Sequence Transformer (SAST) encoder to handle the asynchronous nature of AIS and video data. By introducing a continuous time-aware self-attention mechanism, the encoder explicitly models the dynamic intensity of heterogeneous trajectories and maps them into a unified embedding space. Additionally, a Spatio-Temporal Negative-sample Contrastive Estimation (STNCE) loss function is designed to strengthen the discrimination between physically adjacent and confusing targets. This loss function incorporates physical constraints, such as spatial distance and motion similarity, to penalize hard negative samples in the feature space.  Results and Discussions  Experimental results on the FVessel dataset demonstrate that the STAVFT algorithm significantly improves detection performance, raising the recall rate from 65.6% to 79.4% and the mAP50 to 0.792. The SAST encoder handles asynchronous AIS data effectively, achieving a Multi-Object Fusion Accuracy (MOFA) of 0.934, while the STNCE loss successfully mitigates identity swaps in dense traffic by widening the decision boundaries between confusing targets. Overall, the algorithm maintains a stable MOFA above 95.8% and an identity consistency (IDF1) exceeding 96.8%. Although the integration of AIS guidance and ROI re-detection increases the average processing time to 0.334 seconds per video second, the system remains well within real-time requirements, and its adaptive logic provides superior robustness against clutter and track fragmentation.  Conclusions  This paper presents the STAVFT algorithm, which successfully fuses AIS and video information through spatio-temporal contrastive learning to achieve robust ship tracking. By addressing the issues of environmental interference, data asynchrony, and identity confusion, the algorithm significantly improves the recall rate and identity consistency compared to current mainstream methods. The experimental results on the FVessel dataset validate the scientific validity of the "perception-guided association" logic. Future research will focus on model lightweighting for edge-cloud collaboration to further optimize the deployment performance of the system on embedded terminals.
  • loading
  • [1]
    严新平, 韩亚, 吴兵, 等. 水路交通系统的发展现状与未来展望[J]. 中国航海, 2024, 47(2): 145–152. doi: 10.3969/j.issn.1000-4653.2024.02.019.

    YAN Xinping, HAN Ya, WU Bing, et al. Current development and future prospects of waterborne transportation systems[J]. Navigation of China, 2024, 47(2): 145–152. doi: 10.3969/j.issn.1000-4653.2024.02.019.
    [2]
    BEWLEY A, GE Zongyuan, OTT L, et al. Simple online and realtime tracking[C]. 2016 IEEE International Conference on Image Processing, Phoenix, USA, 2016: 3464–3468. doi: 10.1109/ICIP.2016.7533003.
    [3]
    WOJKE N, BEWLEY A, and PAULUS D. Simple online and realtime tracking with a deep association metric[C]. 2017 IEEE International Conference on Image Processing (ICIP), Beijing, China, 2017: 3645–3649. doi: 10.1109/ICIP.2017.8296962.
    [4]
    WANG Zhongdao, ZHENG Liang, LIU Yixuan, et al. Towards real-time multi-object tracking[C]. Proceedings of the 16th European Conference on Computer Vision-ECCV 2020, Glasgow, UK, 2020: 107–122. doi: 10.1007/978-3-030-58621-8_7.
    [5]
    ZHANG Yifu, WANG Chunyu, WANG Xinggang, et al. FairMOT: On the fairness of detection and re-identification in multiple object tracking[J]. International Journal of Computer Vision, 2021, 129(11): 3069–3087. doi: 10.1007/s11263-021-01513-4.
    [6]
    ZHANG Yifu, SUN Peize, JIANG Yi, et al. ByteTrack: Multi-object tracking by associating every detection box[C]. Proceedings of the 17th European Conference on Computer Vision–ECCV 2022, Tel Aviv, Israel, 2022: 1–21. doi: 10.1007/978-3-031-20047-2_1.
    [7]
    CAO Jinkun, PANG Jiangmiao, WENG Xinshuo, et al. Observation-centric SORT: Rethinking SORT for robust multi-object tracking[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 9686–9696. doi: 10.1109/CVPR52729.2023.00934.
    [8]
    WU Yong, CHU Xiumin, DENG Lei, et al. A new multi-sensor fusion approach for integrated ship motion perception in inland waterways[J]. Measurement, 2022, 200: 111630. doi: 10.1016/j.measurement.2022.111630.
    [9]
    LI Fan, YU Kun, YUAN Chao, et al. Dark ship detection via optical and SAR collaboration: An improved multi-feature association method between remote sensing images and AIS data[J]. Remote Sensing, 2025, 17(13): 2201. doi: 10.3390/rs17132201.
    [10]
    XUE Weibao, AI Jiaqiu, ZHU Yanan, et al. AIS-FCANet: Long-term AIS data assisted frequency-spatial contextual awareness network for salient ship detection in SAR imagery[J]. IEEE Transactions on Aerospace and Electronic Systems, 2025, 61(5): 15166–15171. doi: 10.1109/TAES.2025.3588484.
    [11]
    CHEN Lihang, HU Zhuhua, CHEN Junfei, et al. SVIADF: Small vessel identification and anomaly detection based on wide-area remote sensing imagery and AIS data fusion[J]. Remote Sensing, 2025, 17(5): 868. doi: 10.3390/rs17050868.
    [12]
    QU Jingxiang, LIU R W, GUO Yu, et al. Improving maritime traffic surveillance in inland waterways using the robust fusion of AIS and visual data[J]. Ocean Engineering, 2023, 275: 114198. doi: 10.1016/j.oceaneng.2023.114198.
    [13]
    GUO Yu, LIU R W, QU Jingxiang, et al. Asynchronous trajectory matching-based multimodal maritime data fusion for vessel traffic surveillance in inland waterways[J]. IEEE Transactions on Intelligent Transportation Systems, 2023, 24(11): 12779–12792. doi: 10.1109/TITS.2023.3285415.
    [14]
    杜子俊, 贺益雄, 于德清, 等. 视觉与AIS融合的桥区水域船舶自动监测方法[J]. 中国航海, 2025, 48(1): 34–42. doi: 10.3969/j.issn.1000-4653.2025.01.005.

    DU Zijun, HE Yixiong, YU Deqing, et al. Automatic ship monitoring method in bridge area by fusion of vision and AIS[J]. Navigation of China, 2025, 48(1): 34–42. doi: 10.3969/j.issn.1000-4653.2025.01.005.
    [15]
    MEINHARDT T, KIRILLOV A, LEAL-TAIXÉ L, et al. TrackFormer: Multi-object tracking with transformers[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 8834–8844. doi: 10.1109/CVPR52688.2022.00864.
    [16]
    ZENG Fangao, DONG Bin, ZHANG Yuang, et al. MOTR: End-to-end multiple-object tracking with transformer[C]. Proceedings of the 17th European Conference on Computer Vision–ECCV 2022, Tel Aviv, Israel, 2022: 659–675. doi: 10.1007/978-3-031-19812-0_38.
    [17]
    ZHAO Yian, LV Wenyu, XU Shangliang, et al. DETRs beat YOLOs on real-time object detection[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 16965–16974. doi: 10.1109/CVPR52733.2024.01605.
    [18]
    WANG Han, LI Shengyang, YANG Jian, et al. Cross-modal ship re-identification via optical and SAR imagery: A novel dataset and method[C]. Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, USA, 2025: 7873–7883. doi: 10.1109/ICCV51701.2025.00738.
    [19]
    HU Songtao, CHEN Guanyu, ZHOU Rui, et al. Fishing vessel behavior pattern recognition using AIS sub-trajectory prototype learning based on Gramian Angular Field[J]. Complex & Intelligent Systems, 2026, 12(2): 68. doi: 10.1007/s40747-025-02187-y.
    [20]
    邵延华, 张铎, 楚红雨, 等. 基于深度学习的YOLO目标检测综述[J]. 电子与信息学报, 2022, 44(10): 3697–3708. doi: 10.11999/JEIT210790.

    SHAO Yanhua, ZHANG Duo, CHU Hongyu, et al. A review of YOLO object detection based on deep learning[J]. Journal of Electronics & Information Technology, 2022, 44(10): 3697–3708. doi: 10.11999/JEIT210790.
    [21]
    HE Wei, HE Wenbo, LEI Jinyu, et al. Multi-source perception data fusion of vessels in visual occlusion scenarios: Leveraging prior knowledge of vessel motion[J]. Engineering Applications of Artificial Intelligence, 2025, 156: 111118. doi: 10.1016/j.engappai.2025.111118.
    [22]
    RISTANI E, SOLERA F, ZOU R, et al. Performance measures and a data set for multi-target, multi-camera tracking[C]. Proceedings of the 14th European Conference on Computer Vision-ECCV 2016 Workshops, Amsterdam, The Netherlands, 2016: 17–35. doi: 10.1007/978-3-319-48881-3_2.
    [23]
    REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137–1149. doi: 10.1109/TPAMI.2016.2577031.
    [24]
    LIU Wei, ANGUELOV D, ERHAN D, et al. SSD: Single shot MultiBox detector[C]. Proceedings of the 14th European Conference on Computer Vision–ECCV 2016, Amsterdam, The Netherlands, 2016: 21–37. doi: 10.1007/978-3-319-46448-0_2.
    [25]
    GE Zheng, LIU Songtao, WANG Feng, et al. YOLOX: Exceeding YOLO series in 2021[Z]. arXiv: 2107.08430, 2021. doi: 10.48550/arXiv.2107.08430. (查阅网上资料,请核对文献类型及格式是否正确).
    [26]
    KHANAM R and HUSSAIN M. YOLOv11: An overview of the key architectural enhancements[Z]. arXiv: 2410.17725, 2024. doi: 10.48550/arXiv.2410.17725. (查阅网上资料,请核对文献类型及格式是否正确).
    [27]
    BAI Shaojie, KOLTER J Z, and KOLTUN V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling[Z]. arXiv: 1803.01271, 2018. doi: 10.48550/arXiv.1803.01271. (查阅网上资料,请核对文献类型及格式是否正确).
    [28]
    SUN Peize, CAO Jinkun, JIANG Yi, et al. TransTrack: Multiple object tracking with transformer[Z]. arXiv: 2012.15460, 2020. doi: 10.48550/arXiv.2012.15460. (查阅网上资料,请核对文献类型及格式是否正确).
    [29]
    ZHANG Jiayu, WANG Mei, KAN Ruixiang, et al. Multi-source heterogeneous data fusion algorithm for vessel trajectories in canal scenarios[J]. Electronics, 2025, 14(16): 3223. doi: 10.3390/electronics14163223.
    [30]
    LIU R W, GUO Yu, NIE Jiangtian, et al. Intelligent edge-enabled efficient multi-source data fusion for autonomous surface vehicles in maritime Internet of Things[J]. IEEE Transactions on Green Communications and Networking, 2022, 6(3): 1574–1587. doi: 10.1109/TGCN.2022.3158004.
    [31]
    程伊婷, 董涛, 苏昱玮, 等. 面向通信信号高效接收处理的压缩感知技术综述[J]. 电子与信息学报, 2026, 48(1): 168–182. doi: 10.11999/JEIT250855.

    CHENG Yiting, DONG Tao, SU Yuwei, et al. A review of compressed sensing technology for efficient receiving and processing of communication signal[J]. Journal of Electronics & Information Technology, 2026, 48(1): 168–182. doi: 10.11999/JEIT250855.
    [32]
    余礼苏, 钟润, 吕欣欣, 等. 压缩感知辅助的低复杂度SCMA系统优化设计[J]. 电子与信息学报, 2024, 46(5): 2011–2017. doi: 10.11999/JEIT231226.

    YU Lisu, ZHONG Run, LU Xinxin, et al. Optimized design of low complexity SCMA system assisted by compressed sensing[J]. Journal of Electronics & Information Technology, 2024, 46(5): 2011–2017. doi: 10.11999/JEIT231226.
    [33]
    杨春玲, 梁梓文. 静态与动态域先验增强的两阶段视频压缩感知重构网络[J]. 电子与信息学报, 2024, 46(11): 4247–4258. doi: 10.11999/JEIT240295.

    YANG Chunling and LIANG Ziwen. Static and dynamic-domain prior enhancement two-stage video compressed sensing reconstruction network[J]. Journal of Electronics & Information Technology, 2024, 46(11): 4247–4258. doi: 10.11999/JEIT240295.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(9)  / Tables(10)

    Article Metrics

    Article views (20) PDF downloads(0) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return