高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

STAVFT:AIS与视频信息融合及时空对比学习的船舶跟踪算法

吴优 张铭轩 常仁 杨康 陶诗飞

吴优, 张铭轩, 常仁, 杨康, 陶诗飞. STAVFT:AIS与视频信息融合及时空对比学习的船舶跟踪算法[J]. 电子与信息学报. doi: 10.11999/JEIT260236
引用本文: 吴优, 张铭轩, 常仁, 杨康, 陶诗飞. STAVFT:AIS与视频信息融合及时空对比学习的船舶跟踪算法[J]. 电子与信息学报. doi: 10.11999/JEIT260236
WU You, ZHANG Mingxuan, CHANG Ren, YANG Kang, TAO Shifei. STAVFT: Spatio-Temporal Contrastive Learning for AIS-Video Fusion Tracking[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260236
Citation: WU You, ZHANG Mingxuan, CHANG Ren, YANG Kang, TAO Shifei. STAVFT: Spatio-Temporal Contrastive Learning for AIS-Video Fusion Tracking[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260236

STAVFT:AIS与视频信息融合及时空对比学习的船舶跟踪算法

doi: 10.11999/JEIT260236 cstr: 32379.14.JEIT260236
基金项目: 国家重点实验室基金
详细信息
    作者简介:

    吴优:男,博士生,研究方向为多模态融合目标检测

    张铭轩:男,硕士生,研究方向为多模态融合目标检测

    常仁:男,高级工程师,研究方向为电子对抗技术

    杨康:男,研究员,研究方向为电子侦察与对抗

    陶诗飞:男,研究员,研究方向为电子侦察与识别,多源融合技术

    通讯作者:

    陶诗飞 s.tao@njust.edu.cn

  • 中图分类号: TP391.41; TN95

STAVFT: Spatio-Temporal Contrastive Learning for AIS-Video Fusion Tracking

Funds: State Key Laboratory Fund
  • 摘要: 随着海上船舶密度的日益增加以及智能化海事监管需求的提升,利用多源异构传感器实现目标的持续稳健跟踪已成为提升航行感知能力的有效手段。针对近岸船舶监控中视觉模态易受环境干扰、船舶自动识别系统(AIS)与视频轨迹存在严重异步性及多目标跟踪身份关联不稳定的问题,该文提出一种基于AIS与视频信息深度融合及时空对比学习的船舶跟踪算法(Spatio-Temporal contrastive learning for AIS-Video Fusion Tracking, STAVFT)。首先,算法设计了基于AIS特征注入的视频图像Soft Mask特征增强与ROI重检测补偿策略,利用AIS空间先验引导视觉检测器在低能见度及遮挡环境下锁定目标,解决跟踪前端的目标丢失问题。其次,针对AIS与视频数据的异步匹配难题,提出一种时空感知序列Transformer编码器,通过空洞卷积与时间感知自注意力机制提取跨模态运动特征,实现异构轨迹的统一嵌入表示。最后,设计了时空增强负样本对比估计损失函数,通过引入物理空间约束强化了跟踪过程对邻近易混淆目标的判别能力。在FVessel数据集上的实验结果表明,该算法将船舶跟踪过程中的检测召回率从65.6%提升至79.4%,多目标融合跟踪准确率稳定在95.8%以上,身份一致性指标超过96.8%,显著提升了复杂环境下船舶跟踪的鲁棒性。
  • 图  1  多模态融合跟踪算法STAVFT结构图

    图  2  Soft Mask补偿方法流程图

    图  3  时空感知序列Transformer编码器模型结构图

    图  4  FVessel目标检测数据集示意图

    图  5  加入Soft Mask前后PR曲线图

    图  6  Soft Mask补偿检测方法加入前后热力图结果示意图

    图  7  基于AIS特征融合检测补偿前后检测结果可视化图

    图  8  关键参数敏感性分析

    图  9  不同方法跟踪可视化结果图

    表  1  FVessel数据集AIS信息示例表

    序号MMSILon(°)Lat(°)Speed(kn)Course(°)Heading(°)Timestamp(ms)
    0110000000114.325730.601351149.55111654317561796.00
    1130000000114.319330.615380135.02031654317402824.00
    2180000000114.314630.601553.5213.95111654317458915.00
    3200000000114.317830.614040.1139.11391654317352745.00
    4290000000114.30830.596863211.55111654317437905.00
    5310000000114.322630.607367.838.0381654317553618.00
    6320000000114.318830.613110.157.0571654317275269.00
    7370000000114.318630.613320174.11741654317412483.00
    8440000000114.294330.586610360.05111652181591161.00
    下载: 导出CSV

    表  2  不同算法在FVessel目标检测数据集上对比结果表

    检测模型PRmAP
    50
    参数量(M)FPS
    (帧/秒)
    Faster R-CNN[23]0.6250.4850.51226.2975
    SSD[24]0.7420.5850.63541.3528
    YOLOX[25]0.7180.6420.6729.00145
    YOLOv11[26]0.7230.6560.6882.58170
    下载: 导出CSV

    表  3  AIS融合补偿检测方法消融结果表

    所用方法PRmAP50F1FLOPs(G)
    YOLOv110.7230.6560.6880.6856.3
    本文方法机制A0.6960.7650.7690.7517.3
    机制B0.7350.7150.7390.72915.9
    机制A+B0.7490.7940.7920.77816.9
    下载: 导出CSV

    表  4  误检类型统计表

    误检类型 YOLOv11误检
    比例(%)
    机制A误检
    比例(%)
    水面反光、尾迹及岸线结构 10.8 13.2
    AIS投影偏移区域 3.1 5.5
    邻近或遮挡船舶混淆 6.2 6.5
    同一目标重复框 5.0 5.2
    合计 25.1 30.4
    下载: 导出CSV

    表  5  统计显著性结果

    指标YOLOv11机制A差值95%置信区间
    P0.7230.696-0.027[-0.035,-0.019]
    R0.6560.765+0.109[+0.097,+0.121]
    mAP500.6880.769+0.081[+0.073,+0.089]
    F10.6850.751+0.066[+0.057,+0.075]
    下载: 导出CSV

    表  6  不同方法航迹关联性能对比表

    所用方法Top-1Top-5MOFAFLOPs(G)
    E-Fast DTW0.7940.8950.815-
    TCN0.7650.8620.7820.0012
    Transformer0.8420.9280.8760.0045
    SAST0.8950.9620.9340.0034
    下载: 导出CSV

    表  7  关键模块单变量消融结果

    组别TCN扩张率Top-1Top-5MOFA
    普通残差TCN
    (无时间偏置)
    {1,1,1}0.8720.9120.930
    空洞残差TCN
    (无时间偏置)
    {1,2,3}0.8850.9430.944
    普通残差+时间偏置{1,1,1}0.8780.9460.956
    空洞残差+时间偏置{1,2,3}0.9320.9870.965
    下载: 导出CSV

    表  8  不同损失函数航迹关联性能对比表

    损失函数Top-1Top-5MOFA
    单向InfoNCE0.8420.9320.884
    对称InfoNCE0.8950.9620.934
    STNCE0.9320.9870.965
    下载: 导出CSV

    表  9  不同轨迹融合方法在FVessel数据集上测试性能对比结果表

    视频方法MOFAIDPIDRIDF1
    Clip01
    (夜间低光)
    欧式融合0.6930.8870.8260.855
    MSDF0.6840.8860.8210.852
    DeepSORVF0.9420.9610.9460.953
    本文算法0.9580.9720.9630.968
    Clip02
    (晴天)
    欧式融合0.6970.8940.8850.890
    MSDF0.6950.8920.8820.887
    DeepSORVF0.9710.9740.9690.972
    本文算法0.9800.9810.9730.977
    Clip03
    (多云)
    欧式融合0.7930.9690.8730.918
    MSDF0.8170.9710.8810.924
    DeepSORVF0.9550.9720.9580.965
    本文算法0.9710.9820.9710.976
    Clip04
    (晴天)
    欧式融合0.7240.9010.8920.897
    MSDF0.7190.8990.8890.894
    DeepSORVF0.9730.9750.9700.973
    本文算法0.9800.9820.9750.979
    Clip05
    (严重遮挡)
    欧式融合0.6750.8800.8140.846
    MSDF0.6680.8740.8100.841
    DeepSORVF0.9410.9590.9420.950
    本文算法0.9610.9750.9650.970
    下载: 导出CSV

    表  10  不同轨迹融合方法在FVessel数据集平均每秒处理时间

    视频视频长度DeepSORVF(s)本文方法(s)
    Clip011m51s0.2450.310
    Clip021m36s0.2510.325
    Clip033m42s0.2600.342
    Clip042m42s0.2750.365
    Clip053m05s0.2550.330
    Clip062m38s0.2480.315
    Clip0711m10s0.2650.355
    Clip085m07s0.2520.328
    Clip098m39s0.2580.340
    Clip102m46s0.2530.332
    平均值-0.2560.334
    下载: 导出CSV
  • [1] 严新平, 韩亚, 吴兵, 等. 水路交通系统的发展现状与未来展望[J]. 中国航海, 2024, 47(2): 145–152. doi: 10.3969/j.issn.1000-4653.2024.02.019.

    YAN Xinping, HAN Ya, WU Bing, et al. Current development and future prospects of waterborne transportation systems[J]. Navigation of China, 2024, 47(2): 145–152. doi: 10.3969/j.issn.1000-4653.2024.02.019.
    [2] BEWLEY A, GE Zongyuan, OTT L, et al. Simple online and realtime tracking[C]. 2016 IEEE International Conference on Image Processing, Phoenix, USA, 2016: 3464–3468. doi: 10.1109/ICIP.2016.7533003.
    [3] WOJKE N, BEWLEY A, and PAULUS D. Simple online and realtime tracking with a deep association metric[C]. 2017 IEEE International Conference on Image Processing (ICIP), Beijing, China, 2017: 3645–3649. doi: 10.1109/ICIP.2017.8296962.
    [4] WANG Zhongdao, ZHENG Liang, LIU Yixuan, et al. Towards real-time multi-object tracking[C]. Proceedings of the 16th European Conference on Computer Vision-ECCV 2020, Glasgow, UK, 2020: 107–122. doi: 10.1007/978-3-030-58621-8_7.
    [5] ZHANG Yifu, WANG Chunyu, WANG Xinggang, et al. FairMOT: On the fairness of detection and re-identification in multiple object tracking[J]. International Journal of Computer Vision, 2021, 129(11): 3069–3087. doi: 10.1007/s11263-021-01513-4.
    [6] ZHANG Yifu, SUN Peize, JIANG Yi, et al. ByteTrack: Multi-object tracking by associating every detection box[C]. Proceedings of the 17th European Conference on Computer Vision–ECCV 2022, Tel Aviv, Israel, 2022: 1–21. doi: 10.1007/978-3-031-20047-2_1.
    [7] CAO Jinkun, PANG Jiangmiao, WENG Xinshuo, et al. Observation-centric SORT: Rethinking SORT for robust multi-object tracking[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 9686–9696. doi: 10.1109/CVPR52729.2023.00934.
    [8] WU Yong, CHU Xiumin, DENG Lei, et al. A new multi-sensor fusion approach for integrated ship motion perception in inland waterways[J]. Measurement, 2022, 200: 111630. doi: 10.1016/j.measurement.2022.111630.
    [9] LI Fan, YU Kun, YUAN Chao, et al. Dark ship detection via optical and SAR collaboration: An improved multi-feature association method between remote sensing images and AIS data[J]. Remote Sensing, 2025, 17(13): 2201. doi: 10.3390/rs17132201.
    [10] XUE Weibao, AI Jiaqiu, ZHU Yanan, et al. AIS-FCANet: Long-term AIS data assisted frequency-spatial contextual awareness network for salient ship detection in SAR imagery[J]. IEEE Transactions on Aerospace and Electronic Systems, 2025, 61(5): 15166–15171. doi: 10.1109/TAES.2025.3588484.
    [11] CHEN Lihang, HU Zhuhua, CHEN Junfei, et al. SVIADF: Small vessel identification and anomaly detection based on wide-area remote sensing imagery and AIS data fusion[J]. Remote Sensing, 2025, 17(5): 868. doi: 10.3390/rs17050868.
    [12] QU Jingxiang, LIU R W, GUO Yu, et al. Improving maritime traffic surveillance in inland waterways using the robust fusion of AIS and visual data[J]. Ocean Engineering, 2023, 275: 114198. doi: 10.1016/j.oceaneng.2023.114198.
    [13] GUO Yu, LIU R W, QU Jingxiang, et al. Asynchronous trajectory matching-based multimodal maritime data fusion for vessel traffic surveillance in inland waterways[J]. IEEE Transactions on Intelligent Transportation Systems, 2023, 24(11): 12779–12792. doi: 10.1109/TITS.2023.3285415.
    [14] 杜子俊, 贺益雄, 于德清, 等. 视觉与AIS融合的桥区水域船舶自动监测方法[J]. 中国航海, 2025, 48(1): 34–42. doi: 10.3969/j.issn.1000-4653.2025.01.005.

    DU Zijun, HE Yixiong, YU Deqing, et al. Automatic ship monitoring method in bridge area by fusion of vision and AIS[J]. Navigation of China, 2025, 48(1): 34–42. doi: 10.3969/j.issn.1000-4653.2025.01.005.
    [15] MEINHARDT T, KIRILLOV A, LEAL-TAIXÉ L, et al. TrackFormer: Multi-object tracking with transformers[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 8834–8844. doi: 10.1109/CVPR52688.2022.00864.
    [16] ZENG Fangao, DONG Bin, ZHANG Yuang, et al. MOTR: End-to-end multiple-object tracking with transformer[C]. Proceedings of the 17th European Conference on Computer Vision–ECCV 2022, Tel Aviv, Israel, 2022: 659–675. doi: 10.1007/978-3-031-19812-0_38.
    [17] ZHAO Yian, LV Wenyu, XU Shangliang, et al. DETRs beat YOLOs on real-time object detection[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 16965–16974. doi: 10.1109/CVPR52733.2024.01605.
    [18] WANG Han, LI Shengyang, YANG Jian, et al. Cross-modal ship re-identification via optical and SAR imagery: A novel dataset and method[C]. Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, USA, 2025: 7873–7883. doi: 10.1109/ICCV51701.2025.00738.
    [19] HU Songtao, CHEN Guanyu, ZHOU Rui, et al. Fishing vessel behavior pattern recognition using AIS sub-trajectory prototype learning based on Gramian Angular Field[J]. Complex & Intelligent Systems, 2026, 12(2): 68. doi: 10.1007/s40747-025-02187-y.
    [20] 邵延华, 张铎, 楚红雨, 等. 基于深度学习的YOLO目标检测综述[J]. 电子与信息学报, 2022, 44(10): 3697–3708. doi: 10.11999/JEIT210790.

    SHAO Yanhua, ZHANG Duo, CHU Hongyu, et al. A review of YOLO object detection based on deep learning[J]. Journal of Electronics & Information Technology, 2022, 44(10): 3697–3708. doi: 10.11999/JEIT210790.
    [21] HE Wei, HE Wenbo, LEI Jinyu, et al. Multi-source perception data fusion of vessels in visual occlusion scenarios: Leveraging prior knowledge of vessel motion[J]. Engineering Applications of Artificial Intelligence, 2025, 156: 111118. doi: 10.1016/j.engappai.2025.111118.
    [22] RISTANI E, SOLERA F, ZOU R, et al. Performance measures and a data set for multi-target, multi-camera tracking[C]. Proceedings of the 14th European Conference on Computer Vision-ECCV 2016 Workshops, Amsterdam, The Netherlands, 2016: 17–35. doi: 10.1007/978-3-319-48881-3_2.
    [23] REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137–1149. doi: 10.1109/TPAMI.2016.2577031.
    [24] LIU Wei, ANGUELOV D, ERHAN D, et al. SSD: Single shot MultiBox detector[C]. Proceedings of the 14th European Conference on Computer Vision–ECCV 2016, Amsterdam, The Netherlands, 2016: 21–37. doi: 10.1007/978-3-319-46448-0_2.
    [25] GE Zheng, LIU Songtao, WANG Feng, et al. YOLOX: Exceeding YOLO series in 2021[Z]. arXiv: 2107.08430, 2021. doi: 10.48550/arXiv.2107.08430. (查阅网上资料,请核对文献类型及格式是否正确).
    [26] KHANAM R and HUSSAIN M. YOLOv11: An overview of the key architectural enhancements[Z]. arXiv: 2410.17725, 2024. doi: 10.48550/arXiv.2410.17725. (查阅网上资料,请核对文献类型及格式是否正确).
    [27] BAI Shaojie, KOLTER J Z, and KOLTUN V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling[Z]. arXiv: 1803.01271, 2018. doi: 10.48550/arXiv.1803.01271. (查阅网上资料,请核对文献类型及格式是否正确).
    [28] SUN Peize, CAO Jinkun, JIANG Yi, et al. TransTrack: Multiple object tracking with transformer[Z]. arXiv: 2012.15460, 2020. doi: 10.48550/arXiv.2012.15460. (查阅网上资料,请核对文献类型及格式是否正确).
    [29] ZHANG Jiayu, WANG Mei, KAN Ruixiang, et al. Multi-source heterogeneous data fusion algorithm for vessel trajectories in canal scenarios[J]. Electronics, 2025, 14(16): 3223. doi: 10.3390/electronics14163223.
    [30] LIU R W, GUO Yu, NIE Jiangtian, et al. Intelligent edge-enabled efficient multi-source data fusion for autonomous surface vehicles in maritime Internet of Things[J]. IEEE Transactions on Green Communications and Networking, 2022, 6(3): 1574–1587. doi: 10.1109/TGCN.2022.3158004.
    [31] 程伊婷, 董涛, 苏昱玮, 等. 面向通信信号高效接收处理的压缩感知技术综述[J]. 电子与信息学报, 2026, 48(1): 168–182. doi: 10.11999/JEIT250855.

    CHENG Yiting, DONG Tao, SU Yuwei, et al. A review of compressed sensing technology for efficient receiving and processing of communication signal[J]. Journal of Electronics & Information Technology, 2026, 48(1): 168–182. doi: 10.11999/JEIT250855.
    [32] 余礼苏, 钟润, 吕欣欣, 等. 压缩感知辅助的低复杂度SCMA系统优化设计[J]. 电子与信息学报, 2024, 46(5): 2011–2017. doi: 10.11999/JEIT231226.

    YU Lisu, ZHONG Run, LU Xinxin, et al. Optimized design of low complexity SCMA system assisted by compressed sensing[J]. Journal of Electronics & Information Technology, 2024, 46(5): 2011–2017. doi: 10.11999/JEIT231226.
    [33] 杨春玲, 梁梓文. 静态与动态域先验增强的两阶段视频压缩感知重构网络[J]. 电子与信息学报, 2024, 46(11): 4247–4258. doi: 10.11999/JEIT240295.

    YANG Chunling and LIANG Ziwen. Static and dynamic-domain prior enhancement two-stage video compressed sensing reconstruction network[J]. Journal of Electronics & Information Technology, 2024, 46(11): 4247–4258. doi: 10.11999/JEIT240295.
  • 加载中
图(9) / 表(10)
计量
  • 文章访问数:  20
  • HTML全文浏览量:  5
  • PDF下载量:  0
  • 被引次数: 0
出版历程
  • 修回日期:  2026-09-15
  • 录用日期:  2026-09-15
  • 网络出版日期:  2026-09-20

目录

    /

    返回文章
    返回