高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

融合深度一致性分割与渐进式蒸馏网络的动态视觉SLAM方法

孙进,  徐添一乐,  赵海涛

孙进, 徐添一乐, 赵海涛. 融合深度一致性分割与渐进式蒸馏网络的动态视觉SLAM方法[J]. 电子与信息学报. doi: 10.11999/JEIT260266
引用本文: 孙进, 徐添一乐, 赵海涛. 融合深度一致性分割与渐进式蒸馏网络的动态视觉SLAM方法[J]. 电子与信息学报. doi: 10.11999/JEIT260266
SUN Jin, XU Tianyile, ZHAO Haitao. Dynamic Visual SLAM Integrating Depth-Consistency Segmentation and Progressive Distillation Network[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260266
Citation: SUN Jin, XU Tianyile, ZHAO Haitao. Dynamic Visual SLAM Integrating Depth-Consistency Segmentation and Progressive Distillation Network[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260266

融合深度一致性分割与渐进式蒸馏网络的动态视觉SLAM方法

doi: 10.11999/JEIT260266 cstr: 32379.14.JEIT260266
基金项目: 国家自然科学基金(62203231);海洋工程国家重点实验室(上海交通大学)开放基金资助(GKZD010084);江苏省研究生科研与实践创新计划项目(SJCX25_0374)
详细信息
    作者简介:

    孙进:男,博士,副教授,研究方向为惯性导航与组合导航,邮箱 sunjin@njupt.edu.cn

    徐添一乐:男,硕士研究生,研究方向为惯性导航与组合导航,邮箱 1224077119@njupt.edu.cn

    赵海涛:男,博士,教授,研究方向为多源融合导航,邮箱 zhaoht@njupt.edu.cn

    通讯作者:

    孙进 sunjin@njupt.edu.cn

  • 中图分类号: TN919.85

Dynamic Visual SLAM Integrating Depth-Consistency Segmentation and Progressive Distillation Network

Funds: National Natural Science Foundation of China under Grant 62203231, the Open Fund of State Key Laboratory of Ocean Engineering under Grant GKZD010084, the Postgraduate Research & Practice Innovation Program of Jiangsu Province under Grant SJCX25_0374.
  • 摘要: 针对动态环境下视觉同步定位与地图构建(SLAM)易受运动物体干扰,且现有深度学习方法难以在定位精度与实时性之间取得平衡的问题,本文提出一种轻量化实时动态SLAM系统。首先,提出深度感知掩码算法,利用彩色-深度(RGB-D)深度一致性约束与形态学操作细化前景区域,在一定程度上缓解了基于包围框的目标检测带来的过度掩膜及背景特征误剔除问题。其次,设计轻量化检测网络,引入部分卷积(PConv)与倒残差移动模块重构检测主干,显著降低计算冗余;同时,采用渐进式知识蒸馏策略将教师网络的高维语义特征迁移至学生模型,以补偿模型轻量化带来的精度损失。在TUM与Bonn数据集上的实验表明,本文方法在TUM RGB-D的walking_xyz高动态序列上,绝对轨迹误差(ATE)较ORB-SLAM3降低约95.1%;在输入分辨率为640×640时,模型计算量以109次浮点运算(GFLOPs)计,由7.1降低至5.6,完整SLAM系统在TUM RGB-D和Bonn RGB-D数据集上的平均运行速度分别为24.71 帧/s和27.11 帧/s。该方法在定位鲁棒性与计算效率之间取得了较好的平衡,在中等算力图形处理器(GPU)平台上具备实时动态状态估计能力。
  • 图  1  DCPD-SLAM系统总体框架

    图  2  矩形包围框“过度掩膜”现象与深度感知掩码生成效果对比

    图  3  Light-DEIM轻量化主干网络架构

    图  4  渐进式知识蒸馏框架

    图  5  ATE轨迹误差曲线对比图

    图  6  ATE误差分布箱型图

    表  3  蒸馏策略消融实验对比

    序号截断比例 ρ (T/Tmax)逻辑权重($ {\lambda }_{\text{logic}} $$ {\lambda }_{\text{logic}} $)特征权重($ {\lambda }_{\text{feat}} $$ {\lambda }_{\text{feat}} $)$ \text{mAP}_{50-95}^{\text{val}} $$ \text{mAP}_{50}^{\text{val}} $
    教师模型---42.560.4
    1---36.252.1
    20.71.5039.355.4
    30.702.040.757.2
    41.01.52.040.856.5
    50.51.52.040.154.2
    60.71.02.041.057.1
    70.71.51.040.957.0
    80.71.53.041.257.7
    90.71.52.041.958.7
    下载: 导出CSV

    表  1  各检测模型性能对比

    模型轮数参数量(M)计算量(GFLOPs)$ \text{mAP}_{50-95}^{\text{val}} $$ \text{mAP}_{50}^{\text{val}} $延迟(ms)吞吐率(FPS)
    YOLOv10-N5002.46.737.256.51.85312
    YOLOv10-S50072246.363.02.49282
    YOLOv11-N5002.87.238.557.81.50322
    YOLOv11-S50092247.063.92.50271
    Light-DEIM14835.636.252.12.08323
    Light-DEIM(KD)14835.641.958.72.08323
    DEIM-N1484742.560.42.12304
    DEIM-S120102549.065.93.49254
    注:本表FPS为检测网络独立推理速度,不包含深度掩码生成、ORB特征提取、动态特征剔除和SLAM后端优化;延迟为 TensorRT 16位浮点(Floating Point 16, FP16)、Batch Size=1 条件下仅统计模型前向推理得到的单帧平均延迟;吞吐率为验证过程中依据模型推理与后处理总耗时统计的平均处理帧率。
    下载: 导出CSV

    表  2  Light-DEIM结构消融实验对比

    模型参数量(M)计算量(GFLOPs)$ \text{mAP}_{50-95}^{\text{val}} $$ \text{mAP}_{50}^{\text{val}} $延迟(ms)平均ATE(m)平均FPS
    DEIM-N4.07.042.560.42.120.03824.04
    Ours w/o PConv3.36.242.059.12.180.03025.66
    Ours w/o iRMB2.85.139.255.91.960.04927.31
    Ours w/o Window Attention2.95.740.557.02.020.03626.88
    Ours w/o Width Scaling3.66.542.259.42.200.02725.41
    Ours Full3.05.641.958.72.080.02826.44
    下载: 导出CSV

    表  4  TUM RGB-D数据集ATE和RPE误差对比

    序列ORB-SLAM3DynaSLAMDS-SLAMRDS-SLAMYOLOv8-SLAMOurs ATEOurs RPE
    w/half0.3540.0200.0230.0130.0300.0140.5340.0560.0320.0140.026±0.0020.014±0.001
    w/rpy0.7670.0290.0830.0350.3540.0220.1590.0280.0390.0230.033±0.0030.020±0.002
    w/xyz0.3870.0210.1200.0300.1330.0170.2320.0370.0220.0130.019±0.0020.011±0.001
    s/half0.0630.0080.0170.0140.0160.0100.0270.0120.0840.0200.012±0.0010.013±0.001
    s/xyz0.4370.0080.0340.0100.1100.0090.2760.0070.0140.0110.015±0.0020.006±0.001
    注:表4—表5中Ours列结果均为5次独立运行的均值±标准差,每列左侧数据为ATE,右侧为RPE。
    下载: 导出CSV

    表  5  Bonn RGB-D数据集ATE和RPE误差对比

    序列ORB-SLAM3DynaSLAMDS-SLAMRDS-SLAMYOLOv8-SLAMOurs ATEOurs RPE
    balloon0.0480.0230.0310.0340.0560.0240.1440.0300.0340.0200.034±0.0030.021±0.002
    crowd11.8780.0280.0250.0140.0690.0240.1040.0200.0790.0310.022±0.0020.014±0.001
    crowd20.7190.0990.0360.0200.0820.0330.0810.0220.2160.0620.031±0.0040.018±0.002
    crowd30.1910.0340.0630.0410.0760.0480.0770.0370.0450.0240.037±0.0030.023±0.002
    synchronous1.0130.0270.2120.1650.1230.0230.0360.0160.0290.0160.010±0.0010.013±0.001
    synchronous21.0980.0220.0080.0060.0570.0110.0360.0180.0110.0180.008±0.0010.009±0.001
    下载: 导出CSV

    表  6  标准检测框与深度感知掩码的特征剔除效果对比

    序列 标准框ATE
    (m)
    深度感知掩码ATE
    (m)
    最大连通域筛选ATE
    (m)
    本文ATE
    (m)
    ATE下降率
    (%)
    动态点剔除率
    (%)
    静态点误剔除率
    (%)
    w/xyz 0.027 0.023 0.021 0.019 29.6 96.8 11.2
    crowd1 0.049 0.031 0.026 0.022 55.1 95.9 10.5
    crowd2 0.070 0.045 0.037 0.031 55.7 94.7 12.8
    crowd3 0.049 0.040 0.038 0.037 24.5 94.2 15.1
    平均值 0.049 0.035 0.031 0.027 - 95.4 12.4
    下载: 导出CSV

    表  7  TUM与Bonn RGB-D动态数据集运行速度对比

    模型 w/half w/rpy w/xyz balloon crowd1 crowd2 TUM平均值 Bonn平均值
    ORB-SLAM3 41.65 42.64 38.37 33.66 29.07 31.58 40.89 31.44
    DynaSLAM 0.31 2.11 1.23 0.57 0.54 0.14 1.22 0.42
    DS-SLAM 19.58 20.15 18.12 10.21 11.32 12.35 19.28 11.29
    Crowd-SLAM 35.21 32.41 40.21 29.61 33.43 27.92 35.94 30.32
    RDS-SLAM 42.31 44.63 42.50 44.32 46.21 45.62 43.15 45.38
    YOLOv8-SLAM 8.21 8.68 8.52 9.14 9.37 8.67 8.47 9.06
    Ours 25.01 27.17 21.95 26.93 26.61 27.78 24.71 27.11
    下载: 导出CSV
  • [1] 陈丹, 陈浩, 王子晨, 等. 多层ICP闭环检测下的误差状态卡尔曼滤波多模态融合SLAM[J]. 电子与信息学报, 2025, 47(5): 1517–1528. doi: 10.11999/JEIT240980.

    CHEN Dan, CHEN Hao, WANG Zichen, et al. Error state Kalman filter multimodal fusion SLAM based on MICP closed-loop detection[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1517–1528. doi: 10.11999/JEIT240980.
    [2] CAMPOS C, ELVIRA R, RODRÍGUEZ J J G, et al. ORB-SLAM3: An accurate open-source library for visual, visual-inertial, and multimap SLAM[J]. IEEE Transactions on Robotics, 2021, 37(6): 1874–1890. doi: 10.1109/TRO.2021.3075644.
    [3] 罗元, 沈吉祥, 李方宇. 动态环境下基于深度学习的视觉SLAM研究综述[J]. 半导体光电, 2024, 45(1): 1–10. doi: 10.16818/j.issn1001-5868.2023112202.

    LUO Yuan, SHEN Jixiang, and LI Fangyu. Review of visual SLAM research based on deep learning in dynamic environments[J]. Semiconductor Optoelectronics, 2024, 45(1): 1–10. doi: 10.16818/j.issn1001-5868.2023112202.
    [4] BESCOS B, FÁCIL J M, CIVERA J, et al. DynaSLAM: Tracking, mapping, and inpainting in dynamic scenes[J]. IEEE Robotics and Automation Letters, 2018, 3(4): 4076–4083. doi: 10.1109/LRA.2018.2860039.
    [5] YU Chao, LIU Zuxin, LIU Xinjun, et al. DS-SLAM: A semantic visual SLAM towards dynamic environments[C]. 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems, Madrid, Spain, 2018: 1168–1174. doi: 10.1109/IROS.2018.8593691.
    [6] 梅天灿, 秦宇晟, 杨宏, 等. 动态场景下基于视觉同时定位与地图构建技术的多层次语义地图构建方法[J]. 电子与信息学报, 2023, 45(5): 1737–1746. doi: 10.11999/JEIT220153.

    MEI Tiancan, QIN Yusheng, YANG Hong, et al. Multilevel semantic maps based on visual simultaneous localization and mapping in dynamic scenarios[J]. Journal of Electronics & Information Technology, 2023, 45(5): 1737–1746. doi: 10.11999/JEIT220153.
    [7] ZHANG Yuhao, BUJANCA M, and LUJÁN M. NGD-SLAM: Towards real-time dynamic SLAM without GPU[C]. 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, Hangzhou, China, 2025: 3467–3473. doi: 10.1109/IROS60139.2025.11246202.
    [8] ZHONG Fangwei, WANG Sheng, ZHANG Ziqi, et al. Detect-SLAM: Making object detection and SLAM mutually beneficial[C]. 2018 IEEE Winter Conference on Applications of Computer Vision, Lake Tahoe, USA, 2018: 1001–1010. doi: 10.1109/WACV.2018.00115.
    [9] WU Wenxin, GUO Liang, GAO Hongli, et al. YOLO-SLAM: A semantic SLAM system towards dynamic environment with geometric constraint[J]. Neural Computing and Applications, 2022, 34(8): 6011–6026. doi: 10.1007/s00521-021-06764-3.
    [10] LI Feng, LIU Yuanyuan, ZHANG Kelong, et al. DDETR-SLAM: A transformer-based approach to pose optimisation in dynamic environments[J]. International Journal of Robotics and Automation, 2024, 39(5): 407–421. doi: 10.2316/J.2024.206-1063. (查阅网上资料,未找到本条文献年卷期信息,请确认).
    [11] 余浩扬, 李艳生, 肖凌励, 等. 面向动态环境的巡检机器人轻量级语义视觉SLAM框架[J]. 电子与信息学报, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.

    YU Haoyang, LI Yansheng, XIAO Lingli, et al. A lightweight semantic visual simultaneous localization and mapping framework for inspection robots in dynamic environments[J]. Journal of Electronics & Information Technology, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.
    [12] LIU Yubao and MIURA J. RDS-SLAM: Real-time dynamic SLAM using semantic segmentation methods[J]. IEEE Access, 2021, 9: 23772–23785. doi: 10.1109/ACCESS.2021.3050617.
    [13] LIU Yang, GUO Chi, LUO Yarong, et al. DynaMeshSLAM: A mesh-based dynamic visual SLAMMOT method[J]. IEEE Robotics and Automation Letters, 2024, 9(6): 5791–5798. doi: 10.1109/LRA.2024.3396103.
    [14] HUANG Shihua, LU Zhichao, CUN Xiaodong, et al. DEIM: DETR with improved matching for fast convergence[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2025: 15162–15171.
    [15] CHEN Jierun, KAO S H, HE Hao, et al. Run, don’t walk: Chasing higher FLOPS for faster neural networks[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 12021–12031. doi: 10.1109/CVPR52729.2023.01157.
    [16] ZHANG Jiangning, LI Xiangtai, LI Jian, et al. Rethinking mobile block for efficient attention-based models[C]. Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision, Paris, France, 2023: 1389–1400. doi: 10.1109/ICCV51070.2023.00134.
    [17] LIU Ze, LIN Yutong, CAO Yue, et al. Swin transformer: Hierarchical vision transformer using shifted windows[C]. Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 9992–10002. doi: 10.1109/ICCV48922.2021.00986.
    [18] TAN Mingxing and LE Q V. EfficientNet: Rethinking model scaling for convolutional neural networks[C]. Proceedings of the 36th International Conference on Machine Learning, Long Beach, USA, 2019: 6105–6114.
    [19] 陈雷, 杨吉斌, 曹铁勇, 等. 一种基于Transformer特征金字塔的自蒸馏目标分割方法[J]. 电子与信息学报, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.

    CHEN Lei, YANG Jibin, CAO Tieyong, et al. A self-distillation object segmentation method based on transformer feature pyramid[J]. Journal of Electronics & Information Technology, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.
    [20] YANG Zhendong, LI Zhe, SHAO Mingqi, et al. Masked generative distillation[C]. Proceedings of the 17th European Conference on Computer Vision, Tel Aviv, Israel, 2022: 53–69. doi: 10.1007/978-3-031-20083-0_4.
    [21] CHANG Jiahao, WANG Shuo, XU Haiming, et al. DETRDistill: A universal knowledge distillation framework for DETR-families[C]. Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision, Paris, France, 2023: 6875–6885. doi: 10.1109/ICCV51070.2023.00635.
    [22] STURM J, ENGELHARD N, ENDRES F, et al. A benchmark for the evaluation of RGB-D SLAM systems[C]. 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vilamoura-Algarve, Portugal, 2012: 573–580. doi: 10.1109/IROS.2012.6385773.
    [23] PALAZZOLO E, BEHLEY J, LOTTES P, et al. ReFusion: 3D reconstruction in dynamic environments for RGB-D cameras exploiting residuals[C]. 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems, Macau, China, 2019: 7855–7862. doi: 10.1109/IROS40897.2019.8967590.
    [24] WANG Ao, CHEN Hui, LIU Lihao, et al. YOLOv10: Real-time end-to-end object detection[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 3429.
    [25] SAPKOTA R, FLORES-CALERO M, QURESHI R, et al. YOLO advances to its genesis: A decadal and comprehensive review of the You Only Look Once (YOLO) series[J]. Artificial Intelligence Review, 2025, 58(9): 274. doi: 10.1007/s10462-025-11253-3.
    [26] LI Yanke, SHEN Huabo, FU Yaping, et al. A method of dense point cloud SLAM based on improved YOLOV8 and fused with ORB-SLAM3 to cope with dynamic environments[J]. Expert Systems with Applications, 2024, 255: 124918. doi: 10.1016/j.eswa.2024.124918.
    [27] SOARES J C V, GATTASS M, and MEGGIOLARO M A. Crowd-SLAM: Visual SLAM towards crowded environments using object detection[J]. Journal of Intelligent & Robotic Systems, 2021, 102(2): 50. doi: 10.1007/s10846-021-01414-1.
  • 加载中
图(6) / 表(7)
计量
  • 文章访问数:  16
  • HTML全文浏览量:  2
  • PDF下载量:  0
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-03-11
  • 修回日期:  2026-09-28
  • 录用日期:  2026-09-28
  • 网络出版日期:  2026-10-10

目录

    /

    返回文章
    返回