Dynamic Visual SLAM Integrating Depth-Consistency Segmentation and Progressive Distillation Network
-
摘要: 针对动态环境下视觉同步定位与地图构建(SLAM)易受运动物体干扰,且现有深度学习方法难以在定位精度与实时性之间取得平衡的问题,本文提出一种轻量化实时动态SLAM系统。首先,提出深度感知掩码算法,利用彩色-深度(RGB-D)深度一致性约束与形态学操作细化前景区域,在一定程度上缓解了基于包围框的目标检测带来的过度掩膜及背景特征误剔除问题。其次,设计轻量化检测网络,引入部分卷积(PConv)与倒残差移动模块重构检测主干,显著降低计算冗余;同时,采用渐进式知识蒸馏策略将教师网络的高维语义特征迁移至学生模型,以补偿模型轻量化带来的精度损失。在TUM与Bonn数据集上的实验表明,本文方法在TUM RGB-D的walking_xyz高动态序列上,绝对轨迹误差(ATE)较ORB-SLAM3降低约95.1%;在输入分辨率为640×640时,模型计算量以109次浮点运算(GFLOPs)计,由7.1降低至5.6,完整SLAM系统在TUM RGB-D和Bonn RGB-D数据集上的平均运行速度分别为24.71 帧/s和27.11 帧/s。该方法在定位鲁棒性与计算效率之间取得了较好的平衡,在中等算力图形处理器(GPU)平台上具备实时动态状态估计能力。Abstract:
Objective Visual Simultaneous Localization and Mapping (SLAM) is a fundamental technology for autonomous navigation and environmental perception in Global Navigation Satellite System (GNSS)-denied environments. However, conventional visual SLAM systems generally rely on a static-scene assumption. In dynamic environments containing pedestrians, vehicles, and other moving objects, dynamic feature mismatches may violate geometric constraints and introduce unstable observations into pose estimation, resulting in trajectory drift or tracking degradation. Existing dynamic SLAM methods commonly incorporate semantic priors to suppress such interference. Semantic segmentation-based methods can provide fine-grained pixel-level dynamic regions, but usually require relatively high computational cost. In contrast, object detection-based methods provide higher inference efficiency, but their rectangular detection boxes inevitably contain static background regions, which may lead to the erroneous removal of valid static features and cause an over-masking problem. This issue becomes more pronounced in texture-poor or crowded scenes where sufficient static features are important for maintaining stable geometric constraints. To alleviate the trade-off between dynamic-feature filtering accuracy and computational efficiency, this paper proposes Depth-Consistency and Progressive-Distillation SLAM (DCPD-SLAM), a lightweight dynamic visual SLAM system for real-time state estimation under moderate computational resources. Methods DCPD-SLAM adopts a parallel dual-thread architecture comprising semantic perception and visual SLAM tracking. Semantic inference and feature extraction run concurrently, with frame-level synchronization before dynamic-feature filtering to align semantic results with the current Red-Green-Blue-Depth (RGB-D) frame. First, a depth-aware masking strategy refines coarse detection boxes. For selected non-rigid dynamic targets, such as pedestrians, valid depth pixels within the Region of Interest (ROI) are used to estimate a robust reference depth. The reference depth combines the median and the lower 33rd percentile of valid depth measurements to reduce the influence of background pixels and isolated depth noise. An adaptive depth-consistency threshold is then used to generate an initial foreground mask. When the valid-depth ratio is insufficient, the original detection box is retained as a fallback to maintain filtering reliability. Connected Component Analysis (CCA) and morphological dilation suppress isolated noise and compensate for boundary leakage, yielding a compact pixel-level pseudo-mask. Other predefined dynamic categories retain detection-box-based filtering to avoid additional mask-generation overhead.Second, a lightweight detector, Light-DEIM, reduces semantic perception costs. Partial Convolution (PConv) is incorporated into shallow C2f-FasterBlock stages to reduce spatial redundancy, while deeper C2f-Inverted Residual Mobile Block (iRMB) modules with window attention preserve semantic representations. Balanced width scaling further reduces model complexity while retaining sufficient high-level semantic capacity. Finally, progressive knowledge distillation compensates for representation loss caused by model lightweighting. Masked Generative Distillation (MGD) aligns intermediate features, while DETRDistill-based logits distillation aligns teacher–student predictions. During the first 70% of training epochs, Light-DEIM is trained with ground-truth supervision and feature and output distillation from a frozen DEIM-N teacher. For the remaining epochs, the distillation branches are disabled, and the student is optimized solely with ground-truth labels. Results and Discussions Experiments are conducted on the TUM RGB-D and Bonn RGB-D dynamic datasets. For the detection network, Light-DEIM reduces the parameter count from 4.0 M to 3.0 M and the computational complexity from 7.1 Giga Floating-point Operations (GFLOPs) to 5.6 GFLOPs. Without knowledge distillation, the lightweight model achieves an $ \text{mAP}_{50-95}^{\text{val}} $ of 36.2%. After progressive knowledge distillation, the $ \text{mAP}_{50-95}^{\text{val}} $increases to 41.9%, approaching the 42.5% performance of the DEIM-N teacher model while maintaining a validation throughput of 323 Frames Per Second (FPS) ( Table 1 ). Ablation experiments further show that PConv mainly contributes to computational efficiency, whereas iRMB and Window Attention improve semantic representation. The combination of feature-level and output-level distillation provides better accuracy recovery than either distillation term alone (Table 2 ,Table 3 ).For SLAM evaluation, DCPD-SLAM achieves an Absolute Trajectory Error (ATE) of 0.019 $ \pm $0.002 $ 0.019\pm 0.002 $ m on the highly dynamic walking_xyz sequence, compared with 0.387 m for ORB-SLAM3 (Table 4 ). The trajectory-error curves and error distributions further indicate relatively stable localization performance in highly dynamic scenes (Fig. 5 ,Fig. 6 ). In four representative dynamic sequences, the proposed depth-aware masking strategy achieves an average dynamic-point rejection rate of 95.4% while limiting the static-point erroneous rejection rate to 12.4% (Table 6 ), indicating that it can reduce unnecessary removal of static background features while maintaining effective dynamic-feature filtering. At the complete-system level, DCPD-SLAM achieves average processing rates of 24.71 FPS and 27.11 FPS on the TUM RGB-D and Bonn RGB-D datasets, respectively (Table 7 ), demonstrating real-time processing capability on the moderate-compute Graphics Processing Unit (GPU) platform.Conclusions DCPD-SLAM balances localization accuracy and computational efficiency in dynamic RGB-D environments. Depth-aware masking mitigates over-masking from detection-box filtering, preserving static background features. Light-DEIM and progressive knowledge distillation reduce semantic perception overhead while retaining detection capability for dynamic-feature filtering. Experiments demonstrate real-time processing and stable localization on a moderate-compute GPU platform. However, pseudo-mask quality depends on depth measurements, with potential degradation under severe depth noise, strong illumination changes, reflective or transparent surfaces, distant targets, and severe occlusions. Future work will explore Light Detection and Ranging (LiDAR)-visual fusion to improve adaptability and measurement stability in complex, large-scale environments. -
Key words:
- Visual SLAM /
- Dynamic environments /
- Depth-aware mask /
- Lightweight network /
- Knowledge distillation
-
表 3 蒸馏策略消融实验对比
序号 截断比例 ρ (T/Tmax) 逻辑权重($ {\lambda }_{\text{logic}} $$ {\lambda }_{\text{logic}} $) 特征权重($ {\lambda }_{\text{feat}} $$ {\lambda }_{\text{feat}} $) $ \text{mAP}_{50-95}^{\text{val}} $ $ \text{mAP}_{50}^{\text{val}} $ 教师模型 - - - 42.5 60.4 1 - - - 36.2 52.1 2 0.7 1.5 0 39.3 55.4 3 0.7 0 2.0 40.7 57.2 4 1.0 1.5 2.0 40.8 56.5 5 0.5 1.5 2.0 40.1 54.2 6 0.7 1.0 2.0 41.0 57.1 7 0.7 1.5 1.0 40.9 57.0 8 0.7 1.5 3.0 41.2 57.7 9 0.7 1.5 2.0 41.9 58.7 表 1 各检测模型性能对比
模型 轮数 参数量(M) 计算量(GFLOPs) $ \text{mAP}_{50-95}^{\text{val}} $ $ \text{mAP}_{50}^{\text{val}} $ 延迟(ms) 吞吐率(FPS) YOLOv10-N 500 2.4 6.7 37.2 56.5 1.85 312 YOLOv10-S 500 7 22 46.3 63.0 2.49 282 YOLOv11-N 500 2.8 7.2 38.5 57.8 1.50 322 YOLOv11-S 500 9 22 47.0 63.9 2.50 271 Light-DEIM 148 3 5.6 36.2 52.1 2.08 323 Light-DEIM(KD) 148 3 5.6 41.9 58.7 2.08 323 DEIM-N 148 4 7 42.5 60.4 2.12 304 DEIM-S 120 10 25 49.0 65.9 3.49 254 注:本表FPS为检测网络独立推理速度,不包含深度掩码生成、ORB特征提取、动态特征剔除和SLAM后端优化;延迟为 TensorRT 16位浮点(Floating Point 16, FP16)、Batch Size=1 条件下仅统计模型前向推理得到的单帧平均延迟;吞吐率为验证过程中依据模型推理与后处理总耗时统计的平均处理帧率。 表 2 Light-DEIM结构消融实验对比
模型 参数量(M) 计算量(GFLOPs) $ \text{mAP}_{50-95}^{\text{val}} $ $ \text{mAP}_{50}^{\text{val}} $ 延迟(ms) 平均ATE(m) 平均FPS DEIM-N 4.0 7.0 42.5 60.4 2.12 0.038 24.04 Ours w/o PConv 3.3 6.2 42.0 59.1 2.18 0.030 25.66 Ours w/o iRMB 2.8 5.1 39.2 55.9 1.96 0.049 27.31 Ours w/o Window Attention 2.9 5.7 40.5 57.0 2.02 0.036 26.88 Ours w/o Width Scaling 3.6 6.5 42.2 59.4 2.20 0.027 25.41 Ours Full 3.0 5.6 41.9 58.7 2.08 0.028 26.44 表 4 TUM RGB-D数据集ATE和RPE误差对比
序列 ORB-SLAM3 DynaSLAM DS-SLAM RDS-SLAM YOLOv8-SLAM Ours ATE Ours RPE w/half 0.354 0.020 0.023 0.013 0.030 0.014 0.534 0.056 0.032 0.014 0.026±0.002 0.014±0.001 w/rpy 0.767 0.029 0.083 0.035 0.354 0.022 0.159 0.028 0.039 0.023 0.033±0.003 0.020±0.002 w/xyz 0.387 0.021 0.120 0.030 0.133 0.017 0.232 0.037 0.022 0.013 0.019±0.002 0.011±0.001 s/half 0.063 0.008 0.017 0.014 0.016 0.010 0.027 0.012 0.084 0.020 0.012±0.001 0.013±0.001 s/xyz 0.437 0.008 0.034 0.010 0.110 0.009 0.276 0.007 0.014 0.011 0.015±0.002 0.006±0.001 注:表4—表5中Ours列结果均为5次独立运行的均值±标准差,每列左侧数据为ATE,右侧为RPE。 表 5 Bonn RGB-D数据集ATE和RPE误差对比
序列 ORB-SLAM3 DynaSLAM DS-SLAM RDS-SLAM YOLOv8-SLAM Ours ATE Ours RPE balloon 0.048 0.023 0.031 0.034 0.056 0.024 0.144 0.030 0.034 0.020 0.034±0.003 0.021±0.002 crowd1 1.878 0.028 0.025 0.014 0.069 0.024 0.104 0.020 0.079 0.031 0.022±0.002 0.014±0.001 crowd2 0.719 0.099 0.036 0.020 0.082 0.033 0.081 0.022 0.216 0.062 0.031±0.004 0.018±0.002 crowd3 0.191 0.034 0.063 0.041 0.076 0.048 0.077 0.037 0.045 0.024 0.037±0.003 0.023±0.002 synchronous 1.013 0.027 0.212 0.165 0.123 0.023 0.036 0.016 0.029 0.016 0.010±0.001 0.013±0.001 synchronous2 1.098 0.022 0.008 0.006 0.057 0.011 0.036 0.018 0.011 0.018 0.008±0.001 0.009±0.001 表 6 标准检测框与深度感知掩码的特征剔除效果对比
序列 标准框ATE
(m)深度感知掩码ATE
(m)最大连通域筛选ATE
(m)本文ATE
(m)ATE下降率
(%)动态点剔除率
(%)静态点误剔除率
(%)w/xyz 0.027 0.023 0.021 0.019 29.6 96.8 11.2 crowd1 0.049 0.031 0.026 0.022 55.1 95.9 10.5 crowd2 0.070 0.045 0.037 0.031 55.7 94.7 12.8 crowd3 0.049 0.040 0.038 0.037 24.5 94.2 15.1 平均值 0.049 0.035 0.031 0.027 - 95.4 12.4 表 7 TUM与Bonn RGB-D动态数据集运行速度对比
模型 w/half w/rpy w/xyz balloon crowd1 crowd2 TUM平均值 Bonn平均值 ORB-SLAM3 41.65 42.64 38.37 33.66 29.07 31.58 40.89 31.44 DynaSLAM 0.31 2.11 1.23 0.57 0.54 0.14 1.22 0.42 DS-SLAM 19.58 20.15 18.12 10.21 11.32 12.35 19.28 11.29 Crowd-SLAM 35.21 32.41 40.21 29.61 33.43 27.92 35.94 30.32 RDS-SLAM 42.31 44.63 42.50 44.32 46.21 45.62 43.15 45.38 YOLOv8-SLAM 8.21 8.68 8.52 9.14 9.37 8.67 8.47 9.06 Ours 25.01 27.17 21.95 26.93 26.61 27.78 24.71 27.11 -
[1] 陈丹, 陈浩, 王子晨, 等. 多层ICP闭环检测下的误差状态卡尔曼滤波多模态融合SLAM[J]. 电子与信息学报, 2025, 47(5): 1517–1528. doi: 10.11999/JEIT240980.CHEN Dan, CHEN Hao, WANG Zichen, et al. Error state Kalman filter multimodal fusion SLAM based on MICP closed-loop detection[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1517–1528. doi: 10.11999/JEIT240980. [2] CAMPOS C, ELVIRA R, RODRÍGUEZ J J G, et al. ORB-SLAM3: An accurate open-source library for visual, visual-inertial, and multimap SLAM[J]. IEEE Transactions on Robotics, 2021, 37(6): 1874–1890. doi: 10.1109/TRO.2021.3075644. [3] 罗元, 沈吉祥, 李方宇. 动态环境下基于深度学习的视觉SLAM研究综述[J]. 半导体光电, 2024, 45(1): 1–10. doi: 10.16818/j.issn1001-5868.2023112202.LUO Yuan, SHEN Jixiang, and LI Fangyu. Review of visual SLAM research based on deep learning in dynamic environments[J]. Semiconductor Optoelectronics, 2024, 45(1): 1–10. doi: 10.16818/j.issn1001-5868.2023112202. [4] BESCOS B, FÁCIL J M, CIVERA J, et al. DynaSLAM: Tracking, mapping, and inpainting in dynamic scenes[J]. IEEE Robotics and Automation Letters, 2018, 3(4): 4076–4083. doi: 10.1109/LRA.2018.2860039. [5] YU Chao, LIU Zuxin, LIU Xinjun, et al. DS-SLAM: A semantic visual SLAM towards dynamic environments[C]. 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems, Madrid, Spain, 2018: 1168–1174. doi: 10.1109/IROS.2018.8593691. [6] 梅天灿, 秦宇晟, 杨宏, 等. 动态场景下基于视觉同时定位与地图构建技术的多层次语义地图构建方法[J]. 电子与信息学报, 2023, 45(5): 1737–1746. doi: 10.11999/JEIT220153.MEI Tiancan, QIN Yusheng, YANG Hong, et al. Multilevel semantic maps based on visual simultaneous localization and mapping in dynamic scenarios[J]. Journal of Electronics & Information Technology, 2023, 45(5): 1737–1746. doi: 10.11999/JEIT220153. [7] ZHANG Yuhao, BUJANCA M, and LUJÁN M. NGD-SLAM: Towards real-time dynamic SLAM without GPU[C]. 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, Hangzhou, China, 2025: 3467–3473. doi: 10.1109/IROS60139.2025.11246202. [8] ZHONG Fangwei, WANG Sheng, ZHANG Ziqi, et al. Detect-SLAM: Making object detection and SLAM mutually beneficial[C]. 2018 IEEE Winter Conference on Applications of Computer Vision, Lake Tahoe, USA, 2018: 1001–1010. doi: 10.1109/WACV.2018.00115. [9] WU Wenxin, GUO Liang, GAO Hongli, et al. YOLO-SLAM: A semantic SLAM system towards dynamic environment with geometric constraint[J]. Neural Computing and Applications, 2022, 34(8): 6011–6026. doi: 10.1007/s00521-021-06764-3. [10] LI Feng, LIU Yuanyuan, ZHANG Kelong, et al. DDETR-SLAM: A transformer-based approach to pose optimisation in dynamic environments[J]. International Journal of Robotics and Automation, 2024, 39(5): 407–421. doi: 10.2316/J.2024.206-1063. (查阅网上资料,未找到本条文献年卷期信息,请确认). [11] 余浩扬, 李艳生, 肖凌励, 等. 面向动态环境的巡检机器人轻量级语义视觉SLAM框架[J]. 电子与信息学报, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.YU Haoyang, LI Yansheng, XIAO Lingli, et al. A lightweight semantic visual simultaneous localization and mapping framework for inspection robots in dynamic environments[J]. Journal of Electronics & Information Technology, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301. [12] LIU Yubao and MIURA J. RDS-SLAM: Real-time dynamic SLAM using semantic segmentation methods[J]. IEEE Access, 2021, 9: 23772–23785. doi: 10.1109/ACCESS.2021.3050617. [13] LIU Yang, GUO Chi, LUO Yarong, et al. DynaMeshSLAM: A mesh-based dynamic visual SLAMMOT method[J]. IEEE Robotics and Automation Letters, 2024, 9(6): 5791–5798. doi: 10.1109/LRA.2024.3396103. [14] HUANG Shihua, LU Zhichao, CUN Xiaodong, et al. DEIM: DETR with improved matching for fast convergence[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2025: 15162–15171. [15] CHEN Jierun, KAO S H, HE Hao, et al. Run, don’t walk: Chasing higher FLOPS for faster neural networks[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 12021–12031. doi: 10.1109/CVPR52729.2023.01157. [16] ZHANG Jiangning, LI Xiangtai, LI Jian, et al. Rethinking mobile block for efficient attention-based models[C]. Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision, Paris, France, 2023: 1389–1400. doi: 10.1109/ICCV51070.2023.00134. [17] LIU Ze, LIN Yutong, CAO Yue, et al. Swin transformer: Hierarchical vision transformer using shifted windows[C]. Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 9992–10002. doi: 10.1109/ICCV48922.2021.00986. [18] TAN Mingxing and LE Q V. EfficientNet: Rethinking model scaling for convolutional neural networks[C]. Proceedings of the 36th International Conference on Machine Learning, Long Beach, USA, 2019: 6105–6114. [19] 陈雷, 杨吉斌, 曹铁勇, 等. 一种基于Transformer特征金字塔的自蒸馏目标分割方法[J]. 电子与信息学报, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.CHEN Lei, YANG Jibin, CAO Tieyong, et al. A self-distillation object segmentation method based on transformer feature pyramid[J]. Journal of Electronics & Information Technology, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735. [20] YANG Zhendong, LI Zhe, SHAO Mingqi, et al. Masked generative distillation[C]. Proceedings of the 17th European Conference on Computer Vision, Tel Aviv, Israel, 2022: 53–69. doi: 10.1007/978-3-031-20083-0_4. [21] CHANG Jiahao, WANG Shuo, XU Haiming, et al. DETRDistill: A universal knowledge distillation framework for DETR-families[C]. Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision, Paris, France, 2023: 6875–6885. doi: 10.1109/ICCV51070.2023.00635. [22] STURM J, ENGELHARD N, ENDRES F, et al. A benchmark for the evaluation of RGB-D SLAM systems[C]. 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vilamoura-Algarve, Portugal, 2012: 573–580. doi: 10.1109/IROS.2012.6385773. [23] PALAZZOLO E, BEHLEY J, LOTTES P, et al. ReFusion: 3D reconstruction in dynamic environments for RGB-D cameras exploiting residuals[C]. 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems, Macau, China, 2019: 7855–7862. doi: 10.1109/IROS40897.2019.8967590. [24] WANG Ao, CHEN Hui, LIU Lihao, et al. YOLOv10: Real-time end-to-end object detection[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 3429. [25] SAPKOTA R, FLORES-CALERO M, QURESHI R, et al. YOLO advances to its genesis: A decadal and comprehensive review of the You Only Look Once (YOLO) series[J]. Artificial Intelligence Review, 2025, 58(9): 274. doi: 10.1007/s10462-025-11253-3. [26] LI Yanke, SHEN Huabo, FU Yaping, et al. A method of dense point cloud SLAM based on improved YOLOV8 and fused with ORB-SLAM3 to cope with dynamic environments[J]. Expert Systems with Applications, 2024, 255: 124918. doi: 10.1016/j.eswa.2024.124918. [27] SOARES J C V, GATTASS M, and MEGGIOLARO M A. Crowd-SLAM: Visual SLAM towards crowded environments using object detection[J]. Journal of Intelligent & Robotic Systems, 2021, 102(2): 50. doi: 10.1007/s10846-021-01414-1. -
下载: