DroneRFc-MM: Anti-UAV Multi-modal Detection Measured Dataset
-
摘要: 多模态信息融合技术可有效改善反无人机探测系统的泛化能力、鲁棒性与场景适配能力。针对现有反无人机探测数据集在模态种类、无人机型号和标注信息等方面的缺陷,该文公开了DroneRFc-MM多模态反无人机探测数据集。该数据集同步采集了云台相机、广角相机、射频天线、激光雷达、毫米波雷达和传声器阵列6类传感器数据,覆盖6种消费级无人机机型,并提供机型、定位、姿态和速度等细粒度标注,可支持目标探测、机型识别和运动方向推理等任务,并提供了易用的样本提取工具代码。最后,作为该数据集的使用示范,该文评估了Qwen系列最新模型在无人机飞行方向推理任务的表现。Abstract:
Multimodal data fusion can effectively improve the generalization, robustness, and scene adaptability of counter-unmanned aerial vehicle (UAV) detection systems. To address the limitations of existing counter-UAV datasets in sensing modalities, UAV models, and annotation granularity, the paper releases DroneRFc-MM, a multimodal counter-UAV detection dataset. DroneRFc-MM synchronously collects data from six types of sensors, including a pan-tilt-zoom camera, a wide-angle (fisheye) camera, radio-frequency antennas, LiDAR, millimeter-wave radar, and a microphone array. The dataset covers six consumer UAV models and provides fine-grained annotations such as model category, position, attitude, and velocity. It supports multiple tasks, including target detection, UAV model recognition, and motion-direction reasoning, and is accompanied by user-friendly sample extraction tools. Finally, as a demonstration of dataset usage, the paper evaluates the performance of recent Qwen-series models on UAV flight-direction reasoning. Objective: The work aims to build a more comprehensive benchmark for multimodal counter-UAV detection. Existing datasets often cover limited sensing modalities and UAV models, making it difficult to represent realistic low-altitude scenarios and diverse target characteristics. Their annotations are also typically coarse, such as category labels or bounding boxes, and thus cannot fully support downstream tasks requiring precise spatial, motion, and cross-modal information. DroneRFc-MM addresses these gaps by providing synchronized multimodal data, richer UAV coverage, and fine-grained annotations for detection, model recognition, trajectory analysis, motion reasoning, and multimodal fusion evaluation. Methods: The DroneRFc-MM dataset was synchronously captured via six heterogeneous sensors—including a pan-tilt-zoom(PTZ) camera, fisheye camera, radio frequency(RF) antenna, LiDAR, millimeter-wave radar, and microphone array—on an open rooftop of a university in Zhejiang Province. Featuring a representative urban low-altitude scenario, this dataset contains data recordings of six consumer-grade DJI drones. All devices were time-synchronized via network timestamp, and drones flew in rectangular and vertical reciprocating trajectories within 20–60 meters. Fine-grained annotations including drone type, position, attitude and velocity were provided. For flight direction reasoning task, 5-second multimodal clips were generated: videos for cameras and RF spectrograms, audio for microphones, text coordinates for radar point clouds. Zero-shot inference was conducted on Qwen 3.6-Plus and Qwen 3.5-Omni-Plus models with unified prompts, and accuracy and inference time were evaluated by comparing predicted directions with ground truth calculated from drone positioning data. Conclusions: Experiments on Qwen-series models show that general-purpose multimodal large models can capture weak motion-related features from drone-related videos, audio, RF spectrograms and point clouds, but only achieve limited flight-direction reasoning accuracy ranging from 20% to 30%. Meanwhile, the long inference time and unstable latency make them unable to satisfy the real-time and stability demands of practical low-altitude surveillance systems. These results demonstrate that domain-specific pre-training, supervised fine-tuning, knowledge enhancement and lightweight inference optimization are essential for deploying multi-modal LLMs in real anti-UAV detection scenarios. Future work will focus on expanding the dataset scale and enriching application scenarios to support the development of intelligent and efficient low-altitude airspace management systems. -
Key words:
- Anti-UAV detection /
- multimodal dataset /
- low-altitude economy
-
表 1 常用反无人机探测设备比较
设备 探测范围(km) 优势 劣势 可见光相机 0~6 检测算法成熟,结果直观 容易受光照、遮挡影响 红外成像仪 0~6 可夜间工作,可信度高 成本高,分辨率低 X 波段雷达 3~15 探测距离远,可探知方位和速度 无法探测静止目标,使用受管制,虚警多 射频天线 0.5~10 探测距离远,包含信息丰富 成本高,城市同频干扰多,需指定频段 表 2 本数据集和现有反无人机探测多模态数据集比较
数据集 采集设备 无人机机型 Anti-UAV 可见光相机、红外相机 DJI: Phantom 4, Spreading Wings S1000, Tello MMAUD 立体相机、激光雷达、毫米波雷达、传声器阵列 DJI: Mavic 2, Mavic 3, Phantom 4, Avata, M300 DroneRFc-MM 云台相机、广角相机、射频天线、激光雷达、毫米波雷达、传声器阵列 DJI: Mini 2 SE, Mini 3, Mavic Air 2S, Air 3, Mavic 3, Avata 表 3 采集设备列表
探测设备 品牌型号 主要参数规格 PTZ相机 海康威视DS-2DC6423IW 分辨率2 560×1 440,23倍光学变焦 广角相机 海康威视DS-2CD3346 分辨率2 560×1 440,视场角H180°V98° 全向射频天线 Ettus VERT2450 配合NI USRP-2955,采样率100 MHz,中心频率2.45 GHz,单次连续采样时间50 ms 激光雷达 速腾聚创EM4 线数512,点频2592万点/秒,视场角H120°V27° 毫米波雷达 Arbe Phoenix 探测波段77-81 GHz,点云模式,视场角H100°V30° 电容传声器 爱华AWA14411 4通道阵列,采样率25.6 kHz,动态范围6.5-137 dB 表 4 无人机型号与采集编号
无人机型号 机型编号 DJI Mavic 3 A1 DJI Avata B1 DJI Mini 2 SE C1 DJI Mini 3 E1 DJI Air 3 F1 DJI Air 2s G1 表 5 飞行方向推理任务提示词
种类 提示词 提示词模板
(“{}”内为可变内容)“{data_type} showing a drone. {extra_explanation} Please identify the moving direction of the drone. Respond with only one uppercase letter: F for Forward (approaching the {device}), B for Backward(separating from the {device}), U for Upward, D for Downward, L for Left, R for Right, S for Stopping (moving distance < 2 m).” 云台相机 {data_type}=”Here is a PTZ camera video”, {device}=”camera” 广角相机 {data_type}=”Here is a fisheye camera video”, {device}=”camera” 射频天线 {data_type}=”Here is an RF spectrum video”, {device}=”antenna” 激光雷达 {data_type}=”Here are lidar pointcloud data”, {device}=”lidar”,
{extra_explanation}=” The data are in compact CSV format: each row is one point with three comma-separated values (x,y,z). Frames are separated by a blank line.”毫米波雷达
(视频输入){data_type}=”Here are bird-view, front-view, and right-view videos of a millimeter wave radar pointcloud.”, {device}=”radar” 毫米波雷达
(文本输入){data_type}=”Here are millimeter wave radar pointcloud data”, {device}=”antenna”,
{extra_explanation}=” The data are in compact CSV format: each row is one point with three comma-separated values (x,y,z). Frames are separated by a blank line.”传声器 {data_type}=”Here are 4 audio recordings”, {device}=”microphone array”, {extra_explanation}=” The audio uploading sequence is arranged from left to right in accordance with the mounting positions of the microphone array.” 表 6 各模态无人机飞行方向推理准确率
采集设备 输入模态 总问答 正确回答 准确率(%) 激光雷达 点云位置文本 317 97 30.60 传声器 音频 476 137 28.78 毫米波雷达 点云位置文本 469 125 26.65 广角相机 视频 477 120 25.15 云台相机 视频 476 111 23.31 射频天线 时频谱视频 318 70 22.01 毫米波雷达 点云三视图视频 469 95 20.25 -
[1] DOLATA M and SCHWABE G. Moving beyond privacy and airspace safety: Guidelines for just drones in policing[J]. Government Information Quarterly, 2023, 40(4): 101874. doi: 10.1016/j.giq.2023.101874. [2] LAKHWANI T S, SINJANA Y, and KAPOOR A P. Dynamic medical logistics with drone-truck collaboration and pheromone decay in ant colony optimization[J]. International Journal of Systems Science: Operations & Logistics, 2026, 13(1): 2626506. doi: 10.1080/23302674.2026.2626506. [3] XIE Yingdong, GUO Yan, and GAO Jie. Object detection and UAV inspection for intelligent agriculture-husbandry application research[C]. Proceedings of the 2025 International Conference on Smart Agriculture and Artificial Intelligence, Xi’an, China, 2025: 223–228. doi: 10.1145/3767624.3767656. [4] 钱志鸿, 王义君. 低空经济赋能者: 智能无人机技术体系综述与展望[J]. 电子与信息学报, 2026, 48(1): 1–33. doi: 10.11999/JEIT251246.QIAN Zhihong and WANG Yijun. Intelligent unmanned aerial vehicles for low-altitude economy: A review of the technology framework and future prospects[J]. Journal of Electronics & Information Technology, 2026, 48(1): 1–33. doi: 10.11999/JEIT251246. [5] 王威, 佘丁辰, 王加琪, 等. 多模型融合的无人机异常航迹校正方法[J]. 电子与信息学报, 2025, 47(5): 1332–1344. doi: 10.11999/JEIT241026.WANG Wei, SHE Dingchen, WANG Jiaqi, et al. Multi-model fusion-based abnormal trajectory correction method for unmanned aerial vehicles[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1332–1344. doi: 10.11999/JEIT241026. [6] LIU Ziyi, AN Pei, YANG You, et al. Vision-based drone detection in complex environments: A survey[J]. Drones, 2024, 8(11): 643. doi: 10.3390/drones8110643. [7] 俞宁宁, 毛盛健, 周成伟, 等. DroneRFa: 用于侦测低空无人机的大规模无人机射频信号数据集[J]. 电子与信息学报, 2024, 46(4): 1147–1156. doi: 10.11999/JEIT230570.YU Ningning, MAO Shengjian, ZHOU Chengwei, et al. DroneRFa: A large-scale dataset of drone radio frequency signals for detecting low-altitude drones[J]. Journal of Electronics & Information Technology, 2024, 46(4): 1147–1156. doi: 10.11999/JEIT230570. [8] 任俊宇, 俞宁宁, 周成伟, 等. DroneRFb-DIR: 用于非合作无人机个体识别的射频信号数据集[J]. 电子与信息学报, 2025, 47(3): 573–581. doi: 10.11999/JEIT240804.REN Junyu, YU Ningning, ZHOU Chengwei, et al. DroneRFb-DIR: An RF signal dataset for non-cooperative drone individual identification[J]. Journal of Electronics & Information Technology, 2025, 47(3): 573–581. doi: 10.11999/JEIT240804. [9] YU Ningning, WU Jiajun, ZHOU Chengwei, et al. Open set learning for RF-based drone recognition via signal semantics[J]. IEEE Transactions on Information Forensics and Security, 2024, 19: 9894–9909. doi: 10.1109/TIFS.2024.3463535. [10] YU Ningning, WU Jiajun, ZHOU Chengwei, et al. SMNet: Multi-drone detection and classification via monitoring flight control signals on spectrograms[J]. IEEE Transactions on Cognitive Communications and Networking, 2025, 11(5): 3245–3259. doi: 10.1109/TCCN.2025.3537100. [11] KÜMMRITZ S. The sound of surveillance: Enhancing machine learning-driven drone detection with advanced acoustic augmentation[J]. Drones, 2024, 8(3): 105. doi: 10.3390/drones8030105. [12] SUN Chunlin, MAO Xingpeng, TANG Zhibo, et al. Radar false alarm suppression based on target spatial temporal stationarity for UAV detecting[J]. Drones, 2024, 8(12): 699. doi: 10.3390/drones8120699. [13] JIANG Nan, WANG Kuiran, PENG Xiaoke, et al. Anti-UAV: A large-scale benchmark for vision-based UAV tracking[J]. IEEE Transactions on Multimedia, 2023, 25: 486–500. doi: 10.1109/TMM.2021.3128047. [14] YUAN Shenghai, YANG Yizhuo, NGUYEN T H, et al. MMAUD: A comprehensive multi-modal anti-UAV dataset for modern miniature drone threats[C]. Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 2024: 2745–2751. doi: 10.1109/ICRA57147.2024.10610957. [15] SHI Rui, YU Xiaodong, WANG Shengming, et al. RFUAV: A benchmark dataset for unmanned aerial vehicle detection and identification[J/OL]. arXiv preprint arXiv: 2503.09033, 2025. doi: 10.48550/arXiv.2503.09033. [16] ZHU Chen, ZHAO Zhouxiang, SHAN Zejing, et al. Robust target detection of intelligent integrated optical camera and mmWave radar system[J]. Digital Signal Processing, 2024, 145: 104336. doi: 10.1016/j.dsp.2023.104336. [17] SAKELLARIOU N, LALAS A, VOTIS K, et al. Multi-sensor fusion for UAV classification based on feature maps of image and radar data[J/OL]. arXiv preprint arXiv: 2410.16089, 2024. doi: 10.48550/arXiv.2410.16089. [18] ZHAO W X, ZHOU Kun, LI Junyi, et al. A survey of large language models[J]. Frontiers of Computer Science, 2026, 20(12): 2012627. doi: 10.1007/s11704-026-60308-3. [19] LI Junnan, LI Dongxu, SAVARESE S, et al. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models[C]. Proceedings of the 40th International Conference on Machine Learning, Honolulu, USA, 2023: 814. doi: 10.5555/3618408.3619222. [20] DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 words: Transformers for image recognition at scale[C]. 9th International Conference on Learning Representations, 2021. (查阅网上资料, 未找到本条文献出版地信息, 请确认). [21] RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[C]. Proceedings of the 38th International Conference on Machine Learning, 2021. (查阅网上资料, 未找到本条文献会议举办地信息, 请确认). [22] XI Zhiheng, CHEN Wenxiang, GUO Xin, et al. The rise and potential of large language model based agents: A survey[J]. Science China Information Sciences, 2025, 68(2): 121101. doi: 10.1007/s11432-024-4222-0. -
下载: