高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

UAVREL:无人机视频动态关系理解基准数据集

刘晓瑞 邓楚博 侯钟砚 严启炜 卢宛萱 侯英妍 于泓峰 孙显

刘晓瑞, 邓楚博, 侯钟砚, 严启炜, 卢宛萱, 侯英妍, 于泓峰, 孙显. UAVREL:无人机视频动态关系理解基准数据集[J]. 电子与信息学报. doi: 10.11999/JEIT260221
引用本文: 刘晓瑞, 邓楚博, 侯钟砚, 严启炜, 卢宛萱, 侯英妍, 于泓峰, 孙显. UAVREL:无人机视频动态关系理解基准数据集[J]. 电子与信息学报. doi: 10.11999/JEIT260221
LIU Xiaorui, DENG Chubo, HOU Zhongyan, YAN Qiwei, LU Wanxuan, HOU Yingyan, YU Hongfeng, SUN Xian. UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260221
Citation: LIU Xiaorui, DENG Chubo, HOU Zhongyan, YAN Qiwei, LU Wanxuan, HOU Yingyan, YU Hongfeng, SUN Xian. UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260221

UAVREL:无人机视频动态关系理解基准数据集

doi: 10.11999/JEIT260221 cstr: 32379.14.JEIT260221
基金项目: 国家自然科学基金青年科学基金42301437
详细信息
    作者简介:

    刘晓瑞:男,硕士生,研究方向为遥感图像智能解译

    邓楚博:男,副研究员,研究方向为遥感图像智能解译

    侯钟砚:男,博士生,研究方向为遥感图像智能解译

    严启炜:男,博士生,研究方向为遥感图像智能解译

    卢宛萱:女,副研究员,研究方向为遥感图像智能解译

    侯英妍:女,博士生,研究方向为遥感图像智能解译

    于泓峰:男,副研究员,研究方向为遥感图像智能解译

    孙显:男,研究员,研究方向为遥感图像智能解译

    通讯作者:

    邓楚博 dengcb@aircas.an.cn

  • 中图分类号: TP751.2

UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos

Funds: National Natural Science Foundation of China (Grant Number: 42301437)
  • 摘要: 视频场景图生成是一项从视频中提取目标及其相互关系的高级视觉理解任务。目前相关研究多聚焦于自然场景,在遥感领域仍有待深入探索。无人机视频视角变化明显且目标尺度差异大,使得数据标注成本显著增加,缺乏大规模基准数据集是现阶段无人机视频场景图生成研究的主要障碍。该文构建了一个无人机视频关系理解数据集,在50个无人机视频中标注了12个目标类别,共计876687个目标框;并基于目标框标注了20个关系类别,共计459164个关系三元组。同时,提出了一种时空超图增强关系理解方法,利用超图结构对无人机视频复杂目标关联进行建模,从而实现更有效的特征聚合,增强了无人机视频关系理解能力。为评估主流方法在该数据集上的性能表现,该文对目标检测任务及场景图生成的三个子任务进行了全面测试与对比。结果表明,所提方法在场景图生成三个子任务共18个精度指标中8个优于基线方法。
  • 图  1  动态关系层级定义

    图  2  UAVREL与AeroEye数据集同场景关系层级对比

    图  3  关联半自动标注流程

    图  4  帧级标注数据分布统计图

    图  5  目标标注数量统计

    图  6  目标尺度分布箱线图

    图  7  关系标注数量统计

    图  8  主客体交互强度矩阵

    图  9  时空超图协同增强动态关系理解框架

    图  10  目标检测实验结果可视化示例

    图  11  场景图生成实验结果可视化示例

    表  1  自然场景与遥感场景中数据集统计对比

    数据集名称 视频 帧数 / 图片数 尺寸 目标
    标注
    关系
    标注
    目标
    类别
    目标
    数量
    关系
    类别
    关系
    数量
    年份



    Visual Phrase[14] × 2.8K / 8 3.3K 17 1.8K 2011
    VRD[15] × 5K / 100 / 70 38.0K 2016
    Visual Genome[16] × 108K 72~1280 33877 3.8M 42374 2.3M 2017
    VrR-VG[17] × 59K 72~1280 1600 282.5K 117 203.4K 2019
    VidVRD[18] 296.2K 1920×1080 35 / 132 55.6K 2017
    VidOR[19] 55.4K 640×360 80 38.6K 50 297.4K 2019
    Action Genome[13] 234.3K 1280×720 35 476.2K 25 1.7M 2020
    PVSG[20] 153K 1920×1080 126 / 57 / 2023



    DOTA[21] × 2.8K 800~4000 × 15 188.3K / / 2018
    DIOR-R[22] × 23.5K 800×800 × 20 192.5K / / 2022
    MONET[23] × 53K 800×600 × 2 162K / / 2023
    ReCon1M[24] × 21K 400~10000 60 873.8K 59 1.1M 2024
    UAVDT[25] 80K 1080×540 × 3 841.5K / / 2018
    UAVid[26] 0.3K 4096×2160 × 8 / / / 2020
    Brutal Running[27] 1K 227×227 × 1 / / / 2021
    UIT-ADrone[28] 206.2K 1920×1080 × 8 69.5K / / 2023
    AeroEye[29] 261.5K 3840×2160 57 2.2M 384 43M 2024
    UAVREL(本文) 40K 1024×540 12 884.8K 20 459.7K 2026
    下载: 导出CSV

    表  2  目标检测结果

    目标类别方法
    abcdefgh
    关口59.453.855.557.357.154.457.858.5
    非机动车19.419.115.421.019.817.721.621.5
    公交车53.452.351.454.153.854.056.354.8
    公交车站45.741.942.745.245.247.845.246.5
    汽车50.851.351.955.954.653.457.958.1
    路口58.559.558.860.056.756.163.362.4
    停车点49.150.547.548.046.849.651.150.9
    行人12.29.79.715.718.713.717.218.0
    道路55.155.054.056.254.552.757.057.4
    人行道46.647.645.544.043.044.049.248.9
    收费站57.657.345.748.139.347.854.053.5
    货车44.044.844.447.136.436.548.647.6

    mAP(%)46.045.443.546.143.844.048.348.2
    mAP@50(%)69.367.568.370.868.568.373.072.8
    注:表中(a)Faster R-CNN、(b)Cascade R-CNN、(c)RetinaNet、(d)ATSS、(e)AutoAssign、(f)FCOS、(g)VFNet、(h)DDOD
    下载: 导出CSV

    表  3  UAVREL数据集的场景图生成方法实验结果R@K(%)

    方法PredClsSGClsSGDet
    K=20K=50K=100K=20K=50K=100K=20K=50K=100
    IMP[38]71.777.979.368.573.975.243.148.951.4
    Motifs[39]77.683.585.270.776.578.053.459.060.8
    VCTree[40]74.680.982.671.476.978.548.355.458.0
    STTran[41]78.285.087.274.081.782.151.357.960.3
    下载: 导出CSV

    表  4  UAVREL数据集的场景图生成方法实验结果mR@K(%)

    方法 PreCls SGCls SGDet
    IMP[38] Motifs[39] VCTree[40] STTran[41] IMP[38] Motifs[39] VCTree[40] STTran[41] IMP[38] Motifs[39] VCTree[40] STTran[41]
    阻碍 52.4 66.8 74.9 65.1 48.7 56.4 58.6 61.6 21.1 63.3 29.5 37.9
    并排行驶 67.7 67.7 61.3 68.3 46.8 60.8 61.3 71.7 01.6 61.8 16.7 30.3
    驾驶跟随 58.7 60.9 53.2 76.2 49.0 64.9 55.9 65.4 10.5 66.9 9.2 17.4
    直行 72.6 82.8 81.6 92.2 72.0 81.9 80.6 89.2 58.5 67.0 66.1 72.3
    放行 58.1 91.9 64.5 90.8 48.8 48.8 47.4 49.7 20.9 42.1 37.6 43.8
    超车 84.9 90.9 80.1 91.3 74.4 79.6 76.2 81.3 22.6 57.2 41.3 50.6
    停车等待 60.1 75.1 81.4 82.4 62.2 74.3 76.3 77.2 43.2 70.8 58.6 62.1
    骑行穿过 64.6 97.8 92.8 96.6 70.7 66.0 58.6 68.3 35.9 45.3 37.3 41.8
    …… …… …… …… …… ……. …… …… …… …… …… …… ……
    避让 4.2 41.7 29.2 33.3 18.8 33.3 27.1 12.5 0.0 0.0 0.0 0.0
    汇入 0.0 20.0 10.0 10.0 0.0 5.0 10.0 0.0 0.0 5.0 5.0 0.0
    mR@20 42.7 50.8 48.0 46.8 41.5 42.8 43.2 44.4 19.3 32.8 23.2 26.3
    mR@50 54.2 68.6 65.6 67.2 51.1 62.0 60.0 62.3 25.1 45.5 36.8 39.4
    mR@100 57.7 73.9 70.3 73.6 54.4 66.8 64.1 67.7 28.2 49.9 42.1 45.6
    下载: 导出CSV

    表  5  STHG场景图生成方法实验结果(%)

    STHGPredClsSGClsSGDet
    K=20K=50K=100K=20K=50K=100K=20K=50K=100
    R@K78.484.987.174.881.982.353.559.160.7
    mR@K49.166.872.544.962.668.730.943.647.8
    H@K60.374.879.156.171.074.939.250.253.5
    下载: 导出CSV
  • [1] LIU Fang, WANG Jiahao, JIAO Licheng, et al. Remote sensing video tracking: Current status, challenges, and future[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 14338–14367. doi: 10.1109/JSTARS.2025.3573572.
    [2] MA Bokun, MU Caihong, LIU Yi, et al. RoSENet: Rotation and similarity enhancement network for multimodal remote sensing image land cover classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5511318. doi: 10.1109/TGRS.2025.3561850.
    [3] 金晶, 王峰. 分布式多卫星协同遥感图像场景分类方法[J]. 电子与信息学报, 2025, 47(12): 4677–4688. doi: 10.11999/JEIT250866.

    JIN Jing and WANG Feng. A distributed multi-satellite collaborative framework for remote sensing scene classification[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4677–4688. doi: 10.11999/JEIT250866.
    [4] GAO Feng, JIN Xuepeng, ZHOU Xiaowei, et al. MSFMamba: Multiscale feature fusion state space model for multisource remote sensing image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5504116. doi: 10.1109/TGRS.2025.3535622.
    [5] ZHANG Yin, YE Mu, ZHU Guiyi, et al. FFCA-YOLO for small object detection in remote sensing images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5611215. doi: 10.1109/TGRS.2024.3363057.
    [6] XIAO Yao, XU Tingfa, YU Xin, et al. A lightweight fusion strategy with enhanced interlayer feature correlation for small object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4708011. doi: 10.1109/TGRS.2024.3457155.
    [7] CHEN Tianxiang, YE Zi, TAN Zhentao, et al. MiM-ISTD: Mamba-in-mamba for efficient infrared small-target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007613. doi: 10.1109/TGRS.2024.3485721.
    [8] 姚婷婷, 肇恒鑫, 冯子豪, 等. 上下文感知多感受野融合网络的定向遥感目标检测[J]. 电子与信息学报, 2025, 47(1): 233–243. doi: 10.11999/JEIT240560.

    YAO Tingting, ZHAO Hengxin, FENG Zihao, et al. A context-aware multiple receptive field fusion network for oriented object detection in remote sensing images[J]. Journal of Electronics & Information Technology, 2025, 47(1): 233–243. doi: 10.11999/JEIT240560.
    [9] SUN Xian, WANG Peijin, YAN Zhiyuan, et al. FAIR1M: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 184: 116–130. doi: 10.1016/j.isprsjprs.2021.12.004.
    [10] 周国宇, 张菁, 闫伊, 等. 聚焦注意力与紧致特征融合Transformer的城市遥感影像语义分割[J]. 电子与信息学报, 2025, 47(12): 4790–4800. doi: 10.11999/JEIT250812.

    ZHOU Guoyu, ZHANG Jing, YAN Yi, et al. A focused attention and feature compact fusion Transformer for semantic segmentation of urban remote sensing images[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4790–4800. doi: 10.11999/JEIT250812.
    [11] CHEN Keyan, LIU Chenyang, CHEN Hao, et al. RSPrompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4701117. doi: 10.1109/TGRS.2024.3356074.
    [12] 于国栋, 蒋一纯, 刘云清, 等. 一种空间语义联合感知的红外无人机目标跟踪方法[J]. 电子与信息学报, 2025, 47(11): 4242–4253. doi: 10.11999/JEIT250613.

    YU Guodong, JIANG Yichun, LIU Yunqing, et al. A spatial-semantic combine perception for infrared UAV target tracking[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4242–4253. doi: 10.11999/JEIT250613.
    [13] JI Jingwei, KRISHNA R, LI Feifei, et al. Action genome: Actions as compositions of spatio-temporal scene graphs[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 10233–10244. doi: 10.1109/CVPR42600.2020.01025.
    [14] SADEGHI M A and FARHADI A. Recognition using visual phrases[C]. Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recogniti, Colorado Springs, USA, 2011: 1745–1752. doi: 10.1109/CVPR.2011.5995711.
    [15] LU Cewu, KRISHNA R, BERNSTEIN M, et al. Visual relationship detection with language priors[C]. 14th European Conference Computer Vision-ECCV 2016, Amsterdam, The Netherlands, 2016: 852–869. doi: 10.1007/978-3-319-46448-0_51.
    [16] KRISHNA R, ZHU Yuke, GROTH O, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations[J]. International Journal of Computer Vision, 2017, 123(1): 32–73. doi: 10.1007/s11263-016-0981-7.
    [17] LIANG Yuanzhi, BAI Yalong, ZHANG Wei, et al. VrR-VG: Refocusing visually-relevant relationships[C]. 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Korea, 2019: 10402–10411. doi: 10.1109/ICCV.2019.01050.
    [18] SHANG Xindi, REN Tongwei, GUO Jingfan, et al. Video visual relation detection[C]. Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, USA, 2017: 1300–1308. doi: 10.1145/3123266.3123380.
    [19] SHANG Xindi, DI Donglin, XIAO Junbin, et al. Annotating objects and relations in user-generated videos[C]. Proceedings of the 2019 on International Conference on Multimedia Retrieval, Ottawa, Canada, 2019: 279–287.
    [20] YANG Jingkang, PENG Wenxuan, LI Xiangtai, et al. Panoptic video scene graph generation[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 18675–18685. doi: 10.1109/CVPR52729.2023.01791.
    [21] XIA Guisong, BAI Xiang, DING Jian, et al. DOTA: A large-scale dataset for object detection in aerial images[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 3974–3983. doi: 10.1109/CVPR.2018.00418.
    [22] CHENG Gong, WANG Jiabao, LI Ke, et al. Anchor-free oriented proposal generator for object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5625411. doi: 10.1109/TGRS.2022.3183022.
    [23] RIZ L, CARAFFA A, BORTOLON M, et al. The MONET dataset: Multimodal drone thermal dataset recorded in rural scenarios[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, Canada, 2023: 2546–2554. doi: 10.1109/CVPRW59228.2023.00253.
    [24] YAN Qiwei, DENG Chubo, LIU Chenglong, et al. ReCon1M: A large-scale benchmark dataset for relation comprehension in remote sensing imagery[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 4507022. doi: 10.1109/TGRS.2025.3589986.
    [25] DU Dawei, QI Yuankai, YU Hongyang, et al. The unmanned aerial vehicle benchmark: Object detection and tracking[C]. 15th European Conference Computer Vision-ECCV 2018, Munich, Germany, 2018: 370–386. doi: 10.1007/978-3-030-01249-6_23.
    [26] LYU Ye, VOSSELMAN G, XIA Guisong, et al. UAVid: A semantic segmentation dataset for UAV imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 165: 108–119. doi: 10.1016/j.isprsjprs.2020.05.009.
    [27] HAMDI S, BOUINDOUR S, SNOUSSI H, et al. End-to-end deep one-class learning for anomaly detection in UAV video stream[J]. Journal of Imaging, 2021, 7(5): 90. doi: 10.3390/jimaging7050090.
    [28] TRAN T M, VU T N, NGUYEN T V, et al. UIT-ADrone: A novel drone dataset for traffic anomaly detection[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2023, 16: 5590–5601. doi: 10.1109/JSTARS.2023.3285905.
    [29] NGUYEN T T, NGUYEN P, LI Xin, et al. CYCLO: Cyclic graph transformer approach to multi-object relationship modeling in aerial videos[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 2868.
    [30] REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137–1149. doi: 10.1109/tpami.2016.2577031.
    [31] CAI Zhaowei and VASCONCELOS N. Cascade R-CNN: Delving into high quality object detection[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 6154–6162. doi: 10.1109/CVPR.2018.00644.
    [32] LIN T Y, GOYAL P, GIRSHICK R, et al. Focal loss for dense object detection[C]. 2017 IEEE International Conference on Computer Vision, Venice, Italy, 2017: 2999–3007. doi: 10.1109/ICCV.2017.324.
    [33] ZHANG Shifeng, CHI Cheng, YAO Yongqiang, et al. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 9756–9765. doi: 10.1109/CVPR42600.2020.00978.
    [34] ZHU Benjin, WANG Jianfeng, JIANG Zhengkai, et al. AutoAssign: Differentiable label assignment for dense object detection[EB/OL]. https://arxiv.org/abs/2007.03496, 2020.
    [35] TIAN Zhi, SHEN Chunhua, CHEN Hao, et al. FCOS: A simple and strong anchor-free object detector[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(4): 1922–1933. doi: 10.1109/tpami.2020.3032166.
    [36] ZHANG Haoyang, WANG Ying, DAYOUB F, et al. VarifocalNET: An IoU-aware dense object detector[C]. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 8501–8519. doi: 10.1109/CVPR46437.2021.00841.
    [37] CHEN Zehui, YANG Chenhongyi, LI Qiaofei, et al. Disentangle your dense object detector[C]. Proceedings of the 29th ACM International Conference on Multimedia, 2021: 4939–4948. doi: 10.1145/3474085.3475351. (查阅网上资料,未找到本条文献出版地信息,请确认).
    [38] TANG Kaihua, NIU Yulei, HUANG Jianqiang, et al. Unbiased scene graph generation from biased training[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 3713–3722. doi: 10.1109/CVPR42600.2020.00377.
    [39] ZELLERS R, YATSKAR M, THOMSON S, et al. Neural motifs: Scene graph parsing with global context[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 5831–5840. doi: 10.1109/CVPR.2018.00611.
    [40] TANG Kaihua, ZHANG Hanwang, WU Baoyuan, et al. Learning to compose dynamic tree structures for visual contexts[C]. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 6612–6621. doi: 10.1109/CVPR.2019.00678.
    [41] CONG Yuren, LIAO Wentong, ACKERMANN H, et al. Spatial-temporal transformer for dynamic scene graph generation[C]. 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 16352–16362. doi: 10.1109/ICCV48922.2021.01606.
  • 加载中
图(11) / 表(5)
计量
  • 文章访问数:  27
  • HTML全文浏览量:  7
  • PDF下载量:  2
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-03-01
  • 修回日期:  2026-07-28
  • 录用日期:  2026-07-28
  • 网络出版日期:  2026-09-01

目录

    /

    返回文章
    返回