Advanced Search
Turn off MathJax
Article Contents
LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng. Fusing Global Perspective Rectification and Fine-Grained Semantic Decoupling for Language-Conditioned Robotic Grasping Detection[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260442
Citation: LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng. Fusing Global Perspective Rectification and Fine-Grained Semantic Decoupling for Language-Conditioned Robotic Grasping Detection[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260442

Fusing Global Perspective Rectification and Fine-Grained Semantic Decoupling for Language-Conditioned Robotic Grasping Detection

doi: 10.11999/JEIT260442 cstr: 32379.14.JEIT260442
Funds:  National Key R&D Program of China (Grant No. 2025YFB4712700), and the China Postdoctoral Science Foundation (Grant No. 2025M781675)
  • Received Date: 2026-04-14
  • Accepted Date: 2026-07-14
  • Rev Recd Date: 2026-07-14
  • Available Online: 2026-07-24
  •   Objective  Accurate grasping of target objects from language instructions is an essential prerequisite for service robots to achieve natural human-robot interaction. Existing research primarily relies on massive data-driven training or hierarchical feature fusion schemes for the cross-modal alignment between visual perception and textual instructions. However, these methods commonly overlook the heavy entanglement between target objects and background environments within low-level features, leading to a significant degradation in compositional generalization ability under cross-view and unseen background scenarios. To address these issues, this paper constructs a dual-view cross-scenario synchronous grasp detection and localization dataset, which is used to systematically evaluate and specifically improve the compositional generalization ability of existing models. Building upon this benchmark, we propose a simultaneous detection and localization network to simultaneously output the target location and the optimal grasping pose. Consequently, service robots equipped with the proposed network can perform object manipulation in real-world dynamic environments according to human instructions, providing theoretical foundations and technical support for embodied intelligence.  Methods  The structure of the proposed Simultaneous Grasp and Localization Network (SGL-Net) is illustrated (Fig. 1). Firstly, a cross-modal global context modulation module (CGCMM) is proposed. The module utilizes instruction priors to perform adaptive viewpoint correction and background suppression on visual channels during the early stages of feature extraction. Secondly, a word-pixel cross-modal alignment module (WPCAM) is designed to achieve fine-grained semantic decoupling via a flattened attention mechanism, further enhancing the model's comprehension in dynamic scenes. Finally, a unified decoder is employed to simultaneously output the target location and the optimal grasping pose.  Results and Discussions  Extensive quantitative and qualitative experiments are conducted on a restructured dual-view dataset, including bottom view and top view, and a real-world physical robotic platform. Comparative results demonstrate that SGL-Net significantly outperforms mainstream CNN-based and CLIP-based baseline methods in both grasp detection and object localization metrics (Table 2 and Table 3). Ablation studies explicitly validate the critical contributions of two designed modules in achieving fine-grained semantic alignment and scene decoupling (Table 4 and Table 5). Furthermore, qualitative visualizations (Fig. 4, Fig. 5, and Fig. 6) and real-world physical experiments (Fig. 7) indicate that SGL-Net exhibits strong viability for direct deployment in real-world physical environments. Overall, the network exhibits superior robust generalization capabilities and reliable practical deployment potential when confronting complex physical environments.  Conclusions  To address the challenge of insufficient cross-view and cross-domain generalization capability in robotic grasping operations, this paper constructs a corresponding validation dataset and proposes a simultaneous grasp detection and object localization network. By integrating a cross-modal global context modulation module and a word-pixel cross-modal alignment module, the proposed network achieves accurate localization of instruction-specified objects and prediction of optimal grasp poses. Experimental results from diverse testing datasets and real-world grasping experiments demonstrate the superior performance of the proposed network, highlighting its potential for deployment in practical scenarios. To further enhance robotic grasping capabilities in zero-shot scenarios, future work will focus on integrating Large Multimodal Models and realizing network adaptation through fine-tuning techniques.
  • loading
  • [1]
    张春云, 孟昕曈, 陶陶, 等. 面向机器人螺栓装配的视觉感知与力控协同方法[J]. 电子与信息学报, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.

    ZHANG Chunyun, MENG Xintong, TAO Tao, et al. Vision-guided and force-controlled method for robotic screw assembly[J]. Journal of Electronics & Information Technology, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.
    [2]
    余浩扬, 李艳生, 肖凌励, 等. 面向动态环境的巡检机器人轻量级语义视觉SLAM框架[J]. 电子与信息学报, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.

    YU Haoyang, LI Yansheng, XIAO Lingli, et al. A lightweight semantic visual simultaneous localization and mapping framework for inspection robots in dynamic environments[J]. Journal of Electronics & Information Technology, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.
    [3]
    LI Renjie, DONG Wei, SUN Jiarui, et al. A fast integrated gait, footstep, and motion planning framework for wheeled-legged robots[J]. Journal of Bionic Engineering, 2026, 23(2): 607–621. doi: 10.1007/s42235-025-00830-5.
    [4]
    LIU Zhitai, ZHONG Hanhai, LIN Xiaotian, et al. Integrated motion control framework of wheel-legged biped robot on rugged terrain[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(6): 7383–7394. doi: 10.1109/TMECH.2025.3627438.
    [5]
    崔永成, 田国会, 周昭旭, 等. 智能空间下面向动作序列生成的服务机器人指令解析方法[J]. 机器人, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074.

    CUI Yongcheng, TIAN Guohui, ZHOU Zhaoxu, et al. A service robot instruction parsing method for action sequence generation in intelligent space[J]. Robot, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074.
    [6]
    LENZ I, LEE H, and SAXENA A. Deep learning for detecting robotic grasps[J]. The International Journal of Robotics Research, 2015, 34(4/5): 705–724. doi: 10.1177/0278364914549607.
    [7]
    KUMRA S and KANAN C. Robotic grasp detection using deep convolutional neural networks[C]. 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vancouver, Canada, 2017: 769–776. doi: 10.1109/IROS.2017.8202237.
    [8]
    LAILI Yuanjun, CHEN Zelin, REN Lei, et al. Custom grasping: A region-based robotic grasping detection method in industrial cyber-physical systems[J]. IEEE Transactions on Automation Science and Engineering, 2023, 20(1): 88–100. doi: 10.1109/TASE.2021.3139610.
    [9]
    CHU F J, XU Ruinian, and VELA P A. Real-world multiobject, multigrasp detection[J]. IEEE Robotics and Automation Letters, 2018, 3(4): 3355–3362. doi: 10.1109/LRA.2018.2852777.
    [10]
    BOUSSELHAM W, PETERSEN F, FERRARI V, et al. Grounding everything: Emerging localization properties in vision-language transformers[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 3828–3837. doi: 10.1109/CVPR52733.2024.00367.
    [11]
    YIN Heng, REN Yuqiang, YAN Ke, et al. ROD-MLLM: Towards more reliable object detection in multimodal large language models[C]. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, USA, 2025: 14358–14368. doi: 10.1109/CVPR52734.2025.01339.
    [12]
    ZHOU Yi, SU Hang, WANG Tian, et al. Onet: Twin U-Net architecture for unsupervised binary semantic segmentation in radar and remote sensing images[J]. IEEE Transactions on Image Processing, 2025, 34: 2161–2172. doi: 10.1109/TIP.2025.3530816.
    [13]
    WANG Hao, HU Keyan, GUO Xin, et al. A gift from the integration of discriminative and diffusion-based generative learning: Boundary refinement remote sensing semantic segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5892–5909. doi: 10.1109/TPAMI.2026.3654243.
    [14]
    SHRIDHAR M, MANUELLI L, and FOX D. CLIPort: What and where pathways for robotic manipulation[C]. Proceedings of the 5th Conference on Robot Learning, London, UK, 2021: 894–906.
    [15]
    LIU Jin, XIE Jialong, and XIAO Leibing. Hierarchical multi-modal fusion for language-conditioned robotic grasping detection in clutter[J]. IEEE Robotics and Automation Letters, 2024, 9(10): 8762–8769. doi: 10.1109/LRA.2024.3440833.
    [16]
    VAN VO T, VU M N, HUANG Baoru, et al. Language-driven grasp detection with mask-guided attention[C]. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, Abu Dhabi, UAE, 2024: 7492–7498. doi: 10.1109/IROS58592.2024.10802256.
    [17]
    XIE Jialong, LIU Jin, ZHU Zhenwei, et al. Infusing multisource heterogeneous knowledge for language-conditioned segmentation and grasping[J]. IEEE Transactions on Instrumentation and Measurement, 2024, 73: 5029611. doi: 10.1109/TIM.2024.3446625.
    [18]
    XIE Jialong, ZHOU Fengyu, LIU Jin, et al. Semi-supervised language-conditioned grasping with curriculum-scheduled augmentation and geometric consistency[J]. IEEE Robotics and Automation Letters, 2025, 10(4): 4021–4028. doi: 10.1109/LRA.2025.3547619.
    [19]
    张梅, 金叶, 朱金辉, 等. 特征级语义感知引导的多模态图像融合算法[J]. 电子与信息学报, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.

    ZHANG Mei, JIN Ye, ZHU Jinhui, et al. FSG: Feature-level semantic-aware guidance for multi-modal image fusion algorithm[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.
    [20]
    LI Jindong, LI Yongguang, FU Yali, et al. CLIP-powered domain generalization and domain adaptation: A comprehensive survey[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5405–5424. doi: 10.1109/TPAMI.2026.3651700.
    [21]
    ZHOU Zhenning, ZHU Xiaoxiao, and CAO Qixin. AAGDN: Attention-augmented grasp detection network based on coordinate attention and effective feature fusion method[J]. IEEE Robotics and Automation Letters, 2023, 8(6): 3462–3469. doi: 10.1109/LRA.2023.3268596.
    [22]
    WANG Dexin, LIU Chunsheng, CHANG Faliang, et al. High-performance pixel-level grasp detection based on adaptive grasping and grasp-aware network[J]. IEEE Transactions on Industrial Electronics, 2021, 69(11): 11611–11621. doi: 10.1109/TIE.2021.3120474.
    [23]
    REZATOFIGHI H, TSOI N, GWAK J, et al. Generalized intersection over union: A metric and a loss for bounding box regression[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 658–666. doi: 10.1109/CVPR.2019.00075.
    [24]
    KUMRA S, JOSHI S, and SAHIN F. Antipodal robotic grasping using generative residual convolutional neural network[C]. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems, Las Vegas, USA, 2020: 9626–9633. doi: 10.1109/IROS45743.2020.9340777.
    [25]
    MORRISON D, CORKE P, and LEITNER J. Learning robust, real-time, reactive robotic grasping[J]. The International Journal of Robotics Research, 2020, 39(2/3): 183–201. doi: 10.1177/0278364919859066.
    [26]
    XIE Jialong, LIU Jin, HUANG Saike, et al. Listen, perceive, grasp: CLIP-driven attribute-aware network for language-conditioned visual segmentation and grasping[J]. IEEE Transactions on Automation Science and Engineering, 2025, 22: 9729–9740. doi: 10.1109/TASE.2024.3510777.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(7)  / Tables(6)

    Article Metrics

    Article views (32) PDF downloads(1) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return