VLM-integrated 3D perception model for robust robotic grasping adapted to deformable sacks with arbitrary shapes

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

This study presents a vision-language and 3D perception-integrated system for robust grasping of non-standard logistics parcels, including untrained boxes and deformable sacks. To address the limitations of conventional recognition models in handling unseen or irregularly shaped objects, we propose a zero-shot object recognition framework based on a Vision-Language Model (VLM), combined with confidence threshold optimization and Non-Maximum Suppression (NMS) to improve detection reliability. The system further incorporates 3D point cloud-based post-processing to extract precise grasp points adapted to both rigid and deformable object geometries. Proposed approach achieves a mean Average Precision (mAP) of 80.1 % for detecting untrained boxes and burlap sacks and demonstrates a grasp success rate of 99.0 % for boxes and 73.0 % for deformable sacks in realworld unloading scenarios using collaborative robots. Unlike prior methods that require extensive retraining for new object types, proposed system enables generalizable, real-time grasping without additional dataset preparation. The hybrid integration of rule-based and learning-based strategies in 3D space contributes to sTable and adapTable grasp point selection across variable object types. This work substantiates the feasibility of zero-shot recognition and grasping of non-standard logistics items, offering a practical solution for automated parcel unloading in dynamic industrial environments.

키워드

Unregularized objectZero-shotPoint-cloudGrasp pointVision language model
제목
VLM-integrated 3D perception model for robust robotic grasping adapted to deformable sacks with arbitrary shapes
저자
Yoon, JonghunKim, Seongje
DOI
10.1016/j.robot.2026.105372
발행일
2026-05
유형
정기학술지(Article(Perspective Article포함))
저널명
Robotics and Autonomous Systems
199
페이지
1 ~ 20