Multi-Step Obstruction Reasoning for Target-Oriented Grasp Sequence Generation in Cluttered Scenes
Huixuan Yang, Shiqiang Zhu, Yuhua Zheng, Wei SongRetrieving target objects in severe clutter requires multi-step reasoning to establish valid obstacle removal sequences. While recent Vision–Language Models (VLMs) have advanced instruction-driven clutter grasping, existing paradigms lack explicit construction of graph-constrained grasp sequences encompassing canonical trajectories and valid topological permutations to guide model fine-tuning and evaluation. In addition, comprehensive evaluation requires accounting for the full space of topologically valid clearing sequences while systematically disentangling high-level topological planning errors from low-level physical execution failures. To fulfill these requirements, we introduce a novel DAG-based obstruction reasoning framework coupled with an integrated diagnostic evaluation protocol. Specifically, we model scene-level physical dependencies as Directed Acyclic Graphs (DAGs), explicitly converting graph constraints into topologically feasible sequence permutations to drive VLM fine-tuning. For diagnostic evaluation, we establish a three-part offline protocol comprising Strict Exact Match (EM), Graph-Feasible Accuracy (GFA), and Multi-Reference Normalized Sequence Edit Distance (MR-NSED), paired with online simulation testing. Extensive experiments demonstrate that our topological fine-tuning significantly improves multi-path reasoning performance, outperforming strong foundation model baselines including GPT-4o, GPT-4o-mini, and Llama-3.2-90B-Vision-Instruct, while our diagnostic protocol provides a faithful mechanism to systematically isolate reasoning logic from manipulation mechanics in complex physical clutter.