DOI: 10.3390/robotics15080154 ISSN: 2218-6581

ActivAsk: Free-Energy-Guided Clarification for Robotic Grasping Under Ambiguous Instructions

Haoandong Yang, Gabriel W. Haddon-Hill, Teresa Zielinska, Shingo Murata

Service robots often receive natural language instructions in changing workspaces where multiple visible objects may match one description. Relying on detector confidence, random selection, or direct vision–language model (VLM) prediction can lead to a wrong action. This paper presents ActivAsk, a zero-shot framework for resolving referential ambiguity before robotic grasping. ActivAsk constructs open-vocabulary candidates from red-green-blue-depth (RGB-D) input, asks candidate-grounded yes/no questions when needed, updates the candidate state from the user’s answer, and grasps after target resolution. It selects among VLM-proposed candidate partitions using an expected free energy (EFE) criterion motivated by active inference; with neutral response preferences, this reduces to information gain over candidate partitions. Offline experiments showed that interactive clarification improved target accuracy from about 53–54% for noninteractive baselines to about 90–92%. ActivAsk matched the best interactive accuracy (92.13%) while asking 15.47–19.71% fewer questions on asked trials and 21.43–23.88% fewer for ambiguous instructions. In online real robot experiments, ActivAsk achieved 92.98% target selection accuracy and 87.72% full correct object grasp success; unresolved or wrong targets were not physically executed after operator-controlled verification and were counted as task failures.

More from our Archive