“Task-Based” or “Semantic” robot grasping — where the grasp criteria are based on how the object will be used —has a long history. Humans grasp objects intuitively based on affordances. We believe this is the first paper that applies LLM + CLIP to infer affordances from words.
LLM and CLIP for Robot Grasping Based on Object Affordances
By
–