It’s tool use (PIL/cv2, BFS) debugged with multimodal perception of tool output within the CoT; it’s not 4o-style multimodal output. See the share link in the thread for more info on the tool calls.
Multimodal tool use with CoT and vision libraries
By
–