Tech
Visual Grounding for AI Agents: Set-of-Mark Prompting and Reliable UI Clicks
Ask a frontier vision model for the pixel coordinates of a "Submit" button in a 1920x1080 screenshot and it will often miss by dozens of pixels. That miss is the difference between an agent that completes a checkout flow and one that clicks empty whitespace and loops until it hits a timeout. The fix is not a bigger model. It is a better interface between the model and the screen.
That interface has two parts. Screen painting draws numbered marks onto the screenshot before the model sees it....
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to