A vision-enabled browser agent that looks at annotated screenshots and decides where to click, scroll, and type until the task is done.