REAL Browser Agent – Autonomous Web Tasks
Built with Peixi Xie and Dhruv Patel for the AGI Inc × OpenAI × Lovable REAL Agent Challenge, which it won. An autonomous browser assistant that handles everyday digital chores through vision-language reasoning: applying to jobs, filling forms, sending email, setting calendar events, booking hotels, and multi-page navigation, scored on REAL-Bench. We tested two designs, prompt engineering with reflection loops and an orchestrator coordinating multiple models, and shipped the hybrid, which proved the most stable under time and credit limits.

Project Features
Vision-Language Control
The agent reads screenshots, locates controls, and decides the next action without site-specific scripts.
Reflection Loops
After each step the agent critiques its own progress and recovers from misclicks, which mattered more than raw model size.
Orchestrated Specialists
A coordinator routes sub-tasks to models suited to them, and the hybrid with reflection beat either approach alone.
Benchmarked on REAL-Bench
Sixteen benchmark tasks and an evaluator in the repo, so improvements are measured rather than eyeballed.
Project Info
11/2025
Agent Loop & Evaluation
Team of 3
Winner · The REAL Agent Challenge
Technologies Used
- • Python
- • Qwen3 via OpenRouter
- • REAL-Bench harness and evaluator
- • Reflection loop and orchestrator patterns
- • Playwright-driven browser runtime
Project Gallery

Interested in This Project?
I'm always excited to discuss my projects and share insights about the development process. Feel free to reach out if you'd like to know more!