logoYuvraj Gupta
HomeAboutProjectsSkillsContact
Resume
logoYuvraj Gupta
Back to Projects

Rodex – Multi-Agent Code Review Platform

Static analysers understand syntax and get ignored. Language models understand intent and then describe their own patches as correct whether or not they are. Rodex is built on the position that a model's claim about its own work is worth nothing: four agents review Python code inside a Cursor-style IDE, stream every thought and tool call to the browser, and a patch is recorded only after it passes two mechanical gates the model cannot argue with. On a ten-file benchmark with twenty-three planted defects it reaches 0.913 recall and 0.712 F1, and all 35 proposed patches passed verification. The agent framework is written from scratch in Python, with no LangChain or LangGraph.

Custom agent framework: base agent, event bus, sandbox manager, shared agent driveCoordinator, Security, Bug Detection, and Fix/Patch agentspy_compile gate + AST pattern-removal gate with automatic rollbackPer-agent telemetry: tokens, reasoning tokens, latency, USD cost
View CodeView Documentation

Note: Documentation is available; a live demo is not hosted for this project.

Rodex – Multi-Agent Code Review Platform Screenshot

Project Features

✅

Two-Gate Patch Verification

A fix must compile in the sandbox, then pass an AST check proving the flagged pattern is gone. A patch that changes nothing compiles fine, so compiling alone is not proof. Fail either gate and the patch is rolled back with the reason returned to the coordinator.

🤖

Four Agents, One Bounded Loop

A Coordinator dispatches Security, Bug Detection, and Fix agents concurrently, judges findings, retries failures, and stops after at most 14 iterations. Every step streams into the IDE's agent panel as it happens.

📏

Measured, Not Claimed

Published benchmark: 0.913 recall, 0.583 precision, 0.712 F1, 35 of 35 patches verified, 42.2 s mean latency per file. A finding counts only if the category matches and the line is within two of the defect. Reproducible from a clean checkout.

🔌

Grounded Through MCP

Specialists read repository state from MCP servers (filesystem, git, fetch) before analysing, under per-agent allowlists: the security agent can look up advisories while the fix agent stays offline. Every call is logged with arguments, duration, and output.

Project Info

Duration

04/2026

Role

Sole Developer

Team Size

Solo Project

Status

Production-deployed · Open Source (MIT)

Technologies Used

  • • Custom agent framework: base agent, event bus, sandbox manager, shared agent drive
  • • Coordinator, Security, Bug Detection, and Fix/Patch agents
  • • py_compile gate + AST pattern-removal gate with automatic rollback
  • • Per-agent telemetry: tokens, reasoning tokens, latency, USD cost
  • • Firebase email-link sign-in exchanged for a signed 14-day session cookie
  • • Vertex AI default with short-lived OAuth; provider swap is one env var
  • • Monaco editor, file tree, findings feed, in-editor patch markers
  • • Single Docker image, GitHub Actions CI on two Python versions

Read the Write-up

Rodex: A Multi-Agent Code Review PlatformPDF · 14 pages · April 2026The full write-up: why unverified AI patches are a new source of bugs, the infrastructure and auth model, the agent loop, the two verification gates, and the measured benchmark.Benchmark results, per fileJSON · metrics-gemini-2.5-proRaw precision, recall, F1, and verified-patch counts for each of the ten seeded files, produced by the evaluate.py harness in the repo.

Project Gallery

A Production Review
A Production Review
Seven findings on a 19-line file: one dismissed as a false positive, five verified and patched, one deliberately left alone with the reason stated. Cost of the review: $0.0367.
The Agent Loop
The Agent Loop
Signed-cookie session on every route, a coordinator that dispatches specialists in one turn, and patches that only land after both gates pass.
Two Gates
Two Gates
Compile, then prove the pattern is gone. Rejected patches roll back and the coordinator hears why.
Security Agent, Live
Security Agent, Live
SQL injection and unsafe deserialization flagged as the agent reads the file, streamed over SSE.
Fix Agent Patching
Fix Agent Patching
Parameterised query and a narrowed exception written into the sandbox copy and submitted for verification.
Verdict and Telemetry
Verdict and Telemetry
Plain-language summary of what was fixed and what was left, with tokens, latency, and cost per agent.
Entry Point
Entry Point
Signed-in users upload a folder or file, or paste code, then start a review.

Interested in This Project?

I'm always excited to discuss my projects and share insights about the development process. Feel free to reach out if you'd like to know more!

Get In TouchView More Projects

Yuvraj Gupta

AI Engineer building multi-agent systems, production retrieval, and research that ships. San Francisco.

Quick Links

  • Home
  • About
  • Projects
  • Skills
  • Contact

Get In Touch

📍San Francisco, CA
📧yuvrajgupta1808@gmail.com

© 2026 Yuvraj Gupta. All rights reserved.