logoYuvraj Gupta
HomeAboutProjectsSkillsContact
Resume
logoYuvraj Gupta
Back to Projects

Live Caster – Real-Time Screen Narrator

Born at the Google DeepMind × Cerebral Valley Gemini Hackathon as a live esports commentator, where it placed 3rd and won $20,000 in Gemini credits. Share any window and a single Gemini Live session watches it and speaks: native audio straight from the model, no separate TTS step. It has since been rebuilt as an accessibility tool for blind and low-vision users and taken to a public production deployment on Cloud Run. Screen readers announce the element under the cursor; Live Caster describes the screen as a whole and says which part matters now. Talk over it at any time and it stops mid-word, answers, then resumes. I owned the capture and streaming pipeline at the hackathon and built the production version.

gemini-live-2.5-flash-native-audio on Vertex AIFastAPI bridge: JPEG frames + 16 kHz PCM in, 24 kHz PCM + transcripts outWebSocket bidirectional streamingClient-side frame diffing and 1024px downscaling
View CodeLive DemoView Documentation
Live Caster – Real-Time Screen Narrator Screenshot

Project Features

⚡

Measured Latency

0.58–1.09 seconds to first spoken word on consecutive frames. The UI shows time-to-first-word for every line, so real-time is a number, not an adjective.

🗣️

True Barge-In

Speak and the narrator stops mid-sentence via Gemini's voice-activity detection, answers, then picks the story back up. Both sides appear as live captions.

🧊

Frame Change Detection

Frames are captured every two seconds but diffed on a downscaled grayscale grid. An unchanged screen never hits the API, so there is no cost and no narrating a static board.

⌨️

Fully Keyboard-Operable

Space to start and stop, M to mute, D to describe the screen now, R to repeat, plus quiet mode, slow speech, session history, and self-healing reconnects.

Project Info

Duration

12/2025 – Present

Role

Streaming & Latency Pipeline

Team Size

Team of 3 at the hackathon · solo since

Status

3rd Place · Live in Production

Technologies Used

  • • gemini-live-2.5-flash-native-audio on Vertex AI
  • • FastAPI bridge: JPEG frames + 16 kHz PCM in, 24 kHz PCM + transcripts out
  • • WebSocket bidirectional streaming
  • • Client-side frame diffing and 1024px downscaling
  • • Google Search grounding tool
  • • Session resumption and idle auto-end
  • • Demo mode against a synthetic animated screen

Read the Write-up

Live Caster: A Real-Time Screen Narrator for Blind and Low-Vision UsersPDF · 9 pages · December 2025Why a capture-caption-speak pipeline cannot be interrupted, the single-session design that can, the Cloud Run deployment that scales to zero, quiet mode, and the measured latency.Try it livelivecaster.yuvrajgupta.comThe production deployment. Sign in, press Start, share a window, and talk over the narrator whenever you like.

Project Gallery

Narrating a Shared Window
Narrating a Shared Window
The shared browser on the left, the live transcript and controls on the right.
Answering a Spoken Question
Answering a Spoken Question
The user interrupts, the narrator stops, answers about what is on screen, then resumes.
Searching Mid-Session
Searching Mid-Session
When the screen alone is not enough, the model calls Google Search and the result lands inline.
Time-to-First-Word
Time-to-First-Word
Every line shows its own latency and the counter of frames skipped as unchanged.

Interested in This Project?

I'm always excited to discuss my projects and share insights about the development process. Feel free to reach out if you'd like to know more!

Get In TouchView More Projects

Yuvraj Gupta

AI Engineer building multi-agent systems, production retrieval, and research that ships. San Francisco.

Quick Links

  • Home
  • About
  • Projects
  • Skills
  • Contact

Get In Touch

📍San Francisco, CA
📧yuvrajgupta1808@gmail.com

© 2026 Yuvraj Gupta. All rights reserved.