logoYuvraj Gupta
HomeAboutProjectsSkillsContact
Resume
logoYuvraj Gupta
Back to Projects

JobSkill – Self-RAG Career Advisor

Research with Prof. Tracy Chen at SFSU, written up as a paper in May 2026. Generic LLMs hallucinate course codes and job titles; job boards optimise for keyword recall, not for the gap between what a student knows and what a posting demands. JobSkill reasons jointly over 603 live entry-level job postings and 99 university course descriptions: hybrid BM25 plus vector retrieval on Weaviate with Reciprocal Rank Fusion, an LLM planner that sets a per-turn retrieval budget, a LangGraph workflow that builds three evidence views in parallel, and a deterministic Self-RAG critic that scores each candidate answer for relevance, support, and utility against its citations before one is shown. Across a 30-question gold suite over 7 student personas the pipeline scored 7.93/10 with zero failures.

Python, LangGraph, LangSmithWeaviate hybrid BM25 + dense retrieval, HYBRID_ALPHA tuningOllama: nomic-embed-text, Llama 3.2, DeepSeek-R1OpenRouter and Fireworks AI for the generator bake-off
View CodeView Documentation

Note: Documentation is available; a live demo is not hosted for this project.

JobSkill – Self-RAG Career Advisor Screenshot

Project Features

🧭

Planner-Driven Hybrid Retrieval

A small model classifies query intent and allocates the retrieval budget. Long postings are indexed as 1,400-character overlapping chunks, max-pooled per record, then fused with skill-overlap signals through Reciprocal Rank Fusion.

🧪

Self-RAG Critique

Three candidate answers are generated from three evidence views: job-specific, job-cluster, and course-path. A critic scores relevance, support, and utility, runs deterministic citation checks for invented entities, and triggers a second retrieval pass when support is weak.

⚖️

Ten Generators, One Fixed Judge

With planner and critic held constant, ten generator backbones from 3B to 120B parameters were compared on the gold suite. Qwen 2.5 7B won at 8.53/10, $0.000245 per query, and 16.8 s latency. Bigger models did not buy more quality.

🎯

Persona-Grounded Gold Suite

Thirty realistic questions across seven student profiles, from under-prepared to over-qualified, each with rubric-based expected behaviour. Reported alongside RAGAS faithfulness, response relevancy, and context precision.

Project Info

Duration

12/2025 – 05/2026

Role

AI Engineer & Author

Team Size

Advised by Prof. Tracy Chen

Status

Paper Completed · Code Released

Technologies Used

  • • Python, LangGraph, LangSmith
  • • Weaviate hybrid BM25 + dense retrieval, HYBRID_ALPHA tuning
  • • Ollama: nomic-embed-text, Llama 3.2, DeepSeek-R1
  • • OpenRouter and Fireworks AI for the generator bake-off
  • • Supabase Postgres, ingestion pipeline, gold-question harness
  • • RAGAS evaluation with a neutral judge model
  • • React 18 + TypeScript + shadcn/ui client, Streamlit advisor UI

Read the Write-up

JobSkill: A Self-RAG Career Advisor for Computer Science StudentsPaper · PDF · 6 pages · May 12, 2026Hybrid retrieval, LangGraph orchestration, the persona gold benchmark, and the ten-model generator comparison, with the engineering trade-offs that mattered in practice.Model comparison reportMarkdown · 10 models × 10 questionsEvery generator's actual answer to every gold question, scored, with per-question observations on where each model went right or wrong.Production RAG system slidesPowerPoint deckThe architecture walkthrough presented for the course, from ingestion to critique.

Project Gallery

Job Fit, Grounded
Job Fit, Grounded
An Android Developer posting with its required skills tagged as covered or missing against the student's profile, and the inline advisor ready beside it.
Matched Postings
Matched Postings
Live job cards with skill chips coloured by what the student already has, ranked by the hybrid retriever.
Advisor Chat, With Receipts
Advisor Chat, With Receipts
A grounded answer that names the postings and courses it drew on, then a table of recommended courses against the gap skills each one closes.
Home
Home
Find jobs that fit your skills, learn what is missing. The entry point for a student with a transcript and a profile.
Student Profile
Student Profile
Completed courses, extracted skills, and the profile the planner and critic reason against on every turn.
Quality vs Cost Across Ten Generators
Quality vs Cost Across Ten Generators
Qwen 2.5 7B leads at a fraction of the cost of the 120B and DeepSeek models. The retrieval substrate, not generator size, sets the ceiling.
Generator Leaderboard
Generator Leaderboard
Ten models on the same scaffold: score, speed, and cost per query, with the best value and the fastest called out.
Three Evidence Frames
Three Evidence Frames
Specific job, job cluster, and course path each produce a candidate answer; the critic picks the one the evidence supports.
Hybrid Search With RRF
Hybrid Search With RRF
Query, embed, BM25 plus dense retrieval, pooling per record, and Reciprocal Rank Fusion into a top-k list.
Anatomy of the Gold Set
Anatomy of the Gold Set
Thirty questions across seven personas and nine categories, weighted toward the hard cases.

Interested in This Project?

I'm always excited to discuss my projects and share insights about the development process. Feel free to reach out if you'd like to know more!

Get In TouchView More Projects

Yuvraj Gupta

AI Engineer building multi-agent systems, production retrieval, and research that ships. San Francisco.

Quick Links

  • Home
  • About
  • Projects
  • Skills
  • Contact

Get In Touch

📍San Francisco, CA
📧yuvrajgupta1808@gmail.com

© 2026 Yuvraj Gupta. All rights reserved.