JobSkill – Self-RAG Career Advisor
Research with Prof. Tracy Chen at SFSU, written up as a paper in May 2026. Generic LLMs hallucinate course codes and job titles; job boards optimise for keyword recall, not for the gap between what a student knows and what a posting demands. JobSkill reasons jointly over 603 live entry-level job postings and 99 university course descriptions: hybrid BM25 plus vector retrieval on Weaviate with Reciprocal Rank Fusion, an LLM planner that sets a per-turn retrieval budget, a LangGraph workflow that builds three evidence views in parallel, and a deterministic Self-RAG critic that scores each candidate answer for relevance, support, and utility against its citations before one is shown. Across a 30-question gold suite over 7 student personas the pipeline scored 7.93/10 with zero failures.
Note: Documentation is available; a live demo is not hosted for this project.

Project Features
Planner-Driven Hybrid Retrieval
A small model classifies query intent and allocates the retrieval budget. Long postings are indexed as 1,400-character overlapping chunks, max-pooled per record, then fused with skill-overlap signals through Reciprocal Rank Fusion.
Self-RAG Critique
Three candidate answers are generated from three evidence views: job-specific, job-cluster, and course-path. A critic scores relevance, support, and utility, runs deterministic citation checks for invented entities, and triggers a second retrieval pass when support is weak.
Ten Generators, One Fixed Judge
With planner and critic held constant, ten generator backbones from 3B to 120B parameters were compared on the gold suite. Qwen 2.5 7B won at 8.53/10, $0.000245 per query, and 16.8 s latency. Bigger models did not buy more quality.
Persona-Grounded Gold Suite
Thirty realistic questions across seven student profiles, from under-prepared to over-qualified, each with rubric-based expected behaviour. Reported alongside RAGAS faithfulness, response relevancy, and context precision.
Project Info
12/2025 – 05/2026
AI Engineer & Author
Advised by Prof. Tracy Chen
Paper Completed · Code Released
Technologies Used
- • Python, LangGraph, LangSmith
- • Weaviate hybrid BM25 + dense retrieval, HYBRID_ALPHA tuning
- • Ollama: nomic-embed-text, Llama 3.2, DeepSeek-R1
- • OpenRouter and Fireworks AI for the generator bake-off
- • Supabase Postgres, ingestion pipeline, gold-question harness
- • RAGAS evaluation with a neutral judge model
- • React 18 + TypeScript + shadcn/ui client, Streamlit advisor UI
Read the Write-up
Project Gallery










Interested in This Project?
I'm always excited to discuss my projects and share insights about the development process. Feel free to reach out if you'd like to know more!