logoYuvraj Gupta
HomeAboutProjectsSkillsContact
Resume
logoYuvraj Gupta
Back to Projects

AVT-CA – Audio-Visual Engagement Detection

Research at SFSU advised by Prof. Sanchita Ghose: can adding audio to video help a model tell how engaged a student is? The work builds on the published Audio-Video Token Cross-Attention implementation and produced the first audio-visual fusion result on the EngageNet benchmark (11,206 clips, four ordinal engagement levels, 2,257 held-out test clips), where every published baseline is video-only. The best model reaches 66.36% top-1, 90.96% adjacent accuracy, and 52.03 macro-F1, ahead of the LSTM, CNN-LSTM, MARLIN-Transformer, and TCN baselines and 1.25 points below the best published result. Two findings mattered more than the headline: replacing argmax with ordinal calibration decoding gained +2.03 points on identical weights, and three preprocessing defects, once found and fixed, were worth 2 to 3.7 points on their own.

Python, PyTorch, CUDA DataParallel, AMPEngageNet (4-level ordinal), DAiSEE, RAVDESS loaders; CREMA-D and CMU-MOSEI preprocessedEfficientFace backbone, channel and spatial modulators, cross-attention blocksOrdinal loss, class weighting, label smoothing, synced audio-video temporal crops
View CodeView Documentation

Note: Documentation is available; a live demo is not hosted for this project.

AVT-CA – Audio-Visual Engagement Detection Screenshot

Project Features

🎯

Ordinal Calibration Decoding

Engagement levels are ordered, so argmax throws information away. Decoding the expected level with thresholds fit on validation lifted test top-1 from 64.10 to 66.13 with no retraining, and reproduced on every checkpoint tested. Adjacent accuracy reached 91%.

🔍

Three Preprocessing Defects, Found and Fixed

Audio was truncated to 3.6 s of each 10 s clip. Test and validation were extracted at 15 frames while training used 50. A fifth of the training split was short too. Fixing them moved the headline from 62.68 to 66.36 and reversed an earlier conclusion that audio did not help.

🔀

Does Audio Help? A Narrow, Honest Yes

Audio alone scores exactly the majority-class rate, so it carries no standalone signal. Fused with video it adds about one point, and it does so on all four decoders, which is what makes the gain credible. Single seed so far; the repeat is queued before it goes in a paper.

🧪

Evidence Discipline

Sixty run directories and 105 evaluations consolidated into evidence tables; baselines verified against the papers' PDFs rather than search results; a stopping criterion pre-committed before the retraining runs and triggered as designed at epoch 10. Phase one on RAVDESS emotion recognition reached 71.25% before the pivot.

Project Info

Duration

01/2026 – Present

Role

Lead Researcher

Team Size

Advised by Prof. Sanchita Ghose

Status

Ongoing · retraining on corrected splits

Technologies Used

  • • Python, PyTorch, CUDA DataParallel, AMP
  • • EngageNet (4-level ordinal), DAiSEE, RAVDESS loaders; CREMA-D and CMU-MOSEI preprocessed
  • • EfficientFace backbone, channel and spatial modulators, cross-attention blocks
  • • Ordinal loss, class weighting, label smoothing, synced audio-video temporal crops
  • • Calibration module: argmax, logit bias, expected thresholds, refined expected
  • • Modality ablations, encoder-free audio probes, checkpoint ensembling
  • • Streamlit inference app, unit and smoke tests, GitHub Actions CI

Read the Write-up

Progress report: verified EngageNet resultsMarkdown · development branchThe corrected result set, the decoding ladder, the modality ablation, the preprocessing defects and their fixes, and what is still open. Written to be cited, with caveats attached to every number.Research plan and experiment trackerMarkdown · 1,000+ linesLiterature verified against PDFs, the published baseline table, every experiment from A0 to G02 with its question and outcome, and the pre-committed decision matrix.Phase one: RAVDESS emotion recognitionMarkdown · training run historyThe three bug fixes in the inherited cross-attention, all training runs with their configurations, and the per-emotion breakdown that took test accuracy from 66.67% to 71.25%.TokenFusion AVT-CA fine-tuning reportPDF · 7 pagesThe shared-transformer variant: exact training command, forward-pass algorithm, epoch-by-epoch validation, classification report, and confusion matrix.

Project Gallery

Against the Published Field
Against the Published Field
66.36% top-1 sits ahead of four of the six published EngageNet baselines, all of them video-only, and 1.25 below the best.
The Decoding Ladder
The Decoding Ladder
Same weights, four decoders. Refined expected-value decoding adds two points, and fusion beats video-only at every rung.
Before and After the Preprocessing Fix
Before and After the Preprocessing Fix
Every model gained once the test set carried the same 50 frames as training. Audio-only stayed at the majority-class rate, which is the point.
Phase One: RAVDESS Accuracy Across Runs
Phase One: RAVDESS Accuracy Across Runs
Eight-class emotion recognition, 480 held-out clips. Mel spectrograms and eight heads took the baseline from 66.67% to 71.25%.
Phase One: Per-Class F1
Phase One: Per-Class F1
The gains landed on the hard classes. Fearful, Surprised, Angry, and Sad all rose by more than 0.1.
Proposed Pipeline
Proposed Pipeline
From the research proposal: attention-refined video features and audio features pass through self-attention blocks, then a multi-head cross-attention mechanism, before pooling and classification.

Interested in This Project?

I'm always excited to discuss my projects and share insights about the development process. Feel free to reach out if you'd like to know more!

Get In TouchView More Projects

Yuvraj Gupta

AI Engineer building multi-agent systems, production retrieval, and research that ships. San Francisco.

Quick Links

  • Home
  • About
  • Projects
  • Skills
  • Contact

Get In Touch

📍San Francisco, CA
📧yuvrajgupta1808@gmail.com

© 2026 Yuvraj Gupta. All rights reserved.