AVT-CA – Audio-Visual Engagement Detection
Research at SFSU advised by Prof. Sanchita Ghose: can adding audio to video help a model tell how engaged a student is? The work builds on the published Audio-Video Token Cross-Attention implementation and produced the first audio-visual fusion result on the EngageNet benchmark (11,206 clips, four ordinal engagement levels, 2,257 held-out test clips), where every published baseline is video-only. The best model reaches 66.36% top-1, 90.96% adjacent accuracy, and 52.03 macro-F1, ahead of the LSTM, CNN-LSTM, MARLIN-Transformer, and TCN baselines and 1.25 points below the best published result. Two findings mattered more than the headline: replacing argmax with ordinal calibration decoding gained +2.03 points on identical weights, and three preprocessing defects, once found and fixed, were worth 2 to 3.7 points on their own.
Note: Documentation is available; a live demo is not hosted for this project.

Project Features
Ordinal Calibration Decoding
Engagement levels are ordered, so argmax throws information away. Decoding the expected level with thresholds fit on validation lifted test top-1 from 64.10 to 66.13 with no retraining, and reproduced on every checkpoint tested. Adjacent accuracy reached 91%.
Three Preprocessing Defects, Found and Fixed
Audio was truncated to 3.6 s of each 10 s clip. Test and validation were extracted at 15 frames while training used 50. A fifth of the training split was short too. Fixing them moved the headline from 62.68 to 66.36 and reversed an earlier conclusion that audio did not help.
Does Audio Help? A Narrow, Honest Yes
Audio alone scores exactly the majority-class rate, so it carries no standalone signal. Fused with video it adds about one point, and it does so on all four decoders, which is what makes the gain credible. Single seed so far; the repeat is queued before it goes in a paper.
Evidence Discipline
Sixty run directories and 105 evaluations consolidated into evidence tables; baselines verified against the papers' PDFs rather than search results; a stopping criterion pre-committed before the retraining runs and triggered as designed at epoch 10. Phase one on RAVDESS emotion recognition reached 71.25% before the pivot.
Project Info
01/2026 – Present
Lead Researcher
Advised by Prof. Sanchita Ghose
Ongoing · retraining on corrected splits
Technologies Used
- • Python, PyTorch, CUDA DataParallel, AMP
- • EngageNet (4-level ordinal), DAiSEE, RAVDESS loaders; CREMA-D and CMU-MOSEI preprocessed
- • EfficientFace backbone, channel and spatial modulators, cross-attention blocks
- • Ordinal loss, class weighting, label smoothing, synced audio-video temporal crops
- • Calibration module: argmax, logit bias, expected thresholds, refined expected
- • Modality ablations, encoder-free audio probes, checkpoint ensembling
- • Streamlit inference app, unit and smoke tests, GitHub Actions CI
Read the Write-up
Project Gallery






Interested in This Project?
I'm always excited to discuss my projects and share insights about the development process. Feel free to reach out if you'd like to know more!