Advised by Prof. Sanchita Ghose, SFSUJan 2026 – present
66.36%
top-1 on EngageNet, the first audio-visual result on the benchmark
Audio-visual engagement detection
Does adding audio to a video-only model actually help it read student engagement?
Yes. The fused model beats four of six published video-only baselines. Along the way, fixing the inherited cross-attention lifted RAVDESS emotion accuracy from 66.67% to 71.25%, and swapping argmax for ordinal calibration gained +2.03 points at zero training cost.
PyTorch · Multimodal · EngageNet

