Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jia, Wang, Yang, Qian, Wenhao, Hu, Jialong, Hu, Zhenzhen, Hong, Richang, Wang, Meng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912520777760768
author Li, Jia
Wang, Yang
Qian, Wenhao
Hu, Jialong
Hu, Zhenzhen
Hong, Richang
Wang, Meng
author_facet Li, Jia
Wang, Yang
Qian, Wenhao
Hu, Jialong
Hu, Zhenzhen
Hong, Richang
Wang, Meng
contents Interview performance assessment is essential for determining candidates' suitability for professional positions. To ensure holistic and fair evaluations, we propose a novel and comprehensive framework that explores ``365'' aspects of interview performance by integrating \textit{three} modalities (video, audio, and text), \textit{six} responses per candidate, and \textit{five} key evaluation dimensions. The framework employs modality-specific feature extractors to encode heterogeneous data streams and subsequently fused via a Shared Compression Multilayer Perceptron. This module compresses multimodal embeddings into a unified latent space, facilitating efficient feature interaction. To enhance prediction robustness, we incorporate a two-level ensemble learning strategy: (1) independent regression heads predict scores for each response, and (2) predictions are aggregated across responses using a mean-pooling mechanism to produce final scores for the five target dimensions. By listening to the unspoken, our approach captures both explicit and implicit cues from multimodal data, enabling comprehensive and unbiased assessments. Achieving a multi-dimensional average MSE of 0.1824, our framework secured first place in the AVI Challenge 2025, demonstrating its effectiveness and robustness in advancing automated and multimodal interview performance assessment. The full implementation is available at https://github.com/MSA-LMC/365Aspects.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22676
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
Li, Jia
Wang, Yang
Qian, Wenhao
Hu, Jialong
Hu, Zhenzhen
Hong, Richang
Wang, Meng
Computation and Language
Multimedia
Interview performance assessment is essential for determining candidates' suitability for professional positions. To ensure holistic and fair evaluations, we propose a novel and comprehensive framework that explores ``365'' aspects of interview performance by integrating \textit{three} modalities (video, audio, and text), \textit{six} responses per candidate, and \textit{five} key evaluation dimensions. The framework employs modality-specific feature extractors to encode heterogeneous data streams and subsequently fused via a Shared Compression Multilayer Perceptron. This module compresses multimodal embeddings into a unified latent space, facilitating efficient feature interaction. To enhance prediction robustness, we incorporate a two-level ensemble learning strategy: (1) independent regression heads predict scores for each response, and (2) predictions are aggregated across responses using a mean-pooling mechanism to produce final scores for the five target dimensions. By listening to the unspoken, our approach captures both explicit and implicit cues from multimodal data, enabling comprehensive and unbiased assessments. Achieving a multi-dimensional average MSE of 0.1824, our framework secured first place in the AVI Challenge 2025, demonstrating its effectiveness and robustness in advancing automated and multimodal interview performance assessment. The full implementation is available at https://github.com/MSA-LMC/365Aspects.
title Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
topic Computation and Language
Multimedia
url https://arxiv.org/abs/2507.22676