SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Tianyu, Huang, Jinfa, Ma, Yuexiao, Luo, Rongfang, Yang, Yan, Chen, Wang, Zeng, Yuhui, Fang, Ruize, Zou, Yixuan, Zheng, Xiawu, Luo, Jiebo, Ji, Rongrong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912971253350400
author Xie, Tianyu
Huang, Jinfa
Ma, Yuexiao
Luo, Rongfang
Yang, Yan
Chen, Wang
Zeng, Yuhui
Fang, Ruize
Zou, Yixuan
Zheng, Xiawu
Luo, Jiebo
Ji, Rongrong
author_facet Xie, Tianyu
Huang, Jinfa
Ma, Yuexiao
Luo, Rongfang
Yang, Yan
Chen, Wang
Zeng, Yuhui
Fang, Ruize
Zou, Yixuan
Zheng, Xiawu
Luo, Jiebo
Ji, Rongrong
contents Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a critical gap in assessing social interactivity, the fundamental capacity to navigate dynamic cues in natural dialogues. To this end, we propose SocialOmni, a comprehensive benchmark that operationalizes the evaluation of this conversational interactivity across three core dimensions: (i) speaker separation and identification (who is speaking), (ii) interruption timing control (when to interject), and (iii) natural interruption generation (how to phrase the interruption). SocialOmni features 2,000 perception samples and a quality-controlled diagnostic set of 209 interaction-generation instances with strict temporal and contextual constraints, complemented by controlled audio-visual inconsistency scenarios to test model robustness. We benchmarked 12 leading OLMs, which uncovers significant variance in their social-interaction capabilities across models. Furthermore, our analysis reveals a pronounced decoupling between a model's perceptual accuracy and its ability to generate contextually appropriate interruptions, indicating that understanding-centric metrics alone are insufficient to characterize conversational social competence. More encouragingly, these diagnostics from SocialOmni yield actionable signals for bridging the perception-interaction divide in future OLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16859
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
Xie, Tianyu
Huang, Jinfa
Ma, Yuexiao
Luo, Rongfang
Yang, Yan
Chen, Wang
Zeng, Yuhui
Fang, Ruize
Zou, Yixuan
Zheng, Xiawu
Luo, Jiebo
Ji, Rongrong
Artificial Intelligence
Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a critical gap in assessing social interactivity, the fundamental capacity to navigate dynamic cues in natural dialogues. To this end, we propose SocialOmni, a comprehensive benchmark that operationalizes the evaluation of this conversational interactivity across three core dimensions: (i) speaker separation and identification (who is speaking), (ii) interruption timing control (when to interject), and (iii) natural interruption generation (how to phrase the interruption). SocialOmni features 2,000 perception samples and a quality-controlled diagnostic set of 209 interaction-generation instances with strict temporal and contextual constraints, complemented by controlled audio-visual inconsistency scenarios to test model robustness. We benchmarked 12 leading OLMs, which uncovers significant variance in their social-interaction capabilities across models. Furthermore, our analysis reveals a pronounced decoupling between a model's perceptual accuracy and its ability to generate contextually appropriate interruptions, indicating that understanding-centric metrics alone are insufficient to characterize conversational social competence. More encouragingly, these diagnostics from SocialOmni yield actionable signals for bridging the perception-interaction divide in future OLMs.
title SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
topic Artificial Intelligence
url https://arxiv.org/abs/2603.16859