MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Haitian, Zhou, Yanghao, Huang, Heyan, Chen, Liangji, Cheng, YiMing, Liu, Xu, Jin, Dian, Xu, Jiajun, Liao, Jingyun, Lan, Tian, Zhou, Ziqin, Liu, Yueying, Bai, Yu, Yuan, Changsen, Zhou, Jinxing, Mao, Xian-Ling, Chen, Xuefeng, Feng, Yousheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
von: Jin, Dian, et al.
Veröffentlicht: (2025)
von: Jin, Dian, et al.
Veröffentlicht: (2025)
Look, Listen and Segment: Towards Weakly Supervised Audio-visual Semantic Segmentation
von: Li, Chengzhi, et al.
Veröffentlicht: (2026)
von: Li, Chengzhi, et al.
Veröffentlicht: (2026)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
von: Liu, Hongbo, et al.
Veröffentlicht: (2025)
von: Liu, Hongbo, et al.
Veröffentlicht: (2025)
A Numerical PDEs Approach to Evolution Equations in Shape Analysis Based on Regularized Morphoelasticity
von: Zhou, Ziqin
Veröffentlicht: (2026)
von: Zhou, Ziqin
Veröffentlicht: (2026)
Internet of Things (IoT) of Smart Homes: Privacy and Security
von: Tinashe Magara, et al.
Veröffentlicht: (2024)
von: Tinashe Magara, et al.
Veröffentlicht: (2024)
CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
von: Tian, Baoliang, et al.
Veröffentlicht: (2025)
von: Tian, Baoliang, et al.
Veröffentlicht: (2025)
Waveguiding in two-dimensional Floquet non-Abelian topological insulators
von: Zhou, Yujie, et al.
Veröffentlicht: (2025)
von: Zhou, Yujie, et al.
Veröffentlicht: (2025)
Three-period evolution in a photonic Floquet extended Su-Schrieffer-Heeger waveguide array
von: Li, Changsen, et al.
Veröffentlicht: (2025)
von: Li, Changsen, et al.
Veröffentlicht: (2025)
Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents
von: Zhou, Yunpeng
Veröffentlicht: (2026)
von: Zhou, Yunpeng
Veröffentlicht: (2026)
A Closer Look at Knowledge Distillation in Spiking Neural Network Training
von: Liu, Xu, et al.
Veröffentlicht: (2025)
von: Liu, Xu, et al.
Veröffentlicht: (2025)
A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
von: Zhang, Peixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Peixuan, et al.
Veröffentlicht: (2026)
Nitrogen and phosphorus dynamics and nutrient resorption of rhizophora mangle leaves in south Florida, USA
von: Lin, YiMing
Veröffentlicht: (2007)
von: Lin, YiMing
Veröffentlicht: (2007)
Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization
von: Han, Cailing, et al.
Veröffentlicht: (2026)
von: Han, Cailing, et al.
Veröffentlicht: (2026)
Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation
von: Zhou, Jinxing, et al.
Veröffentlicht: (2026)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2026)
Production of Aromatic Hydrocarbons from Co‐Hydropyrolysis of Biomass Components and HDPE with Application of Modified HZSM‐5 Catalyst
von: Long Ren, et al.
Veröffentlicht: (2024)
von: Long Ren, et al.
Veröffentlicht: (2024)
Effects of Polyphenols and Ascorbic Acid in Honey From Diverse Floral Origins on Liver Alcohol Metabolism
von: Zhiwei Sun, et al.
Veröffentlicht: (2025)
von: Zhiwei Sun, et al.
Veröffentlicht: (2025)
Key Materials for Potassium‐Ion Batteries: Overcoming Challenges and Opening Up Horizons for Commercialization
von: Zhiwang Liu, et al.
Veröffentlicht: (2026)
von: Zhiwang Liu, et al.
Veröffentlicht: (2026)
Diagnosing and Repairing Citation Failures in Generative Engine Optimization
von: Tian, Zhihua, et al.
Veröffentlicht: (2026)
von: Tian, Zhihua, et al.
Veröffentlicht: (2026)
Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
von: Zeng, Xianchao, et al.
Veröffentlicht: (2025)
von: Zeng, Xianchao, et al.
Veröffentlicht: (2025)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2025)
EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans
von: Zhou, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhou, Yingjie, et al.
Veröffentlicht: (2025)
The Posttraumatic Growth Process Experienced by Chinese Patients Newly Diagnosed With Crohn's Disease: A Longitudinal Descriptive Qualitative Study
von: Lingxi Chen, et al.
Veröffentlicht: (2025)
von: Lingxi Chen, et al.
Veröffentlicht: (2025)
Digital Twin-Assisted High-Precision Massive MIMO Localization in Urban Canyons
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
A Multimodal Deep Learning Framework Fusing 1D Spectra, 2DCOS Maps, and Environmental Covariates for Soil Organic Carbon Prediction in Disturbed Mining Areas
von: Zhenhong Tian, et al.
Veröffentlicht: (2026)
von: Zhenhong Tian, et al.
Veröffentlicht: (2026)
MuseBarControl: Enhancing Fine-Grained Control in Symbolic Music Generation through Pre-Training and Counterfactual Loss
von: Shu, Yangyang, et al.
Veröffentlicht: (2024)
von: Shu, Yangyang, et al.
Veröffentlicht: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
The completeness and congruences of quasi-Boolean algebras
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
LLM-Guided Evolution: An Autonomous Model Optimization for Object Detection
von: Yu, YiMing, et al.
Veröffentlicht: (2025)
von: Yu, YiMing, et al.
Veröffentlicht: (2025)
Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations
von: Gao, Huan-ang, et al.
Veröffentlicht: (2025)
von: Gao, Huan-ang, et al.
Veröffentlicht: (2025)
Diagnosing Strong-to-Weak Symmetry Breaking via Wightman Correlators
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
Nonlinear Association Between the C‐Reactive Protein–Triglyceride–Glucose Index and Rheumatoid Arthritis Risk: The Mediating Role of Body Mass Index
von: Haiping Xie, et al.
Veröffentlicht: (2025)
von: Haiping Xie, et al.
Veröffentlicht: (2025)
LLMs learn scientific taste from institutional traces across the social sciences
von: Gong, Ziqin, et al.
Veröffentlicht: (2026)
von: Gong, Ziqin, et al.
Veröffentlicht: (2026)
Sirtuin2 suppresses the polarization of regulatory T cells toward T helper 17 cells through repressing the expression of signal transducer and activator of transcription 3 in a mouse colitis model
von: Liuqing Ge, et al.
Veröffentlicht: (2024)
von: Liuqing Ge, et al.
Veröffentlicht: (2024)
Pardon? Evaluating Conversational Repair in Large Audio-Language Models
von: Huang, Shuanghong, et al.
Veröffentlicht: (2026)
von: Huang, Shuanghong, et al.
Veröffentlicht: (2026)
ExpressivityBench: Can LLMs Communicate Implicitly?
von: Tint, Joshua, et al.
Veröffentlicht: (2024)
von: Tint, Joshua, et al.
Veröffentlicht: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026) -
SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
von: Jin, Dian, et al.
Veröffentlicht: (2025) -
Look, Listen and Segment: Towards Weakly Supervised Audio-visual Semantic Segmentation
von: Li, Chengzhi, et al.
Veröffentlicht: (2026) -
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025) -
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
von: Liu, Hongbo, et al.
Veröffentlicht: (2025)