When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yihuan, Xue, Jun, Jiajun, Liu, Li, Daixian, Zhang, Tong, Yi, Zhuolin, Ren, Yanzhen, Li, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Well Do Current Speech Deepfake Detection Methods Generalize to the Real World?
by: Li, Daixian, et al.
Published: (2026)
by: Li, Daixian, et al.
Published: (2026)
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
by: Huang, Yihuan, et al.
Published: (2025)
by: Huang, Yihuan, et al.
Published: (2025)
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
RTCFake: Speech Deepfake Detection in Real-Time Communication
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
Wireless Semantic Communications for Video Conferencing
by: Jiang, Peiwen, et al.
Published: (2022)
by: Jiang, Peiwen, et al.
Published: (2022)
Video Conferencing
Published: (2024)
Published: (2024)
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
by: Luo, Longjie, et al.
Published: (2025)
by: Luo, Longjie, et al.
Published: (2025)
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Automated, Cross-Layer Root Cause Analysis of 5G Video-Conferencing Quality Degradation
by: Yi, Fan, et al.
Published: (2025)
by: Yi, Fan, et al.
Published: (2025)
Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing
by: Tong, Haonan, et al.
Published: (2024)
by: Tong, Haonan, et al.
Published: (2024)
NTIRE 2025 Challenge on Video Quality Enhancement for Video Conferencing: Datasets, Methods and Results
by: Jain, Varun, et al.
Published: (2025)
by: Jain, Varun, et al.
Published: (2025)
BasicAVSR: Arbitrary-Scale Video Super-Resolution via Image Priors and Enhanced Motion Compensation
by: Shang, Wei, et al.
Published: (2025)
by: Shang, Wei, et al.
Published: (2025)
Reparo: Loss-Resilient Generative Codec for Video Conferencing
by: Li, Tianhong, et al.
Published: (2023)
by: Li, Tianhong, et al.
Published: (2023)
HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR Headsets
by: Jin, Yili, et al.
Published: (2024)
by: Jin, Yili, et al.
Published: (2024)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
Audio-visual Event Localization on Portrait Mode Short Videos
by: Liu, Wuyang, et al.
Published: (2025)
by: Liu, Wuyang, et al.
Published: (2025)
Scalable Video Conferencing Using SDN Principles
by: Michel, Oliver, et al.
Published: (2025)
by: Michel, Oliver, et al.
Published: (2025)
Resolution-Agnostic Neural Compression for High-Fidelity Portrait Video Conferencing via Implicit Radiance Fields
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
Low-Latency Video Conferencing via Optimized Packet Routing and Reordering
by: Xiao, Yao, et al.
Published: (2023)
by: Xiao, Yao, et al.
Published: (2023)
Turn Your Face Into An Attack Surface: Screen Attack Using Facial Reflections in Video Conferencing
by: Huang, Yong, et al.
Published: (2026)
by: Huang, Yong, et al.
Published: (2026)
Stable but Wrong: When More Data Degrades Scientific Conclusions
by: Zhang, Zhipeng, et al.
Published: (2026)
by: Zhang, Zhipeng, et al.
Published: (2026)
Conferencing and care
by: Sara Fuller
Published: (2025)
by: Sara Fuller
Published: (2025)
ICME 2025 Grand Challenge on Video Super-Resolution for Video Conferencing
by: Naderi, Babak, et al.
Published: (2025)
by: Naderi, Babak, et al.
Published: (2025)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
zhuolinqu/Wolbachia-Spatial: Multistage Spatial Model for Wolbachia-Infected Mosquito Release
by: Zhuolin Qu, et al.
Published: (2026)
by: Zhuolin Qu, et al.
Published: (2026)
Multistage spatial model for informing release of Wolbachia-infected mosquitoes as disease control
by: Qu, Zhuolin, et al.
Published: (2024)
by: Qu, Zhuolin, et al.
Published: (2024)
Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Online Conferencing Application
by: Karthiban .R, et al.
Published: (2025)
by: Karthiban .R, et al.
Published: (2025)
When Large Vision-Language Models Meet Person Re-Identification
by: Wang, Qizao, et al.
Published: (2024)
by: Wang, Qizao, et al.
Published: (2024)
ViCES - Video Conferencing Educational Services Main Project Outcomes
Published: (2022)
Published: (2022)
Loss-resilient Coding of Texture and Depth for Free-viewpoint Video Conferencing
by: Macchiavello, Bruno, et al.
Published: (2013)
by: Macchiavello, Bruno, et al.
Published: (2013)
Teledermatology (Live Video Conferencing) Network Devoted to Competition Swimmers in France
by: E. Mahé
Published: (2025)
by: E. Mahé
Published: (2025)
Study on the Mechanism by Which Fe Promotes Toluene Degradation by sp. TG-1.
by: Qiao, Yue, et al.
Published: (2025)
by: Qiao, Yue, et al.
Published: (2025)
When Molecular Similarity Works: Property Cliffs Reveal Hidden Errors
by: Hu, Di, et al.
Published: (2026)
by: Hu, Di, et al.
Published: (2026)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
by: Liu, Wuyang, et al.
Published: (2023)
by: Liu, Wuyang, et al.
Published: (2023)
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Lightweight Call Signaling and Peer-to-Peer Control of WebRTC Video Conferencing
by: Singh, Kundan
Published: (2026)
by: Singh, Kundan
Published: (2026)
Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?
by: Han, Yujin, et al.
Published: (2025)
by: Han, Yujin, et al.
Published: (2025)
Similar Items
-
How Well Do Current Speech Deepfake Detection Methods Generalize to the Real World?
by: Li, Daixian, et al.
Published: (2026) -
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
by: Zhang, Tong, et al.
Published: (2025) -
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
by: Xue, Jun, et al.
Published: (2026) -
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
by: Huang, Yihuan, et al.
Published: (2025) -
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
by: Xue, Jun, et al.
Published: (2026)