Salvato in:
| Autori principali: | Wang, Siyin, Yu, Wenyi, Chen, Xianzhao, Tian, Xiaohai, Zhang, Jun, Lu, Lu, Zhang, Chao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2510.16756 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
di: Wang, Siyin, et al.
Pubblicazione: (2025)
di: Wang, Siyin, et al.
Pubblicazione: (2025)
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
di: Yu, Wenyi, et al.
Pubblicazione: (2024)
di: Yu, Wenyi, et al.
Pubblicazione: (2024)
Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
di: Wang, Siyin, et al.
Pubblicazione: (2024)
di: Wang, Siyin, et al.
Pubblicazione: (2024)
Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities
di: Wang, Siyin, et al.
Pubblicazione: (2024)
di: Wang, Siyin, et al.
Pubblicazione: (2024)
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2026)
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2026)
Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
di: Hori, Chiori, et al.
Pubblicazione: (2025)
di: Hori, Chiori, et al.
Pubblicazione: (2025)
SALMONN: Towards Generic Hearing Abilities for Large Language Models
di: Tang, Changli, et al.
Pubblicazione: (2023)
di: Tang, Changli, et al.
Pubblicazione: (2023)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
di: Korbar, Bruno, et al.
Pubblicazione: (2024)
di: Korbar, Bruno, et al.
Pubblicazione: (2024)
Listening without Looking: Modality Bias in Audio-Visual Captioning
di: Ishikawa, Yuchi, et al.
Pubblicazione: (2025)
di: Ishikawa, Yuchi, et al.
Pubblicazione: (2025)
ANIM-400K: A Large-Scale Dataset for Automated End-To-End Dubbing of Video
di: Cai, Kevin, et al.
Pubblicazione: (2024)
di: Cai, Kevin, et al.
Pubblicazione: (2024)
SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data
di: Lu, Yichen, et al.
Pubblicazione: (2024)
di: Lu, Yichen, et al.
Pubblicazione: (2024)
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
di: Dai, Yuqin, et al.
Pubblicazione: (2025)
di: Dai, Yuqin, et al.
Pubblicazione: (2025)
Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
di: Wang, Siyin, et al.
Pubblicazione: (2025)
di: Wang, Siyin, et al.
Pubblicazione: (2025)
CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation
di: Liu, Xi, et al.
Pubblicazione: (2024)
di: Liu, Xi, et al.
Pubblicazione: (2024)
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
di: Wu, Yihan, et al.
Pubblicazione: (2024)
di: Wu, Yihan, et al.
Pubblicazione: (2024)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
di: Somayazulu, Arjun, et al.
Pubblicazione: (2024)
di: Somayazulu, Arjun, et al.
Pubblicazione: (2024)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
di: Hu, Rui, et al.
Pubblicazione: (2025)
di: Hu, Rui, et al.
Pubblicazione: (2025)
End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation
di: Di Pierno, Andrea, et al.
Pubblicazione: (2025)
di: Di Pierno, Andrea, et al.
Pubblicazione: (2025)
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
di: Fei, Zhengcong, et al.
Pubblicazione: (2023)
di: Fei, Zhengcong, et al.
Pubblicazione: (2023)
Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
di: Huang, Jinhe, et al.
Pubblicazione: (2025)
di: Huang, Jinhe, et al.
Pubblicazione: (2025)
Advancing Automated Spatio-Semantic Analysis in Picture Description Using Language Models
di: Ng, Si-Ioi, et al.
Pubblicazione: (2025)
di: Ng, Si-Ioi, et al.
Pubblicazione: (2025)
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
EchoTrack: Auditory Referring Multi-Object Tracking for Autonomous Driving
di: Lin, Jiacheng, et al.
Pubblicazione: (2024)
di: Lin, Jiacheng, et al.
Pubblicazione: (2024)
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
di: Zhang, Haomin, et al.
Pubblicazione: (2025)
di: Zhang, Haomin, et al.
Pubblicazione: (2025)
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue
di: Lu, Hui, et al.
Pubblicazione: (2026)
di: Lu, Hui, et al.
Pubblicazione: (2026)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
di: Ao, Junyi, et al.
Pubblicazione: (2024)
di: Ao, Junyi, et al.
Pubblicazione: (2024)
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
di: Ao, Junyi, et al.
Pubblicazione: (2025)
di: Ao, Junyi, et al.
Pubblicazione: (2025)
Qwen2.5-Omni Technical Report
di: Xu, Jin, et al.
Pubblicazione: (2025)
di: Xu, Jin, et al.
Pubblicazione: (2025)
Data Augmentation for End-to-end Code-switching Speech Recognition
di: Du, Chenpeng, et al.
Pubblicazione: (2020)
di: Du, Chenpeng, et al.
Pubblicazione: (2020)
Can Whisper perform speech-based in-context learning?
di: Wang, Siyin, et al.
Pubblicazione: (2023)
di: Wang, Siyin, et al.
Pubblicazione: (2023)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
di: Cheng, Shanbo, et al.
Pubblicazione: (2025)
di: Cheng, Shanbo, et al.
Pubblicazione: (2025)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
di: Fernandez-Lopez, Adriana, et al.
Pubblicazione: (2024)
di: Fernandez-Lopez, Adriana, et al.
Pubblicazione: (2024)
Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
di: Ríos-Vila, Antonio, et al.
Pubblicazione: (2024)
di: Ríos-Vila, Antonio, et al.
Pubblicazione: (2024)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
di: Ye, Zhen, et al.
Pubblicazione: (2026)
di: Ye, Zhen, et al.
Pubblicazione: (2026)
Qwen3-Omni Technical Report
di: Xu, Jin, et al.
Pubblicazione: (2025)
di: Xu, Jin, et al.
Pubblicazione: (2025)
Aligned Better, Listen Better for Audio-Visual Large Language Models
di: Guo, Yuxin, et al.
Pubblicazione: (2025)
di: Guo, Yuxin, et al.
Pubblicazione: (2025)
Measuring Sound Symbolism in Audio-visual Models
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024)
Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
di: Ortega-Beltrán, Elena, et al.
Pubblicazione: (2024)
di: Ortega-Beltrán, Elena, et al.
Pubblicazione: (2024)
Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
di: Li, Zhengyang, et al.
Pubblicazione: (2026)
di: Li, Zhengyang, et al.
Pubblicazione: (2026)
Unsupervised Out-of-Distribution Dialect Detection with Mahalanobis Distance
di: Das, Sourya Dipta, et al.
Pubblicazione: (2023)
di: Das, Sourya Dipta, et al.
Pubblicazione: (2023)
Documenti analoghi
-
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
di: Wang, Siyin, et al.
Pubblicazione: (2025) -
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
di: Yu, Wenyi, et al.
Pubblicazione: (2024) -
Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
di: Wang, Siyin, et al.
Pubblicazione: (2024) -
Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities
di: Wang, Siyin, et al.
Pubblicazione: (2024) -
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2026)