Conversational Speech Naturalness Predictor
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Anfeng, Gaur, Yashesh, Kanda, Naoyuki, Ouyang, Zhicheng, Zmolikova, Katerina, Raj, Desh, Merello, Simone, Sun, Anna, Kalinli, Ozlem |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Faster Speech-LLaMA Inference with Multi-token Prediction
di: Raj, Desh, et al.
Pubblicazione: (2024)
di: Raj, Desh, et al.
Pubblicazione: (2024)
Can Speech LLMs Think while Listening?
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025)
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
di: Raj, Desh
Pubblicazione: (2024)
di: Raj, Desh
Pubblicazione: (2024)
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
di: Moritz, Niko, et al.
Pubblicazione: (2024)
di: Moritz, Niko, et al.
Pubblicazione: (2024)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
di: Kang, Wonjune, et al.
Pubblicazione: (2024)
di: Kang, Wonjune, et al.
Pubblicazione: (2024)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
di: Ma, Yingyi, et al.
Pubblicazione: (2024)
di: Ma, Yingyi, et al.
Pubblicazione: (2024)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
di: Zhou, Wei, et al.
Pubblicazione: (2024)
di: Zhou, Wei, et al.
Pubblicazione: (2024)
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
TS3-Codec: Transformer-Based Simple Streaming Single Codec
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
di: Seide, Frank, et al.
Pubblicazione: (2024)
di: Seide, Frank, et al.
Pubblicazione: (2024)
Token-Weighted RNN-T for Learning from Flawed Data
di: Keren, Gil, et al.
Pubblicazione: (2024)
di: Keren, Gil, et al.
Pubblicazione: (2024)
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
di: Pan, Jing, et al.
Pubblicazione: (2023)
di: Pan, Jing, et al.
Pubblicazione: (2023)
Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2025)
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2025)
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2026)
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2026)
Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
Joint ASR and Speaker Role Tagging with Serialized Output Training
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
di: Lian, Zhicheng, et al.
Pubblicazione: (2025)
di: Lian, Zhicheng, et al.
Pubblicazione: (2025)
Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning
di: Ouyang, Zhicheng, et al.
Pubblicazione: (2026)
di: Ouyang, Zhicheng, et al.
Pubblicazione: (2026)
Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
di: Cornell, Samuele, et al.
Pubblicazione: (2025)
di: Cornell, Samuele, et al.
Pubblicazione: (2025)
On Speaker Attribution with SURT
di: Raj, Desh, et al.
Pubblicazione: (2024)
di: Raj, Desh, et al.
Pubblicazione: (2024)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
di: Subramanian, Aswin Shanmugam, et al.
Pubblicazione: (2025)
di: Subramanian, Aswin Shanmugam, et al.
Pubblicazione: (2025)
Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge
di: E, Dhanya, et al.
Pubblicazione: (2026)
di: E, Dhanya, et al.
Pubblicazione: (2026)
Profile-Error-Tolerant Target-Speaker Voice Activity Detection
di: Wang, Dongmei, et al.
Pubblicazione: (2023)
di: Wang, Dongmei, et al.
Pubblicazione: (2023)
Large Language Models based ASR Error Correction for Child Conversations
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
DiariST: Streaming Speech Translation with Speaker Diarization
di: Yang, Mu, et al.
Pubblicazione: (2023)
di: Yang, Mu, et al.
Pubblicazione: (2023)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
UniTalker: Conversational Speech-Visual Synthesis
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
Efficient Streaming LLM for Speech Recognition
di: Jia, Junteng, et al.
Pubblicazione: (2024)
di: Jia, Junteng, et al.
Pubblicazione: (2024)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
di: Cohen, Eyal, et al.
Pubblicazione: (2025)
di: Cohen, Eyal, et al.
Pubblicazione: (2025)
Dynamic Slimmable Networks for Efficient Speech Separation
di: Elminshawi, Mohamed, et al.
Pubblicazione: (2025)
di: Elminshawi, Mohamed, et al.
Pubblicazione: (2025)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2023)
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2023)
Effective Integration of KAN for Keyword Spotting
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
Audio-visual child-adult speaker classification in dyadic interactions
di: Xu, Anfeng, et al.
Pubblicazione: (2023)
di: Xu, Anfeng, et al.
Pubblicazione: (2023)
Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs
di: Behringer, Lyonel, et al.
Pubblicazione: (2026)
di: Behringer, Lyonel, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Faster Speech-LLaMA Inference with Multi-token Prediction
di: Raj, Desh, et al.
Pubblicazione: (2024) -
Can Speech LLMs Think while Listening?
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025) -
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024) -
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
di: Raj, Desh
Pubblicazione: (2024) -
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
di: Moritz, Niko, et al.
Pubblicazione: (2024)