Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Yu-Kuan, Lee, Cheng-Kuang, Wang, Hsiu-Hsuan, Lee, Hung-yi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
por: Zhao, Zhixian, et al.
Publicado: (2026)
por: Zhao, Zhixian, et al.
Publicado: (2026)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
por: Tseng, Liang-Hsuan, et al.
Publicado: (2025)
por: Tseng, Liang-Hsuan, et al.
Publicado: (2025)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
por: Cheng, Xize, et al.
Publicado: (2025)
por: Cheng, Xize, et al.
Publicado: (2025)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
por: Huang, Kuan-Po, et al.
Publicado: (2023)
por: Huang, Kuan-Po, et al.
Publicado: (2023)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
por: Liu, Ailin, et al.
Publicado: (2024)
por: Liu, Ailin, et al.
Publicado: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
por: Arora, Siddhant, et al.
Publicado: (2024)
por: Arora, Siddhant, et al.
Publicado: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
por: Li, Yue, et al.
Publicado: (2024)
por: Li, Yue, et al.
Publicado: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
por: Park, Seohyun, et al.
Publicado: (2025)
por: Park, Seohyun, et al.
Publicado: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
por: Wang, Hongbin, et al.
Publicado: (2025)
por: Wang, Hongbin, et al.
Publicado: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
por: Feng, Tiantian, et al.
Publicado: (2023)
por: Feng, Tiantian, et al.
Publicado: (2023)
VoiceX: A Text-To-Speech Framework for Custom Voices
por: Mertes, Silvan, et al.
Publicado: (2024)
por: Mertes, Silvan, et al.
Publicado: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
por: Xiao, Yi, et al.
Publicado: (2022)
por: Xiao, Yi, et al.
Publicado: (2022)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
por: Kuan, Chun-Yi, et al.
Publicado: (2024)
por: Kuan, Chun-Yi, et al.
Publicado: (2024)
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
por: Kuan, Chun-Yi, et al.
Publicado: (2025)
por: Kuan, Chun-Yi, et al.
Publicado: (2025)
I Know Your Feelings Before You Do: Predicting Future Affective Reactions in Human-Computer Dialogue
por: Li, Yuanchao, et al.
Publicado: (2023)
por: Li, Yuanchao, et al.
Publicado: (2023)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
por: Chen, Youjun, et al.
Publicado: (2025)
por: Chen, Youjun, et al.
Publicado: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
por: Nowrin, Sadia, et al.
Publicado: (2024)
por: Nowrin, Sadia, et al.
Publicado: (2024)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
por: Yu, Luca Jiang-Tao, et al.
Publicado: (2024)
por: Yu, Luca Jiang-Tao, et al.
Publicado: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
por: Dutta, Satwik, et al.
Publicado: (2025)
por: Dutta, Satwik, et al.
Publicado: (2025)
SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality Awareness
por: Chang, Ruei-Che, et al.
Publicado: (2024)
por: Chang, Ruei-Che, et al.
Publicado: (2024)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
por: Yuan, Kuang, et al.
Publicado: (2025)
por: Yuan, Kuang, et al.
Publicado: (2025)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
por: Xue, Liumeng, et al.
Publicado: (2024)
por: Xue, Liumeng, et al.
Publicado: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
por: Sharma, Roshan, et al.
Publicado: (2024)
por: Sharma, Roshan, et al.
Publicado: (2024)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
por: Zhou, Dongliang, et al.
Publicado: (2025)
por: Zhou, Dongliang, et al.
Publicado: (2025)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
por: Ma, Yong, et al.
Publicado: (2025)
por: Ma, Yong, et al.
Publicado: (2025)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
por: Cai, Zhuojiang, et al.
Publicado: (2024)
por: Cai, Zhuojiang, et al.
Publicado: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
por: Nishida, Naoto, et al.
Publicado: (2025)
por: Nishida, Naoto, et al.
Publicado: (2025)
Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours of EEG Data
por: Sato, Motoshige, et al.
Publicado: (2024)
por: Sato, Motoshige, et al.
Publicado: (2024)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
por: Mishra, Ruchik, et al.
Publicado: (2024)
por: Mishra, Ruchik, et al.
Publicado: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
por: Wang, Dingdong, et al.
Publicado: (2025)
por: Wang, Dingdong, et al.
Publicado: (2025)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
por: Kuan, Chun-Yi, et al.
Publicado: (2024)
por: Kuan, Chun-Yi, et al.
Publicado: (2024)
Are Expressions for Music Emotions the Same Across Cultures?
por: Celen, Elif, et al.
Publicado: (2025)
por: Celen, Elif, et al.
Publicado: (2025)
Towards Reliable Large Audio Language Model
por: Ma, Ziyang, et al.
Publicado: (2025)
por: Ma, Ziyang, et al.
Publicado: (2025)
Tailors: New Music Timbre Visualizer to Entertain Music Through Imagery
por: Lee, ChungHa
Publicado: (2024)
por: Lee, ChungHa
Publicado: (2024)
CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
por: Han, Runduo, et al.
Publicado: (2025)
por: Han, Runduo, et al.
Publicado: (2025)
DOO-RE: A dataset of ambient sensors in a meeting room for activity recognition
por: Kim, Hyunju, et al.
Publicado: (2024)
por: Kim, Hyunju, et al.
Publicado: (2024)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
por: Sankey-Olsen, Cuno, et al.
Publicado: (2025)
por: Sankey-Olsen, Cuno, et al.
Publicado: (2025)
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
por: Li, Yinan, et al.
Publicado: (2026)
por: Li, Yinan, et al.
Publicado: (2026)
Ejemplares similares
-
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
por: Zhao, Zhixian, et al.
Publicado: (2026) -
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
por: Tseng, Liang-Hsuan, et al.
Publicado: (2025) -
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
por: Cheng, Xize, et al.
Publicado: (2025) -
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
por: Huang, Kuan-Po, et al.
Publicado: (2023) -
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)