Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Qiongqiong, Sailor, Hardik B., Liu, Tianchi, Aw, Ai Ti |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
von: Pan, Zihan, et al.
Veröffentlicht: (2024)
von: Pan, Zihan, et al.
Veröffentlicht: (2024)
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
von: Huzaifah, Muhammad, et al.
Veröffentlicht: (2024)
von: Huzaifah, Muhammad, et al.
Veröffentlicht: (2024)
Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
von: Liu, Tianchi, et al.
Veröffentlicht: (2024)
von: Liu, Tianchi, et al.
Veröffentlicht: (2024)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
Interpolating Speaker Identities in Embedding Space for Data Expansion
von: Liu, Tianchi, et al.
Veröffentlicht: (2025)
von: Liu, Tianchi, et al.
Veröffentlicht: (2025)
Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2023)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2023)
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
von: Cheng, Yifan, et al.
Veröffentlicht: (2025)
von: Cheng, Yifan, et al.
Veröffentlicht: (2025)
Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
von: Jung, Donghyuk, et al.
Veröffentlicht: (2026)
von: Jung, Donghyuk, et al.
Veröffentlicht: (2026)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
von: Chen, Hongjie, et al.
Veröffentlicht: (2025)
von: Chen, Hongjie, et al.
Veröffentlicht: (2025)
SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
Long-Form Speech Generation with Spoken Language Models
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
Scaling Spoken Language Models with Syllabic Speech Tokenization
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
AudioBench: A Universal Benchmark for Audio Large Language Models
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
von: Wu, Junkai, et al.
Veröffentlicht: (2024)
von: Wu, Junkai, et al.
Veröffentlicht: (2024)
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
Golden Gemini is All You Need: Finding the Sweet Spots for Speaker Verification
von: Liu, Tianchi, et al.
Veröffentlicht: (2023)
von: Liu, Tianchi, et al.
Veröffentlicht: (2023)
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
von: Wang, Bin, et al.
Veröffentlicht: (2025)
von: Wang, Bin, et al.
Veröffentlicht: (2025)
Towards Machine Unlearning for Paralinguistic Speech Processing
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
A Large Dataset of Spontaneous Speech with the Accent Spoken in São Paulo for Automatic Speech Recognition Evaluation
von: Lima, Rodrigo, et al.
Veröffentlicht: (2024)
von: Lima, Rodrigo, et al.
Veröffentlicht: (2024)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
UrduSpeech: A 156-Hour Urdu Speech Corpus with 12-Dimension Paralinguistic Annotations
von: Haq, Attia Nafees ul, et al.
Veröffentlicht: (2026)
von: Haq, Attia Nafees ul, et al.
Veröffentlicht: (2026)
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Textless Speech-to-Speech Translation With Limited Parallel Data
von: Diwan, Anuj, et al.
Veröffentlicht: (2023)
von: Diwan, Anuj, et al.
Veröffentlicht: (2023)
Self-Powered LLM Modality Expansion for Large Speech-Text Models
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
von: Tu, Wenming, et al.
Veröffentlicht: (2025)
von: Tu, Wenming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025) -
Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
von: Pan, Zihan, et al.
Veröffentlicht: (2024) -
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
von: Huzaifah, Muhammad, et al.
Veröffentlicht: (2024) -
Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
von: Liu, Tianchi, et al.
Veröffentlicht: (2024) -
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)