SpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM Development
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Minghan, Bai, Ye, Wang, Yuxia, Vu, Thuy-Trang, Shareghi, Ehsan, Haffari, Gholamreza |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR
von: Wang, Minghan, et al.
Veröffentlicht: (2024)
von: Wang, Minghan, et al.
Veröffentlicht: (2024)
Conversational SimulMT: Efficient Simultaneous Translation with Large Language Models
von: Wang, Minghan, et al.
Veröffentlicht: (2024)
von: Wang, Minghan, et al.
Veröffentlicht: (2024)
Towards Inference-time Scaling for Continuous Space Reasoning
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
Discrete Minds in a Continuous World: Do Language Models Know Time Passes?
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
von: Wang, Minghan, et al.
Veröffentlicht: (2026)
von: Wang, Minghan, et al.
Veröffentlicht: (2026)
Simultaneous Machine Translation with Large Language Models
von: Wang, Minghan, et al.
Veröffentlicht: (2023)
von: Wang, Minghan, et al.
Veröffentlicht: (2023)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
Beyond Imitation: Recovering Dense Rewards from Demonstrations
von: Li, Jiangnan, et al.
Veröffentlicht: (2025)
von: Li, Jiangnan, et al.
Veröffentlicht: (2025)
SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
MAPLE: Multi-Agent Adaptive Planning with Long-Term Memory for Table Reasoning
von: Bai, Ye, et al.
Veröffentlicht: (2025)
von: Bai, Ye, et al.
Veröffentlicht: (2025)
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
Resurfacing Paralinguistic Awareness in Large Audio Language Models
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Continual Speech Learning with Fused Speech Features
von: Wang, Guitao, et al.
Veröffentlicht: (2025)
von: Wang, Guitao, et al.
Veröffentlicht: (2025)
AIPO: Learning to Reason from Active Interaction
von: Liu, Junnan, et al.
Veröffentlicht: (2026)
von: Liu, Junnan, et al.
Veröffentlicht: (2026)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
Adapting Large Language Models for Document-Level Machine Translation
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
von: Feng, Pengchao, et al.
Veröffentlicht: (2025)
von: Feng, Pengchao, et al.
Veröffentlicht: (2025)
Active Continual Learning: On Balancing Knowledge Retention and Learnability
von: Vu, Thuy-Trang, et al.
Veröffentlicht: (2023)
von: Vu, Thuy-Trang, et al.
Veröffentlicht: (2023)
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
von: Hua, Yuncheng, et al.
Veröffentlicht: (2024)
von: Hua, Yuncheng, et al.
Veröffentlicht: (2024)
Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation
von: Sani, Samin Mahdizadeh, et al.
Veröffentlicht: (2024)
von: Sani, Samin Mahdizadeh, et al.
Veröffentlicht: (2024)
CONGRAD:Conflicting Gradient Filtering for Multilingual Preference Alignment
von: Li, Jiangnan, et al.
Veröffentlicht: (2025)
von: Li, Jiangnan, et al.
Veröffentlicht: (2025)
Towards Event Extraction from Speech with Contextual Clues
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
von: Feng, Tao, et al.
Veröffentlicht: (2024)
von: Feng, Tao, et al.
Veröffentlicht: (2024)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Continual Learning for Large Language Models: A Survey
von: Wu, Tongtong, et al.
Veröffentlicht: (2024)
von: Wu, Tongtong, et al.
Veröffentlicht: (2024)
Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
A Full-duplex Speech Dialogue Scheme Based On Large Language Models
von: Wang, Peng, et al.
Veröffentlicht: (2024)
von: Wang, Peng, et al.
Veröffentlicht: (2024)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
von: Riera, Pablo, et al.
Veröffentlicht: (2026)
von: Riera, Pablo, et al.
Veröffentlicht: (2026)
Detecting Mental Manipulation in Speech via Synthetic Multi-Speaker Dialogue
von: Chen, Run, et al.
Veröffentlicht: (2026)
von: Chen, Run, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
von: Wang, Xiong, et al.
Veröffentlicht: (2024)
von: Wang, Xiong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR
von: Wang, Minghan, et al.
Veröffentlicht: (2024) -
Conversational SimulMT: Efficient Simultaneous Translation with Large Language Models
von: Wang, Minghan, et al.
Veröffentlicht: (2024) -
Towards Inference-time Scaling for Continuous Space Reasoning
von: Wang, Minghan, et al.
Veröffentlicht: (2025) -
Discrete Minds in a Continuous World: Do Language Models Know Time Passes?
von: Wang, Minghan, et al.
Veröffentlicht: (2025) -
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
von: Wang, Minghan, et al.
Veröffentlicht: (2026)