EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yuhao, Du, Yuhao, Dai, Zhanchen, Ma, Xiangnan, Kou, Kaiqi, Wang, Benyou, Li, Haizhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
Roadmap towards Superhuman Speech Understanding using Large Language Models
di: Bu, Fan, et al.
Pubblicazione: (2024)
di: Bu, Fan, et al.
Pubblicazione: (2024)
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
di: Wang, Jialing, et al.
Pubblicazione: (2026)
di: Wang, Jialing, et al.
Pubblicazione: (2026)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
di: Du, Yuhao, et al.
Pubblicazione: (2025)
di: Du, Yuhao, et al.
Pubblicazione: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
di: Jiang, Feng, et al.
Pubblicazione: (2025)
di: Jiang, Feng, et al.
Pubblicazione: (2025)
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
di: Lin, Liang, et al.
Pubblicazione: (2026)
di: Lin, Liang, et al.
Pubblicazione: (2026)
A Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech Interactions
di: Jiang, Yiheng, et al.
Pubblicazione: (2025)
di: Jiang, Yiheng, et al.
Pubblicazione: (2025)
EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
di: Zhou, Li, et al.
Pubblicazione: (2025)
di: Zhou, Li, et al.
Pubblicazione: (2025)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
di: Gao, Yan, et al.
Pubblicazione: (2025)
di: Gao, Yan, et al.
Pubblicazione: (2025)
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
di: Ao, Junyi, et al.
Pubblicazione: (2025)
di: Ao, Junyi, et al.
Pubblicazione: (2025)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
di: Zhang, Tong, et al.
Pubblicazione: (2025)
di: Zhang, Tong, et al.
Pubblicazione: (2025)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
di: Wang, Peng, et al.
Pubblicazione: (2026)
di: Wang, Peng, et al.
Pubblicazione: (2026)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
di: Papi, Sara, et al.
Pubblicazione: (2026)
di: Papi, Sara, et al.
Pubblicazione: (2026)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
di: Xu, Ke, et al.
Pubblicazione: (2026)
di: Xu, Ke, et al.
Pubblicazione: (2026)
Towards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection
di: Guo, Xiaoxuan, et al.
Pubblicazione: (2026)
di: Guo, Xiaoxuan, et al.
Pubblicazione: (2026)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
di: Papi, Sara, et al.
Pubblicazione: (2025)
di: Papi, Sara, et al.
Pubblicazione: (2025)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
Efficient Training for Cross-lingual Speech Language Models
di: Zhou, Yan, et al.
Pubblicazione: (2026)
di: Zhou, Yan, et al.
Pubblicazione: (2026)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
di: Song, Yuhan, et al.
Pubblicazione: (2025)
di: Song, Yuhan, et al.
Pubblicazione: (2025)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
di: Ma, Rao, et al.
Pubblicazione: (2025)
di: Ma, Rao, et al.
Pubblicazione: (2025)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
di: Su, Hsuan, et al.
Pubblicazione: (2024)
di: Su, Hsuan, et al.
Pubblicazione: (2024)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
di: Gong, Yitian, et al.
Pubblicazione: (2025)
di: Gong, Yitian, et al.
Pubblicazione: (2025)
EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions
di: Modi, Smit Nautambhai, et al.
Pubblicazione: (2026)
di: Modi, Smit Nautambhai, et al.
Pubblicazione: (2026)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)
di: Futami, Hayato, et al.
Pubblicazione: (2025)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
di: Shi, Jiacheng, et al.
Pubblicazione: (2025)
di: Shi, Jiacheng, et al.
Pubblicazione: (2025)
Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)
di: Wang, Peidong
Pubblicazione: (2026)
di: Wang, Peidong
Pubblicazione: (2026)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
di: Cheng, Ziyang, et al.
Pubblicazione: (2026)
di: Cheng, Ziyang, et al.
Pubblicazione: (2026)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
di: Zhou, Li, et al.
Pubblicazione: (2026)
di: Zhou, Li, et al.
Pubblicazione: (2026)
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
di: Xie, Zhifei, et al.
Pubblicazione: (2026)
di: Xie, Zhifei, et al.
Pubblicazione: (2026)
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
di: Shan, Weiqiao, et al.
Pubblicazione: (2025)
di: Shan, Weiqiao, et al.
Pubblicazione: (2025)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
di: Zhao, Mengjie, et al.
Pubblicazione: (2026)
di: Zhao, Mengjie, et al.
Pubblicazione: (2026)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
di: Zhang, Shaolei, et al.
Pubblicazione: (2024)
di: Zhang, Shaolei, et al.
Pubblicazione: (2024)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
di: Dai, Xunlian, et al.
Pubblicazione: (2025)
di: Dai, Xunlian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
di: Zhang, Yuhao, et al.
Pubblicazione: (2025) -
Soundwave: Less is More for Speech-Text Alignment in LLMs
di: Zhang, Yuhao, et al.
Pubblicazione: (2025) -
Roadmap towards Superhuman Speech Understanding using Large Language Models
di: Bu, Fan, et al.
Pubblicazione: (2024) -
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
di: Wang, Jialing, et al.
Pubblicazione: (2026) -
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
di: Du, Yuhao, et al.
Pubblicazione: (2025)