TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Dang, Trung, Rao, Sharath, Gupta, Ananya, Gagne, Christopher, Tzirakis, Panagiotis, Baird, Alice, Cłapa, Jakub Piotr, Chin, Peter, Cowen, Alan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The NeurIPS 2023 Machine Learning for Audio Workshop: Affective Audio Benchmarks and Novel Data
di: Baird, Alice, et al.
Pubblicazione: (2024)
di: Baird, Alice, et al.
Pubblicazione: (2024)
The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge
di: Tzirakis, Panagiotis, et al.
Pubblicazione: (2026)
di: Tzirakis, Panagiotis, et al.
Pubblicazione: (2026)
Zero-Shot Text-to-Speech from Continuous Text Streams
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
TADA! Tuning Audio Diffusion Models through Activation Steering
di: Staniszewski, Łukasz, et al.
Pubblicazione: (2026)
di: Staniszewski, Łukasz, et al.
Pubblicazione: (2026)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
di: Li, Xuanchen, et al.
Pubblicazione: (2025)
di: Li, Xuanchen, et al.
Pubblicazione: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
The 6th Affective Behavior Analysis in-the-wild (ABAW) Competition
di: Kollias, Dimitrios, et al.
Pubblicazione: (2024)
di: Kollias, Dimitrios, et al.
Pubblicazione: (2024)
A Multilingual Framework for Dysarthria: Detection, Severity Classification, Speech-to-Text, and Clean Speech Generation
di: Raghu, Ananya, et al.
Pubblicazione: (2025)
di: Raghu, Ananya, et al.
Pubblicazione: (2025)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR
di: Magoshi, Ryo, et al.
Pubblicazione: (2026)
di: Magoshi, Ryo, et al.
Pubblicazione: (2026)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
di: Chen, Wenxi, et al.
Pubblicazione: (2025)
di: Chen, Wenxi, et al.
Pubblicazione: (2025)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2025)
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2025)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
di: Vallés-Pérez, Ivan, et al.
Pubblicazione: (2023)
di: Vallés-Pérez, Ivan, et al.
Pubblicazione: (2023)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
di: Zhang, Pei, et al.
Pubblicazione: (2025)
di: Zhang, Pei, et al.
Pubblicazione: (2025)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
di: Ma, Rao, et al.
Pubblicazione: (2025)
di: Ma, Rao, et al.
Pubblicazione: (2025)
Adaptive Duration Model for Text Speech Alignment
di: Cao, Junjie
Pubblicazione: (2025)
di: Cao, Junjie
Pubblicazione: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
di: Du, Chenpeng, et al.
Pubblicazione: (2022)
di: Du, Chenpeng, et al.
Pubblicazione: (2022)
Investigation on the Robustness of Acoustic Foundation Models on Post Exercise Speech
di: Xue, Xiangyuan, et al.
Pubblicazione: (2026)
di: Xue, Xiangyuan, et al.
Pubblicazione: (2026)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
Comparison Performance of Spectrogram and Scalogram as Input of Acoustic Recognition Task
di: Phan, Dang Thoai
Pubblicazione: (2024)
di: Phan, Dang Thoai
Pubblicazione: (2024)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Optimizing Speech Language Models for Acoustic Consistency
di: Rohanian, Morteza, et al.
Pubblicazione: (2025)
di: Rohanian, Morteza, et al.
Pubblicazione: (2025)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
di: Ye, Zongli, et al.
Pubblicazione: (2025)
di: Ye, Zongli, et al.
Pubblicazione: (2025)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
di: Shu, Yuchun, et al.
Pubblicazione: (2024)
di: Shu, Yuchun, et al.
Pubblicazione: (2024)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
di: Ren, Yong, et al.
Pubblicazione: (2026)
di: Ren, Yong, et al.
Pubblicazione: (2026)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
di: Lei, Shun, et al.
Pubblicazione: (2023)
di: Lei, Shun, et al.
Pubblicazione: (2023)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
di: Yang, Jinhyeok, et al.
Pubblicazione: (2024)
di: Yang, Jinhyeok, et al.
Pubblicazione: (2024)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
di: Cao, Junjie, et al.
Pubblicazione: (2025)
di: Cao, Junjie, et al.
Pubblicazione: (2025)
MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
di: Bai, Yatong, et al.
Pubblicazione: (2023)
di: Bai, Yatong, et al.
Pubblicazione: (2023)
Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
di: Seki, Shogo, et al.
Pubblicazione: (2025)
di: Seki, Shogo, et al.
Pubblicazione: (2025)
Acoustic BPE for Speech Generation with Discrete Tokens
di: Shen, Feiyu, et al.
Pubblicazione: (2023)
di: Shen, Feiyu, et al.
Pubblicazione: (2023)
Versatile Symbolic Music-for-Music Modeling via Function Alignment
di: Jiang, Junyan, et al.
Pubblicazione: (2025)
di: Jiang, Junyan, et al.
Pubblicazione: (2025)
CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generation
di: Poon, Crystal Min Hui, et al.
Pubblicazione: (2025)
di: Poon, Crystal Min Hui, et al.
Pubblicazione: (2025)
Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The NeurIPS 2023 Machine Learning for Audio Workshop: Affective Audio Benchmarks and Novel Data
di: Baird, Alice, et al.
Pubblicazione: (2024) -
The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge
di: Tzirakis, Panagiotis, et al.
Pubblicazione: (2026) -
Zero-Shot Text-to-Speech from Continuous Text Streams
di: Dang, Trung, et al.
Pubblicazione: (2024) -
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
di: Dang, Trung, et al.
Pubblicazione: (2024) -
TADA! Tuning Audio Diffusion Models through Activation Steering
di: Staniszewski, Łukasz, et al.
Pubblicazione: (2026)