TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dang, Trung, Rao, Sharath, Gupta, Ananya, Gagne, Christopher, Tzirakis, Panagiotis, Baird, Alice, Cłapa, Jakub Piotr, Chin, Peter, Cowen, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The NeurIPS 2023 Machine Learning for Audio Workshop: Affective Audio Benchmarks and Novel Data
von: Baird, Alice, et al.
Veröffentlicht: (2024)
von: Baird, Alice, et al.
Veröffentlicht: (2024)
The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge
von: Tzirakis, Panagiotis, et al.
Veröffentlicht: (2026)
von: Tzirakis, Panagiotis, et al.
Veröffentlicht: (2026)
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
TADA! Tuning Audio Diffusion Models through Activation Steering
von: Staniszewski, Łukasz, et al.
Veröffentlicht: (2026)
von: Staniszewski, Łukasz, et al.
Veröffentlicht: (2026)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
The 6th Affective Behavior Analysis in-the-wild (ABAW) Competition
von: Kollias, Dimitrios, et al.
Veröffentlicht: (2024)
von: Kollias, Dimitrios, et al.
Veröffentlicht: (2024)
A Multilingual Framework for Dysarthria: Detection, Severity Classification, Speech-to-Text, and Clean Speech Generation
von: Raghu, Ananya, et al.
Veröffentlicht: (2025)
von: Raghu, Ananya, et al.
Veröffentlicht: (2025)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR
von: Magoshi, Ryo, et al.
Veröffentlicht: (2026)
von: Magoshi, Ryo, et al.
Veröffentlicht: (2026)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Adaptive Duration Model for Text Speech Alignment
von: Cao, Junjie
Veröffentlicht: (2025)
von: Cao, Junjie
Veröffentlicht: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
Investigation on the Robustness of Acoustic Foundation Models on Post Exercise Speech
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2026)
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2026)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
Comparison Performance of Spectrogram and Scalogram as Input of Acoustic Recognition Task
von: Phan, Dang Thoai
Veröffentlicht: (2024)
von: Phan, Dang Thoai
Veröffentlicht: (2024)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Optimizing Speech Language Models for Acoustic Consistency
von: Rohanian, Morteza, et al.
Veröffentlicht: (2025)
von: Rohanian, Morteza, et al.
Veröffentlicht: (2025)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
von: Seki, Shogo, et al.
Veröffentlicht: (2025)
von: Seki, Shogo, et al.
Veröffentlicht: (2025)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
Versatile Symbolic Music-for-Music Modeling via Function Alignment
von: Jiang, Junyan, et al.
Veröffentlicht: (2025)
von: Jiang, Junyan, et al.
Veröffentlicht: (2025)
CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generation
von: Poon, Crystal Min Hui, et al.
Veröffentlicht: (2025)
von: Poon, Crystal Min Hui, et al.
Veröffentlicht: (2025)
Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
von: Huynh-Nguyen, Hieu-Nghia, et al.
Veröffentlicht: (2025)
von: Huynh-Nguyen, Hieu-Nghia, et al.
Veröffentlicht: (2025)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The NeurIPS 2023 Machine Learning for Audio Workshop: Affective Audio Benchmarks and Novel Data
von: Baird, Alice, et al.
Veröffentlicht: (2024) -
The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge
von: Tzirakis, Panagiotis, et al.
Veröffentlicht: (2026) -
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024) -
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024) -
TADA! Tuning Audio Diffusion Models through Activation Steering
von: Staniszewski, Łukasz, et al.
Veröffentlicht: (2026)