Gespeichert in:
| 1. Verfasser: | Sakpiboonchit, Siratish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2509.08696 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Gencho: Room Impulse Response Generation from Reverberant Speech and Text via Diffusion Transformers
von: Lin, Jackie, et al.
Veröffentlicht: (2026)
von: Lin, Jackie, et al.
Veröffentlicht: (2026)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
von: Hou, Siyuan, et al.
Veröffentlicht: (2024)
von: Hou, Siyuan, et al.
Veröffentlicht: (2024)
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
von: Zheng, Qixi, et al.
Veröffentlicht: (2025)
von: Zheng, Qixi, et al.
Veröffentlicht: (2025)
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
von: Gu, Yi, et al.
Veröffentlicht: (2026)
von: Gu, Yi, et al.
Veröffentlicht: (2026)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
von: Hasan, Rashedul, et al.
Veröffentlicht: (2025)
von: Hasan, Rashedul, et al.
Veröffentlicht: (2025)
Study of Lightweight Transformer Architectures for Single-Channel Speech Enhancement
von: Zhao, Haixin, et al.
Veröffentlicht: (2025)
von: Zhao, Haixin, et al.
Veröffentlicht: (2025)
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
von: de Groot, Dimme, et al.
Veröffentlicht: (2025)
von: de Groot, Dimme, et al.
Veröffentlicht: (2025)
Post-Training Quantization for Audio Diffusion Transformers
von: Khandelwal, Tanmay, et al.
Veröffentlicht: (2025)
von: Khandelwal, Tanmay, et al.
Veröffentlicht: (2025)
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
von: Kirdey, Stanislav
Veröffentlicht: (2025)
von: Kirdey, Stanislav
Veröffentlicht: (2025)
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
Complex Image-Generative Diffusion Transformer for Audio Denoising
von: Li, Junhui, et al.
Veröffentlicht: (2024)
von: Li, Junhui, et al.
Veröffentlicht: (2024)
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement
von: Zhang, Qiquan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiquan, et al.
Veröffentlicht: (2024)
ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
von: Xin, Detai, et al.
Veröffentlicht: (2026)
von: Xin, Detai, et al.
Veröffentlicht: (2026)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
Improving Musical Accompaniment Co-creation via Diffusion Transformers
von: Nistal, Javier, et al.
Veröffentlicht: (2024)
von: Nistal, Javier, et al.
Veröffentlicht: (2024)
EDSep: An Effective Diffusion-Based Method for Speech Source Separation
von: Dong, Jinwei, et al.
Veröffentlicht: (2025)
von: Dong, Jinwei, et al.
Veröffentlicht: (2025)
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
Wireless Hearables With Programmable Speech AI Accelerators
von: Itani, Malek, et al.
Veröffentlicht: (2025)
von: Itani, Malek, et al.
Veröffentlicht: (2025)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Device Feature based on Graph Fourier Transformation with Logarithmic Processing For Detection of Replay Speech Attacks
von: He, Mingrui, et al.
Veröffentlicht: (2024)
von: He, Mingrui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023) -
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024) -
Gencho: Room Impulse Response Generation from Reverberant Speech and Text via Diffusion Transformers
von: Lin, Jackie, et al.
Veröffentlicht: (2026) -
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025) -
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)