CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Helin, Hai, Jiarui, Chong, Dading, Thakkar, Karan, Feng, Tiantian, Yang, Dongchao, Lee, Junhyeok, Thebaud, Thomas, Velazquez, Laureano Moro, Villalba, Jesus, Qin, Zengyi, Narayanan, Shrikanth, Elhiali, Mounya, Dehak, Najim |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Noise-robust Speech Separation with Fast Generative Correction
by: Wang, Helin, et al.
Published: (2024)
by: Wang, Helin, et al.
Published: (2024)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
by: Wang, Helin, et al.
Published: (2025)
by: Wang, Helin, et al.
Published: (2025)
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
by: Lee, Junhyeok, et al.
Published: (2026)
by: Lee, Junhyeok, et al.
Published: (2026)
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
by: Thebaud, Thomas, et al.
Published: (2026)
by: Thebaud, Thomas, et al.
Published: (2026)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
by: Lee, Junhyeok, et al.
Published: (2025)
by: Lee, Junhyeok, et al.
Published: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
by: Wang, Helin, et al.
Published: (2024)
by: Wang, Helin, et al.
Published: (2024)
DreamVoice: Text-Guided Voice Conversion
by: Hai, Jiarui, et al.
Published: (2024)
by: Hai, Jiarui, et al.
Published: (2024)
Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
by: Chaparala, Kaavya, et al.
Published: (2026)
by: Chaparala, Kaavya, et al.
Published: (2026)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
by: Lu, Yen-Ju, et al.
Published: (2024)
by: Lu, Yen-Ju, et al.
Published: (2024)
DECAF: Dynamic Envelope Context-Aware Fusion for Speech-Envelope Reconstruction from EEG
by: Thakkar, Karan, et al.
Published: (2026)
by: Thakkar, Karan, et al.
Published: (2026)
Demographic Attributes Prediction from Speech Using WavLM Embeddings
by: Yang, Yuchen, et al.
Published: (2025)
by: Yang, Yuchen, et al.
Published: (2025)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
by: Cao, Tianyu, et al.
Published: (2026)
by: Cao, Tianyu, et al.
Published: (2026)
Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings
by: Laouedj, Sarah, et al.
Published: (2025)
by: Laouedj, Sarah, et al.
Published: (2025)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Cross-Corpus and Cross-domain Handwriting Assessment of NeuroDegenerative Diseases via Time-Series-to-Image Conversion
by: Chavez, Gabrielle, et al.
Published: (2025)
by: Chavez, Gabrielle, et al.
Published: (2025)
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
by: Fortier, Alexandrine, et al.
Published: (2025)
by: Fortier, Alexandrine, et al.
Published: (2025)
ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
by: Frummer, Ari, et al.
Published: (2025)
by: Frummer, Ari, et al.
Published: (2025)
Unraveling Adversarial Examples against Speaker Identification -- Techniques for Attack Detection and Victim Model Classification
by: Joshi, Sonal, et al.
Published: (2024)
by: Joshi, Sonal, et al.
Published: (2024)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
by: Feng, Tiantian, et al.
Published: (2023)
by: Feng, Tiantian, et al.
Published: (2023)
Dynamics of Handwriting for Cognitive Assessment
by: Gabrielle Chavez, et al.
Published: (2024)
by: Gabrielle Chavez, et al.
Published: (2024)
Analyzing Attention Focus in the Cookie TheftPicture Description Task Using Word Alignment
by: Anna Favaro, et al.
Published: (2024)
by: Anna Favaro, et al.
Published: (2024)
Cognitive Assessment through Writing Tasks
by: Casey Chen, et al.
Published: (2024)
by: Casey Chen, et al.
Published: (2024)
Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems
by: Joshi, Sonal, et al.
Published: (2021)
by: Joshi, Sonal, et al.
Published: (2021)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
by: Wang, Helin, et al.
Published: (2024)
by: Wang, Helin, et al.
Published: (2024)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
by: Xu, Anfeng, et al.
Published: (2026)
by: Xu, Anfeng, et al.
Published: (2026)
Interpretable Features for the Assessment of Neurodegenerative Diseases through Handwriting Analysis
by: Thebaud, Thomas, et al.
Published: (2024)
by: Thebaud, Thomas, et al.
Published: (2024)
Multimodal characterization of Alzheimer's Disease using speech, eye movement, and handwriting
by: Laureano Moro‐Velazquez, et al.
Published: (2024)
by: Laureano Moro‐Velazquez, et al.
Published: (2024)
Multimodal characterization of Alzheimer’s Disease using speech, eye movement, and handwriting
by: Laureano Moro‐Velazquez, et al.
Published: (2024)
by: Laureano Moro‐Velazquez, et al.
Published: (2024)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Multi-Target Backdoor Attacks Against Speaker Recognition
by: Fortier, Alexandrine, et al.
Published: (2025)
by: Fortier, Alexandrine, et al.
Published: (2025)
Adversarial Attacks and Defenses for Speech Recognition Systems
by: Żelasko, Piotr, et al.
Published: (2021)
by: Żelasko, Piotr, et al.
Published: (2021)
Time Scale Network: A Shallow Neural Network For Time Series Data
by: Meyer, Trevor, et al.
Published: (2023)
by: Meyer, Trevor, et al.
Published: (2023)
Clean Label Attacks against SLU Systems
by: Xinyuan, Henry Li, et al.
Published: (2024)
by: Xinyuan, Henry Li, et al.
Published: (2024)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
by: Lertpetchpun, Thanathai, et al.
Published: (2025)
by: Lertpetchpun, Thanathai, et al.
Published: (2025)
SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
by: Huang, Chien-yu, et al.
Published: (2024)
by: Huang, Chien-yu, et al.
Published: (2024)
FlexSED: Towards Open-Vocabulary Sound Event Detection
by: Hai, Jiarui, et al.
Published: (2025)
by: Hai, Jiarui, et al.
Published: (2025)
Similar Items
-
Noise-robust Speech Separation with Fast Generative Correction
by: Wang, Helin, et al.
Published: (2024) -
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
by: Wang, Helin, et al.
Published: (2025) -
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
by: Lee, Junhyeok, et al.
Published: (2026) -
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
by: Thebaud, Thomas, et al.
Published: (2026) -
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
by: Lee, Junhyeok, et al.
Published: (2025)