DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Tianyu, Wang, Helin, Frummer, Ari, Sieradzki, Yuval, Arbel, Adi, Velazquez, Laureano Moro, Villalba, Jesus, Gal, Oren, Thebaud, Thomas, Dehak, Najim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
von: Frummer, Ari, et al.
Veröffentlicht: (2025)
von: Frummer, Ari, et al.
Veröffentlicht: (2025)
Noise-robust Speech Separation with Fast Generative Correction
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
von: Thebaud, Thomas, et al.
Veröffentlicht: (2026)
von: Thebaud, Thomas, et al.
Veröffentlicht: (2026)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
von: Lee, Junhyeok, et al.
Veröffentlicht: (2025)
von: Lee, Junhyeok, et al.
Veröffentlicht: (2025)
Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
von: Chaparala, Kaavya, et al.
Veröffentlicht: (2026)
von: Chaparala, Kaavya, et al.
Veröffentlicht: (2026)
Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings
von: Laouedj, Sarah, et al.
Veröffentlicht: (2025)
von: Laouedj, Sarah, et al.
Veröffentlicht: (2025)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
Cross-Corpus and Cross-domain Handwriting Assessment of NeuroDegenerative Diseases via Time-Series-to-Image Conversion
von: Chavez, Gabrielle, et al.
Veröffentlicht: (2025)
von: Chavez, Gabrielle, et al.
Veröffentlicht: (2025)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
von: Lee, Junhyeok, et al.
Veröffentlicht: (2026)
von: Lee, Junhyeok, et al.
Veröffentlicht: (2026)
Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems
von: Joshi, Sonal, et al.
Veröffentlicht: (2021)
von: Joshi, Sonal, et al.
Veröffentlicht: (2021)
Dynamics of Handwriting for Cognitive Assessment
von: Gabrielle Chavez, et al.
Veröffentlicht: (2024)
von: Gabrielle Chavez, et al.
Veröffentlicht: (2024)
Analyzing Attention Focus in the Cookie TheftPicture Description Task Using Word Alignment
von: Anna Favaro, et al.
Veröffentlicht: (2024)
von: Anna Favaro, et al.
Veröffentlicht: (2024)
Cognitive Assessment through Writing Tasks
von: Casey Chen, et al.
Veröffentlicht: (2024)
von: Casey Chen, et al.
Veröffentlicht: (2024)
Demographic Attributes Prediction from Speech Using WavLM Embeddings
von: Yang, Yuchen, et al.
Veröffentlicht: (2025)
von: Yang, Yuchen, et al.
Veröffentlicht: (2025)
Unraveling Adversarial Examples against Speaker Identification -- Techniques for Attack Detection and Victim Model Classification
von: Joshi, Sonal, et al.
Veröffentlicht: (2024)
von: Joshi, Sonal, et al.
Veröffentlicht: (2024)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
Interpretable Features for the Assessment of Neurodegenerative Diseases through Handwriting Analysis
von: Thebaud, Thomas, et al.
Veröffentlicht: (2024)
von: Thebaud, Thomas, et al.
Veröffentlicht: (2024)
Multimodal characterization of Alzheimer's Disease using speech, eye movement, and handwriting
von: Laureano Moro‐Velazquez, et al.
Veröffentlicht: (2024)
von: Laureano Moro‐Velazquez, et al.
Veröffentlicht: (2024)
Multimodal characterization of Alzheimer’s Disease using speech, eye movement, and handwriting
von: Laureano Moro‐Velazquez, et al.
Veröffentlicht: (2024)
von: Laureano Moro‐Velazquez, et al.
Veröffentlicht: (2024)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
von: Fortier, Alexandrine, et al.
Veröffentlicht: (2025)
von: Fortier, Alexandrine, et al.
Veröffentlicht: (2025)
Time Scale Network: A Shallow Neural Network For Time Series Data
von: Meyer, Trevor, et al.
Veröffentlicht: (2023)
von: Meyer, Trevor, et al.
Veröffentlicht: (2023)
Multimodal Analysis of Behavior During Stroop Test for Characterization of Alzheimer’s Disease Signs
von: Trevor Meyer, et al.
Veröffentlicht: (2024)
von: Trevor Meyer, et al.
Veröffentlicht: (2024)
Multi-Target Backdoor Attacks Against Speaker Recognition
von: Fortier, Alexandrine, et al.
Veröffentlicht: (2025)
von: Fortier, Alexandrine, et al.
Veröffentlicht: (2025)
Clean Label Attacks against SLU Systems
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2024)
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2024)
Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
von: Thebaud, Thomas, et al.
Veröffentlicht: (2025)
von: Thebaud, Thomas, et al.
Veröffentlicht: (2025)
Latent Speech-Text Transformer
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
Adversarial Attacks and Defenses for Speech Recognition Systems
von: Żelasko, Piotr, et al.
Veröffentlicht: (2021)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2021)
FlowSE: Flow Matching-based Speech Enhancement
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement
von: Cross, Mattias, et al.
Veröffentlicht: (2025)
von: Cross, Mattias, et al.
Veröffentlicht: (2025)
DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing
von: Zhou, Zihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zihan, et al.
Veröffentlicht: (2025)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-based DiTs
von: He, Xianglong, et al.
Veröffentlicht: (2025)
von: He, Xianglong, et al.
Veröffentlicht: (2025)
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
von: Frummer, Ari, et al.
Veröffentlicht: (2025) -
Noise-robust Speech Separation with Fast Generative Correction
von: Wang, Helin, et al.
Veröffentlicht: (2024) -
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
von: Thebaud, Thomas, et al.
Veröffentlicht: (2026) -
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
von: Lee, Junhyeok, et al.
Veröffentlicht: (2025) -
Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
von: Chaparala, Kaavya, et al.
Veröffentlicht: (2026)