Gespeichert in:
| Hauptverfasser: | Cheng, Changhao, Wang, Wei, Zhang, Wangyou, Jia, Dongya, Wu, Jian, Chen, Zhuo, Qian, Yanmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2604.12383 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
Improving Design of Input Condition Invariant Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
von: Wang, Wei, et al.
Veröffentlicht: (2025)
von: Wang, Wei, et al.
Veröffentlicht: (2025)
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
von: Han, Bing, et al.
Veröffentlicht: (2026)
von: Han, Bing, et al.
Veröffentlicht: (2026)
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
UrgentMOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality Assessment
von: Wang, Wei, et al.
Veröffentlicht: (2026)
von: Wang, Wei, et al.
Veröffentlicht: (2026)
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
von: Jia, Dongya, et al.
Veröffentlicht: (2025)
von: Jia, Dongya, et al.
Veröffentlicht: (2025)
ICASSP 2026 URGENT Speech Enhancement Challenge
von: Li, Chenda, et al.
Veröffentlicht: (2026)
von: Li, Chenda, et al.
Veröffentlicht: (2026)
Training Text-to-Speech Model with Purely Synthetic Data: Feasibility, Sensitivity, and Generalization Capability
von: Zhou, Tingxiao, et al.
Veröffentlicht: (2025)
von: Zhou, Tingxiao, et al.
Veröffentlicht: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
SLM-SS: Speech Language Model for Generative Speech Separation
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
von: Song, Yakun, et al.
Veröffentlicht: (2025)
von: Song, Yakun, et al.
Veröffentlicht: (2025)
Cross-Utterance Conditioned VAE for Speech Generation
von: Li, Yang, et al.
Veröffentlicht: (2023)
von: Li, Yang, et al.
Veröffentlicht: (2023)
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
Less is More: Data Curation Matters in Scaling Speech Enhancement
von: Li, Chenda, et al.
Veröffentlicht: (2025)
von: Li, Chenda, et al.
Veröffentlicht: (2025)
URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
Localizing Speech Deepfakes Beyond Transitions via Segment-Aware Learning
von: Mao, Yuchen, et al.
Veröffentlicht: (2026)
von: Mao, Yuchen, et al.
Veröffentlicht: (2026)
URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
From Sharpness to Better Generalization for Speech Deepfake Detection
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
von: Li, Chenda, et al.
Veröffentlicht: (2024)
von: Li, Chenda, et al.
Veröffentlicht: (2024)
A Data-Centric Approach to Generalizable Speech Deepfake Detection
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
Lightweight Front-end Enhancement for Robust ASR via Frame Resampling and Sub-Band Pruning
von: Zhao, Siyi, et al.
Veröffentlicht: (2025)
von: Zhao, Siyi, et al.
Veröffentlicht: (2025)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
von: Han, Bing, et al.
Veröffentlicht: (2024)
von: Han, Bing, et al.
Veröffentlicht: (2024)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation
von: Wang, Wei, et al.
Veröffentlicht: (2025)
von: Wang, Wei, et al.
Veröffentlicht: (2025)
DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
von: Zhang, Leying, et al.
Veröffentlicht: (2026)
von: Zhang, Leying, et al.
Veröffentlicht: (2026)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
von: Gong, Xun, et al.
Veröffentlicht: (2024)
von: Gong, Xun, et al.
Veröffentlicht: (2024)
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
von: Zhang, Leying, et al.
Veröffentlicht: (2026)
von: Zhang, Leying, et al.
Veröffentlicht: (2026)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025) -
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
von: Zhang, Leying, et al.
Veröffentlicht: (2024) -
Improving Design of Input Condition Invariant Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024) -
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025) -
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
von: Wang, Wei, et al.
Veröffentlicht: (2025)