LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Kwak, Doyeop, Jang, Youngjoon, Chung, Joon Son |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
por: Kwak, Doyeop, et al.
Publicado: (2025)
por: Kwak, Doyeop, et al.
Publicado: (2025)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
por: Kwak, Doyeop, et al.
Publicado: (2026)
por: Kwak, Doyeop, et al.
Publicado: (2026)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
por: Jung, Jaemin, et al.
Publicado: (2024)
por: Jung, Jaemin, et al.
Publicado: (2024)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
por: Jung, Chaeyoung, et al.
Publicado: (2024)
por: Jung, Chaeyoung, et al.
Publicado: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
por: Kwak, Doyeop, et al.
Publicado: (2026)
por: Kwak, Doyeop, et al.
Publicado: (2026)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
por: Jung, Jihoo, et al.
Publicado: (2026)
por: Jung, Jihoo, et al.
Publicado: (2026)
VoxSim: A perceptual voice similarity dataset
por: Ahn, Junseok, et al.
Publicado: (2024)
por: Ahn, Junseok, et al.
Publicado: (2024)
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
por: Jang, Youngjoon, et al.
Publicado: (2024)
por: Jang, Youngjoon, et al.
Publicado: (2024)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
por: Nguyen, Tan Dat, et al.
Publicado: (2024)
por: Nguyen, Tan Dat, et al.
Publicado: (2024)
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
por: Kim, Ji-Hoon, et al.
Publicado: (2025)
por: Kim, Ji-Hoon, et al.
Publicado: (2025)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
por: Yang, Da-Hee, et al.
Publicado: (2026)
por: Yang, Da-Hee, et al.
Publicado: (2026)
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
por: Nguyen, Tan Dat, et al.
Publicado: (2026)
por: Nguyen, Tan Dat, et al.
Publicado: (2026)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
por: Choi, Jeongsoo, et al.
Publicado: (2025)
por: Choi, Jeongsoo, et al.
Publicado: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
por: Pham, The Hieu, et al.
Publicado: (2025)
por: Pham, The Hieu, et al.
Publicado: (2025)
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
por: Huynh-Nguyen, Hieu-Nghia, et al.
Publicado: (2025)
por: Huynh-Nguyen, Hieu-Nghia, et al.
Publicado: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
por: Choi, Jeongsoo, et al.
Publicado: (2025)
por: Choi, Jeongsoo, et al.
Publicado: (2025)
SCORE: Scaling audio generation using Standardized COmposite REwards
por: Jung, Jaemin, et al.
Publicado: (2025)
por: Jung, Jaemin, et al.
Publicado: (2025)
Lightweight Audio Segmentation for Long-form Speech Translation
por: Lee, Jaesong, et al.
Publicado: (2024)
por: Lee, Jaesong, et al.
Publicado: (2024)
FlowSE: Flow Matching-based Speech Enhancement
por: Lee, Seonggyu, et al.
Publicado: (2025)
por: Lee, Seonggyu, et al.
Publicado: (2025)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
por: Choi, Jeongsoo, et al.
Publicado: (2024)
por: Choi, Jeongsoo, et al.
Publicado: (2024)
Towards Real-Time Generative Speech Restoration with Flow-Matching
por: Hsieh, Tsun-An, et al.
Publicado: (2025)
por: Hsieh, Tsun-An, et al.
Publicado: (2025)
CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
por: Jiang, Xiao-Hang, et al.
Publicado: (2026)
por: Jiang, Xiao-Hang, et al.
Publicado: (2026)
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
por: Park, Hyun Joon, et al.
Publicado: (2025)
por: Park, Hyun Joon, et al.
Publicado: (2025)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
por: Yang, Jinhyeok, et al.
Publicado: (2026)
por: Yang, Jinhyeok, et al.
Publicado: (2026)
Perceptual Ratings Predict Speech Inversion Articulatory Kinematics in Childhood Speech Sound Disorders
por: Benway, Nina R., et al.
Publicado: (2025)
por: Benway, Nina R., et al.
Publicado: (2025)
FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
por: Wang, Haoxu, et al.
Publicado: (2026)
por: Wang, Haoxu, et al.
Publicado: (2026)
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
por: Das, Shoutrik, et al.
Publicado: (2025)
por: Das, Shoutrik, et al.
Publicado: (2025)
Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
por: Jia, Kaimeng, et al.
Publicado: (2025)
por: Jia, Kaimeng, et al.
Publicado: (2025)
InfiniteAudio: Infinite-Length Audio Generation with Consistency
por: Jung, Chaeyoung, et al.
Publicado: (2025)
por: Jung, Chaeyoung, et al.
Publicado: (2025)
Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
por: Kim, Youkyum, et al.
Publicado: (2024)
por: Kim, Youkyum, et al.
Publicado: (2024)
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
por: Wang, Ziqian, et al.
Publicado: (2025)
por: Wang, Ziqian, et al.
Publicado: (2025)
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
por: Kim, Sungnyun, et al.
Publicado: (2025)
por: Kim, Sungnyun, et al.
Publicado: (2025)
FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech Restoration
por: Yang, Da-Hee, et al.
Publicado: (2025)
por: Yang, Da-Hee, et al.
Publicado: (2025)
Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
por: Ahn, Hyebin, et al.
Publicado: (2025)
por: Ahn, Hyebin, et al.
Publicado: (2025)
Learning to Solve Inverse Problems for Perceptual Sound Matching
por: Han, Han, et al.
Publicado: (2023)
por: Han, Han, et al.
Publicado: (2023)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
por: Zuo, Jialong, et al.
Publicado: (2025)
por: Zuo, Jialong, et al.
Publicado: (2025)
Speech Synthesis along Perceptual Voice Quality Dimensions
por: Rautenberg, Frederik, et al.
Publicado: (2025)
por: Rautenberg, Frederik, et al.
Publicado: (2025)
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
por: Close, George, et al.
Publicado: (2024)
por: Close, George, et al.
Publicado: (2024)
URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
por: Wang, Jiahe, et al.
Publicado: (2025)
por: Wang, Jiahe, et al.
Publicado: (2025)
Ejemplares similares
-
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
por: Kwak, Doyeop, et al.
Publicado: (2025) -
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
por: Kwak, Doyeop, et al.
Publicado: (2026) -
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
por: Jung, Jaemin, et al.
Publicado: (2024) -
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
por: Jung, Chaeyoung, et al.
Publicado: (2024) -
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
por: Kwak, Doyeop, et al.
Publicado: (2026)