STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Jiajun, Wang, Xiaochen, Xiao, Yuhang, Wu, Yulin, Hu, Chenhao, Lv, Xueyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
Transient Noise Removal via Diffusion-based Speech Inpainting
von: Moradi, Mordehay, et al.
Veröffentlicht: (2025)
von: Moradi, Mordehay, et al.
Veröffentlicht: (2025)
A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
von: Xin, Detai, et al.
Veröffentlicht: (2026)
von: Xin, Detai, et al.
Veröffentlicht: (2026)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Long-Context Speech Synthesis with Context-Aware Memory
von: Li, Zhipeng, et al.
Veröffentlicht: (2025)
von: Li, Zhipeng, et al.
Veröffentlicht: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
von: Song, Yulin, et al.
Veröffentlicht: (2024)
von: Song, Yulin, et al.
Veröffentlicht: (2024)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
Spectral Masking with Explicit Time-Context Windowing for Neural Network-Based Monaural Speech Enhancement
von: Fiorio, Luan Vinícius, et al.
Veröffentlicht: (2024)
von: Fiorio, Luan Vinícius, et al.
Veröffentlicht: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
Combined Generative and Predictive Modeling for Speech Super-resolution
von: Wang, Heming, et al.
Veröffentlicht: (2024)
von: Wang, Heming, et al.
Veröffentlicht: (2024)
High-Fidelity Generative Audio Compression at 0.275kbps
von: Ma, Hao, et al.
Veröffentlicht: (2026)
von: Ma, Hao, et al.
Veröffentlicht: (2026)
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024)
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024)
Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
von: Song, Tianyu, et al.
Veröffentlicht: (2025)
von: Song, Tianyu, et al.
Veröffentlicht: (2025)
InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
von: Zeng, Chang, et al.
Veröffentlicht: (2024)
von: Zeng, Chang, et al.
Veröffentlicht: (2024)
High-Fidelity Neural Phonetic Posteriorgrams
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
von: Nawfal, Ismael, et al.
Veröffentlicht: (2025)
von: Nawfal, Ismael, et al.
Veröffentlicht: (2025)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
CogSR: Semantic-Aware Speech Super-Resolution via Chain-of-Thought Guided Flow Matching
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
von: Ma, Ding, et al.
Veröffentlicht: (2026)
von: Ma, Ding, et al.
Veröffentlicht: (2026)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
Phase Repair for Time-Domain Convolutional Neural Networks in Music Super-Resolution
von: Zhang, Yenan, et al.
Veröffentlicht: (2023)
von: Zhang, Yenan, et al.
Veröffentlicht: (2023)
DENSE: Dynamic Embedding Causal Target Speech Extraction
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
HILCodec: High-Fidelity and Lightweight Neural Audio Codec
von: Ahn, Sunghwan, et al.
Veröffentlicht: (2024)
von: Ahn, Sunghwan, et al.
Veröffentlicht: (2024)
Vision-Integrated High-Quality Neural Speech Coding
von: Guo, Yao, et al.
Veröffentlicht: (2025)
von: Guo, Yao, et al.
Veröffentlicht: (2025)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025) -
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025) -
Transient Noise Removal via Diffusion-based Speech Inpainting
von: Moradi, Mordehay, et al.
Veröffentlicht: (2025) -
A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025) -
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
von: Xin, Detai, et al.
Veröffentlicht: (2026)