UniSE: A Unified Framework for Decoder-only Autoregressive LM-based Speech Enhancement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Haoyin, Liu, Chengwei, Xue, Shaofei, Liang, Xiaotao, Xue, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Hybrid Discriminative and Generative System for Universal Speech Enhancement
von: Liu, Yinghao, et al.
Veröffentlicht: (2026)
von: Liu, Yinghao, et al.
Veröffentlicht: (2026)
UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens
von: Liu, Chengwei, et al.
Veröffentlicht: (2025)
von: Liu, Chengwei, et al.
Veröffentlicht: (2025)
QuarkAudio Technical Report
von: Liu, Chengwei, et al.
Veröffentlicht: (2025)
von: Liu, Chengwei, et al.
Veröffentlicht: (2025)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
CaSNet: Compress-and-Send Network Based Multi-Device Speech Enhancement Model for Distributed Microphone Arrays
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement
von: Sajid, M., et al.
Veröffentlicht: (2025)
von: Sajid, M., et al.
Veröffentlicht: (2025)
Decoding Order Matters in Autoregressive Speech Synthesis
von: Zhao, Minghui, et al.
Veröffentlicht: (2026)
von: Zhao, Minghui, et al.
Veröffentlicht: (2026)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
von: Ma, Chengyuan, et al.
Veröffentlicht: (2026)
von: Ma, Chengyuan, et al.
Veröffentlicht: (2026)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
Synaspot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy
von: Li, Kewei, et al.
Veröffentlicht: (2025)
von: Li, Kewei, et al.
Veröffentlicht: (2025)
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
von: Xue, Jun, et al.
Veröffentlicht: (2026)
von: Xue, Jun, et al.
Veröffentlicht: (2026)
ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2025)
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2025)
Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
von: Han, Yichen, et al.
Veröffentlicht: (2025)
von: Han, Yichen, et al.
Veröffentlicht: (2025)
ROSE: A Recognition-Oriented Speech Enhancement Framework in Air Traffic Control Using Multi-Objective Learning
von: Yu, Xincheng, et al.
Veröffentlicht: (2023)
von: Yu, Xincheng, et al.
Veröffentlicht: (2023)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors
von: Bao, Guangyin, et al.
Veröffentlicht: (2026)
von: Bao, Guangyin, et al.
Veröffentlicht: (2026)
PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
von: Xiao, Yujia, et al.
Veröffentlicht: (2025)
von: Xiao, Yujia, et al.
Veröffentlicht: (2025)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
Unifying EEG and Speech for Emotion Recognition: A Two-Step Joint Learning Framework for Handling Missing EEG Data During Inference
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
von: Lee, Jung-Sun, et al.
Veröffentlicht: (2024)
von: Lee, Jung-Sun, et al.
Veröffentlicht: (2024)
Speech Enhancement Based on Drifting Models
von: Xu, Liang, et al.
Veröffentlicht: (2026)
von: Xu, Liang, et al.
Veröffentlicht: (2026)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation
von: Cheng, Yuqing, et al.
Veröffentlicht: (2026)
von: Cheng, Yuqing, et al.
Veröffentlicht: (2026)
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2026)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2026)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training
von: Zhou, Naisong, et al.
Veröffentlicht: (2025)
von: Zhou, Naisong, et al.
Veröffentlicht: (2025)
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model
von: Peng, Shuhai, et al.
Veröffentlicht: (2026)
von: Peng, Shuhai, et al.
Veröffentlicht: (2026)
An Investigation of Incorporating Mamba for Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2024)
von: Chao, Rong, et al.
Veröffentlicht: (2024)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Hybrid Discriminative and Generative System for Universal Speech Enhancement
von: Liu, Yinghao, et al.
Veröffentlicht: (2026) -
UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens
von: Liu, Chengwei, et al.
Veröffentlicht: (2025) -
QuarkAudio Technical Report
von: Liu, Chengwei, et al.
Veröffentlicht: (2025) -
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025) -
CaSNet: Compress-and-Send Network Based Multi-Device Speech Enhancement Model for Distributed Microphone Arrays
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)