TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Mohan, Li, Kai, Chen, Guo, Hu, Xiaolin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation
von: Chen, Guo, et al.
Veröffentlicht: (2025)
von: Chen, Guo, et al.
Veröffentlicht: (2025)
TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion
von: Pegg, Samuel, et al.
Veröffentlicht: (2024)
von: Pegg, Samuel, et al.
Veröffentlicht: (2024)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
SPMamba: State-space model is all you need in speech separation
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Study of the Performance of CEEMDAN in Underdetermined Speech Separation
von: Melhem, Rawad, et al.
Veröffentlicht: (2024)
von: Melhem, Rawad, et al.
Veröffentlicht: (2024)
Advances in Speech Separation: Techniques, Challenges, and Future Trends
von: Li, Kai, et al.
Veröffentlicht: (2025)
von: Li, Kai, et al.
Veröffentlicht: (2025)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
TISDiSS: A Training-Time and Inference-Time Scalable Framework for Discriminative Source Separation
von: Feng, Yongsheng, et al.
Veröffentlicht: (2025)
von: Feng, Yongsheng, et al.
Veröffentlicht: (2025)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
Leveraging Spatial Cues from Cochlear Implant Microphones to Efficiently Enhance Speech Separation in Real-World Listening Scenes
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025)
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025)
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
Music Source Restoration with Ensemble Separation and Targeted Reconstruction
von: Deng, Xinlong, et al.
Veröffentlicht: (2026)
von: Deng, Xinlong, et al.
Veröffentlicht: (2026)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References
von: Jepsen, Simon Dahl, et al.
Veröffentlicht: (2025)
von: Jepsen, Simon Dahl, et al.
Veröffentlicht: (2025)
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis
von: Hu, Xintong, et al.
Veröffentlicht: (2025)
von: Hu, Xintong, et al.
Veröffentlicht: (2025)
VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
von: Choi, Joon-Seung, et al.
Veröffentlicht: (2025)
von: Choi, Joon-Seung, et al.
Veröffentlicht: (2025)
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
von: Li, Bohan, et al.
Veröffentlicht: (2026)
von: Li, Bohan, et al.
Veröffentlicht: (2026)
LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
von: Chou, Benjamin Shiue-Hal, et al.
Veröffentlicht: (2025)
von: Chou, Benjamin Shiue-Hal, et al.
Veröffentlicht: (2025)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
von: Hu, Guoqiang, et al.
Veröffentlicht: (2024)
von: Hu, Guoqiang, et al.
Veröffentlicht: (2024)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
Temporal Information Reconstruction and Non-Aligned Residual in Spiking Neural Networks for Speech Classification
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
Speech-based Clinical Depression Screening: An Empirical Study
von: Chen, Yangbin, et al.
Veröffentlicht: (2024)
von: Chen, Yangbin, et al.
Veröffentlicht: (2024)
One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation
von: Chen, Guo, et al.
Veröffentlicht: (2025) -
TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion
von: Pegg, Samuel, et al.
Veröffentlicht: (2024) -
SepPrune: Structured Pruning for Efficient Deep Speech Separation
von: Li, Yuqi, et al.
Veröffentlicht: (2025) -
SPMamba: State-space model is all you need in speech separation
von: Li, Kai, et al.
Veröffentlicht: (2024) -
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
von: Wang, Helin, et al.
Veröffentlicht: (2025)