Rethinking Mamba in Speech Processing by Self-Supervised Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xiangyu, Ma, Jianbo, Shahin, Mostafa, Ahmed, Beena, Epps, Julien |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
von: Shahin, Mostafa, et al.
Veröffentlicht: (2025)
von: Shahin, Mostafa, et al.
Veröffentlicht: (2025)
Why Pre-trained Models Fail: Feature Entanglement in Multi-modal Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
Mamba in Speech: Towards an Alternative to Self-Attention
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2026)
Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation
von: Ma, Jianbo, et al.
Veröffentlicht: (2026)
von: Ma, Jianbo, et al.
Veröffentlicht: (2026)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
von: C, Shiva Kumar, et al.
Veröffentlicht: (2025)
von: C, Shiva Kumar, et al.
Veröffentlicht: (2025)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
von: Guo, Xin, et al.
Veröffentlicht: (2026)
von: Guo, Xin, et al.
Veröffentlicht: (2026)
Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
von: Wang, Dahan, et al.
Veröffentlicht: (2026)
von: Wang, Dahan, et al.
Veröffentlicht: (2026)
Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks
von: Ma, Duo, et al.
Veröffentlicht: (2024)
von: Ma, Duo, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
von: Ochiai, Tsubasa, et al.
Veröffentlicht: (2024)
von: Ochiai, Tsubasa, et al.
Veröffentlicht: (2024)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
von: Hanilçi, Cemal, et al.
Veröffentlicht: (2026)
von: Hanilçi, Cemal, et al.
Veröffentlicht: (2026)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
Selective State Space Model for Monaural Speech Enhancement
von: Chen, Moran, et al.
Veröffentlicht: (2024)
von: Chen, Moran, et al.
Veröffentlicht: (2024)
SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
von: Shams, Siavash, et al.
Veröffentlicht: (2024)
von: Shams, Siavash, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation
von: Lai, Richard Lee, et al.
Veröffentlicht: (2023)
von: Lai, Richard Lee, et al.
Veröffentlicht: (2023)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
von: Terashima, Ryo, et al.
Veröffentlicht: (2025)
von: Terashima, Ryo, et al.
Veröffentlicht: (2025)
Linear-Complexity Self-Supervised Learning for Speech Processing
von: Zhang, Shucong, et al.
Veröffentlicht: (2024)
von: Zhang, Shucong, et al.
Veröffentlicht: (2024)
Long-Context Modeling Networks for Monaural Speech Enhancement: A Comparative Study
von: Zhang, Qiquan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiquan, et al.
Veröffentlicht: (2025)
Exploring the Capability of Mamba in Speech Applications
von: Miyazaki, Koichi, et al.
Veröffentlicht: (2024)
von: Miyazaki, Koichi, et al.
Veröffentlicht: (2024)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2025)
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2025)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
von: Shahin, Mostafa, et al.
Veröffentlicht: (2025) -
Why Pre-trained Models Fail: Feature Entanglement in Multi-modal Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025) -
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025) -
Mamba in Speech: Towards an Alternative to Self-Attention
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024) -
Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)