SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Xiaoyu, Yang, Yifan, Jin, Zengrui, Cui, Ziyun, Wu, Wen, Li, Baoxiang, Zhang, Chao, Woodland, Phil |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
di: Yang, Xiaoyu, et al.
Pubblicazione: (2024)
di: Yang, Xiaoyu, et al.
Pubblicazione: (2024)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
di: Yang, Yifan, et al.
Pubblicazione: (2024)
di: Yang, Yifan, et al.
Pubblicazione: (2024)
Bayesian Speech Synthesizers Can Learn from Multiple Teachers
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
di: Wang, Mengqi, et al.
Pubblicazione: (2025)
di: Wang, Mengqi, et al.
Pubblicazione: (2025)
Multiplexing Neural Audio Watermarks
di: Yuan, Zheqi, et al.
Pubblicazione: (2025)
di: Yuan, Zheqi, et al.
Pubblicazione: (2025)
ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals
di: E, Ameenudeen P, et al.
Pubblicazione: (2026)
di: E, Ameenudeen P, et al.
Pubblicazione: (2026)
Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
di: Jia, Kaimeng, et al.
Pubblicazione: (2025)
di: Jia, Kaimeng, et al.
Pubblicazione: (2025)
Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
di: Yang, Hsiang-Cheng, et al.
Pubblicazione: (2026)
di: Yang, Hsiang-Cheng, et al.
Pubblicazione: (2026)
Speaker Anonymisation for Speech-based Suicide Risk Detection
di: Cui, Ziyun, et al.
Pubblicazione: (2025)
di: Cui, Ziyun, et al.
Pubblicazione: (2025)
Towards Cross-Task Suicide Risk Detection via Speech LLM
di: Li, Jialun, et al.
Pubblicazione: (2025)
di: Li, Jialun, et al.
Pubblicazione: (2025)
ClariCodec: Optimising Neural Speech Codes for 200bps Communication using Reinforcement Learning
di: Wang, Junyi, et al.
Pubblicazione: (2026)
di: Wang, Junyi, et al.
Pubblicazione: (2026)
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2023)
di: Deng, Keqi, et al.
Pubblicazione: (2023)
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
di: Yang, Yifan, et al.
Pubblicazione: (2026)
di: Yang, Yifan, et al.
Pubblicazione: (2026)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
di: Li, Guinan, et al.
Pubblicazione: (2024)
di: Li, Guinan, et al.
Pubblicazione: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
Can Audio Large Language Models Verify Speaker Identity?
di: Ren, Yiming, et al.
Pubblicazione: (2025)
di: Ren, Yiming, et al.
Pubblicazione: (2025)
The 1st SpeechWellness Challenge: Detecting Suicide Risk Among Adolescents
di: Wu, Wen, et al.
Pubblicazione: (2025)
di: Wu, Wen, et al.
Pubblicazione: (2025)
DNCASR: End-to-End Training for Speaker-Attributed ASR
di: Zheng, Xianrui, et al.
Pubblicazione: (2025)
di: Zheng, Xianrui, et al.
Pubblicazione: (2025)
Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
di: Wu, Wen, et al.
Pubblicazione: (2023)
di: Wu, Wen, et al.
Pubblicazione: (2023)
Distribution-based Emotion Recognition in Conversation
di: Wu, Wen, et al.
Pubblicazione: (2022)
di: Wu, Wen, et al.
Pubblicazione: (2022)
Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
di: Xiao, Yang, et al.
Pubblicazione: (2026)
di: Xiao, Yang, et al.
Pubblicazione: (2026)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
di: Yang, Mu, et al.
Pubblicazione: (2025)
di: Yang, Mu, et al.
Pubblicazione: (2025)
Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
di: N, Rishith Sadashiv T, et al.
Pubblicazione: (2025)
di: N, Rishith Sadashiv T, et al.
Pubblicazione: (2025)
SOT Triggered Neural Clustering for Speaker Attributed ASR
di: Zheng, Xianrui, et al.
Pubblicazione: (2024)
di: Zheng, Xianrui, et al.
Pubblicazione: (2024)
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
di: Lashkarashvili, Nineli, et al.
Pubblicazione: (2024)
di: Lashkarashvili, Nineli, et al.
Pubblicazione: (2024)
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection
di: Kheir, Yassine El, et al.
Pubblicazione: (2025)
di: Kheir, Yassine El, et al.
Pubblicazione: (2025)
Exploring SSL Discrete Tokens for Multilingual ASR
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Non-Causal to Causal SSL-Supported Transfer Learning: Towards a High-Performance Low-Latency Speech Vocoder
di: Shi, Renzheng, et al.
Pubblicazione: (2024)
di: Shi, Renzheng, et al.
Pubblicazione: (2024)
Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
di: Cui, Ziyun, et al.
Pubblicazione: (2024)
di: Cui, Ziyun, et al.
Pubblicazione: (2024)
Unified Audio Event Detection
di: Jiang, Yidi, et al.
Pubblicazione: (2024)
di: Jiang, Yidi, et al.
Pubblicazione: (2024)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
di: Cui, Ziyun, et al.
Pubblicazione: (2023)
di: Cui, Ziyun, et al.
Pubblicazione: (2023)
Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
di: Hsieh, Tsun-An, et al.
Pubblicazione: (2024)
di: Hsieh, Tsun-An, et al.
Pubblicazione: (2024)
Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
di: Li, Zhaoqing, et al.
Pubblicazione: (2025)
di: Li, Zhaoqing, et al.
Pubblicazione: (2025)
Mitigating Language Mismatch in SSL-Based Speaker Anonymization
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
di: Han, Bing, et al.
Pubblicazione: (2026)
di: Han, Bing, et al.
Pubblicazione: (2026)
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
di: Yang, Xiaoyu, et al.
Pubblicazione: (2024) -
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
di: Yang, Yifan, et al.
Pubblicazione: (2024) -
Bayesian Speech Synthesizers Can Learn from Multiple Teachers
di: Zhang, Ziyang, et al.
Pubblicazione: (2025) -
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
di: Wang, Mengqi, et al.
Pubblicazione: (2025) -
Multiplexing Neural Audio Watermarks
di: Yuan, Zheqi, et al.
Pubblicazione: (2025)