DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Tzu-Quan, Lee, Hung-yi, Tang, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Property Neurons in Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
MelHuBERT: A simplified HuBERT on Mel spectrograms
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
von: Shah, Neil, et al.
Veröffentlicht: (2024)
von: Shah, Neil, et al.
Veröffentlicht: (2024)
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
von: Medin, Lucas Block, et al.
Veröffentlicht: (2025)
von: Medin, Lucas Block, et al.
Veröffentlicht: (2025)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
von: Luong, Justin, et al.
Veröffentlicht: (2025)
von: Luong, Justin, et al.
Veröffentlicht: (2025)
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
von: Bradshaw, Louis, et al.
Veröffentlicht: (2025)
von: Bradshaw, Louis, et al.
Veröffentlicht: (2025)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
von: Zhou, Yizhi, et al.
Veröffentlicht: (2025)
von: Zhou, Yizhi, et al.
Veröffentlicht: (2025)
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2026)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2026)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
von: Su, Hsuan, et al.
Veröffentlicht: (2024)
von: Su, Hsuan, et al.
Veröffentlicht: (2024)
Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning
von: Wu, Haibin, et al.
Veröffentlicht: (2021)
von: Wu, Haibin, et al.
Veröffentlicht: (2021)
SUTA-LM: Bridging Test-Time Adaptation and Language Model Rescoring for Robust ASR
von: Huang, Wei-Ping, et al.
Veröffentlicht: (2025)
von: Huang, Wei-Ping, et al.
Veröffentlicht: (2025)
Improving Speech Inversion Through Self-Supervised Embeddings and Enhanced Tract Variables
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
Knowing When to Quit: Probabilistic Early Exits for Speech Separation
von: Olsen, Kenny Falkær, et al.
Veröffentlicht: (2025)
von: Olsen, Kenny Falkær, et al.
Veröffentlicht: (2025)
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2026)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2026)
SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification
von: Lepage, Theo, et al.
Veröffentlicht: (2025)
von: Lepage, Theo, et al.
Veröffentlicht: (2025)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
von: Hyeon, Jonghwan, et al.
Veröffentlicht: (2024)
von: Hyeon, Jonghwan, et al.
Veröffentlicht: (2024)
Adaptive Duration Model for Text Speech Alignment
von: Cao, Junjie
Veröffentlicht: (2025)
von: Cao, Junjie
Veröffentlicht: (2025)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
How Contrastive Decoding Enhances Large Audio Language Models?
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2026)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2026)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Property Neurons in Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024) -
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025) -
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022) -
MelHuBERT: A simplified HuBERT on Mel spectrograms
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022) -
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)