Pruning-aware Loss Functions for STOI-Optimized Pruned Recurrent Autoencoders for the Compression of the Stimulation Patterns of Cochlear Implants at Zero Delay
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hinrichs, Reemt, Ostermann, Jörn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scalable Speech Enhancement with Dynamic Channel Pruning
von: Miccini, Riccardo, et al.
Veröffentlicht: (2024)
von: Miccini, Riccardo, et al.
Veröffentlicht: (2024)
OnDA: On-device Channel Pruning for Efficient Personalized Keyword Spotting
von: Risso, Matteo, et al.
Veröffentlicht: (2026)
von: Risso, Matteo, et al.
Veröffentlicht: (2026)
Revisit Micro-batch Clipping: Adaptive Data Pruning via Gradient Manipulation
von: Wang, Lun
Veröffentlicht: (2024)
von: Wang, Lun
Veröffentlicht: (2024)
Music2Latent: Consistency Autoencoders for Latent Audio Compression
von: Pasini, Marco, et al.
Veröffentlicht: (2024)
von: Pasini, Marco, et al.
Veröffentlicht: (2024)
From Diet to Free Lunch: Estimating Auxiliary Signal Properties using Dynamic Pruning Masks in Speech Enhancement Networks
von: Miccini, Riccardo, et al.
Veröffentlicht: (2026)
von: Miccini, Riccardo, et al.
Veröffentlicht: (2026)
FLToP CTC: Frame-Level Token Pruning via Relative Threshold for Efficient and Memory-Saving Decoding on Diverse Platforms
von: Shree, Atul, et al.
Veröffentlicht: (2025)
von: Shree, Atul, et al.
Veröffentlicht: (2025)
Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions
von: Garcia, Elliot Q C, et al.
Veröffentlicht: (2025)
von: Garcia, Elliot Q C, et al.
Veröffentlicht: (2025)
Can Masked Autoencoders Also Listen to Birds?
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
Data-Driven Room Acoustic Modeling Via Differentiable Feedback Delay Networks With Learnable Delay Lines
von: Mezza, Alessandro Ilic, et al.
Veröffentlicht: (2024)
von: Mezza, Alessandro Ilic, et al.
Veröffentlicht: (2024)
Re-Bottleneck: Latent Re-Structuring for Neural Audio Autoencoders
von: Bralios, Dimitrios, et al.
Veröffentlicht: (2025)
von: Bralios, Dimitrios, et al.
Veröffentlicht: (2025)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
wav2pos: Sound Source Localization using Masked Autoencoders
von: Berg, Axel, et al.
Veröffentlicht: (2024)
von: Berg, Axel, et al.
Veröffentlicht: (2024)
Music Emotion Prediction Using Recurrent Neural Networks
von: Chang, Xinyu, et al.
Veröffentlicht: (2024)
von: Chang, Xinyu, et al.
Veröffentlicht: (2024)
Dynamic Gated Recurrent Neural Network for Compute-efficient Speech Enhancement
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks
von: Jafari, Farshad, et al.
Veröffentlicht: (2024)
von: Jafari, Farshad, et al.
Veröffentlicht: (2024)
DEMONet: Underwater Acoustic Target Recognition based on Multi-Expert Network and Cross-Temporal Variational Autoencoder
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Noise-aware Speech Enhancement using Diffusion Probabilistic Model
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
Investigation of Time-Frequency Feature Combinations with Histogram Layer Time Delay Neural Networks
von: Mohammadi, Amirmohammad, et al.
Veröffentlicht: (2024)
von: Mohammadi, Amirmohammad, et al.
Veröffentlicht: (2024)
Zero-shot Voice Conversion with Diffusion Transformers
von: Liu, Songting
Veröffentlicht: (2024)
von: Liu, Songting
Veröffentlicht: (2024)
Zero-Shot Mono-to-Binaural Speech Synthesis
von: Levkovitch, Alon, et al.
Veröffentlicht: (2024)
von: Levkovitch, Alon, et al.
Veröffentlicht: (2024)
Synthetic data enables context-aware bioacoustic sound event detection
von: Hoffman, Benjamin, et al.
Veröffentlicht: (2025)
von: Hoffman, Benjamin, et al.
Veröffentlicht: (2025)
Recurrence-Based Nonlinear Vocal Dynamics as Digital Biomarkers for Depression Detection from Conversational Speech
von: Samanta, Himadri S
Veröffentlicht: (2026)
von: Samanta, Himadri S
Veröffentlicht: (2026)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
Context-aware child-directed speech detection from long-form recordings
von: Charlot, Théo, et al.
Veröffentlicht: (2026)
von: Charlot, Théo, et al.
Veröffentlicht: (2026)
Transcribing Rhythmic Patterns of the Guitar Track in Polyphonic Music
von: Lukoianov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Lukoianov, Aleksandr, et al.
Veröffentlicht: (2025)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers
von: Biju, Emil, et al.
Veröffentlicht: (2024)
von: Biju, Emil, et al.
Veröffentlicht: (2024)
Zero-Shot Multi-Lingual Speaker Verification in Clinical Trials
von: Akram, Ali, et al.
Veröffentlicht: (2024)
von: Akram, Ali, et al.
Veröffentlicht: (2024)
Embedding-Space Diffusion for Zero-Shot Environmental Sound Classification
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
Information Retrieval for ZeroSpeech 2021: The Submission by University of Wroclaw
von: Chorowski, Jan, et al.
Veröffentlicht: (2021)
von: Chorowski, Jan, et al.
Veröffentlicht: (2021)
Multi-label Zero-Shot Audio Classification with Temporal Attention
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
von: Janiczek, John, et al.
Veröffentlicht: (2024)
von: Janiczek, John, et al.
Veröffentlicht: (2024)
Audio Processing using Pattern Recognition for Music Genre Classification
von: Chatterjee, Sivangi, et al.
Veröffentlicht: (2024)
von: Chatterjee, Sivangi, et al.
Veröffentlicht: (2024)
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scalable Speech Enhancement with Dynamic Channel Pruning
von: Miccini, Riccardo, et al.
Veröffentlicht: (2024) -
OnDA: On-device Channel Pruning for Efficient Personalized Keyword Spotting
von: Risso, Matteo, et al.
Veröffentlicht: (2026) -
Revisit Micro-batch Clipping: Adaptive Data Pruning via Gradient Manipulation
von: Wang, Lun
Veröffentlicht: (2024) -
Music2Latent: Consistency Autoencoders for Latent Audio Compression
von: Pasini, Marco, et al.
Veröffentlicht: (2024) -
From Diet to Free Lunch: Estimating Auxiliary Signal Properties using Dynamic Pruning Masks in Speech Enhancement Networks
von: Miccini, Riccardo, et al.
Veröffentlicht: (2026)