Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Shih-Heng, Feng, Tiantian, Kommineni, Aditya, Lertpetchpun, Thanathai, Yi, Bowen, Shi, Xuan, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
von: Prescott, Jordan, et al.
Veröffentlicht: (2026)
von: Prescott, Jordan, et al.
Veröffentlicht: (2026)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2026)
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2026)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2025)
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2025)
Learning-free L2-Accented Speech Generation using Phonological Rules
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2026)
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2026)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2026)
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2026)
Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
von: Trachu, Thanapat, et al.
Veröffentlicht: (2025)
von: Trachu, Thanapat, et al.
Veröffentlicht: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Toward a Sparse and Interpretable Audio Codec
von: Vinyard, John
Veröffentlicht: (2025)
von: Vinyard, John
Veröffentlicht: (2025)
ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
Audio-visual child-adult speaker classification in dyadic interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2023)
von: Xu, Anfeng, et al.
Veröffentlicht: (2023)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
VoxCare: Studying Natural Communication Behaviors of Hospital Caregivers through Wearable Sensing of Egocentric Audio
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks
von: Tsaprazlis, Efthymios, et al.
Veröffentlicht: (2025)
von: Tsaprazlis, Efthymios, et al.
Veröffentlicht: (2025)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024)
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024)
FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
Emotion-Aligned Contrastive Learning Between Images and Music
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
Towards Neural Audio Codec Source Parsing
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
On the Relationship between Accent Strength and Articulatory Features
von: Huang, Kevin, et al.
Veröffentlicht: (2025)
von: Huang, Kevin, et al.
Veröffentlicht: (2025)
LLM-Codec: Neural Audio Codec Meets Language Model Objectives
von: Chung, Ho-Lam, et al.
Veröffentlicht: (2026)
von: Chung, Ho-Lam, et al.
Veröffentlicht: (2026)
Speech Codec Probing from Semantic and Phonetic Perspectives
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
Code Drift: Towards Idempotent Neural Audio Codecs
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
von: Paek, Nathan, et al.
Veröffentlicht: (2025)
von: Paek, Nathan, et al.
Veröffentlicht: (2025)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
von: Nguyen, Hong, et al.
Veröffentlicht: (2024)
von: Nguyen, Hong, et al.
Veröffentlicht: (2024)
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
von: Luo, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Luo, Xiaoxue, et al.
Veröffentlicht: (2025)
Adapting Neural Audio Codecs to EEG
von: Kastrati, Ard, et al.
Veröffentlicht: (2025)
von: Kastrati, Ard, et al.
Veröffentlicht: (2025)
ADNAC: Audio Denoiser using Neural Audio Codec
von: Jimon, Daniel, et al.
Veröffentlicht: (2025)
von: Jimon, Daniel, et al.
Veröffentlicht: (2025)
Soft Disentanglement in Frequency Bands for Neural Audio Codecs
von: Ginies, Benoit, et al.
Veröffentlicht: (2025)
von: Ginies, Benoit, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
von: Prescott, Jordan, et al.
Veröffentlicht: (2026) -
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
von: Feng, Tiantian, et al.
Veröffentlicht: (2025) -
Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data
von: Lertpetchpun, Thanathai, et al.
Veröffentlicht: (2026) -
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
von: Lee, Jihwan, et al.
Veröffentlicht: (2024) -
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)