Gespeichert in:
| Hauptverfasser: | Fish, Edward, Michieli, Umberto, Ozay, Mete |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2307.12659 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
persoDA: Personalized Data Augmentation for Personalized ASR
von: Parada, Pablo Peso, et al.
Veröffentlicht: (2025)
von: Parada, Pablo Peso, et al.
Veröffentlicht: (2025)
ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization
von: Mehmood, Haaris, et al.
Veröffentlicht: (2025)
von: Mehmood, Haaris, et al.
Veröffentlicht: (2025)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
von: Roth, Amit, et al.
Veröffentlicht: (2024)
von: Roth, Amit, et al.
Veröffentlicht: (2024)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
TOGGL: Transcribing Overlapping Speech with Staggered Labeling
von: Li, Chak-Fai, et al.
Veröffentlicht: (2024)
von: Li, Chak-Fai, et al.
Veröffentlicht: (2024)
MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2025)
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Quantizing Whisper-small: How design choices affect ASR performance
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
von: Tang, Yun, et al.
Veröffentlicht: (2025)
von: Tang, Yun, et al.
Veröffentlicht: (2025)
FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
von: Chen, Tanyu, et al.
Veröffentlicht: (2026)
von: Chen, Tanyu, et al.
Veröffentlicht: (2026)
Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2024)
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2024)
Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems
von: Lin, Jhen-Ke, et al.
Veröffentlicht: (2025)
von: Lin, Jhen-Ke, et al.
Veröffentlicht: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
von: Dao, Alan, et al.
Veröffentlicht: (2024)
von: Dao, Alan, et al.
Veröffentlicht: (2024)
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
von: Cheng, Yifan, et al.
Veröffentlicht: (2025)
von: Cheng, Yifan, et al.
Veröffentlicht: (2025)
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
von: Du, Yichao, et al.
Veröffentlicht: (2024)
von: Du, Yichao, et al.
Veröffentlicht: (2024)
Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
von: Rangappa, Pradeep, et al.
Veröffentlicht: (2025)
von: Rangappa, Pradeep, et al.
Veröffentlicht: (2025)
MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting
von: Ai, Zhiqi, et al.
Veröffentlicht: (2024)
von: Ai, Zhiqi, et al.
Veröffentlicht: (2024)
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
von: Park, Joonyong, et al.
Veröffentlicht: (2024)
von: Park, Joonyong, et al.
Veröffentlicht: (2024)
Beat-Based Rhythm Quantization of MIDI Performances
von: Wachter, Maximilian, et al.
Veröffentlicht: (2025)
von: Wachter, Maximilian, et al.
Veröffentlicht: (2025)
From Weak Labels to Strong Results: Utilizing 5,000 Hours of Noisy Classroom Transcripts with Minimal Accurate Data
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech
von: Dong, Zhongren, et al.
Veröffentlicht: (2024)
von: Dong, Zhongren, et al.
Veröffentlicht: (2024)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
LastResort at SemEval-2024 Task 3: Exploring Multimodal Emotion Cause Pair Extraction as Sequence Labelling Task
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning
von: Shao, Yiwen, et al.
Veröffentlicht: (2025)
von: Shao, Yiwen, et al.
Veröffentlicht: (2025)
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
von: Lei, Ligong, et al.
Veröffentlicht: (2026)
von: Lei, Ligong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025) -
persoDA: Personalized Data Augmentation for Personalized ASR
von: Parada, Pablo Peso, et al.
Veröffentlicht: (2025) -
ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization
von: Mehmood, Haaris, et al.
Veröffentlicht: (2025) -
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
von: Waheed, Abdul, et al.
Veröffentlicht: (2024) -
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)