HydraFormer: One Encoder For All Subsampling Rates
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yaoxun, Song, Xingchen, Wu, Zhiyong, Wu, Di, Peng, Zhendong, Zhang, Binbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
von: Wang, Xi, et al.
Veröffentlicht: (2026)
von: Wang, Xi, et al.
Veröffentlicht: (2026)
One Whisper to Grade Them All
von: Phan, Nhan, et al.
Veröffentlicht: (2025)
von: Phan, Nhan, et al.
Veröffentlicht: (2025)
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue
von: Lu, Hui, et al.
Veröffentlicht: (2026)
von: Lu, Hui, et al.
Veröffentlicht: (2026)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
Training and Inference Efficiency of Encoder-Decoder Speech Models
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
Seamless Language Expansion: Enhancing Multilingual Mastery in Self-Supervised Models
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
WAKE: Watermarking Audio with Key Enrichment
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
MiLorE-SSL: Scaling Multilingual Capabilities in Self-Supervised Models without Forgetting
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
Lamer-SSL: Layer-aware Mixture of LoRA Experts for Continual Multilingual Expansion of Self-supervised Models without Forgetting
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting
von: Ai, Zhiqi, et al.
Veröffentlicht: (2024)
von: Ai, Zhiqi, et al.
Veröffentlicht: (2024)
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
von: Wu, Haibin, et al.
Veröffentlicht: (2025)
von: Wu, Haibin, et al.
Veröffentlicht: (2025)
Speech-Based Depression Prediction Using Encoder-Weight-Only Transfer Learning and a Large Corpus
von: Harati, Amir, et al.
Veröffentlicht: (2024)
von: Harati, Amir, et al.
Veröffentlicht: (2024)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction
von: Xu, Qianheng
Veröffentlicht: (2025)
von: Xu, Qianheng
Veröffentlicht: (2025)
Alethia: A Foundational Encoder for Voice Deepfakes
von: Zhu, Yi, et al.
Veröffentlicht: (2026)
von: Zhu, Yi, et al.
Veröffentlicht: (2026)
Borderless Long Speech Synthesis
von: Song, Xingchen, et al.
Veröffentlicht: (2026)
von: Song, Xingchen, et al.
Veröffentlicht: (2026)
MuCodec: Ultra Low-Bitrate Music Codec
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
von: Tang, Yun, et al.
Veröffentlicht: (2025)
von: Tang, Yun, et al.
Veröffentlicht: (2025)
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Attention or Convolution: Transformer Encoders in Audio Language Models for Inference Efficiency
von: Jeon, Sungho, et al.
Veröffentlicht: (2023)
von: Jeon, Sungho, et al.
Veröffentlicht: (2023)
Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction
von: Hao, Xiang, et al.
Veröffentlicht: (2023)
von: Hao, Xiang, et al.
Veröffentlicht: (2023)
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2026)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2026)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
von: Mdhaffar, Salima, et al.
Veröffentlicht: (2024)
von: Mdhaffar, Salima, et al.
Veröffentlicht: (2024)
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
von: Ling, Shaoshi, et al.
Veröffentlicht: (2023)
von: Ling, Shaoshi, et al.
Veröffentlicht: (2023)
Revisiting Interpolation Augmentation for Speech-to-Text Generation
von: Xu, Chen, et al.
Veröffentlicht: (2024)
von: Xu, Chen, et al.
Veröffentlicht: (2024)
The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
von: Chen, William, et al.
Veröffentlicht: (2025)
von: Chen, William, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024) -
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF
von: Song, Xingchen, et al.
Veröffentlicht: (2024) -
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024) -
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
von: Wang, Xi, et al.
Veröffentlicht: (2026) -
One Whisper to Grade Them All
von: Phan, Nhan, et al.
Veröffentlicht: (2025)