Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Parcollet, Titouan, van Dalen, Rogier, Zhang, Shucong, Batthacharya, Sourav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
Linear-Complexity Self-Supervised Learning for Speech Processing
von: Zhang, Shucong, et al.
Veröffentlicht: (2024)
von: Zhang, Shucong, et al.
Veröffentlicht: (2024)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
von: Zhang, Shucong, et al.
Veröffentlicht: (2025)
von: Zhang, Shucong, et al.
Veröffentlicht: (2025)
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025)
Streaming Speech-to-Text Translation with a SpeechLLM
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
von: Tseng, Yuan, et al.
Veröffentlicht: (2025)
von: Tseng, Yuan, et al.
Veröffentlicht: (2025)
Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
von: van Dalen, Rogier C., et al.
Veröffentlicht: (2025)
von: van Dalen, Rogier C., et al.
Veröffentlicht: (2025)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
Do we really need Self-Attention for Streaming Automatic Speech Recognition?
von: Dkhissi, Youness, et al.
Veröffentlicht: (2026)
von: Dkhissi, Youness, et al.
Veröffentlicht: (2026)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2024)
von: Chen, Jinming, et al.
Veröffentlicht: (2024)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
von: Seide, Frank, et al.
Veröffentlicht: (2024)
von: Seide, Frank, et al.
Veröffentlicht: (2024)
TinyML for Speech Recognition
von: Barovic, Andrew, et al.
Veröffentlicht: (2025)
von: Barovic, Andrew, et al.
Veröffentlicht: (2025)
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
von: Whetten, Ryan, et al.
Veröffentlicht: (2024)
von: Whetten, Ryan, et al.
Veröffentlicht: (2024)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Speech Recognition on TV Series with Video-guided Post-ASR Correction
von: Yang, Haoyuan, et al.
Veröffentlicht: (2025)
von: Yang, Haoyuan, et al.
Veröffentlicht: (2025)
Active Learning with Task Adaptation Pre-training for Speech Emotion Recognition
von: Li, Dongyuan, et al.
Veröffentlicht: (2024)
von: Li, Dongyuan, et al.
Veröffentlicht: (2024)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
Color-based Emotion Representation for Speech Emotion Recognition
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025)
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025)
ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024)
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023) -
Linear-Complexity Self-Supervised Learning for Speech Processing
von: Zhang, Shucong, et al.
Veröffentlicht: (2024) -
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
von: Zhang, Shucong, et al.
Veröffentlicht: (2025) -
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025) -
Streaming Speech-to-Text Translation with a SpeechLLM
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)