LLMs and Speech: Integration vs. Combination
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schmitt, Robin, Zeyer, Albert, Zeineldeen, Mohammad, Schlüter, Ralf, Ney, Hermann |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
The Conformer Encoder May Reverse the Time Dimension
von: Schmitt, Robin, et al.
Veröffentlicht: (2024)
von: Schmitt, Robin, et al.
Veröffentlicht: (2024)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
von: Yang, Zijian, et al.
Veröffentlicht: (2026)
von: Yang, Zijian, et al.
Veröffentlicht: (2026)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
von: Raissi, Tina, et al.
Veröffentlicht: (2024)
von: Raissi, Tina, et al.
Veröffentlicht: (2024)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
Unified Learnable 2D Convolutional Feature Extraction for ASR
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
Label-Context-Dependent Internal Language Model Estimation for CTC
von: Yang, Zijian, et al.
Veröffentlicht: (2025)
von: Yang, Zijian, et al.
Veröffentlicht: (2025)
Alternating Weak Triphone/BPE Alignment Supervision from Hybrid Model Improves End-to-End ASR
von: Jiang, Jintao, et al.
Veröffentlicht: (2024)
von: Jiang, Jintao, et al.
Veröffentlicht: (2024)
Data Augmentation for Pathological Speech Enhancement
von: Hou, Mingchi, et al.
Veröffentlicht: (2026)
von: Hou, Mingchi, et al.
Veröffentlicht: (2026)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
von: Ho, Chun-Wei, et al.
Veröffentlicht: (2025)
von: Ho, Chun-Wei, et al.
Veröffentlicht: (2025)
Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025)
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
von: Vieting, Peter, et al.
Veröffentlicht: (2023)
von: Vieting, Peter, et al.
Veröffentlicht: (2023)
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
von: Sasu, David, et al.
Veröffentlicht: (2025)
von: Sasu, David, et al.
Veröffentlicht: (2025)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2024)
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2024)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
Combined Generative and Predictive Modeling for Speech Super-resolution
von: Wang, Heming, et al.
Veröffentlicht: (2024)
von: Wang, Heming, et al.
Veröffentlicht: (2024)
Privacy Disclosure of Similarity Rank in Speech and Language Processing
von: Bäckström, Tom, et al.
Veröffentlicht: (2025)
von: Bäckström, Tom, et al.
Veröffentlicht: (2025)
Interspeech 2025 URGENT Speech Enhancement Challenge
von: Saijo, Kohei, et al.
Veröffentlicht: (2025)
von: Saijo, Kohei, et al.
Veröffentlicht: (2025)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Universal Score-based Speech Enhancement with High Content Preservation
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
Non-Intrusive Automatic Speech Recognition Refinement: A Survey
von: Peyghan, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Peyghan, Mohammad Reza, et al.
Veröffentlicht: (2025)
Towards EMG-to-Speech with a Necklace Form Factor
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
Integrated Minimum Mean Squared Error Algorithms for Combined Acoustic Echo Cancellation and Noise Reduction
von: Roebben, Arnout, et al.
Veröffentlicht: (2024)
von: Roebben, Arnout, et al.
Veröffentlicht: (2024)
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2025)
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2025)
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
AS-Speech: Adaptive Style For Speech Synthesis
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
von: de Groot, Dimme, et al.
Veröffentlicht: (2026)
von: de Groot, Dimme, et al.
Veröffentlicht: (2026)
Vision-Integrated High-Quality Neural Speech Coding
von: Guo, Yao, et al.
Veröffentlicht: (2025)
von: Guo, Yao, et al.
Veröffentlicht: (2025)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2026)
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2026)
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge
von: Sach, Marvin, et al.
Veröffentlicht: (2025)
von: Sach, Marvin, et al.
Veröffentlicht: (2025)
Enhancing Speech Quality through the Integration of BGRU and Transformer Architectures
von: Alghnam, Souliman, et al.
Veröffentlicht: (2025)
von: Alghnam, Souliman, et al.
Veröffentlicht: (2025)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023) -
The Conformer Encoder May Reverse the Time Dimension
von: Schmitt, Robin, et al.
Veröffentlicht: (2024) -
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
von: Raissi, Tina, et al.
Veröffentlicht: (2025) -
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
von: Yang, Zijian, et al.
Veröffentlicht: (2026) -
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
von: Raissi, Tina, et al.
Veröffentlicht: (2024)