Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saon, George, Dekel, Avihu, Brooks, Alexander, Nagano, Tohru, Daniels, Abraham, Satt, Aharon, Mittal, Ashish, Kingsbury, Brian, Haws, David, Morais, Edmilson, Kurata, Gakuto, Aronowitz, Hagai, Ibrahim, Ibrahim, Kuo, Jeff, Soule, Kate, Lastras, Luis, Suzuki, Masayuki, Hoory, Ron, Thomas, Samuel, Novitasari, Sashi, Fukuda, Takashi, Sunder, Vishal, Cui, Xiaodong, Kons, Zvi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
Spoken question answering for visual queries
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2025)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2025)
Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026)
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026)
Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
von: Turetzky, Arnon, et al.
Veröffentlicht: (2026)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2026)
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
von: Saon, George, et al.
Veröffentlicht: (2026)
von: Saon, George, et al.
Veröffentlicht: (2026)
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
A Non-autoregressive Model for Joint STT and TTS
von: Sunder, Vishal, et al.
Veröffentlicht: (2025)
von: Sunder, Vishal, et al.
Veröffentlicht: (2025)
Exploring the limits of decoder-only models trained on public speech recognition corpora
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
von: Dekel, Avihu, et al.
Veröffentlicht: (2026)
von: Dekel, Avihu, et al.
Veröffentlicht: (2026)
Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
von: Shechtman, Slava, et al.
Veröffentlicht: (2024)
von: Shechtman, Slava, et al.
Veröffentlicht: (2024)
Exploring the Benefits of Tokenization of Discrete Acoustic Units
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Robust ASR Error Correction with Conservative Data Filtering
von: Udagawa, Takuma, et al.
Veröffentlicht: (2024)
von: Udagawa, Takuma, et al.
Veröffentlicht: (2024)
Ice-Breakers to Serve the Elderly
von: Haws, Richard
Veröffentlicht: (1978)
von: Haws, Richard
Veröffentlicht: (1978)
An Attitudinal Study of Students toward a Required Library Course.
von: Haws, Rae
Veröffentlicht: (1987)
von: Haws, Rae
Veröffentlicht: (1987)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
Balanced Thinking: Improving Chain of Thought Training in Vision Language Models
von: Perek, Shaked, et al.
Veröffentlicht: (2026)
von: Perek, Shaked, et al.
Veröffentlicht: (2026)
Prominence-aware automatic speech recognition for conversational speech
von: Linke, Julian, et al.
Veröffentlicht: (2025)
von: Linke, Julian, et al.
Veröffentlicht: (2025)
Granite Embedding Models
von: Awasthy, Parul, et al.
Veröffentlicht: (2025)
von: Awasthy, Parul, et al.
Veröffentlicht: (2025)
The knowledge factory : dismantling the corporate university and creating true higher learning / Stanley Aronowitz
von: Aronowitz, stanley
von: Aronowitz, stanley
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
Acoustic and linguistic effects in synthesized speech augmentation for speech recognition
von: Yohan Lim, et al.
Veröffentlicht: (2025)
von: Yohan Lim, et al.
Veröffentlicht: (2025)
On the Girth of Graph Lifts
von: Hoory, Shlomo
Veröffentlicht: (2024)
von: Hoory, Shlomo
Veröffentlicht: (2024)
The use of severity measures and speech inconsistency in children with speech sound disorders
von: Haydée Fiszbein Wertzner
Veröffentlicht: (2013)
von: Haydée Fiszbein Wertzner
Veröffentlicht: (2013)
The effect of SpeechEasy on stuttering frequency, speech rate and speech naturalness
von: Claudia Regina Furquim de Andrade
Veröffentlicht: (2008)
von: Claudia Regina Furquim de Andrade
Veröffentlicht: (2008)
Text to speech synthesis
von: s, Harini, et al.
Veröffentlicht: (2024)
von: s, Harini, et al.
Veröffentlicht: (2024)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
Limit cycles for speech
von: Gafos, Adamantios I., et al.
Veröffentlicht: (2025)
von: Gafos, Adamantios I., et al.
Veröffentlicht: (2025)
Freedom of speech in Rome
von: José Manuel Díaz de Valdés
Veröffentlicht: (2009)
von: José Manuel Díaz de Valdés
Veröffentlicht: (2009)
Towards explainable reference-free speech intelligibility evaluation of people with pathological speech
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2026)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2026)
What does it take to get state of the art in simultaneous speech-to-speech translation?
von: Wilmet, Vincent, et al.
Veröffentlicht: (2024)
von: Wilmet, Vincent, et al.
Veröffentlicht: (2024)
Language translation, and change of accent for speech-to-speech task using diffusion model
von: Mishra, Abhishek, et al.
Veröffentlicht: (2025)
von: Mishra, Abhishek, et al.
Veröffentlicht: (2025)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
Why Aren't There More Whistleblowers?
von: Robert A. Aronowitz
Veröffentlicht: (2024)
von: Robert A. Aronowitz
Veröffentlicht: (2024)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
Activation of the inner critic increases speech illusions and negative emotional valence of perceived speech
von: June Engeland, et al.
Veröffentlicht: (2025)
von: June Engeland, et al.
Veröffentlicht: (2025)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026) -
Spoken question answering for visual queries
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2025) -
Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026) -
Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
von: Turetzky, Arnon, et al.
Veröffentlicht: (2026) -
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
von: Saon, George, et al.
Veröffentlicht: (2026)