SpeechMapper: Speech-to-text Embedding Projector for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Mohapatra, Biswesh, Boito, Marcely Zanon, Calapodescu, Ioan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NAVER LABS Europe Submission to the Instruction-following Track
by: Lee, Beomseok, et al.
Published: (2025)
by: Lee, Beomseok, et al.
Published: (2025)
mHuBERT-147: A Compact Multilingual HuBERT Model
by: Boito, Marcely Zanon, et al.
Published: (2024)
by: Boito, Marcely Zanon, et al.
Published: (2024)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
by: Ferraz, Thomas Palmeira, et al.
Published: (2023)
by: Ferraz, Thomas Palmeira, et al.
Published: (2023)
StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario
by: Boito, Marcely Zanon, et al.
Published: (2026)
by: Boito, Marcely Zanon, et al.
Published: (2026)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
by: Lee, Beomseok, et al.
Published: (2024)
by: Lee, Beomseok, et al.
Published: (2024)
From TOWER to SPIRE: Adding the Speech Modality to a Translation-Specialist LLM
by: Ambilduke, Kshitij, et al.
Published: (2025)
by: Ambilduke, Kshitij, et al.
Published: (2025)
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection
by: Lee, Beomseok, et al.
Published: (2024)
by: Lee, Beomseok, et al.
Published: (2024)
Conversational Grounding: Annotation and Analysis of Grounding Acts and Grounding Units
by: Mohapatra, Biswesh, et al.
Published: (2024)
by: Mohapatra, Biswesh, et al.
Published: (2024)
Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs
by: Mohapatra, Payal, et al.
Published: (2025)
by: Mohapatra, Payal, et al.
Published: (2025)
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
by: Pandey, Isha, et al.
Published: (2026)
by: Pandey, Isha, et al.
Published: (2026)
An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks
by: Suresh, Varsha, et al.
Published: (2024)
by: Suresh, Varsha, et al.
Published: (2024)
Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
by: Meng, Chutong, et al.
Published: (2025)
by: Meng, Chutong, et al.
Published: (2025)
End-to-end Speech Recognition with similar length speech and text
by: Fan, Peng, et al.
Published: (2025)
by: Fan, Peng, et al.
Published: (2025)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
by: Luu, Nam, et al.
Published: (2025)
by: Luu, Nam, et al.
Published: (2025)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
by: Zhang, Linhao, et al.
Published: (2025)
by: Zhang, Linhao, et al.
Published: (2025)
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
by: Parcollet, Titouan, et al.
Published: (2023)
by: Parcollet, Titouan, et al.
Published: (2023)
It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems
by: Zaitova, Iuliia, et al.
Published: (2025)
by: Zaitova, Iuliia, et al.
Published: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
by: Futami, Hayato, et al.
Published: (2025)
by: Futami, Hayato, et al.
Published: (2025)
Speech LLMs are Contextual Reasoning Transcribers
by: Deng, Keqi, et al.
Published: (2026)
by: Deng, Keqi, et al.
Published: (2026)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
by: Zhao, Mengjie, et al.
Published: (2026)
by: Zhao, Mengjie, et al.
Published: (2026)
Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects
by: Chang, Kalvin, et al.
Published: (2026)
by: Chang, Kalvin, et al.
Published: (2026)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
by: Deng, Keqi, et al.
Published: (2025)
by: Deng, Keqi, et al.
Published: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
by: Song, Yuhan, et al.
Published: (2025)
by: Song, Yuhan, et al.
Published: (2025)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
by: Pareras, Oriol, et al.
Published: (2025)
by: Pareras, Oriol, et al.
Published: (2025)
Chain of Correction for Full-text Speech Recognition with Large Language Models
by: Tang, Zhiyuan, et al.
Published: (2025)
by: Tang, Zhiyuan, et al.
Published: (2025)
Slot Filling as a Reasoning Task for SpeechLLMs
by: Hacioglu, Kadri, et al.
Published: (2025)
by: Hacioglu, Kadri, et al.
Published: (2025)
Comparing Discrete and Continuous Space LLMs for Speech Recognition
by: Xu, Yaoxun, et al.
Published: (2024)
by: Xu, Yaoxun, et al.
Published: (2024)
Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility
by: Wang, Sheng-Fu, et al.
Published: (2025)
by: Wang, Sheng-Fu, et al.
Published: (2025)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
On Importance of Code-Mixed Embeddings for Hate Speech Identification
by: Jagdale, Shruti, et al.
Published: (2024)
by: Jagdale, Shruti, et al.
Published: (2024)
Investigating Decoder-only Large Language Models for Speech-to-text Translation
by: Huang, Chao-Wei, et al.
Published: (2024)
by: Huang, Chao-Wei, et al.
Published: (2024)
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
by: Koneru, Sai, et al.
Published: (2024)
by: Koneru, Sai, et al.
Published: (2024)
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis
by: Attaluri, Kaushal, et al.
Published: (2024)
by: Attaluri, Kaushal, et al.
Published: (2024)
Measuring the Redundancy of Decoder Layers in SpeechLLMs
by: Moumen, Adel, et al.
Published: (2026)
by: Moumen, Adel, et al.
Published: (2026)
Similar Items
-
NAVER LABS Europe Submission to the Instruction-following Track
by: Lee, Beomseok, et al.
Published: (2025) -
mHuBERT-147: A Compact Multilingual HuBERT Model
by: Boito, Marcely Zanon, et al.
Published: (2024) -
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
by: Ferraz, Thomas Palmeira, et al.
Published: (2023) -
StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario
by: Boito, Marcely Zanon, et al.
Published: (2026) -
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
by: Lee, Beomseok, et al.
Published: (2024)