ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mousavi, Pooneh, Wang, Yingzhi, Ravanelli, Mirco, Subakan, Cem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Are They Doing? Joint Audio-Speech Co-Reasoning
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
Autoregressive Speech Enhancement via Acoustic Tokens
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
DASB - Discrete Audio and Speech Benchmark
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
Investigating Faithfulness in Large Audio Language Models
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Listenable Maps for Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
Focal Modulation Networks for Interpretable Sound Classification
von: Della Libera, Luca, et al.
Veröffentlicht: (2024)
von: Della Libera, Luca, et al.
Veröffentlicht: (2024)
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Audio Editing with Non-Rigid Text Prompts
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
ProGRes: Prompted Generative Rescoring on ASR n-Best
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
Toward Faithful Explanations in Acoustic Anomaly Detection
von: Elrashid, Maab, et al.
Veröffentlicht: (2026)
von: Elrashid, Maab, et al.
Veröffentlicht: (2026)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
Long-Form Speech Generation with Spoken Language Models
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
Discrete Audio Tokens: More Than a Survey!
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
Resource-Efficient Separation Transformer
von: Della Libera, Luca, et al.
Veröffentlicht: (2022)
von: Della Libera, Luca, et al.
Veröffentlicht: (2022)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
von: Valente, Martina, et al.
Veröffentlicht: (2024)
von: Valente, Martina, et al.
Veröffentlicht: (2024)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
von: Le, Trang, et al.
Veröffentlicht: (2024)
von: Le, Trang, et al.
Veröffentlicht: (2024)
Computational Narrative Understanding for Expressive Text-to-Speech
von: Michel, Gaspard, et al.
Veröffentlicht: (2025)
von: Michel, Gaspard, et al.
Veröffentlicht: (2025)
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
"Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What Are They Doing? Joint Audio-Speech Co-Reasoning
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024) -
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026) -
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025) -
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024) -
Autoregressive Speech Enhancement via Acoustic Tokens
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)