Cross-Attention is Half Explanation in Speech-to-Text Models
Fuente:
arXiv
Saved in:
| Main Authors: | Papi, Sara, Fucci, Dennis, Gaido, Marco, Negri, Matteo, Bentivogli, Luisa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
by: Conti, Lina, et al.
Published: (2025)
by: Conti, Lina, et al.
Published: (2025)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
by: Papi, Sara, et al.
Published: (2026)
by: Papi, Sara, et al.
Published: (2026)
How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
by: Fucci, Dennis, et al.
Published: (2024)
by: Fucci, Dennis, et al.
Published: (2024)
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
by: Lam, Tsz Kin, et al.
Published: (2025)
by: Lam, Tsz Kin, et al.
Published: (2025)
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
by: Cettolo, Mauro, et al.
Published: (2025)
by: Cettolo, Mauro, et al.
Published: (2025)
Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation
by: Conti, Lina, et al.
Published: (2025)
by: Conti, Lina, et al.
Published: (2025)
Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
by: Fucci, Dennis, et al.
Published: (2025)
by: Fucci, Dennis, et al.
Published: (2025)
Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
by: Djanibekov, Amirbek, et al.
Published: (2026)
by: Djanibekov, Amirbek, et al.
Published: (2026)
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
by: Gaido, Marco, et al.
Published: (2025)
by: Gaido, Marco, et al.
Published: (2025)
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLP
by: Papi, Sara, et al.
Published: (2023)
by: Papi, Sara, et al.
Published: (2023)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
by: Papi, Sara, et al.
Published: (2023)
by: Papi, Sara, et al.
Published: (2023)
Different Speech Translation Models Encode and Translate Speaker Gender Differently
by: Fucci, Dennis, et al.
Published: (2025)
by: Fucci, Dennis, et al.
Published: (2025)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
by: Polák, Peter, et al.
Published: (2025)
by: Polák, Peter, et al.
Published: (2025)
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
by: Gaido, Marco, et al.
Published: (2025)
by: Gaido, Marco, et al.
Published: (2025)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
by: Lee, Beomseok, et al.
Published: (2024)
by: Lee, Beomseok, et al.
Published: (2024)
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection
by: Lee, Beomseok, et al.
Published: (2024)
by: Lee, Beomseok, et al.
Published: (2024)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
Efficient Training for Cross-lingual Speech Language Models
by: Zhou, Yan, et al.
Published: (2026)
by: Zhou, Yan, et al.
Published: (2026)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
by: Lin, Tzu-Quan, et al.
Published: (2025)
by: Lin, Tzu-Quan, et al.
Published: (2025)
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
by: Verdini, Francesco, et al.
Published: (2024)
by: Verdini, Francesco, et al.
Published: (2024)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
by: Shi, Jiacheng, et al.
Published: (2025)
by: Shi, Jiacheng, et al.
Published: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
by: Kotoge, Rikuto, et al.
Published: (2025)
by: Kotoge, Rikuto, et al.
Published: (2025)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
by: Toyin, Hawau Olamide, et al.
Published: (2024)
by: Toyin, Hawau Olamide, et al.
Published: (2024)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
by: Riera, Pablo, et al.
Published: (2026)
by: Riera, Pablo, et al.
Published: (2026)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
by: Lou, Yuxuan, et al.
Published: (2026)
by: Lou, Yuxuan, et al.
Published: (2026)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
RECA-PD: A Robust Explainable Cross-Attention Method for Speech-based Parkinson's Disease Classification
by: Zhong, Terry Yi, et al.
Published: (2025)
by: Zhong, Terry Yi, et al.
Published: (2025)
A Prompt Response to the Demand for Automatic Gender-Neutral Translation
by: Savoldi, Beatrice, et al.
Published: (2024)
by: Savoldi, Beatrice, et al.
Published: (2024)
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
by: Liu, Ruohan, et al.
Published: (2026)
by: Liu, Ruohan, et al.
Published: (2026)
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
by: Attia, Ahmed Adel, et al.
Published: (2025)
by: Attia, Ahmed Adel, et al.
Published: (2025)
Similar Items
-
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
by: Papi, Sara, et al.
Published: (2024) -
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
by: Papi, Sara, et al.
Published: (2024) -
The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
by: Conti, Lina, et al.
Published: (2025) -
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
by: Papi, Sara, et al.
Published: (2026) -
How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena
by: Gaido, Marco, et al.
Published: (2024)