DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Papi, Sara, Bentivogli, Luisa |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
di: Papi, Sara, et al.
Pubblicazione: (2024)
di: Papi, Sara, et al.
Pubblicazione: (2024)
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
di: Papi, Sara, et al.
Pubblicazione: (2024)
di: Papi, Sara, et al.
Pubblicazione: (2024)
Cross-Attention is Half Explanation in Speech-to-Text Models
di: Papi, Sara, et al.
Pubblicazione: (2025)
di: Papi, Sara, et al.
Pubblicazione: (2025)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2026)
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2026)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
di: Polák, Peter, et al.
Pubblicazione: (2025)
di: Polák, Peter, et al.
Pubblicazione: (2025)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
di: Papi, Sara, et al.
Pubblicazione: (2024)
di: Papi, Sara, et al.
Pubblicazione: (2024)
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
di: Papi, Sara, et al.
Pubblicazione: (2025)
di: Papi, Sara, et al.
Pubblicazione: (2025)
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
di: Gaido, Marco, et al.
Pubblicazione: (2024)
di: Gaido, Marco, et al.
Pubblicazione: (2024)
How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena
di: Gaido, Marco, et al.
Pubblicazione: (2024)
di: Gaido, Marco, et al.
Pubblicazione: (2024)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
di: Papi, Sara, et al.
Pubblicazione: (2025)
di: Papi, Sara, et al.
Pubblicazione: (2025)
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
di: Lam, Tsz Kin, et al.
Pubblicazione: (2025)
di: Lam, Tsz Kin, et al.
Pubblicazione: (2025)
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
di: Cettolo, Mauro, et al.
Pubblicazione: (2025)
di: Cettolo, Mauro, et al.
Pubblicazione: (2025)
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
di: Parikh, Aditya Kamlesh, et al.
Pubblicazione: (2026)
di: Parikh, Aditya Kamlesh, et al.
Pubblicazione: (2026)
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
di: Papi, Sara, et al.
Pubblicazione: (2025)
di: Papi, Sara, et al.
Pubblicazione: (2025)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
di: Papi, Sara, et al.
Pubblicazione: (2023)
di: Papi, Sara, et al.
Pubblicazione: (2023)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
di: Satish, Shree Harsha Bokkahalli, et al.
Pubblicazione: (2025)
di: Satish, Shree Harsha Bokkahalli, et al.
Pubblicazione: (2025)
Measuring the Redundancy of Decoder Layers in SpeechLLMs
di: Moumen, Adel, et al.
Pubblicazione: (2026)
di: Moumen, Adel, et al.
Pubblicazione: (2026)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
di: Polák, Peter, et al.
Pubblicazione: (2023)
di: Polák, Peter, et al.
Pubblicazione: (2023)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
di: Zhang, Shaolei, et al.
Pubblicazione: (2024)
di: Zhang, Shaolei, et al.
Pubblicazione: (2024)
Recent Advances in End-to-End Simultaneous Speech Translation
di: Liu, Xiaoqian, et al.
Pubblicazione: (2024)
di: Liu, Xiaoqian, et al.
Pubblicazione: (2024)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
di: Zhao, Mengjie, et al.
Pubblicazione: (2026)
di: Zhao, Mengjie, et al.
Pubblicazione: (2026)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
di: Ma, Zhengrui, et al.
Pubblicazione: (2024)
di: Ma, Zhengrui, et al.
Pubblicazione: (2024)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
di: Song, Yuhan, et al.
Pubblicazione: (2025)
di: Song, Yuhan, et al.
Pubblicazione: (2025)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
di: Gao, Yan, et al.
Pubblicazione: (2025)
di: Gao, Yan, et al.
Pubblicazione: (2025)
Efficient Training for Cross-lingual Speech Language Models
di: Zhou, Yan, et al.
Pubblicazione: (2026)
di: Zhou, Yan, et al.
Pubblicazione: (2026)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
di: Satish, Shree Harsha Bokkahalli, et al.
Pubblicazione: (2026)
di: Satish, Shree Harsha Bokkahalli, et al.
Pubblicazione: (2026)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
di: Ginjala, Srishti, et al.
Pubblicazione: (2026)
di: Ginjala, Srishti, et al.
Pubblicazione: (2026)
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
di: Srivastav, Vaibhav, et al.
Pubblicazione: (2025)
di: Srivastav, Vaibhav, et al.
Pubblicazione: (2025)
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
di: Gaido, Marco, et al.
Pubblicazione: (2024)
di: Gaido, Marco, et al.
Pubblicazione: (2024)
Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)
di: Wang, Peidong
Pubblicazione: (2026)
di: Wang, Peidong
Pubblicazione: (2026)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
di: Chen, Szu-Chi, et al.
Pubblicazione: (2026)
di: Chen, Szu-Chi, et al.
Pubblicazione: (2026)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026)
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026)
Streaming Speech-to-Text Translation with a SpeechLLM
di: Parcollet, Titouan, et al.
Pubblicazione: (2026)
di: Parcollet, Titouan, et al.
Pubblicazione: (2026)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
di: Le, Chenyang, et al.
Pubblicazione: (2024)
di: Le, Chenyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
di: Papi, Sara, et al.
Pubblicazione: (2024) -
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
di: Papi, Sara, et al.
Pubblicazione: (2024) -
Cross-Attention is Half Explanation in Speech-to-Text Models
di: Papi, Sara, et al.
Pubblicazione: (2025) -
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2026) -
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
di: Polák, Peter, et al.
Pubblicazione: (2025)