DTW-Align: Bridging the Modality Gap in End-to-End Speech Translation with Dynamic Time Warping Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Issam, Abderrahmane, Semerci, Yusuf Can, Scholtes, Jan, Spanakis, Gerasimos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text
by: Issam, Abderrahmane, et al.
Published: (2026)
by: Issam, Abderrahmane, et al.
Published: (2026)
Fixed and Adaptive Simultaneous Machine Translation Strategies Using Adapters
by: Issam, Abderrahmane, et al.
Published: (2024)
by: Issam, Abderrahmane, et al.
Published: (2024)
A Representation Level Analysis of NMT Model Robustness to Grammatical Errors
by: Issam, Abderrahmane, et al.
Published: (2025)
by: Issam, Abderrahmane, et al.
Published: (2025)
Language Models as Artificial Learners: Investigating Crosslinguistic Influence
by: Issam, Abderrahmane, et al.
Published: (2026)
by: Issam, Abderrahmane, et al.
Published: (2026)
Sequence Shortening for Context-Aware Machine Translation
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2025)
by: Mąka, Paweł, et al.
Published: (2025)
From FusHa to Folk: Exploring Cross-Lingual Transfer in Arabic Language Models
by: Khalak, Abdulmuizz, et al.
Published: (2026)
by: Khalak, Abdulmuizz, et al.
Published: (2026)
Dutch CrowS-Pairs: Adapting a Challenge Dataset for Measuring Social Biases in Language Models for Dutch
by: Strazda, Elza, et al.
Published: (2025)
by: Strazda, Elza, et al.
Published: (2025)
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
by: Hsu, Ming-Hao, et al.
Published: (2026)
by: Hsu, Ming-Hao, et al.
Published: (2026)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
A Case Study on Filtering for End-to-End Speech Translation
by: Alam, Md Mahfuz Ibn, et al.
Published: (2024)
by: Alam, Md Mahfuz Ibn, et al.
Published: (2024)
Pushing the Limits of Zero-shot End-to-End Speech Translation
by: Tsiamas, Ioannis, et al.
Published: (2024)
by: Tsiamas, Ioannis, et al.
Published: (2024)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
by: Polák, Peter, et al.
Published: (2023)
by: Polák, Peter, et al.
Published: (2023)
Representation Purification for End-to-End Speech Translation
by: Zhang, Chengwei, et al.
Published: (2024)
by: Zhang, Chengwei, et al.
Published: (2024)
End-to-End Speech-to-Text Translation: A Survey
by: Sethiya, Nivedita, et al.
Published: (2023)
by: Sethiya, Nivedita, et al.
Published: (2023)
Maastricht University at AMIYA: Adapting LLMs for Dialectal Arabic using Fine-tuning and MBR Decoding
by: Alali, Abdulhai, et al.
Published: (2026)
by: Alali, Abdulhai, et al.
Published: (2026)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
by: Vendrame, Katia, et al.
Published: (2025)
by: Vendrame, Katia, et al.
Published: (2025)
Recent Advances in End-to-End Simultaneous Speech Translation
by: Liu, Xiaoqian, et al.
Published: (2024)
by: Liu, Xiaoqian, et al.
Published: (2024)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
by: Alastruey, Belen, et al.
Published: (2023)
by: Alastruey, Belen, et al.
Published: (2023)
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
by: Yusuf, Bolaji, et al.
Published: (2024)
by: Yusuf, Bolaji, et al.
Published: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
by: Huang, Wuwei, et al.
Published: (2025)
by: Huang, Wuwei, et al.
Published: (2025)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
by: Huang, Wuwei, et al.
Published: (2025)
by: Huang, Wuwei, et al.
Published: (2025)
Know When to Fuse: Investigating Non-English Hybrid Retrieval in the Legal Domain
by: Louis, Antoine, et al.
Published: (2024)
by: Louis, Antoine, et al.
Published: (2024)
Computational Studies in Influencer Marketing: A Systematic Literature Review
by: Gui, Haoyang, et al.
Published: (2025)
by: Gui, Haoyang, et al.
Published: (2025)
Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis
by: Bertaglia, Thales, et al.
Published: (2026)
by: Bertaglia, Thales, et al.
Published: (2026)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
by: Ma, Zhengrui, et al.
Published: (2024)
by: Ma, Zhengrui, et al.
Published: (2024)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
by: Moslem, Yasmin
Published: (2024)
by: Moslem, Yasmin
Published: (2024)
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
by: Pothula, Aishwarya, et al.
Published: (2025)
by: Pothula, Aishwarya, et al.
Published: (2025)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
by: Luu, Nam, et al.
Published: (2025)
by: Luu, Nam, et al.
Published: (2025)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
by: Zhang, Linhao, et al.
Published: (2025)
by: Zhang, Linhao, et al.
Published: (2025)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
by: Min, Anna, et al.
Published: (2025)
by: Min, Anna, et al.
Published: (2025)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
by: Wang, Peidong, et al.
Published: (2024)
by: Wang, Peidong, et al.
Published: (2024)
MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data
by: Saxena, Vageesh, et al.
Published: (2024)
by: Saxena, Vageesh, et al.
Published: (2024)
ColBERT-XM: A Modular Multi-Vector Representation Model for Zero-Shot Multilingual Information Retrieval
by: Louis, Antoine, et al.
Published: (2024)
by: Louis, Antoine, et al.
Published: (2024)
Cross-modality Data Augmentation for End-to-End Sign Language Translation
by: Ye, Jinhui, et al.
Published: (2023)
by: Ye, Jinhui, et al.
Published: (2023)
End-to-End Intracortical Speech Decoding from Neural Activity
by: Khanday, Owais Mujtaba, et al.
Published: (2026)
by: Khanday, Owais Mujtaba, et al.
Published: (2026)
Aligning Stuttered-Speech Research with End-User Needs: Scoping Review, Survey, and Guidelines
by: Toyin, Hawau Olamide, et al.
Published: (2026)
by: Toyin, Hawau Olamide, et al.
Published: (2026)
HITSZ's End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track
by: Wei, Xuchen, et al.
Published: (2025)
by: Wei, Xuchen, et al.
Published: (2025)
End-to-End Training for Back-Translation with Categorical Reparameterization Trick
by: Heo, DongNyeong, et al.
Published: (2022)
by: Heo, DongNyeong, et al.
Published: (2022)
Similar Items
-
Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text
by: Issam, Abderrahmane, et al.
Published: (2026) -
Fixed and Adaptive Simultaneous Machine Translation Strategies Using Adapters
by: Issam, Abderrahmane, et al.
Published: (2024) -
A Representation Level Analysis of NMT Model Robustness to Grammatical Errors
by: Issam, Abderrahmane, et al.
Published: (2025) -
Language Models as Artificial Learners: Investigating Crosslinguistic Influence
by: Issam, Abderrahmane, et al.
Published: (2026) -
Sequence Shortening for Context-Aware Machine Translation
by: Mąka, Paweł, et al.
Published: (2024)