RASST: Fast Cross-modal Retrieval-Augmented Simultaneous Speech Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Jiaxuan, Ouyang, Siqi, Li, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FASST: Fast LLM-based Simultaneous Speech Translation
by: Ouyang, Siqi, et al.
Published: (2024)
by: Ouyang, Siqi, et al.
Published: (2024)
CMU's IWSLT 2025 Simultaneous Speech Translation System
by: Ouyang, Siqi, et al.
Published: (2025)
by: Ouyang, Siqi, et al.
Published: (2025)
InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model
by: Ouyang, Siqi, et al.
Published: (2025)
by: Ouyang, Siqi, et al.
Published: (2025)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
by: Xu, Xi, et al.
Published: (2024)
by: Xu, Xi, et al.
Published: (2024)
Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
by: Ouyang, Siqi, et al.
Published: (2026)
by: Ouyang, Siqi, et al.
Published: (2026)
CMU's IWSLT 2024 Simultaneous Speech Translation System
by: Xu, Xi, et al.
Published: (2024)
by: Xu, Xi, et al.
Published: (2024)
Anticipating Future with Large Language Model for Simultaneous Machine Translation
by: Ouyang, Siqi, et al.
Published: (2024)
by: Ouyang, Siqi, et al.
Published: (2024)
Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration Approach
by: Li, Siqi, et al.
Published: (2024)
by: Li, Siqi, et al.
Published: (2024)
Translation Canvas: An Explainable Interface to Pinpoint and Analyze Translation Systems
by: Dandekar, Chinmay, et al.
Published: (2024)
by: Dandekar, Chinmay, et al.
Published: (2024)
Cross-modality Data Augmentation for End-to-End Sign Language Translation
by: Ye, Jinhui, et al.
Published: (2023)
by: Ye, Jinhui, et al.
Published: (2023)
Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation
by: Liu, Yifeng, et al.
Published: (2026)
by: Liu, Yifeng, et al.
Published: (2026)
Contrastive Feedback Mechanism for Simultaneous Speech Translation
by: Tan, Haotian, et al.
Published: (2024)
by: Tan, Haotian, et al.
Published: (2024)
Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture
by: Fu, Biao, et al.
Published: (2025)
by: Fu, Biao, et al.
Published: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
by: Labiausse, Tom, et al.
Published: (2025)
by: Labiausse, Tom, et al.
Published: (2025)
How Does Knowledge Selection Help Retrieval Augmented Generation?
by: Li, Xiangci, et al.
Published: (2024)
by: Li, Xiangci, et al.
Published: (2024)
Simultaneous Speech-to-Speech Translation Without Aligned Data
by: Labiausse, Tom, et al.
Published: (2026)
by: Labiausse, Tom, et al.
Published: (2026)
Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation
by: Javorský, Dávid, et al.
Published: (2022)
by: Javorský, Dávid, et al.
Published: (2022)
MLLP-VRAIN UPV system for the IWSLT 2025 Simultaneous Speech Translation Translation task
by: Iranzo-Sánchez, Jorge, et al.
Published: (2025)
by: Iranzo-Sánchez, Jorge, et al.
Published: (2025)
Recent Advances in End-to-End Simultaneous Speech Translation
by: Liu, Xiaoqian, et al.
Published: (2024)
by: Liu, Xiaoqian, et al.
Published: (2024)
Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs
by: Conia, Simone, et al.
Published: (2024)
by: Conia, Simone, et al.
Published: (2024)
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation
by: Tan, Haotian, et al.
Published: (2025)
by: Tan, Haotian, et al.
Published: (2025)
DPO-Tuned Large Language Models for Segmentation in Simultaneous Speech Translation
by: Yang, Zeyu, et al.
Published: (2025)
by: Yang, Zeyu, et al.
Published: (2025)
Exploring the Correlation between Human and Machine Evaluation of Simultaneous Speech Translation
by: Wang, Xiaoman, et al.
Published: (2024)
by: Wang, Xiaoman, et al.
Published: (2024)
SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation
by: Yang, Zeyu, et al.
Published: (2025)
by: Yang, Zeyu, et al.
Published: (2025)
SimulTron: On-Device Simultaneous Speech to Speech Translation
by: Agranovich, Alex, et al.
Published: (2024)
by: Agranovich, Alex, et al.
Published: (2024)
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
by: Moritz, Niko, et al.
Published: (2024)
by: Moritz, Niko, et al.
Published: (2024)
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
by: Xie, Jingyou, et al.
Published: (2024)
by: Xie, Jingyou, et al.
Published: (2024)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
by: Ma, Zi-Ao, et al.
Published: (2024)
by: Ma, Zi-Ao, et al.
Published: (2024)
Cross-Lingual Transfer Learning for Speech Translation
by: Ma, Rao, et al.
Published: (2024)
by: Ma, Rao, et al.
Published: (2024)
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
by: Zhu, Mengdan, et al.
Published: (2025)
by: Zhu, Mengdan, et al.
Published: (2025)
Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025
by: Macháček, Dominik, et al.
Published: (2025)
by: Macháček, Dominik, et al.
Published: (2025)
Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems
by: Feng, Tao, et al.
Published: (2026)
by: Feng, Tao, et al.
Published: (2026)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
by: Zhang, Shaolei, et al.
Published: (2024)
by: Zhang, Shaolei, et al.
Published: (2024)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
by: Wei, Kun, et al.
Published: (2023)
by: Wei, Kun, et al.
Published: (2023)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
by: Li, Xuanchen, et al.
Published: (2025)
by: Li, Xuanchen, et al.
Published: (2025)
NAIST Simultaneous Speech Translation System for IWSLT 2024
by: Ko, Yuka, et al.
Published: (2024)
by: Ko, Yuka, et al.
Published: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
by: Deng, Keqi, et al.
Published: (2025)
by: Deng, Keqi, et al.
Published: (2025)
Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Similar Items
-
FASST: Fast LLM-based Simultaneous Speech Translation
by: Ouyang, Siqi, et al.
Published: (2024) -
CMU's IWSLT 2025 Simultaneous Speech Translation System
by: Ouyang, Siqi, et al.
Published: (2025) -
InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model
by: Ouyang, Siqi, et al.
Published: (2025) -
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
by: Xu, Xi, et al.
Published: (2024) -
Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
by: Ouyang, Siqi, et al.
Published: (2026)