KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koneru, Sai, Züfle, Maike, Nguyen, Thai-Binh, Akti, Seymanur, Niehues, Jan, Waibel, Alexander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918025191489536
author Koneru, Sai
Züfle, Maike
Nguyen, Thai-Binh
Akti, Seymanur
Niehues, Jan
Waibel, Alexander
author_facet Koneru, Sai
Züfle, Maike
Nguyen, Thai-Binh
Akti, Seymanur
Niehues, Jan
Waibel, Alexander
contents The scope of the International Workshop on Spoken Language Translation (IWSLT) has recently broadened beyond traditional Speech Translation (ST) to encompass a wider array of tasks, including Speech Question Answering and Summarization. This shift is partly driven by the growing capabilities of modern systems, particularly with the success of Large Language Models (LLMs). In this paper, we present the Karlsruhe Institute of Technology's submissions for the Offline ST and Instruction Following (IF) tracks, where we leverage LLMs to enhance performance across all tasks. For the Offline ST track, we propose a pipeline that employs multiple automatic speech recognition systems, whose outputs are fused using an LLM with document-level context. This is followed by a two-step translation process, incorporating additional refinement step to improve translation quality. For the IF track, we develop an end-to-end model that integrates a speech encoder with an LLM to perform a wide range of instruction-following tasks. We complement it with a final document-level refinement stage to further enhance output quality by using contextual information.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13036
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
Koneru, Sai
Züfle, Maike
Nguyen, Thai-Binh
Akti, Seymanur
Niehues, Jan
Waibel, Alexander
Computation and Language
Artificial Intelligence
The scope of the International Workshop on Spoken Language Translation (IWSLT) has recently broadened beyond traditional Speech Translation (ST) to encompass a wider array of tasks, including Speech Question Answering and Summarization. This shift is partly driven by the growing capabilities of modern systems, particularly with the success of Large Language Models (LLMs). In this paper, we present the Karlsruhe Institute of Technology's submissions for the Offline ST and Instruction Following (IF) tracks, where we leverage LLMs to enhance performance across all tasks. For the Offline ST track, we propose a pipeline that employs multiple automatic speech recognition systems, whose outputs are fused using an LLM with document-level context. This is followed by a two-step translation process, incorporating additional refinement step to improve translation quality. For the IF track, we develop an end-to-end model that integrates a speech encoder with an LLM to perform a wide range of instruction-following tasks. We complement it with a final document-level refinement stage to further enhance output quality by using contextual information.
title KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.13036