KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
Fuente:
arXiv
Saved in:
| Main Authors: | Koneru, Sai, Züfle, Maike, Nguyen, Thai-Binh, Akti, Seymanur, Niehues, Jan, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
by: Koneru, Sai, et al.
Published: (2024)
by: Koneru, Sai, et al.
Published: (2024)
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
by: Li, Zhaolin, et al.
Published: (2025)
by: Li, Zhaolin, et al.
Published: (2025)
BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
by: Retkowski, Fabian, et al.
Published: (2026)
by: Retkowski, Fabian, et al.
Published: (2026)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
by: Akti, Seymanur, et al.
Published: (2025)
by: Akti, Seymanur, et al.
Published: (2025)
MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations
by: Scott, Aaron, et al.
Published: (2025)
by: Scott, Aaron, et al.
Published: (2025)
Contextual Refinement of Translations: Large Language Models for Sentence and Document-Level Post-Editing
by: Koneru, Sai, et al.
Published: (2023)
by: Koneru, Sai, et al.
Published: (2023)
Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025
by: Macháček, Dominik, et al.
Published: (2025)
by: Macháček, Dominik, et al.
Published: (2025)
Contrastive Learning for Task-Independent SpeechLLM-Pretraining
by: Züfle, Maike, et al.
Published: (2024)
by: Züfle, Maike, et al.
Published: (2024)
Plug, Play, and Fuse: Zero-Shot Joint Decoding via Word-Level Re-ranking Across Diverse Vocabularies
by: Koneru, Sai, et al.
Published: (2024)
by: Koneru, Sai, et al.
Published: (2024)
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
CMU's IWSLT 2024 Simultaneous Speech Translation System
by: Xu, Xi, et al.
Published: (2024)
by: Xu, Xi, et al.
Published: (2024)
Do What I Say: A Spoken Prompt Dataset for Instruction-Following
by: Züfle, Maike, et al.
Published: (2026)
by: Züfle, Maike, et al.
Published: (2026)
When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR
by: Züfle, Maike, et al.
Published: (2026)
by: Züfle, Maike, et al.
Published: (2026)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
Summarizing Speech: A Comprehensive Survey
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
Cocktail-Party Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
In-context Language Learning for Endangered Languages in Speech Recognition
by: Li, Zhaolin, et al.
Published: (2025)
by: Li, Zhaolin, et al.
Published: (2025)
How Transferable are Attribute Controllers on Pretrained Multilingual Translation Models?
by: Liu, Danni, et al.
Published: (2023)
by: Liu, Danni, et al.
Published: (2023)
Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
by: Attanasio, Giuseppe, et al.
Published: (2025)
by: Attanasio, Giuseppe, et al.
Published: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024)
by: Nguyen, Thai-Binh, et al.
Published: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
by: Nguyen, Thai-Binh, et al.
Published: (2023)
by: Nguyen, Thai-Binh, et al.
Published: (2023)
Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
by: Broy, Frederik, et al.
Published: (2025)
by: Broy, Frederik, et al.
Published: (2025)
MUSCAT: MUltilingual, SCientific ConversATion Benchmark
by: Sinhamahapatra, Supriti, et al.
Published: (2026)
by: Sinhamahapatra, Supriti, et al.
Published: (2026)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
Augmenting Automatic Speech Recognition Models with Disfluency Detection
by: Amann, Robin, et al.
Published: (2024)
by: Amann, Robin, et al.
Published: (2024)
Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences
by: Koneru, Sai, et al.
Published: (2023)
by: Koneru, Sai, et al.
Published: (2023)
Early-Exit and Instant Confidence Translation Quality Estimation
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
CMU's IWSLT 2025 Simultaneous Speech Translation System
by: Ouyang, Siqi, et al.
Published: (2025)
by: Ouyang, Siqi, et al.
Published: (2025)
Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
by: Sinhamahapatra, Supriti, et al.
Published: (2025)
by: Sinhamahapatra, Supriti, et al.
Published: (2025)
Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs
by: Liu, Danni, et al.
Published: (2025)
by: Liu, Danni, et al.
Published: (2025)
Multimodal In-context Learning for ASR of Low-resource Languages
by: Li, Zhaolin, et al.
Published: (2026)
by: Li, Zhaolin, et al.
Published: (2026)
Context Selection for Hypothesis and Statistical Evidence Extraction from Full-Text Scientific Articles
by: Koneru, Sai, et al.
Published: (2026)
by: Koneru, Sai, et al.
Published: (2026)
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
by: Sperber, Matthias, et al.
Published: (2024)
by: Sperber, Matthias, et al.
Published: (2024)
Similar Items
-
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
by: Koneru, Sai, et al.
Published: (2024) -
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
by: Li, Zhaolin, et al.
Published: (2025) -
BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
by: Koneru, Sai, et al.
Published: (2025) -
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026) -
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
by: Retkowski, Fabian, et al.
Published: (2026)