BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
Fuente:
arXiv
Saved in:
| Main Authors: | Koneru, Sai, Retkowski, Fabian, Huber, Christian, Hilgert, Lukas, Akti, Seymanur, Ugan, Enes Yavuz, Waibel, Alexander, Niehues, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
by: Akti, Seymanur, et al.
Published: (2025)
by: Akti, Seymanur, et al.
Published: (2025)
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
by: Li, Zhaolin, et al.
Published: (2025)
by: Li, Zhaolin, et al.
Published: (2025)
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
by: Retkowski, Fabian, et al.
Published: (2026)
by: Retkowski, Fabian, et al.
Published: (2026)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Bayesian Low-Rank Factorization for Robust Model Adaptation
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Adapting Language Balance in Code-Switching Speech
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
by: Koneru, Sai, et al.
Published: (2024)
by: Koneru, Sai, et al.
Published: (2024)
Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions
by: Retkowski, Fabian, et al.
Published: (2024)
by: Retkowski, Fabian, et al.
Published: (2024)
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Zero-Shot Strategies for Length-Controllable Summarization
by: Retkowski, Fabian, et al.
Published: (2024)
by: Retkowski, Fabian, et al.
Published: (2024)
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
MUSCAT: MUltilingual, SCientific ConversATion Benchmark
by: Sinhamahapatra, Supriti, et al.
Published: (2026)
by: Sinhamahapatra, Supriti, et al.
Published: (2026)
Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
The AI Co-Ethnographer: How Far Can Automation Take Qualitative Research?
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Summarizing Speech: A Comprehensive Survey
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024)
by: Yaman, Dogucan, et al.
Published: (2024)
Do What I Say: A Spoken Prompt Dataset for Instruction-Following
by: Züfle, Maike, et al.
Published: (2026)
by: Züfle, Maike, et al.
Published: (2026)
Plug, Play, and Fuse: Zero-Shot Joint Decoding via Word-Level Re-ranking Across Diverse Vocabularies
by: Koneru, Sai, et al.
Published: (2024)
by: Koneru, Sai, et al.
Published: (2024)
Quality-Aware Decoding: Unifying Quality Estimation and Decoding
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Contextual Refinement of Translations: Large Language Models for Sentence and Document-Level Post-Editing
by: Koneru, Sai, et al.
Published: (2023)
by: Koneru, Sai, et al.
Published: (2023)
Towards continually learning new languages
by: Pham, Ngoc-Quan, et al.
Published: (2022)
by: Pham, Ngoc-Quan, et al.
Published: (2022)
Conditions for Catastrophic Forgetting in Multilingual Translation
by: Liu, Danni, et al.
Published: (2025)
by: Liu, Danni, et al.
Published: (2025)
How Transferable are Attribute Controllers on Pretrained Multilingual Translation Models?
by: Liu, Danni, et al.
Published: (2023)
by: Liu, Danni, et al.
Published: (2023)
MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations
by: Scott, Aaron, et al.
Published: (2025)
by: Scott, Aaron, et al.
Published: (2025)
Incremental Learning of Humanoid Robot Behavior from Natural Interaction and Large Language Models
by: Bärmann, Leonard, et al.
Published: (2023)
by: Bärmann, Leonard, et al.
Published: (2023)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Context Biasing for Pronunciation-Orthography Mismatch in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2025)
by: Huber, Christian, et al.
Published: (2025)
Continuously Learning New Words in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Multimodal In-context Learning for ASR of Low-resource Languages
by: Li, Zhaolin, et al.
Published: (2026)
by: Li, Zhaolin, et al.
Published: (2026)
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations
by: Lee, Hyunji, et al.
Published: (2024)
by: Lee, Hyunji, et al.
Published: (2024)
Reseña de "Stable Isotopes and Archaeology in Southern South America. Hunter-Gatherers, Pastoralism and Agriculture" Editado por R. Barberena, A. Gil, G. Neme y R. Tykot
by: Andrew Ugan
Published: (2009)
by: Andrew Ugan
Published: (2009)
Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
by: Retkowski, Jan, et al.
Published: (2024)
by: Retkowski, Jan, et al.
Published: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024)
by: Nguyen, Thai-Binh, et al.
Published: (2024)
Similar Items
-
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
by: Koneru, Sai, et al.
Published: (2025) -
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026) -
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023) -
OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion
by: Koneru, Sai, et al.
Published: (2025) -
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
by: Akti, Seymanur, et al.
Published: (2025)