MUSCAT: MUltilingual, SCientific ConversATion Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Sinhamahapatra, Supriti, Nguyen, Thai-Binh, Oğuz, Yiğit, Ugan, Enes, Niehues, Jan, Waibel, Alexander |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
por: Sinhamahapatra, Supriti, et al.
Publicado: (2025)
por: Sinhamahapatra, Supriti, et al.
Publicado: (2025)
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations
por: Lee, Hyunji, et al.
Publicado: (2024)
por: Lee, Hyunji, et al.
Publicado: (2024)
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
por: Retkowski, Fabian, et al.
Publicado: (2026)
por: Retkowski, Fabian, et al.
Publicado: (2026)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
por: Nguyen, Thai-Binh, et al.
Publicado: (2023)
por: Nguyen, Thai-Binh, et al.
Publicado: (2023)
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
por: Koneru, Sai, et al.
Publicado: (2025)
por: Koneru, Sai, et al.
Publicado: (2025)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
por: Eyiokur, Fevziye Irem, et al.
Publicado: (2024)
por: Eyiokur, Fevziye Irem, et al.
Publicado: (2024)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
por: Koneru, Sai, et al.
Publicado: (2025)
por: Koneru, Sai, et al.
Publicado: (2025)
Bayesian Low-Rank Factorization for Robust Model Adaptation
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
Adapting Language Balance in Code-Switching Speech
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
Cocktail-Party Audio-Visual Speech Recognition
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
por: Koneru, Sai, et al.
Publicado: (2024)
por: Koneru, Sai, et al.
Publicado: (2024)
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
por: Li, Zhaolin, et al.
Publicado: (2025)
por: Li, Zhaolin, et al.
Publicado: (2025)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
por: Huber, Christian, et al.
Publicado: (2023)
por: Huber, Christian, et al.
Publicado: (2023)
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
Towards continually learning new languages
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions
por: Retkowski, Fabian, et al.
Publicado: (2024)
por: Retkowski, Fabian, et al.
Publicado: (2024)
Do What I Say: A Spoken Prompt Dataset for Instruction-Following
por: Züfle, Maike, et al.
Publicado: (2026)
por: Züfle, Maike, et al.
Publicado: (2026)
Summarizing Speech: A Comprehensive Survey
por: Retkowski, Fabian, et al.
Publicado: (2025)
por: Retkowski, Fabian, et al.
Publicado: (2025)
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
por: Retkowski, Fabian, et al.
Publicado: (2025)
por: Retkowski, Fabian, et al.
Publicado: (2025)
Zero-Shot Strategies for Length-Controllable Summarization
por: Retkowski, Fabian, et al.
Publicado: (2024)
por: Retkowski, Fabian, et al.
Publicado: (2024)
When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR
por: Züfle, Maike, et al.
Publicado: (2026)
por: Züfle, Maike, et al.
Publicado: (2026)
Conditions for Catastrophic Forgetting in Multilingual Translation
por: Liu, Danni, et al.
Publicado: (2025)
por: Liu, Danni, et al.
Publicado: (2025)
SARA: Stress Test Reasoning in Audio Deepfake Detection
por: Nguyen, Binh, et al.
Publicado: (2026)
por: Nguyen, Binh, et al.
Publicado: (2026)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
por: Akti, Seymanur, et al.
Publicado: (2026)
por: Akti, Seymanur, et al.
Publicado: (2026)
Contrastive Learning for Task-Independent SpeechLLM-Pretraining
por: Züfle, Maike, et al.
Publicado: (2024)
por: Züfle, Maike, et al.
Publicado: (2024)
Context Biasing for Pronunciation-Orthography Mismatch in Automatic Speech Recognition
por: Huber, Christian, et al.
Publicado: (2025)
por: Huber, Christian, et al.
Publicado: (2025)
Continuously Learning New Words in Automatic Speech Recognition
por: Huber, Christian, et al.
Publicado: (2024)
por: Huber, Christian, et al.
Publicado: (2024)
Multimodal In-context Learning for ASR of Low-resource Languages
por: Li, Zhaolin, et al.
Publicado: (2026)
por: Li, Zhaolin, et al.
Publicado: (2026)
The Silent Curriculum: How Does LLM Monoculture Shape Educational Content and Its Accessibility?
por: Priyanshu, Aman, et al.
Publicado: (2024)
por: Priyanshu, Aman, et al.
Publicado: (2024)
Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs
por: Liu, Danni, et al.
Publicado: (2025)
por: Liu, Danni, et al.
Publicado: (2025)
In-context Language Learning for Endangered Languages in Speech Recognition
por: Li, Zhaolin, et al.
Publicado: (2025)
por: Li, Zhaolin, et al.
Publicado: (2025)
How Transferable are Attribute Controllers on Pretrained Multilingual Translation Models?
por: Liu, Danni, et al.
Publicado: (2023)
por: Liu, Danni, et al.
Publicado: (2023)
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
por: Priyanshu, Aman, et al.
Publicado: (2024)
por: Priyanshu, Aman, et al.
Publicado: (2024)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
por: Nguyen, Thai-Binh, et al.
Publicado: (2025)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
por: Broy, Frederik, et al.
Publicado: (2025)
por: Broy, Frederik, et al.
Publicado: (2025)
Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration Approach
por: Li, Siqi, et al.
Publicado: (2024)
por: Li, Siqi, et al.
Publicado: (2024)
Ejemplares similares
-
Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
por: Sinhamahapatra, Supriti, et al.
Publicado: (2025) -
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations
por: Lee, Hyunji, et al.
Publicado: (2024) -
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
por: Retkowski, Fabian, et al.
Publicado: (2026) -
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024) -
Convoifilter: A case study of doing cocktail party speech recognition
por: Nguyen, Thai-Binh, et al.
Publicado: (2023)