A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Thai-Binh, Zmolikova, Katerina, Ma, Pingchuan, Pham, Ngoc Quan, Fuegen, Christian, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cocktail-Party Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024)
by: Nguyen, Thai-Binh, et al.
Published: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
by: Nguyen, Thai-Binh, et al.
Published: (2023)
by: Nguyen, Thai-Binh, et al.
Published: (2023)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
by: Koneru, Sai, et al.
Published: (2024)
by: Koneru, Sai, et al.
Published: (2024)
Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen Languages
by: Mullov, Carlos, et al.
Published: (2024)
by: Mullov, Carlos, et al.
Published: (2024)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
MUSCAT: MUltilingual, SCientific ConversATion Benchmark
by: Sinhamahapatra, Supriti, et al.
Published: (2026)
by: Sinhamahapatra, Supriti, et al.
Published: (2026)
Towards continually learning new languages
by: Pham, Ngoc-Quan, et al.
Published: (2022)
by: Pham, Ngoc-Quan, et al.
Published: (2022)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
by: Retkowski, Fabian, et al.
Published: (2026)
by: Retkowski, Fabian, et al.
Published: (2026)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Bayesian Low-Rank Factorization for Robust Model Adaptation
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Adapting Language Balance in Code-Switching Speech
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
by: Loc, Ngoc Phan Phuoc, et al.
Published: (2026)
by: Loc, Ngoc Phan Phuoc, et al.
Published: (2026)
From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions
by: Retkowski, Fabian, et al.
Published: (2024)
by: Retkowski, Fabian, et al.
Published: (2024)
Comparing Without Saying: A Dataset and Benchmark for Implicit Comparative Opinion Mining from Same-User Reviews
by: Nguyen, Thanh-Lam T., et al.
Published: (2026)
by: Nguyen, Thanh-Lam T., et al.
Published: (2026)
Context Biasing for Pronunciation-Orthography Mismatch in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2025)
by: Huber, Christian, et al.
Published: (2025)
Continuously Learning New Words in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction
by: Hao, Xiang, et al.
Published: (2023)
by: Hao, Xiang, et al.
Published: (2023)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Zero-Shot Strategies for Length-Controllable Summarization
by: Retkowski, Fabian, et al.
Published: (2024)
by: Retkowski, Fabian, et al.
Published: (2024)
A Benchmark Dataset and Evaluation Framework for Vietnamese Large Language Models in Customer Support
by: Nguyen, Long S. T., et al.
Published: (2025)
by: Nguyen, Long S. T., et al.
Published: (2025)
SARA: Stress Test Reasoning in Audio Deepfake Detection
by: Nguyen, Binh, et al.
Published: (2026)
by: Nguyen, Binh, et al.
Published: (2026)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
VM14K: First Vietnamese Medical Benchmark
by: Nguyen, Thong, et al.
Published: (2025)
by: Nguyen, Thong, et al.
Published: (2025)
BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots
by: Svikhnushina, Ekaterina, et al.
Published: (2024)
by: Svikhnushina, Ekaterina, et al.
Published: (2024)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
by: Dong, Nguyen Tien, et al.
Published: (2025)
by: Dong, Nguyen Tien, et al.
Published: (2025)
Swift Cross-Dataset Pruning: Enhancing Fine-Tuning Efficiency in Natural Language Understanding
by: Nguyen, Binh-Nguyen, et al.
Published: (2025)
by: Nguyen, Binh-Nguyen, et al.
Published: (2025)
MPCEval: A Benchmark for Multi-Party Conversation Generation
by: Zhang, Minxing, et al.
Published: (2026)
by: Zhang, Minxing, et al.
Published: (2026)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
by: Nguyen, Binh, et al.
Published: (2025)
by: Nguyen, Binh, et al.
Published: (2025)
FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
by: Luo, Junyu, et al.
Published: (2025)
by: Luo, Junyu, et al.
Published: (2025)
Leveraging Sentence-oriented Augmentation and Transformer-Based Architecture for Vietnamese-Bahnaric Translation
by: Nguyen, Tan Sang, et al.
Published: (2026)
by: Nguyen, Tan Sang, et al.
Published: (2026)
Similar Items
-
Cocktail-Party Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025) -
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024) -
Convoifilter: A case study of doing cocktail party speech recognition
by: Nguyen, Thai-Binh, et al.
Published: (2023) -
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025) -
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
by: Koneru, Sai, et al.
Published: (2024)