CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Heyang, Wang, Yuhao, Cheng, Ziyang, Wu, Ronghua, Gu, Qunshan, Wang, Yanfeng, Wang, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
di: Liu, Heyang, et al.
Pubblicazione: (2025)
di: Liu, Heyang, et al.
Pubblicazione: (2025)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
di: Cheng, Ziyang, et al.
Pubblicazione: (2026)
di: Cheng, Ziyang, et al.
Pubblicazione: (2026)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
di: Liu, Heyang, et al.
Pubblicazione: (2025)
di: Liu, Heyang, et al.
Pubblicazione: (2025)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
LaSR: Context-Aware Speech Recognition via Latent Reasoning
di: Liu, Heyang, et al.
Pubblicazione: (2026)
di: Liu, Heyang, et al.
Pubblicazione: (2026)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
di: Hou, Yixuan, et al.
Pubblicazione: (2025)
di: Hou, Yixuan, et al.
Pubblicazione: (2025)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
di: Liu, Hongcheng, et al.
Pubblicazione: (2025)
di: Liu, Hongcheng, et al.
Pubblicazione: (2025)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
di: Quang, Trung Nguyen, et al.
Pubblicazione: (2026)
di: Quang, Trung Nguyen, et al.
Pubblicazione: (2026)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
di: Liu, Heyang, et al.
Pubblicazione: (2024)
di: Liu, Heyang, et al.
Pubblicazione: (2024)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
di: Yan, Brian, et al.
Pubblicazione: (2025)
di: Yan, Brian, et al.
Pubblicazione: (2025)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
di: Wang, Yuhao, et al.
Pubblicazione: (2024)
di: Wang, Yuhao, et al.
Pubblicazione: (2024)
Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech
di: Wotherspoon, Shannon, et al.
Pubblicazione: (2024)
di: Wotherspoon, Shannon, et al.
Pubblicazione: (2024)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
di: Wang, Yuhao, et al.
Pubblicazione: (2024)
di: Wang, Yuhao, et al.
Pubblicazione: (2024)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
Multilingual Stutter Event Detection for English, German, and Mandarin Speech
di: Haas, Felix, et al.
Pubblicazione: (2026)
di: Haas, Felix, et al.
Pubblicazione: (2026)
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
di: Liu, Hongcheng, et al.
Pubblicazione: (2024)
di: Liu, Hongcheng, et al.
Pubblicazione: (2024)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
di: Du, Yuhao, et al.
Pubblicazione: (2025)
di: Du, Yuhao, et al.
Pubblicazione: (2025)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
di: Yang, Chih-Kai, et al.
Pubblicazione: (2023)
di: Yang, Chih-Kai, et al.
Pubblicazione: (2023)
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
di: Paik, Gio, et al.
Pubblicazione: (2025)
di: Paik, Gio, et al.
Pubblicazione: (2025)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
di: Heakl, Ahmed, et al.
Pubblicazione: (2024)
di: Heakl, Ahmed, et al.
Pubblicazione: (2024)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
di: Xu, Jing, et al.
Pubblicazione: (2026)
di: Xu, Jing, et al.
Pubblicazione: (2026)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models
di: Zhao, Zihan, et al.
Pubblicazione: (2023)
di: Zhao, Zihan, et al.
Pubblicazione: (2023)
Soundwave: Less is More for Speech-Text Alignment in LLMs
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
di: Feng, Sheng, et al.
Pubblicazione: (2024)
di: Feng, Sheng, et al.
Pubblicazione: (2024)
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
di: Feng, Pengchao, et al.
Pubblicazione: (2025)
di: Feng, Pengchao, et al.
Pubblicazione: (2025)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
di: Huang, Hukai, et al.
Pubblicazione: (2024)
di: Huang, Hukai, et al.
Pubblicazione: (2024)
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
di: Zhang, Haoyang, et al.
Pubblicazione: (2025)
di: Zhang, Haoyang, et al.
Pubblicazione: (2025)
Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language
di: Sharma, Yash, et al.
Pubblicazione: (2024)
di: Sharma, Yash, et al.
Pubblicazione: (2024)
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
di: Yang, Tzu-Ting, et al.
Pubblicazione: (2024)
di: Yang, Tzu-Ting, et al.
Pubblicazione: (2024)
DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech
di: Guo, Haotian, et al.
Pubblicazione: (2025)
di: Guo, Haotian, et al.
Pubblicazione: (2025)
Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
di: Liao, Yusheng, et al.
Pubblicazione: (2024)
di: Liao, Yusheng, et al.
Pubblicazione: (2024)
Code-Mixed Telugu-English Hate Speech Detection
di: Kakarla, Santhosh, et al.
Pubblicazione: (2025)
di: Kakarla, Santhosh, et al.
Pubblicazione: (2025)
Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility
di: Wang, Sheng-Fu, et al.
Pubblicazione: (2025)
di: Wang, Sheng-Fu, et al.
Pubblicazione: (2025)
Methods of Automatic Matrix Language Determination for Code-Switched Speech
di: Iakovenko, Olga, et al.
Pubblicazione: (2024)
di: Iakovenko, Olga, et al.
Pubblicazione: (2024)
Enhancing Voice Wake-Up for Dysarthria: Mandarin Dysarthria Speech Corpus Release and Customized System Design
di: Gao, Ming, et al.
Pubblicazione: (2024)
di: Gao, Ming, et al.
Pubblicazione: (2024)
M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
di: Chen, Zhe, et al.
Pubblicazione: (2024)
di: Chen, Zhe, et al.
Pubblicazione: (2024)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
di: Jiang, Feng, et al.
Pubblicazione: (2025)
di: Jiang, Feng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
di: Liu, Heyang, et al.
Pubblicazione: (2025) -
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
di: Cheng, Ziyang, et al.
Pubblicazione: (2026) -
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
di: Liu, Heyang, et al.
Pubblicazione: (2025) -
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
di: Wang, Yuhao, et al.
Pubblicazione: (2025) -
LaSR: Context-Aware Speech Recognition via Latent Reasoning
di: Liu, Heyang, et al.
Pubblicazione: (2026)