WESR: Scaling and Evaluating Word-level Event-Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Chenchen, Huang, Kexin, Fan, Liwei, Tu, Qian, Jiang, Botian, Zhang, Dong, Yin, Linqi, Li, Shimin, Fei, Zhaoye, Cheng, Qinyuan, Qiu, Xipeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
di: Huang, Kexin, et al.
Pubblicazione: (2025)
di: Huang, Kexin, et al.
Pubblicazione: (2025)
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
di: Huang, Kexin, et al.
Pubblicazione: (2026)
di: Huang, Kexin, et al.
Pubblicazione: (2026)
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
di: Deng, Ruifan, et al.
Pubblicazione: (2025)
di: Deng, Ruifan, et al.
Pubblicazione: (2025)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
di: Gong, Yitian, et al.
Pubblicazione: (2025)
di: Gong, Yitian, et al.
Pubblicazione: (2025)
SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
di: Zhang, Dong, et al.
Pubblicazione: (2024)
di: Zhang, Dong, et al.
Pubblicazione: (2024)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
di: Gong, Yitian, et al.
Pubblicazione: (2026)
di: Gong, Yitian, et al.
Pubblicazione: (2026)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
di: Zhang, Xin, et al.
Pubblicazione: (2023)
di: Zhang, Xin, et al.
Pubblicazione: (2023)
MOSS-TTSD: Text to Spoken Dialogue Generation
di: Zhang, Yuqian, et al.
Pubblicazione: (2026)
di: Zhang, Yuqian, et al.
Pubblicazione: (2026)
SpeechAlign: Aligning Speech Generation to Human Preferences
di: Zhang, Dong, et al.
Pubblicazione: (2024)
di: Zhang, Dong, et al.
Pubblicazione: (2024)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
di: Wang, Peng, et al.
Pubblicazione: (2026)
di: Wang, Peng, et al.
Pubblicazione: (2026)
MOSS-TTS Technical Report
di: Gong, Yitian, et al.
Pubblicazione: (2026)
di: Gong, Yitian, et al.
Pubblicazione: (2026)
VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
di: Zhan, Jun, et al.
Pubblicazione: (2025)
di: Zhan, Jun, et al.
Pubblicazione: (2025)
MOSS-Audio Technical Report
di: Yang, Chen, et al.
Pubblicazione: (2026)
di: Yang, Chen, et al.
Pubblicazione: (2026)
MOSS Transcribe Diarize Technical Report
di: AI, MOSI., et al.
Pubblicazione: (2026)
di: AI, MOSI., et al.
Pubblicazione: (2026)
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
di: Huang, Dawei, et al.
Pubblicazione: (2026)
di: Huang, Dawei, et al.
Pubblicazione: (2026)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
di: Jing, Ruihao, et al.
Pubblicazione: (2025)
di: Jing, Ruihao, et al.
Pubblicazione: (2025)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
di: Hu, Ke, et al.
Pubblicazione: (2025)
di: Hu, Ke, et al.
Pubblicazione: (2025)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
di: Hu, Cheng-Hung, et al.
Pubblicazione: (2025)
di: Hu, Cheng-Hung, et al.
Pubblicazione: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
di: Park, Chanho, et al.
Pubblicazione: (2024)
di: Park, Chanho, et al.
Pubblicazione: (2024)
Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
di: Yin, Han, et al.
Pubblicazione: (2025)
di: Yin, Han, et al.
Pubblicazione: (2025)
Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition
di: Ahmed, Tauseef, et al.
Pubblicazione: (2026)
di: Ahmed, Tauseef, et al.
Pubblicazione: (2026)
Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition
di: Mei, Yuxiang, et al.
Pubblicazione: (2026)
di: Mei, Yuxiang, et al.
Pubblicazione: (2026)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
di: Zhang, Hanlin, et al.
Pubblicazione: (2026)
di: Zhang, Hanlin, et al.
Pubblicazione: (2026)
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
di: Wang, Tianduo, et al.
Pubblicazione: (2025)
di: Wang, Tianduo, et al.
Pubblicazione: (2025)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2026)
di: Chen, Youjun, et al.
Pubblicazione: (2026)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
di: Nozawa, Kento, et al.
Pubblicazione: (2024)
di: Nozawa, Kento, et al.
Pubblicazione: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
di: Akinrintoyo, Emmanuel, et al.
Pubblicazione: (2025)
di: Akinrintoyo, Emmanuel, et al.
Pubblicazione: (2025)
Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
RAS: a Reliability Oriented Metric for Automatic Speech Recognition
di: Huang, Wenbin, et al.
Pubblicazione: (2026)
di: Huang, Wenbin, et al.
Pubblicazione: (2026)
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
di: Bhogale, Kaushal Santosh, et al.
Pubblicazione: (2026)
di: Bhogale, Kaushal Santosh, et al.
Pubblicazione: (2026)
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
di: Guo, Jiaxin, et al.
Pubblicazione: (2024)
di: Guo, Jiaxin, et al.
Pubblicazione: (2024)
Multi-blank Transducers for Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2022)
di: Xu, Hainan, et al.
Pubblicazione: (2022)
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
di: Sundar, Anirudh S., et al.
Pubblicazione: (2023)
di: Sundar, Anirudh S., et al.
Pubblicazione: (2023)
ParaGSE: Parallel Generative Speech Enhancement with Group-Vector-Quantization-based Neural Speech Codec
di: Liu, Fei, et al.
Pubblicazione: (2026)
di: Liu, Fei, et al.
Pubblicazione: (2026)
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
di: Zhang, Leying, et al.
Pubblicazione: (2024)
di: Zhang, Leying, et al.
Pubblicazione: (2024)
Bypassing Direct Reconstruction: Speech Detection from MEG via Large-Scale Audio Retrieval
di: Xiao, Boda, et al.
Pubblicazione: (2026)
di: Xiao, Boda, et al.
Pubblicazione: (2026)
Documenti analoghi
-
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
di: Huang, Kexin, et al.
Pubblicazione: (2025) -
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
di: Huang, Kexin, et al.
Pubblicazione: (2026) -
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
di: Deng, Ruifan, et al.
Pubblicazione: (2025) -
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
di: Gong, Yitian, et al.
Pubblicazione: (2025) -
SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
di: Zhang, Dong, et al.
Pubblicazione: (2024)