Salvato in:
| Autori principali: | Wang, Mengzhi, Xiong, Shifu, Wan, Genshun, Chen, Hang, Gao, Jianqing, Dai, Lirong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2409.17603 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Lightweight Transducer Based on Frame-Level Criterion
di: Wan, Genshun, et al.
Pubblicazione: (2024)
di: Wan, Genshun, et al.
Pubblicazione: (2024)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
di: Wan, Genshun, et al.
Pubblicazione: (2026)
di: Wan, Genshun, et al.
Pubblicazione: (2026)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices
di: Velikovich, Leonid, et al.
Pubblicazione: (2024)
di: Velikovich, Leonid, et al.
Pubblicazione: (2024)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
di: Wang, He, et al.
Pubblicazione: (2025)
di: Wang, He, et al.
Pubblicazione: (2025)
OCR-Enhanced Multimodal ASR Can Read While Listening
di: Chen, Junli, et al.
Pubblicazione: (2026)
di: Chen, Junli, et al.
Pubblicazione: (2026)
The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge
di: Niu, Shutong, et al.
Pubblicazione: (2024)
di: Niu, Shutong, et al.
Pubblicazione: (2024)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
di: Wei, Kun, et al.
Pubblicazione: (2023)
di: Wei, Kun, et al.
Pubblicazione: (2023)
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
di: Wei, Sheng-Lun, et al.
Pubblicazione: (2026)
di: Wei, Sheng-Lun, et al.
Pubblicazione: (2026)
Multi-class Decoding of Attended Speaker Direction Using Electroencephalogram and Audio Spatial Spectrum
di: Zhang, Yuanming, et al.
Pubblicazione: (2024)
di: Zhang, Yuanming, et al.
Pubblicazione: (2024)
Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience
di: Jiang, Xilin, et al.
Pubblicazione: (2024)
di: Jiang, Xilin, et al.
Pubblicazione: (2024)
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
Deep Learning for Assessment of Oral Reading Fluency
di: Vaidya, Mithilesh, et al.
Pubblicazione: (2024)
di: Vaidya, Mithilesh, et al.
Pubblicazione: (2024)
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
di: Liu, Zihan, et al.
Pubblicazione: (2025)
di: Liu, Zihan, et al.
Pubblicazione: (2025)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
Speech Separation based on Contrastive Learning and Deep Modularization
di: Ochieng, Peter
Pubblicazione: (2023)
di: Ochieng, Peter
Pubblicazione: (2023)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
di: Zhao, Zhixian, et al.
Pubblicazione: (2025)
di: Zhao, Zhixian, et al.
Pubblicazione: (2025)
Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
di: Wu, Wen, et al.
Pubblicazione: (2023)
di: Wu, Wen, et al.
Pubblicazione: (2023)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
di: Lin, Zhennan, et al.
Pubblicazione: (2025)
di: Lin, Zhennan, et al.
Pubblicazione: (2025)
A Deep Learning Automatic Speech Recognition Model for Shona Language
di: Sirora, Leslie Wellington, et al.
Pubblicazione: (2025)
di: Sirora, Leslie Wellington, et al.
Pubblicazione: (2025)
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
di: Gu, Yue, et al.
Pubblicazione: (2025)
di: Gu, Yue, et al.
Pubblicazione: (2025)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
di: Shakeel, Muhammad, et al.
Pubblicazione: (2024)
di: Shakeel, Muhammad, et al.
Pubblicazione: (2024)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
di: Sudo, Yui, et al.
Pubblicazione: (2025)
di: Sudo, Yui, et al.
Pubblicazione: (2025)
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
di: Dowerah, Sandipana, et al.
Pubblicazione: (2025)
di: Dowerah, Sandipana, et al.
Pubblicazione: (2025)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
di: Sun, Guangzhi, et al.
Pubblicazione: (2022)
di: Sun, Guangzhi, et al.
Pubblicazione: (2022)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
di: Shakeel, Muhammad, et al.
Pubblicazione: (2026)
di: Shakeel, Muhammad, et al.
Pubblicazione: (2026)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her
di: Gao, Heting, et al.
Pubblicazione: (2025)
di: Gao, Heting, et al.
Pubblicazione: (2025)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
di: Nozawa, Kento, et al.
Pubblicazione: (2024)
di: Nozawa, Kento, et al.
Pubblicazione: (2024)
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
di: Sudo, Yui, et al.
Pubblicazione: (2025)
di: Sudo, Yui, et al.
Pubblicazione: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
di: Shao, Hang, et al.
Pubblicazione: (2023)
di: Shao, Hang, et al.
Pubblicazione: (2023)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
di: Nakagome, Yu, et al.
Pubblicazione: (2025)
di: Nakagome, Yu, et al.
Pubblicazione: (2025)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
di: Wang, Haoyu, et al.
Pubblicazione: (2022)
di: Wang, Haoyu, et al.
Pubblicazione: (2022)
Task-Agnostic Structured Pruning of Speech Representation Models
di: Wang, Haoyu, et al.
Pubblicazione: (2023)
di: Wang, Haoyu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Lightweight Transducer Based on Frame-Level Criterion
di: Wan, Genshun, et al.
Pubblicazione: (2024) -
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
di: Wan, Genshun, et al.
Pubblicazione: (2026) -
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025) -
Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices
di: Velikovich, Leonid, et al.
Pubblicazione: (2024) -
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
di: Wang, He, et al.
Pubblicazione: (2025)