Salvato in:
| Autori principali: | Liu, Zefang, Zhu, Chenyang, Cho, Sangwoo, Zhang, Shi-Xiong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.18721 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
di: Sun, Haoqin, et al.
Pubblicazione: (2024)
di: Sun, Haoqin, et al.
Pubblicazione: (2024)
Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization
di: Guo, Yuxin, et al.
Pubblicazione: (2024)
di: Guo, Yuxin, et al.
Pubblicazione: (2024)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
di: Li, Yangze, et al.
Pubblicazione: (2024)
di: Li, Yangze, et al.
Pubblicazione: (2024)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
di: Cohen, Eyal, et al.
Pubblicazione: (2025)
di: Cohen, Eyal, et al.
Pubblicazione: (2025)
Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
di: Tsutsumi, Ayuto, et al.
Pubblicazione: (2026)
di: Tsutsumi, Ayuto, et al.
Pubblicazione: (2026)
Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids
di: Saleem, Nasir, et al.
Pubblicazione: (2025)
di: Saleem, Nasir, et al.
Pubblicazione: (2025)
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
Unified Semi-Supervised Pipeline for Automatic Speech Recognition
di: Tadevosyan, Nune, et al.
Pubblicazione: (2025)
di: Tadevosyan, Nune, et al.
Pubblicazione: (2025)
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
di: de Groot, Dimme, et al.
Pubblicazione: (2026)
di: de Groot, Dimme, et al.
Pubblicazione: (2026)
Interpreting the Role of Visemes in Audio-Visual Speech Recognition
di: Papadopoulos, Aristeidis, et al.
Pubblicazione: (2025)
di: Papadopoulos, Aristeidis, et al.
Pubblicazione: (2025)
Audio-Visual Speech Separation via Bottleneck Iterative Network
di: Zhang, Sidong, et al.
Pubblicazione: (2025)
di: Zhang, Sidong, et al.
Pubblicazione: (2025)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
di: Luo, Longjie, et al.
Pubblicazione: (2025)
di: Luo, Longjie, et al.
Pubblicazione: (2025)
BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
di: Ratnarajah, Anton, et al.
Pubblicazione: (2023)
di: Ratnarajah, Anton, et al.
Pubblicazione: (2023)
Non-Intrusive Automatic Speech Recognition Refinement: A Survey
di: Peyghan, Mohammad Reza, et al.
Pubblicazione: (2025)
di: Peyghan, Mohammad Reza, et al.
Pubblicazione: (2025)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
di: Wan, Genshun, et al.
Pubblicazione: (2026)
di: Wan, Genshun, et al.
Pubblicazione: (2026)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
di: Zhao, Ya, et al.
Pubblicazione: (2026)
di: Zhao, Ya, et al.
Pubblicazione: (2026)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
di: Lin, Yuke, et al.
Pubblicazione: (2025)
di: Lin, Yuke, et al.
Pubblicazione: (2025)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection
di: Zhang, Yuliang, et al.
Pubblicazione: (2025)
di: Zhang, Yuliang, et al.
Pubblicazione: (2025)
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
di: Kim, Jongsuk, et al.
Pubblicazione: (2025)
di: Kim, Jongsuk, et al.
Pubblicazione: (2025)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
di: Shan, Weiqiao, et al.
Pubblicazione: (2025)
di: Shan, Weiqiao, et al.
Pubblicazione: (2025)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
di: Lin, Zhaofeng, et al.
Pubblicazione: (2024)
di: Lin, Zhaofeng, et al.
Pubblicazione: (2024)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
di: Yang, Qingran, et al.
Pubblicazione: (2026)
di: Yang, Qingran, et al.
Pubblicazione: (2026)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
di: Yu, Fan, et al.
Pubblicazione: (2024)
di: Yu, Fan, et al.
Pubblicazione: (2024)
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
di: Shi, Mohan, et al.
Pubblicazione: (2025)
di: Shi, Mohan, et al.
Pubblicazione: (2025)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
di: Ok, Seaone, et al.
Pubblicazione: (2026)
di: Ok, Seaone, et al.
Pubblicazione: (2026)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
di: Su, Fei, et al.
Pubblicazione: (2026)
di: Su, Fei, et al.
Pubblicazione: (2026)
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
di: Bhogale, Kaushal Santosh, et al.
Pubblicazione: (2024)
di: Bhogale, Kaushal Santosh, et al.
Pubblicazione: (2024)
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
di: Luo, Longjie, et al.
Pubblicazione: (2025)
di: Luo, Longjie, et al.
Pubblicazione: (2025)
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
di: Huang, Wen, et al.
Pubblicazione: (2025)
di: Huang, Wen, et al.
Pubblicazione: (2025)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
di: Li, Hengyu, et al.
Pubblicazione: (2024)
di: Li, Hengyu, et al.
Pubblicazione: (2024)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
di: Liu, Hexin, et al.
Pubblicazione: (2024)
di: Liu, Hexin, et al.
Pubblicazione: (2024)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
di: Zhao, Xiaohan, et al.
Pubblicazione: (2025)
di: Zhao, Xiaohan, et al.
Pubblicazione: (2025)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
Identifying Hearing Difficulty Moments in Conversational Audio
di: Collins, Jack, et al.
Pubblicazione: (2025)
di: Collins, Jack, et al.
Pubblicazione: (2025)
Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
di: Ihori, Mana, et al.
Pubblicazione: (2025)
di: Ihori, Mana, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
di: Sun, Haoqin, et al.
Pubblicazione: (2024) -
Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization
di: Guo, Yuxin, et al.
Pubblicazione: (2024) -
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
di: Li, Yangze, et al.
Pubblicazione: (2024) -
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
di: Cohen, Eyal, et al.
Pubblicazione: (2025) -
Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)