Salvato in:
| Autori principali: | Xie, Yudong, Han, Zhifeng, Xiao, Qinfan, Liang, Liwei, Tao, Lu-Qi, Ren, Tian-Ling |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.17829 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
di: Benster, Tyler, et al.
Pubblicazione: (2024)
di: Benster, Tyler, et al.
Pubblicazione: (2024)
Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning
di: Dong, Xuefu, et al.
Pubblicazione: (2025)
di: Dong, Xuefu, et al.
Pubblicazione: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
di: Hou, Junfeng, et al.
Pubblicazione: (2024)
di: Hou, Junfeng, et al.
Pubblicazione: (2024)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2025)
di: Chen, Youjun, et al.
Pubblicazione: (2025)
Unimodal Aggregation for CTC-based Speech Recognition
di: Fang, Ying, et al.
Pubblicazione: (2023)
di: Fang, Ying, et al.
Pubblicazione: (2023)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
Automatic Speech Recognition with BERT and CTC Transformers: A Review
di: Djeffal, Noussaiba, et al.
Pubblicazione: (2024)
di: Djeffal, Noussaiba, et al.
Pubblicazione: (2024)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
di: Sakuma, Asahi, et al.
Pubblicazione: (2025)
di: Sakuma, Asahi, et al.
Pubblicazione: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
di: Nowrin, Sadia, et al.
Pubblicazione: (2024)
di: Nowrin, Sadia, et al.
Pubblicazione: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
di: Dutta, Satwik, et al.
Pubblicazione: (2025)
di: Dutta, Satwik, et al.
Pubblicazione: (2025)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
di: Chang, Yi, et al.
Pubblicazione: (2024)
di: Chang, Yi, et al.
Pubblicazione: (2024)
Enhancing CTC-Based Visual Speech Recognition
di: Laux, Hendrik, et al.
Pubblicazione: (2024)
di: Laux, Hendrik, et al.
Pubblicazione: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
di: Eom, SooHwan, et al.
Pubblicazione: (2024)
di: Eom, SooHwan, et al.
Pubblicazione: (2024)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
di: Mishra, Ruchik, et al.
Pubblicazione: (2024)
di: Mishra, Ruchik, et al.
Pubblicazione: (2024)
Timbre-Aware LLM-based Direct Speech-to-Speech Translation Extendable to Multiple Language Pairs
di: Arya, Lalaram, et al.
Pubblicazione: (2026)
di: Arya, Lalaram, et al.
Pubblicazione: (2026)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
di: Xiao, Yi, et al.
Pubblicazione: (2022)
di: Xiao, Yi, et al.
Pubblicazione: (2022)
Human Feedback Driven Dynamic Speech Emotion Recognition
di: Fedorov, Ilya, et al.
Pubblicazione: (2025)
di: Fedorov, Ilya, et al.
Pubblicazione: (2025)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
di: Zhao, Zhixian, et al.
Pubblicazione: (2024)
di: Zhao, Zhixian, et al.
Pubblicazione: (2024)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
di: Cui, Mingyu, et al.
Pubblicazione: (2025)
di: Cui, Mingyu, et al.
Pubblicazione: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
di: Burchi, Maxime, et al.
Pubblicazione: (2024)
di: Burchi, Maxime, et al.
Pubblicazione: (2024)
Cluster-to-Predict Affect Contours from Speech
di: Kuşçu, Gökhan, et al.
Pubblicazione: (2024)
di: Kuşçu, Gökhan, et al.
Pubblicazione: (2024)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
di: Yu, Luca Jiang-Tao, et al.
Pubblicazione: (2024)
di: Yu, Luca Jiang-Tao, et al.
Pubblicazione: (2024)
Toward using Speech to Sense Student Emotion in Remote Learning Environments
di: Vyas, Sargam, et al.
Pubblicazione: (2026)
di: Vyas, Sargam, et al.
Pubblicazione: (2026)
Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
di: Hou, Yuanbo, et al.
Pubblicazione: (2024)
di: Hou, Yuanbo, et al.
Pubblicazione: (2024)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
di: Eom, SooHwan, et al.
Pubblicazione: (2026)
di: Eom, SooHwan, et al.
Pubblicazione: (2026)
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
di: Bartelds, Martijn, et al.
Pubblicazione: (2025)
di: Bartelds, Martijn, et al.
Pubblicazione: (2025)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
di: Fujita, Yuya, et al.
Pubblicazione: (2024)
di: Fujita, Yuya, et al.
Pubblicazione: (2024)
IR-UWB Radar-Based Contactless Silent Speech Recognition with Attention-Enhanced Temporal Convolutional Networks
di: Lee, Sunghwa, et al.
Pubblicazione: (2025)
di: Lee, Sunghwa, et al.
Pubblicazione: (2025)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
di: Li, Yue, et al.
Pubblicazione: (2024)
di: Li, Yue, et al.
Pubblicazione: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
di: Park, Seohyun, et al.
Pubblicazione: (2025)
di: Park, Seohyun, et al.
Pubblicazione: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
di: Wang, Hongbin, et al.
Pubblicazione: (2025)
di: Wang, Hongbin, et al.
Pubblicazione: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
di: Mertes, Silvan, et al.
Pubblicazione: (2024)
di: Mertes, Silvan, et al.
Pubblicazione: (2024)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
di: Guo, Yiwei, et al.
Pubblicazione: (2023)
di: Guo, Yiwei, et al.
Pubblicazione: (2023)
Documenti analoghi
-
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
di: Benster, Tyler, et al.
Pubblicazione: (2024) -
Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning
di: Dong, Xuefu, et al.
Pubblicazione: (2025) -
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025) -
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
di: Hou, Junfeng, et al.
Pubblicazione: (2024) -
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2025)