Gespeichert in:
| Hauptverfasser: | Xie, Yudong, Han, Zhifeng, Xiao, Qinfan, Liang, Liwei, Tao, Lu-Qi, Ren, Tian-Ling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.17829 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
von: Benster, Tyler, et al.
Veröffentlicht: (2024)
von: Benster, Tyler, et al.
Veröffentlicht: (2024)
Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning
von: Dong, Xuefu, et al.
Veröffentlicht: (2025)
von: Dong, Xuefu, et al.
Veröffentlicht: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Unimodal Aggregation for CTC-based Speech Recognition
von: Fang, Ying, et al.
Veröffentlicht: (2023)
von: Fang, Ying, et al.
Veröffentlicht: (2023)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Automatic Speech Recognition with BERT and CTC Transformers: A Review
von: Djeffal, Noussaiba, et al.
Veröffentlicht: (2024)
von: Djeffal, Noussaiba, et al.
Veröffentlicht: (2024)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
Enhancing CTC-Based Visual Speech Recognition
von: Laux, Hendrik, et al.
Veröffentlicht: (2024)
von: Laux, Hendrik, et al.
Veröffentlicht: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024)
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024)
Timbre-Aware LLM-based Direct Speech-to-Speech Translation Extendable to Multiple Language Pairs
von: Arya, Lalaram, et al.
Veröffentlicht: (2026)
von: Arya, Lalaram, et al.
Veröffentlicht: (2026)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
von: Cui, Mingyu, et al.
Veröffentlicht: (2025)
von: Cui, Mingyu, et al.
Veröffentlicht: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
Cluster-to-Predict Affect Contours from Speech
von: Kuşçu, Gökhan, et al.
Veröffentlicht: (2024)
von: Kuşçu, Gökhan, et al.
Veröffentlicht: (2024)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
Toward using Speech to Sense Student Emotion in Remote Learning Environments
von: Vyas, Sargam, et al.
Veröffentlicht: (2026)
von: Vyas, Sargam, et al.
Veröffentlicht: (2026)
Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
von: Eom, SooHwan, et al.
Veröffentlicht: (2026)
von: Eom, SooHwan, et al.
Veröffentlicht: (2026)
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
von: Bartelds, Martijn, et al.
Veröffentlicht: (2025)
von: Bartelds, Martijn, et al.
Veröffentlicht: (2025)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
von: Fujita, Yuya, et al.
Veröffentlicht: (2024)
von: Fujita, Yuya, et al.
Veröffentlicht: (2024)
IR-UWB Radar-Based Contactless Silent Speech Recognition with Attention-Enhanced Temporal Convolutional Networks
von: Lee, Sunghwa, et al.
Veröffentlicht: (2025)
von: Lee, Sunghwa, et al.
Veröffentlicht: (2025)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
von: Benster, Tyler, et al.
Veröffentlicht: (2024) -
Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning
von: Dong, Xuefu, et al.
Veröffentlicht: (2025) -
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025) -
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024) -
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)