AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Eom, SooHwan, Yoon, Eunseop, Yoon, Hee Suk, Kim, Chanwoo, Hasegawa-Johnson, Mark, Yoo, Chang D. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
by: Eom, SooHwan, et al.
Published: (2026)
by: Eom, SooHwan, et al.
Published: (2026)
LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
by: Yoon, Eunseop, et al.
Published: (2024)
by: Yoon, Eunseop, et al.
Published: (2024)
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
by: Yoon, Eunseop, et al.
Published: (2025)
by: Yoon, Eunseop, et al.
Published: (2025)
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
by: Yoon, Hee Suk, et al.
Published: (2026)
by: Yoon, Hee Suk, et al.
Published: (2026)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
by: Yoon, Hee Suk, et al.
Published: (2026)
by: Yoon, Hee Suk, et al.
Published: (2026)
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
by: Hong, Ji Woo, et al.
Published: (2026)
by: Hong, Ji Woo, et al.
Published: (2026)
TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
by: Yoon, Eunseop, et al.
Published: (2024)
by: Yoon, Eunseop, et al.
Published: (2024)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
by: Yoon, Eunseop, et al.
Published: (2025)
by: Yoon, Eunseop, et al.
Published: (2025)
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
by: Yoon, Hee Suk, et al.
Published: (2025)
by: Yoon, Hee Suk, et al.
Published: (2025)
SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation
by: Ryu, Hyun, et al.
Published: (2023)
by: Ryu, Hyun, et al.
Published: (2023)
C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion
by: Yoon, Hee Suk, et al.
Published: (2024)
by: Yoon, Hee Suk, et al.
Published: (2024)
HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue
by: Yoon, Sunjae, et al.
Published: (2023)
by: Yoon, Sunjae, et al.
Published: (2023)
Graph Connectionist Temporal Classification for Phoneme Recognition
by: Grafé, Henry, et al.
Published: (2025)
by: Grafé, Henry, et al.
Published: (2025)
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
by: Yoon, Sunjae, et al.
Published: (2022)
by: Yoon, Sunjae, et al.
Published: (2022)
ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
by: Yoon, Hee Suk, et al.
Published: (2023)
by: Yoon, Hee Suk, et al.
Published: (2023)
Self-distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective Approach
by: Zhang, Ziyin, et al.
Published: (2023)
by: Zhang, Ziyin, et al.
Published: (2023)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
by: Lee, Hyeonseung, et al.
Published: (2024)
by: Lee, Hyeonseung, et al.
Published: (2024)
Zero-Shot Dual-Path Integration Framework for Open-Vocabulary 3D Instance Segmentation
by: Ton, Tri, et al.
Published: (2024)
by: Ton, Tri, et al.
Published: (2024)
BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation
by: Yoon, Hee Suk, et al.
Published: (2024)
by: Yoon, Hee Suk, et al.
Published: (2024)
Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
by: Salvi, Giampiero
Published: (2024)
by: Salvi, Giampiero
Published: (2024)
Automatic Speech Recognition with BERT and CTC Transformers: A Review
by: Djeffal, Noussaiba, et al.
Published: (2024)
by: Djeffal, Noussaiba, et al.
Published: (2024)
Segment Boundary Detection via Class Entropy Measurements in Connectionist Phoneme Recognition
by: Salvi, Giampiero
Published: (2024)
by: Salvi, Giampiero
Published: (2024)
MER-DG: Modality-Entropy Regularization for Multimodal Domain Generalization
by: Yarici, Yavuz, et al.
Published: (2026)
by: Yarici, Yavuz, et al.
Published: (2026)
Focused Discriminative Training For Streaming CTC-Trained Automatic Speech Recognition Models
by: Haider, Adnan, et al.
Published: (2024)
by: Haider, Adnan, et al.
Published: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
by: Zheng, Xiuwen, et al.
Published: (2024)
by: Zheng, Xiuwen, et al.
Published: (2024)
Physics Informed Distillation for Diffusion Models
by: Tee, Joshua Tian Jin, et al.
Published: (2024)
by: Tee, Joshua Tian Jin, et al.
Published: (2024)
Segmentation-free Connectionist Temporal Classification loss based OCR Model for Text Captcha Classification
by: Khatavkar, Vaibhav, et al.
Published: (2024)
by: Khatavkar, Vaibhav, et al.
Published: (2024)
Unimodal Aggregation for CTC-based Speech Recognition
by: Fang, Ying, et al.
Published: (2023)
by: Fang, Ying, et al.
Published: (2023)
Enhancing CTC-Based Visual Speech Recognition
by: Laux, Hendrik, et al.
Published: (2024)
by: Laux, Hendrik, et al.
Published: (2024)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
by: Hou, Junfeng, et al.
Published: (2024)
by: Hou, Junfeng, et al.
Published: (2024)
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
by: Yoon, Ji Won, et al.
Published: (2022)
by: Yoon, Ji Won, et al.
Published: (2022)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
by: Sakuma, Asahi, et al.
Published: (2025)
by: Sakuma, Asahi, et al.
Published: (2025)
Programmable spectral shaping to improve the measurement precision of frequency comb mode-resolved spectral interferometric ranging
by: Jang, Yoon-Soo, et al.
Published: (2023)
by: Jang, Yoon-Soo, et al.
Published: (2023)
AdaTKG: Adaptive Memory for Temporal Knowledge Graph Reasoning
by: Lee, Seunghan, et al.
Published: (2026)
by: Lee, Seunghan, et al.
Published: (2026)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
by: Lee, Wonjun, et al.
Published: (2026)
by: Lee, Wonjun, et al.
Published: (2026)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
by: Vieting, Peter, et al.
Published: (2025)
by: Vieting, Peter, et al.
Published: (2025)
Highly Ordered Mesoporous Polymer‐Supported Phosphine as the Ligand for Organometallic Reaction: Suzuki‐Miyaura Cross‐Coupling of Aryl Chlorides at Room Temperature
by: Hwang Suk Kim, et al.
Published: (2024)
by: Hwang Suk Kim, et al.
Published: (2024)
TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
by: Zheng, Haolong, et al.
Published: (2025)
by: Zheng, Haolong, et al.
Published: (2025)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
by: Zheng, Xiuwen, et al.
Published: (2026)
by: Zheng, Xiuwen, et al.
Published: (2026)
The study on the multiplicity dependence of ridge behavior in $pp$ collisions at $\sqrt{s}=13$ TeV at the LHC
by: Yoon, Jeongseok, et al.
Published: (2023)
by: Yoon, Jeongseok, et al.
Published: (2023)
Similar Items
-
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
by: Eom, SooHwan, et al.
Published: (2026) -
LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
by: Yoon, Eunseop, et al.
Published: (2024) -
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
by: Yoon, Eunseop, et al.
Published: (2025) -
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
by: Yoon, Hee Suk, et al.
Published: (2026) -
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
by: Yoon, Hee Suk, et al.
Published: (2026)