Enhancing ASR Performance through OCR Word Frequency Analysis: Theoretical Foundations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jung, Kyudan, Kim, Nam-Joon, Ryu, Hyun Gon, Lee, Hyuk-Jae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MathReader : Text-to-Speech for Mathematical Documents
von: Hyeon, Sieun, et al.
Veröffentlicht: (2025)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2025)
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025)
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025)
SonicRadiation: A Hybrid Numerical Solution for Sound Radiation without Ghost Cells
von: Jin, Xutong, et al.
Veröffentlicht: (2025)
von: Jin, Xutong, et al.
Veröffentlicht: (2025)
Signal-Theoretic Characterization of Waveguide Mesh Geometries for Models of Two--Dimensional Wave Propagation in Elastic Media
von: Fontana, Federico, et al.
Veröffentlicht: (2001)
von: Fontana, Federico, et al.
Veröffentlicht: (2001)
Personal Sound Zones and Shielded Localized Communication through Active Acoustic Control
von: Egarguin, Neil Jerome A., et al.
Veröffentlicht: (2024)
von: Egarguin, Neil Jerome A., et al.
Veröffentlicht: (2024)
Keep the beat going: Automatic drum transcription with momentum
von: Foster, Alisha L., et al.
Veröffentlicht: (2025)
von: Foster, Alisha L., et al.
Veröffentlicht: (2025)
Online Correction of Dispersion Error in 2D Waveguide Meshes
von: Fontana, Federico, et al.
Veröffentlicht: (2000)
von: Fontana, Federico, et al.
Veröffentlicht: (2000)
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
von: Hyeon, Sieun, et al.
Veröffentlicht: (2024)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2024)
SLICE: Speech Enhancement via Layer-wise Injection of Conditioning Embeddings
von: Moon, Seokhoon, et al.
Veröffentlicht: (2026)
von: Moon, Seokhoon, et al.
Veröffentlicht: (2026)
MathBridge: A Large Corpus Dataset for Translating Spoken Mathematical Expressions into $LaTeX$ Formulas for Improved Readability
von: Jung, Kyudan, et al.
Veröffentlicht: (2024)
von: Jung, Kyudan, et al.
Veröffentlicht: (2024)
Numerically Informed Convolutional Operator Network with Subproblem Decomposition for Poisson Equations
von: Jung, Kyoungjin, et al.
Veröffentlicht: (2026)
von: Jung, Kyoungjin, et al.
Veröffentlicht: (2026)
SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
OCR-Enhanced Multimodal ASR Can Read While Listening
von: Chen, Junli, et al.
Veröffentlicht: (2026)
von: Chen, Junli, et al.
Veröffentlicht: (2026)
Explicit radial basis function Runge-Kutta methods
von: Gu, Jiaxi, et al.
Veröffentlicht: (2024)
von: Gu, Jiaxi, et al.
Veröffentlicht: (2024)
Schwartz duality for singularly perturbed nonlinear differential equations with Chebyshev spectral method
von: Heo, Eunwoo, et al.
Veröffentlicht: (2025)
von: Heo, Eunwoo, et al.
Veröffentlicht: (2025)
A spatial-temporal weight analysis and novel nonlinear weights of weighted essentially non-oscillatory schemes for hyperbolic conservation laws
von: Chen, Xinjuan, et al.
Veröffentlicht: (2023)
von: Chen, Xinjuan, et al.
Veröffentlicht: (2023)
A Numerical Study of WENO Approximations to Sharp Propagating Fronts for Reaction-Diffusion Systems
von: Gu, Jiaxi, et al.
Veröffentlicht: (2024)
von: Gu, Jiaxi, et al.
Veröffentlicht: (2024)
Performance Analysis of Multi-Hop Networks at Terahertz Frequencies
von: Cavallero, Sara, et al.
Veröffentlicht: (2025)
von: Cavallero, Sara, et al.
Veröffentlicht: (2025)
JS-type and Z-type weights for fourth-order central-upwind weighted essentially non-oscillatory schemes
von: Gu, Jiaxi, et al.
Veröffentlicht: (2025)
von: Gu, Jiaxi, et al.
Veröffentlicht: (2025)
Inversion of Arctic dual-channel sound speed profile based on random airgun signal
von: Weng, Jinbao, et al.
Veröffentlicht: (2025)
von: Weng, Jinbao, et al.
Veröffentlicht: (2025)
Acoustic source depth estimation method based on a single hydrophone in Arctic underwater
von: Weng, Jinbao, et al.
Veröffentlicht: (2025)
von: Weng, Jinbao, et al.
Veröffentlicht: (2025)
A Family of Even-Order Central-Upwind WENO Schemes with Averaged Downwind and Novel Global Smoothness Indicators
von: Gu, Jiaxi, et al.
Veröffentlicht: (2026)
von: Gu, Jiaxi, et al.
Veröffentlicht: (2026)
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
A two-level overlapping Schwarz method with energy-minimizing multiscale coarse basis functions
von: Wang, Junxian, et al.
Veröffentlicht: (2019)
von: Wang, Junxian, et al.
Veröffentlicht: (2019)
Iterative Contact-resolving Hybrid Methods for Multiscale Contact Mechanics
von: Chung, Eric T., et al.
Veröffentlicht: (2025)
von: Chung, Eric T., et al.
Veröffentlicht: (2025)
Finite element approximation to the non-stationary quasi-geostrophic equation
von: Kim, Dohyun, et al.
Veröffentlicht: (2024)
von: Kim, Dohyun, et al.
Veröffentlicht: (2024)
Theoretical analysis and numerical solution to a vector equation $Ax-\|x\|_1x=b$
von: Wang, Yuezhi, et al.
Veröffentlicht: (2025)
von: Wang, Yuezhi, et al.
Veröffentlicht: (2025)
LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence
von: Lee, Hyeongkeun, et al.
Veröffentlicht: (2026)
von: Lee, Hyeongkeun, et al.
Veröffentlicht: (2026)
Puzzle Game: Prediction and Classification of Wordle Solution Words
von: Xin, Haidong, et al.
Veröffentlicht: (2024)
von: Xin, Haidong, et al.
Veröffentlicht: (2024)
Numerical study on hyper parameter settings for neural network approximation to partial differential equations
von: Yang, Hee Jun, et al.
Veröffentlicht: (2025)
von: Yang, Hee Jun, et al.
Veröffentlicht: (2025)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
Structure-Preserving Operator Learning: Modeling the Collision Operator of Kinetic Equations
von: Lee, Jae Yong, et al.
Veröffentlicht: (2024)
von: Lee, Jae Yong, et al.
Veröffentlicht: (2024)
Discretization error analysis for a radially symmetric harmonic map heat flow problem
von: Nguyen, Nam Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Nam Anh, et al.
Veröffentlicht: (2025)
A comparative study of finite element methods for a class of harmonic map heat flow problems
von: Nguyen, Nam Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Nam Anh, et al.
Veröffentlicht: (2025)
Error analysis for a Finite Element Discretization of a corotational harmonic map heat flow problem
von: Nguyen, Nam Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Nam Anh, et al.
Veröffentlicht: (2025)
SLIPT for Underwater IoT: System Modeling and Performance Analysis
von: Shang, Shunyuan, et al.
Veröffentlicht: (2025)
von: Shang, Shunyuan, et al.
Veröffentlicht: (2025)
FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech Restoration
von: Yang, Da-Hee, et al.
Veröffentlicht: (2025)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2025)
A Precise Performance Analysis of the Randomized Singular Value Decomposition
von: Akhtiamov, Danil, et al.
Veröffentlicht: (2025)
von: Akhtiamov, Danil, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
Performance Analysis of Double Reconfigurable Intelligent Surfaces Assisted NOMA Networks
von: Li, Xuehua, et al.
Veröffentlicht: (2024)
von: Li, Xuehua, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MathReader : Text-to-Speech for Mathematical Documents
von: Hyeon, Sieun, et al.
Veröffentlicht: (2025) -
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025) -
SonicRadiation: A Hybrid Numerical Solution for Sound Radiation without Ghost Cells
von: Jin, Xutong, et al.
Veröffentlicht: (2025) -
Signal-Theoretic Characterization of Waveguide Mesh Geometries for Models of Two--Dimensional Wave Propagation in Elastic Media
von: Fontana, Federico, et al.
Veröffentlicht: (2001) -
Personal Sound Zones and Shielded Localized Communication through Active Acoustic Control
von: Egarguin, Neil Jerome A., et al.
Veröffentlicht: (2024)