Let's Fuse Step by Step: A Generative Fusion Decoding Algorithm with LLMs for Robust and Instruction-Aware ASR and OCR
Fuente:
arXiv
Saved in:
| Main Authors: | Hsu, Chan-Jan, Chen, Yi-Chang, Liao, Feng-Ting, Ho, Pei-Chen, Wang, Yu-Hsiang, Hsu, Po-Chun, Shiu, Da-shan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
by: Chen, Yi-Chang, et al.
Published: (2024)
by: Chen, Yi-Chang, et al.
Published: (2024)
Breeze-7B Technical Report
by: Hsu, Chan-Jan, et al.
Published: (2024)
by: Hsu, Chan-Jan, et al.
Published: (2024)
Taiwan Safety Benchmark and Breeze Guard: Toward Trustworthy AI for Taiwanese Mandarin
by: Hsu, Po-Chun, et al.
Published: (2026)
by: Hsu, Po-Chun, et al.
Published: (2026)
The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities
by: Research, MediaTek, et al.
Published: (2025)
by: Research, MediaTek, et al.
Published: (2025)
Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity
by: Hsu, Chan-Jan, et al.
Published: (2025)
by: Hsu, Chan-Jan, et al.
Published: (2025)
Latent Flow Transformer
by: Wu, Yen-Chen, et al.
Published: (2025)
by: Wu, Yen-Chen, et al.
Published: (2025)
Revisiting the Shape Convention of Transformer Language Models
by: Liao, Feng-Ting, et al.
Published: (2026)
by: Liao, Feng-Ting, et al.
Published: (2026)
Rethinking the shape convention of an MLP
by: Chen, Meng-Hsi, et al.
Published: (2025)
by: Chen, Meng-Hsi, et al.
Published: (2025)
Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts
by: Chen, Yi-Chang, et al.
Published: (2026)
by: Chen, Yi-Chang, et al.
Published: (2026)
Breeze Taigi: Benchmarks and Models for Taiwanese Hokkien Speech Recognition and Synthesis
by: Lan, Yu-Siang, et al.
Published: (2026)
by: Lan, Yu-Siang, et al.
Published: (2026)
StepFun-Prover Preview: Let's Think and Verify Step by Step
by: Shang, Shijie, et al.
Published: (2025)
by: Shang, Shijie, et al.
Published: (2025)
STAR: Semantic Table Representation with Header-Aware Clustering and Adaptive Weighted Fusion
by: Hsu, Shui-Hsiang, et al.
Published: (2026)
by: Hsu, Shui-Hsiang, et al.
Published: (2026)
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
by: Kuo, Tzu-Lin, et al.
Published: (2024)
by: Kuo, Tzu-Lin, et al.
Published: (2024)
Let's Verify Math Questions Step by Step
by: Shen, Chengyu, et al.
Published: (2025)
by: Shen, Chengyu, et al.
Published: (2025)
Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs
by: Lyu, Zhiyi, et al.
Published: (2025)
by: Lyu, Zhiyi, et al.
Published: (2025)
Let's Rectify Step by Step: Improving Aspect-based Sentiment Analysis with Diffusion Models
by: Liu, Shunyu, et al.
Published: (2024)
by: Liu, Shunyu, et al.
Published: (2024)
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
by: Hsu, Hsiang, et al.
Published: (2026)
by: Hsu, Hsiang, et al.
Published: (2026)
Editorial: Tenofovir Alafenamide Versus Entecavir for Hepatitis B Surface Antigen Decline–A Step Toward Functional Cure or a Marginal Margin?
by: Yao‐Chun Hsu, et al.
Published: (2026)
by: Yao‐Chun Hsu, et al.
Published: (2026)
Robust Multi-Modal Face Anti-Spoofing with Domain Adaptation: Tackling Missing Modalities, Noisy Pseudo-Labels, and Model Degradation
by: Hsu, Ming-Tsung, et al.
Published: (2025)
by: Hsu, Ming-Tsung, et al.
Published: (2025)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
by: Chou, Cheng-Kang, et al.
Published: (2025)
by: Chou, Cheng-Kang, et al.
Published: (2025)
OVOR: OnePrompt with Virtual Outlier Regularization for Rehearsal-Free Class-Incremental Learning
by: Huang, Wei-Cheng, et al.
Published: (2024)
by: Huang, Wei-Cheng, et al.
Published: (2024)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
by: Robertson, Zachary, et al.
Published: (2025)
by: Robertson, Zachary, et al.
Published: (2025)
Let's Learn Step by Step: Enhancing In-Context Learning Ability with Curriculum Learning
by: Liu, Yinpeng, et al.
Published: (2024)
by: Liu, Yinpeng, et al.
Published: (2024)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
by: Xu, Guowei, et al.
Published: (2024)
by: Xu, Guowei, et al.
Published: (2024)
A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
by: Lazzaro, Joseph, et al.
Published: (2026)
by: Lazzaro, Joseph, et al.
Published: (2026)
From Tokens to Steps: Verification-Aware Speculative Decoding for Efficient Multi-Step Reasoning
by: Purohit, Kiran, et al.
Published: (2026)
by: Purohit, Kiran, et al.
Published: (2026)
EFL Learners' Processing of Written Corrective Feedback With Think‐Alouds for L2 Pragmatic Learning
by: Hui‐Tzu Hsu, et al.
Published: (2025)
by: Hui‐Tzu Hsu, et al.
Published: (2025)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Bibliographic Instruction in a Step-by-Step Approach.
by: Soash, Richard L.
Published: (1992)
by: Soash, Richard L.
Published: (1992)
Hepatocellular Carcinoma With Small Bowel Metastasis and Intussusception: A Case Report and Literature Review
by: Meng‐Kai Hsu, et al.
Published: (2025)
by: Meng‐Kai Hsu, et al.
Published: (2025)
Stability of three-dimensional stochastic Navier-Stokes equation with Markov switching
by: Hsu, Po-Han
Published: (2022)
by: Hsu, Po-Han
Published: (2022)
Machine Unlearning for Image-to-Image Generative Models
by: Li, Guihong, et al.
Published: (2024)
by: Li, Guihong, et al.
Published: (2024)
Dropout-Based Rashomon Set Exploration for Efficient Predictive Multiplicity Estimation
by: Hsu, Hsiang, et al.
Published: (2024)
by: Hsu, Hsiang, et al.
Published: (2024)
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
by: Nichani, Arjun, et al.
Published: (2026)
by: Nichani, Arjun, et al.
Published: (2026)
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
by: Andrade, Moises, et al.
Published: (2025)
by: Andrade, Moises, et al.
Published: (2025)
Multi-Agent Deep Reinforcement Learning for Energy Efficient Multi-Hop STAR-RIS-Assisted Transmissions
by: Liao, Pei-Hsiang, et al.
Published: (2024)
by: Liao, Pei-Hsiang, et al.
Published: (2024)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
by: Zhang, Jiahao, et al.
Published: (2023)
by: Zhang, Jiahao, et al.
Published: (2023)
Two Step SOVA-Based Decoding Algorithm for Tailbiting Codes
by: Ortin, Jorge, et al.
Published: (2025)
by: Ortin, Jorge, et al.
Published: (2025)
Similar Items
-
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
by: Chen, Yi-Chang, et al.
Published: (2024) -
Breeze-7B Technical Report
by: Hsu, Chan-Jan, et al.
Published: (2024) -
Taiwan Safety Benchmark and Breeze Guard: Toward Trustworthy AI for Taiwanese Mandarin
by: Hsu, Po-Chun, et al.
Published: (2026) -
The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities
by: Research, MediaTek, et al.
Published: (2025) -
Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity
by: Hsu, Chan-Jan, et al.
Published: (2025)