PACR: Progressively Ascending Confidence Reward for LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yoon, Eunseop, Yoon, Hee Suk, Jang, Jaehyun, Eom, SooHwan, Dai, Qi, Luo, Chong, Hasegawa-Johnson, Mark A., Yoo, Chang D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
by: Yoon, Hee Suk, et al.
Published: (2026)
by: Yoon, Hee Suk, et al.
Published: (2026)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
by: Yoon, Hee Suk, et al.
Published: (2026)
by: Yoon, Hee Suk, et al.
Published: (2026)
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
by: Eom, SooHwan, et al.
Published: (2024)
by: Eom, SooHwan, et al.
Published: (2024)
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
by: Hong, Ji Woo, et al.
Published: (2026)
by: Hong, Ji Woo, et al.
Published: (2026)
TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
by: Yoon, Eunseop, et al.
Published: (2024)
by: Yoon, Eunseop, et al.
Published: (2024)
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
by: Yoon, Hee Suk, et al.
Published: (2025)
by: Yoon, Hee Suk, et al.
Published: (2025)
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
by: Eom, SooHwan, et al.
Published: (2026)
by: Eom, SooHwan, et al.
Published: (2026)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
by: Yoon, Eunseop, et al.
Published: (2025)
by: Yoon, Eunseop, et al.
Published: (2025)
LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
by: Yoon, Eunseop, et al.
Published: (2024)
by: Yoon, Eunseop, et al.
Published: (2024)
C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion
by: Yoon, Hee Suk, et al.
Published: (2024)
by: Yoon, Hee Suk, et al.
Published: (2024)
SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation
by: Ryu, Hyun, et al.
Published: (2023)
by: Ryu, Hyun, et al.
Published: (2023)
HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue
by: Yoon, Sunjae, et al.
Published: (2023)
by: Yoon, Sunjae, et al.
Published: (2023)
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
by: Yoon, Sunjae, et al.
Published: (2022)
by: Yoon, Sunjae, et al.
Published: (2022)
ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
by: Yoon, Hee Suk, et al.
Published: (2023)
by: Yoon, Hee Suk, et al.
Published: (2023)
BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation
by: Yoon, Hee Suk, et al.
Published: (2024)
by: Yoon, Hee Suk, et al.
Published: (2024)
Zero-Shot Dual-Path Integration Framework for Open-Vocabulary 3D Instance Segmentation
by: Ton, Tri, et al.
Published: (2024)
by: Ton, Tri, et al.
Published: (2024)
Programmable spectral shaping to improve the measurement precision of frequency comb mode-resolved spectral interferometric ranging
by: Jang, Yoon-Soo, et al.
Published: (2023)
by: Jang, Yoon-Soo, et al.
Published: (2023)
Approaching the quantum-limited precision in frequency-comb-based spectral interferometry for length measurements
by: Jang, Yoon-Soo, et al.
Published: (2025)
by: Jang, Yoon-Soo, et al.
Published: (2025)
Approaching the Quantum‐Limited Precision in Frequency‐Comb‐Based Spectral Interferometric Ranging
by: Yoon‐Soo Jang, et al.
Published: (2025)
by: Yoon‐Soo Jang, et al.
Published: (2025)
Approaching the Quantum‐Limited Precision in Frequency‐Comb‐Based Spectral Interferometric Ranging (Laser Photonics Rev. 19(11)/2025)
by: Yoon‐Soo Jang, et al.
Published: (2025)
by: Yoon‐Soo Jang, et al.
Published: (2025)
73‐2: Impact on the Observer Metameric Failure by Adding a White Channel to RGB‐Primary Display
by: Jang Jin Yoo, et al.
Published: (2025)
by: Jang Jin Yoo, et al.
Published: (2025)
OID-PPO: Optimal Interior Design using Proximal Policy Optimization by Transforming Design Guidelines into Reward Functions
by: Yoon, Chanyoung, et al.
Published: (2025)
by: Yoon, Chanyoung, et al.
Published: (2025)
Highly Ordered Mesoporous Polymer‐Supported Phosphine as the Ligand for Organometallic Reaction: Suzuki‐Miyaura Cross‐Coupling of Aryl Chlorides at Room Temperature
by: Hwang Suk Kim, et al.
Published: (2024)
by: Hwang Suk Kim, et al.
Published: (2024)
The study on the multiplicity dependence of ridge behavior in $pp$ collisions at $\sqrt{s}=13$ TeV at the LHC
by: Yoon, Jeongseok, et al.
Published: (2023)
by: Yoon, Jeongseok, et al.
Published: (2023)
Impact of Higher-order Tidal Corrections on the Measurement Accuracy of Neutron Star Tidal Deformability
by: Park, Gyeongbin, et al.
Published: (2025)
by: Park, Gyeongbin, et al.
Published: (2025)
Systematic bias due to eccentricity in parameter estimation for merging binary neutron stars : Spinning case
by: Lee, Eunjung, et al.
Published: (2025)
by: Lee, Eunjung, et al.
Published: (2025)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
by: Zheng, Xiuwen, et al.
Published: (2026)
by: Zheng, Xiuwen, et al.
Published: (2026)
Fully stabilized 25 GHz frequency comb for frequency calibration of optical spectrum analyzer
by: On, Yoonkwon, et al.
Published: (2025)
by: On, Yoonkwon, et al.
Published: (2025)
Confidence-guided Refinement Reasoning for Zero-shot Question Answering
by: Jang, Youwon, et al.
Published: (2025)
by: Jang, Youwon, et al.
Published: (2025)
Versatile and Fast Location-Based Private Information Retrieval with Fully Homomorphic Encryption over the Torus
by: Yoo, Joon Soo, et al.
Published: (2025)
by: Yoo, Joon Soo, et al.
Published: (2025)
Physics Informed Distillation for Diffusion Models
by: Tee, Joshua Tian Jin, et al.
Published: (2024)
by: Tee, Joshua Tian Jin, et al.
Published: (2024)
On the Improvement of the "Copyright Law" of Korea for Library Services for Persons with Disabilities
by: Yoon, Hee-Yoon, et al.
Published: (2013)
by: Yoon, Hee-Yoon, et al.
Published: (2013)
Adaptive Testing for LLM-Based Applications: A Diversity-based Approach
by: Yoon, Juyeon, et al.
Published: (2025)
by: Yoon, Juyeon, et al.
Published: (2025)
World Model Implanting for Test-time Adaptation of Embodied Agents
by: Yoo, Minjong, et al.
Published: (2025)
by: Yoo, Minjong, et al.
Published: (2025)
CLST: Cold-Start Mitigation in Knowledge Tracing by Aligning a Generative Language Model as a Students' Knowledge Tracer
by: Jung, Heeseok, et al.
Published: (2024)
by: Jung, Heeseok, et al.
Published: (2024)
Test-Time Mixture of World Models for Embodied Agents in Dynamic Environments
by: Jang, Jinwoo, et al.
Published: (2026)
by: Jang, Jinwoo, et al.
Published: (2026)
Progressive Fourier Neural Representation for Sequential Video Compilation
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
Bidirectional Biometric Authentication Using Transciphering and (T)FHE
by: Yoo, Joon Soo, et al.
Published: (2025)
by: Yoo, Joon Soo, et al.
Published: (2025)
Off–On fluorescent benzothiazole‐fused coumarin for sensitive detection of nitroreductases and hydrogen sulfide
by: Song Yi Yoo, et al.
Published: (2024)
by: Song Yi Yoo, et al.
Published: (2024)
SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors
by: Yegorova, Yekaterina, et al.
Published: (2026)
by: Yegorova, Yekaterina, et al.
Published: (2026)
Similar Items
-
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
by: Yoon, Hee Suk, et al.
Published: (2026) -
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
by: Yoon, Hee Suk, et al.
Published: (2026) -
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
by: Eom, SooHwan, et al.
Published: (2024) -
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
by: Hong, Ji Woo, et al.
Published: (2026) -
TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
by: Yoon, Eunseop, et al.
Published: (2024)