Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Manh, Nguyen, Dung, Do, Dai, Venkatesh, Svetha, Le, Hung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Under 1 Billion: Memory-Augmented Reinforcement Learning for Large Language Models
by: Le, Hung, et al.
Published: (2025)
by: Le, Hung, et al.
Published: (2025)
SPaCe: Unlocking Sample-Efficient Large Language Models Training With Self-Pace Curriculum Learning
by: Do, Dai, et al.
Published: (2025)
by: Do, Dai, et al.
Published: (2025)
Beyond Surprise: Improving Exploration Through Surprise Novelty
by: Le, Hung, et al.
Published: (2023)
by: Le, Hung, et al.
Published: (2023)
Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement Learning
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
Enhancing Length Extrapolation in Sequential Models with Pointer-Augmented Neural Memory
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
Continual Fine-Tuning of Large Language Models via Program Memory
by: Le, Hung, et al.
Published: (2026)
by: Le, Hung, et al.
Published: (2026)
Multi-Reference Preference Optimization for Large Language Models
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
Variable-Agnostic Causal Exploration for Reinforcement Learning
by: Nguyen, Minh Hoang, et al.
Published: (2024)
by: Nguyen, Minh Hoang, et al.
Published: (2024)
Generating Realistic Tabular Data with Large Language Models
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
by: Nguyen, Manh, et al.
Published: (2025)
by: Nguyen, Manh, et al.
Published: (2025)
Probabilities Are All You Need: A Probability-Only Approach to Uncertainty Estimation in Large Language Models
by: Nguyen, Manh, et al.
Published: (2025)
by: Nguyen, Manh, et al.
Published: (2025)
Revisiting the Dataset Bias Problem from a Statistical Perspective
by: Do, Kien, et al.
Published: (2024)
by: Do, Kien, et al.
Published: (2024)
Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention
by: Nguyen, Manh, et al.
Published: (2026)
by: Nguyen, Manh, et al.
Published: (2026)
Large Language Models for Imbalanced Classification: Diversity makes the difference
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Adaptive Acquisition Selection for Bayesian Optimization with Large Language Models
by: Ngo, Giang, et al.
Published: (2026)
by: Ngo, Giang, et al.
Published: (2026)
GRAD: Graph-Retrieved Adaptive Decoding for Hallucination Mitigation
by: Nguyen, Manh, et al.
Published: (2025)
by: Nguyen, Manh, et al.
Published: (2025)
Reviving Error Correction in Modern Deep Time-Series Forecasting
by: Nguyen, Minh Hoang, et al.
Published: (2026)
by: Nguyen, Minh Hoang, et al.
Published: (2026)
Large Language Models Prompting With Episodic Memory
by: Do, Dai, et al.
Published: (2024)
by: Do, Dai, et al.
Published: (2024)
Automatic Prompt Selection for Large Language Models
by: Do, Viet-Tung, et al.
Published: (2024)
by: Do, Viet-Tung, et al.
Published: (2024)
Variational Flow Models: Flowing in Your Style
by: Do, Kien, et al.
Published: (2024)
by: Do, Kien, et al.
Published: (2024)
Accelerating Long-Term Molecular Dynamics with Physics-Informed Time-Series Forecasting
by: Le, Hung, et al.
Published: (2025)
by: Le, Hung, et al.
Published: (2025)
Federated Domain Generalization with Latent Space Inversion
by: Palakkadavath, Ragja, et al.
Published: (2025)
by: Palakkadavath, Ragja, et al.
Published: (2025)
Retrieval-augmented Decoding for Improving Truthfulness in Open-ended Generation
by: Nguyen, Manh, et al.
Published: (2025)
by: Nguyen, Manh, et al.
Published: (2025)
Score-based Integrated Gradient for Root Cause Explanations of Outliers
by: Nguyen, Phuoc, et al.
Published: (2026)
by: Nguyen, Phuoc, et al.
Published: (2026)
Leveraging Large Language Models for Information Verification -- an Engineering Approach
by: Hung, Nguyen Nang, et al.
Published: (2025)
by: Hung, Nguyen Nang, et al.
Published: (2025)
Bayesian Optimistic Optimisation with Exponentially Decaying Regret
by: Tran-The, Hung, et al.
Published: (2021)
by: Tran-The, Hung, et al.
Published: (2021)
Trading Convergence Rate with Computational Budget in High Dimensional Bayesian Optimization
by: Tran-The, Hung, et al.
Published: (2019)
by: Tran-The, Hung, et al.
Published: (2019)
Regret Bounds for Expected Improvement Algorithms in Gaussian Process Bandit Optimization
by: Tran-The, Hung, et al.
Published: (2022)
by: Tran-The, Hung, et al.
Published: (2022)
Spectral Text Fusion: A Frequency-Aware Approach to Multimodal Time-Series Forecasting
by: Nguyen, Huu Hiep, et al.
Published: (2026)
by: Nguyen, Huu Hiep, et al.
Published: (2026)
ChargeFlow: Flow-Matching Refinement of Charge-Conditioned Electron Densities
by: Nguyen, Tri Minh, et al.
Published: (2026)
by: Nguyen, Tri Minh, et al.
Published: (2026)
Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
by: Pham, Cuong, et al.
Published: (2025)
by: Pham, Cuong, et al.
Published: (2025)
Optimizing Electric Vehicle Charging Station Placement Using Reinforcement Learning and Agent-Based Simulations
by: Nguyen, Minh-Duc, et al.
Published: (2025)
by: Nguyen, Minh-Duc, et al.
Published: (2025)
RL-Guided Data Selection for Language Model Finetuning
by: Jha, Animesh, et al.
Published: (2025)
by: Jha, Animesh, et al.
Published: (2025)
Diffusion-Inspired Reconfiguration of Transformers for Uncertainty Calibration
by: Dao, Manh Cuong, et al.
Published: (2026)
by: Dao, Manh Cuong, et al.
Published: (2026)
Meta-Learning from Learning Curves for Budget-Limited Algorithm Selection
by: Nguyen, Manh Hung, et al.
Published: (2024)
by: Nguyen, Manh Hung, et al.
Published: (2024)
Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
by: Pham, Cuong, et al.
Published: (2025)
by: Pham, Cuong, et al.
Published: (2025)
Graph Contrastive Learning via Spectral Graph Alignment
by: Nguyen, Manh
Published: (2025)
by: Nguyen, Manh
Published: (2025)
Finding the Trigger: Causal Abductive Reasoning on Video Events
by: Le, Thao Minh, et al.
Published: (2025)
by: Le, Thao Minh, et al.
Published: (2025)
Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
by: Hoang, Dung Anh, et al.
Published: (2025)
by: Hoang, Dung Anh, et al.
Published: (2025)
Sub-linear Regret Bounds for Bayesian Optimisation in Unknown Search Spaces
by: Tran-The, Hung, et al.
Published: (2020)
by: Tran-The, Hung, et al.
Published: (2020)
Similar Items
-
Reasoning Under 1 Billion: Memory-Augmented Reinforcement Learning for Large Language Models
by: Le, Hung, et al.
Published: (2025) -
SPaCe: Unlocking Sample-Efficient Large Language Models Training With Self-Pace Curriculum Learning
by: Do, Dai, et al.
Published: (2025) -
Beyond Surprise: Improving Exploration Through Surprise Novelty
by: Le, Hung, et al.
Published: (2023) -
Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement Learning
by: Le, Hung, et al.
Published: (2024) -
Enhancing Length Extrapolation in Sequential Models with Pointer-Augmented Neural Memory
by: Le, Hung, et al.
Published: (2024)