Maximizing Confidence Alone Improves Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Prabhudesai, Mihir, Chen, Lili, Ippoliti, Alex, Fragkiadaki, Katerina, Liu, Hao, Pathak, Deepak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Questioning Language Models
by: Chen, Lili, et al.
Published: (2025)
by: Chen, Lili, et al.
Published: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
by: Prabhudesai, Mihir, et al.
Published: (2023)
by: Prabhudesai, Mihir, et al.
Published: (2023)
Diffusion Beats Autoregressive in Data-Constrained Settings
by: Prabhudesai, Mihir, et al.
Published: (2025)
by: Prabhudesai, Mihir, et al.
Published: (2025)
Unified Multimodal Discrete Diffusion
by: Swerdlow, Alexander, et al.
Published: (2025)
by: Swerdlow, Alexander, et al.
Published: (2025)
Video Diffusion Alignment via Reward Gradients
by: Prabhudesai, Mihir, et al.
Published: (2024)
by: Prabhudesai, Mihir, et al.
Published: (2024)
Iterative Refinement Improves Compositional Image Generation
by: Jaiswal, Shantanu, et al.
Published: (2026)
by: Jaiswal, Shantanu, et al.
Published: (2026)
Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
by: Prabhudesai, Mihir, et al.
Published: (2026)
by: Prabhudesai, Mihir, et al.
Published: (2026)
Can LLMs Lie? Investigation beyond Hallucination
by: Huan, Haoran, et al.
Published: (2025)
by: Huan, Haoran, et al.
Published: (2025)
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
by: Ke, Tsung-Wei, et al.
Published: (2024)
by: Ke, Tsung-Wei, et al.
Published: (2024)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Intrinsic Explainability of Multimodal Learning for Crop Yield Prediction
by: Najjar, Hiba, et al.
Published: (2025)
by: Najjar, Hiba, et al.
Published: (2025)
Marginals Before Conditionals
by: Sahasrabudhe, Mihir
Published: (2026)
by: Sahasrabudhe, Mihir
Published: (2026)
Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
by: Vishwakarma, Harit, et al.
Published: (2024)
by: Vishwakarma, Harit, et al.
Published: (2024)
OT Score: An OT based Confidence Score for Prototype-Assisted Source Free Unsupervised Domain Adaptation
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
Meta-Evolve: Continuous Robot Evolution for One-to-many Policy Transfer
by: Liu, Xingyu, et al.
Published: (2024)
by: Liu, Xingyu, et al.
Published: (2024)
Evolutionary Policy Optimization
by: Wang, Jianren, et al.
Published: (2025)
by: Wang, Jianren, et al.
Published: (2025)
HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models
by: Sarch, Gabriel, et al.
Published: (2024)
by: Sarch, Gabriel, et al.
Published: (2024)
Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following
by: Yang, Brian, et al.
Published: (2024)
by: Yang, Brian, et al.
Published: (2024)
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
by: Liu, Yuliang, et al.
Published: (2025)
by: Liu, Yuliang, et al.
Published: (2025)
Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic
by: Mao, Zhenjiang, et al.
Published: (2025)
by: Mao, Zhenjiang, et al.
Published: (2025)
VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
by: Sarch, Gabriel, et al.
Published: (2024)
by: Sarch, Gabriel, et al.
Published: (2024)
Standard Neural Computation Alone Is Insufficient for Logical Intelligence
by: Kim, Youngsung
Published: (2025)
by: Kim, Youngsung
Published: (2025)
HEARTS: Benchmarking LLM Reasoning on Health Time Series
by: Li, Sirui, et al.
Published: (2026)
by: Li, Sirui, et al.
Published: (2026)
Reasoning-Aware Training for Time Series Forecasting
by: Ahamed, Md Atik, et al.
Published: (2026)
by: Ahamed, Md Atik, et al.
Published: (2026)
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
by: Sharma, Aman, et al.
Published: (2025)
by: Sharma, Aman, et al.
Published: (2025)
Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
by: He, Qianxi, et al.
Published: (2025)
by: He, Qianxi, et al.
Published: (2025)
Certified Robustness via Dynamic Margin Maximization and Improved Lipschitz Regularization
by: Fazlyab, Mahyar, et al.
Published: (2023)
by: Fazlyab, Mahyar, et al.
Published: (2023)
GradientSpace: Unsupervised Data Clustering for Improved Instruction Tuning
by: Sridharan, Shrihari, et al.
Published: (2025)
by: Sridharan, Shrihari, et al.
Published: (2025)
Sanity Checks for Explanation Uncertainty
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
Uncertainty Quantification for Gradient-based Explanations in Neural Networks
by: Mulye, Mihir, et al.
Published: (2024)
by: Mulye, Mihir, et al.
Published: (2024)
LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
by: Akgül, Ömer Faruk, et al.
Published: (2025)
by: Akgül, Ömer Faruk, et al.
Published: (2025)
Flowing with Confidence
by: de Kruiff, Friso, et al.
Published: (2026)
by: de Kruiff, Friso, et al.
Published: (2026)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
by: Gkanatsios, Nikolaos, et al.
Published: (2023)
by: Gkanatsios, Nikolaos, et al.
Published: (2023)
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
by: Pandita, Deepak, et al.
Published: (2026)
by: Pandita, Deepak, et al.
Published: (2026)
Flora: Low-Rank Adapters Are Secretly Gradient Compressors
by: Hao, Yongchang, et al.
Published: (2024)
by: Hao, Yongchang, et al.
Published: (2024)
NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
by: Hao, Yongchang, et al.
Published: (2024)
by: Hao, Yongchang, et al.
Published: (2024)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
by: Cai, Xin-Qiang, et al.
Published: (2026)
by: Cai, Xin-Qiang, et al.
Published: (2026)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
Early Stopping for Large Reasoning Models via Confidence Dynamics
by: Hosseini, Parsa, et al.
Published: (2026)
by: Hosseini, Parsa, et al.
Published: (2026)
Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems
by: Xia, Song, et al.
Published: (2025)
by: Xia, Song, et al.
Published: (2025)
Similar Items
-
Self-Questioning Language Models
by: Chen, Lili, et al.
Published: (2025) -
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
by: Prabhudesai, Mihir, et al.
Published: (2023) -
Diffusion Beats Autoregressive in Data-Constrained Settings
by: Prabhudesai, Mihir, et al.
Published: (2025) -
Unified Multimodal Discrete Diffusion
by: Swerdlow, Alexander, et al.
Published: (2025) -
Video Diffusion Alignment via Reward Gradients
by: Prabhudesai, Mihir, et al.
Published: (2024)