Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Cordero-Encinar, Paula, Duncan, Andrew B. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Soft Specialists: $α$-Rényi Ensembles for Uncertainty-Aware LLM Post-Training
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2026)
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2026)
Non-asymptotic Analysis of Diffusion Annealed Langevin Monte Carlo for Generative Modelling
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2025)
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2025)
Deep Optimal Sensor Placement for Black Box Stochastic Simulations
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2024)
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2024)
Sampling by averaging: A multiscale approach to score estimation
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2025)
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2025)
Diffusion Path Samplers via Sequential Monte Carlo
di: Young, James Matthew, et al.
Pubblicazione: (2026)
di: Young, James Matthew, et al.
Pubblicazione: (2026)
Diffusion annealed Langevin dynamics: a theoretical study
di: Cattiaux, Patrick, et al.
Pubblicazione: (2025)
di: Cattiaux, Patrick, et al.
Pubblicazione: (2025)
Proximal Interacting Particle Langevin Algorithms
di: Encinar, Paula Cordero, et al.
Pubblicazione: (2024)
di: Encinar, Paula Cordero, et al.
Pubblicazione: (2024)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
di: Lee, Jaehyeok, et al.
Pubblicazione: (2024)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2024)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
di: Kim, Junseok, et al.
Pubblicazione: (2026)
di: Kim, Junseok, et al.
Pubblicazione: (2026)
Reliable Statistical Guarantees for Conformal Predictors with Small Datasets
di: Sánchez-Domínguez, Miguel, et al.
Pubblicazione: (2025)
di: Sánchez-Domínguez, Miguel, et al.
Pubblicazione: (2025)
Round-trip Reinforcement Learning: Self-Consistent Training for Better Chemical LLMs
di: Kong, Lecheng, et al.
Pubblicazione: (2025)
di: Kong, Lecheng, et al.
Pubblicazione: (2025)
Certainty in Uncertainty: Reasoning over Uncertain Knowledge Graphs with Statistical Guarantees
di: Zhu, Yuqicheng, et al.
Pubblicazione: (2025)
di: Zhu, Yuqicheng, et al.
Pubblicazione: (2025)
Causal Discovery for Irregularly Time Series with Consistency Guarantees
di: Li, Weihong, et al.
Pubblicazione: (2025)
di: Li, Weihong, et al.
Pubblicazione: (2025)
$H$-Consistency Guarantees for Regression
di: Mao, Anqi, et al.
Pubblicazione: (2024)
di: Mao, Anqi, et al.
Pubblicazione: (2024)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
di: Zhou, Zenghui, et al.
Pubblicazione: (2026)
di: Zhou, Zenghui, et al.
Pubblicazione: (2026)
Self-Trained Verification for Training- and Test-Time Self-Improvement
di: Wu, Chen Henry, et al.
Pubblicazione: (2026)
di: Wu, Chen Henry, et al.
Pubblicazione: (2026)
Ranking Reasoning LLMs under Test-Time Scaling
di: Hariri, Mohsen, et al.
Pubblicazione: (2026)
di: Hariri, Mohsen, et al.
Pubblicazione: (2026)
Compression Aware Certified Training
di: Xu, Changming, et al.
Pubblicazione: (2025)
di: Xu, Changming, et al.
Pubblicazione: (2025)
GeoCert: Certified Geometric AI for Reliable Forecasting
di: Zhang, Regina, et al.
Pubblicazione: (2026)
di: Zhang, Regina, et al.
Pubblicazione: (2026)
Open-World Test-Time Training: Self-Training with Contrast Learning
di: Su, Houcheng, et al.
Pubblicazione: (2024)
di: Su, Houcheng, et al.
Pubblicazione: (2024)
Test-Time Training on Graphs with Large Language Models (LLMs)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
Training Guarantees of Neural Network Classification Two-Sample Tests by Kernel Analysis
di: Khurana, Varun, et al.
Pubblicazione: (2024)
di: Khurana, Varun, et al.
Pubblicazione: (2024)
Multi-Label Learning with Stronger Consistency Guarantees
di: Mao, Anqi, et al.
Pubblicazione: (2024)
di: Mao, Anqi, et al.
Pubblicazione: (2024)
Learning Neural Networks with Distribution Shift: Efficiently Certifiable Guarantees
di: Chandrasekaran, Gautam, et al.
Pubblicazione: (2025)
di: Chandrasekaran, Gautam, et al.
Pubblicazione: (2025)
Certifying Counterfactual Bias in LLMs
di: Chaudhary, Isha, et al.
Pubblicazione: (2024)
di: Chaudhary, Isha, et al.
Pubblicazione: (2024)
Towards Certified Malware Detection: Provable Guarantees Against Evasion Attacks
di: Giri, Nandakrishna, et al.
Pubblicazione: (2026)
di: Giri, Nandakrishna, et al.
Pubblicazione: (2026)
GF-Score: Certified Class-Conditional Robustness Evaluation with Fairness Guarantees
di: Shah, Arya, et al.
Pubblicazione: (2026)
di: Shah, Arya, et al.
Pubblicazione: (2026)
Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
di: Zhai, Zhiyuan, et al.
Pubblicazione: (2026)
di: Zhai, Zhiyuan, et al.
Pubblicazione: (2026)
Statistical Guarantees for Reasoning Probes on Looped Boolean Circuits
di: Kratsios, Anastasis, et al.
Pubblicazione: (2026)
di: Kratsios, Anastasis, et al.
Pubblicazione: (2026)
Certifiable Boolean Reasoning Is Universal
di: Li, Wenhao, et al.
Pubblicazione: (2026)
di: Li, Wenhao, et al.
Pubblicazione: (2026)
Computational and Statistical Guarantees for Tensor-on-Tensor Regression with Tensor Train Decomposition
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
Statistical Guarantees for Offline Domain Randomization
di: Fickinger, Arnaud, et al.
Pubblicazione: (2025)
di: Fickinger, Arnaud, et al.
Pubblicazione: (2025)
Efficient Evaluation of LLM Performance with Statistical Guarantees
di: Wu, Skyler, et al.
Pubblicazione: (2026)
di: Wu, Skyler, et al.
Pubblicazione: (2026)
Symmetry Guarantees Statistic Recovery in Variational Inference
di: Marks, Daniel, et al.
Pubblicazione: (2026)
di: Marks, Daniel, et al.
Pubblicazione: (2026)
TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
di: Zhang, Junru, et al.
Pubblicazione: (2025)
di: Zhang, Junru, et al.
Pubblicazione: (2025)
Decoupled Prototype Learning for Reliable Test-Time Adaptation
di: Wang, Guowei, et al.
Pubblicazione: (2024)
di: Wang, Guowei, et al.
Pubblicazione: (2024)
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
di: Shah, Avidan, et al.
Pubblicazione: (2026)
di: Shah, Avidan, et al.
Pubblicazione: (2026)
Sample Compression for Self Certified Continual Learning
di: Comeau, Jacob, et al.
Pubblicazione: (2025)
di: Comeau, Jacob, et al.
Pubblicazione: (2025)
Statistical Guarantees for High-Dimensional Stochastic Gradient Descent
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
Active Seriation: Efficient Ordering Recovery with Statistical Guarantees
di: Cheshire, James, et al.
Pubblicazione: (2026)
di: Cheshire, James, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Soft Specialists: $α$-Rényi Ensembles for Uncertainty-Aware LLM Post-Training
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2026) -
Non-asymptotic Analysis of Diffusion Annealed Langevin Monte Carlo for Generative Modelling
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2025) -
Deep Optimal Sensor Placement for Black Box Stochastic Simulations
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2024) -
Sampling by averaging: A multiscale approach to score estimation
di: Cordero-Encinar, Paula, et al.
Pubblicazione: (2025) -
Diffusion Path Samplers via Sequential Monte Carlo
di: Young, James Matthew, et al.
Pubblicazione: (2026)