POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Menet, Nicolas, Krause, Andreas, Rahimi, Abbas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thompson Sampling via Fine-Tuning of LLMs
von: Menet, Nicolas, et al.
Veröffentlicht: (2025)
von: Menet, Nicolas, et al.
Veröffentlicht: (2025)
A Theoretical Analysis of Test-Driven Code Generation
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
Locally Coherent Parallel Decoding in Diffusion Language Models
von: Hersche, Michael, et al.
Veröffentlicht: (2026)
von: Hersche, Michael, et al.
Veröffentlicht: (2026)
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
UCPO: Uncertainty-Aware Policy Optimization
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots
von: Li, Chenhao, et al.
Veröffentlicht: (2025)
von: Li, Chenhao, et al.
Veröffentlicht: (2025)
Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics
von: Li, Chenhao, et al.
Veröffentlicht: (2025)
von: Li, Chenhao, et al.
Veröffentlicht: (2025)
Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025)
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
Safe Exploration via Policy Priors
von: Wendl, Manuel, et al.
Veröffentlicht: (2026)
von: Wendl, Manuel, et al.
Veröffentlicht: (2026)
Safe Exploration Using Bayesian World Models and Log-Barrier Optimization
von: As, Yarden, et al.
Veröffentlicht: (2024)
von: As, Yarden, et al.
Veröffentlicht: (2024)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
von: Li, Xuan, et al.
Veröffentlicht: (2026)
von: Li, Xuan, et al.
Veröffentlicht: (2026)
Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
von: Sharma, Mridul, et al.
Veröffentlicht: (2025)
von: Sharma, Mridul, et al.
Veröffentlicht: (2025)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
von: Luo, Yu, et al.
Veröffentlicht: (2026)
von: Luo, Yu, et al.
Veröffentlicht: (2026)
Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
von: Yang, Daniel, et al.
Veröffentlicht: (2026)
von: Yang, Daniel, et al.
Veröffentlicht: (2026)
LITE: Efficiently Estimating Gaussian Probability of Maximality
von: Menet, Nicolas, et al.
Veröffentlicht: (2025)
von: Menet, Nicolas, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Generative Oversampling Using an Entropy-Guided Conditional Variational Autoencoder
von: Zare, Amirhossein, et al.
Veröffentlicht: (2025)
von: Zare, Amirhossein, et al.
Veröffentlicht: (2025)
Efficient Skill Discovery via Regret-Aware Optimization
von: Zhang, He, et al.
Veröffentlicht: (2025)
von: Zhang, He, et al.
Veröffentlicht: (2025)
Probabilistic Abduction for Visual Abstract Reasoning via Learning Rules in Vector-symbolic Architectures
von: Hersche, Michael, et al.
Veröffentlicht: (2024)
von: Hersche, Michael, et al.
Veröffentlicht: (2024)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Credal Ensemble Distillation for Uncertainty Quantification
von: Wang, Kaizheng, et al.
Veröffentlicht: (2025)
von: Wang, Kaizheng, et al.
Veröffentlicht: (2025)
Efficient Epistemic Uncertainty Estimation in Regression Ensemble Models Using Pairwise-Distance Estimators
von: Berry, Lucas, et al.
Veröffentlicht: (2023)
von: Berry, Lucas, et al.
Veröffentlicht: (2023)
On the Role of Noise in Factorizers for Disentangling Distributed Representations
von: Karunaratne, Geethan, et al.
Veröffentlicht: (2024)
von: Karunaratne, Geethan, et al.
Veröffentlicht: (2024)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
von: Hübotter, Jonas, et al.
Veröffentlicht: (2024)
von: Hübotter, Jonas, et al.
Veröffentlicht: (2024)
Fine-tuning Pocket-Aware Diffusion Models via Denoising Policy Optimization
von: Xue, Yuan, et al.
Veröffentlicht: (2026)
von: Xue, Yuan, et al.
Veröffentlicht: (2026)
Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency Awareness
von: Chen, Zizhao, et al.
Veröffentlicht: (2026)
von: Chen, Zizhao, et al.
Veröffentlicht: (2026)
STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
von: Chen, Yuhan, et al.
Veröffentlicht: (2025)
von: Chen, Yuhan, et al.
Veröffentlicht: (2025)
Calibration-Aware Policy Optimization for Reasoning LLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
Twin-Boot: Uncertainty-Aware Optimization via Online Two-Sample Bootstrapping
von: Brito, Carlos Stein
Veröffentlicht: (2025)
von: Brito, Carlos Stein
Veröffentlicht: (2025)
FineFT: Efficient and Risk-Aware Ensemble Reinforcement Learning for Futures Trading
von: Qin, Molei, et al.
Veröffentlicht: (2025)
von: Qin, Molei, et al.
Veröffentlicht: (2025)
Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning
von: Manchingal, Shireen Kudukkil, et al.
Veröffentlicht: (2026)
von: Manchingal, Shireen Kudukkil, et al.
Veröffentlicht: (2026)
Probabilistic Artificial Intelligence
von: Krause, Andreas, et al.
Veröffentlicht: (2025)
von: Krause, Andreas, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Fusion: An Ensemble Framework for Mitigating Hallucinations in Large Language Models
von: Dey, Prasenjit, et al.
Veröffentlicht: (2025)
von: Dey, Prasenjit, et al.
Veröffentlicht: (2025)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
Group-in-Group Policy Optimization for LLM Agent Training
von: Feng, Lang, et al.
Veröffentlicht: (2025)
von: Feng, Lang, et al.
Veröffentlicht: (2025)
Learning Efficient and Fair Policies for Uncertainty-Aware Collaborative Human-Robot Order Picking
von: Smit, Igor G., et al.
Veröffentlicht: (2024)
von: Smit, Igor G., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Thompson Sampling via Fine-Tuning of LLMs
von: Menet, Nicolas, et al.
Veröffentlicht: (2025) -
A Theoretical Analysis of Test-Driven Code Generation
von: Menet, Nicolas, et al.
Veröffentlicht: (2026) -
Locally Coherent Parallel Decoding in Diffusion Language Models
von: Hersche, Michael, et al.
Veröffentlicht: (2026) -
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025) -
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)