Trust the Model Where It Trusts Itself -- Model-Based Actor-Critic with Uncertainty-Aware Rollout Adaption
Fuente:
arXiv
Guardado en:
| Autores principales: | Frauenknecht, Bernd, Eisele, Artur, Subhasish, Devdutt, Solowjow, Friedrich, Trimpe, Sebastian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On Rollouts in Model-Based Reinforcement Learning
por: Frauenknecht, Bernd, et al.
Publicado: (2025)
por: Frauenknecht, Bernd, et al.
Publicado: (2025)
Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty
por: Eisele, Artur, et al.
Publicado: (2026)
por: Eisele, Artur, et al.
Publicado: (2026)
Learning to Race in Minutes: Infoprop Dyna on the Mini Wheelbot
por: Subhasish, Devdutt, et al.
Publicado: (2026)
por: Subhasish, Devdutt, et al.
Publicado: (2026)
Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Space Models
por: Berger, Julia, et al.
Publicado: (2026)
por: Berger, Julia, et al.
Publicado: (2026)
Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics
por: Frauenknecht, Bernd, et al.
Publicado: (2026)
por: Frauenknecht, Bernd, et al.
Publicado: (2026)
Contextualized Hybrid Ensemble Q-learning: Learning Fast with Control Priors
por: Cramer, Emma, et al.
Publicado: (2024)
por: Cramer, Emma, et al.
Publicado: (2024)
On the Consistency of Kernel Methods with Dependent Observations
por: Massiani, Pierre-François, et al.
Publicado: (2024)
por: Massiani, Pierre-François, et al.
Publicado: (2024)
On Foundation Models for Dynamical Systems from Purely Synthetic Data
por: Ziegler, Martin, et al.
Publicado: (2024)
por: Ziegler, Martin, et al.
Publicado: (2024)
Event-Triggered Time-Varying Bayesian Optimization
por: Brunzema, Paul, et al.
Publicado: (2022)
por: Brunzema, Paul, et al.
Publicado: (2022)
Sailing Towards Zero-Shot State Estimation using Foundation Models Combined with a UKF
por: Holtmann, Tobin, et al.
Publicado: (2025)
por: Holtmann, Tobin, et al.
Publicado: (2025)
Safe Value Functions
por: Massiani, Pierre-François, et al.
Publicado: (2021)
por: Massiani, Pierre-François, et al.
Publicado: (2021)
Kernel conditional tests from learning-theoretic bounds
por: Massiani, Pierre-François, et al.
Publicado: (2025)
por: Massiani, Pierre-François, et al.
Publicado: (2025)
The Mini Wheelbot Dataset: High-Fidelity Data for Robot Learning
por: Hose, Henrik, et al.
Publicado: (2026)
por: Hose, Henrik, et al.
Publicado: (2026)
Data-Driven Observability Analysis for Nonlinear Stochastic Systems
por: Massiani, Pierre-François, et al.
Publicado: (2023)
por: Massiani, Pierre-François, et al.
Publicado: (2023)
When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal
por: Phalod, Aditya Ajay
Publicado: (2026)
por: Phalod, Aditya Ajay
Publicado: (2026)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
por: Wang, Tao, et al.
Publicado: (2026)
por: Wang, Tao, et al.
Publicado: (2026)
Trust Me, I Know the Way: Predictive Uncertainty in the Presence of Shortcut Learning
por: Wimmer, Lisa, et al.
Publicado: (2025)
por: Wimmer, Lisa, et al.
Publicado: (2025)
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
por: Yoo, Seungwoo, et al.
Publicado: (2026)
por: Yoo, Seungwoo, et al.
Publicado: (2026)
Learning Hybrid Dynamics Models With Simulator-Informed Latent States
por: Ensinger, Katharina, et al.
Publicado: (2023)
por: Ensinger, Katharina, et al.
Publicado: (2023)
MPX: Mixed Precision Training for JAX
por: Gräfe, Alexander, et al.
Publicado: (2025)
por: Gräfe, Alexander, et al.
Publicado: (2025)
BayeSQP: Bayesian Optimization through Sequential Quadratic Programming
por: Brunzema, Paul, et al.
Publicado: (2026)
por: Brunzema, Paul, et al.
Publicado: (2026)
Bounded Exploration with World Model Uncertainty in Soft Actor-Critic Reinforcement Learning Algorithm
por: Qiao, Ting, et al.
Publicado: (2024)
por: Qiao, Ting, et al.
Publicado: (2024)
Actor-Critic without Actor
por: Ki, Donghyeon, et al.
Publicado: (2025)
por: Ki, Donghyeon, et al.
Publicado: (2025)
To Trust Or Not To Trust Your Vision-Language Model's Prediction
por: Dong, Hao, et al.
Publicado: (2025)
por: Dong, Hao, et al.
Publicado: (2025)
Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning
por: Ishfaq, Haque, et al.
Publicado: (2025)
por: Ishfaq, Haque, et al.
Publicado: (2025)
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
por: Krishnan, Ranganath, et al.
Publicado: (2024)
por: Krishnan, Ranganath, et al.
Publicado: (2024)
Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
From Uncertainty to Trust: Kernel Dropout for AI-Powered Medical Predictions
por: Azam, Ubaid, et al.
Publicado: (2024)
por: Azam, Ubaid, et al.
Publicado: (2024)
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
por: Nauman, Michal, et al.
Publicado: (2023)
por: Nauman, Michal, et al.
Publicado: (2023)
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
por: Bhatnagar, Shalabh, et al.
Publicado: (2022)
por: Bhatnagar, Shalabh, et al.
Publicado: (2022)
D2 Actor Critic: Diffusion Actor Meets Distributional Critic
por: Zhang, Lunjun, et al.
Publicado: (2025)
por: Zhang, Lunjun, et al.
Publicado: (2025)
Actor-Critic Reinforcement Learning with Phased Actor
por: Wu, Ruofan, et al.
Publicado: (2024)
por: Wu, Ruofan, et al.
Publicado: (2024)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
por: Moalla, Skander, et al.
Publicado: (2024)
por: Moalla, Skander, et al.
Publicado: (2024)
In Trust We Survive: Emergent Trust Learning
por: Chen, Qianpu, et al.
Publicado: (2026)
por: Chen, Qianpu, et al.
Publicado: (2026)
Distributed Event-Based Learning via ADMM
por: Er, Guner Dilsad, et al.
Publicado: (2024)
por: Er, Guner Dilsad, et al.
Publicado: (2024)
Generative Actor Critic
por: Qin, Aoyang, et al.
Publicado: (2025)
por: Qin, Aoyang, et al.
Publicado: (2025)
Compressed Models are NOT Trust-equivalent to Their Large Counterparts
por: Rai, Rohit Raj, et al.
Publicado: (2025)
por: Rai, Rohit Raj, et al.
Publicado: (2025)
Evidential Trust-Aware Model Personalization in Decentralized Federated Learning for Wearable IoT
por: Rangwala, Murtaza, et al.
Publicado: (2025)
por: Rangwala, Murtaza, et al.
Publicado: (2025)
TRAM: Bridging Trust Regions and Sharpness Aware Minimization
por: Sherborne, Tom, et al.
Publicado: (2023)
por: Sherborne, Tom, et al.
Publicado: (2023)
Scaling Effects and Uncertainty Quantification in Neural Actor Critic Algorithms
por: Georgoudios, Nikos, et al.
Publicado: (2026)
por: Georgoudios, Nikos, et al.
Publicado: (2026)
Ejemplares similares
-
On Rollouts in Model-Based Reinforcement Learning
por: Frauenknecht, Bernd, et al.
Publicado: (2025) -
Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty
por: Eisele, Artur, et al.
Publicado: (2026) -
Learning to Race in Minutes: Infoprop Dyna on the Mini Wheelbot
por: Subhasish, Devdutt, et al.
Publicado: (2026) -
Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Space Models
por: Berger, Julia, et al.
Publicado: (2026) -
Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics
por: Frauenknecht, Bernd, et al.
Publicado: (2026)