Why long model-based rollouts are no reason for bad Q-value estimates
Fuente:
arXiv
Saved in:
| Main Authors: | Wissmann, Philipp, Hein, Daniel, Udluft, Steffen, Tresp, Volker |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is Q-learning an Ill-posed Problem?
by: Wissmann, Philipp, et al.
Published: (2025)
by: Wissmann, Philipp, et al.
Published: (2025)
Model-based Offline Quantum Reinforcement Learning
by: Eisenmann, Simon, et al.
Published: (2024)
by: Eisenmann, Simon, et al.
Published: (2024)
Learning Control Policies for Variable Objectives from Offline Data
by: Weber, Marc, et al.
Published: (2023)
by: Weber, Marc, et al.
Published: (2023)
Why Uncertainty Calibration Matters for Reliable Perturbation-based Explanations
by: Decker, Thomas, et al.
Published: (2025)
by: Decker, Thomas, et al.
Published: (2025)
Neural-ANOVA: Analytical Model Decomposition using Automatic Integration
by: Limmer, Steffen, et al.
Published: (2024)
by: Limmer, Steffen, et al.
Published: (2024)
Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
by: Decker, Thomas, et al.
Published: (2025)
by: Decker, Thomas, et al.
Published: (2025)
To bootstrap or to rollout? An optimal and adaptive interpolation
by: Mou, Wenlong, et al.
Published: (2024)
by: Mou, Wenlong, et al.
Published: (2024)
First Experience with Real-Time Control Using Simulated VQC-Based Quantum Policies
by: Sun, Yize, et al.
Published: (2025)
by: Sun, Yize, et al.
Published: (2025)
TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning
by: Ormancı, Batıkan Bora, et al.
Published: (2024)
by: Ormancı, Batıkan Bora, et al.
Published: (2024)
How the (Tensor-) Brain uses Embeddings and Embodiment to Encode Senses and Symbols
by: Tresp, Volker, et al.
Published: (2024)
by: Tresp, Volker, et al.
Published: (2024)
FedPop: Federated Population-based Hyperparameter Tuning
by: Chen, Haokun, et al.
Published: (2023)
by: Chen, Haokun, et al.
Published: (2023)
Quantum Architecture Search with Unsupervised Representation Learning
by: Sun, Yize, et al.
Published: (2024)
by: Sun, Yize, et al.
Published: (2024)
A Comparative Study on How Data Normalization Affects Zero-Shot Generalization in Time Series Foundation Models
by: Ahmed, Ihab, et al.
Published: (2025)
by: Ahmed, Ihab, et al.
Published: (2025)
Incremental Uncertainty-aware Performance Monitoring with Active Labeling Intervention
by: Koebler, Alexander, et al.
Published: (2025)
by: Koebler, Alexander, et al.
Published: (2025)
Temporal Fact Reasoning over Hyper-Relational Knowledge Graphs
by: Ding, Zifeng, et al.
Published: (2023)
by: Ding, Zifeng, et al.
Published: (2023)
Explanatory Model Monitoring to Understand the Effects of Feature Shifts on Performance
by: Decker, Thomas, et al.
Published: (2024)
by: Decker, Thomas, et al.
Published: (2024)
Multi-event Video-Text Retrieval
by: Zhang, Gengyuan, et al.
Published: (2023)
by: Zhang, Gengyuan, et al.
Published: (2023)
Bayes or Heisenberg: Who(se) Rules?
by: Tresp, Volker, et al.
Published: (2025)
by: Tresp, Volker, et al.
Published: (2025)
Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model
by: Zhang, Jing, et al.
Published: (2024)
by: Zhang, Jing, et al.
Published: (2024)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
Agentic Neural Networks: Self-Evolving Multi-Agent Systems via Textual Backpropagation
by: Ma, Xiaowen, et al.
Published: (2025)
by: Ma, Xiaowen, et al.
Published: (2025)
GenTKG: Generative Forecasting on Temporal Knowledge Graph with Large Language Models
by: Liao, Ruotong, et al.
Published: (2023)
by: Liao, Ruotong, et al.
Published: (2023)
FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
LookupViT: Compressing visual information to a limited number of tokens
by: Koner, Rajat, et al.
Published: (2024)
by: Koner, Rajat, et al.
Published: (2024)
Variational Quantum Circuits in Offline Contextual Bandit Problems
by: Schulte, Lukas, et al.
Published: (2025)
by: Schulte, Lukas, et al.
Published: (2025)
Entropy After </Think> for reasoning model early exiting
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
by: Liu, Yilun, et al.
Published: (2025)
by: Liu, Yilun, et al.
Published: (2025)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
by: Liu, Yilun, et al.
Published: (2024)
by: Liu, Yilun, et al.
Published: (2024)
Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents
by: Yan, Sikuan, et al.
Published: (2026)
by: Yan, Sikuan, et al.
Published: (2026)
Provably Better Explanations with Optimized Aggregation of Feature Attributions
by: Decker, Thomas, et al.
Published: (2024)
by: Decker, Thomas, et al.
Published: (2024)
Model predictive control-based value estimation for efficient reinforcement learning
by: Wu, Qizhen, et al.
Published: (2023)
by: Wu, Qizhen, et al.
Published: (2023)
Q-value Regularized Transformer for Offline Reinforcement Learning
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Make LLMs better zero-shot reasoners: Structure-orientated autonomous reasoning
by: He, Pengfei, et al.
Published: (2024)
by: He, Pengfei, et al.
Published: (2024)
Neural topology optimization: the good, the bad, and the ugly
by: Sanu, Suryanarayanan Manoj, et al.
Published: (2024)
by: Sanu, Suryanarayanan Manoj, et al.
Published: (2024)
DyGMamba: Efficiently Modeling Long-Term Temporal Dependency on Continuous-Time Dynamic Graphs with State Space Models
by: Ding, Zifeng, et al.
Published: (2024)
by: Ding, Zifeng, et al.
Published: (2024)
zrLLM: Zero-Shot Relational Learning on Temporal Knowledge Graphs with Large Language Models
by: Ding, Zifeng, et al.
Published: (2023)
by: Ding, Zifeng, et al.
Published: (2023)
Universal approximation property of Banach space-valued random feature models including random neural networks
by: Neufeld, Ariel, et al.
Published: (2023)
by: Neufeld, Ariel, et al.
Published: (2023)
Relational reasoning and inductive bias in transformers and large language models
by: Geerts, Jesse, et al.
Published: (2025)
by: Geerts, Jesse, et al.
Published: (2025)
The good, the bad and the ugly sides of data augmentation: An implicit spectral regularization perspective
by: Lin, Chi-Heng, et al.
Published: (2022)
by: Lin, Chi-Heng, et al.
Published: (2022)
Performance of AI agents based on reasoning language models on ALD process optimization tasks
by: Yanguas-Gil, Angel
Published: (2026)
by: Yanguas-Gil, Angel
Published: (2026)
Similar Items
-
Is Q-learning an Ill-posed Problem?
by: Wissmann, Philipp, et al.
Published: (2025) -
Model-based Offline Quantum Reinforcement Learning
by: Eisenmann, Simon, et al.
Published: (2024) -
Learning Control Policies for Variable Objectives from Offline Data
by: Weber, Marc, et al.
Published: (2023) -
Why Uncertainty Calibration Matters for Reliable Perturbation-based Explanations
by: Decker, Thomas, et al.
Published: (2025) -
Neural-ANOVA: Analytical Model Decomposition using Automatic Integration
by: Limmer, Steffen, et al.
Published: (2024)