Ergodicity in reinforcement learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Baumann, Dominik, Noorani, Erfaun, Mustafin, Arsenii, Sheng, Xinyi, Verbruggen, Bert, Vanhoyweghen, Arne, Ginis, Vincent, Schön, Thomas B. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Model-Agnostic Solutions for Deep Reinforcement Learning in Non-Ergodic Contexts
por: Verbruggen, Bert, et al.
Publicado: (2026)
por: Verbruggen, Bert, et al.
Publicado: (2026)
Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases
por: Mustafin, Arsenii, et al.
Publicado: (2025)
por: Mustafin, Arsenii, et al.
Publicado: (2025)
Reinforcement learning with non-ergodic reward increments: robustness via ergodicity transformations
por: Baumann, Dominik, et al.
Publicado: (2023)
por: Baumann, Dominik, et al.
Publicado: (2023)
Safe reinforcement learning in uncertain contexts
por: Baumann, Dominik, et al.
Publicado: (2024)
por: Baumann, Dominik, et al.
Publicado: (2024)
Lexical Hints of Accuracy in LLM Reasoning Chains
por: Vanhoyweghen, Arne, et al.
Publicado: (2025)
por: Vanhoyweghen, Arne, et al.
Publicado: (2025)
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning
por: Sheng, Xinyi, et al.
Publicado: (2025)
por: Sheng, Xinyi, et al.
Publicado: (2025)
Distributed Risk-Sensitive Safety Filters for Uncertain Discrete-Time Systems
por: Lederer, Armin, et al.
Publicado: (2025)
por: Lederer, Armin, et al.
Publicado: (2025)
Risk-Sensitive Reinforcement Learning with Exponential Criteria
por: Noorani, Erfaun, et al.
Publicado: (2022)
por: Noorani, Erfaun, et al.
Publicado: (2022)
Metro 3 in Brussels under uncertainty: scenario-based public transport accessibility analysis
por: Verbeken, Brecht, et al.
Publicado: (2025)
por: Verbeken, Brecht, et al.
Publicado: (2025)
On Value Iteration Convergence in Connected MDPs
por: Mustafin, Arsenii, et al.
Publicado: (2024)
por: Mustafin, Arsenii, et al.
Publicado: (2024)
Closing the gap between SVRG and TD-SVRG with Gradient Splitting
por: Mustafin, Arsenii, et al.
Publicado: (2022)
por: Mustafin, Arsenii, et al.
Publicado: (2022)
Towards Efficient Risk-Sensitive Policy Gradient: An Iteration Complexity Analysis
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Counterfactual Explanations for Model Ensembles Using Entropic Risk Measures
por: Noorani, Erfaun, et al.
Publicado: (2025)
por: Noorani, Erfaun, et al.
Publicado: (2025)
Safe learning-based control via function-based uncertainty quantification
por: Tokmak, Abdullah, et al.
Publicado: (2026)
por: Tokmak, Abdullah, et al.
Publicado: (2026)
From Abstraction to Reality: DARPA's Vision for Robust Sim-to-Real Autonomy
por: Noorani, Erfaun, et al.
Publicado: (2025)
por: Noorani, Erfaun, et al.
Publicado: (2025)
Analysis of Value Iteration Through Absolute Probability Sequences
por: Mustafin, Arsenii, et al.
Publicado: (2025)
por: Mustafin, Arsenii, et al.
Publicado: (2025)
Geometric Re-Analysis of Classical MDP Solving Algorithms
por: Mustafin, Arsenii, et al.
Publicado: (2025)
por: Mustafin, Arsenii, et al.
Publicado: (2025)
Towards safe control parameter tuning in distributed multi-agent systems
por: Tokmak, Abdullah, et al.
Publicado: (2025)
por: Tokmak, Abdullah, et al.
Publicado: (2025)
Safe Bayesian optimization across noise models via scenario programming
por: Tokmak, Abdullah, et al.
Publicado: (2025)
por: Tokmak, Abdullah, et al.
Publicado: (2025)
PACSBO: Probably approximately correct safe Bayesian optimization
por: Tokmak, Abdullah, et al.
Publicado: (2024)
por: Tokmak, Abdullah, et al.
Publicado: (2024)
MDP Geometry, Normalization and Reward Balancing Solvers
por: Mustafin, Arsenii, et al.
Publicado: (2024)
por: Mustafin, Arsenii, et al.
Publicado: (2024)
Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees
por: Hamman, Faisal, et al.
Publicado: (2023)
por: Hamman, Faisal, et al.
Publicado: (2023)
The Effectiveness of Curvature-Based Rewiring and the Role of Hyperparameters in GNNs Revisited
por: Tori, Floriano, et al.
Publicado: (2024)
por: Tori, Floriano, et al.
Publicado: (2024)
Safe exploration in reproducing kernel Hilbert spaces
por: Tokmak, Abdullah, et al.
Publicado: (2025)
por: Tokmak, Abdullah, et al.
Publicado: (2025)
The Relationship Between Reasoning and Performance in Large Language Models -- o3 (mini) Thinks Harder, Not Longer
por: Ballon, Marthe, et al.
Publicado: (2025)
por: Ballon, Marthe, et al.
Publicado: (2025)
Structurally Human, Semantically Biased: Detecting LLM-Generated References with Embeddings and GNNs
por: Mobini, Melika, et al.
Publicado: (2026)
por: Mobini, Melika, et al.
Publicado: (2026)
Probing the Trajectories of Reasoning Traces in Large Language Models
por: Ballon, Marthe, et al.
Publicado: (2026)
por: Ballon, Marthe, et al.
Publicado: (2026)
Estimating problem difficulty without ground truth using Large Language Model comparisons
por: Ballon, Marthe, et al.
Publicado: (2025)
por: Ballon, Marthe, et al.
Publicado: (2025)
A Rigorous, Tractable Measure of Model Complexity
por: Allerbo, Oskar, et al.
Publicado: (2026)
por: Allerbo, Oskar, et al.
Publicado: (2026)
Is Supervised Learning Really That Different from Unsupervised?
por: Allerbo, Oskar, et al.
Publicado: (2025)
por: Allerbo, Oskar, et al.
Publicado: (2025)
Flexible Counterfactual Explanations with Generative Models
por: Hellemans, Stig, et al.
Publicado: (2025)
por: Hellemans, Stig, et al.
Publicado: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
por: Ballon, Marthe, et al.
Publicado: (2026)
por: Ballon, Marthe, et al.
Publicado: (2026)
Distill2Explain: Differentiable decision trees for explainable reinforcement learning in energy application controllers
por: Gokhale, Gargya, et al.
Publicado: (2024)
por: Gokhale, Gargya, et al.
Publicado: (2024)
Operating critical machine learning models in resource constrained regimes
por: Selvan, Raghavendra, et al.
Publicado: (2023)
por: Selvan, Raghavendra, et al.
Publicado: (2023)
The hidden risks of temporal resampling in clinical reinforcement learning
por: Frost, Thomas, et al.
Publicado: (2026)
por: Frost, Thomas, et al.
Publicado: (2026)
Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)
por: Verbeken, Brecht, et al.
Publicado: (2026)
por: Verbeken, Brecht, et al.
Publicado: (2026)
Probing Graph Neural Network Activation Patterns Through Graph Topology
por: Tori, Floriano, et al.
Publicado: (2026)
por: Tori, Floriano, et al.
Publicado: (2026)
Leveraging weights signals -- Predicting and improving generalizability in reinforcement learning
por: Moulin, Olivier, et al.
Publicado: (2025)
por: Moulin, Olivier, et al.
Publicado: (2025)
Transfer Learning in Latent Contextual Bandits with Covariate Shift Through Causal Transportability
por: Deng, Mingwei, et al.
Publicado: (2025)
por: Deng, Mingwei, et al.
Publicado: (2025)
Multi-Round Human-AI Collaboration with User-Specified Requirements
por: Noorani, Sima, et al.
Publicado: (2026)
por: Noorani, Sima, et al.
Publicado: (2026)
Ejemplares similares
-
Model-Agnostic Solutions for Deep Reinforcement Learning in Non-Ergodic Contexts
por: Verbruggen, Bert, et al.
Publicado: (2026) -
Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases
por: Mustafin, Arsenii, et al.
Publicado: (2025) -
Reinforcement learning with non-ergodic reward increments: robustness via ergodicity transformations
por: Baumann, Dominik, et al.
Publicado: (2023) -
Safe reinforcement learning in uncertain contexts
por: Baumann, Dominik, et al.
Publicado: (2024) -
Lexical Hints of Accuracy in LLM Reasoning Chains
por: Vanhoyweghen, Arne, et al.
Publicado: (2025)