What is the objective of reasoning with reinforcement learning?
Fuente:
arXiv
Guardado en:
| Autores principales: | Davis, Damek, Recht, Benjamin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When do spectral gradient updates help in deep learning?
por: Davis, Damek, et al.
Publicado: (2025)
por: Davis, Damek, et al.
Publicado: (2025)
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
por: Davis, Damek, et al.
Publicado: (2026)
por: Davis, Damek, et al.
Publicado: (2026)
Separating Geometry from Probability in the Analysis of Generalization
por: Raginsky, Maxim, et al.
Publicado: (2026)
por: Raginsky, Maxim, et al.
Publicado: (2026)
Iteratively reweighted kernel machines efficiently learn sparse functions
por: Zhu, Libin, et al.
Publicado: (2025)
por: Zhu, Libin, et al.
Publicado: (2025)
Online Covariance Estimation in Nonsmooth Stochastic Approximation
por: Jiang, Liwei, et al.
Publicado: (2025)
por: Jiang, Liwei, et al.
Publicado: (2025)
Meta-reinforcement learning with minimum attention
por: Gupta, Shashank, et al.
Publicado: (2025)
por: Gupta, Shashank, et al.
Publicado: (2025)
Stochastic Halpern iteration in normed spaces and applications to reinforcement learning
por: Bravo, Mario, et al.
Publicado: (2024)
por: Bravo, Mario, et al.
Publicado: (2024)
Flowsheet synthesis through hierarchical reinforcement learning and graph neural networks
por: Stops, Laura, et al.
Publicado: (2022)
por: Stops, Laura, et al.
Publicado: (2022)
Randomized algorithms and PAC bounds for inverse reinforcement learning in continuous spaces
por: Kamoutsi, Angeliki, et al.
Publicado: (2024)
por: Kamoutsi, Angeliki, et al.
Publicado: (2024)
Online reinforcement learning via sparse Gaussian mixture model Q-functions
por: Vu, Minh, et al.
Publicado: (2025)
por: Vu, Minh, et al.
Publicado: (2025)
A reinforcement learning agent for maintenance of deteriorating systems with increasingly imperfect repairs
por: Marugán, Alberto Pliego, et al.
Publicado: (2025)
por: Marugán, Alberto Pliego, et al.
Publicado: (2025)
Independent policy gradient-based reinforcement learning for economic and reliable energy management of multi-microgrid systems
por: Hu, Junkai, et al.
Publicado: (2025)
por: Hu, Junkai, et al.
Publicado: (2025)
Scalable spectral representations for multi-agent reinforcement learning in network MDPs
por: Ren, Zhaolin, et al.
Publicado: (2024)
por: Ren, Zhaolin, et al.
Publicado: (2024)
Beating level-set methods for 3D seismic data interpolation: a primal-dual alternating approach
por: Kumar, Rajiv, et al.
Publicado: (2016)
por: Kumar, Rajiv, et al.
Publicado: (2016)
Gradient descent with adaptive stepsize converges (nearly) linearly under fourth-order growth
por: Davis, Damek, et al.
Publicado: (2024)
por: Davis, Damek, et al.
Publicado: (2024)
Continuous-time reinforcement learning for optimal switching over multiple regimes
por: Huang, Yijie, et al.
Publicado: (2025)
por: Huang, Yijie, et al.
Publicado: (2025)
Exploiting inter-agent coupling information for efficient reinforcement learning of cooperative LQR
por: Syed, Shahbaz P Qadri, et al.
Publicado: (2025)
por: Syed, Shahbaz P Qadri, et al.
Publicado: (2025)
Taming "data-hungry" reinforcement learning? Stability in continuous state-action spaces
por: Duan, Yaqi, et al.
Publicado: (2024)
por: Duan, Yaqi, et al.
Publicado: (2024)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
por: Yang, Yan, et al.
Publicado: (2024)
por: Yang, Yan, et al.
Publicado: (2024)
Continuous-time reinforcement learning: ellipticity enables model-free value function approximation
por: Mou, Wenlong
Publicado: (2026)
por: Mou, Wenlong
Publicado: (2026)
Stabilizing reinforcement learning control: A modular framework for optimizing over all stable behavior
por: Lawrence, Nathan P., et al.
Publicado: (2023)
por: Lawrence, Nathan P., et al.
Publicado: (2023)
Learning a local trading strategy: deep reinforcement learning for grid-scale renewable energy integration
por: Ju, Caleb, et al.
Publicado: (2024)
por: Ju, Caleb, et al.
Publicado: (2024)
Randomization Inference When N Equals One
por: Liang, Tengyuan, et al.
Publicado: (2023)
por: Liang, Tengyuan, et al.
Publicado: (2023)
On characterizing optimal learning trajectories in a class of learning problems
por: Befekadu, Getachew K
Publicado: (2025)
por: Befekadu, Getachew K
Publicado: (2025)
Analysing heavy-tail properties of Stochastic Gradient Descent by means of Stochastic Recurrence Equations
por: Damek, Ewa, et al.
Publicado: (2024)
por: Damek, Ewa, et al.
Publicado: (2024)
What Data Enables Optimal Decisions? An Exact Characterization for Linear Optimization
por: Bennouna, Omar, et al.
Publicado: (2025)
por: Bennouna, Omar, et al.
Publicado: (2025)
Restarted contractive operators to learn at equilibrium
por: Davy, Leo, et al.
Publicado: (2025)
por: Davy, Leo, et al.
Publicado: (2025)
Distributed optimization: designed for federated learning
por: Guo, Wenyou, et al.
Publicado: (2025)
por: Guo, Wenyou, et al.
Publicado: (2025)
Finite sample learning of moving targets
por: Vertovec, Nikolaus, et al.
Publicado: (2024)
por: Vertovec, Nikolaus, et al.
Publicado: (2024)
Risk-averse learning with delayed feedback
por: Wang, Siyi, et al.
Publicado: (2024)
por: Wang, Siyi, et al.
Publicado: (2024)
A survey on secure decentralized optimization and learning
por: Liu, Changxin, et al.
Publicado: (2024)
por: Liu, Changxin, et al.
Publicado: (2024)
Regularized Q-learning through Robust Averaging
por: Schmitt-Förster, Peter, et al.
Publicado: (2024)
por: Schmitt-Förster, Peter, et al.
Publicado: (2024)
Apprenticeship learning with prior beliefs using inverse optimization
por: Junca, Mauricio, et al.
Publicado: (2025)
por: Junca, Mauricio, et al.
Publicado: (2025)
Convergence of gradient flow for learning convolutional neural networks
por: Diederen, Jona-Maria, et al.
Publicado: (2026)
por: Diederen, Jona-Maria, et al.
Publicado: (2026)
Data augmentation for machine learning of chemical process flowsheets
por: Balhorn, Lukas Schulze, et al.
Publicado: (2023)
por: Balhorn, Lukas Schulze, et al.
Publicado: (2023)
How to unlearn a learned Machine Learning model ?
por: Achour, Seifeddine
Publicado: (2024)
por: Achour, Seifeddine
Publicado: (2024)
Novel clustered federated learning based on local loss
por: Gu, Endong, et al.
Publicado: (2024)
por: Gu, Endong, et al.
Publicado: (2024)
A brief note on learning problem with global perspectives
por: Befekadu, Getachew K.
Publicado: (2026)
por: Befekadu, Getachew K.
Publicado: (2026)
A primal-dual perspective for distributed TD-learning
por: Lim, Han-Dong, et al.
Publicado: (2023)
por: Lim, Han-Dong, et al.
Publicado: (2023)
DistrictNet: Decision-aware learning for geographical districting
por: Ahmed, Cheikh, et al.
Publicado: (2024)
por: Ahmed, Cheikh, et al.
Publicado: (2024)
Ejemplares similares
-
When do spectral gradient updates help in deep learning?
por: Davis, Damek, et al.
Publicado: (2025) -
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
por: Davis, Damek, et al.
Publicado: (2026) -
Separating Geometry from Probability in the Analysis of Generalization
por: Raginsky, Maxim, et al.
Publicado: (2026) -
Iteratively reweighted kernel machines efficiently learn sparse functions
por: Zhu, Libin, et al.
Publicado: (2025) -
Online Covariance Estimation in Nonsmooth Stochastic Approximation
por: Jiang, Liwei, et al.
Publicado: (2025)