Reinforcement Learning in MDPs with Information-Ordered Policies
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zhongjun, Agrawal, Shipra, Lobel, Ilan, Sinclair, Sean R., Yu, Christina Lee |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive Discretization in Online Reinforcement Learning
por: Sinclair, Sean R., et al.
Publicado: (2021)
por: Sinclair, Sean R., et al.
Publicado: (2021)
Black-Box Uniform Stability for Non-Euclidean Empirical Risk Minimization
por: Vary, Simon, et al.
Publicado: (2024)
por: Vary, Simon, et al.
Publicado: (2024)
Models Parametric Analysis via Adaptive Kernel Learning
por: Norkin, Vladimir, et al.
Publicado: (2025)
por: Norkin, Vladimir, et al.
Publicado: (2025)
Ambiguous Online Learning
por: Kosoy, Vanessa
Publicado: (2025)
por: Kosoy, Vanessa
Publicado: (2025)
Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
por: Chen, Qiyu, et al.
Publicado: (2025)
por: Chen, Qiyu, et al.
Publicado: (2025)
Regret Bounds for Robust Online Decision Making
por: Appel, Alexander, et al.
Publicado: (2025)
por: Appel, Alexander, et al.
Publicado: (2025)
Agnostic Learning under Targeted Poisoning: Optimal Rates and the Role of Randomness
por: Chornomaz, Bogdan, et al.
Publicado: (2025)
por: Chornomaz, Bogdan, et al.
Publicado: (2025)
On-Average Stability of Multipass Preconditioned SGD and Effective Dimension
por: Vary, Simon, et al.
Publicado: (2026)
por: Vary, Simon, et al.
Publicado: (2026)
Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
por: Xu, Zhi-Qin John, et al.
Publicado: (2019)
por: Xu, Zhi-Qin John, et al.
Publicado: (2019)
Backpropagation Through Time For Networks With Long-Term Dependencies
por: Bird, George, et al.
Publicado: (2021)
por: Bird, George, et al.
Publicado: (2021)
The Current and Future Perspectives of Zinc Oxide Nanoparticles in the Treatment of Diabetes Mellitus
por: Yousaf, Iqra
Publicado: (2024)
por: Yousaf, Iqra
Publicado: (2024)
Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality
por: Palasamudram, Amogh, et al.
Publicado: (2026)
por: Palasamudram, Amogh, et al.
Publicado: (2026)
Margin in Abstract Spaces
por: Ashlagi, Yair, et al.
Publicado: (2026)
por: Ashlagi, Yair, et al.
Publicado: (2026)
Sparse Knowledge Distillation: A Mathematical Framework for Probability-Domain Temperature Scaling and Multi-Stage Compression
por: Flouro, Aaron R., et al.
Publicado: (2026)
por: Flouro, Aaron R., et al.
Publicado: (2026)
Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory
por: Zhang, Zhi, et al.
Publicado: (2024)
por: Zhang, Zhi, et al.
Publicado: (2024)
Golden Handcuffs make safer AI agents
por: Ebtekar, Aram, et al.
Publicado: (2026)
por: Ebtekar, Aram, et al.
Publicado: (2026)
Retrieval-Augmented Memory for Online Learning
por: Du, Wenzhang
Publicado: (2025)
por: Du, Wenzhang
Publicado: (2025)
Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
por: Du, Wenzhang
Publicado: (2025)
por: Du, Wenzhang
Publicado: (2025)
Superior Scoring Rules for Probabilistic Evaluation of Single-Label Multi-Class Classification Tasks
por: Ahmadian, Rouhollah, et al.
Publicado: (2024)
por: Ahmadian, Rouhollah, et al.
Publicado: (2024)
Understanding the Nature of Generative AI as Threshold Logic in High-Dimensional Space
por: Levin, Ilya
Publicado: (2026)
por: Levin, Ilya
Publicado: (2026)
Integration of Deep Reinforcement Learning and Agent-based Simulation to Explore Strategies Counteracting Information Disorder
por: Lomasto, Luigi, et al.
Publicado: (2026)
por: Lomasto, Luigi, et al.
Publicado: (2026)
Tricks and Plug-ins for Gradient Boosting with Transformers
por: Fang, Biyi, et al.
Publicado: (2025)
por: Fang, Biyi, et al.
Publicado: (2025)
Energy-Efficient Information Representation in MNIST Classification Using Biologically Inspired Learning
por: Stricker, Patrick, et al.
Publicado: (2026)
por: Stricker, Patrick, et al.
Publicado: (2026)
MMD-Balls as Credal Sets: A PAC-Bayesian Framework for Epistemic Uncertainty in Test-Time Adaptation
por: Ariq, Ahanaf Hasan
Publicado: (2026)
por: Ariq, Ahanaf Hasan
Publicado: (2026)
Inductive Venn-Abers and related regressors
por: Petej, Ivan, et al.
Publicado: (2026)
por: Petej, Ivan, et al.
Publicado: (2026)
Aggregation in conformal e-classification
por: Vovk, Vladimir
Publicado: (2026)
por: Vovk, Vladimir
Publicado: (2026)
Does DQN Learn?
por: Gopalan, Aditya, et al.
Publicado: (2022)
por: Gopalan, Aditya, et al.
Publicado: (2022)
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
por: BC, Samiksha
Publicado: (2025)
por: BC, Samiksha
Publicado: (2025)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
por: Ma, Minghui, et al.
Publicado: (2026)
por: Ma, Minghui, et al.
Publicado: (2026)
Beyond Discreteness: Sample Complexity Analysis of Straight-Through Estimator for 1-bit Quantization
por: Jeong, Halyun, et al.
Publicado: (2025)
por: Jeong, Halyun, et al.
Publicado: (2025)
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
por: Li, Yin
Publicado: (2025)
por: Li, Yin
Publicado: (2025)
Optimistic Feasible Search for Closed-Loop Fair Threshold Decision-Making
por: Du, Wenzhang
Publicado: (2025)
por: Du, Wenzhang
Publicado: (2025)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
por: Da, Longchao, et al.
Publicado: (2025)
por: Da, Longchao, et al.
Publicado: (2025)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
por: Chen, Tiejin, et al.
Publicado: (2026)
por: Chen, Tiejin, et al.
Publicado: (2026)
From Score Matching to Diffusion: A Fine-Grained Error Analysis in the Gaussian Setting
por: Hurault, Samuel, et al.
Publicado: (2025)
por: Hurault, Samuel, et al.
Publicado: (2025)
The two clocks and the innovation window: When and how generative models learn rules
por: Wang, Binxu, et al.
Publicado: (2026)
por: Wang, Binxu, et al.
Publicado: (2026)
Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
por: Sakabe, Eduardo Y., et al.
Publicado: (2025)
por: Sakabe, Eduardo Y., et al.
Publicado: (2025)
Advances in Set Function Learning: A Survey of Techniques and Applications
por: Xie, Jiahao, et al.
Publicado: (2025)
por: Xie, Jiahao, et al.
Publicado: (2025)
Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation
por: Gharami, Kanchon, et al.
Publicado: (2025)
por: Gharami, Kanchon, et al.
Publicado: (2025)
Hallucinations Live in Variance
por: Flouro, Aaron R., et al.
Publicado: (2026)
por: Flouro, Aaron R., et al.
Publicado: (2026)
Ejemplares similares
-
Adaptive Discretization in Online Reinforcement Learning
por: Sinclair, Sean R., et al.
Publicado: (2021) -
Black-Box Uniform Stability for Non-Euclidean Empirical Risk Minimization
por: Vary, Simon, et al.
Publicado: (2024) -
Models Parametric Analysis via Adaptive Kernel Learning
por: Norkin, Vladimir, et al.
Publicado: (2025) -
Ambiguous Online Learning
por: Kosoy, Vanessa
Publicado: (2025) -
Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
por: Chen, Qiyu, et al.
Publicado: (2025)