Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhi, Chow, Chris, Zhang, Yasi, Sun, Yanchao, Zhang, Haochen, Jiang, Eric Hanchen, Liu, Han, Huang, Furong, Cui, Yuchen, Padilla, Oscar Hernan Madrid |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Current and Future Perspectives of Zinc Oxide Nanoparticles in the Treatment of Diabetes Mellitus
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery
by: Su, Jiarui, et al.
Published: (2026)
by: Su, Jiarui, et al.
Published: (2026)
The two clocks and the innovation window: When and how generative models learn rules
by: Wang, Binxu, et al.
Published: (2026)
by: Wang, Binxu, et al.
Published: (2026)
Chain-Oriented Objective Logic with Neural Network Feedback Control and Cascade Filtering for Dynamic Multi-DSL Regulation
by: Han, Jipeng
Published: (2024)
by: Han, Jipeng
Published: (2024)
Task and Motion Planning in Hierarchical 3D Scene Graphs
by: Ray, Aaron, et al.
Published: (2024)
by: Ray, Aaron, et al.
Published: (2024)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
by: Platzer, André
Published: (2024)
by: Platzer, André
Published: (2024)
Optimistic Feasible Search for Closed-Loop Fair Threshold Decision-Making
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Bridging the Gap Between Theoretical and Practical Reinforcement Learning in Undergraduate Education
by: Atif, Muhammad Ahmed, et al.
Published: (2025)
by: Atif, Muhammad Ahmed, et al.
Published: (2025)
Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
by: Koubaa, Anis, et al.
Published: (2025)
by: Koubaa, Anis, et al.
Published: (2025)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)
by: Saini, Saurabh, et al.
Published: (2026)
MSMixer: Learned Multi-Scale Temporal Mixing with Complementary Linear Shortcut for Long-Term Time Series Forecasting
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
AI Agents: Evolution, Architecture, and Real-World Applications
by: Krishnan, Naveen
Published: (2025)
by: Krishnan, Naveen
Published: (2025)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
by: Shekar, Pavan C, et al.
Published: (2025)
by: Shekar, Pavan C, et al.
Published: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
by: Klačan, Ján, et al.
Published: (2026)
by: Klačan, Ján, et al.
Published: (2026)
Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
by: Xu, Zhi-Qin John, et al.
Published: (2019)
by: Xu, Zhi-Qin John, et al.
Published: (2019)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
Superior Scoring Rules for Probabilistic Evaluation of Single-Label Multi-Class Classification Tasks
by: Ahmadian, Rouhollah, et al.
Published: (2024)
by: Ahmadian, Rouhollah, et al.
Published: (2024)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
Scalable Heterogeneous Graph Foundation Models for Data-Driven Optimal Power Flow in Smart Grids
by: Pasini, Massimiliano Lupo, et al.
Published: (2026)
by: Pasini, Massimiliano Lupo, et al.
Published: (2026)
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
Cooperative Patrol Routing: Optimizing Urban Crime Surveillance through Multi-Agent Reinforcement Learning
by: Palma-Borda, Juan, et al.
Published: (2025)
by: Palma-Borda, Juan, et al.
Published: (2025)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
AI Model for Predicting Binding Affinity of Antidiabetic Compounds Targeting PPAR
by: Aman, La Ode, et al.
Published: (2024)
by: Aman, La Ode, et al.
Published: (2024)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
Kodezi Chronos: A Debugging-First Language Model for Repository-Scale Code Understanding
by: Khan, Ishraq, et al.
Published: (2025)
by: Khan, Ishraq, et al.
Published: (2025)
Murphys Laws of AI Alignment: Why the Gap Always Wins
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
Connectivity-Aware Representations for Constrained Motion Planning via Multi-Scale Contrastive Learning
by: Jeon, Suhyun, et al.
Published: (2026)
by: Jeon, Suhyun, et al.
Published: (2026)
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
by: Gulati, Aryan, et al.
Published: (2025)
by: Gulati, Aryan, et al.
Published: (2025)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
by: Ma, Minghui, et al.
Published: (2026)
by: Ma, Minghui, et al.
Published: (2026)
Improving Efficiency of Sampling-based Motion Planning via Message-Passing Monte Carlo
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
Navigational Thinking as an Emerging Paradigm of Computer Science in the Age of Generative AI
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
by: Liu, Xiaoou, et al.
Published: (2026)
by: Liu, Xiaoou, et al.
Published: (2026)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Similar Items
-
The Current and Future Perspectives of Zinc Oxide Nanoparticles in the Treatment of Diabetes Mellitus
by: Yousaf, Iqra
Published: (2024) -
Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
by: Du, Wenzhang
Published: (2025) -
STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery
by: Su, Jiarui, et al.
Published: (2026) -
The two clocks and the innovation window: When and how generative models learn rules
by: Wang, Binxu, et al.
Published: (2026) -
Chain-Oriented Objective Logic with Neural Network Feedback Control and Cascade Filtering for Dynamic Multi-DSL Regulation
by: Han, Jipeng
Published: (2024)