Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
Fuente:
arXiv
Saved in:
| Main Authors: | Roknilamouki, Amirhossein, Ghosh, Arnob, Shi, Ming, Nourzad, Fatemeh, Ekici, Eylem, Shroff, Ness B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FIRM: Federated In-client Regularized Multi-objective Alignment for Large Language Models
by: Nourzad, Fatemeh, et al.
Published: (2025)
by: Nourzad, Fatemeh, et al.
Published: (2025)
Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration
by: Roknilamouki, Amirhossein, et al.
Published: (2026)
by: Roknilamouki, Amirhossein, et al.
Published: (2026)
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
by: Kitamura, Toshinori, et al.
Published: (2025)
by: Kitamura, Toshinori, et al.
Published: (2025)
Online Learning for Optimizing AoI-Energy Tradeoff under Unknown Channel Statistics
by: Abd-Elmagid, Mohamed A., et al.
Published: (2025)
by: Abd-Elmagid, Mohamed A., et al.
Published: (2025)
Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization
by: Cao, Linfeng, et al.
Published: (2025)
by: Cao, Linfeng, et al.
Published: (2025)
Performing Load Balancing under Constraints
by: Fox, Andrea, et al.
Published: (2025)
by: Fox, Andrea, et al.
Published: (2025)
Provably Efficient Sample Complexity for Robust CMDP
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Beyond Freshness and Semantics: A Coupon-Collector Framework for Effective Status Updates
by: Ahmed, Youssef, et al.
Published: (2026)
by: Ahmed, Youssef, et al.
Published: (2026)
How to Find the Exact Pareto Front for Multi-Objective MDPs?
by: Li, Yining, et al.
Published: (2024)
by: Li, Yining, et al.
Published: (2024)
Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual
by: Li, Yining, et al.
Published: (2026)
by: Li, Yining, et al.
Published: (2026)
Constraint-Rectified Training for Efficient Chain-of-Thought
by: Wu, Qinhang, et al.
Published: (2026)
by: Wu, Qinhang, et al.
Published: (2026)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
by: Lin, Zhiqiang, et al.
Published: (2025)
by: Lin, Zhiqiang, et al.
Published: (2025)
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
by: Shi, Ming, et al.
Published: (2023)
by: Shi, Ming, et al.
Published: (2023)
Separation is Optimal for LQR under Intermittent Feedback
by: Etcibasi, Abdullah Y., et al.
Published: (2026)
by: Etcibasi, Abdullah Y., et al.
Published: (2026)
Absorb and Converge: Provable Convergence Guarantee for Absorbing Discrete Diffusion Models
by: Liang, Yuchen, et al.
Published: (2025)
by: Liang, Yuchen, et al.
Published: (2025)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Cruising the Spectrum: Joint Spectrum Mobility and Antenna Array Management for Mobile (cm/mm)Wave Connectivity
by: Bingöl, Ece, et al.
Published: (2025)
by: Bingöl, Ece, et al.
Published: (2025)
Provably Efficient Algorithms for S- and Non-Rectangular Robust MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences
by: Shi, Ming, et al.
Published: (2026)
by: Shi, Ming, et al.
Published: (2026)
Optimal Parallel Scheduling under Concave Speedup Functions
by: Li, Chengzhang, et al.
Published: (2025)
by: Li, Chengzhang, et al.
Published: (2025)
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
by: Kitamura, Toshinori, et al.
Published: (2026)
by: Kitamura, Toshinori, et al.
Published: (2026)
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026)
by: Mhammedi, Zakaria, et al.
Published: (2026)
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
by: Liang, Yuchen, et al.
Published: (2026)
by: Liang, Yuchen, et al.
Published: (2026)
Minimal Intervention Shared Control with Guaranteed Safety under Non-Convex Constraints
by: Chaubey, Shivam, et al.
Published: (2025)
by: Chaubey, Shivam, et al.
Published: (2025)
Look Once, Beam Twice: Camera-Primed Real-Time Double-Directional mmWave Beam Management for Vehicular Connectivity
by: Biswas, Avhishek, et al.
Published: (2026)
by: Biswas, Avhishek, et al.
Published: (2026)
Optimistic Safety for Online Convex Optimization with Unknown Linear Constraints
by: Hutchinson, Spencer, et al.
Published: (2024)
by: Hutchinson, Spencer, et al.
Published: (2024)
Monitoring State Transitions in Markovian Systems with Sampling Cost
by: Saurav, Kumar, et al.
Published: (2025)
by: Saurav, Kumar, et al.
Published: (2025)
A Programmable Linear Optical Quantum Reservoir with Measurement Feedback for Time Series Analysis
by: Ekici, Çağın
Published: (2026)
by: Ekici, Çağın
Published: (2026)
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
by: Ganguly, Sourav, et al.
Published: (2026)
by: Ganguly, Sourav, et al.
Published: (2026)
Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
by: Zhang, Zuyuan, et al.
Published: (2025)
by: Zhang, Zuyuan, et al.
Published: (2025)
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
by: Huang, Ruiquan, et al.
Published: (2026)
by: Huang, Ruiquan, et al.
Published: (2026)
Instantaneous Planning, Control and Safety for Navigation in Unknown Underwater Spaces
by: Karthik, Veejay, et al.
Published: (2026)
by: Karthik, Veejay, et al.
Published: (2026)
Discrete Diffusion Models: Novel Analysis and New Sampler Guarantees
by: Liang, Yuchen, et al.
Published: (2025)
by: Liang, Yuchen, et al.
Published: (2025)
Broadening Target Distributions for Accelerated Diffusion Models via a Novel Analysis Approach
by: Liang, Yuchen, et al.
Published: (2024)
by: Liang, Yuchen, et al.
Published: (2024)
Theory on Score-Mismatched Diffusion Models and Zero-Shot Conditional Samplers
by: Liang, Yuchen, et al.
Published: (2024)
by: Liang, Yuchen, et al.
Published: (2024)
Sharp Convergence Rates for Masked Diffusion Models
by: Liang, Yuchen, et al.
Published: (2026)
by: Liang, Yuchen, et al.
Published: (2026)
AI‐EDGE: An NSF AI institute for future edge networks and distributed intelligence
by: Peizhong Ju, et al.
Published: (2024)
by: Peizhong Ju, et al.
Published: (2024)
Model-Free Change Point Detection for Mixing Processes
by: Chen, Hao, et al.
Published: (2023)
by: Chen, Hao, et al.
Published: (2023)
Finding Super-spreaders in SIS Epidemics
by: Sridhar, Anirudh, et al.
Published: (2026)
by: Sridhar, Anirudh, et al.
Published: (2026)
Similar Items
-
FIRM: Federated In-client Regularized Multi-objective Alignment for Large Language Models
by: Nourzad, Fatemeh, et al.
Published: (2025) -
Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration
by: Roknilamouki, Amirhossein, et al.
Published: (2026) -
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
by: Kitamura, Toshinori, et al.
Published: (2025) -
Online Learning for Optimizing AoI-Energy Tradeoff under Unknown Channel Statistics
by: Abd-Elmagid, Mohamed A., et al.
Published: (2025) -
Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization
by: Cao, Linfeng, et al.
Published: (2025)