Saved in:
| Main Authors: | Smirnov, Ivan, Gu, Shangding |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.15040 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving
by: Zheng, Zhi, et al.
Published: (2024)
by: Zheng, Zhi, et al.
Published: (2024)
SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning
by: Guo, Zijian, et al.
Published: (2026)
by: Guo, Zijian, et al.
Published: (2026)
Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization
by: Gu, Shangding
Published: (2026)
by: Gu, Shangding
Published: (2026)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
Commit to the Bit: Reactive Reinforcement Learning Done Right
by: Eberhard, Onno, et al.
Published: (2026)
by: Eberhard, Onno, et al.
Published: (2026)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
by: Liu, Xianyang, et al.
Published: (2026)
by: Liu, Xianyang, et al.
Published: (2026)
A Review of Safe Reinforcement Learning: Methods, Theory and Applications
by: Gu, Shangding, et al.
Published: (2022)
by: Gu, Shangding, et al.
Published: (2022)
Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning
by: Gu, Shangding, et al.
Published: (2025)
by: Gu, Shangding, et al.
Published: (2025)
CombOptNet: Fit the Right NP-Hard Problem by Learning Integer Programming Constraints
by: Paulus, Anselm, et al.
Published: (2021)
by: Paulus, Anselm, et al.
Published: (2021)
LLMs as Assessors: Right for the Right Reason?
by: Saha, Sourav, et al.
Published: (2026)
by: Saha, Sourav, et al.
Published: (2026)
Right Reward Right Time for Federated Learning
by: Nguyen, Thanh Linh, et al.
Published: (2025)
by: Nguyen, Thanh Linh, et al.
Published: (2025)
What is the Right Notion of Distance between Predict-then-Optimize Tasks?
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
by: Tang, Yinghao, et al.
Published: (2025)
by: Tang, Yinghao, et al.
Published: (2025)
Right for the Right Reasons: Avoiding Reasoning Shortcuts via Prototypical Neurosymbolic AI
by: Andolfi, Luca, et al.
Published: (2025)
by: Andolfi, Luca, et al.
Published: (2025)
Finer is Better (with the Right Scaling)
by: Schaefer, Clemens, et al.
Published: (2026)
by: Schaefer, Clemens, et al.
Published: (2026)
Using Petri Nets as an Integrated Constraint Mechanism for Reinforcement Learning Tasks
by: Sachweh, Timon, et al.
Published: (2024)
by: Sachweh, Timon, et al.
Published: (2024)
StyleBench: Evaluating thinking styles in Large Language Models
by: Guo, Junyu, et al.
Published: (2025)
by: Guo, Junyu, et al.
Published: (2025)
LLMs Should Express Uncertainty Explicitly
by: Guo, Junyu, et al.
Published: (2026)
by: Guo, Junyu, et al.
Published: (2026)
Mutual Enhancement of Large Language and Reinforcement Learning Models through Bi-Directional Feedback Mechanisms: A Planning Case Study
by: Gu, Shangding
Published: (2024)
by: Gu, Shangding
Published: (2024)
MoRight: Motion Control Done Right
by: Liu, Shaowei, et al.
Published: (2026)
by: Liu, Shaowei, et al.
Published: (2026)
Learning Survival Models with Right-Censored Reporting Delays
by: Shikuri, Yuta, et al.
Published: (2025)
by: Shikuri, Yuta, et al.
Published: (2025)
Neural Network Compression for Reinforcement Learning Tasks
by: Ivanov, Dmitry A., et al.
Published: (2024)
by: Ivanov, Dmitry A., et al.
Published: (2024)
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
by: Guo, Junyu, et al.
Published: (2025)
by: Guo, Junyu, et al.
Published: (2025)
Many Ways to be Right: Rashomon Sets for Concept-Based Neural Networks
by: Feng, Shihan, et al.
Published: (2025)
by: Feng, Shihan, et al.
Published: (2025)
From Rights to Rites: Expectations Management in Smart-Home AI
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
by: Ye, Qilin, et al.
Published: (2025)
by: Ye, Qilin, et al.
Published: (2025)
Tuning the Right Foundation Models is What you Need for Partial Label Learning
by: He, Kuang, et al.
Published: (2025)
by: He, Kuang, et al.
Published: (2025)
Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond
by: Gu, Shangding, et al.
Published: (2025)
by: Gu, Shangding, et al.
Published: (2025)
Predict Confidently, Predict Right: Abstention in Dynamic Graph Learning
by: Gayen, Jayadratha, et al.
Published: (2025)
by: Gayen, Jayadratha, et al.
Published: (2025)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
by: Moulin, Antoine, et al.
Published: (2025)
by: Moulin, Antoine, et al.
Published: (2025)
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
by: Liu, Shih-Yang, et al.
Published: (2025)
by: Liu, Shih-Yang, et al.
Published: (2025)
Stochastic Gradient Descent for Gaussian Processes Done Right
by: Lin, Jihao Andreas, et al.
Published: (2023)
by: Lin, Jihao Andreas, et al.
Published: (2023)
On Fixing the Right Problems in Predictive Analytics: AUC Is Not the Problem
by: Baker, Ryan S., et al.
Published: (2024)
by: Baker, Ryan S., et al.
Published: (2024)
Learning from the Right Rollouts: Data Attribution for PPO-based LLM Post-Training
by: Shu, Dong, et al.
Published: (2026)
by: Shu, Dong, et al.
Published: (2026)
Transformers Can Do Arithmetic with the Right Embeddings
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
Elastic Weight Consolidation Done Right for Continual Learning
by: Liu, Xuan, et al.
Published: (2026)
by: Liu, Xuan, et al.
Published: (2026)
Similar Items
-
Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving
by: Zheng, Zhi, et al.
Published: (2024) -
SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning
by: Guo, Zijian, et al.
Published: (2026) -
Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization
by: Gu, Shangding
Published: (2026) -
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
by: Wang, Yuqing, et al.
Published: (2025) -
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
by: Gu, Shangding, et al.
Published: (2024)