Gespeichert in:
| Hauptverfasser: | Smirnov, Ivan, Gu, Shangding |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2505.15040 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving
von: Zheng, Zhi, et al.
Veröffentlicht: (2024)
von: Zheng, Zhi, et al.
Veröffentlicht: (2024)
SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning
von: Guo, Zijian, et al.
Veröffentlicht: (2026)
von: Guo, Zijian, et al.
Veröffentlicht: (2026)
Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization
von: Gu, Shangding
Veröffentlicht: (2026)
von: Gu, Shangding
Veröffentlicht: (2026)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
Commit to the Bit: Reactive Reinforcement Learning Done Right
von: Eberhard, Onno, et al.
Veröffentlicht: (2026)
von: Eberhard, Onno, et al.
Veröffentlicht: (2026)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
von: Liu, Xianyang, et al.
Veröffentlicht: (2026)
von: Liu, Xianyang, et al.
Veröffentlicht: (2026)
A Review of Safe Reinforcement Learning: Methods, Theory and Applications
von: Gu, Shangding, et al.
Veröffentlicht: (2022)
von: Gu, Shangding, et al.
Veröffentlicht: (2022)
Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning
von: Gu, Shangding, et al.
Veröffentlicht: (2025)
von: Gu, Shangding, et al.
Veröffentlicht: (2025)
CombOptNet: Fit the Right NP-Hard Problem by Learning Integer Programming Constraints
von: Paulus, Anselm, et al.
Veröffentlicht: (2021)
von: Paulus, Anselm, et al.
Veröffentlicht: (2021)
LLMs as Assessors: Right for the Right Reason?
von: Saha, Sourav, et al.
Veröffentlicht: (2026)
von: Saha, Sourav, et al.
Veröffentlicht: (2026)
Right Reward Right Time for Federated Learning
von: Nguyen, Thanh Linh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thanh Linh, et al.
Veröffentlicht: (2025)
What is the Right Notion of Distance between Predict-then-Optimize Tasks?
von: Rodriguez-Diaz, Paula, et al.
Veröffentlicht: (2024)
von: Rodriguez-Diaz, Paula, et al.
Veröffentlicht: (2024)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
von: Tang, Yinghao, et al.
Veröffentlicht: (2025)
von: Tang, Yinghao, et al.
Veröffentlicht: (2025)
Right for the Right Reasons: Avoiding Reasoning Shortcuts via Prototypical Neurosymbolic AI
von: Andolfi, Luca, et al.
Veröffentlicht: (2025)
von: Andolfi, Luca, et al.
Veröffentlicht: (2025)
Finer is Better (with the Right Scaling)
von: Schaefer, Clemens, et al.
Veröffentlicht: (2026)
von: Schaefer, Clemens, et al.
Veröffentlicht: (2026)
Using Petri Nets as an Integrated Constraint Mechanism for Reinforcement Learning Tasks
von: Sachweh, Timon, et al.
Veröffentlicht: (2024)
von: Sachweh, Timon, et al.
Veröffentlicht: (2024)
StyleBench: Evaluating thinking styles in Large Language Models
von: Guo, Junyu, et al.
Veröffentlicht: (2025)
von: Guo, Junyu, et al.
Veröffentlicht: (2025)
LLMs Should Express Uncertainty Explicitly
von: Guo, Junyu, et al.
Veröffentlicht: (2026)
von: Guo, Junyu, et al.
Veröffentlicht: (2026)
Mutual Enhancement of Large Language and Reinforcement Learning Models through Bi-Directional Feedback Mechanisms: A Planning Case Study
von: Gu, Shangding
Veröffentlicht: (2024)
von: Gu, Shangding
Veröffentlicht: (2024)
MoRight: Motion Control Done Right
von: Liu, Shaowei, et al.
Veröffentlicht: (2026)
von: Liu, Shaowei, et al.
Veröffentlicht: (2026)
Learning Survival Models with Right-Censored Reporting Delays
von: Shikuri, Yuta, et al.
Veröffentlicht: (2025)
von: Shikuri, Yuta, et al.
Veröffentlicht: (2025)
Neural Network Compression for Reinforcement Learning Tasks
von: Ivanov, Dmitry A., et al.
Veröffentlicht: (2024)
von: Ivanov, Dmitry A., et al.
Veröffentlicht: (2024)
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
von: Guo, Junyu, et al.
Veröffentlicht: (2025)
von: Guo, Junyu, et al.
Veröffentlicht: (2025)
Many Ways to be Right: Rashomon Sets for Concept-Based Neural Networks
von: Feng, Shihan, et al.
Veröffentlicht: (2025)
von: Feng, Shihan, et al.
Veröffentlicht: (2025)
From Rights to Rites: Expectations Management in Smart-Home AI
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
von: Ye, Qilin, et al.
Veröffentlicht: (2025)
von: Ye, Qilin, et al.
Veröffentlicht: (2025)
Tuning the Right Foundation Models is What you Need for Partial Label Learning
von: He, Kuang, et al.
Veröffentlicht: (2025)
von: He, Kuang, et al.
Veröffentlicht: (2025)
Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond
von: Gu, Shangding, et al.
Veröffentlicht: (2025)
von: Gu, Shangding, et al.
Veröffentlicht: (2025)
Predict Confidently, Predict Right: Abstention in Dynamic Graph Learning
von: Gayen, Jayadratha, et al.
Veröffentlicht: (2025)
von: Gayen, Jayadratha, et al.
Veröffentlicht: (2025)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2025)
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2025)
Stochastic Gradient Descent for Gaussian Processes Done Right
von: Lin, Jihao Andreas, et al.
Veröffentlicht: (2023)
von: Lin, Jihao Andreas, et al.
Veröffentlicht: (2023)
On Fixing the Right Problems in Predictive Analytics: AUC Is Not the Problem
von: Baker, Ryan S., et al.
Veröffentlicht: (2024)
von: Baker, Ryan S., et al.
Veröffentlicht: (2024)
Learning from the Right Rollouts: Data Attribution for PPO-based LLM Post-Training
von: Shu, Dong, et al.
Veröffentlicht: (2026)
von: Shu, Dong, et al.
Veröffentlicht: (2026)
Transformers Can Do Arithmetic with the Right Embeddings
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
Elastic Weight Consolidation Done Right for Continual Learning
von: Liu, Xuan, et al.
Veröffentlicht: (2026)
von: Liu, Xuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving
von: Zheng, Zhi, et al.
Veröffentlicht: (2024) -
SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning
von: Guo, Zijian, et al.
Veröffentlicht: (2026) -
Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization
von: Gu, Shangding
Veröffentlicht: (2026) -
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
von: Wang, Yuqing, et al.
Veröffentlicht: (2025) -
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
von: Gu, Shangding, et al.
Veröffentlicht: (2024)