Enhancing RL Safety with Counterfactual LLM Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Gross, Dennis, Spieker, Helge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Safety-Oriented Pruning and Interpretation of Reinforcement Learning Policies
di: Gross, Dennis, et al.
Pubblicazione: (2024)
di: Gross, Dennis, et al.
Pubblicazione: (2024)
Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning Policies
di: Gross, Dennis, et al.
Pubblicazione: (2025)
di: Gross, Dennis, et al.
Pubblicazione: (2025)
Translating the Rashomon Effect to Sequential Decision-Making Tasks
di: Gross, Dennis, et al.
Pubblicazione: (2025)
di: Gross, Dennis, et al.
Pubblicazione: (2025)
Enhancing Manufacturing Quality Prediction Models through the Integration of Explainability Methods
di: Gross, Dennis, et al.
Pubblicazione: (2024)
di: Gross, Dennis, et al.
Pubblicazione: (2024)
Efficient Milling Quality Prediction with Explainable Machine Learning
di: Gross, Dennis, et al.
Pubblicazione: (2024)
di: Gross, Dennis, et al.
Pubblicazione: (2024)
COOL-MC: Verifying and Explaining RL Policies for Platelet Inventory Management
di: Gross, Dennis
Pubblicazione: (2026)
di: Gross, Dennis
Pubblicazione: (2026)
Semi-supervised CAPP Transformer Learning via Pseudo-labeling
di: Gross, Dennis, et al.
Pubblicazione: (2026)
di: Gross, Dennis, et al.
Pubblicazione: (2026)
COOL-MC: Verifying and Explaining RL Policies for Multi-bridge Network Maintenance
di: Gross, Dennis
Pubblicazione: (2026)
di: Gross, Dennis
Pubblicazione: (2026)
Probabilistic Model Checking of Stochastic Reinforcement Learning Policies
di: Gross, Dennis, et al.
Pubblicazione: (2024)
di: Gross, Dennis, et al.
Pubblicazione: (2024)
Budgeting Counterfactual for Offline RL
di: Liu, Yao, et al.
Pubblicazione: (2023)
di: Liu, Yao, et al.
Pubblicazione: (2023)
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning
di: Wang, Jingyao, et al.
Pubblicazione: (2026)
di: Wang, Jingyao, et al.
Pubblicazione: (2026)
Token-Efficient RL for LLM Reasoning
di: Lee, Alan, et al.
Pubblicazione: (2025)
di: Lee, Alan, et al.
Pubblicazione: (2025)
Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows
di: Cho, Minjae, et al.
Pubblicazione: (2024)
di: Cho, Minjae, et al.
Pubblicazione: (2024)
On Minimizing Adversarial Counterfactual Error in Adversarial RL
di: Belaire, Roman, et al.
Pubblicazione: (2024)
di: Belaire, Roman, et al.
Pubblicazione: (2024)
CFLight: Enhancing Safety with Traffic Signal Control through Counterfactual Learning
di: Li, Mingyuan, et al.
Pubblicazione: (2025)
di: Li, Mingyuan, et al.
Pubblicazione: (2025)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
di: Brantley, Kianté, et al.
Pubblicazione: (2025)
di: Brantley, Kianté, et al.
Pubblicazione: (2025)
Verifying Memoryless Sequential Decision-making of Large Language Models
di: Gross, Dennis, et al.
Pubblicazione: (2025)
di: Gross, Dennis, et al.
Pubblicazione: (2025)
Bounded PCTL Model Checking of Large Language Model Outputs
di: Gross, Dennis, et al.
Pubblicazione: (2025)
di: Gross, Dennis, et al.
Pubblicazione: (2025)
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
di: Sun, Ke, et al.
Pubblicazione: (2026)
di: Sun, Ke, et al.
Pubblicazione: (2026)
Formally Verifying and Explaining Sepsis Treatment Policies with COOL-MC
di: Gross, Dennis
Pubblicazione: (2026)
di: Gross, Dennis
Pubblicazione: (2026)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
di: Liu, Haolin, et al.
Pubblicazione: (2026)
di: Liu, Haolin, et al.
Pubblicazione: (2026)
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
di: Zhou, Yifei, et al.
Pubblicazione: (2025)
di: Zhou, Yifei, et al.
Pubblicazione: (2025)
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
di: Wu, Xian, et al.
Pubblicazione: (2026)
di: Wu, Xian, et al.
Pubblicazione: (2026)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
di: Dou, Shihan, et al.
Pubblicazione: (2025)
di: Dou, Shihan, et al.
Pubblicazione: (2025)
Rashomon in the Streets: Explanation Ambiguity in Scene Understanding
di: Spieker, Helge, et al.
Pubblicazione: (2025)
di: Spieker, Helge, et al.
Pubblicazione: (2025)
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
di: Joishy, Anish R, et al.
Pubblicazione: (2025)
di: Joishy, Anish R, et al.
Pubblicazione: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
di: Zhu, Xuekai, et al.
Pubblicazione: (2025)
di: Zhu, Xuekai, et al.
Pubblicazione: (2025)
ACTER: Diverse and Actionable Counterfactual Sequences for Explaining and Diagnosing RL Policies
di: Gajcin, Jasmina, et al.
Pubblicazione: (2024)
di: Gajcin, Jasmina, et al.
Pubblicazione: (2024)
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
di: Wu, Jinyang, et al.
Pubblicazione: (2025)
di: Wu, Jinyang, et al.
Pubblicazione: (2025)
Offline Imitation Learning with Variational Counterfactual Reasoning
di: He, Bowei, et al.
Pubblicazione: (2023)
di: He, Bowei, et al.
Pubblicazione: (2023)
Turn-based Multi-Agent Reinforcement Learning Model Checking
di: Gross, Dennis
Pubblicazione: (2025)
di: Gross, Dennis
Pubblicazione: (2025)
PickLLM: Context-Aware RL-Assisted Large Language Model Routing
di: Sikeridis, Dimitrios, et al.
Pubblicazione: (2024)
di: Sikeridis, Dimitrios, et al.
Pubblicazione: (2024)
On Designing Effective RL Reward at Training Time for LLM Reasoning
di: Gao, Jiaxuan, et al.
Pubblicazione: (2024)
di: Gao, Jiaxuan, et al.
Pubblicazione: (2024)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
di: Limozin, Alexis, et al.
Pubblicazione: (2026)
di: Limozin, Alexis, et al.
Pubblicazione: (2026)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
di: Hu, Zhengding, et al.
Pubblicazione: (2026)
di: Hu, Zhengding, et al.
Pubblicazione: (2026)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
di: Samadi, Amir, et al.
Pubblicazione: (2024)
di: Samadi, Amir, et al.
Pubblicazione: (2024)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
di: Liu, Zihe, et al.
Pubblicazione: (2025)
di: Liu, Zihe, et al.
Pubblicazione: (2025)
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
di: Howe, Nikolaus, et al.
Pubblicazione: (2025)
di: Howe, Nikolaus, et al.
Pubblicazione: (2025)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
RAGEN-2: Reasoning Collapse in Agentic RL
di: Wang, Zihan, et al.
Pubblicazione: (2026)
di: Wang, Zihan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Safety-Oriented Pruning and Interpretation of Reinforcement Learning Policies
di: Gross, Dennis, et al.
Pubblicazione: (2024) -
Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning Policies
di: Gross, Dennis, et al.
Pubblicazione: (2025) -
Translating the Rashomon Effect to Sequential Decision-Making Tasks
di: Gross, Dennis, et al.
Pubblicazione: (2025) -
Enhancing Manufacturing Quality Prediction Models through the Integration of Explainability Methods
di: Gross, Dennis, et al.
Pubblicazione: (2024) -
Efficient Milling Quality Prediction with Explainable Machine Learning
di: Gross, Dennis, et al.
Pubblicazione: (2024)