Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Le, Tuan, Shefin, Risal, Gupta, Debashis, Le, Thai, Alqahtani, Sarra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretable Failure Analysis in Multi-Agent Reinforcement Learning Systems
von: Shefin, Risal Shahriar, et al.
Veröffentlicht: (2026)
von: Shefin, Risal Shahriar, et al.
Veröffentlicht: (2026)
xSRL: Safety-Aware Explainable Reinforcement Learning -- Safety as a Product of Explainability
von: Shefin, Risal Shahriar, et al.
Veröffentlicht: (2024)
von: Shefin, Risal Shahriar, et al.
Veröffentlicht: (2024)
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
von: Park, Seohong, et al.
Veröffentlicht: (2023)
von: Park, Seohong, et al.
Veröffentlicht: (2023)
Reinforcement Learning by Guided Safe Exploration
von: Yang, Qisong, et al.
Veröffentlicht: (2023)
von: Yang, Qisong, et al.
Veröffentlicht: (2023)
Enhance Exploration in Safe Reinforcement Learning with Contrastive Representation Learning
von: Doan, Duc Kien, et al.
Veröffentlicht: (2025)
von: Doan, Duc Kien, et al.
Veröffentlicht: (2025)
Policy Learning for Off-Dynamics RL with Deficient Support
von: Van, Linh Le Pham, et al.
Veröffentlicht: (2024)
von: Van, Linh Le Pham, et al.
Veröffentlicht: (2024)
Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in Conversation
von: Mai, Anh-Tuan, et al.
Veröffentlicht: (2026)
von: Mai, Anh-Tuan, et al.
Veröffentlicht: (2026)
ASGM-KG: Unveiling Alluvial Gold Mining Through Knowledge Graphs
von: Gupta, Debashis, et al.
Veröffentlicht: (2024)
von: Gupta, Debashis, et al.
Veröffentlicht: (2024)
SafeAR: Safe Algorithmic Recourse by Risk-Aware Policies
von: Wu, Haochen, et al.
Veröffentlicht: (2023)
von: Wu, Haochen, et al.
Veröffentlicht: (2023)
Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
Safe Exploration via Policy Priors
von: Wendl, Manuel, et al.
Veröffentlicht: (2026)
von: Wendl, Manuel, et al.
Veröffentlicht: (2026)
Precise Verification of Transformers through ReLU-Catalyzed Abstraction Refinement
von: Liu, Hengjie, et al.
Veröffentlicht: (2026)
von: Liu, Hengjie, et al.
Veröffentlicht: (2026)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction
von: GX-Chen, Anthony, et al.
Veröffentlicht: (2024)
von: GX-Chen, Anthony, et al.
Veröffentlicht: (2024)
Neural Network Verification with PyRAT
von: Lemesle, Augustin, et al.
Veröffentlicht: (2024)
von: Lemesle, Augustin, et al.
Veröffentlicht: (2024)
Preference-Conditioned Language-Guided Abstraction
von: Peng, Andi, et al.
Veröffentlicht: (2024)
von: Peng, Andi, et al.
Veröffentlicht: (2024)
Learning with Language-Guided State Abstractions
von: Peng, Andi, et al.
Veröffentlicht: (2024)
von: Peng, Andi, et al.
Veröffentlicht: (2024)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
Revisiting Safe Exploration in Safe Reinforcement learning
von: Eckel, David, et al.
Veröffentlicht: (2024)
von: Eckel, David, et al.
Veröffentlicht: (2024)
Enhancing Robustness of Offline Reinforcement Learning Under Data Corruption via Sharpness-Aware Minimization
von: Xu, Le, et al.
Veröffentlicht: (2025)
von: Xu, Le, et al.
Veröffentlicht: (2025)
In-Context Planning with Latent Temporal Abstractions
von: Luo, Baiting, et al.
Veröffentlicht: (2026)
von: Luo, Baiting, et al.
Veröffentlicht: (2026)
Sharpness-Aware Teleportation on Riemannian Manifolds
von: Truong, Tuan, et al.
Veröffentlicht: (2023)
von: Truong, Tuan, et al.
Veröffentlicht: (2023)
Meta-RL Induces Exploration in Language Agents
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
Beyond Individual Facts: Investigating Categorical Knowledge Locality of Taxonomy and Meronomy Concepts in GPT Models
von: Burger, Christopher, et al.
Veröffentlicht: (2024)
von: Burger, Christopher, et al.
Veröffentlicht: (2024)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
von: Cai, Xin-Qiang, et al.
Veröffentlicht: (2026)
von: Cai, Xin-Qiang, et al.
Veröffentlicht: (2026)
Variable-Agnostic Causal Exploration for Reinforcement Learning
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2024)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2024)
Guided Game Level Repair via Explainable AI
von: Bazzaz, Mahsa, et al.
Veröffentlicht: (2024)
von: Bazzaz, Mahsa, et al.
Veröffentlicht: (2024)
ART: Adaptive Reasoning Trees for Explainable Claim Verification
von: Wadhwa, Sahil, et al.
Veröffentlicht: (2026)
von: Wadhwa, Sahil, et al.
Veröffentlicht: (2026)
Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
von: Gupta, Prakhar, et al.
Veröffentlicht: (2025)
von: Gupta, Prakhar, et al.
Veröffentlicht: (2025)
Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing
von: Khan, Azal Ahmad, et al.
Veröffentlicht: (2026)
von: Khan, Azal Ahmad, et al.
Veröffentlicht: (2026)
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
von: Zheng, Huaixiu Steven, et al.
Veröffentlicht: (2023)
von: Zheng, Huaixiu Steven, et al.
Veröffentlicht: (2023)
AIGS: Generating Science from AI-Powered Automated Falsification
von: Liu, Zijun, et al.
Veröffentlicht: (2024)
von: Liu, Zijun, et al.
Veröffentlicht: (2024)
Causal-Aware Generative Adversarial Networks with Reinforcement Learning
von: Nguyen, Tu Anh Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Tu Anh Hoang, et al.
Veröffentlicht: (2025)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
von: Li, Ziniu, et al.
Veröffentlicht: (2025)
von: Li, Ziniu, et al.
Veröffentlicht: (2025)
RL for Mitigating Cascading Failures: Targeted Exploration via Sensitivity Factors
von: Dwivedi, Anmol, et al.
Veröffentlicht: (2024)
von: Dwivedi, Anmol, et al.
Veröffentlicht: (2024)
Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk
von: Ying, Chengyang, et al.
Veröffentlicht: (2022)
von: Ying, Chengyang, et al.
Veröffentlicht: (2022)
Uniformly Safe RL with Objective Suppression for Multi-Constraint Safety-Critical Applications
von: Zhou, Zihan, et al.
Veröffentlicht: (2024)
von: Zhou, Zihan, et al.
Veröffentlicht: (2024)
TLDR: Unsupervised Goal-Conditioned RL via Temporal Distance-Aware Representations
von: Bae, Junik, et al.
Veröffentlicht: (2024)
von: Bae, Junik, et al.
Veröffentlicht: (2024)
FLAME: Towards Federated Fine-Tuning Large Language Models Through Adaptive SMoE
von: Le, Khiem, et al.
Veröffentlicht: (2025)
von: Le, Khiem, et al.
Veröffentlicht: (2025)
Sample-efficient and Scalable Exploration in Continuous-Time RL
von: Iten, Klemens, et al.
Veröffentlicht: (2025)
von: Iten, Klemens, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interpretable Failure Analysis in Multi-Agent Reinforcement Learning Systems
von: Shefin, Risal Shahriar, et al.
Veröffentlicht: (2026) -
xSRL: Safety-Aware Explainable Reinforcement Learning -- Safety as a Product of Explainability
von: Shefin, Risal Shahriar, et al.
Veröffentlicht: (2024) -
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
von: Park, Seohong, et al.
Veröffentlicht: (2023) -
Reinforcement Learning by Guided Safe Exploration
von: Yang, Qisong, et al.
Veröffentlicht: (2023) -
Enhance Exploration in Safe Reinforcement Learning with Contrastive Representation Learning
von: Doan, Duc Kien, et al.
Veröffentlicht: (2025)