XRL-Bench: A Benchmark for Evaluating and Comparing Explainable Reinforcement Learning Techniques
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiong, Yu, Hu, Zhipeng, Huang, Ye, Wu, Runze, Guan, Kai, Fang, Xingchen, Jiang, Ji, Zhou, Tianze, Hu, Yujing, Liu, Haoyu, Lyu, Tangjie, Fan, Changjie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards
von: Jiang, Zhaohui, et al.
Veröffentlicht: (2024)
von: Jiang, Zhaohui, et al.
Veröffentlicht: (2024)
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
von: Liu, Jinyi, et al.
Veröffentlicht: (2023)
von: Liu, Jinyi, et al.
Veröffentlicht: (2023)
Bayesian Design Principles for Offline-to-Online Reinforcement Learning
von: Hu, Hao, et al.
Veröffentlicht: (2024)
von: Hu, Hao, et al.
Veröffentlicht: (2024)
Digital Player: Evaluating Large Language Models based Human-like Agent in Games
von: Wang, Jiawei, et al.
Veröffentlicht: (2025)
von: Wang, Jiawei, et al.
Veröffentlicht: (2025)
AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model
von: Dong, Zibin, et al.
Veröffentlicht: (2023)
von: Dong, Zibin, et al.
Veröffentlicht: (2023)
SymbXRL: Symbolic Explainable Deep Reinforcement Learning for Mobile Networks
von: Duttagupta, Abhishek, et al.
Veröffentlicht: (2026)
von: Duttagupta, Abhishek, et al.
Veröffentlicht: (2026)
Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling
von: Xu, Bihan, et al.
Veröffentlicht: (2025)
von: Xu, Bihan, et al.
Veröffentlicht: (2025)
vMFER: Von Mises-Fisher Experience Resampling Based on Uncertainty of Gradient Directions for Policy Improvement
von: Zhu, Yiwen, et al.
Veröffentlicht: (2024)
von: Zhu, Yiwen, et al.
Veröffentlicht: (2024)
Towards a Simultaneous and Granular Identity-Expression Control in Personalized Face Generation
von: Liu, Renshuai, et al.
Veröffentlicht: (2024)
von: Liu, Renshuai, et al.
Veröffentlicht: (2024)
StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads
von: Wang, Suzhen, et al.
Veröffentlicht: (2024)
von: Wang, Suzhen, et al.
Veröffentlicht: (2024)
A New Baseline Assumption of Integated Gradients Based on Shaply value
von: Liu, Shuyang, et al.
Veröffentlicht: (2023)
von: Liu, Shuyang, et al.
Veröffentlicht: (2023)
The Effects of Data Augmentation on Confidence Estimation for LLMs
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
A Comparative User Evaluation of XRL Explanations using Goal Identification
von: Towers, Mark, et al.
Veröffentlicht: (2025)
von: Towers, Mark, et al.
Veröffentlicht: (2025)
TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
von: Ma, Yifeng, et al.
Veröffentlicht: (2023)
von: Ma, Yifeng, et al.
Veröffentlicht: (2023)
ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry
von: Xu, Tianze, et al.
Veröffentlicht: (2025)
von: Xu, Tianze, et al.
Veröffentlicht: (2025)
A Dataset for the Validation of Truth Inference Algorithms Suitable for Online Deployment
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
An inexact primal-dual method with correction step for a saddle point problem in image debluring
von: Fang, Changjie, et al.
Veröffentlicht: (2021)
von: Fang, Changjie, et al.
Veröffentlicht: (2021)
CharacterBench: Benchmarking Character Customization of Large Language Models
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2024)
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2024)
ViroBench: Benchmarking Nucleotide Foundation Models on Viral Genomics Tasks
von: Ye, Dongxin, et al.
Veröffentlicht: (2026)
von: Ye, Dongxin, et al.
Veröffentlicht: (2026)
WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
Hypothalamic malate dehydrogenase 2 (MDH2) modulates systemic glucose metabolism through oxytocin-mediated thermogenesis
von: Tianze, Xiong
Veröffentlicht: (2025)
von: Tianze, Xiong
Veröffentlicht: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
Fast-DataShapley: Neural Modeling for Training Data Valuation
von: Sun, Haifeng, et al.
Veröffentlicht: (2025)
von: Sun, Haifeng, et al.
Veröffentlicht: (2025)
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
von: Hu, Caiyu, et al.
Veröffentlicht: (2025)
von: Hu, Caiyu, et al.
Veröffentlicht: (2025)
FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment
von: Xiong, Betty, et al.
Veröffentlicht: (2026)
von: Xiong, Betty, et al.
Veröffentlicht: (2026)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
von: Xiong, Zixin, et al.
Veröffentlicht: (2026)
von: Xiong, Zixin, et al.
Veröffentlicht: (2026)
ICE: Interactive 3D Game Character Editing via Dialogue
von: Wu, Haoqian, et al.
Veröffentlicht: (2024)
von: Wu, Haoqian, et al.
Veröffentlicht: (2024)
Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection
von: Ma, Yuhang, et al.
Veröffentlicht: (2024)
von: Ma, Yuhang, et al.
Veröffentlicht: (2024)
Character-Adapter: Prompt-Guided Region Control for High-Fidelity Character Customization
von: Ma, Yuhang, et al.
Veröffentlicht: (2024)
von: Ma, Yuhang, et al.
Veröffentlicht: (2024)
Study on Aramid Nanofibers‐Reinforced Epoxy All‐Organic Composites With Enhanced Thermal Conductivity and Suppressed Thermal Expansion
von: Fan Yang, et al.
Veröffentlicht: (2025)
von: Fan Yang, et al.
Veröffentlicht: (2025)
VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
von: Li, Xuzhao, et al.
Veröffentlicht: (2025)
von: Li, Xuzhao, et al.
Veröffentlicht: (2025)
ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
von: Cheng, Ao, et al.
Veröffentlicht: (2026)
von: Cheng, Ao, et al.
Veröffentlicht: (2026)
CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs
von: Fang, Zhengru, et al.
Veröffentlicht: (2026)
von: Fang, Zhengru, et al.
Veröffentlicht: (2026)
Comparative Evaluation of High‐Frequency Microneedling Using a Layering Technique Versus Conventional Technique for Facial Rejuvenation
von: Ye Tang, et al.
Veröffentlicht: (2026)
von: Ye Tang, et al.
Veröffentlicht: (2026)
F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration
von: Liu, Lu, et al.
Veröffentlicht: (2024)
von: Liu, Lu, et al.
Veröffentlicht: (2024)
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
von: Li, Haoran, et al.
Veröffentlicht: (2023)
von: Li, Haoran, et al.
Veröffentlicht: (2023)
XRL: An FMM-Accelerated SIE Simulator for Resistance and Inductance Extraction of Complicated 3-D Geometries
von: Wang, Mingyu, et al.
Veröffentlicht: (2024)
von: Wang, Mingyu, et al.
Veröffentlicht: (2024)
FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation Model
von: Qiu, Feng, et al.
Veröffentlicht: (2024)
von: Qiu, Feng, et al.
Veröffentlicht: (2024)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards
von: Jiang, Zhaohui, et al.
Veröffentlicht: (2024) -
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
von: Liu, Jinyi, et al.
Veröffentlicht: (2023) -
Bayesian Design Principles for Offline-to-Online Reinforcement Learning
von: Hu, Hao, et al.
Veröffentlicht: (2024) -
Digital Player: Evaluating Large Language Models based Human-like Agent in Games
von: Wang, Jiawei, et al.
Veröffentlicht: (2025) -
AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model
von: Dong, Zibin, et al.
Veröffentlicht: (2023)