Tell me why: Training preferences-based RL with human preferences and step-level explanations
Fuente:
arXiv
Saved in:
| Main Author: | Karalus, Jakob |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The future of human-centric eXplainable Artificial Intelligence (XAI) is not post-hoc explanations
by: Swamy, Vinitra, et al.
Published: (2023)
by: Swamy, Vinitra, et al.
Published: (2023)
Everyone prefers human writers, including AI
by: Haverals, Wouter, et al.
Published: (2025)
by: Haverals, Wouter, et al.
Published: (2025)
Why do explanations fail? A typology and discussion on failures in XAI
by: Bove, Clara, et al.
Published: (2024)
by: Bove, Clara, et al.
Published: (2024)
Interactive dense pixel visualizations for time series and model attribution explanations
by: Schlegel, Udo, et al.
Published: (2024)
by: Schlegel, Udo, et al.
Published: (2024)
Explainable agency: human preferences for simple or complex explanations
by: Blom, Michelle, et al.
Published: (2024)
by: Blom, Michelle, et al.
Published: (2024)
Making RL with Preference-based Feedback Efficient via Randomization
by: Wu, Runzhe, et al.
Published: (2023)
by: Wu, Runzhe, et al.
Published: (2023)
Humans learn to prefer trustworthy AI over human partners
by: Jiang, Yaomin, et al.
Published: (2025)
by: Jiang, Yaomin, et al.
Published: (2025)
In defence of post-hoc explanations in medical AI
by: Hatherley, Joshua, et al.
Published: (2025)
by: Hatherley, Joshua, et al.
Published: (2025)
Policy alone is probably not the solution: A large-scale experiment on how developers struggle to design meaningful end-user explanations
by: Nahar, Nadia, et al.
Published: (2025)
by: Nahar, Nadia, et al.
Published: (2025)
ExplainReduce: Generating global explanations from many local explanations
by: Seppäläinen, Lauri, et al.
Published: (2025)
by: Seppäläinen, Lauri, et al.
Published: (2025)
Compositional learning of functions in humans and machines
by: Zhou, Yanli, et al.
Published: (2024)
by: Zhou, Yanli, et al.
Published: (2024)
Tell me more: Intent Fulfilment Framework for Enhancing User Experiences in Conversational XAI
by: Wijekoon, Anjana, et al.
Published: (2024)
by: Wijekoon, Anjana, et al.
Published: (2024)
TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health
by: Fan, Yuang, et al.
Published: (2026)
by: Fan, Yuang, et al.
Published: (2026)
The two-way knowledge interaction interface between humans and neural networks
by: He, Zhanliang, et al.
Published: (2024)
by: He, Zhanliang, et al.
Published: (2024)
WebXAII: an open-source web framework to study human-XAI interaction
by: Leguy, Jules, et al.
Published: (2025)
by: Leguy, Jules, et al.
Published: (2025)
AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
by: Agarwal, Shyam, et al.
Published: (2025)
by: Agarwal, Shyam, et al.
Published: (2025)
PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories
by: Aroca-Ouellette, Stephane, et al.
Published: (2024)
by: Aroca-Ouellette, Stephane, et al.
Published: (2024)
Limited but consistent gains in adversarial robustness by co-training object recognition models with human EEG
by: Guo, Manshan, et al.
Published: (2024)
by: Guo, Manshan, et al.
Published: (2024)
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
by: Su, Yiheng, et al.
Published: (2023)
by: Su, Yiheng, et al.
Published: (2023)
Alignment-Based Adversarial Training (ABAT) for Improving the Robustness and Accuracy of EEG-Based BCIs
by: Chen, Xiaoqing, et al.
Published: (2024)
by: Chen, Xiaoqing, et al.
Published: (2024)
Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human Alignment
by: Zhang, Chen, et al.
Published: (2024)
by: Zhang, Chen, et al.
Published: (2024)
Towards User-Focused Research in Training Data Attribution for Human-Centered Explainable AI
by: Nguyen, Elisa, et al.
Published: (2024)
by: Nguyen, Elisa, et al.
Published: (2024)
Perils of Label Indeterminacy: A Case Study on Prediction of Neurological Recovery After Cardiac Arrest
by: Schoeffer, Jakob, et al.
Published: (2025)
by: Schoeffer, Jakob, et al.
Published: (2025)
Evaluating graph-based explanations for AI-based recommender systems
by: Delarue, Simon, et al.
Published: (2024)
by: Delarue, Simon, et al.
Published: (2024)
"Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations
by: Wu, Yang, et al.
Published: (2026)
by: Wu, Yang, et al.
Published: (2026)
Towards a copilot in BIM authoring tool using a large language model-based agent for intelligent human-machine interaction
by: Du, Changyu, et al.
Published: (2024)
by: Du, Changyu, et al.
Published: (2024)
A shortest-path based clustering algorithm for joint human-machine analysis of complex datasets
by: Pizzagalli, Diego Ulisse, et al.
Published: (2018)
by: Pizzagalli, Diego Ulisse, et al.
Published: (2018)
Introducing User Feedback-based Counterfactual Explanations (UFCE)
by: Suffian, Muhammad, et al.
Published: (2024)
by: Suffian, Muhammad, et al.
Published: (2024)
LookALike: Human Mimicry based collaborative decision making
by: Karanjai, Rabimba, et al.
Published: (2024)
by: Karanjai, Rabimba, et al.
Published: (2024)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Imitation of human motion achieves natural head movements for humanoid robots in an active-speaker detection task
by: Ding, Bosong, et al.
Published: (2024)
by: Ding, Bosong, et al.
Published: (2024)
Protecting Multiple Types of Privacy Simultaneously in EEG-based Brain-Computer Interfaces
by: Meng, Lubin, et al.
Published: (2024)
by: Meng, Lubin, et al.
Published: (2024)
Spatial Distillation based Distribution Alignment (SDDA) for Cross-Headset EEG Classification
by: Liu, Dingkun, et al.
Published: (2025)
by: Liu, Dingkun, et al.
Published: (2025)
DiConStruct: Causal Concept-based Explanations through Black-Box Distillation
by: Moreira, Ricardo, et al.
Published: (2024)
by: Moreira, Ricardo, et al.
Published: (2024)
Explainability of Recurrent Neural Networks for Enhancing P300-based Brain-Computer Interfaces
by: Oliva, Christian, et al.
Published: (2026)
by: Oliva, Christian, et al.
Published: (2026)
A Pain Assessment Framework based on multimodal data and Deep Machine Learning methods
by: Gkikas, Stefanos
Published: (2025)
by: Gkikas, Stefanos
Published: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
by: Cook, Jonathan, et al.
Published: (2024)
by: Cook, Jonathan, et al.
Published: (2024)
InFiConD: Interactive No-code Fine-tuning with Concept-based Knowledge Distillation
by: Huang, Jinbin, et al.
Published: (2024)
by: Huang, Jinbin, et al.
Published: (2024)
Bidirectional human-AI collaboration in brain tumour assessments improves both expert human and AI agent performance
by: Ruffle, James K, et al.
Published: (2025)
by: Ruffle, James K, et al.
Published: (2025)
Adversarial Domain Adaptation for Cross-user Activity Recognition Using Diffusion-based Noise-centred Learning
by: Ye, Xiaozhou, et al.
Published: (2024)
by: Ye, Xiaozhou, et al.
Published: (2024)
Similar Items
-
The future of human-centric eXplainable Artificial Intelligence (XAI) is not post-hoc explanations
by: Swamy, Vinitra, et al.
Published: (2023) -
Everyone prefers human writers, including AI
by: Haverals, Wouter, et al.
Published: (2025) -
Why do explanations fail? A typology and discussion on failures in XAI
by: Bove, Clara, et al.
Published: (2024) -
Interactive dense pixel visualizations for time series and model attribution explanations
by: Schlegel, Udo, et al.
Published: (2024) -
Explainable agency: human preferences for simple or complex explanations
by: Blom, Michelle, et al.
Published: (2024)