Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Cousins, Cyrus, Keswani, Vijay, Conitzer, Vincent, Heidari, Hoda, Borg, Jana Schaich, Sinnott-Armstrong, Walter |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Pros and Cons of Active Learning for Moral Preference Elicitation
di: Keswani, Vijay, et al.
Pubblicazione: (2024)
di: Keswani, Vijay, et al.
Pubblicazione: (2024)
Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
di: Keswani, Vijay, et al.
Pubblicazione: (2025)
di: Keswani, Vijay, et al.
Pubblicazione: (2025)
Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
di: Keswani, Vijay, et al.
Pubblicazione: (2025)
di: Keswani, Vijay, et al.
Pubblicazione: (2025)
On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
di: Boerstler, Kyle, et al.
Pubblicazione: (2024)
di: Boerstler, Kyle, et al.
Pubblicazione: (2024)
Shutdown Safety Valves for Advanced AI
di: Conitzer, Vincent
Pubblicazione: (2026)
di: Conitzer, Vincent
Pubblicazione: (2026)
What Is Required for Empathic AI? It Depends, and Why That Matters for AI Developers and Users
di: Borg, Jana Schaich, et al.
Pubblicazione: (2024)
di: Borg, Jana Schaich, et al.
Pubblicazione: (2024)
To Pool or Not To Pool: Analyzing the Regularizing Effects of Group-Fair Training on Shared Models
di: Cousins, Cyrus, et al.
Pubblicazione: (2024)
di: Cousins, Cyrus, et al.
Pubblicazione: (2024)
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
di: Buckley, Warren, et al.
Pubblicazione: (2023)
di: Buckley, Warren, et al.
Pubblicazione: (2023)
Towards Stable Preferences for Stakeholder-aligned Machine Learning
di: Sheraz, Haleema, et al.
Pubblicazione: (2024)
di: Sheraz, Haleema, et al.
Pubblicazione: (2024)
Towards Neural Network based Cognitive Models of Dynamic Decision-Making by Humans
di: Chen, Changyu, et al.
Pubblicazione: (2024)
di: Chen, Changyu, et al.
Pubblicazione: (2024)
Fair Classification with Partial Feedback: An Exploration-Based Data Collection Approach
di: Keswani, Vijay, et al.
Pubblicazione: (2024)
di: Keswani, Vijay, et al.
Pubblicazione: (2024)
Percentile Criterion Optimization in Offline Reinforcement Learning
di: Lobo, Elita A., et al.
Pubblicazione: (2024)
di: Lobo, Elita A., et al.
Pubblicazione: (2024)
Fair and Welfare-Efficient Constrained Multi-matchings under Uncertainty
di: Lobo, Elita, et al.
Pubblicazione: (2024)
di: Lobo, Elita, et al.
Pubblicazione: (2024)
DoubleTake: Contrastive Reasoning for Faithful Decision-Making in Medical Imaging
di: Patel, Daivik, et al.
Pubblicazione: (2026)
di: Patel, Daivik, et al.
Pubblicazione: (2026)
FedGTEA: Federated Class-Incremental Learning with Gaussian Task Embedding and Alignment
di: Li, Haolin, et al.
Pubblicazione: (2025)
di: Li, Haolin, et al.
Pubblicazione: (2025)
An Interpretable Automated Mechanism Design Framework with Large Language Models
di: Liu, Jiayuan, et al.
Pubblicazione: (2025)
di: Liu, Jiayuan, et al.
Pubblicazione: (2025)
Learning Safe Autonomous Driving Policies Using Predictive Safety Representations
di: Keswani, Mahesh, et al.
Pubblicazione: (2025)
di: Keswani, Mahesh, et al.
Pubblicazione: (2025)
Rethinking Distance Metrics for Counterfactual Explainability
di: Williams, Joshua Nathaniel, et al.
Pubblicazione: (2024)
di: Williams, Joshua Nathaniel, et al.
Pubblicazione: (2024)
Supporting Deep Learning Solutions for Diverse Application Domains
di: Sinnott, Richard
Pubblicazione: (2024)
di: Sinnott, Richard
Pubblicazione: (2024)
Towards Responsible AI in Banking: Addressing Bias for Fair Decision-Making
di: Castelnovo, Alessandro
Pubblicazione: (2024)
di: Castelnovo, Alessandro
Pubblicazione: (2024)
Incorporating Cognitive Biases into Reinforcement Learning for Financial Decision-Making
di: He, Liu
Pubblicazione: (2026)
di: He, Liu
Pubblicazione: (2026)
Towards Cost Sensitive Decision Making
di: Li, Yang, et al.
Pubblicazione: (2024)
di: Li, Yang, et al.
Pubblicazione: (2024)
Red-Teaming for Generative AI: Silver Bullet or Security Theater?
di: Feffer, Michael, et al.
Pubblicazione: (2024)
di: Feffer, Michael, et al.
Pubblicazione: (2024)
Fairness in Criminal Justice Risk Assessments: The State of the Art
di: Berk, Richard A., et al.
Pubblicazione: (2017)
di: Berk, Richard A., et al.
Pubblicazione: (2017)
Are LLM Decisions Faithful to Verbal Confidence?
di: Wang, Jiawei, et al.
Pubblicazione: (2026)
di: Wang, Jiawei, et al.
Pubblicazione: (2026)
FaithLM: Towards Faithful Explanations for Large Language Models
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants
di: Kapania, Shivani, et al.
Pubblicazione: (2024)
di: Kapania, Shivani, et al.
Pubblicazione: (2024)
Improved Compression Bounds for Scenario Decision Making
di: Berger, Guillaume O.
Pubblicazione: (2025)
di: Berger, Guillaume O.
Pubblicazione: (2025)
Deep Time Warping for Multiple Time Series Alignment
di: Nourbakhsh, Alireza, et al.
Pubblicazione: (2025)
di: Nourbakhsh, Alireza, et al.
Pubblicazione: (2025)
Decision Making under Imperfect Recall: Algorithms and Benchmarks
di: Tewolde, Emanuel, et al.
Pubblicazione: (2026)
di: Tewolde, Emanuel, et al.
Pubblicazione: (2026)
Efficiently Solving Turn-Taking Stochastic Games with Extensive-Form Correlation
di: Zhang, Hanrui, et al.
Pubblicazione: (2024)
di: Zhang, Hanrui, et al.
Pubblicazione: (2024)
Improving Transformers using Faithful Positional Encoding
di: Idé, Tsuyoshi, et al.
Pubblicazione: (2024)
di: Idé, Tsuyoshi, et al.
Pubblicazione: (2024)
Improving Health Professionals' Onboarding with AI and XAI for Trustworthy Human-AI Collaborative Decision Making
di: Lee, Min Hun, et al.
Pubblicazione: (2024)
di: Lee, Min Hun, et al.
Pubblicazione: (2024)
CAMAL: Improving Attention Alignment and Faithfulness with Segmentation Masks
di: Hundal, Rajdeep Singh, et al.
Pubblicazione: (2026)
di: Hundal, Rajdeep Singh, et al.
Pubblicazione: (2026)
Towards Establishing Guaranteed Error for Learned Database Operations
di: Zeighami, Sepanta, et al.
Pubblicazione: (2024)
di: Zeighami, Sepanta, et al.
Pubblicazione: (2024)
Towards Uncertainty Aware Task Delegation and Human-AI Collaborative Decision-Making
di: Lee, Min Hun, et al.
Pubblicazione: (2025)
di: Lee, Min Hun, et al.
Pubblicazione: (2025)
Safe Langevin Soft Actor Critic
di: Keswani, Mahesh, et al.
Pubblicazione: (2026)
di: Keswani, Mahesh, et al.
Pubblicazione: (2026)
Improving Human Sequential Decision-Making with Reinforcement Learning
di: Bastani, Hamsa, et al.
Pubblicazione: (2021)
di: Bastani, Hamsa, et al.
Pubblicazione: (2021)
Conformal Prediction Sets Improve Human Decision Making
di: Cresswell, Jesse C., et al.
Pubblicazione: (2024)
di: Cresswell, Jesse C., et al.
Pubblicazione: (2024)
Towards Faithful Multimodal Concept Bottleneck Models
di: Moreau, Pierre, et al.
Pubblicazione: (2026)
di: Moreau, Pierre, et al.
Pubblicazione: (2026)
Documenti analoghi
-
On the Pros and Cons of Active Learning for Moral Preference Elicitation
di: Keswani, Vijay, et al.
Pubblicazione: (2024) -
Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
di: Keswani, Vijay, et al.
Pubblicazione: (2025) -
Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
di: Keswani, Vijay, et al.
Pubblicazione: (2025) -
On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
di: Boerstler, Kyle, et al.
Pubblicazione: (2024) -
Shutdown Safety Valves for Advanced AI
di: Conitzer, Vincent
Pubblicazione: (2026)