"A Good Bot Always Knows Its Limitations": Assessing Autonomous System Decision-making Competencies through Factorized Machine Self-confidence
Fuente:
arXiv
Saved in:
| Main Authors: | Israelsen, Brett W., Ahmed, Nisar R., Aitken, Matthew, Frew, Eric W., Lawrence, Dale A., Argrow, Brian M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance
by: Israelsen, Brett, et al.
Published: (2026)
by: Israelsen, Brett, et al.
Published: (2026)
Expected Moral Shortfall for Ethical Competence in Decision-making Models
by: Aijaz, Aisha, et al.
Published: (2026)
by: Aijaz, Aisha, et al.
Published: (2026)
Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
by: Ahmed, Nisar, et al.
Published: (2025)
by: Ahmed, Nisar, et al.
Published: (2025)
Assessing confidence in frontier AI safety cases
by: Barrett, Stephen, et al.
Published: (2025)
by: Barrett, Stephen, et al.
Published: (2025)
Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents
by: Xu, Rongwu, et al.
Published: (2025)
by: Xu, Rongwu, et al.
Published: (2025)
Using Surprise Index for Competency Assessment in Autonomous Decision-Making
by: Ratheesh, Akash, et al.
Published: (2023)
by: Ratheesh, Akash, et al.
Published: (2023)
RadProPoser: Probabilistic Radar Tensor Human Pose Estimation That Knows Its Limits
by: Mueller, Jonas Leo, et al.
Published: (2025)
by: Mueller, Jonas Leo, et al.
Published: (2025)
Tidynote: Always-Clear Notebook Authoring
by: Huang, Ruanqianqian, et al.
Published: (2026)
by: Huang, Ruanqianqian, et al.
Published: (2026)
Human Decision-making is Susceptible to AI-driven Manipulation
by: Sabour, Sahand, et al.
Published: (2025)
by: Sabour, Sahand, et al.
Published: (2025)
Bot Wars Evolved: Orchestrating Competing LLMs in a Counterstrike Against Phone Scams
by: Basta, Nardine, et al.
Published: (2025)
by: Basta, Nardine, et al.
Published: (2025)
FHIRPath-QA: Executable Question Answering over FHIR Electronic Health Records
by: Frew, Michael, et al.
Published: (2026)
by: Frew, Michael, et al.
Published: (2026)
D2E-An Autonomous Decision-making Dataset involving Driver States and Human Evaluation
by: Ke, Zehong, et al.
Published: (2024)
by: Ke, Zehong, et al.
Published: (2024)
Knowing What's Missing: Assessing Information Sufficiency in Question Answering
by: Jain, Akriti, et al.
Published: (2025)
by: Jain, Akriti, et al.
Published: (2025)
FATe of Bots: Ethical Considerations of Social Bot Detection
by: Ng, Lynnette Hui Xian, et al.
Published: (2026)
by: Ng, Lynnette Hui Xian, et al.
Published: (2026)
Deep Learning-Based Age Estimation and Gender Deep Learning-Based Age Estimation and Gender Classification for Targeted Advertisement
by: Zaman, Muhammad Imran, et al.
Published: (2025)
by: Zaman, Muhammad Imran, et al.
Published: (2025)
BotDetect: A Decentralized Federated Learning Framework for Detecting Financial Bots on the EVM Blockchains
by: Bendada, Ahmed Mounsf Rafik, et al.
Published: (2025)
by: Bendada, Ahmed Mounsf Rafik, et al.
Published: (2025)
As Good as We Are, We Can Always Get Better
by: Harvey, Carl A., II
Published: (2006)
by: Harvey, Carl A., II
Published: (2006)
Deep Networks Always Grok and Here is Why
by: Humayun, Ahmed Imtiaz, et al.
Published: (2024)
by: Humayun, Ahmed Imtiaz, et al.
Published: (2024)
Good Teaching Is Good Teaching. An Emerging Set of Guiding Principles and Practices for the Design and Development of Distance Education.
by: Ragan, Lawrence C.
Published: (1999)
by: Ragan, Lawrence C.
Published: (1999)
Assistive AI for Augmenting Human Decision-making
by: Gyöngyössy, Natabara Máté, et al.
Published: (2024)
by: Gyöngyössy, Natabara Máté, et al.
Published: (2024)
Know Your Limits: A Survey of Abstention in Large Language Models
by: Wen, Bingbing, et al.
Published: (2024)
by: Wen, Bingbing, et al.
Published: (2024)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Optimal confidence interval for the difference of proportions
by: Peer, Almog, et al.
Published: (2023)
by: Peer, Almog, et al.
Published: (2023)
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
by: Tsuchiya, Yuta, et al.
Published: (2026)
by: Tsuchiya, Yuta, et al.
Published: (2026)
Assessing novice programmers' perception of ChatGPT:performance, risk, decision-making, and intentions
by: Miranda, John Paul P., et al.
Published: (2025)
by: Miranda, John Paul P., et al.
Published: (2025)
Learning to Detect Baked Goods with Limited Supervision
by: Schmitt, Thomas H., et al.
Published: (2026)
by: Schmitt, Thomas H., et al.
Published: (2026)
A Crowdsourced Study of ChatBot Influence in Value-Driven Decision Making Scenarios
by: Wise, Anthony, et al.
Published: (2025)
by: Wise, Anthony, et al.
Published: (2025)
Are Large Random Graphs Always Safe to Hide?
by: Chakraborty, Sourav, et al.
Published: (2025)
by: Chakraborty, Sourav, et al.
Published: (2025)
Implementing MCMC: Multivariate estimation with confidence
by: Flegal, James M., et al.
Published: (2024)
by: Flegal, James M., et al.
Published: (2024)
Collaboration or Corporate Capture? Quantifying NLP's Reliance on Industry Artifacts and Contributions
by: Aitken, Will, et al.
Published: (2023)
by: Aitken, Will, et al.
Published: (2023)
Are Bigger Encoders Always Better in Vision Large Models?
by: Li, Bozhou, et al.
Published: (2024)
by: Li, Bozhou, et al.
Published: (2024)
AI Agents May Always Fall for Prompt Injections
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
As Confidence Aligns: Exploring the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making
by: Li, Jingshu, et al.
Published: (2025)
by: Li, Jingshu, et al.
Published: (2025)
The Road to the Closest Point is Paved by Good Neighbors
by: Har-Peled, Sariel, et al.
Published: (2025)
by: Har-Peled, Sariel, et al.
Published: (2025)
You are a Bot! -- Studying the Development of Bot Accusations on Twitter
by: Assenmacher, Dennis, et al.
Published: (2023)
by: Assenmacher, Dennis, et al.
Published: (2023)
When Are Two Lists Better than One?: Benefits and Harms in Joint Decision-making
by: Donahue, Kate, et al.
Published: (2023)
by: Donahue, Kate, et al.
Published: (2023)
What Do People Want to Know About Artificial Intelligence (AI)? The Importance of Answering End-User Questions to Explain Autonomous Vehicle (AV) Decisions
by: Molaei, Somayeh, et al.
Published: (2025)
by: Molaei, Somayeh, et al.
Published: (2025)
Hyperspectral Sensors and Autonomous Driving: Technologies, Limitations, and Opportunities
by: Shah, Imad Ali, et al.
Published: (2025)
by: Shah, Imad Ali, et al.
Published: (2025)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
by: Samragh, Mohammad, et al.
Published: (2025)
by: Samragh, Mohammad, et al.
Published: (2025)
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
Similar Items
-
When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance
by: Israelsen, Brett, et al.
Published: (2026) -
Expected Moral Shortfall for Ethical Competence in Decision-making Models
by: Aijaz, Aisha, et al.
Published: (2026) -
Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
by: Ahmed, Nisar, et al.
Published: (2025) -
Assessing confidence in frontier AI safety cases
by: Barrett, Stephen, et al.
Published: (2025) -
Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents
by: Xu, Rongwu, et al.
Published: (2025)