AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
Fuente:
arXiv
Salvato in:
| Autori principali: | Gor, Maharshi, Sung, Yoo Yeon, Hou, Yu, Fleisig, Eve, Ying, Irene, Zhou, Tianyi, Boyd-Graber, Jordan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
di: Gor, Maharshi, et al.
Pubblicazione: (2024)
di: Gor, Maharshi, et al.
Pubblicazione: (2024)
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
di: Srikanth, Neha, et al.
Pubblicazione: (2026)
di: Srikanth, Neha, et al.
Pubblicazione: (2026)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
di: Li, Zongxia, et al.
Pubblicazione: (2024)
di: Li, Zongxia, et al.
Pubblicazione: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
di: Srikanth, Neha, et al.
Pubblicazione: (2023)
di: Srikanth, Neha, et al.
Pubblicazione: (2023)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Mapping Social Choice Theory to RLHF
di: Dai, Jessica, et al.
Pubblicazione: (2024)
di: Dai, Jessica, et al.
Pubblicazione: (2024)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
di: Calvo-Bartolomé, Lorena, et al.
Pubblicazione: (2025)
di: Calvo-Bartolomé, Lorena, et al.
Pubblicazione: (2025)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
di: Fleisig, Eve, et al.
Pubblicazione: (2023)
di: Fleisig, Eve, et al.
Pubblicazione: (2023)
Labeled Interactive Topic Models
di: Seelman, Kyle, et al.
Pubblicazione: (2023)
di: Seelman, Kyle, et al.
Pubblicazione: (2023)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
di: Saphra, Naomi, et al.
Pubblicazione: (2023)
di: Saphra, Naomi, et al.
Pubblicazione: (2023)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
di: Jaggi, Harbani, et al.
Pubblicazione: (2024)
di: Jaggi, Harbani, et al.
Pubblicazione: (2024)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
Compositional Consistency-Guided Decoding for Three-Way Logical Question Answering
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
di: Shu, Matthew, et al.
Pubblicazione: (2024)
di: Shu, Matthew, et al.
Pubblicazione: (2024)
Should I Trust You? Detecting Deception in Negotiations using Counterfactual RL
di: Wongkamjan, Wichayaporn, et al.
Pubblicazione: (2025)
di: Wongkamjan, Wichayaporn, et al.
Pubblicazione: (2025)
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
di: Verma, Vivek, et al.
Pubblicazione: (2023)
di: Verma, Vivek, et al.
Pubblicazione: (2023)
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
di: Fleisig, Eve, et al.
Pubblicazione: (2025)
di: Fleisig, Eve, et al.
Pubblicazione: (2025)
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
di: Hoyle, Alexander, et al.
Pubblicazione: (2025)
di: Hoyle, Alexander, et al.
Pubblicazione: (2025)
CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
Standard Language Ideology in AI-Generated Language
di: Smith, Genevieve, et al.
Pubblicazione: (2024)
di: Smith, Genevieve, et al.
Pubblicazione: (2024)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
DriveLM: Driving with Graph Visual Question Answering
di: Sima, Chonghao, et al.
Pubblicazione: (2023)
di: Sima, Chonghao, et al.
Pubblicazione: (2023)
TrustUQA: A Trustful Framework for Unified Structured Data Question Answering
di: Zhang, Wen, et al.
Pubblicazione: (2024)
di: Zhang, Wen, et al.
Pubblicazione: (2024)
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
di: Tao, Mingzhe, et al.
Pubblicazione: (2026)
di: Tao, Mingzhe, et al.
Pubblicazione: (2026)
More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play
di: Wongkamjan, Wichayaporn, et al.
Pubblicazione: (2024)
di: Wongkamjan, Wichayaporn, et al.
Pubblicazione: (2024)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025)
Structured List-Grounded Question Answering
di: Sung, Mujeen, et al.
Pubblicazione: (2024)
di: Sung, Mujeen, et al.
Pubblicazione: (2024)
Cooperation and Control in Delegation Games
di: Sourbut, Oliver, et al.
Pubblicazione: (2024)
di: Sourbut, Oliver, et al.
Pubblicazione: (2024)
MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering
di: Ning, Yingpeng, et al.
Pubblicazione: (2025)
di: Ning, Yingpeng, et al.
Pubblicazione: (2025)
Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering Benchmark
di: Xu, Zhikun, et al.
Pubblicazione: (2024)
di: Xu, Zhikun, et al.
Pubblicazione: (2024)
GROUNDEDKG-RAG: Grounded Knowledge Graph Index for Long-document Question Answering
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
Bias Behind the Wheel: Fairness Testing of Autonomous Driving Systems
di: Li, Xinyue, et al.
Pubblicazione: (2023)
di: Li, Xinyue, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024) -
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
di: Gor, Maharshi, et al.
Pubblicazione: (2024) -
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025) -
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024) -
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
di: Srikanth, Neha, et al.
Pubblicazione: (2026)