Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
Fuente:
arXiv
Saved in:
| Main Authors: | Gor, Maharshi, Daumé III, Hal, Zhou, Tianyi, Boyd-Graber, Jordan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026)
by: Gor, Maharshi, et al.
Published: (2026)
A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
by: Goyal, Navita, et al.
Published: (2026)
by: Goyal, Navita, et al.
Published: (2026)
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections
by: Zhao, Lingjun, et al.
Published: (2024)
by: Zhao, Lingjun, et al.
Published: (2024)
"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
by: Hou, Yu, et al.
Published: (2025)
by: Hou, Yu, et al.
Published: (2025)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025)
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
by: Sung, Yoo Yeon, et al.
Published: (2024)
by: Sung, Yoo Yeon, et al.
Published: (2024)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026)
by: Zhao, Lingjun, et al.
Published: (2026)
When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing
by: Baumler, Connor, et al.
Published: (2024)
by: Baumler, Connor, et al.
Published: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
by: Srikanth, Neha, et al.
Published: (2026)
by: Srikanth, Neha, et al.
Published: (2026)
Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind Users
by: Zeraati, Farnaz Zamiri, et al.
Published: (2026)
by: Zeraati, Farnaz Zamiri, et al.
Published: (2026)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features
by: Goyal, Navita, et al.
Published: (2023)
by: Goyal, Navita, et al.
Published: (2023)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026)
by: Nghiem, Huy, et al.
Published: (2026)
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
by: Mondal, Ishani, et al.
Published: (2025)
by: Mondal, Ishani, et al.
Published: (2025)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
by: Si, Chenglei, et al.
Published: (2023)
by: Si, Chenglei, et al.
Published: (2023)
Can Hallucination Correction Improve Video-Language Alignment?
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
by: Cao, Yang Trista, et al.
Published: (2023)
by: Cao, Yang Trista, et al.
Published: (2023)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
by: Gu, Feng, et al.
Published: (2025)
by: Gu, Feng, et al.
Published: (2025)
GROUNDEDKG-RAG: Grounded Knowledge Graph Index for Long-document Question Answering
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
Augmenting Question Answering with A Hybrid RAG Approach
by: Yang, Tianyi, et al.
Published: (2026)
by: Yang, Tianyi, et al.
Published: (2026)
Exploring Human-AI Complementarity in CPS Diagnosis Using Unimodal and Multimodal BERT Models
by: Wong, Kester, et al.
Published: (2025)
by: Wong, Kester, et al.
Published: (2025)
Do not think about pink elephant!
by: Hwang, Kyomin, et al.
Published: (2024)
by: Hwang, Kyomin, et al.
Published: (2024)
Seamful XAI: Operationalizing Seamful Design in Explainable AI
by: Ehsan, Upol, et al.
Published: (2022)
by: Ehsan, Upol, et al.
Published: (2022)
HPE:Answering Complex Questions over Text by Hybrid Question Parsing and Execution
by: Liu, Ye, et al.
Published: (2023)
by: Liu, Ye, et al.
Published: (2023)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
by: Keluskar, Aryan, et al.
Published: (2024)
by: Keluskar, Aryan, et al.
Published: (2024)
Uncertainty Estimation of Large Language Models in Medical Question Answering
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
by: Hao, Yuexing, et al.
Published: (2025)
by: Hao, Yuexing, et al.
Published: (2025)
Compositional Consistency-Guided Decoding for Three-Way Logical Question Answering
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
To Reason or Not to: Selective Chain-of-Thought in Medical Question Answering
by: Zhan, Zaifu, et al.
Published: (2026)
by: Zhan, Zaifu, et al.
Published: (2026)
Efficient Medical Question Answering with Knowledge-Augmented Question Generation
by: Khlaut, Julien, et al.
Published: (2024)
by: Khlaut, Julien, et al.
Published: (2024)
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
by: Balamurali, Sai Shridhar, et al.
Published: (2025)
by: Balamurali, Sai Shridhar, et al.
Published: (2025)
Similar Items
-
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026) -
A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
by: Zhao, Lingjun, et al.
Published: (2025) -
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
by: Goyal, Navita, et al.
Published: (2026) -
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024) -
Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections
by: Zhao, Lingjun, et al.
Published: (2024)