What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Malaviya, Chaitanya, Lee, Subin, Roth, Dan, Yatskar, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ExpertQA: Expert-Curated Questions and Attributed Answers
by: Malaviya, Chaitanya, et al.
Published: (2023)
by: Malaviya, Chaitanya, et al.
Published: (2023)
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
by: Bharadwaj, Anirudh, et al.
Published: (2025)
by: Bharadwaj, Anirudh, et al.
Published: (2025)
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
by: Malaviya, Chaitanya, et al.
Published: (2024)
by: Malaviya, Chaitanya, et al.
Published: (2024)
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
by: Yifei, Li S., et al.
Published: (2025)
by: Yifei, Li S., et al.
Published: (2025)
A Simple Joint Model for Improved Contextual Neural Lemmatization
by: Malaviya, Chaitanya, et al.
Published: (2019)
by: Malaviya, Chaitanya, et al.
Published: (2019)
On Reference (In-)Determinacy in Natural Language Inference
by: Chen, Sihao, et al.
Published: (2025)
by: Chen, Sihao, et al.
Published: (2025)
DOLOMITES: Domain-Specific Long-Form Methodical Tasks
by: Malaviya, Chaitanya, et al.
Published: (2024)
by: Malaviya, Chaitanya, et al.
Published: (2024)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
by: Fu, Xingyu, et al.
Published: (2023)
by: Fu, Xingyu, et al.
Published: (2023)
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification
by: Qu, Renyi, et al.
Published: (2024)
by: Qu, Renyi, et al.
Published: (2024)
Fakes of Varying Shades: How Warning Affects Human Perception and Engagement Regarding LLM Hallucinations
by: Nahar, Mahjabin, et al.
Published: (2024)
by: Nahar, Mahjabin, et al.
Published: (2024)
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
by: Yoran, Ori, et al.
Published: (2024)
by: Yoran, Ori, et al.
Published: (2024)
How Does the Disclosure of AI Assistance Affect the Perceptions of Writing?
by: Li, Zhuoyan, et al.
Published: (2024)
by: Li, Zhuoyan, et al.
Published: (2024)
What to Format and How: A Benchmark and Workflow Approach for Document Formatting
by: Rao, Shihao, et al.
Published: (2026)
by: Rao, Shihao, et al.
Published: (2026)
Comparing How a Chatbot References User Utterances from Previous Chatting Sessions: An Investigation of Users' Privacy Concerns and Perceptions
by: Cox, Samuel Rhys, et al.
Published: (2023)
by: Cox, Samuel Rhys, et al.
Published: (2023)
From Lists to Emojis: How Format Bias Affects Model Alignment
by: Zhang, Xuanchang, et al.
Published: (2024)
by: Zhang, Xuanchang, et al.
Published: (2024)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
by: Hong, Pingjun, et al.
Published: (2025)
by: Hong, Pingjun, et al.
Published: (2025)
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
by: Yang, Yahan, et al.
Published: (2024)
by: Yang, Yahan, et al.
Published: (2024)
Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
Calibrating Large Language Models with Sample Consistency
by: Lyu, Qing, et al.
Published: (2024)
by: Lyu, Qing, et al.
Published: (2024)
Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
by: Meincke, Lennart, et al.
Published: (2025)
by: Meincke, Lennart, et al.
Published: (2025)
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
by: Trinley, Katharina, et al.
Published: (2025)
by: Trinley, Katharina, et al.
Published: (2025)
User eXperience Perception Insights Dataset (UXPID): Synthetic User Feedback from Public Industrial Forums
by: Kulyabin, Mikhail, et al.
Published: (2025)
by: Kulyabin, Mikhail, et al.
Published: (2025)
HumanOmni-Speaker: Identifying Who said What and When
by: Bai, Detao, et al.
Published: (2026)
by: Bai, Detao, et al.
Published: (2026)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck
by: Ludan, Josh Magnus, et al.
Published: (2023)
by: Ludan, Josh Magnus, et al.
Published: (2023)
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
by: Shi, Taiwei, et al.
Published: (2024)
by: Shi, Taiwei, et al.
Published: (2024)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
by: Yang, Yahan, et al.
Published: (2025)
by: Yang, Yahan, et al.
Published: (2025)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
Reasoning is about giving reasons
by: Shah, Krunal, et al.
Published: (2025)
by: Shah, Krunal, et al.
Published: (2025)
Conflicts in Texts: Data, Implications and Challenges
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
What talking you?: Translating Code-Mixed Messaging Texts to English
by: Ng, Lynnette Hui Xian, et al.
Published: (2024)
by: Ng, Lynnette Hui Xian, et al.
Published: (2024)
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
by: Movva, Rajiv, et al.
Published: (2025)
by: Movva, Rajiv, et al.
Published: (2025)
Compact Example-Based Explanations for Language Models
by: Schoenegger, Loris, et al.
Published: (2026)
by: Schoenegger, Loris, et al.
Published: (2026)
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
by: Lee, Yunseung, et al.
Published: (2026)
by: Lee, Yunseung, et al.
Published: (2026)
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
by: Schoenegger, Loris, et al.
Published: (2024)
by: Schoenegger, Loris, et al.
Published: (2024)
What Affects the Effective Depth of Large Language Models?
by: Hu, Yi, et al.
Published: (2025)
by: Hu, Yi, et al.
Published: (2025)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
by: Elangovan, Aparna, et al.
Published: (2024)
by: Elangovan, Aparna, et al.
Published: (2024)
Similar Items
-
ExpertQA: Expert-Curated Questions and Attributed Answers
by: Malaviya, Chaitanya, et al.
Published: (2023) -
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
by: Bharadwaj, Anirudh, et al.
Published: (2025) -
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
by: Malaviya, Chaitanya, et al.
Published: (2024) -
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
by: Yifei, Li S., et al.
Published: (2025) -
A Simple Joint Model for Improved Contextual Neural Lemmatization
by: Malaviya, Chaitanya, et al.
Published: (2019)