Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
Fuente:
arXiv
Saved in:
| Main Authors: | Seegmiller, Parker, Gatto, Joseph, Sharif, Omar, Basak, Madhusudan, Preum, Sarah Masud |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models for Document-Level Event-Argument Data Augmentation for Challenging Role Types
by: Gatto, Joseph, et al.
Published: (2024)
by: Gatto, Joseph, et al.
Published: (2024)
Depth $F_1$: Improving Evaluation of Cross-Domain Text Classification by Measuring Semantic Generalizability
by: Seegmiller, Parker, et al.
Published: (2024)
by: Seegmiller, Parker, et al.
Published: (2024)
REGen: A Reliable Evaluation Framework for Generative Event Argument Extraction
by: Sharif, Omar, et al.
Published: (2025)
by: Sharif, Omar, et al.
Published: (2025)
Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
by: Sharif, Omar, et al.
Published: (2024)
by: Sharif, Omar, et al.
Published: (2024)
Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance
by: Seegmiller, Parker, et al.
Published: (2026)
by: Seegmiller, Parker, et al.
Published: (2026)
In-Context Learning for Preserving Patient Privacy: A Framework for Synthesizing Realistic Patient Portal Messages
by: Gatto, Joseph, et al.
Published: (2024)
by: Gatto, Joseph, et al.
Published: (2024)
Follow-up Question Generation For Enhanced Patient-Provider Conversations
by: Gatto, Joseph, et al.
Published: (2025)
by: Gatto, Joseph, et al.
Published: (2025)
Scope of Large Language Models for Mining Emerging Opinions in Online Health Discourse
by: Gatto, Joseph, et al.
Published: (2024)
by: Gatto, Joseph, et al.
Published: (2024)
How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
by: Seegmiller, Parker, et al.
Published: (2026)
by: Seegmiller, Parker, et al.
Published: (2026)
Medical Triage as Pairwise Ranking: A Benchmark for Urgency in Patient Portal Messages
by: Gatto, Joseph, et al.
Published: (2026)
by: Gatto, Joseph, et al.
Published: (2026)
Socially Constructed Treatment Plans: Analyzing Online Peer Interactions to Understand How Patients Navigate Complex Medical Conditions
by: Basak, Madhusudan, et al.
Published: (2025)
by: Basak, Madhusudan, et al.
Published: (2025)
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
by: Cheng, Letian, et al.
Published: (2026)
by: Cheng, Letian, et al.
Published: (2026)
Deciphering Hate: Identifying Hateful Memes and Their Targets
by: Hossain, Eftekhar, et al.
Published: (2024)
by: Hossain, Eftekhar, et al.
Published: (2024)
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
by: Hossain, Eftekhar, et al.
Published: (2024)
by: Hossain, Eftekhar, et al.
Published: (2024)
Low-Perplexity LLM-Generated Sequences and Where To Find Them
by: Wuhrmann, Arthur, et al.
Published: (2025)
by: Wuhrmann, Arthur, et al.
Published: (2025)
Do LLMs and Humans Find the Same Questions Difficult? A Case Study on Japanese Quiz Answering
by: Sugiura, Naoya, et al.
Published: (2025)
by: Sugiura, Naoya, et al.
Published: (2025)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
by: Ankner, Zachary, et al.
Published: (2024)
by: Ankner, Zachary, et al.
Published: (2024)
A Thematic Framework for Analyzing Large-scale Self-reported Social Media Data on Opioid Use Disorder Treatment Using Buprenorphine Product
by: Basak, Madhusudan, et al.
Published: (2024)
by: Basak, Madhusudan, et al.
Published: (2024)
Can LLMs Solve My Grandma's Riddle? Evaluating Multilingual Large Language Models on Reasoning Traditional Bangla Tricky Riddles
by: Sayeedi, Nurul Labib, et al.
Published: (2025)
by: Sayeedi, Nurul Labib, et al.
Published: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
by: Liu, Lei, et al.
Published: (2025)
by: Liu, Lei, et al.
Published: (2025)
Explainable Fact-checking through Question Answering
by: Yang, Jing, et al.
Published: (2021)
by: Yang, Jing, et al.
Published: (2021)
Facts Do Care About Your Language: Assessing Answer Quality of Multilingual LLMs
by: Kansal, Yuval, et al.
Published: (2025)
by: Kansal, Yuval, et al.
Published: (2025)
AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering
by: Kuan, Chun-Yi, et al.
Published: (2026)
by: Kuan, Chun-Yi, et al.
Published: (2026)
Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions
by: Sajadi, Amirali, et al.
Published: (2025)
by: Sajadi, Amirali, et al.
Published: (2025)
Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering
by: Barbu, Eduard, et al.
Published: (2025)
by: Barbu, Eduard, et al.
Published: (2025)
Improving Pretraining Data Using Perplexity Correlations
by: Thrush, Tristan, et al.
Published: (2024)
by: Thrush, Tristan, et al.
Published: (2024)
CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval
by: Putta, Akshith Reddy, et al.
Published: (2026)
by: Putta, Akshith Reddy, et al.
Published: (2026)
HealthNLP_Retrievers at ArchEHR-QA 2026: Cascaded LLM Pipeline for Grounded Clinical Question Answering
by: Hosen, Md Biplob, et al.
Published: (2026)
by: Hosen, Md Biplob, et al.
Published: (2026)
Rethinking GSPO: The Perplexity-Entropy Equivalence
by: Liu, Chi
Published: (2025)
by: Liu, Chi
Published: (2025)
What is Wrong with Perplexity for Long-context Language Modeling?
by: Fang, Lizhe, et al.
Published: (2024)
by: Fang, Lizhe, et al.
Published: (2024)
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
by: Fernandes, Patrick, et al.
Published: (2025)
by: Fernandes, Patrick, et al.
Published: (2025)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
by: Shivagunde, Namrata, et al.
Published: (2026)
by: Shivagunde, Namrata, et al.
Published: (2026)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
by: Sawczyn, Albert, et al.
Published: (2025)
by: Sawczyn, Albert, et al.
Published: (2025)
TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
by: Iranmanesh, Reihaneh, et al.
Published: (2026)
by: Iranmanesh, Reihaneh, et al.
Published: (2026)
Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning
by: Gado, Elena Grazia, et al.
Published: (2024)
by: Gado, Elena Grazia, et al.
Published: (2024)
Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
by: Seegmiller, Parker, et al.
Published: (2025)
by: Seegmiller, Parker, et al.
Published: (2025)
The Moral Foundations Reddit Corpus
by: Trager, Jackson, et al.
Published: (2022)
by: Trager, Jackson, et al.
Published: (2022)
Similar Items
-
Large Language Models for Document-Level Event-Argument Data Augmentation for Challenging Role Types
by: Gatto, Joseph, et al.
Published: (2024) -
Depth $F_1$: Improving Evaluation of Cross-Domain Text Classification by Measuring Semantic Generalizability
by: Seegmiller, Parker, et al.
Published: (2024) -
REGen: A Reliable Evaluation Framework for Generative Event Argument Extraction
by: Sharif, Omar, et al.
Published: (2025) -
Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
by: Sharif, Omar, et al.
Published: (2024) -
Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance
by: Seegmiller, Parker, et al.
Published: (2026)