Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hagar, Nick, Agustianto, Wilma, Diakopoulos, Nicholas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
Using Generative Agents to Create Tip Sheets for Investigative Data Reporting
von: Veerbeek, Joris, et al.
Veröffentlicht: (2024)
von: Veerbeek, Joris, et al.
Veröffentlicht: (2024)
Towards Leveraging News Media to Support Impact Assessment of AI Technologies
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
Simulating Policy Impacts: Developing a Generative Scenario Writing Method to Evaluate the Perceived Effects of Regulation
von: Barnett, Julia, et al.
Veröffentlicht: (2024)
von: Barnett, Julia, et al.
Veröffentlicht: (2024)
Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners
von: Li, Charlotte, et al.
Veröffentlicht: (2025)
von: Li, Charlotte, et al.
Veröffentlicht: (2025)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
What's Wrong? Refining Meeting Summaries with LLM Feedback
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
von: Yao, Jihan, et al.
Veröffentlicht: (2024)
von: Yao, Jihan, et al.
Veröffentlicht: (2024)
Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
von: Chhikara, Prateek
Veröffentlicht: (2025)
von: Chhikara, Prateek
Veröffentlicht: (2025)
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
von: He, Yanjie
Veröffentlicht: (2026)
von: He, Yanjie
Veröffentlicht: (2026)
Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2026)
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2026)
LLM-Assisted News Discovery in High-Volume Information Streams: A Case Study
von: Hagar, Nick, et al.
Veröffentlicht: (2025)
von: Hagar, Nick, et al.
Veröffentlicht: (2025)
LLM Augmentations to support Analytical Reasoning over Multiple Documents
von: Yousuf, Raquib Bin, et al.
Veröffentlicht: (2024)
von: Yousuf, Raquib Bin, et al.
Veröffentlicht: (2024)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
von: Curran, Damian, et al.
Veröffentlicht: (2025)
von: Curran, Damian, et al.
Veröffentlicht: (2025)
An LLM Maturity Model for Reliable and Transparent Text-to-Query
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
von: Fleisig, Eve, et al.
Veröffentlicht: (2023)
von: Fleisig, Eve, et al.
Veröffentlicht: (2023)
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying
von: Castagna, Federico, et al.
Veröffentlicht: (2024)
von: Castagna, Federico, et al.
Veröffentlicht: (2024)
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
On-Premise AI for the Newsroom: Evaluating Small Language Models for Investigative Document Search
von: Hagar, Nick, et al.
Veröffentlicht: (2025)
von: Hagar, Nick, et al.
Veröffentlicht: (2025)
Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2025)
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2025)
Easy Problems That LLMs Get Wrong
von: Williams, Sean, et al.
Veröffentlicht: (2024)
von: Williams, Sean, et al.
Veröffentlicht: (2024)
How Transformers Reject Wrong Answers: Rotational Dynamics of Factual Constraint Processing
von: Marín, Javier
Veröffentlicht: (2026)
von: Marín, Javier
Veröffentlicht: (2026)
Query-Based Adversarial Prompt Generation
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
QueryNER: Segmentation of E-commerce Queries
von: Palen-Michel, Chester, et al.
Veröffentlicht: (2024)
von: Palen-Michel, Chester, et al.
Veröffentlicht: (2024)
Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict
von: Chen, Yihang, et al.
Veröffentlicht: (2026)
von: Chen, Yihang, et al.
Veröffentlicht: (2026)
Multi-Lingual Cyber Threat Detection in Tweets/X Using ML, DL, and LLM: A Comparative Analysis
von: Murad, Saydul Akbar, et al.
Veröffentlicht: (2025)
von: Murad, Saydul Akbar, et al.
Veröffentlicht: (2025)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
von: Deng, Wei
Veröffentlicht: (2026)
von: Deng, Wei
Veröffentlicht: (2026)
Perplexity Cannot Always Tell Right from Wrong
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning
von: Singh, Vikash, et al.
Veröffentlicht: (2026)
von: Singh, Vikash, et al.
Veröffentlicht: (2026)
No Query, No Access
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
CompassLLM: A Multi-Agent Approach toward Geo-Spatial Reasoning for Popular Path Query
von: Ananto, Md. Nazmul Islam, et al.
Veröffentlicht: (2025)
von: Ananto, Md. Nazmul Islam, et al.
Veröffentlicht: (2025)
The Effect of Document Summarization on LLM-Based Relevance Judgments
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2025)
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2025)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
von: Watson, William, et al.
Veröffentlicht: (2026)
von: Watson, William, et al.
Veröffentlicht: (2026)
Humans Perceive Wrong Narratives from AI Reasoning Texts
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents
von: Schiepanski, Thassilo M., et al.
Veröffentlicht: (2025)
von: Schiepanski, Thassilo M., et al.
Veröffentlicht: (2025)
Improved LLM Agents for Financial Document Question Answering
von: Tan, Nelvin, et al.
Veröffentlicht: (2025)
von: Tan, Nelvin, et al.
Veröffentlicht: (2025)
Generating Query-Focused Summarization Datasets from Query-Free Summarization Datasets
von: Chali, Yllias, et al.
Veröffentlicht: (2026)
von: Chali, Yllias, et al.
Veröffentlicht: (2026)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
von: Ding, Wenxuan, et al.
Veröffentlicht: (2026)
von: Ding, Wenxuan, et al.
Veröffentlicht: (2026)
Ask-EDA: A Design Assistant Empowered by LLM, Hybrid RAG and Abbreviation De-hallucination
von: Shi, Luyao, et al.
Veröffentlicht: (2024)
von: Shi, Luyao, et al.
Veröffentlicht: (2024)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
von: Ding, Dujian, et al.
Veröffentlicht: (2024)
von: Ding, Dujian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024) -
Using Generative Agents to Create Tip Sheets for Investigative Data Reporting
von: Veerbeek, Joris, et al.
Veröffentlicht: (2024) -
Towards Leveraging News Media to Support Impact Assessment of AI Technologies
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024) -
Simulating Policy Impacts: Developing a Generative Scenario Writing Method to Evaluate the Perceived Effects of Regulation
von: Barnett, Julia, et al.
Veröffentlicht: (2024) -
Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners
von: Li, Charlotte, et al.
Veröffentlicht: (2025)