WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Wenting, Goyal, Tanya, Chiu, Yu Ying, Jiang, Liwei, Newman, Benjamin, Ravichander, Abhilasha, Chandu, Khyathi, Bras, Ronan Le, Cardie, Claire, Deng, Yuntian, Choi, Yejin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
WildChat: 1M ChatGPT Interaction Logs in the Wild
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023)
von: Yin, Da, et al.
Veröffentlicht: (2023)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
von: Newman, Benjamin, et al.
Veröffentlicht: (2025)
von: Newman, Benjamin, et al.
Veröffentlicht: (2025)
MASH: Modeling Abstention via Selective Help-Seeking
von: Gul, Mustafa Omer, et al.
Veröffentlicht: (2025)
von: Gul, Mustafa Omer, et al.
Veröffentlicht: (2025)
MacGyver: Are Large Language Models Creative Problem Solvers?
von: Tian, Yufei, et al.
Veröffentlicht: (2023)
von: Tian, Yufei, et al.
Veröffentlicht: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
von: Brahman, Faeze, et al.
Veröffentlicht: (2024)
von: Brahman, Faeze, et al.
Veröffentlicht: (2024)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
von: Mendelsohn, Julia, et al.
Veröffentlicht: (2023)
von: Mendelsohn, Julia, et al.
Veröffentlicht: (2023)
What Has Been Lost with Synthetic Evaluation?
von: Gill, Alexander, et al.
Veröffentlicht: (2025)
von: Gill, Alexander, et al.
Veröffentlicht: (2025)
Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions
von: Su, Jinyan, et al.
Veröffentlicht: (2026)
von: Su, Jinyan, et al.
Veröffentlicht: (2026)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
Revisiting the Past: Data Unlearning with Model State History
von: Rezaei, Keivan, et al.
Veröffentlicht: (2025)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2025)
RefreshKV: Updating Small KV Cache During Long-form Generation
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
I Could've Asked That: Reformulating Unanswerable Questions
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
von: Chandu, Khyathi Raghavi, et al.
Veröffentlicht: (2024)
von: Chandu, Khyathi Raghavi, et al.
Veröffentlicht: (2024)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
Commit0: Library Generation from Scratch
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
von: Han, Seungju, et al.
Veröffentlicht: (2024)
von: Han, Seungju, et al.
Veröffentlicht: (2024)
Reasoning Court: Combining Reasoning, Action, and Judgment for Multi-Hop Reasoning
von: Wu, Jingtian, et al.
Veröffentlicht: (2025)
von: Wu, Jingtian, et al.
Veröffentlicht: (2025)
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
von: Park, Hansol, et al.
Veröffentlicht: (2025)
von: Park, Hansol, et al.
Veröffentlicht: (2025)
Challenges in Trustworthy Human Evaluation of Chatbots
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
A Call for Clarity in Beam Search: How It Works and When It Stops
von: Kasai, Jungo, et al.
Veröffentlicht: (2022)
von: Kasai, Jungo, et al.
Veröffentlicht: (2022)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
von: Hallinan, Skyler, et al.
Veröffentlicht: (2025)
von: Hallinan, Skyler, et al.
Veröffentlicht: (2025)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation
von: Xu, Fan, et al.
Veröffentlicht: (2025)
von: Xu, Fan, et al.
Veröffentlicht: (2025)
Are Triggers Needed for Document-Level Event Extraction?
von: Shaar, Shaden, et al.
Veröffentlicht: (2024)
von: Shaar, Shaden, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024) -
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024) -
WildChat: 1M ChatGPT Interaction Logs in the Wild
von: Zhao, Wenting, et al.
Veröffentlicht: (2024) -
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
von: Deng, Yuntian, et al.
Veröffentlicht: (2024) -
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)