The Pitfalls of Defining Hallucination
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | van Deemter, Kees |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
My Life in Artificial Intelligence: People, anecdotes, and some lessons learnt
von: van Deemter, Kees
Veröffentlicht: (2025)
von: van Deemter, Kees
Veröffentlicht: (2025)
Intrinsic Task-based Evaluation for Referring Expression Generation
von: Chen, Guanyi, et al.
Veröffentlicht: (2024)
von: Chen, Guanyi, et al.
Veröffentlicht: (2024)
Computational Modelling of Plurality and Definiteness in Chinese Noun Phrases
von: Liu, Yuqi, et al.
Veröffentlicht: (2024)
von: Liu, Yuqi, et al.
Veröffentlicht: (2024)
Reference-free Evaluation Metrics for Text Generation: A Survey
von: Ito, Takumi, et al.
Veröffentlicht: (2025)
von: Ito, Takumi, et al.
Veröffentlicht: (2025)
Textual Summarisation of Large Sets: Towards a General Approach
von: Kuptavanich, Kittipitch, et al.
Veröffentlicht: (2024)
von: Kuptavanich, Kittipitch, et al.
Veröffentlicht: (2024)
Pitfalls and Outlooks in Using COMET
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
Bias in News Summarization: Measures, Pitfalls and Corpora
von: Steen, Julius, et al.
Veröffentlicht: (2023)
von: Steen, Julius, et al.
Veröffentlicht: (2023)
Pitfalls of Evaluating Language Models with Open Benchmarks
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2025)
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2025)
Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals
von: van Dijk, Gijs
Veröffentlicht: (2026)
von: van Dijk, Gijs
Veröffentlicht: (2026)
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
von: Chan, Jason, et al.
Veröffentlicht: (2025)
von: Chan, Jason, et al.
Veröffentlicht: (2025)
Pitfalls of Conversational LLMs on News Debiasing
von: Schlicht, Ipek Baris, et al.
Veröffentlicht: (2024)
von: Schlicht, Ipek Baris, et al.
Veröffentlicht: (2024)
Analysing the Safety Pitfalls of Steering Vectors
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
Mind the Gap: Pitfalls of LLM Alignment with Asian Public Opinion
von: Shankar, Hari, et al.
Veröffentlicht: (2026)
von: Shankar, Hari, et al.
Veröffentlicht: (2026)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
von: Anzenberg, Eitan, et al.
Veröffentlicht: (2025)
von: Anzenberg, Eitan, et al.
Veröffentlicht: (2025)
Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP
von: Khurana, Urja, et al.
Veröffentlicht: (2026)
von: Khurana, Urja, et al.
Veröffentlicht: (2026)
Mitigating Hallucination on Hallucination in RAG via Ensemble Voting
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning
von: Guo, Jiaxing, et al.
Veröffentlicht: (2025)
von: Guo, Jiaxing, et al.
Veröffentlicht: (2025)
Pitfalls of Scale: Investigating the Inverse Task of Redefinition in Large Language Models
von: Stringli, Elena, et al.
Veröffentlicht: (2025)
von: Stringli, Elena, et al.
Veröffentlicht: (2025)
Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
von: Guan, Jiannan, et al.
Veröffentlicht: (2025)
von: Guan, Jiannan, et al.
Veröffentlicht: (2025)
Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training
von: Chen, Hang, et al.
Veröffentlicht: (2026)
von: Chen, Hang, et al.
Veröffentlicht: (2026)
Large Language Models for Stemming: Promises, Pitfalls and Failures
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
von: Godbole, Ameya, et al.
Veröffentlicht: (2025)
von: Godbole, Ameya, et al.
Veröffentlicht: (2025)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
von: Ousidhoum, Nedjma, et al.
Veröffentlicht: (2024)
von: Ousidhoum, Nedjma, et al.
Veröffentlicht: (2024)
What's Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs
von: Wegmann, Anna, et al.
Veröffentlicht: (2024)
von: Wegmann, Anna, et al.
Veröffentlicht: (2024)
Hallucination Detection and Hallucination Mitigation: An Investigation
von: Luo, Junliang, et al.
Veröffentlicht: (2024)
von: Luo, Junliang, et al.
Veröffentlicht: (2024)
Editing the Mind of Giants: An In-Depth Exploration of Pitfalls of Knowledge Editing in Large Language Models
von: Hsueh, Cheng-Hsun, et al.
Veröffentlicht: (2024)
von: Hsueh, Cheng-Hsun, et al.
Veröffentlicht: (2024)
A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024)
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024)
PPA-Plan: Proactive Pitfall Avoidance for Reliable Planning in Long-Context LLM Reasoning
von: Kim, Byeongjin, et al.
Veröffentlicht: (2026)
von: Kim, Byeongjin, et al.
Veröffentlicht: (2026)
Is a Document Educational or Just Wikipedia-Style? -- Pitfalls of Classifier-Based Quality Filtering
von: Klimaszewski, Mateusz, et al.
Veröffentlicht: (2026)
von: Klimaszewski, Mateusz, et al.
Veröffentlicht: (2026)
Detecting Hallucination and Coverage Errors in Retrieval Augmented Generation for Controversial Topics
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
Removal of Hallucination on Hallucination: Debate-Augmented RAG
von: Hu, Wentao, et al.
Veröffentlicht: (2025)
von: Hu, Wentao, et al.
Veröffentlicht: (2025)
Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning
von: Liao, Yiming, et al.
Veröffentlicht: (2026)
von: Liao, Yiming, et al.
Veröffentlicht: (2026)
Investigating the Role of Prompting and External Tools in Hallucination Rates of Large Language Models
von: Barkley, Liam, et al.
Veröffentlicht: (2024)
von: Barkley, Liam, et al.
Veröffentlicht: (2024)
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
von: Horych, Tomas, et al.
Veröffentlicht: (2024)
von: Horych, Tomas, et al.
Veröffentlicht: (2024)
BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2023)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2023)
Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
von: Wang, Ante, et al.
Veröffentlicht: (2025)
von: Wang, Ante, et al.
Veröffentlicht: (2025)
On the Hallucination in Simultaneous Machine Translation
von: Zhong, Meizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Meizhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
My Life in Artificial Intelligence: People, anecdotes, and some lessons learnt
von: van Deemter, Kees
Veröffentlicht: (2025) -
Intrinsic Task-based Evaluation for Referring Expression Generation
von: Chen, Guanyi, et al.
Veröffentlicht: (2024) -
Computational Modelling of Plurality and Definiteness in Chinese Noun Phrases
von: Liu, Yuqi, et al.
Veröffentlicht: (2024) -
Reference-free Evaluation Metrics for Text Generation: A Survey
von: Ito, Takumi, et al.
Veröffentlicht: (2025) -
Textual Summarisation of Large Sets: Towards a General Approach
von: Kuptavanich, Kittipitch, et al.
Veröffentlicht: (2024)