How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Roig, JV |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
par: Roig, JV
Publié: (2025)
par: Roig, JV
Publié: (2025)
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
par: Islam, Saad Obaid ul, et autres
Publié: (2025)
par: Islam, Saad Obaid ul, et autres
Publié: (2025)
Scalable and Reliable Evaluation of AI Knowledge Retrieval Systems: RIKER and the Coherent Simulated Universe
par: Roig, JV
Publié: (2025)
par: Roig, JV
Publié: (2025)
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
par: Zhou, Yang, et autres
Publié: (2025)
par: Zhou, Yang, et autres
Publié: (2025)
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
par: Vakilian, Vala, et autres
Publié: (2025)
par: Vakilian, Vala, et autres
Publié: (2025)
Do Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths
par: Nair, Inderjeet, et autres
Publié: (2025)
par: Nair, Inderjeet, et autres
Publié: (2025)
How Do LLMs Perform Two-Hop Reasoning in Context?
par: Guo, Tianyu, et autres
Publié: (2025)
par: Guo, Tianyu, et autres
Publié: (2025)
Do Large Language Models Know How Much They Know?
par: Prato, Gabriele, et autres
Publié: (2025)
par: Prato, Gabriele, et autres
Publié: (2025)
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
par: Silva, Enzo S. N., et autres
Publié: (2026)
par: Silva, Enzo S. N., et autres
Publié: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
par: Wu, Wei, et autres
Publié: (2024)
par: Wu, Wei, et autres
Publié: (2024)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
par: Wang, Zengzhi, et autres
Publié: (2023)
par: Wang, Zengzhi, et autres
Publié: (2023)
Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
par: Sato, Makoto
Publié: (2025)
par: Sato, Makoto
Publié: (2025)
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning
par: Chen, Haoyang, et autres
Publié: (2026)
par: Chen, Haoyang, et autres
Publié: (2026)
Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs
par: Vaddi, Snehit, et autres
Publié: (2026)
par: Vaddi, Snehit, et autres
Publié: (2026)
How Do LLMs Use Their Depth?
par: Gupta, Akshat, et autres
Publié: (2025)
par: Gupta, Akshat, et autres
Publié: (2025)
BertaQA: How Much Do Language Models Know About Local Culture?
par: Etxaniz, Julen, et autres
Publié: (2024)
par: Etxaniz, Julen, et autres
Publié: (2024)
How Much Can RAG Help the Reasoning of LLM?
par: Liu, Jingyu, et autres
Publié: (2024)
par: Liu, Jingyu, et autres
Publié: (2024)
How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent
par: Jung, Sungwoo, et autres
Publié: (2026)
par: Jung, Sungwoo, et autres
Publié: (2026)
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
par: Moslonka, Charles, et autres
Publié: (2025)
par: Moslonka, Charles, et autres
Publié: (2025)
A Geometric Taxonomy of Hallucinations in LLMs
par: Marín, Javier
Publié: (2026)
par: Marín, Javier
Publié: (2026)
Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations
par: Roig, JV
Publié: (2025)
par: Roig, JV
Publié: (2025)
How Much Data is Enough Data? Fine-Tuning Large Language Models for In-House Translation: Performance Evaluation Across Multiple Dataset Sizes
par: Vieira, Inacio, et autres
Publié: (2024)
par: Vieira, Inacio, et autres
Publié: (2024)
AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
par: Huang, Haoyu, et autres
Publié: (2025)
par: Huang, Haoyu, et autres
Publié: (2025)
Tokenization Constraints in LLMs: A Study of Symbolic and Arithmetic Reasoning Limits
par: Zhang, Xiang, et autres
Publié: (2025)
par: Zhang, Xiang, et autres
Publié: (2025)
How Well Do LLMs Understand Tunisian Arabic?
par: Mahdi, Mohamed
Publié: (2025)
par: Mahdi, Mohamed
Publié: (2025)
How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
par: Lee, Ayeong, et autres
Publié: (2025)
par: Lee, Ayeong, et autres
Publié: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
par: Singh, Janvijay, et autres
Publié: (2026)
par: Singh, Janvijay, et autres
Publié: (2026)
How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control
par: Li, Kunhang, et autres
Publié: (2025)
par: Li, Kunhang, et autres
Publié: (2025)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
par: Mahabadi, Rabeeh Karimi, et autres
Publié: (2025)
par: Mahabadi, Rabeeh Karimi, et autres
Publié: (2025)
Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding
par: Zhu, Yifan, et autres
Publié: (2026)
par: Zhu, Yifan, et autres
Publié: (2026)
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Literal Extraction, Logical Inference, and Hallucination Risks in Long-Context LLMs
par: Ebrahimzadeh, Amirali, et autres
Publié: (2026)
par: Ebrahimzadeh, Amirali, et autres
Publié: (2026)
Towards Long Context Hallucination Detection
par: Liu, Siyi, et autres
Publié: (2025)
par: Liu, Siyi, et autres
Publié: (2025)
Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision
par: Kadasi, Pritam, et autres
Publié: (2026)
par: Kadasi, Pritam, et autres
Publié: (2026)
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
par: Zhou, Yuxuan, et autres
Publié: (2025)
par: Zhou, Yuxuan, et autres
Publié: (2025)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
par: Su, Jinyan, et autres
Publié: (2025)
par: Su, Jinyan, et autres
Publié: (2025)
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
par: Deng, Jiaqi, et autres
Publié: (2025)
par: Deng, Jiaqi, et autres
Publié: (2025)
Look Within, Why LLMs Hallucinate: A Causal Perspective
par: Li, He, et autres
Publié: (2024)
par: Li, He, et autres
Publié: (2024)
Hallucination Detection with the Internal Layers of LLMs
par: Preiß, Martin
Publié: (2025)
par: Preiß, Martin
Publié: (2025)
The First Token Knows: Single-Decode Confidence for Hallucination Detection
par: Gabriel, Mina
Publié: (2026)
par: Gabriel, Mina
Publié: (2026)
TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG
par: Lu, Pengqian, et autres
Publié: (2025)
par: Lu, Pengqian, et autres
Publié: (2025)
Documents similaires
-
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
par: Roig, JV
Publié: (2025) -
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
par: Islam, Saad Obaid ul, et autres
Publié: (2025) -
Scalable and Reliable Evaluation of AI Knowledge Retrieval Systems: RIKER and the Coherent Simulated Universe
par: Roig, JV
Publié: (2025) -
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
par: Zhou, Yang, et autres
Publié: (2025) -
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
par: Vakilian, Vala, et autres
Publié: (2025)