Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Belyi, Masha, Friel, Robert, Shao, Shuai, Sanyal, Atindriyo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
von: Friel, Robert, et al.
Veröffentlicht: (2024)
von: Friel, Robert, et al.
Veröffentlicht: (2024)
Luna-2: Scalable Single-Token Evaluation with Small Language Models
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
Theoretical Foundations and Mitigation of Hallucination in Large Language Models
von: Gumaan, Esmail
Veröffentlicht: (2025)
von: Gumaan, Esmail
Veröffentlicht: (2025)
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
von: Doshi, Savan
Veröffentlicht: (2026)
von: Doshi, Savan
Veröffentlicht: (2026)
Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
von: Na, Injae, et al.
Veröffentlicht: (2025)
von: Na, Injae, et al.
Veröffentlicht: (2025)
Pruning Foundation Models for High Accuracy without Retraining
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
von: Kim, Yubin, et al.
Veröffentlicht: (2025)
von: Kim, Yubin, et al.
Veröffentlicht: (2025)
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
Cost-of-Pass: An Economic Framework for Evaluating Language Models
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2025)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2025)
Calibrated Language Models Must Hallucinate
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2023)
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2023)
Hallucination Detection with Small Language Models
von: Cheung, Ming
Veröffentlicht: (2025)
von: Cheung, Ming
Veröffentlicht: (2025)
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
von: Benkirane, Kenza, et al.
Veröffentlicht: (2024)
von: Benkirane, Kenza, et al.
Veröffentlicht: (2024)
Lynx: An Open Source Hallucination Evaluation Model
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
Catching Chameleons: Detecting Evolving Disinformation Generated using Large Language Models
von: Jiang, Bohan, et al.
Veröffentlicht: (2024)
von: Jiang, Bohan, et al.
Veröffentlicht: (2024)
Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA
von: Baqar, Mohammad, et al.
Veröffentlicht: (2025)
von: Baqar, Mohammad, et al.
Veröffentlicht: (2025)
Quantifying Hallucinations in Language Language Models on Medical Textbooks
von: Colelough, Brandon C., et al.
Veröffentlicht: (2026)
von: Colelough, Brandon C., et al.
Veröffentlicht: (2026)
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
von: Chen, Kedi, et al.
Veröffentlicht: (2024)
von: Chen, Kedi, et al.
Veröffentlicht: (2024)
Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
An Evolutionary Large Language Model for Hallucination Mitigation
von: Boulesnane, Abdennour, et al.
Veröffentlicht: (2024)
von: Boulesnane, Abdennour, et al.
Veröffentlicht: (2024)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
An Empirical Study on Large Language Models in Accuracy and Robustness under Chinese Industrial Scenarios
von: Li, Zongjie, et al.
Veröffentlicht: (2024)
von: Li, Zongjie, et al.
Veröffentlicht: (2024)
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks
von: Zhao, Justin, et al.
Veröffentlicht: (2024)
von: Zhao, Justin, et al.
Veröffentlicht: (2024)
Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology
von: Liu, Dou, et al.
Veröffentlicht: (2025)
von: Liu, Dou, et al.
Veröffentlicht: (2025)
The System Hallucination Scale (SHS): A Minimal yet Effective Human-Centered Instrument for Evaluating Hallucination-Related Behavior in Large Language Models
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models
von: Liu, Chonghan, et al.
Veröffentlicht: (2026)
von: Liu, Chonghan, et al.
Veröffentlicht: (2026)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
Confabulation: The Surprising Value of Large Language Model Hallucinations
von: Sui, Peiqi, et al.
Veröffentlicht: (2024)
von: Sui, Peiqi, et al.
Veröffentlicht: (2024)
The Science of Evaluating Foundation Models
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
Copy-Paste to Mitigate Large Language Model Hallucinations
von: Long, Yongchao, et al.
Veröffentlicht: (2025)
von: Long, Yongchao, et al.
Veröffentlicht: (2025)
The Impact of Negated Text on Hallucination with Large Language Models
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025)
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025)
Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
von: Sato, Makoto
Veröffentlicht: (2025)
von: Sato, Makoto
Veröffentlicht: (2025)
Free Lunch for Pass@$k$? Low Cost Diverse Sampling for Diffusion Language Models
von: Lamont, Sean, et al.
Veröffentlicht: (2026)
von: Lamont, Sean, et al.
Veröffentlicht: (2026)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
How Large Language Models are Designed to Hallucinate
von: Ackermann, Richard, et al.
Veröffentlicht: (2025)
von: Ackermann, Richard, et al.
Veröffentlicht: (2025)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
HICD: Hallucination-Inducing via Attention Dispersion for Contrastive Decoding to Mitigate Hallucinations in Large Language Models
von: Jiang, Xinyan, et al.
Veröffentlicht: (2025)
von: Jiang, Xinyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
von: Friel, Robert, et al.
Veröffentlicht: (2024) -
Luna-2: Scalable Single-Token Evaluation with Small Language Models
von: Goel, Vatsal, et al.
Veröffentlicht: (2026) -
Theoretical Foundations and Mitigation of Hallucination in Large Language Models
von: Gumaan, Esmail
Veröffentlicht: (2025) -
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
von: Doshi, Savan
Veröffentlicht: (2026) -
Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
von: Na, Injae, et al.
Veröffentlicht: (2025)