Atasoy, I. F., Mutlu, B., Sezer, E. A., & Wahdan, A. (2026). Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment.
Chicago Style (17th ed.) CitationAtasoy, I. F., B. Mutlu, E. A. Sezer, and A. Wahdan. Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment. 2026.
MLA (9th ed.) CitationAtasoy, I. F., et al. Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment. 2026.
Warning: These citations may not always be 100% accurate.