HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Zhiying, Yang, Yiming, Sun, Zhiqing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
HaluMem: Evaluating Hallucinations in Memory Systems of Agents
di: Chen, Ding, et al.
Pubblicazione: (2025)
di: Chen, Ding, et al.
Pubblicazione: (2025)
MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning
di: Huang, Ruijun, et al.
Pubblicazione: (2026)
di: Huang, Ruijun, et al.
Pubblicazione: (2026)
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
di: Chen, Kedi, et al.
Pubblicazione: (2024)
di: Chen, Kedi, et al.
Pubblicazione: (2024)
MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models
di: Agarwal, Vibhor, et al.
Pubblicazione: (2024)
di: Agarwal, Vibhor, et al.
Pubblicazione: (2024)
Halu-J: Critique-Based Hallucination Judge
di: Wang, Binjie, et al.
Pubblicazione: (2024)
di: Wang, Binjie, et al.
Pubblicazione: (2024)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
di: Lu, Yujie, et al.
Pubblicazione: (2024)
di: Lu, Yujie, et al.
Pubblicazione: (2024)
The Mirage of Model Editing: Revisiting Evaluation in the Wild
di: Yang, Wanli, et al.
Pubblicazione: (2025)
di: Yang, Wanli, et al.
Pubblicazione: (2025)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
di: Tian, Yuchen, et al.
Pubblicazione: (2024)
di: Tian, Yuchen, et al.
Pubblicazione: (2024)
OnionEval: An Unified Evaluation of Fact-conflicting Hallucination for Small-Large Language Models
di: Sun, Chongren, et al.
Pubblicazione: (2025)
di: Sun, Chongren, et al.
Pubblicazione: (2025)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
di: Tong, Chaodong, et al.
Pubblicazione: (2025)
di: Tong, Chaodong, et al.
Pubblicazione: (2025)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
di: Urlana, Ashok, et al.
Pubblicazione: (2025)
di: Urlana, Ashok, et al.
Pubblicazione: (2025)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
di: Hosseini, Mohammad, et al.
Pubblicazione: (2025)
di: Hosseini, Mohammad, et al.
Pubblicazione: (2025)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
di: Wang, Pengyu, et al.
Pubblicazione: (2026)
di: Wang, Pengyu, et al.
Pubblicazione: (2026)
FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation
di: Bayat, Farima Fatahi, et al.
Pubblicazione: (2024)
di: Bayat, Farima Fatahi, et al.
Pubblicazione: (2024)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
di: Liu, Tengxiao, et al.
Pubblicazione: (2026)
di: Liu, Tengxiao, et al.
Pubblicazione: (2026)
Benchmarking Table Comprehension In The Wild
di: Pan, Yikang, et al.
Pubblicazione: (2024)
di: Pan, Yikang, et al.
Pubblicazione: (2024)
WildIFEval: Instruction Following in the Wild
di: Lior, Gili, et al.
Pubblicazione: (2025)
di: Lior, Gili, et al.
Pubblicazione: (2025)
Translation in the Wild
di: Balashov, Yuri
Pubblicazione: (2025)
di: Balashov, Yuri
Pubblicazione: (2025)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
di: Peng, Hao, et al.
Pubblicazione: (2026)
di: Peng, Hao, et al.
Pubblicazione: (2026)
Evaluation Framework for AI Systems in "the Wild"
di: Jabbour, Sarah, et al.
Pubblicazione: (2025)
di: Jabbour, Sarah, et al.
Pubblicazione: (2025)
WildChat: 1M ChatGPT Interaction Logs in the Wild
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models
di: Lei, Fangyu, et al.
Pubblicazione: (2023)
di: Lei, Fangyu, et al.
Pubblicazione: (2023)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
di: Tang, Liyan, et al.
Pubblicazione: (2024)
di: Tang, Liyan, et al.
Pubblicazione: (2024)
Tool Learning in the Wild: Empowering Language Models as Automatic Tool Agents
di: Shi, Zhengliang, et al.
Pubblicazione: (2024)
di: Shi, Zhengliang, et al.
Pubblicazione: (2024)
Can LLMs Reason in the Wild with Programs?
di: Yang, Yuan, et al.
Pubblicazione: (2024)
di: Yang, Yuan, et al.
Pubblicazione: (2024)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
GuessBench: Sensemaking Multimodal Creativity in the Wild
di: Zhu, Zifeng, et al.
Pubblicazione: (2025)
di: Zhu, Zifeng, et al.
Pubblicazione: (2025)
A Systematic Study of In-the-Wild Model Merging for Large Language Models
di: Hitit, Oğuz Kağan, et al.
Pubblicazione: (2025)
di: Hitit, Oğuz Kağan, et al.
Pubblicazione: (2025)
Membership Inference on LLMs in the Wild
di: Yi, Jiatong, et al.
Pubblicazione: (2026)
di: Yi, Jiatong, et al.
Pubblicazione: (2026)
MMFormalizer: Multimodal Autoformalization in the Wild
di: Xiong, Jing, et al.
Pubblicazione: (2026)
di: Xiong, Jing, et al.
Pubblicazione: (2026)
Reconstructing Animals and the Wild
di: Kulits, Peter, et al.
Pubblicazione: (2024)
di: Kulits, Peter, et al.
Pubblicazione: (2024)
CocoaBench: Evaluating Unified Digital Agents in the Wild
di: CocoaBench Team, et al.
Pubblicazione: (2026)
di: CocoaBench Team, et al.
Pubblicazione: (2026)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025) -
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025) -
HaluMem: Evaluating Hallucinations in Memory Systems of Agents
di: Chen, Ding, et al.
Pubblicazione: (2025) -
MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning
di: Huang, Ruijun, et al.
Pubblicazione: (2026) -
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
di: Chen, Kedi, et al.
Pubblicazione: (2024)