Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Chengxu, Yuan, Jingling, Cai, Siqi, Jiang, Jiawei, Hu, Chuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
by: Yang, Chengxu, et al.
Published: (2026)
by: Yang, Chengxu, et al.
Published: (2026)
Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided memory for Document-level Machine Translation
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
Heaven & Hell II: Scale Laws and Robustness in One-Step Heaven-Hell Consensus
by: Aghanya, Nnamdi Daniel, et al.
Published: (2025)
by: Aghanya, Nnamdi Daniel, et al.
Published: (2025)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
by: Liu, Xuannan, et al.
Published: (2026)
by: Liu, Xuannan, et al.
Published: (2026)
S2Sent: Nested Selectivity Aware Sentence Representation Learning
by: Zang, Jianxiang, et al.
Published: (2025)
by: Zang, Jianxiang, et al.
Published: (2025)
Heaven, Hell, and Everything In-Between
Published: (2026)
Published: (2026)
Halfway between Heaven and Hell
by: Montgomery, Richard
Published: (2026)
by: Montgomery, Richard
Published: (2026)
ProtSent: Protein Sentence Transformers
by: Ofer, Dan, et al.
Published: (2026)
by: Ofer, Dan, et al.
Published: (2026)
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
by: Luo, Meng, et al.
Published: (2024)
by: Luo, Meng, et al.
Published: (2024)
LLM Hallucination Detection: HSAD
by: Li, JinXin, et al.
Published: (2025)
by: Li, JinXin, et al.
Published: (2025)
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
by: Yu, Jiaqi, et al.
Published: (2026)
by: Yu, Jiaqi, et al.
Published: (2026)
The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings
by: Mochtak, Michal, et al.
Published: (2023)
by: Mochtak, Michal, et al.
Published: (2023)
Lex2Sent: A bagging approach to unsupervised sentiment analysis
by: Lange, Kai-Robin, et al.
Published: (2022)
by: Lange, Kai-Robin, et al.
Published: (2022)
Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once
by: Dhingra, Harnoor
Published: (2026)
by: Dhingra, Harnoor
Published: (2026)
DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning
by: Wang, Xinghao, et al.
Published: (2024)
by: Wang, Xinghao, et al.
Published: (2024)
MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination
by: Li, Zhuo, et al.
Published: (2026)
by: Li, Zhuo, et al.
Published: (2026)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
by: Sun, Siqi, et al.
Published: (2026)
by: Sun, Siqi, et al.
Published: (2026)
CausalSent: Interpretable Sentiment Classification with RieszNet
by: Frees, Daniel, et al.
Published: (2025)
by: Frees, Daniel, et al.
Published: (2025)
AdaptiSent: Context-Aware Adaptive Attention for Multimodal Aspect-Based Sentiment Analysis
by: Rafiuddin, S M, et al.
Published: (2025)
by: Rafiuddin, S M, et al.
Published: (2025)
RobustSentEmbed: Robust Sentence Embeddings Using Adversarial Self-Supervised Contrastive Learning
by: Asl, Javad Rafiei, et al.
Published: (2024)
by: Asl, Javad Rafiei, et al.
Published: (2024)
The Wisdom of Partisan Crowds: Comparing Collective Intelligence in Humans and LLM-based Agents
by: Chuang, Yun-Shiuan, et al.
Published: (2023)
by: Chuang, Yun-Shiuan, et al.
Published: (2023)
BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis
by: Alam, Sadia, et al.
Published: (2024)
by: Alam, Sadia, et al.
Published: (2024)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
by: Wang, Junyang, et al.
Published: (2023)
by: Wang, Junyang, et al.
Published: (2023)
TriCon-Fair: Triplet Contrastive Learning for Mitigating Social Bias in Pre-trained Language Models
by: Lyu, Chong, et al.
Published: (2025)
by: Lyu, Chong, et al.
Published: (2025)
OAEI-LLM: A Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching
by: Qiang, Zhangcheng, et al.
Published: (2024)
by: Qiang, Zhangcheng, et al.
Published: (2024)
Bi'an: A Bilingual Benchmark and Model for Hallucination Detection in Retrieval-Augmented Generation
by: Jiang, Zhouyu, et al.
Published: (2025)
by: Jiang, Zhouyu, et al.
Published: (2025)
LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
by: Yao, Jia-Yu, et al.
Published: (2023)
by: Yao, Jia-Yu, et al.
Published: (2023)
Heaven & Hell: One-Step Hub Consensus
by: Aghanya, Nnamdi Daniel
Published: (2025)
by: Aghanya, Nnamdi Daniel
Published: (2025)
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
by: Chuang, Yu-Neng, et al.
Published: (2025)
by: Chuang, Yu-Neng, et al.
Published: (2025)
DiffuSent: Towards a Unified Diffusion Framework for Aspect-Based Sentiment Analysis
by: Long, Shu, et al.
Published: (2026)
by: Long, Shu, et al.
Published: (2026)
SentGraph: Hierarchical Sentence Graph for Multi-hop Retrieval-Augmented Question Answering
by: Liang, Junli, et al.
Published: (2026)
by: Liang, Junli, et al.
Published: (2026)
Optimal Transport Guided Correlation Assignment for Multimodal Entity Linking
by: Zhang, Zefeng, et al.
Published: (2024)
by: Zhang, Zefeng, et al.
Published: (2024)
HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
by: Tong, Chaodong, et al.
Published: (2025)
by: Tong, Chaodong, et al.
Published: (2025)
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
by: Jha, Rishi, et al.
Published: (2026)
by: Jha, Rishi, et al.
Published: (2026)
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
by: Atasoy, I. F., et al.
Published: (2026)
by: Atasoy, I. F., et al.
Published: (2026)
Removal of Hallucination on Hallucination: Debate-Augmented RAG
by: Hu, Wentao, et al.
Published: (2025)
by: Hu, Wentao, et al.
Published: (2025)
Similar Items
-
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
by: Yang, Chengxu, et al.
Published: (2026) -
Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided memory for Document-level Machine Translation
by: Guo, Jiaxin, et al.
Published: (2025) -
Heaven & Hell II: Scale Laws and Robustness in One-Step Heaven-Hell Consensus
by: Aghanya, Nnamdi Daniel, et al.
Published: (2025) -
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025) -
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
by: Liu, Xuannan, et al.
Published: (2026)