How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Hao, Lv, Liuzhenghao, Cao, He, Liu, Zijing, Yan, Zhiyuan, Wang, Yu, Tian, Yonghong, Li, Yu, Yuan, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916693501018112
author Li, Hao
Lv, Liuzhenghao
Cao, He
Liu, Zijing
Yan, Zhiyuan
Wang, Yu
Tian, Yonghong
Li, Yu
Yuan, Li
author_facet Li, Hao
Lv, Liuzhenghao
Cao, He
Liu, Zijing
Yan, Zhiyuan
Wang, Yu
Tian, Yonghong
Li, Yu
Yuan, Li
contents Large language models are increasingly used in scientific domains, especially for molecular understanding and analysis. However, existing models are affected by hallucination issues, resulting in errors in drug design and utilization. In this paper, we first analyze the sources of hallucination in LLMs for molecular comprehension tasks, specifically the knowledge shortcut phenomenon observed in the PubChem dataset. To evaluate hallucination in molecular comprehension tasks with computational efficiency, we introduce \textbf{Mol-Hallu}, a novel free-form evaluation metric that quantifies the degree of hallucination based on the scientific entailment relationship between generated text and actual molecular properties. Utilizing the Mol-Hallu metric, we reassess and analyze the extent of hallucination in various LLMs performing molecular comprehension tasks. Furthermore, the Hallucination Reduction Post-processing stage~(HRPP) is proposed to alleviate molecular hallucinations, Experiments show the effectiveness of HRPP on decoder-only and encoder-decoder molecular LLMs. Our findings provide critical insights into mitigating hallucination and improving the reliability of LLMs in scientific applications.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12314
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
Li, Hao
Lv, Liuzhenghao
Cao, He
Liu, Zijing
Yan, Zhiyuan
Wang, Yu
Tian, Yonghong
Li, Yu
Yuan, Li
Computation and Language
Artificial Intelligence
Large language models are increasingly used in scientific domains, especially for molecular understanding and analysis. However, existing models are affected by hallucination issues, resulting in errors in drug design and utilization. In this paper, we first analyze the sources of hallucination in LLMs for molecular comprehension tasks, specifically the knowledge shortcut phenomenon observed in the PubChem dataset. To evaluate hallucination in molecular comprehension tasks with computational efficiency, we introduce \textbf{Mol-Hallu}, a novel free-form evaluation metric that quantifies the degree of hallucination based on the scientific entailment relationship between generated text and actual molecular properties. Utilizing the Mol-Hallu metric, we reassess and analyze the extent of hallucination in various LLMs performing molecular comprehension tasks. Furthermore, the Hallucination Reduction Post-processing stage~(HRPP) is proposed to alleviate molecular hallucinations, Experiments show the effectiveness of HRPP on decoder-only and encoder-decoder molecular LLMs. Our findings provide critical insights into mitigating hallucination and improving the reliability of LLMs in scientific applications.
title How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.12314