Can Hallucinations Help? Boosting LLMs for Drug Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Shuzhou, Qu, Zhan, Kangen, Ashish Yashwanth, Färber, Michael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916911610068992
author Yuan, Shuzhou
Qu, Zhan
Kangen, Ashish Yashwanth
Färber, Michael
author_facet Yuan, Shuzhou
Qu, Zhan
Kangen, Ashish Yashwanth
Färber, Michael
contents Hallucinations in large language models (LLMs), plausible but factually inaccurate text, are often viewed as undesirable. However, recent work suggests that such outputs may hold creative potential. In this paper, we investigate whether hallucinations can improve LLMs on molecule property prediction, a key task in early-stage drug discovery. We prompt LLMs to generate natural language descriptions from molecular SMILES strings and incorporate these often hallucinated descriptions into downstream classification tasks. Evaluating seven instruction-tuned LLMs across five datasets, we find that hallucinations significantly improve predictive accuracy for some models. Notably, Falcon3-Mamba-7B outperforms all baselines when hallucinated text is included, while hallucinations generated by GPT-4o consistently yield the greatest gains between models. We further identify and categorize over 18,000 beneficial hallucinations, with structural misdescriptions emerging as the most impactful type, suggesting that hallucinated statements about molecular structure may increase model confidence. Ablation studies show that larger models benefit more from hallucinations, while temperature has a limited effect. Our findings challenge conventional views of hallucination as purely problematic and suggest new directions for leveraging hallucinations as a useful signal in scientific modeling tasks like drug discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13824
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Hallucinations Help? Boosting LLMs for Drug Discovery
Yuan, Shuzhou
Qu, Zhan
Kangen, Ashish Yashwanth
Färber, Michael
Computation and Language
Artificial Intelligence
Hallucinations in large language models (LLMs), plausible but factually inaccurate text, are often viewed as undesirable. However, recent work suggests that such outputs may hold creative potential. In this paper, we investigate whether hallucinations can improve LLMs on molecule property prediction, a key task in early-stage drug discovery. We prompt LLMs to generate natural language descriptions from molecular SMILES strings and incorporate these often hallucinated descriptions into downstream classification tasks. Evaluating seven instruction-tuned LLMs across five datasets, we find that hallucinations significantly improve predictive accuracy for some models. Notably, Falcon3-Mamba-7B outperforms all baselines when hallucinated text is included, while hallucinations generated by GPT-4o consistently yield the greatest gains between models. We further identify and categorize over 18,000 beneficial hallucinations, with structural misdescriptions emerging as the most impactful type, suggesting that hallucinated statements about molecular structure may increase model confidence. Ablation studies show that larger models benefit more from hallucinations, while temperature has a limited effect. Our findings challenge conventional views of hallucination as purely problematic and suggest new directions for leveraging hallucinations as a useful signal in scientific modeling tasks like drug discovery.
title Can Hallucinations Help? Boosting LLMs for Drug Discovery
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.13824