Saved in:
Bibliographic Details
Main Authors: Han, Kyubeen, Jang, Junseo, Kim, Hongjin, Jeong, Geunyeong, Kim, Harksoo
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.18203
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908464331096064
author Han, Kyubeen
Jang, Junseo
Kim, Hongjin
Jeong, Geunyeong
Kim, Harksoo
author_facet Han, Kyubeen
Jang, Junseo
Kim, Hongjin
Jeong, Geunyeong
Kim, Harksoo
contents Instruction-tuning enhances the ability of large language models (LLMs) to follow user instructions more accurately, improving usability while reducing harmful outputs. However, this process may increase the model's dependence on user input, potentially leading to the unfiltered acceptance of misinformation and the generation of hallucinations. Existing studies primarily highlight that LLMs are receptive to external information that contradict their parametric knowledge, but little research has been conducted on the direct impact of instruction-tuning on this phenomenon. In our study, we investigate the impact of instruction-tuning on LLM's susceptibility to misinformation. Our analysis reveals that instruction-tuned LLMs are significantly more likely to accept misinformation when it is presented by the user. A comparison with base models shows that instruction-tuning increases reliance on user-provided information, shifting susceptibility from the assistant role to the user role. Furthermore, we explore additional factors influencing misinformation susceptibility, such as the role of the user in prompt structure, misinformation length, and the presence of warnings in the system prompt. Our findings underscore the need for systematic approaches to mitigate unintended consequences of instruction-tuning and enhance the reliability of LLMs in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18203
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Impact of Instruction-Tuning on LLM's Susceptibility to Misinformation
Han, Kyubeen
Jang, Junseo
Kim, Hongjin
Jeong, Geunyeong
Kim, Harksoo
Computation and Language
Instruction-tuning enhances the ability of large language models (LLMs) to follow user instructions more accurately, improving usability while reducing harmful outputs. However, this process may increase the model's dependence on user input, potentially leading to the unfiltered acceptance of misinformation and the generation of hallucinations. Existing studies primarily highlight that LLMs are receptive to external information that contradict their parametric knowledge, but little research has been conducted on the direct impact of instruction-tuning on this phenomenon. In our study, we investigate the impact of instruction-tuning on LLM's susceptibility to misinformation. Our analysis reveals that instruction-tuned LLMs are significantly more likely to accept misinformation when it is presented by the user. A comparison with base models shows that instruction-tuning increases reliance on user-provided information, shifting susceptibility from the assistant role to the user role. Furthermore, we explore additional factors influencing misinformation susceptibility, such as the role of the user in prompt structure, misinformation length, and the presence of warnings in the system prompt. Our findings underscore the need for systematic approaches to mitigate unintended consequences of instruction-tuning and enhance the reliability of LLMs in real-world applications.
title Exploring the Impact of Instruction-Tuning on LLM's Susceptibility to Misinformation
topic Computation and Language
url https://arxiv.org/abs/2507.18203