Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Zhibo, Wang, Chen, Shu, Yanfeng, Paik, Hye-young, Zhu, Liming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918282786766848
author Hu, Zhibo
Wang, Chen
Shu, Yanfeng
Paik, Hye-young
Zhu, Liming
author_facet Hu, Zhibo
Wang, Chen
Shu, Yanfeng
Paik, Hye-young
Zhu, Liming
contents Earlier research has shown that metaphors influence human's decision making, which raises the question of whether metaphors also influence large language models (LLMs)' reasoning pathways, considering their training data contain a large number of metaphors. In this work, we investigate the problem in the scope of the emergent misalignment problem where LLMs can generalize patterns learned from misaligned content in one domain to another domain. We discover a strong causal relationship between metaphors in training data and the misalignment degree of LLMs' reasoning contents. With interventions using metaphors in pre-training, fine-tuning and re-alignment phases, models' cross-domain misalignment degrees change significantly. As we delve deeper into the causes behind this phenomenon, we observe that there is a connection between metaphors and the activation of global and local latent features of large reasoning models. By monitoring these latent features, we design a detector that predict misaligned content with high accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03388
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models
Hu, Zhibo
Wang, Chen
Shu, Yanfeng
Paik, Hye-young
Zhu, Liming
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
Earlier research has shown that metaphors influence human's decision making, which raises the question of whether metaphors also influence large language models (LLMs)' reasoning pathways, considering their training data contain a large number of metaphors. In this work, we investigate the problem in the scope of the emergent misalignment problem where LLMs can generalize patterns learned from misaligned content in one domain to another domain. We discover a strong causal relationship between metaphors in training data and the misalignment degree of LLMs' reasoning contents. With interventions using metaphors in pre-training, fine-tuning and re-alignment phases, models' cross-domain misalignment degrees change significantly. As we delve deeper into the causes behind this phenomenon, we observe that there is a connection between metaphors and the activation of global and local latent features of large reasoning models. By monitoring these latent features, we design a detector that predict misaligned content with high accuracy.
title Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
url https://arxiv.org/abs/2601.03388