From Data to Behavior: Predicting Unintended Model Behaviors Before Training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Mengru, Xu, Zhenqian, Fang, Junfeng, Yao, Yunzhi, Deng, Shumin, Chen, Huajun, Zhang, Ningyu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911421888987136
author Wang, Mengru
Xu, Zhenqian
Fang, Junfeng
Yao, Yunzhi
Deng, Shumin
Chen, Huajun
Zhang, Ningyu
author_facet Wang, Mengru
Xu, Zhenqian
Fang, Junfeng
Yao, Yunzhi
Deng, Shumin
Chen, Huajun
Zhang, Ningyu
contents Large Language Models (LLMs) can acquire unintended biases from seemingly benign training data even without explicit cues or malicious content. Existing methods struggle to detect such risks before fine-tuning, making post hoc evaluation costly and inefficient. To address this challenge, we introduce Data2Behavior, a new task for predicting unintended model behaviors prior to training. We also propose Manipulating Data Features (MDF), a lightweight approach that summarizes candidate data through their mean representations and injects them into the forward pass of a base model, allowing latent statistical signals in the data to shape model activations and reveal potential biases and safety risks without updating any parameters. MDF achieves reliable prediction while consuming only about 20% of the GPU resources required for fine-tuning. Experiments on Qwen3-14B, Qwen2.5-32B-Instruct, and Gemma-3-12b-it confirm that MDF can anticipate unintended behaviors and provide insight into pre-training vulnerabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04735
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Data to Behavior: Predicting Unintended Model Behaviors Before Training
Wang, Mengru
Xu, Zhenqian
Fang, Junfeng
Yao, Yunzhi
Deng, Shumin
Chen, Huajun
Zhang, Ningyu
Machine Learning
Artificial Intelligence
Computation and Language
Computers and Society
Information Retrieval
Large Language Models (LLMs) can acquire unintended biases from seemingly benign training data even without explicit cues or malicious content. Existing methods struggle to detect such risks before fine-tuning, making post hoc evaluation costly and inefficient. To address this challenge, we introduce Data2Behavior, a new task for predicting unintended model behaviors prior to training. We also propose Manipulating Data Features (MDF), a lightweight approach that summarizes candidate data through their mean representations and injects them into the forward pass of a base model, allowing latent statistical signals in the data to shape model activations and reveal potential biases and safety risks without updating any parameters. MDF achieves reliable prediction while consuming only about 20% of the GPU resources required for fine-tuning. Experiments on Qwen3-14B, Qwen2.5-32B-Instruct, and Gemma-3-12b-it confirm that MDF can anticipate unintended behaviors and provide insight into pre-training vulnerabilities.
title From Data to Behavior: Predicting Unintended Model Behaviors Before Training
topic Machine Learning
Artificial Intelligence
Computation and Language
Computers and Society
Information Retrieval
url https://arxiv.org/abs/2602.04735