Saved in:
Bibliographic Details
Main Authors: Martin, Sam, Roger, Fabien
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.12265
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911675553153024
author Martin, Sam
Roger, Fabien
author_facet Martin, Sam
Roger, Fabien
contents Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tuning can bring. We study whether training on multiple classification tasks, each with its own prompt, improves performance on new domains with new classification prompts. We show that such training partially generalizes to adjacent domains, improving classification performance on tasks that are unseen during training. However, we identify specific edge cases where the fine-tuned models fail to follow prompts, such as when the classification prompt changes completely while the data domain remains the same as during training. We show that classification training can be mixed with general instruction following training, and that (when done well) such training keeps the benefits of classification training and mitigates its generalization failures. Surprisingly, we see that this no-thinking supervised classification training can generalize to with-thinking classification and summarization, suggesting that no-thinking classification training might be instrumentally useful in building other kinds of classifiers and monitoring systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12265
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Useful Is Cross-Domain Generalization for Training LLM Monitors?
Martin, Sam
Roger, Fabien
Artificial Intelligence
Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tuning can bring. We study whether training on multiple classification tasks, each with its own prompt, improves performance on new domains with new classification prompts. We show that such training partially generalizes to adjacent domains, improving classification performance on tasks that are unseen during training. However, we identify specific edge cases where the fine-tuned models fail to follow prompts, such as when the classification prompt changes completely while the data domain remains the same as during training. We show that classification training can be mixed with general instruction following training, and that (when done well) such training keeps the benefits of classification training and mitigates its generalization failures. Surprisingly, we see that this no-thinking supervised classification training can generalize to with-thinking classification and summarization, suggesting that no-thinking classification training might be instrumentally useful in building other kinds of classifiers and monitoring systems.
title How Useful Is Cross-Domain Generalization for Training LLM Monitors?
topic Artificial Intelligence
url https://arxiv.org/abs/2605.12265