Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Xin, Li, Letian, Wuerkaixi, Abudukelimu, Cheng, Xuxin, Liu, Cao, Zeng, Ke, Cai, Xunliang, Jiang, Wenyuan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914367459557376
author Yang, Xin
Li, Letian
Wuerkaixi, Abudukelimu
Cheng, Xuxin
Liu, Cao
Zeng, Ke
Cai, Xunliang
Jiang, Wenyuan
author_facet Yang, Xin
Li, Letian
Wuerkaixi, Abudukelimu
Cheng, Xuxin
Liu, Cao
Zeng, Ke
Cai, Xunliang
Jiang, Wenyuan
contents Large language models (LLMs) have demonstrated remarkable and steadily improving performance across a wide range of tasks. However, LLM performance may be highly sensitive to prompt variations especially in scenarios with limited openness or strict output formatting requirements, indicating insufficient robustness. In real-world applications, user prompts provided to LLMs often contain imperfections, which may undermine the quality of the model's responses. To address this issue, previous work has primarily focused on preprocessing prompts, employing external tools or even LLMs to refine prompt formulations in advance. However, these approaches overlook the intrinsic robustness of LLMs, and their reliance on external components introduces additional computational overhead and uncertainty. In this work, we propose a Contrastive Learning-based Inverse Direct Preference Optimization (CoIPO) method that minimizes the discrepancy between the label-aligned logits produced by the model under a clean prompt and its noisy counterpart, and conduct a detailed analysis using mutual information theory. We augment the FLAN dataset by constructing paired prompts, each consisting of a clean prompt and its corresponding noisy version for training. Additionally, to evaluate the effectiveness, we develop NoisyPromptBench, a benchmark enhanced and derived from the existing PromptBench. Experimental results conducted on NoisyPromptBench demonstrate that our proposed method achieves a significant improvement in average accuracy over the current state-of-the-art approaches. The source code of CoIPO, pair-wise FLAN datasets, and NoisyPromptBench have already been released on https://github.com/vegetable-yx/CoIPO.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03314
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
Yang, Xin
Li, Letian
Wuerkaixi, Abudukelimu
Cheng, Xuxin
Liu, Cao
Zeng, Ke
Cai, Xunliang
Jiang, Wenyuan
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) have demonstrated remarkable and steadily improving performance across a wide range of tasks. However, LLM performance may be highly sensitive to prompt variations especially in scenarios with limited openness or strict output formatting requirements, indicating insufficient robustness. In real-world applications, user prompts provided to LLMs often contain imperfections, which may undermine the quality of the model's responses. To address this issue, previous work has primarily focused on preprocessing prompts, employing external tools or even LLMs to refine prompt formulations in advance. However, these approaches overlook the intrinsic robustness of LLMs, and their reliance on external components introduces additional computational overhead and uncertainty. In this work, we propose a Contrastive Learning-based Inverse Direct Preference Optimization (CoIPO) method that minimizes the discrepancy between the label-aligned logits produced by the model under a clean prompt and its noisy counterpart, and conduct a detailed analysis using mutual information theory. We augment the FLAN dataset by constructing paired prompts, each consisting of a clean prompt and its corresponding noisy version for training. Additionally, to evaluate the effectiveness, we develop NoisyPromptBench, a benchmark enhanced and derived from the existing PromptBench. Experimental results conducted on NoisyPromptBench demonstrate that our proposed method achieves a significant improvement in average accuracy over the current state-of-the-art approaches. The source code of CoIPO, pair-wise FLAN datasets, and NoisyPromptBench have already been released on https://github.com/vegetable-yx/CoIPO.
title Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.03314