Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hejabi, Parsa, Rahmati, Elnaz, Ziabari, Alireza S., Dehghani, Morteza |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoCo-CoLa: Evaluating and Improving Language Adherence in Multilingual LLMs
by: Rahmati, Elnaz, et al.
Published: (2025)
by: Rahmati, Elnaz, et al.
Published: (2025)
Structural Abstraction as an Inductive Bias for Non-Stationary Language Model Training
by: Rahmati, Elnaz, et al.
Published: (2026)
by: Rahmati, Elnaz, et al.
Published: (2026)
Evaluating Creativity and Deception in Large Language Models: A Simulation Framework for Multi-Agent Balderdash
by: Hejabi, Parsa, et al.
Published: (2024)
by: Hejabi, Parsa, et al.
Published: (2024)
The Homogenizing Effect of Large Language Models on Human Expression and Thought
by: Sourati, Zhivar, et al.
Published: (2025)
by: Sourati, Zhivar, et al.
Published: (2025)
Prompt Perturbation Consistency Learning for Robust Language Models
by: Qiang, Yao, et al.
Published: (2024)
by: Qiang, Yao, et al.
Published: (2024)
The UNDO Flip-Flop: A Controlled Probe for Reversible Semantic State Management in State Space Model
by: Zhou, Hongxu
Published: (2026)
by: Zhou, Hongxu
Published: (2026)
Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation
by: Golazizian, Preni, et al.
Published: (2024)
by: Golazizian, Preni, et al.
Published: (2024)
The Moral Foundations Reddit Corpus
by: Trager, Jackson, et al.
Published: (2022)
by: Trager, Jackson, et al.
Published: (2022)
Reasoning on a Spectrum: Aligning LLMs to System 1 and System 2 Thinking
by: Ziabari, Alireza S., et al.
Published: (2025)
by: Ziabari, Alireza S., et al.
Published: (2025)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
by: Shang, Xianpeng, et al.
Published: (2026)
by: Shang, Xianpeng, et al.
Published: (2026)
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs
by: Liu, Hongliang, et al.
Published: (2026)
by: Liu, Hongliang, et al.
Published: (2026)
Secret Keepers: The Impact of LLMs on Linguistic Markers of Personal Traits
by: Sourati, Zhivar, et al.
Published: (2024)
by: Sourati, Zhivar, et al.
Published: (2024)
Unsupervised Contrast-Consistent Ranking with Language Models
by: Stoehr, Niklas, et al.
Published: (2023)
by: Stoehr, Niklas, et al.
Published: (2023)
The Subjectivity of Respect in Police Traffic Stops: Modeling Community Perspectives in Body-Worn Camera Footage
by: Golazizian, Preni, et al.
Published: (2026)
by: Golazizian, Preni, et al.
Published: (2026)
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
by: Laban, Philippe, et al.
Published: (2023)
by: Laban, Philippe, et al.
Published: (2023)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
by: Lee, Jaehyeok, et al.
Published: (2024)
by: Lee, Jaehyeok, et al.
Published: (2024)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
by: Lee, Unggi, et al.
Published: (2026)
by: Lee, Unggi, et al.
Published: (2026)
Consolidating Rewarded Perturbations for LLM Post-Training
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
Enough Coin Flips Can Make LLMs Act Bayesian
by: Gupta, Ritwik, et al.
Published: (2025)
by: Gupta, Ritwik, et al.
Published: (2025)
Ensemble Self-Training for Unsupervised Machine Translation
by: Aharon, Ido, et al.
Published: (2026)
by: Aharon, Ido, et al.
Published: (2026)
The Impact of Quantization on the Robustness of Transformer-based Text Classifiers
by: Neshaei, Seyed Parsa, et al.
Published: (2024)
by: Neshaei, Seyed Parsa, et al.
Published: (2024)
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
by: Wang, Jialu, et al.
Published: (2026)
by: Wang, Jialu, et al.
Published: (2026)
Relational Graph Convolutional Networks for Sentiment Analysis
by: Khosravi, Asal, et al.
Published: (2024)
by: Khosravi, Asal, et al.
Published: (2024)
Unsupervised Data Validation Methods for Efficient Model Training
by: Paniv, Yurii
Published: (2024)
by: Paniv, Yurii
Published: (2024)
Flipping Against All Odds: Reducing LLM Coin Flip Bias via Verbalized Rejection Sampling
by: Xiao, Tim Z., et al.
Published: (2025)
by: Xiao, Tim Z., et al.
Published: (2025)
SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs
by: Jayasuriya, Dinithi, et al.
Published: (2025)
by: Jayasuriya, Dinithi, et al.
Published: (2025)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
by: Shetty, Pranav, et al.
Published: (2025)
by: Shetty, Pranav, et al.
Published: (2025)
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
Review, Remask, Refine (R3): Process-Guided Block Diffusion for Text Generation
by: Mounier, Nikita, et al.
Published: (2025)
by: Mounier, Nikita, et al.
Published: (2025)
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
How Far Can Unsupervised RLVR Scale LLM Training?
by: He, Bingxiang, et al.
Published: (2026)
by: He, Bingxiang, et al.
Published: (2026)
Improving Self Consistency in LLMs through Probabilistic Tokenization
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
by: Wang, Wenxiao, et al.
Published: (2025)
by: Wang, Wenxiao, et al.
Published: (2025)
Auto-Prompt Generation is Not Robust: Prompt Optimization Driven by Pseudo Gradient
by: Shi, Zeru, et al.
Published: (2024)
by: Shi, Zeru, et al.
Published: (2024)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
by: Huang, Kaixuan, et al.
Published: (2025)
by: Huang, Kaixuan, et al.
Published: (2025)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
by: Dai, Hui, et al.
Published: (2026)
by: Dai, Hui, et al.
Published: (2026)
LLM Circuit Analyses Are Consistent Across Training and Scale
by: Tigges, Curt, et al.
Published: (2024)
by: Tigges, Curt, et al.
Published: (2024)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
by: Yang, Xin, et al.
Published: (2026)
by: Yang, Xin, et al.
Published: (2026)
Similar Items
-
CoCo-CoLa: Evaluating and Improving Language Adherence in Multilingual LLMs
by: Rahmati, Elnaz, et al.
Published: (2025) -
Structural Abstraction as an Inductive Bias for Non-Stationary Language Model Training
by: Rahmati, Elnaz, et al.
Published: (2026) -
Evaluating Creativity and Deception in Large Language Models: A Simulation Framework for Multi-Agent Balderdash
by: Hejabi, Parsa, et al.
Published: (2024) -
The Homogenizing Effect of Large Language Models on Human Expression and Thought
by: Sourati, Zhivar, et al.
Published: (2025) -
Prompt Perturbation Consistency Learning for Robust Language Models
by: Qiang, Yao, et al.
Published: (2024)