Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
Fuente:
arXiv
Saved in:
| Main Authors: | Roytburg, Dani, Bozoukov, Matthew, Nguyen, Matthew, Barzdukas, Jou, Fu, Simon, Oozeer, Narmeen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
by: Bozoukov, Matthew, et al.
Published: (2025)
by: Bozoukov, Matthew, et al.
Published: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Approximating Human Preferences Using a Multi-Judge Learned System
by: Sprejer, Eitán, et al.
Published: (2025)
by: Sprejer, Eitán, et al.
Published: (2025)
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024)
by: Oozeer, Narmeen, et al.
Published: (2024)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)
by: Yang, Jinming, et al.
Published: (2026)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
Measuring Reasoning Trace Legibility: Can Those Who Understand Teach?
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026)
by: Ivanov, Georgi, et al.
Published: (2026)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
by: Finlayson, Matthew, et al.
Published: (2025)
by: Finlayson, Matthew, et al.
Published: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
by: Yang, Matthew Y. R., et al.
Published: (2026)
by: Yang, Matthew Y. R., et al.
Published: (2026)
Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
by: Ding, Yiwen, et al.
Published: (2024)
by: Ding, Yiwen, et al.
Published: (2024)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026)
by: Karlekar, Sweta, et al.
Published: (2026)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
by: Mündler, Niels, et al.
Published: (2023)
by: Mündler, Niels, et al.
Published: (2023)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024)
by: Liu, Ollie, et al.
Published: (2024)
Self-Consistency Preference Optimization
by: Prasad, Archiki, et al.
Published: (2024)
by: Prasad, Archiki, et al.
Published: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025)
by: Gupta, Taneesh, et al.
Published: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
by: Mohamed, Anas, et al.
Published: (2025)
by: Mohamed, Anas, et al.
Published: (2025)
Advancing LLM Reasoning Generalists with Preference Trees
by: Yuan, Lifan, et al.
Published: (2024)
by: Yuan, Lifan, et al.
Published: (2024)
Mitigating Hallucinated Translations in Large Language Models with Hallucination-focused Preference Optimization
by: Tang, Zilu, et al.
Published: (2025)
by: Tang, Zilu, et al.
Published: (2025)
TSO: Self-Training with Scaled Preference Optimization
by: Chen, Kaihui, et al.
Published: (2024)
by: Chen, Kaihui, et al.
Published: (2024)
Turning LLM Activations Quantization-Friendly
by: Czakó, Patrik, et al.
Published: (2025)
by: Czakó, Patrik, et al.
Published: (2025)
RouteLLM: Learning to Route LLMs with Preference Data
by: Ong, Isaac, et al.
Published: (2024)
by: Ong, Isaac, et al.
Published: (2024)
Mitigating LLM Hallucinations via Conformal Abstention
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
On Mitigating Code LLM Hallucinations with API Documentation
by: Jain, Nihal, et al.
Published: (2024)
by: Jain, Nihal, et al.
Published: (2024)
Self-Play Preference Optimization for Language Model Alignment
by: Wu, Yue, et al.
Published: (2024)
by: Wu, Yue, et al.
Published: (2024)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
by: Yuan, Yurun, et al.
Published: (2026)
by: Yuan, Yurun, et al.
Published: (2026)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
DreamReader: An Interpretability Toolkit for Text-to-Image Models
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
by: Kundurthy, Srivatsa, et al.
Published: (2026)
by: Kundurthy, Srivatsa, et al.
Published: (2026)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
by: Hu, Zhengyu, et al.
Published: (2024)
by: Hu, Zhengyu, et al.
Published: (2024)
West-of-N: Synthetic Preferences for Self-Improving Reward Models
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
by: Kutakh, Matthew
Published: (2026)
by: Kutakh, Matthew
Published: (2026)
SMART: Self-Aware Agent for Tool Overuse Mitigation
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Similar Items
-
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026) -
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026) -
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
by: Bozoukov, Matthew, et al.
Published: (2025) -
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025) -
Approximating Human Preferences Using a Multi-Judge Learned System
by: Sprejer, Eitán, et al.
Published: (2025)