Recontextualization Mitigates Specification Gaming without Modifying the Specification
Fuente:
arXiv
Saved in:
| Main Authors: | Azarbal, Ariana, Gillioz, Victor, Ivanov, Vladimir, Woodworth, Bryce, Drori, Jacob, Wichers, Nevan, Ebtekar, Aram, Cloud, Alex, Turner, Alexander Matt |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
by: Wichers, Nevan, et al.
Published: (2025)
by: Wichers, Nevan, et al.
Published: (2025)
Output Supervision Can Obfuscate the Chain of Thought
by: Drori, Jacob, et al.
Published: (2025)
by: Drori, Jacob, et al.
Published: (2025)
Gradient-Based Language Model Red Teaming
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Distillation Robustifies Unlearning
by: Lee, Bruce W., et al.
Published: (2025)
by: Lee, Bruce W., et al.
Published: (2025)
Visualizing Neural Network Imagination
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Foundations of algorithmic thermodynamics
by: Ebtekar, Aram, et al.
Published: (2023)
by: Ebtekar, Aram, et al.
Published: (2023)
Golden Handcuffs make safer AI agents
by: Ebtekar, Aram, et al.
Published: (2026)
by: Ebtekar, Aram, et al.
Published: (2026)
Toward Universal Laws of Outlier Propagation
by: Ebtekar, Aram, et al.
Published: (2025)
by: Ebtekar, Aram, et al.
Published: (2025)
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024)
by: Cloud, Alex, et al.
Published: (2024)
Model Spec Midtraining: Improving How Alignment Training Generalizes
by: Li, Chloe, et al.
Published: (2026)
by: Li, Chloe, et al.
Published: (2026)
The momentum-space conformal bootstrap in 2d
by: Gillioz, Marc
Published: (2025)
by: Gillioz, Marc
Published: (2025)
Dynamic Relation Inference via Verb Embeddings
by: Suissa, Omri, et al.
Published: (2025)
by: Suissa, Omri, et al.
Published: (2025)
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
Manipulating ferroelectricity without electrical bias: A perspective
by: Yan, Bixin, et al.
Published: (2026)
by: Yan, Bixin, et al.
Published: (2026)
Recontextualized Knowledge and Narrative Coalitions on Telegram
by: Willaert, Tom
Published: (2024)
by: Willaert, Tom
Published: (2024)
Intensive language contact in the Caucasus
by: Wichers Schreur, Jesse
Published: (2026)
by: Wichers Schreur, Jesse
Published: (2026)
Problems of Associations Accrediting Preparation Programs for School Media Professionals.
by: Wichers, Jean Elaine
Published: (1982)
by: Wichers, Jean Elaine
Published: (1982)
Heart of the Humanities Program
by: Wichers, Jean Elaine
Published: (1976)
by: Wichers, Jean Elaine
Published: (1976)
VIGIL: Tackling Hallucination Detection in Image Recontextualization
by: Wojciechowicz, Joanna, et al.
Published: (2026)
by: Wojciechowicz, Joanna, et al.
Published: (2026)
Recontextualizing Famous Quotes for Brand Slogan Generation
by: Yang, Ziao, et al.
Published: (2026)
by: Yang, Ziao, et al.
Published: (2026)
Recontextualization in Multilingual Science Teacher Professional Learning
by: Kerry Soo Von Esch, et al.
Published: (2025)
by: Kerry Soo Von Esch, et al.
Published: (2025)
Spawning induction in common carp (Cyprinus carpio) using pituitary extract or GnRH superactive analogue combined with metoclopramide : nalysis of hormone profile, progress of oocyte maturation and dependence on temperature / sigal Drori
by: Drori, Sigal
Published: (1994)
by: Drori, Sigal
Published: (1994)
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
by: Gallego, Víctor
Published: (2025)
by: Gallego, Víctor
Published: (2025)
Adventures in FRET and Specification
by: Farrell, Marie, et al.
Published: (2025)
by: Farrell, Marie, et al.
Published: (2025)
What is Formal Verification without Specifications? A Survey on mining LTL Specifications
by: Neider, Daniel, et al.
Published: (2025)
by: Neider, Daniel, et al.
Published: (2025)
AI Blob! LLM-Driven Recontextualization of Italian Television Archives
by: Balestri, Roberto
Published: (2025)
by: Balestri, Roberto
Published: (2025)
Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
by: Grant, Satchel, et al.
Published: (2026)
by: Grant, Satchel, et al.
Published: (2026)
Site-Specific Outdoor Propagation Assessment and Ray-Tracing Analysis for Wireless Digital Twins
by: Aram, Morteza Ghaderi, et al.
Published: (2024)
by: Aram, Morteza Ghaderi, et al.
Published: (2024)
A large synthetic dataset for machine learning applications in power transmission grids
by: Gillioz, Marc, et al.
Published: (2024)
by: Gillioz, Marc, et al.
Published: (2024)
Ludax: A GPU-Accelerated Domain Specific Language for Board Games
by: Todd, Graham, et al.
Published: (2025)
by: Todd, Graham, et al.
Published: (2025)
Sentinel REACT
by: Woodworth, Michael
Published: (2025)
by: Woodworth, Michael
Published: (2025)
Sentinel REACT
by: Woodworth, Michael
Published: (2025)
by: Woodworth, Michael
Published: (2025)
Towards Understanding Specification Gaming in Reasoning Models
by: Nishimura-Gasparian, Kei, et al.
Published: (2026)
by: Nishimura-Gasparian, Kei, et al.
Published: (2026)
Sector-Specific Substitution and the Effect of Sectoral Shocks
by: Gosselin, Jacob Toner
Published: (2025)
by: Gosselin, Jacob Toner
Published: (2025)
Preserving Product Fidelity in Large Scale Image Recontextualization with Diffusion Models
by: Malhi, Ishaan, et al.
Published: (2025)
by: Malhi, Ishaan, et al.
Published: (2025)
Towards a unified and verified understanding of group-operation networks
by: Wu, Wilson, et al.
Published: (2024)
by: Wu, Wilson, et al.
Published: (2024)
Reality Bites
by: Cloud, Dana L.
Published: (2022)
by: Cloud, Dana L.
Published: (2022)
Information and Arbitrage: Applications of Quantum Groups in Mathematical Finance
by: McCloud, Paul
Published: (2017)
by: McCloud, Paul
Published: (2017)
Entender el comic : el arte invisible / Scott McCloud ; traducción Enrique S. Abulí
by: McCloud, Scott
by: McCloud, Scott
Similar Items
-
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
by: Wichers, Nevan, et al.
Published: (2025) -
Output Supervision Can Obfuscate the Chain of Thought
by: Drori, Jacob, et al.
Published: (2025) -
Gradient-Based Language Model Red Teaming
by: Wichers, Nevan, et al.
Published: (2024) -
Distillation Robustifies Unlearning
by: Lee, Bruce W., et al.
Published: (2025) -
Visualizing Neural Network Imagination
by: Wichers, Nevan, et al.
Published: (2024)