Monitoring Emergent Reward Hacking During Generation via Internal Activations
Fuente:
arXiv
Guardado en:
| Autores principales: | Wilhelm, Patrick, Wittkopp, Thorsten, Kao, Odej |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference
por: Wilhelm, Patrick, et al.
Publicado: (2026)
por: Wilhelm, Patrick, et al.
Publicado: (2026)
Revisiting Gradient Staleness: Evaluating Distance Metrics for Asynchronous Federated Learning Aggregation
por: Wilhelm, Patrick, et al.
Publicado: (2026)
por: Wilhelm, Patrick, et al.
Publicado: (2026)
Noise-aware Client Selection for carbon-efficient Federated Learning via Gradient Norm Thresholding
por: Wilhelm, Patrick, et al.
Publicado: (2026)
por: Wilhelm, Patrick, et al.
Publicado: (2026)
Spontaneous Reward Hacking in Iterative Self-Refinement
por: Pan, Jane, et al.
Publicado: (2024)
por: Pan, Jane, et al.
Publicado: (2024)
Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions
por: Flerlage, Justus, et al.
Publicado: (2025)
por: Flerlage, Justus, et al.
Publicado: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
por: Turpin, Miles, et al.
Publicado: (2025)
por: Turpin, Miles, et al.
Publicado: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
por: Wang, Xinpeng, et al.
Publicado: (2025)
por: Wang, Xinpeng, et al.
Publicado: (2025)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
por: Khalifa, Muhammad, et al.
Publicado: (2026)
por: Khalifa, Muhammad, et al.
Publicado: (2026)
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
por: Gallego, Víctor
Publicado: (2025)
por: Gallego, Víctor
Publicado: (2025)
Feedback Loops With Language Models Drive In-Context Reward Hacking
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
por: Ji-An, Li, et al.
Publicado: (2025)
por: Ji-An, Li, et al.
Publicado: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
por: Zhao, Bingchen, et al.
Publicado: (2026)
por: Zhao, Bingchen, et al.
Publicado: (2026)
A layered architecture for log analysis in complex IT systems
por: Wittkopp, Thorsten
Publicado: (2025)
por: Wittkopp, Thorsten
Publicado: (2025)
Towards Machine-Generated Code for the Resolution of User Intentions
por: Flerlage, Justus, et al.
Publicado: (2025)
por: Flerlage, Justus, et al.
Publicado: (2025)
VQEL: Enabling Self-Play in Emergent Language Games via Agent-Internal Vector Quantization
por: Paqaleh, Mohammad Mahdi Samiei, et al.
Publicado: (2025)
por: Paqaleh, Mohammad Mahdi Samiei, et al.
Publicado: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
LogRCA: Log-based Root Cause Analysis for Distributed Services
por: Wittkopp, Thorsten, et al.
Publicado: (2024)
por: Wittkopp, Thorsten, et al.
Publicado: (2024)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
por: Sahoo, Subramanyam
Publicado: (2026)
por: Sahoo, Subramanyam
Publicado: (2026)
GRAM: A Generative Foundation Reward Model for Reward Generalization
por: Wang, Chenglong, et al.
Publicado: (2025)
por: Wang, Chenglong, et al.
Publicado: (2025)
Generative Prompt Internalization
por: Shin, Haebin, et al.
Publicado: (2024)
por: Shin, Haebin, et al.
Publicado: (2024)
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
por: Torrielli, Federico, et al.
Publicado: (2026)
por: Torrielli, Federico, et al.
Publicado: (2026)
Co-Evolution of Policy and Internal Reward for Language Agents
por: Wang, Xinyu, et al.
Publicado: (2026)
por: Wang, Xinyu, et al.
Publicado: (2026)
Negotiating with LLMS: Prompt Hacks, Skill Gaps, and Reasoning Deficits
por: Schneider, Johannes, et al.
Publicado: (2023)
por: Schneider, Johannes, et al.
Publicado: (2023)
Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions
por: Luo, Shunyang, et al.
Publicado: (2026)
por: Luo, Shunyang, et al.
Publicado: (2026)
On Teacher Hacking in Language Model Distillation
por: Tiapkin, Daniil, et al.
Publicado: (2025)
por: Tiapkin, Daniil, et al.
Publicado: (2025)
Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards
por: Lara, Luis, et al.
Publicado: (2026)
por: Lara, Luis, et al.
Publicado: (2026)
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
por: Liu, Runheng, et al.
Publicado: (2026)
por: Liu, Runheng, et al.
Publicado: (2026)
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
por: Schulhoff, Sander, et al.
Publicado: (2023)
por: Schulhoff, Sander, et al.
Publicado: (2023)
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
por: Li, Jiahui, et al.
Publicado: (2024)
por: Li, Jiahui, et al.
Publicado: (2024)
Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study
por: Wiesner, Philipp, et al.
Publicado: (2026)
por: Wiesner, Philipp, et al.
Publicado: (2026)
Automated Rewards via LLM-Generated Progress Functions
por: Sarukkai, Vishnu, et al.
Publicado: (2024)
por: Sarukkai, Vishnu, et al.
Publicado: (2024)
Unleashing Scientific Reasoning for Bio-experimental Protocol Generation via Structured Component-based Reward Mechanism
por: Sun, Haoran, et al.
Publicado: (2025)
por: Sun, Haoran, et al.
Publicado: (2025)
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
por: Elle
Publicado: (2025)
por: Elle
Publicado: (2025)
MT-RewardTree: A Comprehensive Framework for Advancing LLM-Based Machine Translation via Reward Modeling
por: Feng, Zhaopeng, et al.
Publicado: (2025)
por: Feng, Zhaopeng, et al.
Publicado: (2025)
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
por: Qin, Kai, et al.
Publicado: (2026)
por: Qin, Kai, et al.
Publicado: (2026)
Generative Emergent Communication: Large Language Model is a Collective World Model
por: Taniguchi, Tadahiro, et al.
Publicado: (2024)
por: Taniguchi, Tadahiro, et al.
Publicado: (2024)
Ejemplares similares
-
Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference
por: Wilhelm, Patrick, et al.
Publicado: (2026) -
Revisiting Gradient Staleness: Evaluating Distance Metrics for Asynchronous Federated Learning Aggregation
por: Wilhelm, Patrick, et al.
Publicado: (2026) -
Noise-aware Client Selection for carbon-efficient Federated Learning via Gradient Norm Thresholding
por: Wilhelm, Patrick, et al.
Publicado: (2026) -
Spontaneous Reward Hacking in Iterative Self-Refinement
por: Pan, Jane, et al.
Publicado: (2024) -
Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions
por: Flerlage, Justus, et al.
Publicado: (2025)