Exploration Hacking: Can LLMs Learn to Resist RL Training?
Fuente:
arXiv
Salvato in:
| Autori principali: | Jang, Eyon, Falck, Damon, Braun, Joschka, Kirch, Nathalie, Menon, Achu, Moodley, Perusha, Emmons, Scott, Zimmermann, Roland S., Lindner, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Stress-Testing Alignment Audits With Prompt-Level Strategic Deception
di: Daniels, Oliver, et al.
Pubblicazione: (2026)
di: Daniels, Oliver, et al.
Pubblicazione: (2026)
Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals
di: Menon, Achyutha, et al.
Pubblicazione: (2026)
di: Menon, Achyutha, et al.
Pubblicazione: (2026)
A Pragmatic Way to Measure Chain-of-Thought Monitorability
di: Emmons, Scott, et al.
Pubblicazione: (2025)
di: Emmons, Scott, et al.
Pubblicazione: (2025)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
di: Braun, Joschka
Pubblicazione: (2026)
di: Braun, Joschka
Pubblicazione: (2026)
Asymmetric Goal Drift in Coding Agents Under Value Conflict
di: Saebo, Magnus, et al.
Pubblicazione: (2026)
di: Saebo, Magnus, et al.
Pubblicazione: (2026)
Early Signs of Steganographic Capabilities in Frontier LLMs
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
Multi-State-Action Tokenisation in Decision Transformers for Multi-Discrete Action Spaces
di: Moodley, Perusha, et al.
Pubblicazione: (2024)
di: Moodley, Perusha, et al.
Pubblicazione: (2024)
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
di: Kaufmann, Max, et al.
Pubblicazione: (2026)
di: Kaufmann, Max, et al.
Pubblicazione: (2026)
Revisiting Safe Exploration in Safe Reinforcement learning
di: Eckel, David, et al.
Pubblicazione: (2024)
di: Eckel, David, et al.
Pubblicazione: (2024)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
di: Kirch, Nathalie, et al.
Pubblicazione: (2024)
di: Kirch, Nathalie, et al.
Pubblicazione: (2024)
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
di: McGuinness, Max, et al.
Pubblicazione: (2025)
di: McGuinness, Max, et al.
Pubblicazione: (2025)
TRIAGE: Ethical Benchmarking of AI Models Through Mass Casualty Simulations
di: Kirch, Nathalie Maria, et al.
Pubblicazione: (2024)
di: Kirch, Nathalie Maria, et al.
Pubblicazione: (2024)
Can LLMs Hack Enterprise Networks? -- Replicated Computational Results (RCR) Report
di: Happe, Andreas, et al.
Pubblicazione: (2026)
di: Happe, Andreas, et al.
Pubblicazione: (2026)
Welcome First--Books Later; The Service Center Branch, Richmond Public Library, December 1967 - June 1971.
di: Emmons, Karen
Pubblicazione: (1971)
di: Emmons, Karen
Pubblicazione: (1971)
She's Practiced What She Teaches
di: Emmons, Julia
Pubblicazione: (1976)
di: Emmons, Julia
Pubblicazione: (1976)
The Impact of Off-Policy Training Data on Probe Generalisation
di: Kirch, Nathalie, et al.
Pubblicazione: (2025)
di: Kirch, Nathalie, et al.
Pubblicazione: (2025)
Natural Emergent Misalignment from Reward Hacking in Production RL
di: MacDiarmid, Monte, et al.
Pubblicazione: (2025)
di: MacDiarmid, Monte, et al.
Pubblicazione: (2025)
Logit Reweighting for Topic-Focused Summarization
di: Braun, Joschka, et al.
Pubblicazione: (2025)
di: Braun, Joschka, et al.
Pubblicazione: (2025)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
di: Braun, Joschka, et al.
Pubblicazione: (2025)
di: Braun, Joschka, et al.
Pubblicazione: (2025)
Land of 10,000 Publishers: Minnesota Children's Book Publishing
di: Kirch, Claire
Pubblicazione: (2008)
di: Kirch, Claire
Pubblicazione: (2008)
Financial constraints and the interdependence of corporate financial decisions A cross-country study
di: Guilherme Kirch
Pubblicazione: (2020)
di: Guilherme Kirch
Pubblicazione: (2020)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
di: Taylor, Mia, et al.
Pubblicazione: (2025)
di: Taylor, Mia, et al.
Pubblicazione: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
di: Wu, Runzhe, et al.
Pubblicazione: (2025)
di: Wu, Runzhe, et al.
Pubblicazione: (2025)
Rapid Integration of LLMs in Healthcare Raises Ethical Concerns: An Investigation into Deceptive Patterns in Social Robots
di: Ranisch, Robert, et al.
Pubblicazione: (2024)
di: Ranisch, Robert, et al.
Pubblicazione: (2024)
The Ethics of ChatGPT in Medicine and Healthcare: A Systematic Review on Large Language Models (LLMs)
di: Haltaufderheide, Joschka, et al.
Pubblicazione: (2024)
di: Haltaufderheide, Joschka, et al.
Pubblicazione: (2024)
Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks
di: Happe, Andreas, et al.
Pubblicazione: (2025)
di: Happe, Andreas, et al.
Pubblicazione: (2025)
Efficient RL Training for LLMs with Experience Replay
di: Arnal, Charles, et al.
Pubblicazione: (2026)
di: Arnal, Charles, et al.
Pubblicazione: (2026)
We Can Imagine the Future, but Are We Equipped to Create It?
di: Jaggars, Damon E.
Pubblicazione: (2014)
di: Jaggars, Damon E.
Pubblicazione: (2014)
Language Models Can Autonomously Hack and Self-Replicate
di: Air, Alena, et al.
Pubblicazione: (2026)
di: Air, Alena, et al.
Pubblicazione: (2026)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
di: Helff, Lukas, et al.
Pubblicazione: (2026)
di: Helff, Lukas, et al.
Pubblicazione: (2026)
Criança e adolescente: a problemática da adoção e posterior devolução às casas de acolhimento
di: Aline Taiane Kirch
Pubblicazione: (2014)
di: Aline Taiane Kirch
Pubblicazione: (2014)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
di: Li, Ziniu, et al.
Pubblicazione: (2025)
di: Li, Ziniu, et al.
Pubblicazione: (2025)
ALMANACS: A Simulatability Benchmark for Language Model Explainability
di: Mills, Edmund, et al.
Pubblicazione: (2023)
di: Mills, Edmund, et al.
Pubblicazione: (2023)
Observation Interference in Partially Observable Assistance Games
di: Emmons, Scott, et al.
Pubblicazione: (2024)
di: Emmons, Scott, et al.
Pubblicazione: (2024)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
di: Bailey, Luke, et al.
Pubblicazione: (2023)
di: Bailey, Luke, et al.
Pubblicazione: (2023)
LLMs Can Learn to Reason Via Off-Policy RL
di: Ritter, Daniel, et al.
Pubblicazione: (2026)
di: Ritter, Daniel, et al.
Pubblicazione: (2026)
Understanding (Un)Reliability of Steering Vectors in Language Models
di: Braun, Joschka, et al.
Pubblicazione: (2025)
di: Braun, Joschka, et al.
Pubblicazione: (2025)
Vendor-Aware Industrial Agents: RAG-Enhanced LLMs for Secure On-Premise PLC Code Generation
di: Kersting, Joschka, et al.
Pubblicazione: (2025)
di: Kersting, Joschka, et al.
Pubblicazione: (2025)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
di: Chen, Wanyi, et al.
Pubblicazione: (2026)
di: Chen, Wanyi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Stress-Testing Alignment Audits With Prompt-Level Strategic Deception
di: Daniels, Oliver, et al.
Pubblicazione: (2026) -
Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals
di: Menon, Achyutha, et al.
Pubblicazione: (2026) -
A Pragmatic Way to Measure Chain-of-Thought Monitorability
di: Emmons, Scott, et al.
Pubblicazione: (2025) -
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
di: Braun, Joschka
Pubblicazione: (2026) -
Asymmetric Goal Drift in Coding Agents Under Value Conflict
di: Saebo, Magnus, et al.
Pubblicazione: (2026)