The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Taufeeque, Mohammad, Heimersheim, Stefan, Gleave, Adam, Cundy, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Preference Learning with Lie Detectors can Induce Honesty or Evasion
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
Planning in a recurrent neural network that plays Sokoban
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
Benchmarking Deception Probes via Black-to-White Performance Boosts
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
Exploiting Novel GPT-4 APIs
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023)
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023)
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
von: Cundy, Chris, et al.
Veröffentlicht: (2023)
von: Cundy, Chris, et al.
Veröffentlicht: (2023)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
Detecting Strategic Deception Using Linear Probes
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
Training LLMs for Honesty via Confessions
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
von: He, Yuhang, et al.
Veröffentlicht: (2026)
von: He, Yuhang, et al.
Veröffentlicht: (2026)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
von: Kim, Soeun, et al.
Veröffentlicht: (2026)
von: Kim, Soeun, et al.
Veröffentlicht: (2026)
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
von: Dombrowski, Ann-Kathrin, et al.
Veröffentlicht: (2025)
von: Dombrowski, Ann-Kathrin, et al.
Veröffentlicht: (2025)
Deception Abilities Emerged in Large Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2023)
von: Hagendorff, Thilo
Veröffentlicht: (2023)
Building Better Deception Probes Using Targeted Instruction Pairs
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
Can Go AIs be adversarially robust?
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
von: Mitsuhashi, Ryo, et al.
Veröffentlicht: (2026)
von: Mitsuhashi, Ryo, et al.
Veröffentlicht: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
STARC: A General Framework For Quantifying Differences Between Reward Functions
von: Skalse, Joar, et al.
Veröffentlicht: (2023)
von: Skalse, Joar, et al.
Veröffentlicht: (2023)
Think Before You Lie: How Reasoning Leads to Honesty
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
von: Ye, Hao, et al.
Veröffentlicht: (2026)
von: Ye, Hao, et al.
Veröffentlicht: (2026)
The Path Not Taken: RLVR Provably Learns Off the Principals
von: Zhu, Hanqing, et al.
Veröffentlicht: (2025)
von: Zhu, Hanqing, et al.
Veröffentlicht: (2025)
RLVR-World: Training World Models with Reinforcement Learning
von: Wu, Jialong, et al.
Veröffentlicht: (2025)
von: Wu, Jialong, et al.
Veröffentlicht: (2025)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
von: Hao, Zhezheng, et al.
Veröffentlicht: (2025)
von: Hao, Zhezheng, et al.
Veröffentlicht: (2025)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025)
von: Ren, Richard, et al.
Veröffentlicht: (2025)
Scaling Trends for Data Poisoning in LLMs
von: Bowen, Dillon, et al.
Veröffentlicht: (2024)
von: Bowen, Dillon, et al.
Veröffentlicht: (2024)
Where to Measure: Epistemic Uncertainty-Based Sensor Placement with ConvCNPs
von: Eksen, Feyza, et al.
Veröffentlicht: (2025)
von: Eksen, Feyza, et al.
Veröffentlicht: (2025)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
von: Helff, Lukas, et al.
Veröffentlicht: (2026)
von: Helff, Lukas, et al.
Veröffentlicht: (2026)
The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR
von: Adewuyi, Israel, et al.
Veröffentlicht: (2026)
von: Adewuyi, Israel, et al.
Veröffentlicht: (2026)
GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR
von: Zhang, Jiaying, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaying, et al.
Veröffentlicht: (2026)
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
von: Mou, Chaoli, et al.
Veröffentlicht: (2026)
von: Mou, Chaoli, et al.
Veröffentlicht: (2026)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
von: Belakaria, Syrine, et al.
Veröffentlicht: (2025)
von: Belakaria, Syrine, et al.
Veröffentlicht: (2025)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
von: Sun, Guangzhi, et al.
Veröffentlicht: (2025)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Preference Learning with Lie Detectors can Induce Honesty or Evasion
von: Cundy, Chris, et al.
Veröffentlicht: (2025) -
Planning in a recurrent neural network that plays Sokoban
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024) -
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025) -
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
von: Bowen, Dillon, et al.
Veröffentlicht: (2025) -
Benchmarking Deception Probes via Black-to-White Performance Boosts
von: Parrack, Avi, et al.
Veröffentlicht: (2025)