Auditing language models for hidden objectives
Fuente:
arXiv
Guardado en:
| Autores principales: | Marks, Samuel, Treutlein, Johannes, Bricken, Trenton, Lindsey, Jack, Marcus, Jonathan, Mishra-Sharma, Siddharth, Ziegler, Daniel, Ameisen, Emmanuel, Batson, Joshua, Belonax, Tim, Bowman, Samuel R., Carter, Shan, Chen, Brian, Cunningham, Hoagy, Denison, Carson, Dietz, Florian, Golechha, Satvik, Khan, Akbir, Kirchner, Jan, Leike, Jan, Meek, Austin, Nishimura-Gasparian, Kei, Ong, Euan, Olah, Christopher, Pearce, Adam, Roger, Fabien, Salle, Jeanne, Shih, Andy, Tong, Meg, Thomas, Drake, Rivoire, Kelley, Jermyn, Adam, MacDiarmid, Monte, Henighan, Tom, Hubinger, Evan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
por: Templeton, Adly, et al.
Publicado: (2026)
por: Templeton, Adly, et al.
Publicado: (2026)
Alignment faking in large language models
por: Greenblatt, Ryan, et al.
Publicado: (2024)
por: Greenblatt, Ryan, et al.
Publicado: (2024)
Progress Measures for Grokking on Real-world Tasks
por: Golechha, Satvik
Publicado: (2024)
por: Golechha, Satvik
Publicado: (2024)
When Models Manipulate Manifolds: The Geometry of a Counting Task
por: Gurnee, Wes, et al.
Publicado: (2026)
por: Gurnee, Wes, et al.
Publicado: (2026)
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
por: Jiralerspong, Thomas, et al.
Publicado: (2026)
por: Jiralerspong, Thomas, et al.
Publicado: (2026)
Natural Emergent Misalignment from Reward Hacking in Production RL
por: MacDiarmid, Monte, et al.
Publicado: (2025)
por: MacDiarmid, Monte, et al.
Publicado: (2025)
Challenges in Mechanistically Interpreting Model Representations
por: Golechha, Satvik, et al.
Publicado: (2024)
por: Golechha, Satvik, et al.
Publicado: (2024)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
por: Golechha, Satvik, et al.
Publicado: (2025)
por: Golechha, Satvik, et al.
Publicado: (2025)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
por: Denison, Carson, et al.
Publicado: (2024)
por: Denison, Carson, et al.
Publicado: (2024)
Training Neural Networks for Modularity aids Interpretability
por: Golechha, Satvik, et al.
Publicado: (2024)
por: Golechha, Satvik, et al.
Publicado: (2024)
Emotion Concepts and their Function in a Large Language Model
por: Sofroniew, Nicholas, et al.
Publicado: (2026)
por: Sofroniew, Nicholas, et al.
Publicado: (2026)
Scottish meat consumption survey (Feb–Jul 2023): attitudes, COM-B measures, and perceived effectiveness of meat-reduction policies (Best-Worst Scaling)
por: McBey, David, et al.
Publicado: (2026)
por: McBey, David, et al.
Publicado: (2026)
NICE: To Optimize In-Context Examples or Not?
por: Srivastava, Pragya, et al.
Publicado: (2024)
por: Srivastava, Pragya, et al.
Publicado: (2024)
Circumbinary discs for stellar population models
por: Izzard, Robert G., et al.
Publicado: (2024)
por: Izzard, Robert G., et al.
Publicado: (2024)
A Post‐Pandemic Bail System: Lessons Learned From Supervising Accused During Covid‐19
por: Laura MacDiarmid, et al.
Publicado: (2025)
por: Laura MacDiarmid, et al.
Publicado: (2025)
Building Better Deception Probes Using Targeted Instruction Pairs
por: Natarajan, Vikram, et al.
Publicado: (2026)
por: Natarajan, Vikram, et al.
Publicado: (2026)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
por: Lidayan, Aly, et al.
Publicado: (2025)
por: Lidayan, Aly, et al.
Publicado: (2025)
Studying Cross-cluster Modularity in Neural Networks
por: Golechha, Satvik, et al.
Publicado: (2025)
por: Golechha, Satvik, et al.
Publicado: (2025)
Das Berlin Max Webers
por: Aldenhoff-Hübinger, Rita, et al.
Publicado: (2026)
por: Aldenhoff-Hübinger, Rita, et al.
Publicado: (2026)
Polysemanticity and Capacity in Neural Networks
por: Scherlis, Adam, et al.
Publicado: (2022)
por: Scherlis, Adam, et al.
Publicado: (2022)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
por: Balappanawar, Ishwar, et al.
Publicado: (2025)
por: Balappanawar, Ishwar, et al.
Publicado: (2025)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
por: Treutlein, Johannes, et al.
Publicado: (2024)
por: Treutlein, Johannes, et al.
Publicado: (2024)
Latent Planning Emerges with Scale
por: Hanna, Michael, et al.
Publicado: (2026)
por: Hanna, Michael, et al.
Publicado: (2026)
Descriptions of new fishes from Panama
por: Meek, Seth E., et al.
Publicado: (1912)
por: Meek, Seth E., et al.
Publicado: (1912)
New species of fishes from Panama
por: Meek, Seth E., et al.
Publicado: (1913)
por: Meek, Seth E., et al.
Publicado: (1913)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2024)
por: Chanin, David, et al.
Publicado: (2024)
Europa desde 1918 hasta hoy / Mario Rivoire ; Traducción al español por Carlos Gerhard
por: Rivoire, Mario
Publicado: (1918)
por: Rivoire, Mario
Publicado: (1918)
Continuous data assimilation for the Richards equation of unsaturated flow: A computational study
por: Amanda Rowley, et al.
Publicado: (2026)
por: Amanda Rowley, et al.
Publicado: (2026)
Limit-Computable Grains of Truth for Arbitrary Computable Extensive-Form (Un)Known Games
por: Wyeth, Cole, et al.
Publicado: (2025)
por: Wyeth, Cole, et al.
Publicado: (2025)
Excess Description Length of Learning Generalizable Predictors
por: Donoway, Elizabeth, et al.
Publicado: (2026)
por: Donoway, Elizabeth, et al.
Publicado: (2026)
Government Publications in Kansas Public Libraries.
por: Batson, Donald
Publicado: (1975)
por: Batson, Donald
Publicado: (1975)
Steering Language Models With Activation Engineering
por: Turner, Alexander Matt, et al.
Publicado: (2023)
por: Turner, Alexander Matt, et al.
Publicado: (2023)
CataractBot: An LLM-Powered Expert-in-the-Loop Chatbot for Cataract Patients
por: Ramjee, Pragnya, et al.
Publicado: (2024)
por: Ramjee, Pragnya, et al.
Publicado: (2024)
Descriptions of new Fishes from Panama
por: Meek, Seth E. (Seth Eugene), et al.
Publicado: (1912)
por: Meek, Seth E. (Seth Eugene), et al.
Publicado: (1912)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
por: Taylor, Mia, et al.
Publicado: (2025)
por: Taylor, Mia, et al.
Publicado: (2025)
Pendant appearances and components in random graphs from structured classes
por: McDiarmid, Colin
Publicado: (2021)
por: McDiarmid, Colin
Publicado: (2021)
The Savage Worlds of Henry Drummond (1851–1897): Science, Racism and Religion in the Work of a Popular Evolutionist
por: Diarmid A. Finnegan
Publicado: (2025)
por: Diarmid A. Finnegan
Publicado: (2025)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
por: Hubinger, Evan, et al.
Publicado: (2024)
por: Hubinger, Evan, et al.
Publicado: (2024)
State Government Publications: Selection, Acquisition, and Reference Service.
por: Batson, Donald W.
Publicado: (1990)
por: Batson, Donald W.
Publicado: (1990)
AD genes and microglia phenotypes: does microglia phenotypic heterogeneity matter?
por: Marta Olah
Publicado: (2024)
por: Marta Olah
Publicado: (2024)
Ejemplares similares
-
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
por: Templeton, Adly, et al.
Publicado: (2026) -
Alignment faking in large language models
por: Greenblatt, Ryan, et al.
Publicado: (2024) -
Progress Measures for Grokking on Real-world Tasks
por: Golechha, Satvik
Publicado: (2024) -
When Models Manipulate Manifolds: The Geometry of a Counting Task
por: Gurnee, Wes, et al.
Publicado: (2026) -
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
por: Jiralerspong, Thomas, et al.
Publicado: (2026)