Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Fuente:
arXiv
Guardado en:
| Autores principales: | Baker, Bowen, Huizinga, Joost, Gao, Leo, Dou, Zehao, Guan, Melody Y., Madry, Aleksander, Zaremba, Wojciech, Pachocki, Jakub, Farhi, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Monitoring Monitorability
por: Guan, Melody Y., et al.
Publicado: (2025)
por: Guan, Melody Y., et al.
Publicado: (2025)
Hodoscope: Unsupervised Monitoring for AI Misbehaviors
por: Zhong, Ziqian, et al.
Publicado: (2026)
por: Zhong, Ziqian, et al.
Publicado: (2026)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
por: Korbak, Tomek, et al.
Publicado: (2025)
por: Korbak, Tomek, et al.
Publicado: (2025)
Small-to-Large Generalization: Data Influences Models Consistently Across Scale
por: Khaddaj, Alaa, et al.
Publicado: (2025)
por: Khaddaj, Alaa, et al.
Publicado: (2025)
DsDm: Model-Aware Dataset Selection with Datamodels
por: Engstrom, Logan, et al.
Publicado: (2024)
por: Engstrom, Logan, et al.
Publicado: (2024)
Decomposing and Editing Predictions by Modeling Model Computation
por: Shah, Harshay, et al.
Publicado: (2024)
por: Shah, Harshay, et al.
Publicado: (2024)
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
por: Yang, Shu, et al.
Publicado: (2026)
por: Yang, Shu, et al.
Publicado: (2026)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
por: Zolkowski, Artur, et al.
Publicado: (2025)
por: Zolkowski, Artur, et al.
Publicado: (2025)
Ask Your Distribution Shift if Pre-Training is Right for You
por: Cohen-Wang, Benjamin, et al.
Publicado: (2024)
por: Cohen-Wang, Benjamin, et al.
Publicado: (2024)
User Strategization and Trustworthy Algorithms
por: Cen, Sarah H., et al.
Publicado: (2023)
por: Cen, Sarah H., et al.
Publicado: (2023)
Do Large Language Model Benchmarks Test Reliability?
por: Vendrow, Joshua, et al.
Publicado: (2025)
por: Vendrow, Joshua, et al.
Publicado: (2025)
Learning to Attribute with Attention
por: Cohen-Wang, Benjamin, et al.
Publicado: (2025)
por: Cohen-Wang, Benjamin, et al.
Publicado: (2025)
De hand van Huizinga
por: Huizinga, Johan
Publicado: (2010)
por: Huizinga, Johan
Publicado: (2010)
Privatization, public investment, and capital income taxation / Harry Huizinga, Søren Bo Nielsen
por: Huizinga, Harry
Publicado: (1997)
por: Huizinga, Harry
Publicado: (1997)
Homo ludens / Johan Huizinga ; traductor Eugenio Imaz
por: Huizinga, Johan
por: Huizinga, Johan
El otoño de la edad media : estudios sobre la forma de la vida y del espíritu durante los siglos XIV y XV en Francia en los Países Bajos / Johan Huizinga ; versión española de José Gaos
por: Huizinga, Johan
Publicado: (1981)
por: Huizinga, Johan
Publicado: (1981)
El elemento estético de las representaciones históricas
por: Johan Huizinga
Publicado: (2005)
por: Johan Huizinga
Publicado: (2005)
Existe uma metamorfose da História? Resposta à pergunta: como o presente se torna passado? (Berliner Tageblatt, 31 de maio de 1936)
por: Johan Huizinga
Publicado: (2015)
por: Johan Huizinga
Publicado: (2015)
ContextCite: Attributing Model Generation to Context
por: Cohen-Wang, Benjamin, et al.
Publicado: (2024)
por: Cohen-Wang, Benjamin, et al.
Publicado: (2024)
The modal theory of linear orders
por: Wołoszyn, Wojciech Aleksander
Publicado: (2026)
por: Wołoszyn, Wojciech Aleksander
Publicado: (2026)
Modal group theory
por: Wołoszyn, Wojciech Aleksander
Publicado: (2026)
por: Wołoszyn, Wojciech Aleksander
Publicado: (2026)
The modal theory of the category of sets
por: Wołoszyn, Wojciech Aleksander
Publicado: (2026)
por: Wołoszyn, Wojciech Aleksander
Publicado: (2026)
Grzegorczyk Logic Unlocked
por: Wołoszyn, Wojciech Aleksander
Publicado: (2025)
por: Wołoszyn, Wojciech Aleksander
Publicado: (2025)
Modal group theory: homomorphisms
por: Wołoszyn, Wojciech Aleksander
Publicado: (2026)
por: Wołoszyn, Wojciech Aleksander
Publicado: (2026)
Atoms and Molecules as Quantum Attosecond Processors
por: Farhi, Asaf
Publicado: (2025)
por: Farhi, Asaf
Publicado: (2025)
On refinements of two-term Machin-like formulas
por: Farhi, Bakir
Publicado: (2026)
por: Farhi, Bakir
Publicado: (2026)
Expansion of a bivariate symmetric mean in the neighborhood of the first bisector
por: Farhi, Bakir
Publicado: (2025)
por: Farhi, Bakir
Publicado: (2025)
R2L: Reliable Reinforcement Learning: Guaranteed Return & Reliable Policies in Reinforcement Learning
por: Farhi, Nadir
Publicado: (2025)
por: Farhi, Nadir
Publicado: (2025)
A measure of intelligence of an approximation of a real number in a given model
por: Farhi, Bakir
Publicado: (2017)
por: Farhi, Bakir
Publicado: (2017)
$q$-analogues of sums of consecutive powers of natural numbers and extended Carlitz $q$-Bernoulli numbers and polynomials
por: Farhi, Bakir
Publicado: (2025)
por: Farhi, Bakir
Publicado: (2025)
New formulas involving Bernoulli and Stirling numbers of both kinds
por: Farhi, Bakir
Publicado: (2024)
por: Farhi, Bakir
Publicado: (2024)
Pilgrim interaction with services provided by the general presidency of Alharamain affairs
por: F. Farhi
Publicado: (2020)
por: F. Farhi
Publicado: (2020)
Training Agents to Self-Report Misbehavior
por: Lee, Bruce W., et al.
Publicado: (2026)
por: Lee, Bruce W., et al.
Publicado: (2026)
Wink: Recovering from Misbehaviors in Coding Agents
por: Nanda, Rahul, et al.
Publicado: (2026)
por: Nanda, Rahul, et al.
Publicado: (2026)
Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
por: Turtayev, Rustem, et al.
Publicado: (2025)
por: Turtayev, Rustem, et al.
Publicado: (2025)
LLMScan: Causal Scan for LLM Misbehavior Detection
por: Zhang, Mengdi, et al.
Publicado: (2024)
por: Zhang, Mengdi, et al.
Publicado: (2024)
Measuring Strategization in Recommendation: Users Adapt Their Behavior to Shape Future Content
por: Cen, Sarah H., et al.
Publicado: (2024)
por: Cen, Sarah H., et al.
Publicado: (2024)
Training on Documents About Monitoring Leads to CoT Obfuscation
por: Haskins, Reilly, et al.
Publicado: (2026)
por: Haskins, Reilly, et al.
Publicado: (2026)
Misbehavior Forecasting for Focused Autonomous Driving Systems Testing
por: Naziri, M M Abid, et al.
Publicado: (2025)
por: Naziri, M M Abid, et al.
Publicado: (2025)
Communism – Legitimacy – Nationalism
por: Zaremba, Marcin
Publicado: (2024)
por: Zaremba, Marcin
Publicado: (2024)
Ejemplares similares
-
Monitoring Monitorability
por: Guan, Melody Y., et al.
Publicado: (2025) -
Hodoscope: Unsupervised Monitoring for AI Misbehaviors
por: Zhong, Ziqian, et al.
Publicado: (2026) -
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
por: Korbak, Tomek, et al.
Publicado: (2025) -
Small-to-Large Generalization: Data Influences Models Consistently Across Scale
por: Khaddaj, Alaa, et al.
Publicado: (2025) -
DsDm: Model-Aware Dataset Selection with Datamodels
por: Engstrom, Logan, et al.
Publicado: (2024)