Calibrating Conservatism for Scalable Oversight
Fuente:
arXiv
Saved in:
| Main Authors: | Overman, William, Bayati, Mohsen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
by: Overman, William, et al.
Published: (2025)
by: Overman, William, et al.
Published: (2025)
Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
by: Overman, William, et al.
Published: (2025)
by: Overman, William, et al.
Published: (2025)
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
by: Overman, William, et al.
Published: (2026)
by: Overman, William, et al.
Published: (2026)
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
by: Ye, Junze, et al.
Published: (2025)
by: Ye, Junze, et al.
Published: (2025)
Simulating and Experimenting with Social Media Mobilization Using LLM Agents
by: Shirani, Sadegh, et al.
Published: (2025)
by: Shirani, Sadegh, et al.
Published: (2025)
A Benchmark for Scalable Oversight Protocols
by: Sudhir, Abhimanyu Pallavi, et al.
Published: (2025)
by: Sudhir, Abhimanyu Pallavi, et al.
Published: (2025)
Causal Effects with Unobserved Unit Types in Interacting Human-AI Systems
by: Overman, William, et al.
Published: (2026)
by: Overman, William, et al.
Published: (2026)
Scaling Laws For Scalable Oversight
by: Engels, Joshua, et al.
Published: (2025)
by: Engels, Joshua, et al.
Published: (2025)
Steering LLMs via Scalable Interactive Oversight
by: Zhou, Enyu, et al.
Published: (2026)
by: Zhou, Enyu, et al.
Published: (2026)
Aligning Model Properties via Conformal Risk Control
by: Overman, William, et al.
Published: (2024)
by: Overman, William, et al.
Published: (2024)
Modeling Human Beliefs about AI Behavior for Scalable Oversight
by: Lang, Leon, et al.
Published: (2025)
by: Lang, Leon, et al.
Published: (2025)
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
Towards Scalable Oversight via Partitioned Human Supervision
by: Yin, Ren, et al.
Published: (2025)
by: Yin, Ren, et al.
Published: (2025)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
by: Wen, Xueru, et al.
Published: (2025)
by: Wen, Xueru, et al.
Published: (2025)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
by: Recchia, Gabriel, et al.
Published: (2025)
by: Recchia, Gabriel, et al.
Published: (2025)
Analysis of Thompson Sampling for Controlling Unknown Linear Diffusion Processes
by: Faradonbeh, Mohamad Kazem Shirani, et al.
Published: (2022)
by: Faradonbeh, Mohamad Kazem Shirani, et al.
Published: (2022)
Quantile Regression with Large Language Models for Price Prediction
by: Vedula, Nikhita, et al.
Published: (2025)
by: Vedula, Nikhita, et al.
Published: (2025)
Sparsity-based Safety Conservatism for Constrained Offline Reinforcement Learning
by: Cho, Minjae, et al.
Published: (2024)
by: Cho, Minjae, et al.
Published: (2024)
Compositional Conservatism: A Transductive Approach in Offline Reinforcement Learning
by: Song, Yeda, et al.
Published: (2024)
by: Song, Yeda, et al.
Published: (2024)
Higher-Order Causal Message Passing for Experimentation with Complex Interference
by: Bayati, Mohsen, et al.
Published: (2024)
by: Bayati, Mohsen, et al.
Published: (2024)
Can We Validate Counterfactual Estimations in the Presence of General Network Interference?
by: Shirani, Sadegh, et al.
Published: (2025)
by: Shirani, Sadegh, et al.
Published: (2025)
FormalJudge: A Neuro-Symbolic Paradigm for Agentic Oversight
by: Zhou, Jiayi, et al.
Published: (2026)
by: Zhou, Jiayi, et al.
Published: (2026)
Scalable Utility-Aware Multiclass Calibration
by: Hegazy, Mahmoud, et al.
Published: (2025)
by: Hegazy, Mahmoud, et al.
Published: (2025)
Oversight Structures for Agentic AI in Public-Sector Organizations
by: Schmitz, Chris, et al.
Published: (2025)
by: Schmitz, Chris, et al.
Published: (2025)
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight
by: Cui, Christopher Z., et al.
Published: (2026)
by: Cui, Christopher Z., et al.
Published: (2026)
Unbiasing on the Fly: Explanation-Guided Human Oversight of Machine Learning System Decisions
by: Mamman, Hussaini, et al.
Published: (2024)
by: Mamman, Hussaini, et al.
Published: (2024)
Tiered Agentic Oversight: A Hierarchical Multi-Agent System for Healthcare Safety
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
by: Beltoft, Stine Lyngsø, et al.
Published: (2026)
by: Beltoft, Stine Lyngsø, et al.
Published: (2026)
AI and Human Oversight: A Risk-Based Framework for Alignment
by: Kandikatla, Laxmiraju, et al.
Published: (2025)
by: Kandikatla, Laxmiraju, et al.
Published: (2025)
Overseeing Agents Without Constant Oversight: Challenges and Opportunities
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2026)
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2026)
Human-AI Complementarity: A Goal for Amplified Oversight
by: Jain, Rishub, et al.
Published: (2025)
by: Jain, Rishub, et al.
Published: (2025)
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
by: Lermen, Simon, et al.
Published: (2025)
by: Lermen, Simon, et al.
Published: (2025)
PICO: Secure Transformers via Robust Prompt Isolation and Cybersecurity Oversight
by: Goertzel, Ben, et al.
Published: (2025)
by: Goertzel, Ben, et al.
Published: (2025)
Great Models Think Alike and this Undermines AI Oversight
by: Goel, Shashwat, et al.
Published: (2025)
by: Goel, Shashwat, et al.
Published: (2025)
When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel
by: Li, Wenkai, et al.
Published: (2026)
by: Li, Wenkai, et al.
Published: (2026)
Quantifying Automation Risk in High-Automation AI Systems: A Bayesian Framework for Failure Propagation and Optimal Oversight
by: Srivastava, Vishal, et al.
Published: (2026)
by: Srivastava, Vishal, et al.
Published: (2026)
Hierarchical Pedagogical Oversight: A Multi-Agent Adversarial Framework for Reliable AI Tutoring
by: Sadhu, Saisab, et al.
Published: (2025)
by: Sadhu, Saisab, et al.
Published: (2025)
Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers
by: Li, Albus Yizhuo, et al.
Published: (2026)
by: Li, Albus Yizhuo, et al.
Published: (2026)
Exploring Moral Exercises for Human Oversight of AI systems: Insights from Three Pilot Studies
by: Crafa, Silvia, et al.
Published: (2025)
by: Crafa, Silvia, et al.
Published: (2025)
Intelligent support for Human Oversight: Integrating Reinforcement Learning with Gaze Simulation to Personalize Highlighting
by: Klößner, Thorsten, et al.
Published: (2026)
by: Klößner, Thorsten, et al.
Published: (2026)
Similar Items
-
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
by: Overman, William, et al.
Published: (2025) -
Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
by: Overman, William, et al.
Published: (2025) -
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
by: Overman, William, et al.
Published: (2026) -
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
by: Ye, Junze, et al.
Published: (2025) -
Simulating and Experimenting with Social Media Mobilization Using LLM Agents
by: Shirani, Sadegh, et al.
Published: (2025)