Training Deliberative Monitors for Black-Box Scheming Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Sinha, Aditya, Naik, Akshat, Gillioz, Victor, Storf, Simon, Merkelbach, Kilian, Barton-Cooper, Rich, Højmark, Axel, Hobbhahn, Marius |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025)
by: Schoen, Bronson, et al.
Published: (2025)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
by: Højmark, Axel, et al.
Published: (2024)
by: Højmark, Axel, et al.
Published: (2024)
A Chain-of-Thought Approach to Semantic Query Categorization in e-Commerce Taxonomies
by: Duraj, Jetlir, et al.
Published: (2026)
by: Duraj, Jetlir, et al.
Published: (2026)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Large Language Models Often Know When They Are Being Evaluated
by: Needham, Joe, et al.
Published: (2025)
by: Needham, Joe, et al.
Published: (2025)
Correct Black-Box Monitors for Distributed Deadlock Detection: Formalisation and Implementation (Technical Report)
by: Rowicki, Radosław Jan, et al.
Published: (2025)
by: Rowicki, Radosław Jan, et al.
Published: (2025)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
Memory-Augmented Agent Training for Business Document Understanding
by: Liu, Jiale, et al.
Published: (2024)
by: Liu, Jiale, et al.
Published: (2024)
$k$NNProxy: Efficient Training-Free Proxy Alignment for Black-Box Zero-Shot LLM-Generated Text Detection
by: Wong, Kahim, et al.
Published: (2026)
by: Wong, Kahim, et al.
Published: (2026)
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
by: Rao, Abhinav, et al.
Published: (2023)
by: Rao, Abhinav, et al.
Published: (2023)
DORY: Deliberative Prompt Recovery for LLM
by: Gao, Lirong, et al.
Published: (2024)
by: Gao, Lirong, et al.
Published: (2024)
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
by: Cheng, Jiale, et al.
Published: (2023)
by: Cheng, Jiale, et al.
Published: (2023)
Inside the Black Box: Detecting Data Leakage in Pre-trained Language Encoders
by: Xin, Yuan, et al.
Published: (2024)
by: Xin, Yuan, et al.
Published: (2024)
Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
Will we run out of data? Limits of LLM scaling based on human-generated data
by: Villalobos, Pablo, et al.
Published: (2022)
by: Villalobos, Pablo, et al.
Published: (2022)
Does It Make Sense to Explain a Black Box With Another Black Box?
by: Delaunay, Julien, et al.
Published: (2024)
by: Delaunay, Julien, et al.
Published: (2024)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
by: Asawa, Parth, et al.
Published: (2025)
by: Asawa, Parth, et al.
Published: (2025)
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024)
by: Meinke, Alexander, et al.
Published: (2024)
You've Changed: Detecting Modification of Black-Box Large Language Models
by: Dima, Alden, et al.
Published: (2025)
by: Dima, Alden, et al.
Published: (2025)
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
by: Xue, Yihao, et al.
Published: (2025)
by: Xue, Yihao, et al.
Published: (2025)
Black-Box Segmentation of Electronic Medical Records
by: Yuan, Hongyi, et al.
Published: (2024)
by: Yuan, Hongyi, et al.
Published: (2024)
M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
The momentum-space conformal bootstrap in 2d
by: Gillioz, Marc
Published: (2025)
by: Gillioz, Marc
Published: (2025)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
by: Zhu, He, et al.
Published: (2026)
by: Zhu, He, et al.
Published: (2026)
Vbox: Efficient Black-Box Serializability Verification
by: Sun, Weihua, et al.
Published: (2025)
by: Sun, Weihua, et al.
Published: (2025)
Lowest Span Confidence: A Zero-Shot Metric for Efficient and Black-Box Hallucination Detection in LLMs
by: Qiao, Yitong, et al.
Published: (2026)
by: Qiao, Yitong, et al.
Published: (2026)
Knowledge Distillation of Black-Box Large Language Models
by: Chen, Hongzhan, et al.
Published: (2024)
by: Chen, Hongzhan, et al.
Published: (2024)
Universal statistical laws governing culinary design
by: Bagler, Ganesh, et al.
Published: (2026)
by: Bagler, Ganesh, et al.
Published: (2026)
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
by: Sinha, Aarush
Published: (2025)
by: Sinha, Aarush
Published: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
by: Sawczyn, Albert, et al.
Published: (2025)
by: Sawczyn, Albert, et al.
Published: (2025)
Black-Box Guardrail Reverse-engineering Attack
by: Yao, Hongwei, et al.
Published: (2025)
by: Yao, Hongwei, et al.
Published: (2025)
NaN-Propagation: A Novel Method for Sparsity Detection in Black-Box Computational Functions
by: Sharpe, Peter
Published: (2025)
by: Sharpe, Peter
Published: (2025)
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
by: Moslonka, Charles, et al.
Published: (2025)
by: Moslonka, Charles, et al.
Published: (2025)
M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
It Takes Two: A Dual Stage Approach for Terminology-Aware Translation
by: Jaswal, Akshat Singh
Published: (2025)
by: Jaswal, Akshat Singh
Published: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024)
by: Laine, Rudolf, et al.
Published: (2024)
SAC3: Reliable Hallucination Detection in Black-Box Language Models via Semantic-aware Cross-check Consistency
by: Zhang, Jiaxin, et al.
Published: (2023)
by: Zhang, Jiaxin, et al.
Published: (2023)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
by: Wang, Zhuoshang, et al.
Published: (2026)
by: Wang, Zhuoshang, et al.
Published: (2026)
DELTA: Deliberative Multi-Agent Reasoning with Reinforcement Learning for Multimodal Psychological Counseling
by: Yang, Jiangnan, et al.
Published: (2026)
by: Yang, Jiangnan, et al.
Published: (2026)
Similar Items
-
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025) -
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025) -
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
by: Højmark, Axel, et al.
Published: (2024) -
A Chain-of-Thought Approach to Semantic Query Categorization in e-Commerce Taxonomies
by: Duraj, Jetlir, et al.
Published: (2026) -
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)