Gram: Assessing sabotage propensities via automated alignment auditing
Fuente:
arXiv
Saved in:
| Main Authors: | Lindner, David, Krakovna, Victoria, Farquhar, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Realistic honeypot evaluations for scheming propensity
by: Krakovna, Victoria, et al.
Published: (2026)
by: Krakovna, Victoria, et al.
Published: (2026)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
by: Farquhar, Sebastian, et al.
Published: (2025)
by: Farquhar, Sebastian, et al.
Published: (2025)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
CLUE: Neural Networks Calibration via Learning Uncertainty-Error alignment
by: Mendes, Pedro, et al.
Published: (2025)
by: Mendes, Pedro, et al.
Published: (2025)
Learning Safety Constraints from Demonstrations with Unknown Rewards
by: Lindner, David, et al.
Published: (2023)
by: Lindner, David, et al.
Published: (2023)
Scalable Meta-Learning via Mixed-Mode Differentiation
by: Kemaev, Iurii, et al.
Published: (2025)
by: Kemaev, Iurii, et al.
Published: (2025)
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
by: Cha, Taehun, et al.
Published: (2026)
by: Cha, Taehun, et al.
Published: (2026)
GRC-Net: Gram Residual Co-attention Net for epilepsy prediction
by: You, Bihao, et al.
Published: (2025)
by: You, Bihao, et al.
Published: (2025)
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
by: Kaufmann, Max, et al.
Published: (2026)
by: Kaufmann, Max, et al.
Published: (2026)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
FLoRG: Federated Fine-tuning with Low-rank Gram Matrices and Procrustes Alignment
by: Meng, Chuiyang, et al.
Published: (2026)
by: Meng, Chuiyang, et al.
Published: (2026)
Adaptive auditing of AI systems with anytime-valid guarantees
by: Zhou, Siyu, et al.
Published: (2026)
by: Zhou, Siyu, et al.
Published: (2026)
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
by: Rocamonde, Juan, et al.
Published: (2023)
by: Rocamonde, Juan, et al.
Published: (2023)
Large-Scale Multipurpose Benchmark Datasets For Assessing Data-Driven Deep Learning Approaches For Water Distribution Networks
by: Tello, Andres, et al.
Published: (2024)
by: Tello, Andres, et al.
Published: (2024)
Human-aligned Chess with a Bit of Search
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Neuro-inspired automated lens design
by: Gao, Yao, et al.
Published: (2025)
by: Gao, Yao, et al.
Published: (2025)
Knowledge distillation through geometry-aware representational alignment
by: Bhattarai, Prajjwal, et al.
Published: (2025)
by: Bhattarai, Prajjwal, et al.
Published: (2025)
Towards Stable Preferences for Stakeholder-aligned Machine Learning
by: Sheraz, Haleema, et al.
Published: (2024)
by: Sheraz, Haleema, et al.
Published: (2024)
ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
by: Ranganath, Aditya, et al.
Published: (2026)
by: Ranganath, Aditya, et al.
Published: (2026)
Ethics2vec: aligning automatic agents and human preferences
by: Bontempi, Gianluca
Published: (2025)
by: Bontempi, Gianluca
Published: (2025)
Understanding the performance gap between online and offline alignment algorithms
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Graders should cheat: privileged information enables expert-level automated evaluations
by: Zhou, Jin Peng, et al.
Published: (2025)
by: Zhou, Jin Peng, et al.
Published: (2025)
Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society
by: Hartmann, David, et al.
Published: (2024)
by: Hartmann, David, et al.
Published: (2024)
Learning-augmented robotic automation for real-world manufacturing
by: Kim, Yunho, et al.
Published: (2026)
by: Kim, Yunho, et al.
Published: (2026)
Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models
by: Jain, Vineet, et al.
Published: (2025)
by: Jain, Vineet, et al.
Published: (2025)
Fitness aligned structural modeling enables scalable virtual screening with AuroBind
by: Zhang, Zhongyue, et al.
Published: (2025)
by: Zhang, Zhongyue, et al.
Published: (2025)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
by: Fronsdal, Kai, et al.
Published: (2024)
by: Fronsdal, Kai, et al.
Published: (2024)
Quantum automated learning with provable and explainable trainability
by: Ye, Qi, et al.
Published: (2025)
by: Ye, Qi, et al.
Published: (2025)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
by: Tice, Cameron, et al.
Published: (2026)
by: Tice, Cameron, et al.
Published: (2026)
Quantifying stability of non-power-seeking in artificial agents
by: Gunter, Evan Ryan, et al.
Published: (2024)
by: Gunter, Evan Ryan, et al.
Published: (2024)
On Goodhart's law, with an application to value alignment
by: El-Mhamdi, El-Mahdi, et al.
Published: (2024)
by: El-Mhamdi, El-Mahdi, et al.
Published: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Getting aligned on representational alignment
by: Sucholutsky, Ilia, et al.
Published: (2023)
by: Sucholutsky, Ilia, et al.
Published: (2023)
A modular framework for automated evaluation of procedural content generation in serious games with deep reinforcement learning agents
by: Kalafatis, Eleftherios, et al.
Published: (2025)
by: Kalafatis, Eleftherios, et al.
Published: (2025)
Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning
by: Black, James R. M., et al.
Published: (2025)
by: Black, James R. M., et al.
Published: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Exploring the impact of traffic signal control and connected and automated vehicles on intersections safety: A deep reinforcement learning approach
by: Karbasi, Amir Hossein, et al.
Published: (2024)
by: Karbasi, Amir Hossein, et al.
Published: (2024)
Decoding Latent Spaces: Assessing the Interpretability of Time Series Foundation Models for Visual Analytics
by: Santamaria-Valenzuela, Inmaculada, et al.
Published: (2025)
by: Santamaria-Valenzuela, Inmaculada, et al.
Published: (2025)
Patentformer: A demonstration of AI-assisted automated patent drafting
by: Mudhiganti, Sai Krishna Reddy, et al.
Published: (2025)
by: Mudhiganti, Sai Krishna Reddy, et al.
Published: (2025)
Similar Items
-
Realistic honeypot evaluations for scheming propensity
by: Krakovna, Victoria, et al.
Published: (2026) -
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
by: Farquhar, Sebastian, et al.
Published: (2025) -
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025) -
CLUE: Neural Networks Calibration via Learning Uncertainty-Error alignment
by: Mendes, Pedro, et al.
Published: (2025) -
Learning Safety Constraints from Demonstrations with Unknown Rewards
by: Lindner, David, et al.
Published: (2023)