Behavioural Analysis of Alignment Faking
Fuente:
arXiv
Saved in:
| Main Authors: | Hadida, Nathaniel Mitrani, Karty, Rhea, Williams-King, David, Cooney, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
by: Le, Linh, et al.
Published: (2026)
by: Le, Linh, et al.
Published: (2026)
xInv: Explainable Optimization of Inverse Problems
by: Memery, Sean, et al.
Published: (2025)
by: Memery, Sean, et al.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
by: Rath, Plawan Kumar, et al.
Published: (2026)
by: Rath, Plawan Kumar, et al.
Published: (2026)
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
by: Oliveira, Rafael C. T.
Published: (2026)
by: Oliveira, Rafael C. T.
Published: (2026)
Extreme AutoML: Analysis of Classification, Regression, and NLP Performance
by: Ratner, Edward, et al.
Published: (2024)
by: Ratner, Edward, et al.
Published: (2024)
Disposition Distillation at Small Scale: A Three-Arc Negative Result
by: Sadasivan, Hari
Published: (2026)
by: Sadasivan, Hari
Published: (2026)
Identity as Attractor: Geometric Evidence for Persistent Agent Architecture in LLM Activation Space
by: Vasilenko, Vladimir
Published: (2026)
by: Vasilenko, Vladimir
Published: (2026)
Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
by: Zhang, Zhiwei, et al.
Published: (2026)
by: Zhang, Zhiwei, et al.
Published: (2026)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
by: Vankov, Daniil, et al.
Published: (2026)
by: Vankov, Daniil, et al.
Published: (2026)
GNN for Structural Displacement Prediction
by: Chang, Hung-Fu, et al.
Published: (2026)
by: Chang, Hung-Fu, et al.
Published: (2026)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
by: Eldenk, Doğaç, et al.
Published: (2026)
by: Eldenk, Doğaç, et al.
Published: (2026)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
by: Mircea, Andrei, et al.
Published: (2025)
by: Mircea, Andrei, et al.
Published: (2025)
CoDA: Coding LM via Diffusion Adaptation
by: Chen, Haolin, et al.
Published: (2025)
by: Chen, Haolin, et al.
Published: (2025)
End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning
by: Chen, Guanzhong, et al.
Published: (2025)
by: Chen, Guanzhong, et al.
Published: (2025)
Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
by: Zhang, Zhengxin, et al.
Published: (2024)
by: Zhang, Zhengxin, et al.
Published: (2024)
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
by: Lasby, Mike, et al.
Published: (2025)
by: Lasby, Mike, et al.
Published: (2025)
Random Rule Forest (RRF): Interpretable Ensembles of LLM-Generated Questions for Predicting Startup Success
by: Griffin, Ben, et al.
Published: (2025)
by: Griffin, Ben, et al.
Published: (2025)
Beyond Memorization: Violating Privacy Via Inference with Large Language Models
by: Staab, Robin, et al.
Published: (2023)
by: Staab, Robin, et al.
Published: (2023)
Meta-Learning at Scale for Large Language Models via Low-Rank Amortized Bayesian Meta-Learning
by: Zhang, Liyi, et al.
Published: (2025)
by: Zhang, Liyi, et al.
Published: (2025)
Beyond the Black Box: A Statistical Model for LLM Reasoning and Inference
by: Dalal, Siddhartha, et al.
Published: (2024)
by: Dalal, Siddhartha, et al.
Published: (2024)
InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
by: Gao, Heyang, et al.
Published: (2025)
by: Gao, Heyang, et al.
Published: (2025)
Explainable AI for Smart Greenhouse Control: Interpretability of Temporal Fusion Transformer in the Internet of Robotic Things
by: Bashir, Muhammad Jawad, et al.
Published: (2025)
by: Bashir, Muhammad Jawad, et al.
Published: (2025)
Balancing Efficiency and Effectiveness: An LLM-Infused Approach for Optimized CTR Prediction
by: Zhang, Guoxiao, et al.
Published: (2024)
by: Zhang, Guoxiao, et al.
Published: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
CortexCompile: Harnessing Cortical-Inspired Architectures for Enhanced Multi-Agent NLP Code Synthesis
by: Ramachandran, Gautham, et al.
Published: (2024)
by: Ramachandran, Gautham, et al.
Published: (2024)
Solve it with EASE
by: Viktorin, Adam, et al.
Published: (2025)
by: Viktorin, Adam, et al.
Published: (2025)
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025)
by: Williams-King, David, et al.
Published: (2025)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
by: Zhang, Zhaowei, et al.
Published: (2026)
by: Zhang, Zhaowei, et al.
Published: (2026)
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
by: Vegner, Ivan, et al.
Published: (2025)
by: Vegner, Ivan, et al.
Published: (2025)
Going All-In on LLM Accuracy: Fake Prediction Markets, Real Confidence Signals
by: Todasco, Michael
Published: (2025)
by: Todasco, Michael
Published: (2025)
How Does Unfaithful Reasoning Emerge from Autoregressive Training? A Study of Synthetic Experiments
by: Wang, Fuxin, et al.
Published: (2026)
by: Wang, Fuxin, et al.
Published: (2026)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
by: Panuganti, Rajkiran
Published: (2026)
by: Panuganti, Rajkiran
Published: (2026)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
by: Deshmukh, Pratik, et al.
Published: (2026)
by: Deshmukh, Pratik, et al.
Published: (2026)
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
Similar Items
-
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
by: Le, Linh, et al.
Published: (2026) -
xInv: Explainable Optimization of Inverse Problems
by: Memery, Sean, et al.
Published: (2025) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023) -
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
by: Rath, Plawan Kumar, et al.
Published: (2026)