Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies
Fuente:
arXiv
Saved in:
| Main Author: | Mittal, Avni |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
by: Mittal, Avni
Published: (2026)
by: Mittal, Avni
Published: (2026)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
by: Jeong, Hyejun, et al.
Published: (2024)
by: Jeong, Hyejun, et al.
Published: (2024)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
by: Wang, Chenyang, et al.
Published: (2025)
by: Wang, Chenyang, et al.
Published: (2025)
PLDR-LLMs Learn A Generalizable Tensor Operator That Can Replace Its Own Deep Neural Net At Inference
by: Gokden, Burc
Published: (2025)
by: Gokden, Burc
Published: (2025)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
LongSafety: Enhance Safety for Long-Context LLMs
by: Huang, Mianqiu, et al.
Published: (2024)
by: Huang, Mianqiu, et al.
Published: (2024)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
Training Language Models to Explain Their Own Computations
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Language Models Can Predict Their Own Behavior
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
by: Choi, Yunho, et al.
Published: (2026)
by: Choi, Yunho, et al.
Published: (2026)
Bring Your Own KG: Self-Supervised Program Synthesis for Zero-Shot KGQA
by: Agarwal, Dhruv, et al.
Published: (2023)
by: Agarwal, Dhruv, et al.
Published: (2023)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
FALCON: Autonomous Cyber Threat Intelligence Mining with LLMs for IDS Rule Generation
by: Mitra, Shaswata, et al.
Published: (2025)
by: Mitra, Shaswata, et al.
Published: (2025)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
by: Yang, Xikang, et al.
Published: (2025)
by: Yang, Xikang, et al.
Published: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
Do LLMs Encode Functional Importance of Reasoning Tokens?
by: Singh, Janvijay, et al.
Published: (2026)
by: Singh, Janvijay, et al.
Published: (2026)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
by: Guliani, Keerat, et al.
Published: (2026)
by: Guliani, Keerat, et al.
Published: (2026)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
Transformer-Squared: Self-adaptive LLMs
by: Sun, Qi, et al.
Published: (2025)
by: Sun, Qi, et al.
Published: (2025)
(How) Do Language Models Track State?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
by: Jan, Essa, et al.
Published: (2024)
by: Jan, Essa, et al.
Published: (2024)
Do LLMs Benefit From Their Own Words?
by: Huang, Jenny Y., et al.
Published: (2026)
by: Huang, Jenny Y., et al.
Published: (2026)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
Rule by Rule: Learning with Confidence through Vocabulary Expansion
by: Nössig, Albert, et al.
Published: (2024)
by: Nössig, Albert, et al.
Published: (2024)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
by: Yang, Zhihe, et al.
Published: (2025)
by: Yang, Zhihe, et al.
Published: (2025)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
by: Xu, Tianyang, et al.
Published: (2024)
by: Xu, Tianyang, et al.
Published: (2024)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
by: Xia, Wei, et al.
Published: (2026)
by: Xia, Wei, et al.
Published: (2026)
Auditing language models for hidden objectives
by: Marks, Samuel, et al.
Published: (2025)
by: Marks, Samuel, et al.
Published: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
by: Kim, Minsang, et al.
Published: (2026)
by: Kim, Minsang, et al.
Published: (2026)
Similar Items
-
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
by: Mittal, Avni
Published: (2026) -
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023) -
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025) -
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
by: Jeong, Hyejun, et al.
Published: (2024) -
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
by: Wang, Chenyang, et al.
Published: (2025)