Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies
Fuente:
arXiv
Guardado en:
| Autor principal: | Mittal, Avni |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
por: Mittal, Avni
Publicado: (2026)
por: Mittal, Avni
Publicado: (2026)
Can LLMs Follow Simple Rules?
por: Mu, Norman, et al.
Publicado: (2023)
por: Mu, Norman, et al.
Publicado: (2023)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
por: Mayne, Harry, et al.
Publicado: (2025)
por: Mayne, Harry, et al.
Publicado: (2025)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
por: Jeong, Hyejun, et al.
Publicado: (2024)
por: Jeong, Hyejun, et al.
Publicado: (2024)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
por: Wang, Chenyang, et al.
Publicado: (2025)
por: Wang, Chenyang, et al.
Publicado: (2025)
PLDR-LLMs Learn A Generalizable Tensor Operator That Can Replace Its Own Deep Neural Net At Inference
por: Gokden, Burc
Publicado: (2025)
por: Gokden, Burc
Publicado: (2025)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
por: Mittal, Avni, et al.
Publicado: (2026)
por: Mittal, Avni, et al.
Publicado: (2026)
LongSafety: Enhance Safety for Long-Context LLMs
por: Huang, Mianqiu, et al.
Publicado: (2024)
por: Huang, Mianqiu, et al.
Publicado: (2024)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
por: Zhao, Hao, et al.
Publicado: (2024)
por: Zhao, Hao, et al.
Publicado: (2024)
Training Language Models to Explain Their Own Computations
por: Li, Belinda Z., et al.
Publicado: (2025)
por: Li, Belinda Z., et al.
Publicado: (2025)
Language Models Can Predict Their Own Behavior
por: Ashok, Dhananjay, et al.
Publicado: (2025)
por: Ashok, Dhananjay, et al.
Publicado: (2025)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
por: Choi, Yunho, et al.
Publicado: (2026)
por: Choi, Yunho, et al.
Publicado: (2026)
Bring Your Own KG: Self-Supervised Program Synthesis for Zero-Shot KGQA
por: Agarwal, Dhruv, et al.
Publicado: (2023)
por: Agarwal, Dhruv, et al.
Publicado: (2023)
Do Multilingual LLMs Think In English?
por: Schut, Lisa, et al.
Publicado: (2025)
por: Schut, Lisa, et al.
Publicado: (2025)
Why Do Safety Guardrails Degrade Across Languages?
por: Zhang, Max, et al.
Publicado: (2026)
por: Zhang, Max, et al.
Publicado: (2026)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
por: Garcia-Gasulla, Dario, et al.
Publicado: (2025)
por: Garcia-Gasulla, Dario, et al.
Publicado: (2025)
FALCON: Autonomous Cyber Threat Intelligence Mining with LLMs for IDS Rule Generation
por: Mitra, Shaswata, et al.
Publicado: (2025)
por: Mitra, Shaswata, et al.
Publicado: (2025)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
por: Yang, Xikang, et al.
Publicado: (2025)
por: Yang, Xikang, et al.
Publicado: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
por: Ren, Richard, et al.
Publicado: (2024)
por: Ren, Richard, et al.
Publicado: (2024)
Do LLMs Encode Functional Importance of Reasoning Tokens?
por: Singh, Janvijay, et al.
Publicado: (2026)
por: Singh, Janvijay, et al.
Publicado: (2026)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
por: Wei, Boyi, et al.
Publicado: (2024)
por: Wei, Boyi, et al.
Publicado: (2024)
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
por: Guliani, Keerat, et al.
Publicado: (2026)
por: Guliani, Keerat, et al.
Publicado: (2026)
Multilingual Safety Alignment via Self-Distillation
por: Qin, Ruiyang, et al.
Publicado: (2026)
por: Qin, Ruiyang, et al.
Publicado: (2026)
How Likely Do LLMs with CoT Mimic Human Reasoning?
por: Bao, Guangsheng, et al.
Publicado: (2024)
por: Bao, Guangsheng, et al.
Publicado: (2024)
Transformer-Squared: Self-adaptive LLMs
por: Sun, Qi, et al.
Publicado: (2025)
por: Sun, Qi, et al.
Publicado: (2025)
(How) Do Language Models Track State?
por: Li, Belinda Z., et al.
Publicado: (2025)
por: Li, Belinda Z., et al.
Publicado: (2025)
Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
por: Jan, Essa, et al.
Publicado: (2024)
por: Jan, Essa, et al.
Publicado: (2024)
Do LLMs Benefit From Their Own Words?
por: Huang, Jenny Y., et al.
Publicado: (2026)
por: Huang, Jenny Y., et al.
Publicado: (2026)
Questionnaire Responses Do not Capture the Safety of AI Agents
por: Hellrigel-Holderbaum, Max, et al.
Publicado: (2026)
por: Hellrigel-Holderbaum, Max, et al.
Publicado: (2026)
Rule by Rule: Learning with Confidence through Vocabulary Expansion
por: Nössig, Albert, et al.
Publicado: (2024)
por: Nössig, Albert, et al.
Publicado: (2024)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
por: Yang, Zhihe, et al.
Publicado: (2025)
por: Yang, Zhihe, et al.
Publicado: (2025)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
por: Xu, Tianyang, et al.
Publicado: (2024)
por: Xu, Tianyang, et al.
Publicado: (2024)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
por: Xi, Zhiheng, et al.
Publicado: (2025)
por: Xi, Zhiheng, et al.
Publicado: (2025)
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
por: Xia, Wei, et al.
Publicado: (2026)
por: Xia, Wei, et al.
Publicado: (2026)
Auditing language models for hidden objectives
por: Marks, Samuel, et al.
Publicado: (2025)
por: Marks, Samuel, et al.
Publicado: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
por: Chaudhary, Maheep, et al.
Publicado: (2025)
por: Chaudhary, Maheep, et al.
Publicado: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
por: Wu, Xuansheng, et al.
Publicado: (2023)
por: Wu, Xuansheng, et al.
Publicado: (2023)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
por: Mohammadi, Seyedali, et al.
Publicado: (2025)
por: Mohammadi, Seyedali, et al.
Publicado: (2025)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
por: Kim, Minsang, et al.
Publicado: (2026)
por: Kim, Minsang, et al.
Publicado: (2026)
Ejemplares similares
-
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
por: Mittal, Avni
Publicado: (2026) -
Can LLMs Follow Simple Rules?
por: Mu, Norman, et al.
Publicado: (2023) -
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
por: Mayne, Harry, et al.
Publicado: (2025) -
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
por: Jeong, Hyejun, et al.
Publicado: (2024) -
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
por: Wang, Chenyang, et al.
Publicado: (2025)