The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind
Fuente:
arXiv
Saved in:
| Main Authors: | DeLeeuw, Caleb, Chawla, Gaurav, Sharma, Aniket, Dietze, Vanessa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
by: DeLeeuw, Caleb
Published: (2026)
by: DeLeeuw, Caleb
Published: (2026)
Real-World En Call Center Transcripts Dataset with PII Redaction
by: Dao, Ha, et al.
Published: (2025)
by: Dao, Ha, et al.
Published: (2025)
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Autospeculation
by: Hu, Hengyuan, et al.
Published: (2025)
by: Hu, Hengyuan, et al.
Published: (2025)
Position: Model Collapse Does Not Mean What You Think
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power
by: LaCroix, Travis, et al.
Published: (2026)
by: LaCroix, Travis, et al.
Published: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
Improving the Fairness of Deep-Learning, Short-term Crime Prediction with Under-reporting-aware Models
by: Wu, Jiahui, et al.
Published: (2024)
by: Wu, Jiahui, et al.
Published: (2024)
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025)
by: Haldar, Rajdeep, et al.
Published: (2025)
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)
by: Barez, Fazl, et al.
Published: (2025)
Defining and Evaluating Physical Safety for Large Language Models
by: Tang, Yung-Chen, et al.
Published: (2024)
by: Tang, Yung-Chen, et al.
Published: (2024)
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Safety challenges of AI in medicine in the era of large language models
by: Wang, Xiaoye, et al.
Published: (2024)
by: Wang, Xiaoye, et al.
Published: (2024)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
by: Bowen, Dillon, et al.
Published: (2025)
by: Bowen, Dillon, et al.
Published: (2025)
Faster Results from a Smarter Schedule: Reframing Collegiate Cross Country through Analysis of the National Running Club Database
by: Karr Jr, Jonathan A., et al.
Published: (2025)
by: Karr Jr, Jonathan A., et al.
Published: (2025)
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
by: Peng, Xiyue, et al.
Published: (2024)
by: Peng, Xiyue, et al.
Published: (2024)
AI Fairness Beyond Complete Demographics: Current Achievements and Future Directions
by: Wang, Zichong, et al.
Published: (2025)
by: Wang, Zichong, et al.
Published: (2025)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
by: Zhang, Eva, et al.
Published: (2024)
by: Zhang, Eva, et al.
Published: (2024)
Beyond Algorithmic Fairness: A Guide to Develop and Deploy Ethical AI-Enabled Decision-Support Tools
by: Gonzalez, Rosemarie Santa, et al.
Published: (2024)
by: Gonzalez, Rosemarie Santa, et al.
Published: (2024)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
by: Chen, Yimeng, et al.
Published: (2025)
by: Chen, Yimeng, et al.
Published: (2025)
Quantifying Prediction Consistency Under Fine-Tuning Multiplicity in Tabular LLMs
by: Hamman, Faisal, et al.
Published: (2024)
by: Hamman, Faisal, et al.
Published: (2024)
DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management
by: Cardei, Maria Ana, et al.
Published: (2025)
by: Cardei, Maria Ana, et al.
Published: (2025)
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
by: Potham, Ram
Published: (2025)
by: Potham, Ram
Published: (2025)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Green AI in Action: Strategic Model Selection for Ensembles in Production
by: Nijkamp, Nienke, et al.
Published: (2024)
by: Nijkamp, Nienke, et al.
Published: (2024)
Explainable AI for Mental Health Emergency Returns: Integrating LLMs with Predictive Modeling
by: Ahmed, Abdulaziz, et al.
Published: (2025)
by: Ahmed, Abdulaziz, et al.
Published: (2025)
Automated Feedback in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
by: Baral, Sami, et al.
Published: (2024)
by: Baral, Sami, et al.
Published: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
Breaking the ICE: Exploring promises and challenges of benchmarks for Inference Carbon & Energy estimation for LLMs
by: Sikand, Samarth, et al.
Published: (2025)
by: Sikand, Samarth, et al.
Published: (2025)
Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
by: Radmehr, Bahar, et al.
Published: (2024)
by: Radmehr, Bahar, et al.
Published: (2024)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Cost Efficient Fairness Audit Under Partial Feedback
by: Das, Nirjhar, et al.
Published: (2025)
by: Das, Nirjhar, et al.
Published: (2025)
Patentformer: A demonstration of AI-assisted automated patent drafting
by: Mudhiganti, Sai Krishna Reddy, et al.
Published: (2025)
by: Mudhiganti, Sai Krishna Reddy, et al.
Published: (2025)
AI and Remote Sensing for Resilient and Sustainable Built Environments: A Review of Current Methods, Open Data and Future Directions
by: Joulani, Ubada El, et al.
Published: (2025)
by: Joulani, Ubada El, et al.
Published: (2025)
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
When the Domain Expert Has No Time and the LLM Developer Has No Clinical Expertise: Real-World Lessons from LLM Co-Design in a Safety-Net Hospital
by: Kothari, Avni, et al.
Published: (2025)
by: Kothari, Avni, et al.
Published: (2025)
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
by: Verma, Richa, et al.
Published: (2026)
by: Verma, Richa, et al.
Published: (2026)
Similar Items
-
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
by: DeLeeuw, Caleb
Published: (2026) -
Real-World En Call Center Transcripts Dataset with PII Redaction
by: Dao, Ha, et al.
Published: (2025) -
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Autospeculation
by: Hu, Hengyuan, et al.
Published: (2025) -
Position: Model Collapse Does Not Mean What You Think
by: Schaeffer, Rylan, et al.
Published: (2025) -
Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power
by: LaCroix, Travis, et al.
Published: (2026)