LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Chloe, Phuong, Mary, Siegel, Noah Y. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
di: Tice, Cameron, et al.
Pubblicazione: (2024)
di: Tice, Cameron, et al.
Pubblicazione: (2024)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
Benchmarking Misuse Mitigation Against Covert Adversaries
di: Brown, Davis, et al.
Pubblicazione: (2025)
di: Brown, Davis, et al.
Pubblicazione: (2025)
Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning
di: Chang, Zhiyuan, et al.
Pubblicazione: (2026)
di: Chang, Zhiyuan, et al.
Pubblicazione: (2026)
Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning
di: Draganov, Andrew, et al.
Pubblicazione: (2026)
di: Draganov, Andrew, et al.
Pubblicazione: (2026)
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
di: MacDermott, Matt, et al.
Pubblicazione: (2025)
di: MacDermott, Matt, et al.
Pubblicazione: (2025)
Output Supervision Can Obfuscate the Chain of Thought
di: Drori, Jacob, et al.
Pubblicazione: (2025)
di: Drori, Jacob, et al.
Pubblicazione: (2025)
SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration
di: Pan, Yu, et al.
Pubblicazione: (2026)
di: Pan, Yu, et al.
Pubblicazione: (2026)
An Application-Layer Multi-Modal Covert-Channel Reference Monitor for LLM Agent Egress
di: Metere, Alfredo
Pubblicazione: (2026)
di: Metere, Alfredo
Pubblicazione: (2026)
Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework
di: Zhou, Jiling, et al.
Pubblicazione: (2026)
di: Zhou, Jiling, et al.
Pubblicazione: (2026)
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
di: Wu, Daoyuan, et al.
Pubblicazione: (2024)
di: Wu, Daoyuan, et al.
Pubblicazione: (2024)
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
di: Ren, Zhenzhen, et al.
Pubblicazione: (2025)
di: Ren, Zhenzhen, et al.
Pubblicazione: (2025)
Evaluation of ChatGPT's Smart Contract Auditing Capabilities Based on Chain of Thought
di: Du, Yuying, et al.
Pubblicazione: (2024)
di: Du, Yuying, et al.
Pubblicazione: (2024)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
di: Liu, Shuaitong, et al.
Pubblicazione: (2025)
di: Liu, Shuaitong, et al.
Pubblicazione: (2025)
Large Language Model-driven Security Assistant for Internet of Things via Chain-of-Thought
di: Zeng, Mingfei, et al.
Pubblicazione: (2025)
di: Zeng, Mingfei, et al.
Pubblicazione: (2025)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
di: Lu, Yu-An, et al.
Pubblicazione: (2026)
di: Lu, Yu-An, et al.
Pubblicazione: (2026)
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
di: Yueh-Han, Chen, et al.
Pubblicazione: (2025)
di: Yueh-Han, Chen, et al.
Pubblicazione: (2025)
ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation
di: Jiang, Xiaochong, et al.
Pubblicazione: (2026)
di: Jiang, Xiaochong, et al.
Pubblicazione: (2026)
Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
di: Hammadia, Taha, et al.
Pubblicazione: (2026)
di: Hammadia, Taha, et al.
Pubblicazione: (2026)
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
di: Du, Hao, et al.
Pubblicazione: (2025)
di: Du, Hao, et al.
Pubblicazione: (2025)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
di: Najt, Elle, et al.
Pubblicazione: (2026)
di: Najt, Elle, et al.
Pubblicazione: (2026)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
Stealing AI Model Weights Through Covert Communication Channels
di: Barbaza, Valentin, et al.
Pubblicazione: (2025)
di: Barbaza, Valentin, et al.
Pubblicazione: (2025)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
di: Guan, Shaowei, et al.
Pubblicazione: (2025)
di: Guan, Shaowei, et al.
Pubblicazione: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
di: Ma, Haokai, et al.
Pubblicazione: (2025)
di: Ma, Haokai, et al.
Pubblicazione: (2025)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
di: Li, Siyuan, et al.
Pubblicazione: (2026)
di: Li, Siyuan, et al.
Pubblicazione: (2026)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
di: Deason, Lauren, et al.
Pubblicazione: (2025)
di: Deason, Lauren, et al.
Pubblicazione: (2025)
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
di: Rodriguez, Mikel, et al.
Pubblicazione: (2025)
di: Rodriguez, Mikel, et al.
Pubblicazione: (2025)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
ACF: A Collaborative Framework for Agent Covert Communication under Cognitive Asymmetry
di: Wu, Wansheng, et al.
Pubblicazione: (2026)
di: Wu, Wansheng, et al.
Pubblicazione: (2026)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
di: Liu, Xiaoqun, et al.
Pubblicazione: (2024)
di: Liu, Xiaoqun, et al.
Pubblicazione: (2024)
Can LLMs Make (Personalized) Access Control Decisions?
di: Groschupp, Friederike, et al.
Pubblicazione: (2025)
di: Groschupp, Friederike, et al.
Pubblicazione: (2025)
Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
di: Li, Qinfeng, et al.
Pubblicazione: (2025)
di: Li, Qinfeng, et al.
Pubblicazione: (2025)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
di: Wang, Xunguang, et al.
Pubblicazione: (2024)
di: Wang, Xunguang, et al.
Pubblicazione: (2024)
CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge Deployment
di: Li, Qinfeng, et al.
Pubblicazione: (2024)
di: Li, Qinfeng, et al.
Pubblicazione: (2024)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
di: Che, Zora, et al.
Pubblicazione: (2025)
di: Che, Zora, et al.
Pubblicazione: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
di: Black, Sid, et al.
Pubblicazione: (2025)
di: Black, Sid, et al.
Pubblicazione: (2025)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
di: Kouremetis, Michael, et al.
Pubblicazione: (2025)
di: Kouremetis, Michael, et al.
Pubblicazione: (2025)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
di: Liu, Zicheng, et al.
Pubblicazione: (2025)
di: Liu, Zicheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
di: Tice, Cameron, et al.
Pubblicazione: (2024) -
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
di: Zolkowski, Artur, et al.
Pubblicazione: (2025) -
Benchmarking Misuse Mitigation Against Covert Adversaries
di: Brown, Davis, et al.
Pubblicazione: (2025) -
Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning
di: Chang, Zhiyuan, et al.
Pubblicazione: (2026) -
Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning
di: Draganov, Andrew, et al.
Pubblicazione: (2026)