Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
Fuente:
arXiv
Saved in:
| Main Authors: | Ji, Zimo, Li, Zongjie, Jiang, Wenyuan, Gao, Yudong, Wang, Shuai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Empirical Study of Code Large Language Models for Binary Security Patch Detection
by: Li, Qingyuan, et al.
Published: (2025)
by: Li, Qingyuan, et al.
Published: (2025)
Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?
by: Qian, Yi, et al.
Published: (2026)
by: Qian, Yi, et al.
Published: (2026)
Poisoned Identifiers Survive LLM Deobfuscation: A Case Study on Claude Opus 4.6
by: Lorenzo, Luis Guzmán
Published: (2026)
by: Lorenzo, Luis Guzmán
Published: (2026)
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review
by: Mitropoulos, Dimitris, et al.
Published: (2026)
by: Mitropoulos, Dimitris, et al.
Published: (2026)
CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming Solutions
by: Shi, Jingwei, et al.
Published: (2026)
by: Shi, Jingwei, et al.
Published: (2026)
SecCodeBench-V2 Technical Report
by: Chen, Longfei, et al.
Published: (2026)
by: Chen, Longfei, et al.
Published: (2026)
Testing Storage-System Correctness: Challenges, Fuzzing Limitations, and AI-Augmented Opportunities
by: Wang, Ying, et al.
Published: (2026)
by: Wang, Ying, et al.
Published: (2026)
AutoStub: Genetic Programming-Based Stub Creation for Symbolic Execution
by: Mächtle, Felix, et al.
Published: (2025)
by: Mächtle, Felix, et al.
Published: (2025)
PatUntrack: Automated Generating Patch Examples for Issue Reports without Tracked Insecure Code
by: Jiang, Ziyou, et al.
Published: (2024)
by: Jiang, Ziyou, et al.
Published: (2024)
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
by: Hu, Junze, et al.
Published: (2025)
by: Hu, Junze, et al.
Published: (2025)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026)
by: Yang, Ruozhao, et al.
Published: (2026)
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Protect Your Secrets: Understanding and Measuring Data Exposure in VSCode Extensions
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
Towards Privacy-Preserving Code Generation: Differentially Private Code Language Models
by: Catal, Melih, et al.
Published: (2025)
by: Catal, Melih, et al.
Published: (2025)
SPDZCoder: Combining Expert Knowledge with LLMs for Generating Privacy-Computing Code
by: Dong, Xiaoning, et al.
Published: (2024)
by: Dong, Xiaoning, et al.
Published: (2024)
Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
by: Patir, Rupam, et al.
Published: (2025)
by: Patir, Rupam, et al.
Published: (2025)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
by: Wang, Su, et al.
Published: (2026)
by: Wang, Su, et al.
Published: (2026)
DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation
by: Huang, Li, et al.
Published: (2026)
by: Huang, Li, et al.
Published: (2026)
Learning to Generate Secure Code via Token-Level Rewards
by: Quan, Jiazheng, et al.
Published: (2026)
by: Quan, Jiazheng, et al.
Published: (2026)
MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
We Urgently Need Privilege Management in MCP: A Measurement of API Usage in MCP Ecosystems
by: Li, Zhihao, et al.
Published: (2025)
by: Li, Zhihao, et al.
Published: (2025)
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023)
by: Al-Kaswan, Ali, et al.
Published: (2023)
Reflection-Driven Control for Trustworthy Code Agents
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
AI Code Generators for Security: Friend or Foe?
by: Natella, Roberto, et al.
Published: (2024)
by: Natella, Roberto, et al.
Published: (2024)
Rethinking and Exploring String-Based Malware Family Classification in the Era of LLMs and RAG
by: Chen, Yufan, et al.
Published: (2025)
by: Chen, Yufan, et al.
Published: (2025)
VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Security of LLM-generated Code: A Comparative Analysis
by: Morkonda, Srivathsan G, et al.
Published: (2026)
by: Morkonda, Srivathsan G, et al.
Published: (2026)
Harnessing the Power of LLMs in Source Code Vulnerability Detection
by: Mahyari, Andrew A
Published: (2024)
by: Mahyari, Andrew A
Published: (2024)
Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
by: Gagnon, Charles E., et al.
Published: (2025)
by: Gagnon, Charles E., et al.
Published: (2025)
An Empirical Study of Code Obfuscation Practices in the Google Play Store
by: Niroshan, Akila, et al.
Published: (2025)
by: Niroshan, Akila, et al.
Published: (2025)
Cochise: A Reference Harness for Autonomous Penetration Testing
by: Happe, Andreas, et al.
Published: (2026)
by: Happe, Andreas, et al.
Published: (2026)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
Beyond Trusting Trust: Multi-Model Validation for Robust Code Generation
by: McDanel, Bradley
Published: (2025)
by: McDanel, Bradley
Published: (2025)
Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models
by: Bruni, Marc, et al.
Published: (2025)
by: Bruni, Marc, et al.
Published: (2025)
DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation
by: Pathak, Abhijeet, et al.
Published: (2025)
by: Pathak, Abhijeet, et al.
Published: (2025)
From Detection to Prevention: Explaining Security-Critical Code to Avoid Vulnerabilities
by: Krishnamurthy, Ranjith, et al.
Published: (2026)
by: Krishnamurthy, Ranjith, et al.
Published: (2026)
Similar Items
-
Empirical Study of Code Large Language Models for Binary Security Patch Detection
by: Li, Qingyuan, et al.
Published: (2025) -
Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?
by: Qian, Yi, et al.
Published: (2026) -
Poisoned Identifiers Survive LLM Deobfuscation: A Case Study on Claude Opus 4.6
by: Lorenzo, Luis Guzmán
Published: (2026) -
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
by: Steinberg, Jonathan, et al.
Published: (2026) -
Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review
by: Mitropoulos, Dimitris, et al.
Published: (2026)