Localizing Malicious Outputs from CodeLLM
Fuente:
arXiv
Saved in:
| Main Authors: | Borana, Mayukh, Liang, Junyi, Rajan, Sai Sathiesh, Chattopadhyay, Sudipta |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge-based Consistency Testing of Large Language Models
by: Rajan, Sai Sathiesh, et al.
Published: (2024)
by: Rajan, Sai Sathiesh, et al.
Published: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
by: Wahed, Muntasir, et al.
Published: (2025)
by: Wahed, Muntasir, et al.
Published: (2025)
Reformulation is All You Need: Addressing Malicious Text Features in DNNs
by: Jiang, Yi, et al.
Published: (2025)
by: Jiang, Yi, et al.
Published: (2025)
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
by: Rabin, Rafiqul, et al.
Published: (2025)
by: Rabin, Rafiqul, et al.
Published: (2025)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
by: Noah, Amit Finkman, et al.
Published: (2024)
by: Noah, Amit Finkman, et al.
Published: (2024)
TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
by: Sander, Tom, et al.
Published: (2026)
by: Sander, Tom, et al.
Published: (2026)
Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language Models
by: Huo, Mingjia, et al.
Published: (2024)
by: Huo, Mingjia, et al.
Published: (2024)
Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models
by: Zhang, Tianchen, et al.
Published: (2024)
by: Zhang, Tianchen, et al.
Published: (2024)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
by: Liang, Jiacheng, et al.
Published: (2024)
by: Liang, Jiacheng, et al.
Published: (2024)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
Watermarking Language Models with Error Correcting Codes
by: Chao, Patrick, et al.
Published: (2024)
by: Chao, Patrick, et al.
Published: (2024)
Localizing Paragraph Memorization in Language Models
by: Stoehr, Niklas, et al.
Published: (2024)
by: Stoehr, Niklas, et al.
Published: (2024)
LLM Unlearning Should Be Form-Independent
by: Ye, Xiaotian, et al.
Published: (2025)
by: Ye, Xiaotian, et al.
Published: (2025)
GCG Attack On A Diffusion LLM
by: Neyroud, Ruben, et al.
Published: (2025)
by: Neyroud, Ruben, et al.
Published: (2025)
Enhancing Robustness of AI Offensive Code Generators via Data Augmentation
by: Improta, Cristina, et al.
Published: (2023)
by: Improta, Cristina, et al.
Published: (2023)
LLMGuard: Guarding Against Unsafe LLM Behavior
by: Goyal, Shubh, et al.
Published: (2024)
by: Goyal, Shubh, et al.
Published: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
by: Assogba, Yannick, et al.
Published: (2026)
by: Assogba, Yannick, et al.
Published: (2026)
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
by: Zhang, Ruisi, et al.
Published: (2025)
by: Zhang, Ruisi, et al.
Published: (2025)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
by: Krishna, Arjun, et al.
Published: (2025)
by: Krishna, Arjun, et al.
Published: (2025)
PVMark: Enabling Public Verifiability for LLM Watermarking Schemes
by: Duan, Haohua, et al.
Published: (2025)
by: Duan, Haohua, et al.
Published: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
by: Liu, Ruixuan, et al.
Published: (2026)
by: Liu, Ruixuan, et al.
Published: (2026)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
by: Wang, Tianchun, et al.
Published: (2024)
by: Wang, Tianchun, et al.
Published: (2024)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
by: Zhang, Boyang, et al.
Published: (2026)
by: Zhang, Boyang, et al.
Published: (2026)
Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
by: Shafee, Samaneh, et al.
Published: (2024)
by: Shafee, Samaneh, et al.
Published: (2024)
LLM Dataset Inference: Did you train on my dataset?
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
by: Liu, Hongfu, et al.
Published: (2024)
by: Liu, Hongfu, et al.
Published: (2024)
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Proving membership in LLM pretraining data via data watermarks
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
Multi-Agent Systems Execute Arbitrary Malicious Code
by: Triedman, Harold, et al.
Published: (2025)
by: Triedman, Harold, et al.
Published: (2025)
Towards Quantum Machine Learning for Malicious Code Analysis
by: Lopez, Jesus, et al.
Published: (2025)
by: Lopez, Jesus, et al.
Published: (2025)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
by: Meeus, Matthieu, et al.
Published: (2025)
by: Meeus, Matthieu, et al.
Published: (2025)
No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
by: Pang, Qi, et al.
Published: (2024)
by: Pang, Qi, et al.
Published: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
by: Yang, Xikang, et al.
Published: (2024)
by: Yang, Xikang, et al.
Published: (2024)
WaterMax: breaking the LLM watermark detectability-robustness-quality trade-off
by: Giboulot, Eva, et al.
Published: (2024)
by: Giboulot, Eva, et al.
Published: (2024)
A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
by: Chen, Zhiyu, et al.
Published: (2024)
by: Chen, Zhiyu, et al.
Published: (2024)
Similar Items
-
Knowledge-based Consistency Testing of Large Language Models
by: Rajan, Sai Sathiesh, et al.
Published: (2024) -
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024) -
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
by: Wahed, Muntasir, et al.
Published: (2025) -
Reformulation is All You Need: Addressing Malicious Text Features in DNNs
by: Jiang, Yi, et al.
Published: (2025) -
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
by: Rabin, Rafiqul, et al.
Published: (2025)