Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Xiaopeng, Yan, Ming, Zhou, Xiwen, Zhao, Chenlong, Wang, Suli, Zhang, Yong, Zhou, Joey Tianyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
A Semantic Invariant Robust Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023)
by: Liu, Aiwei, et al.
Published: (2023)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
by: Hill, Brennen, et al.
Published: (2025)
by: Hill, Brennen, et al.
Published: (2025)
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)
by: Wang, Yanshu, et al.
Published: (2026)
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
Logits of API-Protected LLMs Leak Proprietary Information
by: Finlayson, Matthew, et al.
Published: (2024)
by: Finlayson, Matthew, et al.
Published: (2024)
Monotonicity as an Architectural Bias for Robust Language Models
by: Cooper, Patrick, et al.
Published: (2026)
by: Cooper, Patrick, et al.
Published: (2026)
AVEC: Bootstrapping Privacy for Local LLMs
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
by: Anam, Rizal Khoirul
Published: (2025)
by: Anam, Rizal Khoirul
Published: (2025)
Dark LLMs: The Growing Threat of Unaligned AI Models
by: Fire, Michael, et al.
Published: (2025)
by: Fire, Michael, et al.
Published: (2025)
Generative AI Models: Opportunities and Risks for Industry and Authorities
by: Alt, Tobias, et al.
Published: (2024)
by: Alt, Tobias, et al.
Published: (2024)
From Attack Descriptions to Vulnerabilities: A Sentence Transformer-Based Approach
by: Othman, Refat, et al.
Published: (2025)
by: Othman, Refat, et al.
Published: (2025)
Exploiting Web Search Tools of AI Agents for Data Exfiltration
by: Rall, Dennis, et al.
Published: (2025)
by: Rall, Dennis, et al.
Published: (2025)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
by: Merves, Tyler H., et al.
Published: (2026)
by: Merves, Tyler H., et al.
Published: (2026)
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
by: Seo, Yeongbin, et al.
Published: (2025)
by: Seo, Yeongbin, et al.
Published: (2025)
BreakFun: Jailbreaking LLMs via Schema Exploitation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
by: Zhao, Lepeng, et al.
Published: (2026)
by: Zhao, Lepeng, et al.
Published: (2026)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments
by: Nathanson, Samuel, et al.
Published: (2025)
by: Nathanson, Samuel, et al.
Published: (2025)
CRUPL: A Semi-Supervised Cyber Attack Detection with Consistency Regularization and Uncertainty-aware Pseudo-Labeling in Smart Grid
by: Dash, Smruti P., et al.
Published: (2025)
by: Dash, Smruti P., et al.
Published: (2025)
A Method for Quantifying Human Risk and a Blueprint for LLM Integration
by: Canale, Giuseppe
Published: (2025)
by: Canale, Giuseppe
Published: (2025)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
by: Wang, Xinhai, et al.
Published: (2026)
by: Wang, Xinhai, et al.
Published: (2026)
TinyML NLP Scheme for Semantic Wireless Sentiment Classification with Privacy Preservation
by: Radwan, Ahmed Y., et al.
Published: (2024)
by: Radwan, Ahmed Y., et al.
Published: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
Nested Named Entity Recognition as Single-Pass Sequence Labeling
by: Muñoz-Ortiz, Alberto, et al.
Published: (2025)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2025)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023)
GSDFuse: Capturing Cognitive Inconsistencies from Multi-Dimensional Weak Signals in Social Media Steganalysis
by: Huang, Kaibo, et al.
Published: (2025)
by: Huang, Kaibo, et al.
Published: (2025)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
by: Buchner, Valentin Leonhard, et al.
Published: (2023)
by: Buchner, Valentin Leonhard, et al.
Published: (2023)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
by: Zhang, Zhehao, et al.
Published: (2025)
by: Zhang, Zhehao, et al.
Published: (2025)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
by: Johnson, Warren
Published: (2026)
by: Johnson, Warren
Published: (2026)
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
by: Johnson, Warren, et al.
Published: (2026)
by: Johnson, Warren, et al.
Published: (2026)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning
by: Qiu, Xiaoqi, et al.
Published: (2024)
by: Qiu, Xiaoqi, et al.
Published: (2024)
Similar Items
-
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
by: Liu, Aiwei, et al.
Published: (2024) -
MarkLLM: An Open-Source Toolkit for LLM Watermarking
by: Pan, Leyi, et al.
Published: (2024) -
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025) -
A Semantic Invariant Robust Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023) -
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
by: Hill, Brennen, et al.
Published: (2025)