On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Bahar, Atmane Ayoub Mansour, Wazan, Ahmad Samer |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AVEC: Bootstrapping Privacy for Local LLMs
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)
by: Wang, Yanshu, et al.
Published: (2026)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
by: Gautam, Sushant, et al.
Published: (2026)
by: Gautam, Sushant, et al.
Published: (2026)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
by: Hill, Brennen, et al.
Published: (2025)
by: Hill, Brennen, et al.
Published: (2025)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
by: Merves, Tyler H., et al.
Published: (2026)
by: Merves, Tyler H., et al.
Published: (2026)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
Dark LLMs: The Growing Threat of Unaligned AI Models
by: Fire, Michael, et al.
Published: (2025)
by: Fire, Michael, et al.
Published: (2025)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
A Method for Quantifying Human Risk and a Blueprint for LLM Integration
by: Canale, Giuseppe
Published: (2025)
by: Canale, Giuseppe
Published: (2025)
AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
A Semantic Invariant Robust Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023)
by: Liu, Aiwei, et al.
Published: (2023)
BreakFun: Jailbreaking LLMs via Schema Exploitation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Watermarking for AI Content Detection: A Review on Text, Visual, and Audio Modalities
by: Cao, Lele
Published: (2025)
by: Cao, Lele
Published: (2025)
Exploiting Web Search Tools of AI Agents for Data Exfiltration
by: Rall, Dennis, et al.
Published: (2025)
by: Rall, Dennis, et al.
Published: (2025)
Generative AI Models: Opportunities and Risks for Industry and Authorities
by: Alt, Tobias, et al.
Published: (2024)
by: Alt, Tobias, et al.
Published: (2024)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
RAR: Setting Knowledge Tripwires for Retrieval Augmented Rejection
by: Buonocore, Tommaso Mario, et al.
Published: (2025)
by: Buonocore, Tommaso Mario, et al.
Published: (2025)
Securing Federated Sensitive Topic Classification against Poisoning Attacks
by: Chu, Tianyue, et al.
Published: (2022)
by: Chu, Tianyue, et al.
Published: (2022)
ChatGPT Based Data Augmentation for Improved Parameter-Efficient Debiasing of LLMs
by: Han, Pengrui, et al.
Published: (2024)
by: Han, Pengrui, et al.
Published: (2024)
Logits of API-Protected LLMs Leak Proprietary Information
by: Finlayson, Matthew, et al.
Published: (2024)
by: Finlayson, Matthew, et al.
Published: (2024)
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
by: Xie, Xiaopeng, et al.
Published: (2024)
by: Xie, Xiaopeng, et al.
Published: (2024)
Revisiting Third-Party Library Detection: A Ground Truth Dataset and Its Implications Across Security Tasks
by: Gu, Jintao, et al.
Published: (2025)
by: Gu, Jintao, et al.
Published: (2025)
Semantically Guided Adversarial Testing of Vision Models Using Language Models
by: Filus, Katarzyna, et al.
Published: (2025)
by: Filus, Katarzyna, et al.
Published: (2025)
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering
by: Stantchev, Vladimir
Published: (2026)
by: Stantchev, Vladimir
Published: (2026)
Monotonicity as an Architectural Bias for Robust Language Models
by: Cooper, Patrick, et al.
Published: (2026)
by: Cooper, Patrick, et al.
Published: (2026)
From Attack Descriptions to Vulnerabilities: A Sentence Transformer-Based Approach
by: Othman, Refat, et al.
Published: (2025)
by: Othman, Refat, et al.
Published: (2025)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
by: Nguyen, Minh Hoang, et al.
Published: (2026)
by: Nguyen, Minh Hoang, et al.
Published: (2026)
STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
by: Firc, Anton, et al.
Published: (2025)
by: Firc, Anton, et al.
Published: (2025)
Identity Deepfake Threats to Biometric Authentication Systems: Public and Expert Perspectives
by: He, Shijing, et al.
Published: (2025)
by: He, Shijing, et al.
Published: (2025)
Sure! Here's a short and concise title for your paper: "Contamination in Generated Text Detection Benchmarks"
by: Dingfelder, Philipp, et al.
Published: (2025)
by: Dingfelder, Philipp, et al.
Published: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
by: Zhao, Lepeng, et al.
Published: (2026)
by: Zhao, Lepeng, et al.
Published: (2026)
Beyond Human Judgment: A Bayesian Evaluation of LLMs' Moral Values Understanding
by: Skorski, Maciej, et al.
Published: (2025)
by: Skorski, Maciej, et al.
Published: (2025)
The Ethics Engine: A Modular Pipeline for Accessible Psychometric Assessment of Large Language Models
by: Van Clief, Jake, et al.
Published: (2025)
by: Van Clief, Jake, et al.
Published: (2025)
Similar Items
-
AVEC: Bootstrapping Privacy for Local LLMs
by: Gaikwad, Madhava
Published: (2025) -
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026) -
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026) -
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025) -
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
by: Gautam, Sushant, et al.
Published: (2026)