Fast Proxies for LLM Robustness Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Beyer, Tim, Schuchardt, Jan, Schwinn, Leo, Günnemann, Stephan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Provable Adversarial Robustness for Group Equivariant Tasks: Graphs, Point Clouds, Molecules, and More
by: Schuchardt, Jan, et al.
Published: (2023)
by: Schuchardt, Jan, et al.
Published: (2023)
Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
by: Saxena, Aman, et al.
Published: (2026)
by: Saxena, Aman, et al.
Published: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Unified Mechanism-Specific Amplification by Subsampling and Group Privacy Amplification
by: Schuchardt, Jan, et al.
Published: (2024)
by: Schuchardt, Jan, et al.
Published: (2024)
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
Revisiting the Robust Alignment of Circuit Breakers
by: Schwinn, Leo, et al.
Published: (2024)
by: Schwinn, Leo, et al.
Published: (2024)
AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Amplified Patch-Level Differential Privacy for Free via Random Cropping
by: Durmaz, Kaan, et al.
Published: (2026)
by: Durmaz, Kaan, et al.
Published: (2026)
MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers
by: Ahmadi, Arash, et al.
Published: (2025)
by: Ahmadi, Arash, et al.
Published: (2025)
Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control
by: Uppala, Rohith
Published: (2026)
by: Uppala, Rohith
Published: (2026)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026)
by: Schwinn, Leo, et al.
Published: (2026)
ShadowLogic: Backdoors in Any Whitebox LLM
by: Schulz, Kasimir, et al.
Published: (2025)
by: Schulz, Kasimir, et al.
Published: (2025)
Privacy Amplification by Structured Subsampling for Deep Differentially Private Time Series Forecasting
by: Schuchardt, Jan, et al.
Published: (2025)
by: Schuchardt, Jan, et al.
Published: (2025)
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
by: Zhao, Jin, et al.
Published: (2026)
by: Zhao, Jin, et al.
Published: (2026)
Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks
by: Isbarov, Jafar, et al.
Published: (2026)
by: Isbarov, Jafar, et al.
Published: (2026)
FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference
by: Lin, Chenqi, et al.
Published: (2024)
by: Lin, Chenqi, et al.
Published: (2024)
TED-LaST: Towards Robust Backdoor Defense Against Adaptive Attacks
by: Mo, Xiaoxing, et al.
Published: (2025)
by: Mo, Xiaoxing, et al.
Published: (2025)
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
by: Wang, Shida, et al.
Published: (2025)
by: Wang, Shida, et al.
Published: (2025)
Character-Level Perturbations Disrupt LLM Watermarks
by: Zhang, Zhaoxi, et al.
Published: (2025)
by: Zhang, Zhaoxi, et al.
Published: (2025)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
by: Chia-Pei, et al.
Published: (2026)
by: Chia-Pei, et al.
Published: (2026)
NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples
by: Pu, Shi, et al.
Published: (2025)
by: Pu, Shi, et al.
Published: (2025)
SAGE: A Generic Framework for LLM Safety Evaluation
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
Adaptive and Robust Cost-Aware Proof of Quality for Decentralized LLM Inference Networks
by: Tian, Arther, et al.
Published: (2026)
by: Tian, Arther, et al.
Published: (2026)
FedLiTeCAN : A Federated Lightweight Transformer for Fast and Robust CAN Bus Intrusion Detection
by: S, Devika, et al.
Published: (2025)
by: S, Devika, et al.
Published: (2025)
SafeGenes: Evaluating the Adversarial Robustness of Genomic Foundation Models
by: Zhan, Huixin, et al.
Published: (2025)
by: Zhan, Huixin, et al.
Published: (2025)
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
by: Gao, Jiaxin, et al.
Published: (2025)
by: Gao, Jiaxin, et al.
Published: (2025)
Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection
by: Halder, Subho, et al.
Published: (2026)
by: Halder, Subho, et al.
Published: (2026)
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
by: Wang, Junyu, et al.
Published: (2025)
by: Wang, Junyu, et al.
Published: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
by: Ji, Zimo, et al.
Published: (2025)
by: Ji, Zimo, et al.
Published: (2025)
Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study
by: Chen, Zhihao, et al.
Published: (2026)
by: Chen, Zhihao, et al.
Published: (2026)
Are Robust LLM Fingerprints Adversarially Robust?
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
by: Wang, Shouju, et al.
Published: (2025)
by: Wang, Shouju, et al.
Published: (2025)
Enhancing Security in LLM Applications: A Performance Evaluation of Early Detection Systems
by: Gakh, Valerii, et al.
Published: (2025)
by: Gakh, Valerii, et al.
Published: (2025)
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks
by: Sahu, Anubhab, et al.
Published: (2026)
by: Sahu, Anubhab, et al.
Published: (2026)
SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks
by: Sivaroopan, Nirhoshan, et al.
Published: (2026)
by: Sivaroopan, Nirhoshan, et al.
Published: (2026)
ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense
by: Lau, Nancy, et al.
Published: (2026)
by: Lau, Nancy, et al.
Published: (2026)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
by: Fu, Yuchuan, et al.
Published: (2025)
by: Fu, Yuchuan, et al.
Published: (2025)
Similar Items
-
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025) -
Provable Adversarial Robustness for Group Equivariant Tasks: Graphs, Point Clouds, Molecules, and More
by: Schuchardt, Jan, et al.
Published: (2023) -
Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
by: Saxena, Aman, et al.
Published: (2026) -
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026) -
Unified Mechanism-Specific Amplification by Subsampling and Group Privacy Amplification
by: Schuchardt, Jan, et al.
Published: (2024)