LLM-Safety Evaluations Lack Robustness
Fuente:
arXiv
Guardado en:
| Autores principales: | Beyer, Tim, Xhonneux, Sophie, Geisler, Simon, Gidel, Gauthier, Schwinn, Leo, Günnemann, Stephan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fast Proxies for LLM Robustness Evaluation
por: Beyer, Tim, et al.
Publicado: (2025)
por: Beyer, Tim, et al.
Publicado: (2025)
Efficient Adversarial Training in LLMs with Continuous Attacks
por: Xhonneux, Sophie, et al.
Publicado: (2024)
por: Xhonneux, Sophie, et al.
Publicado: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
por: Dobre, David, et al.
Publicado: (2025)
por: Dobre, David, et al.
Publicado: (2025)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
por: Schwinn, Leo, et al.
Publicado: (2026)
por: Schwinn, Leo, et al.
Publicado: (2026)
Revisiting the Robust Alignment of Circuit Breakers
por: Schwinn, Leo, et al.
Publicado: (2024)
por: Schwinn, Leo, et al.
Publicado: (2024)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
por: Schwinn, Leo, et al.
Publicado: (2024)
por: Schwinn, Leo, et al.
Publicado: (2024)
In-Context Learning Can Re-learn Forbidden Tasks
por: Xhonneux, Sophie, et al.
Publicado: (2024)
por: Xhonneux, Sophie, et al.
Publicado: (2024)
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
por: Scholten, Yan, et al.
Publicado: (2025)
por: Scholten, Yan, et al.
Publicado: (2025)
Closing the Distribution Gap in Adversarial Training for LLMs
por: Hu, Chengzhi, et al.
Publicado: (2026)
por: Hu, Chengzhi, et al.
Publicado: (2026)
SAGE: A Generic Framework for LLM Safety Evaluation
por: Jindal, Madhur, et al.
Publicado: (2025)
por: Jindal, Madhur, et al.
Publicado: (2025)
AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
por: Beyer, Tim, et al.
Publicado: (2025)
por: Beyer, Tim, et al.
Publicado: (2025)
Adaptive and Robust Cost-Aware Proof of Quality for Decentralized LLM Inference Networks
por: Tian, Arther, et al.
Publicado: (2026)
por: Tian, Arther, et al.
Publicado: (2026)
ShadowLogic: Backdoors in Any Whitebox LLM
por: Schulz, Kasimir, et al.
Publicado: (2025)
por: Schulz, Kasimir, et al.
Publicado: (2025)
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
por: Zhao, Jin, et al.
Publicado: (2026)
por: Zhao, Jin, et al.
Publicado: (2026)
aiXamine: Simplified LLM Safety and Security
por: Deniz, Fatih, et al.
Publicado: (2025)
por: Deniz, Fatih, et al.
Publicado: (2025)
FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
por: Kuznetsov, Daniel, et al.
Publicado: (2026)
por: Kuznetsov, Daniel, et al.
Publicado: (2026)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
por: Li, Nanxi, et al.
Publicado: (2025)
por: Li, Nanxi, et al.
Publicado: (2025)
Safety Layers in Aligned Large Language Models: The Key to LLM Security
por: Li, Shen, et al.
Publicado: (2024)
por: Li, Shen, et al.
Publicado: (2024)
QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
por: Lee, Taegyeong, et al.
Publicado: (2025)
por: Lee, Taegyeong, et al.
Publicado: (2025)
Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis
por: Topol, Zvi
Publicado: (2026)
por: Topol, Zvi
Publicado: (2026)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
por: Hossain, Saad, et al.
Publicado: (2026)
por: Hossain, Saad, et al.
Publicado: (2026)
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
por: Chen, Jizhou, et al.
Publicado: (2025)
por: Chen, Jizhou, et al.
Publicado: (2025)
SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration
por: Pan, Yu, et al.
Publicado: (2026)
por: Pan, Yu, et al.
Publicado: (2026)
TED-LaST: Towards Robust Backdoor Defense Against Adaptive Attacks
por: Mo, Xiaoxing, et al.
Publicado: (2025)
por: Mo, Xiaoxing, et al.
Publicado: (2025)
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
por: Wang, Shida, et al.
Publicado: (2025)
por: Wang, Shida, et al.
Publicado: (2025)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
por: Yang, Chenglin
Publicado: (2026)
por: Yang, Chenglin
Publicado: (2026)
Character-Level Perturbations Disrupt LLM Watermarks
por: Zhang, Zhaoxi, et al.
Publicado: (2025)
por: Zhang, Zhaoxi, et al.
Publicado: (2025)
LLM-based event log analysis techniques: A survey
por: Akhtar, Siraaj, et al.
Publicado: (2025)
por: Akhtar, Siraaj, et al.
Publicado: (2025)
NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples
por: Pu, Shi, et al.
Publicado: (2025)
por: Pu, Shi, et al.
Publicado: (2025)
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
por: Zheng, Baolin, et al.
Publicado: (2025)
por: Zheng, Baolin, et al.
Publicado: (2025)
Efficient LLM Safety Evaluation through Multi-Agent Debate
por: Lin, Dachuan, et al.
Publicado: (2025)
por: Lin, Dachuan, et al.
Publicado: (2025)
SafeGenes: Evaluating the Adversarial Robustness of Genomic Foundation Models
por: Zhan, Huixin, et al.
Publicado: (2025)
por: Zhan, Huixin, et al.
Publicado: (2025)
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
por: Gao, Jiaxin, et al.
Publicado: (2025)
por: Gao, Jiaxin, et al.
Publicado: (2025)
Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection
por: Halder, Subho, et al.
Publicado: (2026)
por: Halder, Subho, et al.
Publicado: (2026)
Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
por: Li, Xurui, et al.
Publicado: (2025)
por: Li, Xurui, et al.
Publicado: (2025)
CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning
por: ElZemity, Adel, et al.
Publicado: (2025)
por: ElZemity, Adel, et al.
Publicado: (2025)
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
por: Wang, Junyu, et al.
Publicado: (2025)
por: Wang, Junyu, et al.
Publicado: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
por: Che, Zora, et al.
Publicado: (2025)
por: Che, Zora, et al.
Publicado: (2025)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
por: Ji, Zimo, et al.
Publicado: (2025)
por: Ji, Zimo, et al.
Publicado: (2025)
Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers
por: Heye, David, et al.
Publicado: (2026)
por: Heye, David, et al.
Publicado: (2026)
Ejemplares similares
-
Fast Proxies for LLM Robustness Evaluation
por: Beyer, Tim, et al.
Publicado: (2025) -
Efficient Adversarial Training in LLMs with Continuous Attacks
por: Xhonneux, Sophie, et al.
Publicado: (2024) -
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
por: Dobre, David, et al.
Publicado: (2025) -
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
por: Schwinn, Leo, et al.
Publicado: (2026) -
Revisiting the Robust Alignment of Circuit Breakers
por: Schwinn, Leo, et al.
Publicado: (2024)