Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity
Fuente:
arXiv
Guardado en:
| Autores principales: | An, Hongjun, Song, Yiliang, Chen, Jiangan, Shao, Jiawei, Zhang, Chi, Li, Xuelong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring Advanced Methodologies in Security Evaluation for LLMs
por: Huang, Jun, et al.
Publicado: (2024)
por: Huang, Jun, et al.
Publicado: (2024)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
por: Fastowski, Alina, et al.
Publicado: (2025)
por: Fastowski, Alina, et al.
Publicado: (2025)
PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
por: Wei, Zichao, et al.
Publicado: (2025)
por: Wei, Zichao, et al.
Publicado: (2025)
QUIC-Exfil: Exploiting QUIC's Server Preferred Address Feature to Perform Data Exfiltration Attacks
por: Grübl, Thomas, et al.
Publicado: (2025)
por: Grübl, Thomas, et al.
Publicado: (2025)
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
por: Li, Hongyi, et al.
Publicado: (2024)
por: Li, Hongyi, et al.
Publicado: (2024)
MPMA: Preference Manipulation Attack Against Model Context Protocol
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
SoK: Root Cause of $1 Billion Loss in Smart Contract Real-World Attacks via a Systematic Literature Review of Vulnerabilities
por: Rezaei, Hadis, et al.
Publicado: (2025)
por: Rezaei, Hadis, et al.
Publicado: (2025)
BACKRUNNER: Mitigating Smart Contract Attacks in the Real World
por: Shou, Chaofan, et al.
Publicado: (2024)
por: Shou, Chaofan, et al.
Publicado: (2024)
KingsGuard: Enclave Data Protection Under Real-World TEE Vulnerabilities
por: Allaqband, Saltanat Firdous, et al.
Publicado: (2026)
por: Allaqband, Saltanat Firdous, et al.
Publicado: (2026)
AutoPatch: Multi-Agent Framework for Patching Real-World CVE Vulnerabilities
por: Seo, Minjae, et al.
Publicado: (2025)
por: Seo, Minjae, et al.
Publicado: (2025)
Revisiting Locally Differentially Private Protocols: Towards Better Trade-offs in Privacy, Utility, and Attack Resistance
por: Arcolezi, Héber H., et al.
Publicado: (2025)
por: Arcolezi, Héber H., et al.
Publicado: (2025)
Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models
por: Yang, Yijun, et al.
Publicado: (2025)
por: Yang, Yijun, et al.
Publicado: (2025)
AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering
por: Gungor, Onat, et al.
Publicado: (2025)
por: Gungor, Onat, et al.
Publicado: (2025)
Comprehensive Evaluation of Cloaking Backdoor Attacks on Object Detector in Real-World
por: Ma, Hua, et al.
Publicado: (2025)
por: Ma, Hua, et al.
Publicado: (2025)
Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training
por: Wang, Yisen, et al.
Publicado: (2025)
por: Wang, Yisen, et al.
Publicado: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
por: Yin, Yu, et al.
Publicado: (2026)
por: Yin, Yu, et al.
Publicado: (2026)
Agentic Discovery and Validation of Android App Vulnerabilities
por: Wang, Ziyue, et al.
Publicado: (2025)
por: Wang, Ziyue, et al.
Publicado: (2025)
LAAF: Logic-layer Automated Attack Framework A Systematic Red-Teaming Methodology for LPCI Vulnerabilities in Agentic Large Language Model Systems
por: Atta, Hammad, et al.
Publicado: (2026)
por: Atta, Hammad, et al.
Publicado: (2026)
Shadow-Free Membership Inference Attacks: Recommender Systems Are More Vulnerable Than You Thought
por: Chi, Xiaoxiao, et al.
Publicado: (2024)
por: Chi, Xiaoxiao, et al.
Publicado: (2024)
Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
por: Tan, Rui Yang, et al.
Publicado: (2026)
por: Tan, Rui Yang, et al.
Publicado: (2026)
The Privacy-Utility Trade-off in the Topics API
por: Alvim, Mário S., et al.
Publicado: (2024)
por: Alvim, Mário S., et al.
Publicado: (2024)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
por: Xi, Haoran, et al.
Publicado: (2025)
por: Xi, Haoran, et al.
Publicado: (2025)
HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
por: Dai, Fangqi, et al.
Publicado: (2025)
por: Dai, Fangqi, et al.
Publicado: (2025)
JNI Global References Are Still Vulnerable: Attacks and Defenses
por: He, Yi, et al.
Publicado: (2024)
por: He, Yi, et al.
Publicado: (2024)
Zero Day Attacks: Novel Behaviour or Novel Vulnerability?
por: Jibunoh, Nnamdi, et al.
Publicado: (2026)
por: Jibunoh, Nnamdi, et al.
Publicado: (2026)
A Novel Classification of Attacks on Blockchain Layers: Vulnerabilities, Attacks, Mitigations, and Research Directions
por: Dwivedi, Kaustubh, et al.
Publicado: (2024)
por: Dwivedi, Kaustubh, et al.
Publicado: (2024)
Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences
por: Hasan, Mohammad Saqib, et al.
Publicado: (2025)
por: Hasan, Mohammad Saqib, et al.
Publicado: (2025)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
por: Ding, Ruomeng, et al.
Publicado: (2026)
por: Ding, Ruomeng, et al.
Publicado: (2026)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
por: Wang, Xinkai, et al.
Publicado: (2025)
por: Wang, Xinkai, et al.
Publicado: (2025)
Cryptanalysis of a Lightweight RFID Authentication Protocol Based on a Variable Matrix Encryption Algorithm
por: Wu, Hongjun
Publicado: (2026)
por: Wu, Hongjun
Publicado: (2026)
Uncovering Hidden Inclusions of Vulnerable Dependencies in Real-World Java Projects
por: Schott, Stefan, et al.
Publicado: (2026)
por: Schott, Stefan, et al.
Publicado: (2026)
Real-World Usability of Vulnerability Proof-of-Concepts: A Comprehensive Study
por: Dang, Wenjing, et al.
Publicado: (2025)
por: Dang, Wenjing, et al.
Publicado: (2025)
Uncovering Privacy Vulnerabilities through Analytical Gradient Inversion Attacks
por: Eltaras, Tamer Ahmed, et al.
Publicado: (2025)
por: Eltaras, Tamer Ahmed, et al.
Publicado: (2025)
Validating Threat Modeling Results with the Help of Vulnerable Test Applications
por: Adamov, Oleksandr, et al.
Publicado: (2026)
por: Adamov, Oleksandr, et al.
Publicado: (2026)
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
por: Zhang, Shuoming, et al.
Publicado: (2025)
por: Zhang, Shuoming, et al.
Publicado: (2025)
Analyzing Vector Register Usage in Linux Packages to Understand Real-World Impact of Downfall Attack
por: Harata, Yohei, et al.
Publicado: (2026)
por: Harata, Yohei, et al.
Publicado: (2026)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
por: Shao, Zedian, et al.
Publicado: (2024)
por: Shao, Zedian, et al.
Publicado: (2024)
Tool Preferences in Agentic LLMs are Unreliable
por: Faghih, Kazem, et al.
Publicado: (2025)
por: Faghih, Kazem, et al.
Publicado: (2025)
Vanishing Watermarks: Diffusion-Based Image Editing Undermines Robust Invisible Watermarking
por: Guo, Fan, et al.
Publicado: (2026)
por: Guo, Fan, et al.
Publicado: (2026)
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
por: Halloran, John
Publicado: (2025)
por: Halloran, John
Publicado: (2025)
Ejemplares similares
-
Exploring Advanced Methodologies in Security Evaluation for LLMs
por: Huang, Jun, et al.
Publicado: (2024) -
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
por: Fastowski, Alina, et al.
Publicado: (2025) -
PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
por: Wei, Zichao, et al.
Publicado: (2025) -
QUIC-Exfil: Exploiting QUIC's Server Preferred Address Feature to Perform Data Exfiltration Attacks
por: Grübl, Thomas, et al.
Publicado: (2025) -
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
por: Li, Hongyi, et al.
Publicado: (2024)