Decoupled Alignment for Robust Plug-and-Play Adaptation
Fuente:
arXiv
Salvato in:
| Autori principali: | Luo, Haozheng, Yu, Jiahao, Zhang, Wenxin, Li, Jialong, Hu, Jerry Yao-Chieh, Xing, Xinyu, Liu, Han |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
di: Chen, Sixu, et al.
Pubblicazione: (2026)
di: Chen, Sixu, et al.
Pubblicazione: (2026)
What Matters For Safety Alignment?
di: Li, Xing, et al.
Pubblicazione: (2026)
di: Li, Xing, et al.
Pubblicazione: (2026)
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
di: Huo, Jiahao, et al.
Pubblicazione: (2026)
di: Huo, Jiahao, et al.
Pubblicazione: (2026)
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
di: Xu, Yixiao, et al.
Pubblicazione: (2025)
di: Xu, Yixiao, et al.
Pubblicazione: (2025)
BlockScan: Detecting Anomalies in Blockchain Transactions
di: Yu, Jiahao, et al.
Pubblicazione: (2024)
di: Yu, Jiahao, et al.
Pubblicazione: (2024)
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem
di: Liu, Hongyi, et al.
Pubblicazione: (2024)
di: Liu, Hongyi, et al.
Pubblicazione: (2024)
Soft-Label Integration for Robust Toxicity Classification
di: Cheng, Zelei, et al.
Pubblicazione: (2024)
di: Cheng, Zelei, et al.
Pubblicazione: (2024)
CATMark: A Context-Aware Thresholding Framework for Robust Cross-Task Watermarking in Large Language Models
di: Zhang, Yu, et al.
Pubblicazione: (2025)
di: Zhang, Yu, et al.
Pubblicazione: (2025)
XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts
di: Xu, Jiahao, et al.
Pubblicazione: (2026)
di: Xu, Jiahao, et al.
Pubblicazione: (2026)
Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position
di: Xie, Zhixin, et al.
Pubblicazione: (2025)
di: Xie, Zhixin, et al.
Pubblicazione: (2025)
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
di: Fan, Kaisheng, et al.
Pubblicazione: (2026)
di: Fan, Kaisheng, et al.
Pubblicazione: (2026)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
di: Pisano, Matthew, et al.
Pubblicazione: (2023)
di: Pisano, Matthew, et al.
Pubblicazione: (2023)
Assessing Prompt Injection Risks in 200+ Custom GPTs
di: Yu, Jiahao, et al.
Pubblicazione: (2023)
di: Yu, Jiahao, et al.
Pubblicazione: (2023)
AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
di: Zhang, Jinchuan, et al.
Pubblicazione: (2025)
di: Zhang, Jinchuan, et al.
Pubblicazione: (2025)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
di: Luo, Weidi, et al.
Pubblicazione: (2024)
di: Luo, Weidi, et al.
Pubblicazione: (2024)
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
di: Liang, Zi, et al.
Pubblicazione: (2025)
di: Liang, Zi, et al.
Pubblicazione: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
di: Mu, Junjie, et al.
Pubblicazione: (2025)
di: Mu, Junjie, et al.
Pubblicazione: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
di: Huang, Chao, et al.
Pubblicazione: (2025)
di: Huang, Chao, et al.
Pubblicazione: (2025)
Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents
di: Liang, Zhibo, et al.
Pubblicazione: (2025)
di: Liang, Zhibo, et al.
Pubblicazione: (2025)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
di: Sel, Bilgehan, et al.
Pubblicazione: (2026)
di: Sel, Bilgehan, et al.
Pubblicazione: (2026)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
di: Wang, Kun, et al.
Pubblicazione: (2026)
di: Wang, Kun, et al.
Pubblicazione: (2026)
Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training
di: Yu, Yao-Ching, et al.
Pubblicazione: (2025)
di: Yu, Yao-Ching, et al.
Pubblicazione: (2025)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
di: Ma, Siyuan, et al.
Pubblicazione: (2024)
di: Ma, Siyuan, et al.
Pubblicazione: (2024)
On Differentially Private String Distances
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
PROMPTFUZZ: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
di: Yu, Jiahao, et al.
Pubblicazione: (2024)
di: Yu, Jiahao, et al.
Pubblicazione: (2024)
EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
di: Zhou, Xinyun, et al.
Pubblicazione: (2025)
di: Zhou, Xinyun, et al.
Pubblicazione: (2025)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
di: Li, Tianhao, et al.
Pubblicazione: (2024)
di: Li, Tianhao, et al.
Pubblicazione: (2024)
SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From
di: Tong, Yao, et al.
Pubblicazione: (2025)
di: Tong, Yao, et al.
Pubblicazione: (2025)
Generalization-Enhanced Code Vulnerability Detection via Multi-Task Instruction Fine-Tuning
di: Du, Xiaohu, et al.
Pubblicazione: (2024)
di: Du, Xiaohu, et al.
Pubblicazione: (2024)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
di: Lv, Lijia, et al.
Pubblicazione: (2024)
di: Lv, Lijia, et al.
Pubblicazione: (2024)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
di: Kumar, Anurakt, et al.
Pubblicazione: (2024)
di: Kumar, Anurakt, et al.
Pubblicazione: (2024)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
di: Kim, Jinhwa, et al.
Pubblicazione: (2025)
di: Kim, Jinhwa, et al.
Pubblicazione: (2025)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
di: Chae, Kyubyung, et al.
Pubblicazione: (2025)
di: Chae, Kyubyung, et al.
Pubblicazione: (2025)
Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
di: Luo, Zhifan, et al.
Pubblicazione: (2025)
di: Luo, Zhifan, et al.
Pubblicazione: (2025)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
di: Nawal, Aditya, et al.
Pubblicazione: (2026)
di: Nawal, Aditya, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
di: Chen, Sixu, et al.
Pubblicazione: (2026) -
What Matters For Safety Alignment?
di: Li, Xing, et al.
Pubblicazione: (2026) -
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
di: Huo, Jiahao, et al.
Pubblicazione: (2026) -
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
di: Xu, Yixiao, et al.
Pubblicazione: (2025) -
BlockScan: Detecting Anomalies in Blockchain Transactions
di: Yu, Jiahao, et al.
Pubblicazione: (2024)