A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhiyu, Li, Yu, Zhang, Suochao, Zhou, Jingbo, Zhou, Jiwen, Bao, Chenfu, Yu, Dianhai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
von: Jin, Zhibo, et al.
Veröffentlicht: (2024)
von: Jin, Zhibo, et al.
Veröffentlicht: (2024)
Robust LLM safeguarding via refusal feature adversarial training
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
Low-Cost Hard-Label Adversarial Attack with Theoretical Foundations
von: Liu, Jun, et al.
Veröffentlicht: (2026)
von: Liu, Jun, et al.
Veröffentlicht: (2026)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
von: Hsiung, Lei, et al.
Veröffentlicht: (2025)
von: Hsiung, Lei, et al.
Veröffentlicht: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
Panther: A Cost-Effective Privacy-Preserving Framework for GNN Training and Inference Services in Cloud Environments
von: Chen, Congcong, et al.
Veröffentlicht: (2025)
von: Chen, Congcong, et al.
Veröffentlicht: (2025)
LLM Unlearning Should Be Form-Independent
von: Ye, Xiaotian, et al.
Veröffentlicht: (2025)
von: Ye, Xiaotian, et al.
Veröffentlicht: (2025)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
PVMark: Enabling Public Verifiability for LLM Watermarking Schemes
von: Duan, Haohua, et al.
Veröffentlicht: (2025)
von: Duan, Haohua, et al.
Veröffentlicht: (2025)
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
Attack and defense techniques in large language models: A survey and new perspectives
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
FGAD: Self-boosted Knowledge Distillation for An Effective Federated Graph Anomaly Detection Framework
von: Cai, Jinyu, et al.
Veröffentlicht: (2024)
von: Cai, Jinyu, et al.
Veröffentlicht: (2024)
InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
von: Vinod, Vishnu, et al.
Veröffentlicht: (2025)
von: Vinod, Vishnu, et al.
Veröffentlicht: (2025)
VERA: Variational Inference Framework for Jailbreaking Large Language Models
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
FreqMark: Frequency-Based Watermark for Sentence-Level Detection of LLM-Generated Text
von: Xu, Zhenyu, et al.
Veröffentlicht: (2024)
von: Xu, Zhenyu, et al.
Veröffentlicht: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
Shake to Leak: Fine-tuning Diffusion Models Can Amplify the Generative Privacy Risk
von: Li, Zhangheng, et al.
Veröffentlicht: (2024)
von: Li, Zhangheng, et al.
Veröffentlicht: (2024)
The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
von: Makroo, Owais, et al.
Veröffentlicht: (2025)
von: Makroo, Owais, et al.
Veröffentlicht: (2025)
WaterVIB: Learning Minimal Sufficient Watermark Representations via Variational Information Bottleneck
von: He, Haoyuan, et al.
Veröffentlicht: (2026)
von: He, Haoyuan, et al.
Veröffentlicht: (2026)
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
GCG Attack On A Diffusion LLM
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
von: Sahili, Ali Al, et al.
Veröffentlicht: (2025)
von: Sahili, Ali Al, et al.
Veröffentlicht: (2025)
SecFormer: Fast and Accurate Privacy-Preserving Inference for Transformer Models via SMPC
von: Luo, Jinglong, et al.
Veröffentlicht: (2024)
von: Luo, Jinglong, et al.
Veröffentlicht: (2024)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
LLMGuard: Guarding Against Unsafe LLM Behavior
von: Goyal, Shubh, et al.
Veröffentlicht: (2024)
von: Goyal, Shubh, et al.
Veröffentlicht: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
Localizing Malicious Outputs from CodeLLM
von: Borana, Mayukh, et al.
Veröffentlicht: (2025)
von: Borana, Mayukh, et al.
Veröffentlicht: (2025)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
Quantifying Policy Administration Cost in an Active Learning Framework
von: Zhang, Si, et al.
Veröffentlicht: (2023)
von: Zhang, Si, et al.
Veröffentlicht: (2023)
Copyright-Protected Language Generation via Adaptive Model Fusion
von: Abad, Javier, et al.
Veröffentlicht: (2024)
von: Abad, Javier, et al.
Veröffentlicht: (2024)
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
VEXA: Evidence-Grounded and Persona-Adaptive Explanations for Scam Risk Sensemaking
von: An, Heajun, et al.
Veröffentlicht: (2026)
von: An, Heajun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
von: Jin, Zhibo, et al.
Veröffentlicht: (2024) -
Robust LLM safeguarding via refusal feature adversarial training
von: Yu, Lei, et al.
Veröffentlicht: (2024) -
Low-Cost Hard-Label Adversarial Attack with Theoretical Foundations
von: Liu, Jun, et al.
Veröffentlicht: (2026) -
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
von: Wang, Tianchun, et al.
Veröffentlicht: (2024) -
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
von: Hsiung, Lei, et al.
Veröffentlicht: (2025)