Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Yihao, Wei, Zeming, Sun, Jun, Sun, Meng |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Boosting Jailbreak Attack with Momentum
par: Zhang, Yihao, et autres
Publié: (2024)
par: Zhang, Yihao, et autres
Publié: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
par: Wu, Chengcan, et autres
Publié: (2025)
par: Wu, Chengcan, et autres
Publié: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
par: Zhang, Zhixin, et autres
Publié: (2025)
par: Zhang, Zhixin, et autres
Publié: (2025)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
par: Zhang, Yihao, et autres
Publié: (2024)
par: Zhang, Yihao, et autres
Publié: (2024)
Exploring the Robustness of In-Context Learning with Noisy Labels
par: Cheng, Chen, et autres
Publié: (2024)
par: Cheng, Chen, et autres
Publié: (2024)
Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks
par: Wang, Rui, et autres
Publié: (2025)
par: Wang, Rui, et autres
Publié: (2025)
Automata-Based Steering of Large Language Models for Diverse Structured Generation
par: Luan, Xiaokun, et autres
Publié: (2025)
par: Luan, Xiaokun, et autres
Publié: (2025)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
par: Wei, Zeming, et autres
Publié: (2026)
par: Wei, Zeming, et autres
Publié: (2026)
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
par: Wu, Chengcan, et autres
Publié: (2025)
par: Wu, Chengcan, et autres
Publié: (2025)
MILE: A Mutation Testing Framework of In-Context Learning Systems
par: Wei, Zeming, et autres
Publié: (2024)
par: Wei, Zeming, et autres
Publié: (2024)
Identifying and Understanding Cross-Class Features in Adversarial Training
par: Wei, Zeming, et autres
Publié: (2025)
par: Wei, Zeming, et autres
Publié: (2025)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
par: Wang, Kai, et autres
Publié: (2025)
par: Wang, Kai, et autres
Publié: (2025)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
par: Wei, Zeming, et autres
Publié: (2026)
par: Wei, Zeming, et autres
Publié: (2026)
SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing
par: Sun, Jialong, et autres
Publié: (2026)
par: Sun, Jialong, et autres
Publié: (2026)
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction
par: Wei, Zeming, et autres
Publié: (2025)
par: Wei, Zeming, et autres
Publié: (2025)
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
par: Liu, Fuqiang, et autres
Publié: (2024)
par: Liu, Fuqiang, et autres
Publié: (2024)
On Model Protection in Federated Learning against Eavesdropping Attacks
par: Maity, Dipankar, et autres
Publié: (2025)
par: Maity, Dipankar, et autres
Publié: (2025)
Clip-and-Verify: Linear Constraint-Driven Domain Clipping for Accelerating Neural Network Verification
par: Zhou, Duo, et autres
Publié: (2025)
par: Zhou, Duo, et autres
Publié: (2025)
A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense
par: Yun, Jihyeon, et autres
Publié: (2026)
par: Yun, Jihyeon, et autres
Publié: (2026)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
par: Yuan, Hongbang, et autres
Publié: (2024)
par: Yuan, Hongbang, et autres
Publié: (2024)
Approximate and Weighted Data Reconstruction Attack in Federated Learning
par: Song, Yongcun, et autres
Publié: (2023)
par: Song, Yongcun, et autres
Publié: (2023)
Characterizing the Training Dynamics of Private Fine-tuning with Langevin diffusion
par: Ke, Shuqi, et autres
Publié: (2024)
par: Ke, Shuqi, et autres
Publié: (2024)
Correlated Noise Provably Beats Independent Noise for Differentially Private Learning
par: Choquette-Choo, Christopher A., et autres
Publié: (2023)
par: Choquette-Choo, Christopher A., et autres
Publié: (2023)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
par: Mo, Yichuan, et autres
Publié: (2024)
par: Mo, Yichuan, et autres
Publié: (2024)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
par: Wei, Zeming, et autres
Publié: (2023)
par: Wei, Zeming, et autres
Publié: (2023)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
par: Liu, Qin, et autres
Publié: (2024)
par: Liu, Qin, et autres
Publié: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
par: Tan, Yuting, et autres
Publié: (2025)
par: Tan, Yuting, et autres
Publié: (2025)
SVIP: Towards Verifiable Inference of Open-source Large Language Models
par: Sun, Yifan, et autres
Publié: (2024)
par: Sun, Yifan, et autres
Publié: (2024)
DPZero: Private Fine-Tuning of Language Models without Backpropagation
par: Zhang, Liang, et autres
Publié: (2023)
par: Zhang, Liang, et autres
Publié: (2023)
Adversarial Text Purification: A Large Language Model Approach for Defense
par: Moraffah, Raha, et autres
Publié: (2024)
par: Moraffah, Raha, et autres
Publié: (2024)
Reverse-Engineering Model Editing on Language Models
par: Sun, Zhiyu, et autres
Publié: (2026)
par: Sun, Zhiyu, et autres
Publié: (2026)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
par: Hua, Peichun, et autres
Publié: (2025)
par: Hua, Peichun, et autres
Publié: (2025)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
par: Chugh, Rishit
Publié: (2026)
par: Chugh, Rishit
Publié: (2026)
Was it Slander? Towards Exact Inversion of Generative Language Models
par: Skapars, Adrians, et autres
Publié: (2024)
par: Skapars, Adrians, et autres
Publié: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
par: Reddy, Aashray, et autres
Publié: (2025)
par: Reddy, Aashray, et autres
Publié: (2025)
Efficient Optimization Algorithms for Linear Adversarial Training
par: RIbeiro, Antônio H., et autres
Publié: (2024)
par: RIbeiro, Antônio H., et autres
Publié: (2024)
On Adversarial Robustness of Language Models in Transfer Learning
par: Turbal, Bohdan, et autres
Publié: (2024)
par: Turbal, Bohdan, et autres
Publié: (2024)
Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization
par: Ribeiro, Antônio H., et autres
Publié: (2025)
par: Ribeiro, Antônio H., et autres
Publié: (2025)
The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems
par: Zhang, Yihao, et autres
Publié: (2026)
par: Zhang, Yihao, et autres
Publié: (2026)
Instructional Fingerprinting of Large Language Models
par: Xu, Jiashu, et autres
Publié: (2024)
par: Xu, Jiashu, et autres
Publié: (2024)
Documents similaires
-
Boosting Jailbreak Attack with Momentum
par: Zhang, Yihao, et autres
Publié: (2024) -
Secure LLM Fine-Tuning via Safety-Aware Probing
par: Wu, Chengcan, et autres
Publié: (2025) -
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
par: Zhang, Zhixin, et autres
Publié: (2025) -
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
par: Zhang, Yihao, et autres
Publié: (2024) -
Exploring the Robustness of In-Context Learning with Noisy Labels
par: Cheng, Chen, et autres
Publié: (2024)