IF-GUIDE: Influence Function-Guided Detoxification of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Coalson, Zachary, Bae, Juhan, Carlini, Nicholas, Hong, Sanghyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs
by: Coalson, Zachary, et al.
Published: (2026)
by: Coalson, Zachary, et al.
Published: (2026)
Fail-Closed Alignment for Large Language Models
by: Coalson, Zachary, et al.
Published: (2026)
by: Coalson, Zachary, et al.
Published: (2026)
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
by: Coalson, Zachary, et al.
Published: (2024)
by: Coalson, Zachary, et al.
Published: (2024)
Discovering Universal Activation Directions for PII Leakage in Language Models
by: Marchyok, Leo, et al.
Published: (2026)
by: Marchyok, Leo, et al.
Published: (2026)
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
by: Hong, Sanghyun, et al.
Published: (2024)
by: Hong, Sanghyun, et al.
Published: (2024)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
Cutting through buggy adversarial example defenses: fixing 1 line of code breaks Sabre
by: Carlini, Nicholas
Published: (2024)
by: Carlini, Nicholas
Published: (2024)
Remote Timing Attacks on Efficient Language Model Inference
by: Carlini, Nicholas, et al.
Published: (2024)
by: Carlini, Nicholas, et al.
Published: (2024)
PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips
by: Coalson, Zachary, et al.
Published: (2024)
by: Coalson, Zachary, et al.
Published: (2024)
Evading Black-box Classifiers Without Breaking Eggs
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
by: Tramèr, Florian, et al.
Published: (2022)
by: Tramèr, Florian, et al.
Published: (2022)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
by: Rando, Javier, et al.
Published: (2025)
by: Rando, Javier, et al.
Published: (2025)
Modeling Neural Networks with Privacy Using Neural Stochastic Differential Equations
by: Hong, Sanghyun, et al.
Published: (2025)
by: Hong, Sanghyun, et al.
Published: (2025)
Large-scale online deanonymization with LLMs
by: Lermen, Simon, et al.
Published: (2026)
by: Lermen, Simon, et al.
Published: (2026)
Hessian-aware Training for Enhancing DNNs Resilience to Parameter Corruptions
by: Prato, Tahmid Hasan, et al.
Published: (2025)
by: Prato, Tahmid Hasan, et al.
Published: (2025)
MADCAT: Combating Malware Detection Under Concept Drift with Test-Time Adaptation
by: Roh, Eunjin, et al.
Published: (2025)
by: Roh, Eunjin, et al.
Published: (2025)
Understanding Deep Gradient Leakage via Inversion Influence Functions
by: Zhang, Haobo, et al.
Published: (2023)
by: Zhang, Haobo, et al.
Published: (2023)
Entropy-Guided Attention for Private LLMs
by: Jha, Nandan Kumar, et al.
Published: (2025)
by: Jha, Nandan Kumar, et al.
Published: (2025)
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
SoK: Watermarking for AI-Generated Content
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Detecting Instruction Fine-tuning Attacks using Influence Function
by: Li, Jiawei
Published: (2025)
by: Li, Jiawei
Published: (2025)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Poisoning Web-Scale Training Datasets is Practical
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Backdoor Learning Curves: Explaining Backdoor Poisoning Beyond Influence Functions
by: Cinà, Antonio Emanuele, et al.
Published: (2021)
by: Cinà, Antonio Emanuele, et al.
Published: (2021)
Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks
by: Nasr, Milad, et al.
Published: (2025)
by: Nasr, Milad, et al.
Published: (2025)
Can LLMs Handle WebShell Detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework
by: Han, Feijiang, et al.
Published: (2025)
by: Han, Feijiang, et al.
Published: (2025)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
by: Lee, Dongjun, et al.
Published: (2026)
by: Lee, Dongjun, et al.
Published: (2026)
Stealthy and Adjustable Text-Guided Backdoor Attacks on Multimodal Pretrained Models
by: Zhang, Yiyang, et al.
Published: (2026)
by: Zhang, Yiyang, et al.
Published: (2026)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
Can LLMs be Fooled? Investigating Vulnerabilities in LLMs
by: Abdali, Sara, et al.
Published: (2024)
by: Abdali, Sara, et al.
Published: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
by: Nasr, Milad, et al.
Published: (2025)
by: Nasr, Milad, et al.
Published: (2025)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
by: Huang, Yangsibo, et al.
Published: (2025)
by: Huang, Yangsibo, et al.
Published: (2025)
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs
by: Jha, Piyush, et al.
Published: (2024)
by: Jha, Piyush, et al.
Published: (2024)
Query-Based Adversarial Prompt Generation
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
Excessive Reasoning Attack on Reasoning LLMs
by: Si, Wai Man, et al.
Published: (2025)
by: Si, Wai Man, et al.
Published: (2025)
Towards Watermarking of Open-Source LLMs
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
Can LLMs Patch Security Issues?
by: Alrashedy, Kamel, et al.
Published: (2023)
by: Alrashedy, Kamel, et al.
Published: (2023)
Lap2: Revisiting Laplace DP-SGD for High Dimensions via Majorization Theory
by: Mohammady, Meisam, et al.
Published: (2026)
by: Mohammady, Meisam, et al.
Published: (2026)
Similar Items
-
Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs
by: Coalson, Zachary, et al.
Published: (2026) -
Fail-Closed Alignment for Large Language Models
by: Coalson, Zachary, et al.
Published: (2026) -
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
by: Coalson, Zachary, et al.
Published: (2024) -
Discovering Universal Activation Directions for PII Leakage in Language Models
by: Marchyok, Leo, et al.
Published: (2026) -
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
by: Hong, Sanghyun, et al.
Published: (2024)