Immunization against harmful fine-tuning attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Rosati, Domenic, Wehner, Jan, Williams, Kai, Bartoszcze, Łukasz, Batzner, Jan, Sajjad, Hassan, Rudzicz, Frank |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Representation Noising: A Defence Mechanism Against Harmful Finetuning
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
Resolving Lexical Bias in Model Editing
by: Rizwan, Hammad, et al.
Published: (2024)
by: Rizwan, Hammad, et al.
Published: (2024)
Dependency Parsing is More Parameter-Efficient with Normalization
by: Gajo, Paolo, et al.
Published: (2025)
by: Gajo, Paolo, et al.
Published: (2025)
Long-form evaluation of model editing
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
LLMs Underperform Graph-Based Parsers on Supervised Relation Extraction for Complex Graphs
by: Gajo, Paolo, et al.
Published: (2026)
by: Gajo, Paolo, et al.
Published: (2026)
One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization
by: Weeber, Franziska, et al.
Published: (2026)
by: Weeber, Franziska, et al.
Published: (2026)
Sycophancy Claims about Language Models: The Missing Human-in-the-Loop
by: Batzner, Jan, et al.
Published: (2025)
by: Batzner, Jan, et al.
Published: (2025)
Limits of Convergence-Rate Control for Open-Weight Safety
by: Rosati, Domenic, et al.
Published: (2026)
by: Rosati, Domenic, et al.
Published: (2026)
Semantic Consistency for Assuring Reliability of Large Language Models
by: Raj, Harsh, et al.
Published: (2023)
by: Raj, Harsh, et al.
Published: (2023)
GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
by: Batzner, Jan, et al.
Published: (2024)
by: Batzner, Jan, et al.
Published: (2024)
Understanding Language Model Circuits through Knowledge Editing
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
Improving Consistency in Large Language Models through Chain of Guidance
by: Raj, Harsh, et al.
Published: (2025)
by: Raj, Harsh, et al.
Published: (2025)
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
by: Huang, Xi Yu, et al.
Published: (2024)
by: Huang, Xi Yu, et al.
Published: (2024)
Auxiliary Knowledge-Induced Learning for Automatic Multi-Label Medical Document Classification
by: Wang, Xindi, et al.
Published: (2024)
by: Wang, Xindi, et al.
Published: (2024)
Scenarios and Approaches for Situated Natural Language Explanations
by: Qiu, Pengshuo, et al.
Published: (2024)
by: Qiu, Pengshuo, et al.
Published: (2024)
Exploring the features used for summary evaluation by Human and GPT
by: Sadeghi, Zahra, et al.
Published: (2025)
by: Sadeghi, Zahra, et al.
Published: (2025)
Consistency in Language Models: Current Landscape, Challenges, and Future Directions
by: Novikova, Jekaterina, et al.
Published: (2025)
by: Novikova, Jekaterina, et al.
Published: (2025)
Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning
by: Netík, Jan, et al.
Published: (2026)
by: Netík, Jan, et al.
Published: (2026)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
by: Berlot-Attwell, Ian, et al.
Published: (2024)
by: Berlot-Attwell, Ian, et al.
Published: (2024)
Multi-stage Retrieve and Re-rank Model for Automatic Medical Coding Recommendation
by: Wang, Xindi, et al.
Published: (2024)
by: Wang, Xindi, et al.
Published: (2024)
LLM Library Learning Fails: A LEGO-Prover Case Study
by: Berlot-Attwell, Ian, et al.
Published: (2025)
by: Berlot-Attwell, Ian, et al.
Published: (2025)
Whose Personae? Synthetic Persona Experiments in LLM Research and Pathways to Transparency
by: Batzner, Jan, et al.
Published: (2025)
by: Batzner, Jan, et al.
Published: (2025)
Can adversarial attacks by large language models be attributed?
by: Cebrian, Manuel, et al.
Published: (2024)
by: Cebrian, Manuel, et al.
Published: (2024)
LLM-Generated Black-box Explanations Can Be Adversarially Helpful
by: Ajwani, Rohan, et al.
Published: (2024)
by: Ajwani, Rohan, et al.
Published: (2024)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
by: Jaipersaud, Brandon, et al.
Published: (2024)
by: Jaipersaud, Brandon, et al.
Published: (2024)
Complexity-aware fine-tuning
by: Goncharov, Andrey, et al.
Published: (2025)
by: Goncharov, Andrey, et al.
Published: (2025)
Temporal fine-tuning for early risk detection
by: Thompson, Horacio, et al.
Published: (2025)
by: Thompson, Horacio, et al.
Published: (2025)
Graph-tree Fusion Model with Bidirectional Information Propagation for Long Document Classification
by: Roy, Sudipta Singha, et al.
Published: (2024)
by: Roy, Sudipta Singha, et al.
Published: (2024)
Labeling supervised fine-tuning data with the scaling law
by: Kong, Huanjun
Published: (2024)
by: Kong, Huanjun
Published: (2024)
Understanding Syntactic Generalization in Structure-inducing Language Models
by: Arps, David, et al.
Published: (2025)
by: Arps, David, et al.
Published: (2025)
Discovering Salient Neurons in Deep NLP Models
by: Durrani, Nadir, et al.
Published: (2022)
by: Durrani, Nadir, et al.
Published: (2022)
Quantifying the Capabilities of LLMs across Scale and Precision
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
Interpreting the Effects of Quantization on LLMs
by: Singh, Manpreet, et al.
Published: (2025)
by: Singh, Manpreet, et al.
Published: (2025)
Understanding the effects of language-specific class imbalance in multilingual fine-tuning
by: Jung, Vincent, et al.
Published: (2024)
by: Jung, Vincent, et al.
Published: (2024)
Improving LLM-based Ontology Matching with fine-tuning on synthetic data
by: Sousa, Guilherme, et al.
Published: (2025)
by: Sousa, Guilherme, et al.
Published: (2025)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
by: Ajwani, Rohan Deepak, et al.
Published: (2024)
by: Ajwani, Rohan Deepak, et al.
Published: (2024)
Retention analysis of edited knowledge after fine-tuning
by: Wen, Fufang, et al.
Published: (2025)
by: Wen, Fufang, et al.
Published: (2025)
Replaying pre-training data improves fine-tuning
by: Kotha, Suhas, et al.
Published: (2026)
by: Kotha, Suhas, et al.
Published: (2026)
Similar Items
-
Representation Noising: A Defence Mechanism Against Harmful Finetuning
by: Rosati, Domenic, et al.
Published: (2024) -
Evaluating Defences against Unsafe Feedback in RLHF
by: Rosati, Domenic, et al.
Published: (2024) -
Resolving Lexical Bias in Model Editing
by: Rizwan, Hammad, et al.
Published: (2024) -
Dependency Parsing is More Parameter-Efficient with Normalization
by: Gajo, Paolo, et al.
Published: (2025) -
Long-form evaluation of model editing
by: Rosati, Domenic, et al.
Published: (2024)