Representation Noising: A Defence Mechanism Against Harmful Finetuning
Fuente:
arXiv
Saved in:
| Main Authors: | Rosati, Domenic, Wehner, Jan, Williams, Kai, Bartoszcze, Łukasz, Atanasov, David, Gonzales, Robie, Majumdar, Subhabrata, Maple, Carsten, Sajjad, Hassan, Rudzicz, Frank |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Immunization against harmful fine-tuning attacks
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
Limits of Convergence-Rate Control for Open-Weight Safety
by: Rosati, Domenic, et al.
Published: (2026)
by: Rosati, Domenic, et al.
Published: (2026)
Long-form evaluation of model editing
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
Improving Consistency in Large Language Models through Chain of Guidance
by: Raj, Harsh, et al.
Published: (2025)
by: Raj, Harsh, et al.
Published: (2025)
Semantic Consistency for Assuring Reliability of Large Language Models
by: Raj, Harsh, et al.
Published: (2023)
by: Raj, Harsh, et al.
Published: (2023)
Representation Engineering for Large-Language Models: Survey and Research Challenges
by: Bartoszcze, Lukasz, et al.
Published: (2025)
by: Bartoszcze, Lukasz, et al.
Published: (2025)
Consistency in Language Models: Current Landscape, Challenges, and Future Directions
by: Novikova, Jekaterina, et al.
Published: (2025)
by: Novikova, Jekaterina, et al.
Published: (2025)
Resolving Lexical Bias in Model Editing
by: Rizwan, Hammad, et al.
Published: (2024)
by: Rizwan, Hammad, et al.
Published: (2024)
Dependency Parsing is More Parameter-Efficient with Normalization
by: Gajo, Paolo, et al.
Published: (2025)
by: Gajo, Paolo, et al.
Published: (2025)
LLMs Underperform Graph-Based Parsers on Supervised Relation Extraction for Complex Graphs
by: Gajo, Paolo, et al.
Published: (2026)
by: Gajo, Paolo, et al.
Published: (2026)
A Game-Theoretic Approach for PMU Deployment Against False Data Injection Attacks
by: Maleki, Sajjad, et al.
Published: (2024)
by: Maleki, Sajjad, et al.
Published: (2024)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025)
by: Davies, Xander, et al.
Published: (2025)
Embedding-based classifiers can detect prompt injection attacks
by: Ayub, Md. Ahsan, et al.
Published: (2024)
by: Ayub, Md. Ahsan, et al.
Published: (2024)
Voluntary Collusion with Secret Tools in Competing LLM Agents
by: Zeng, Xijie, et al.
Published: (2026)
by: Zeng, Xijie, et al.
Published: (2026)
Individualised Counterfactual Examples Using Conformal Prediction Intervals
by: Adams, James M., et al.
Published: (2025)
by: Adams, James M., et al.
Published: (2025)
Testing Realism in Quantum Mechanics Through Charge Conservation
by: Atanasov, Victor
Published: (2025)
by: Atanasov, Victor
Published: (2025)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
Quantum Advantage for Coordinated Frequency Selection Against Distributed Jammers
by: Wehner, Stephanie
Published: (2026)
by: Wehner, Stephanie
Published: (2026)
Elicitor Specific Mechanisms of Defence Priming in Oak Seedlings Against Powdery Mildew
by: Rosa Sanchez‐Lucas, et al.
Published: (2025)
by: Rosa Sanchez‐Lucas, et al.
Published: (2025)
Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
by: Maple, Carsten, et al.
Published: (2026)
by: Maple, Carsten, et al.
Published: (2026)
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges
by: Tapwal, Riya, et al.
Published: (2026)
by: Tapwal, Riya, et al.
Published: (2026)
PRISM: Generation-Time Detection and Mitigation of Secret Leakage in Multi-Agent LLM Pipelines
by: Tapwal, Riya, et al.
Published: (2026)
by: Tapwal, Riya, et al.
Published: (2026)
Justified Evidence Collection for Argument-based AI Fairness Assurance
by: Sabuncuoglu, Alpay, et al.
Published: (2025)
by: Sabuncuoglu, Alpay, et al.
Published: (2025)
Towards Robust Federated Analytics via Differentially Private Measurements of Statistical Heterogeneity
by: Scott, Mary, et al.
Published: (2024)
by: Scott, Mary, et al.
Published: (2024)
Private Federated Multiclass Post-hoc Calibration
by: Maddock, Samuel, et al.
Published: (2025)
by: Maddock, Samuel, et al.
Published: (2025)
FLAIM: AIM-based Synthetic Data Generation in the Federated Setting
by: Maddock, Samuel, et al.
Published: (2023)
by: Maddock, Samuel, et al.
Published: (2023)
DriveSafe: A Hierarchical Risk Taxonomy for Safety-Critical LLM-Based Driving Assistants
by: Kumar, Abhishek, et al.
Published: (2026)
by: Kumar, Abhishek, et al.
Published: (2026)
Wavelet Representation and Sampling of Complex Luminaires
by: A. Atanasov, et al.
Published: (2025)
by: A. Atanasov, et al.
Published: (2025)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Exploring the features used for summary evaluation by Human and GPT
by: Sadeghi, Zahra, et al.
Published: (2025)
by: Sadeghi, Zahra, et al.
Published: (2025)
Scenarios and Approaches for Situated Natural Language Explanations
by: Qiu, Pengshuo, et al.
Published: (2024)
by: Qiu, Pengshuo, et al.
Published: (2024)
Understanding Language Model Circuits through Knowledge Editing
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
SoftAdaClip: A Smooth Clipping Strategy for Fair and Private Model Training
by: Soleymani, Dorsa, et al.
Published: (2025)
by: Soleymani, Dorsa, et al.
Published: (2025)
On the Limitations of Speaker Diarization
by: Joana Amorim, et al.
Published: (2026)
by: Joana Amorim, et al.
Published: (2026)
Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities
by: Nguyen, Ilana, et al.
Published: (2026)
by: Nguyen, Ilana, et al.
Published: (2026)
Nachruf: Hans‐Herbert Schmidtke
by: Mihail Atanasov, et al.
Published: (2024)
by: Mihail Atanasov, et al.
Published: (2024)
Threat, Risk and Mitigation Taxonomy for Digital Identity Systems
by: SHEIK, AL TARIQ, et al.
Published: (2024)
by: SHEIK, AL TARIQ, et al.
Published: (2024)
Unveiling the Interplay: Salinity‐Modulated Defence Mechanisms in Medicago truncatula Against Phoma medicaginis Infection
by: Manel Chaouachi, et al.
Published: (2025)
by: Manel Chaouachi, et al.
Published: (2025)
Similar Items
-
Immunization against harmful fine-tuning attacks
by: Rosati, Domenic, et al.
Published: (2024) -
Evaluating Defences against Unsafe Feedback in RLHF
by: Rosati, Domenic, et al.
Published: (2024) -
Limits of Convergence-Rate Control for Open-Weight Safety
by: Rosati, Domenic, et al.
Published: (2026) -
Long-form evaluation of model editing
by: Rosati, Domenic, et al.
Published: (2024) -
Improving Consistency in Large Language Models through Chain of Guidance
by: Raj, Harsh, et al.
Published: (2025)