What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jain, Samyak, Lubana, Ekdeep Singh, Oksuz, Kemal, Joy, Tom, Torr, Philip H. S., Sanyal, Amartya, Dokania, Puneet K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoCaE: Mixture of Calibrated Experts Significantly Improves Object Detection
von: Oksuz, Kemal, et al.
Veröffentlicht: (2023)
von: Oksuz, Kemal, et al.
Veröffentlicht: (2023)
Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
Fine-tuning can cripple your foundation model; preserving features may be the solution
von: Mukhoti, Jishnu, et al.
Veröffentlicht: (2023)
von: Mukhoti, Jishnu, et al.
Veröffentlicht: (2023)
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
On Calibration of Object Detectors: Pitfalls, Evaluation and Baselines
von: Kuzucu, Selim, et al.
Veröffentlicht: (2024)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2024)
Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
von: Oksuz, Kemal, et al.
Veröffentlicht: (2025)
von: Oksuz, Kemal, et al.
Veröffentlicht: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2024)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2024)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
Placing Objects in Context via Inpainting for Out-of-distribution Segmentation
von: de Jorge, Pau, et al.
Veröffentlicht: (2024)
von: de Jorge, Pau, et al.
Veröffentlicht: (2024)
Random Representations Outperform Online Continually Learned Representations
von: Prabhu, Ameya, et al.
Veröffentlicht: (2024)
von: Prabhu, Ameya, et al.
Veröffentlicht: (2024)
Corrective Machine Unlearning
von: Goel, Shashwat, et al.
Veröffentlicht: (2024)
von: Goel, Shashwat, et al.
Veröffentlicht: (2024)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
von: Zur, Amir, et al.
Veröffentlicht: (2025)
von: Zur, Amir, et al.
Veröffentlicht: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
von: Pres, Itamar, et al.
Veröffentlicht: (2024)
von: Pres, Itamar, et al.
Veröffentlicht: (2024)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
von: Menon, Abhinav, et al.
Veröffentlicht: (2024)
von: Menon, Abhinav, et al.
Veröffentlicht: (2024)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
von: Ramesh, Rahul, et al.
Veröffentlicht: (2023)
von: Ramesh, Rahul, et al.
Veröffentlicht: (2023)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
von: Lubana, Ekdeep Singh, et al.
Veröffentlicht: (2024)
von: Lubana, Ekdeep Singh, et al.
Veröffentlicht: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
von: Okawa, Maya, et al.
Veröffentlicht: (2023)
von: Okawa, Maya, et al.
Veröffentlicht: (2023)
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
von: Nishi, Kento, et al.
Veröffentlicht: (2024)
von: Nishi, Kento, et al.
Veröffentlicht: (2024)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
Mixture of Experts Made Intrinsically Interpretable
von: Yang, Xingyi, et al.
Veröffentlicht: (2025)
von: Yang, Xingyi, et al.
Veröffentlicht: (2025)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
An Iterative Algorithm for Differentially Private $k$-PCA with Adaptive Noise
von: Düngler, Johanna, et al.
Veröffentlicht: (2025)
von: Düngler, Johanna, et al.
Veröffentlicht: (2025)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
Modeling Gene Expression Distributional Shifts for Unseen Genetic Perturbations
von: Ramakrishnan, Kalyan, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Kalyan, et al.
Veröffentlicht: (2025)
Mechanistic Fine-tuning for In-context Learning
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
AnnoCaseLaw: A Richly-Annotated Dataset For Benchmarking Explainable Legal Judgment Prediction
von: Sesodia, Magnus, et al.
Veröffentlicht: (2025)
von: Sesodia, Magnus, et al.
Veröffentlicht: (2025)
Force Dipole Interactions in Tubular Fluid Membranes
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
Nuclear stability and the Fold Catastrophe
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
Tunneling half-lives in macroscopic-microscopic picture
von: Jain, Samyak, et al.
Veröffentlicht: (2024)
von: Jain, Samyak, et al.
Veröffentlicht: (2024)
Catastrophe theoretic approach to the Higgs Mechanism
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
In-Context Learning Dynamics with Random Binary Sequences
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023)
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023)
Swing-by Dynamics in Concept Learning and Compositional Generalization
von: Yang, Yongyi, et al.
Veröffentlicht: (2024)
von: Yang, Yongyi, et al.
Veröffentlicht: (2024)
Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
Attention-Enhanced Hybrid Feature Aggregation Network for 3D Brain Tumor Segmentation
von: Yazıcı, Ziya Ata, et al.
Veröffentlicht: (2024)
von: Yazıcı, Ziya Ata, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MoCaE: Mixture of Calibrated Experts Significantly Improves Object Detection
von: Oksuz, Kemal, et al.
Veröffentlicht: (2023) -
Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
von: Eiras, Francisco, et al.
Veröffentlicht: (2023) -
Fine-tuning can cripple your foundation model; preserving features may be the solution
von: Mukhoti, Jishnu, et al.
Veröffentlicht: (2023) -
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
von: Jain, Samyak, et al.
Veröffentlicht: (2023) -
On Calibration of Object Detectors: Pitfalls, Evaluation and Baselines
von: Kuzucu, Selim, et al.
Veröffentlicht: (2024)