SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Djuhera, Aladin, Kadhe, Swanand Ravindra, Ahmed, Farhan, Zawad, Syed, Boche, Holger |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
par: Djuhera, Aladin, et autres
Publié: (2025)
par: Djuhera, Aladin, et autres
Publié: (2025)
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
par: Djuhera, Aladin, et autres
Publié: (2025)
par: Djuhera, Aladin, et autres
Publié: (2025)
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
par: Djuhera, Aladin, et autres
Publié: (2026)
par: Djuhera, Aladin, et autres
Publié: (2026)
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
par: Djuhera, Aladin, et autres
Publié: (2025)
par: Djuhera, Aladin, et autres
Publié: (2025)
Against the Monolithic Wireless World Model: Why NextG Needs Composable and Agentic Intelligence
par: Djuhera, Aladin, et autres
Publié: (2026)
par: Djuhera, Aladin, et autres
Publié: (2026)
MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction
par: Djuhera, Aladin, et autres
Publié: (2026)
par: Djuhera, Aladin, et autres
Publié: (2026)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
par: Kadhe, Swanand Ravindra, et autres
Publié: (2024)
par: Kadhe, Swanand Ravindra, et autres
Publié: (2024)
"Don't Do That!": Guiding Embodied Systems through Large Language Model-based Constraint Generation
par: Seffo, Amin, et autres
Publié: (2025)
par: Seffo, Amin, et autres
Publié: (2025)
In-Context Probing for Membership Inference in Fine-Tuned Language Models
par: Lu, Zhexi, et autres
Publié: (2025)
par: Lu, Zhexi, et autres
Publié: (2025)
Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
par: Jiang, Shuli, et autres
Publié: (2024)
par: Jiang, Shuli, et autres
Publié: (2024)
R-MTLLMF: Resilient Multi-Task Large Language Model Fusion at the Wireless Edge
par: Djuhera, Aladin, et autres
Publié: (2024)
par: Djuhera, Aladin, et autres
Publié: (2024)
R-SFLLM: Jamming Resilient Framework for Split Federated Learning with Large Language Models
par: Djuhera, Aladin, et autres
Publié: (2024)
par: Djuhera, Aladin, et autres
Publié: (2024)
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
par: Koch, Fernando, et autres
Publié: (2025)
par: Koch, Fernando, et autres
Publié: (2025)
SCoTT: Strategic Chain-of-Thought Tasking for Wireless-Aware Robot Navigation in Digital Twins
par: Djuhera, Aladin, et autres
Publié: (2024)
par: Djuhera, Aladin, et autres
Publié: (2024)
An Analysis of Capacity-Distortion Trade-Offs in Memoryless ISAC Systems
par: Li, Xinyang, et autres
Publié: (2024)
par: Li, Xinyang, et autres
Publié: (2024)
Joint Partitioning and Placement of Foundation Models for Real-Time Edge AI
par: Djuhera, Aladin, et autres
Publié: (2025)
par: Djuhera, Aladin, et autres
Publié: (2025)
Resilient-By-Design Framework for MIMO-OFDM Communications under Smart Jamming
par: Andrei, Vlad C., et autres
Publié: (2024)
par: Andrei, Vlad C., et autres
Publié: (2024)
A Simultaneous Decoding Approach to Joint State and Message Communications
par: Li, Xinyang, et autres
Publié: (2025)
par: Li, Xinyang, et autres
Publié: (2025)
Evaluating the Dynamics of Membership Privacy in Deep Learning
par: Chen, Yuetian, et autres
Publié: (2025)
par: Chen, Yuetian, et autres
Publié: (2025)
STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
par: An, Sungeun, et autres
Publié: (2026)
par: An, Sungeun, et autres
Publié: (2026)
MERGE$^3$: Efficient Evolutionary Merging on Consumer-grade GPUs
par: Mencattini, Tommaso, et autres
Publié: (2025)
par: Mencattini, Tommaso, et autres
Publié: (2025)
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
par: Hossain, Saad, et autres
Publié: (2025)
par: Hossain, Saad, et autres
Publié: (2025)
Towards a Re-evaluation of Data Forging Attacks in Practice
par: Suliman, Mohamed, et autres
Publié: (2024)
par: Suliman, Mohamed, et autres
Publié: (2024)
Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation
par: Liu, Guozhi, et autres
Publié: (2024)
par: Liu, Guozhi, et autres
Publié: (2024)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
par: Lu, Ning, et autres
Publié: (2025)
par: Lu, Ning, et autres
Publié: (2025)
GR-SAP: Generative Replay for Safety Alignment Preservation during Fine-Tuning
par: Fang, Zhouxiang, et autres
Publié: (2026)
par: Fang, Zhouxiang, et autres
Publié: (2026)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
par: Roy, Aniruddha, et autres
Publié: (2025)
par: Roy, Aniruddha, et autres
Publié: (2025)
Know When To Fold 'Em: Token-Efficient LLM Synthetic Data Generation via Multi-Stage In-Flight Rejection
par: Chowdhury, Anjir Ahmed, et autres
Publié: (2026)
par: Chowdhury, Anjir Ahmed, et autres
Publié: (2026)
Hierarchical Alignment: Surgical Fine-Tuning via Functional Layer Specialization in Large Language Models
par: Zhang, Yukun, et autres
Publié: (2025)
par: Zhang, Yukun, et autres
Publié: (2025)
Membership Inference Attacks as Privacy Tools: Reliability, Disparity and Ensemble
par: Wang, Zhiqi, et autres
Publié: (2025)
par: Wang, Zhiqi, et autres
Publié: (2025)
Safety-Aware Fine-Tuning of Large Language Models
par: Choi, Hyeong Kyu, et autres
Publié: (2024)
par: Choi, Hyeong Kyu, et autres
Publié: (2024)
Understanding and Preserving Safety in Fine-Tuned LLMs
par: Zhang, Jiawen, et autres
Publié: (2026)
par: Zhang, Jiawen, et autres
Publié: (2026)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
par: Alssum, Lama, et autres
Publié: (2025)
par: Alssum, Lama, et autres
Publié: (2025)
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
par: Wang, Zhaoxin, et autres
Publié: (2026)
par: Wang, Zhaoxin, et autres
Publié: (2026)
Preserving Diversity in Supervised Fine-Tuning of Large Language Models
par: Li, Ziniu, et autres
Publié: (2024)
par: Li, Ziniu, et autres
Publié: (2024)
Non-Uniform Parameter-Wise Model Merging
par: Camacho, Albert Manuel Orozco, et autres
Publié: (2024)
par: Camacho, Albert Manuel Orozco, et autres
Publié: (2024)
PoliTune: Analyzing the Impact of Data Selection and Fine-Tuning on Economic and Political Biases in Large Language Models
par: Agiza, Ahmed, et autres
Publié: (2024)
par: Agiza, Ahmed, et autres
Publié: (2024)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
par: Hammoud, Hasan Abed Al Kader, et autres
Publié: (2024)
par: Hammoud, Hasan Abed Al Kader, et autres
Publié: (2024)
PPFS: Predictive Permutation Feature Selection
par: Hassan, Atif, et autres
Publié: (2021)
par: Hassan, Atif, et autres
Publié: (2021)
Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
par: Zhang, Dengming, et autres
Publié: (2025)
par: Zhang, Dengming, et autres
Publié: (2025)
Documents similaires
-
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
par: Djuhera, Aladin, et autres
Publié: (2025) -
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
par: Djuhera, Aladin, et autres
Publié: (2025) -
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
par: Djuhera, Aladin, et autres
Publié: (2026) -
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
par: Djuhera, Aladin, et autres
Publié: (2025) -
Against the Monolithic Wireless World Model: Why NextG Needs Composable and Agentic Intelligence
par: Djuhera, Aladin, et autres
Publié: (2026)