Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hammoud, Hasan Abed Al Kader, Michieli, Umberto, Pizzati, Fabio, Torr, Philip, Bibi, Adel, Ghanem, Bernard, Ozay, Mete |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
On Pretraining Data Diversity for Self-Supervised Learning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
DiffCLIP: Differential Attention Meets CLIP
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
On the Importance of Pretraining Data Alignment for Atomic Property Prediction
von: Ghunaim, Yasir, et al.
Veröffentlicht: (2025)
von: Ghunaim, Yasir, et al.
Veröffentlicht: (2025)
From Categories to Classifiers: Name-Only Continual Learning by Exploring the Web
von: Prabhu, Ameya, et al.
Veröffentlicht: (2023)
von: Prabhu, Ameya, et al.
Veröffentlicht: (2023)
Randomized Asymmetric Chain of LoRA: The First Meaningful Theoretical Framework for Low-Rank Adaptation
von: Malinovsky, Grigory, et al.
Veröffentlicht: (2024)
von: Malinovsky, Grigory, et al.
Veröffentlicht: (2024)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
HOP to the Next Tasks and Domains for Continual Learning in NLP
von: Michieli, Umberto, et al.
Veröffentlicht: (2024)
von: Michieli, Umberto, et al.
Veröffentlicht: (2024)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
von: Zbeeb, Mohammad, et al.
Veröffentlicht: (2025)
von: Zbeeb, Mohammad, et al.
Veröffentlicht: (2025)
Feature-Space Generative Models for One-Shot Class-Incremental Learning
von: Foster, Jack, et al.
Veröffentlicht: (2026)
von: Foster, Jack, et al.
Veröffentlicht: (2026)
A Model for Every User and Budget: Label-Free and Personalized Mixed-Precision Quantization
von: Fish, Edward, et al.
Veröffentlicht: (2023)
von: Fish, Edward, et al.
Veröffentlicht: (2023)
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
von: Shenaj, Donald, et al.
Veröffentlicht: (2025)
von: Shenaj, Donald, et al.
Veröffentlicht: (2025)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation
von: Yang, Yibo, et al.
Veröffentlicht: (2024)
von: Yang, Yibo, et al.
Veröffentlicht: (2024)
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation
von: Shenaj, Donald, et al.
Veröffentlicht: (2024)
von: Shenaj, Donald, et al.
Veröffentlicht: (2024)
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
von: Camuffo, Elena, et al.
Veröffentlicht: (2025)
von: Camuffo, Elena, et al.
Veröffentlicht: (2025)
TAPS: Task Aware Proposal Distributions for Speculative Sampling
von: Zbib, Mohamad, et al.
Veröffentlicht: (2026)
von: Zbib, Mohamad, et al.
Veröffentlicht: (2026)
Object-conditioned Bag of Instances for Few-Shot Personalized Instance Recognition
von: Michieli, Umberto, et al.
Veröffentlicht: (2024)
von: Michieli, Umberto, et al.
Veröffentlicht: (2024)
Data-driven Clustering and Merging of Adapters for On-device Large Language Models
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2026)
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2026)
Clustering-driven Memory Compression for On-device Large Language Models
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2026)
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2026)
Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization Statistics
von: Camuffo, Elena, et al.
Veröffentlicht: (2024)
von: Camuffo, Elena, et al.
Veröffentlicht: (2024)
Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging
von: An, Bang, et al.
Veröffentlicht: (2025)
von: An, Bang, et al.
Veröffentlicht: (2025)
Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
von: Ceritli, Taha, et al.
Veröffentlicht: (2025)
von: Ceritli, Taha, et al.
Veröffentlicht: (2025)
FFT-based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted Images
von: Camuffo, Elena, et al.
Veröffentlicht: (2024)
von: Camuffo, Elena, et al.
Veröffentlicht: (2024)
Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object Detection
von: Barbato, Francesco, et al.
Veröffentlicht: (2024)
von: Barbato, Francesco, et al.
Veröffentlicht: (2024)
DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching
von: Aiello, Emanuele, et al.
Veröffentlicht: (2024)
von: Aiello, Emanuele, et al.
Veröffentlicht: (2024)
Controllable Forgetting Mechanism for Few-Shot Class-Incremental Learning
von: Paramonov, Kirill, et al.
Veröffentlicht: (2025)
von: Paramonov, Kirill, et al.
Veröffentlicht: (2025)
Deep Neural Network Models Trained With A Fixed Random Classifier Transfer Better Across Domains
von: Ali, Hafiz Tiomoko, et al.
Veröffentlicht: (2024)
von: Ali, Hafiz Tiomoko, et al.
Veröffentlicht: (2024)
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
von: Zbeeb, Mohammad, et al.
Veröffentlicht: (2025)
von: Zbeeb, Mohammad, et al.
Veröffentlicht: (2025)
Efficient Compositional Multi-tasking for On-device Large Language Models
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2025)
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2025)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
A Modular System for Enhanced Robustness of Multimedia Understanding Networks via Deep Parametric Estimation
von: Barbato, Francesco, et al.
Veröffentlicht: (2024)
von: Barbato, Francesco, et al.
Veröffentlicht: (2024)
Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object Search
von: Paramonov, Kirill, et al.
Veröffentlicht: (2024)
von: Paramonov, Kirill, et al.
Veröffentlicht: (2024)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
von: Oldfield, James, et al.
Veröffentlicht: (2025)
von: Oldfield, James, et al.
Veröffentlicht: (2025)
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
von: Hosseini, Peyman, et al.
Veröffentlicht: (2025)
von: Hosseini, Peyman, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
von: Alssum, Lama, et al.
Veröffentlicht: (2025) -
On Pretraining Data Diversity for Self-Supervised Learning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024) -
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024) -
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025) -
DiffCLIP: Differential Attention Meets CLIP
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)