Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
Fuente:
arXiv
Saved in:
| Main Authors: | Shairah, Harethah Abu, Hammoud, Hasan Abed Al Kader, Turkiyyah, George, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
by: Alssum, Lama, et al.
Published: (2025)
by: Alssum, Lama, et al.
Published: (2025)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
On the Importance of Pretraining Data Alignment for Atomic Property Prediction
by: Ghunaim, Yasir, et al.
Published: (2025)
by: Ghunaim, Yasir, et al.
Published: (2025)
Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
by: Zbeeb, Mohammad, et al.
Published: (2025)
by: Zbeeb, Mohammad, et al.
Published: (2025)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
TAPS: Task Aware Proposal Distributions for Speculative Sampling
by: Zbib, Mohamad, et al.
Published: (2026)
by: Zbib, Mohamad, et al.
Published: (2026)
ALARB: An Arabic Legal Argument Reasoning Benchmark
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
by: Hijazi, Faris, et al.
Published: (2024)
by: Hijazi, Faris, et al.
Published: (2024)
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
by: Zbeeb, Mohammad, et al.
Published: (2025)
by: Zbeeb, Mohammad, et al.
Published: (2025)
QuanBench+: A Unified Multi-Framework Benchmark for LLM-Based Quantum Code Generation
by: Slim, Ali, et al.
Published: (2026)
by: Slim, Ali, et al.
Published: (2026)
On Pretraining Data Diversity for Self-Supervised Learning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
by: Alssum, Lama, et al.
Published: (2025)
by: Alssum, Lama, et al.
Published: (2025)
AraSpell: A Deep Learning Approach for Arabic Spelling Correction
by: Salhab, Mahmoud, et al.
Published: (2024)
by: Salhab, Mahmoud, et al.
Published: (2024)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
by: Hinojosa, Carlos, et al.
Published: (2026)
by: Hinojosa, Carlos, et al.
Published: (2026)
From Categories to Classifiers: Name-Only Continual Learning by Exploring the Web
by: Prabhu, Ameya, et al.
Published: (2023)
by: Prabhu, Ameya, et al.
Published: (2023)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
by: Wang, Dianyun, et al.
Published: (2025)
by: Wang, Dianyun, et al.
Published: (2025)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
SpellForger: Prompting Custom Spell Properties In-Game using BERT supervised-trained model
by: Silva, Emanuel C., et al.
Published: (2025)
by: Silva, Emanuel C., et al.
Published: (2025)
Don't Pay Attention
by: Hammoud, Mohammad, et al.
Published: (2025)
by: Hammoud, Mohammad, et al.
Published: (2025)
Avey-B
by: Acharya, Devang, et al.
Published: (2026)
by: Acharya, Devang, et al.
Published: (2026)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Multimodal Safety Evaluation in Generative Agent Social Simulations
by: Vera, Alhim, et al.
Published: (2025)
by: Vera, Alhim, et al.
Published: (2025)
Safety Alignment via Constrained Knowledge Unlearning
by: Shi, Zesheng, et al.
Published: (2025)
by: Shi, Zesheng, et al.
Published: (2025)
Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
by: Agarwal, Rajan, et al.
Published: (2025)
by: Agarwal, Rajan, et al.
Published: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
Retrieval Augmented Spelling Correction for E-Commerce Applications
by: Guo, Xuan, et al.
Published: (2024)
by: Guo, Xuan, et al.
Published: (2024)
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment
by: Noël, Valentin, et al.
Published: (2025)
by: Noël, Valentin, et al.
Published: (2025)
Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study
by: Yaacoub, Antoun, et al.
Published: (2025)
by: Yaacoub, Antoun, et al.
Published: (2025)
OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations
by: Chen, Jiangwang, et al.
Published: (2026)
by: Chen, Jiangwang, et al.
Published: (2026)
Randomized Asymmetric Chain of LoRA: The First Meaningful Theoretical Framework for Low-Rank Adaptation
by: Malinovsky, Grigory, et al.
Published: (2024)
by: Malinovsky, Grigory, et al.
Published: (2024)
Preference Ranking Optimization for Human Alignment
by: Song, Feifan, et al.
Published: (2023)
by: Song, Feifan, et al.
Published: (2023)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
Language Ranker: A Lightweight Ranking framework for LLM Decoding
by: Zhang, Chenheng, et al.
Published: (2025)
by: Zhang, Chenheng, et al.
Published: (2025)
Similar Items
-
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
by: Shairah, Harethah Abu, et al.
Published: (2025) -
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
by: Alssum, Lama, et al.
Published: (2025) -
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024) -
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025) -
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)