RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Asif, Sadia, Amiri, Mohammad Mohammadi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Trust and Safety in Digital Payments: An LLM-Powered Approach
by: Dahiphale, Devendra, et al.
Published: (2024)
by: Dahiphale, Devendra, et al.
Published: (2024)
Embedding Trust at Scale: Physics-Aware Neural Watermarking for Secure and Verifiable Data Pipelines
by: Tallam, Krti
Published: (2025)
by: Tallam, Krti
Published: (2025)
Thought Purity: A Defense Framework For Chain-of-Thought Attack
by: Xue, Zihao, et al.
Published: (2025)
by: Xue, Zihao, et al.
Published: (2025)
Malware Classification using Diluted Convolutional Neural Network with Fast Gradient Sign Method
by: Anand, Ashish, et al.
Published: (2026)
by: Anand, Ashish, et al.
Published: (2026)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
by: Lan, Wenhao, et al.
Published: (2026)
by: Lan, Wenhao, et al.
Published: (2026)
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
by: Manczak, Blazej, et al.
Published: (2024)
by: Manczak, Blazej, et al.
Published: (2024)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)
by: Hubinger, Evan, et al.
Published: (2024)
Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward
by: Xie, Xuan, et al.
Published: (2024)
by: Xie, Xuan, et al.
Published: (2024)
AMLNet: A Knowledge-Based Multi-Agent Framework to Generate and Detect Realistic Money Laundering Transactions
by: Huda, Sabin, et al.
Published: (2025)
by: Huda, Sabin, et al.
Published: (2025)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
by: Poppi, Samuele, et al.
Published: (2024)
by: Poppi, Samuele, et al.
Published: (2024)
Federated Graph AGI for Cross-Border Insider Threat Intelligence in Government Financial Schemes
by: Nayak, Srikumar, et al.
Published: (2026)
by: Nayak, Srikumar, et al.
Published: (2026)
Systematic Review of Cybersecurity in Banking: Evolution from Pre-Industry 4.0 to Post-Industry 4.0 in Artificial Intelligence, Blockchain, Policies and Practice
by: Tran, Tue Nhi
Published: (2025)
by: Tran, Tue Nhi
Published: (2025)
COALESCE: Economic and Security Dynamics of Skill-Based Task Outsourcing Among Team of Autonomous LLM Agents
by: Bhatt, Manish, et al.
Published: (2025)
by: Bhatt, Manish, et al.
Published: (2025)
Zero-Knowledge Proof (ZKP) Authentication for Offline CBDC Payment System Using IoT Devices
by: Mondal, Santanu, et al.
Published: (2026)
by: Mondal, Santanu, et al.
Published: (2026)
Decoupling Identity from Utility: Privacy-by-Design Frameworks for Financial Ecosystems
by: Ibikunle, Ifayoyinsola, et al.
Published: (2026)
by: Ibikunle, Ifayoyinsola, et al.
Published: (2026)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
by: Hasan, Adib, et al.
Published: (2024)
by: Hasan, Adib, et al.
Published: (2024)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
by: Frikha, Ahmed, et al.
Published: (2024)
by: Frikha, Ahmed, et al.
Published: (2024)
Tempora-Fusion: Time-Lock Puzzle with Efficient Verifiable Homomorphic Linear Combination
by: Abadi, Aydin
Published: (2024)
by: Abadi, Aydin
Published: (2024)
SolRPDS: A Dataset for Analyzing Rug Pulls in Solana Decentralized Finance
by: Alhaidari, Abdulrahman, et al.
Published: (2025)
by: Alhaidari, Abdulrahman, et al.
Published: (2025)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
Continual Gradient Low-Rank Projection Fine-Tuning for LLMs
by: Wang, Chenxu, et al.
Published: (2025)
by: Wang, Chenxu, et al.
Published: (2025)
Improving the Accuracy of Transaction-Based Ponzi Detection on Ethereum
by: Huynh, Phuong Duy, et al.
Published: (2023)
by: Huynh, Phuong Duy, et al.
Published: (2023)
An adaptive network-based approach for advanced forecasting of cryptocurrency values
by: Mehrban, Ali, et al.
Published: (2024)
by: Mehrban, Ali, et al.
Published: (2024)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
Secure Code Generation at Scale with Reflexion
by: Datta, Arup, et al.
Published: (2025)
by: Datta, Arup, et al.
Published: (2025)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
SentinelSphere: Integrating AI-Powered Real-Time Threat Detection with Cybersecurity Awareness Training
by: Tantaroudas, Nikolaos D., et al.
Published: (2026)
by: Tantaroudas, Nikolaos D., et al.
Published: (2026)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Bitcoin's Edge: Embedded Sentiment in Blockchain Transactional Data
by: Kleitsikas, Charalampos, et al.
Published: (2025)
by: Kleitsikas, Charalampos, et al.
Published: (2025)
Power to the Clients: Federated Learning in a Dictatorship Setting
by: Alipour, Mohammadsajad, et al.
Published: (2025)
by: Alipour, Mohammadsajad, et al.
Published: (2025)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023)
by: Vega, Jason, et al.
Published: (2023)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation
by: Li, Xianzhi, et al.
Published: (2024)
by: Li, Xianzhi, et al.
Published: (2024)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
by: Yang, Xiaoxue, et al.
Published: (2025)
by: Yang, Xiaoxue, et al.
Published: (2025)
Demystifying Domain-adaptive Post-training for Financial LLMs
by: Ke, Zixuan, et al.
Published: (2025)
by: Ke, Zixuan, et al.
Published: (2025)
CoSMeTIC: Zero-Knowledge Computational Sparse Merkle Trees with Inclusion-Exclusion Proofs for Clinical Research
by: Shahid, Mohammad, et al.
Published: (2026)
by: Shahid, Mohammad, et al.
Published: (2026)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
by: Wen, Xiaofei, et al.
Published: (2025)
by: Wen, Xiaofei, et al.
Published: (2025)
Similar Items
-
Enhancing Trust and Safety in Digital Payments: An LLM-Powered Approach
by: Dahiphale, Devendra, et al.
Published: (2024) -
Embedding Trust at Scale: Physics-Aware Neural Watermarking for Secure and Verifiable Data Pipelines
by: Tallam, Krti
Published: (2025) -
Thought Purity: A Defense Framework For Chain-of-Thought Attack
by: Xue, Zihao, et al.
Published: (2025) -
Malware Classification using Diluted Convolutional Neural Network with Fast Gradient Sign Method
by: Anand, Ashish, et al.
Published: (2026) -
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)