Secure LLM Fine-Tuning via Safety-Aware Probing
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Chengcan, Zhang, Zhixin, Wei, Zeming, Zhang, Yihao, Luan, Xiaokun, Sun, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
by: Zhang, Zhixin, et al.
Published: (2025)
by: Zhang, Zhixin, et al.
Published: (2025)
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Boosting Jailbreak Attack with Momentum
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
Automata-Based Steering of Large Language Models for Diverse Structured Generation
by: Luan, Xiaokun, et al.
Published: (2025)
by: Luan, Xiaokun, et al.
Published: (2025)
Exploring the Robustness of In-Context Learning with Noisy Labels
by: Cheng, Chen, et al.
Published: (2024)
by: Cheng, Chen, et al.
Published: (2024)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Absorber LLM: Harnessing Causal Synchronization for Test-Time Training
by: Zhang, Zhixin, et al.
Published: (2026)
by: Zhang, Zhixin, et al.
Published: (2026)
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
by: Zhang, Yihao, et al.
Published: (2026)
by: Zhang, Yihao, et al.
Published: (2026)
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction
by: Wei, Zeming, et al.
Published: (2025)
by: Wei, Zeming, et al.
Published: (2025)
Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
MILE: A Mutation Testing Framework of In-Context Learning Systems
by: Wei, Zeming, et al.
Published: (2024)
by: Wei, Zeming, et al.
Published: (2024)
DPZero: Private Fine-Tuning of Language Models without Backpropagation
by: Zhang, Liang, et al.
Published: (2023)
by: Zhang, Liang, et al.
Published: (2023)
Improving the Security of United States Elections with Robust Optimization
by: Crimmins, Braden L., et al.
Published: (2023)
by: Crimmins, Braden L., et al.
Published: (2023)
Identifying and Understanding Cross-Class Features in Adversarial Training
by: Wei, Zeming, et al.
Published: (2025)
by: Wei, Zeming, et al.
Published: (2025)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
by: Zhang, Junbo, et al.
Published: (2026)
by: Zhang, Junbo, et al.
Published: (2026)
Characterizing the Training Dynamics of Private Fine-tuning with Langevin diffusion
by: Ke, Shuqi, et al.
Published: (2024)
by: Ke, Shuqi, et al.
Published: (2024)
Impossibility Results of Card-Based Protocols via Mathematical Optimization
by: Ikeda, Shunnosuke, et al.
Published: (2025)
by: Ikeda, Shunnosuke, et al.
Published: (2025)
ViSTR-GP: Online Cyberattack Detection via Vision-to-State Tensor Regression and Gaussian Processes in Automated Robotic Operations
by: Aftabi, Navid, et al.
Published: (2025)
by: Aftabi, Navid, et al.
Published: (2025)
Graph Attention Network-based Block Propagation with Optimal AoI and Reputation in Web 3.0
by: Liao, Jiana, et al.
Published: (2024)
by: Liao, Jiana, et al.
Published: (2024)
Clip-and-Verify: Linear Constraint-Driven Domain Clipping for Accelerating Neural Network Verification
by: Zhou, Duo, et al.
Published: (2025)
by: Zhou, Duo, et al.
Published: (2025)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing
by: Sun, Jialong, et al.
Published: (2026)
by: Sun, Jialong, et al.
Published: (2026)
VOW: Verifiable and Oblivious Watermark Detection for Large Language Models
by: Luan, Xiaokun, et al.
Published: (2026)
by: Luan, Xiaokun, et al.
Published: (2026)
Secure Semantic Communications via AI Defenses: Fundamentals, Solutions, and Future Directions
by: Zhang, Lan, et al.
Published: (2026)
by: Zhang, Lan, et al.
Published: (2026)
The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems
by: Zhang, Yihao, et al.
Published: (2026)
by: Zhang, Yihao, et al.
Published: (2026)
Optimizing Scalar Selection in Elliptic Curve Cryptography Using Differential Evolution for Enhanced Security
by: Haider, Takreem
Published: (2025)
by: Haider, Takreem
Published: (2025)
Information-Theoretic Digital Twins for Stealthy Attack Detection in Industrial Control Systems: A Closed-Form KL Divergence Approach
by: Kreso, Inda, et al.
Published: (2026)
by: Kreso, Inda, et al.
Published: (2026)
Fuzzy Mathematical Model For Optimizing Success Criteria Of Projects: A Project Management Application
by: Sammany, Mohammad, et al.
Published: (2024)
by: Sammany, Mohammad, et al.
Published: (2024)
Differentially Private Decentralized Optimization with Relay Communication
by: Wang, Luqing, et al.
Published: (2022)
by: Wang, Luqing, et al.
Published: (2022)
Advanced Kernel Search approach for the MST Problem with conflicts involving affinity detection and initial solution construction
by: Carrabs, Francesco, et al.
Published: (2024)
by: Carrabs, Francesco, et al.
Published: (2024)
A Defender-Attacker-Defender Model for Optimizing the Resilience of Hospital Networks to Cyberattacks
by: Helfrich, Stephan, et al.
Published: (2026)
by: Helfrich, Stephan, et al.
Published: (2026)
Differentially Private Linear Optimization for Multi-Party Resource Sharing
by: Karaca, Utku, et al.
Published: (2021)
by: Karaca, Utku, et al.
Published: (2021)
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
by: Wang, Kai, et al.
Published: (2025)
by: Wang, Kai, et al.
Published: (2025)
Approximate and Weighted Data Reconstruction Attack in Federated Learning
by: Song, Yongcun, et al.
Published: (2023)
by: Song, Yongcun, et al.
Published: (2023)
Correlated Noise Provably Beats Independent Noise for Differentially Private Learning
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
by: Lu, Guoxin, et al.
Published: (2026)
by: Lu, Guoxin, et al.
Published: (2026)
Similar Items
-
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026) -
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
by: Wu, Chengcan, et al.
Published: (2025) -
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
by: Zhang, Zhixin, et al.
Published: (2025) -
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
by: Zhang, Yihao, et al.
Published: (2024) -
Boosting Jailbreak Attack with Momentum
by: Zhang, Yihao, et al.
Published: (2024)