Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Youstra, Jack, Mahfoud, Mohammed, Yan, Yang, Sleight, Henry, Perez, Ethan, Sharma, Mrinank |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
by: Guo, Shiyuan, et al.
Published: (2025)
by: Guo, Shiyuan, et al.
Published: (2025)
Forecasting Rare Language Model Behaviors
by: Jones, Erik, et al.
Published: (2025)
by: Jones, Erik, et al.
Published: (2025)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
by: Peng, Alwin, et al.
Published: (2024)
by: Peng, Alwin, et al.
Published: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment
by: Wang, Haozhong, et al.
Published: (2026)
by: Wang, Haozhong, et al.
Published: (2026)
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
by: Wang, Tony T., et al.
Published: (2024)
by: Wang, Tony T., et al.
Published: (2024)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
by: Ensign, Danielle, et al.
Published: (2025)
by: Ensign, Danielle, et al.
Published: (2025)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
by: Zhou, Huichi, et al.
Published: (2025)
by: Zhou, Huichi, et al.
Published: (2025)
Safeguarding Graph Neural Networks against Topology Inference Attacks
by: Fu, Jie, et al.
Published: (2025)
by: Fu, Jie, et al.
Published: (2025)
Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
by: Chen, Runjin, et al.
Published: (2025)
by: Chen, Runjin, et al.
Published: (2025)
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
Log Probability Tracking of LLM APIs
by: Chauvin, Timothée, et al.
Published: (2025)
by: Chauvin, Timothée, et al.
Published: (2025)
Analyzing the Effect of Noise in LLM Fine-tuning
by: Li, Lingfang, et al.
Published: (2026)
by: Li, Lingfang, et al.
Published: (2026)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Are LLM-Enhanced Graph Neural Networks Robust against Poisoning Attacks?
by: Ma, Yuhang, et al.
Published: (2026)
by: Ma, Yuhang, et al.
Published: (2026)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
by: Poppi, Samuele, et al.
Published: (2024)
by: Poppi, Samuele, et al.
Published: (2024)
Adversarial Inception Backdoor Attacks against Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
Incorporating Unlabelled Data into Bayesian Neural Networks
by: Sharma, Mrinank, et al.
Published: (2023)
by: Sharma, Mrinank, et al.
Published: (2023)
Token-Efficient Change Detection in LLM APIs
by: Chauvin, Timothée, et al.
Published: (2026)
by: Chauvin, Timothée, et al.
Published: (2026)
Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
by: Zhang, Jiale, et al.
Published: (2025)
by: Zhang, Jiale, et al.
Published: (2025)
Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Hierarchical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM
by: Yao, Yongqiang, et al.
Published: (2025)
by: Yao, Yongqiang, et al.
Published: (2025)
Detecting Instruction Fine-tuning Attacks using Influence Function
by: Li, Jiawei
Published: (2025)
by: Li, Jiawei
Published: (2025)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
by: Fu, Wenjie, et al.
Published: (2023)
by: Fu, Wenjie, et al.
Published: (2023)
Deep Learning-Assisted Improved Differential Fault Attacks on Lightweight Stream Ciphers
by: Lim, Kok Ping, et al.
Published: (2026)
by: Lim, Kok Ping, et al.
Published: (2026)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025)
by: Davies, Xander, et al.
Published: (2025)
Understanding Fine-tuning in Approximate Unlearning: A Theoretical Perspective
by: Ding, Meng, et al.
Published: (2024)
by: Ding, Meng, et al.
Published: (2024)
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
by: Du, Hao, et al.
Published: (2025)
by: Du, Hao, et al.
Published: (2025)
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Membership Inference Attacks Against Fine-tuned Diffusion Language Models
by: Chen, Yuetian, et al.
Published: (2026)
by: Chen, Yuetian, et al.
Published: (2026)
BoostLLM: Boosting-inspired LLM Fine-tuning for Few-shot Tabular Classification
by: Wang, Yi-Siang, et al.
Published: (2026)
by: Wang, Yi-Siang, et al.
Published: (2026)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization
by: Su, Yang, et al.
Published: (2024)
by: Su, Yang, et al.
Published: (2024)
Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
by: Li, Boheng, et al.
Published: (2025)
by: Li, Boheng, et al.
Published: (2025)
LoRA vs Full Fine-tuning: An Illusion of Equivalence
by: Shuttleworth, Reece, et al.
Published: (2024)
by: Shuttleworth, Reece, et al.
Published: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025)
by: Kurian, Kevin, et al.
Published: (2025)
Similar Items
-
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
by: Guo, Shiyuan, et al.
Published: (2025) -
Forecasting Rare Language Model Behaviors
by: Jones, Erik, et al.
Published: (2025) -
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
by: Peng, Alwin, et al.
Published: (2024) -
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026) -
Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment
by: Wang, Haozhong, et al.
Published: (2026)