TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
Fuente:
arXiv
Saved in:
| Main Authors: | Hossain, Saad, Tseng, Tom, Pandey, Punya Syon, Vajpayee, Samanvay, Kowal, Matthew, Nonta, Nayeema, Simko, Samuel, Casper, Stephen, Jin, Zhijing, Pelrine, Kellin, Rambhatla, Sirisha |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Randomized Gradient Subspaces for Efficient Large Language Model Training
by: Rajabi, Sahar, et al.
Published: (2025)
by: Rajabi, Sahar, et al.
Published: (2025)
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
by: Hossain, Saad, et al.
Published: (2025)
by: Hossain, Saad, et al.
Published: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
SubTrack++ : Gradient Subspace Tracking for Scalable LLM Training
by: Rajabi, Sahar, et al.
Published: (2025)
by: Rajabi, Sahar, et al.
Published: (2025)
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026)
by: Pandey, Punya Syon, et al.
Published: (2026)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
by: Harrasse, Abir, et al.
Published: (2025)
by: Harrasse, Abir, et al.
Published: (2025)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
by: Vennemeyer, Daniel, et al.
Published: (2026)
by: Vennemeyer, Daniel, et al.
Published: (2026)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
International. Tampering with Jerusalem
Published: (1998)
Published: (1998)
Can Go AIs be adversarially robust?
by: Tseng, Tom, et al.
Published: (2024)
by: Tseng, Tom, et al.
Published: (2024)
Anti-Tamper Radio meets Reconfigurable Intelligent Surface for System-Level Tamper Detection
by: Tabar, Maryam Shaygan, et al.
Published: (2025)
by: Tabar, Maryam Shaygan, et al.
Published: (2025)
BioTamperNet: Affinity-Guided State-Space Model Detecting Tampered Biomedical Images
by: Nandi, Soumyaroop, et al.
Published: (2026)
by: Nandi, Soumyaroop, et al.
Published: (2026)
Tamper-Proofing with Self-Modifying Code
by: Morse, Gregory, et al.
Published: (2026)
by: Morse, Gregory, et al.
Published: (2026)
Towards Universal Quantum Tamper Detection
by: Broadbent, Anne, et al.
Published: (2025)
by: Broadbent, Anne, et al.
Published: (2025)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026)
by: Struppek, Lukas, et al.
Published: (2026)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
by: Ortu, Francesco, et al.
Published: (2026)
by: Ortu, Francesco, et al.
Published: (2026)
LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing
by: Soni, Achint, et al.
Published: (2025)
by: Soni, Achint, et al.
Published: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
by: Kowal, Matthew, et al.
Published: (2026)
by: Kowal, Matthew, et al.
Published: (2026)
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
A Technique for the Detection of PDF Tampering or Forgery
by: Grobler, Gabriel, et al.
Published: (2025)
by: Grobler, Gabriel, et al.
Published: (2025)
Evidence Tampering and Chain of Custody in Layered Attestations
by: Kretz, Ian D., et al.
Published: (2024)
by: Kretz, Ian D., et al.
Published: (2024)
Emergent Persuasion: Will LLMs Persuade Without Being Prompted?
by: Chang, Vincent, et al.
Published: (2025)
by: Chang, Vincent, et al.
Published: (2025)
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
by: O'Brien, Kyle, et al.
Published: (2025)
by: O'Brien, Kyle, et al.
Published: (2025)
Anti-Tamper Protection for Unauthorized Individual Image Generation
by: Li, Zelin, et al.
Published: (2025)
by: Li, Zelin, et al.
Published: (2025)
Research about the Ability of LLM in the Tamper-Detection Area
by: Yang, Xinyu, et al.
Published: (2024)
by: Yang, Xinyu, et al.
Published: (2024)
Information Theoretic Analysis of PUF-Based Tamper Protection
by: Maringer, Georg, et al.
Published: (2025)
by: Maringer, Georg, et al.
Published: (2025)
TextSleuth: Towards Explainable Tampered Text Detection
by: Qu, Chenfan, et al.
Published: (2024)
by: Qu, Chenfan, et al.
Published: (2024)
Poster: Camera Tampering Detection for Outdoor IoT Systems
by: Attarha, Shadi, et al.
Published: (2026)
by: Attarha, Shadi, et al.
Published: (2026)
On Split-State Quantum Tamper Detection and Non-Malleability
by: Bergamaschi, Thiago, et al.
Published: (2023)
by: Bergamaschi, Thiago, et al.
Published: (2023)
Tamper-evident Image using JPEG Fixed Points
by: Si, Zhaofeng, et al.
Published: (2025)
by: Si, Zhaofeng, et al.
Published: (2025)
Management of Innovation in Academia: A Case Study in Tampere
by: Tuomo Heinonen
Published: (2015)
by: Tuomo Heinonen
Published: (2015)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
by: Ceraolo, Roberto, et al.
Published: (2024)
by: Ceraolo, Roberto, et al.
Published: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025)
by: Simko, Samuel, et al.
Published: (2025)
Recursive Binary Identification with Differential Privacy and Data Tampering Attacks
by: Wang, Jimin, et al.
Published: (2026)
by: Wang, Jimin, et al.
Published: (2026)
Detection of Physiological Data Tampering Attacks with Quantum Machine Learning
by: Onim, Md. Saif Hassan, et al.
Published: (2025)
by: Onim, Md. Saif Hassan, et al.
Published: (2025)
How to Tamper with a Parliament: Strategic Campaigns in Apportionment Elections
by: Bredereck, Robert, et al.
Published: (2026)
by: Bredereck, Robert, et al.
Published: (2026)
UVL2: A Unified Framework for Video Tampering Localization
by: Pei, Pengfei
Published: (2023)
by: Pei, Pengfei
Published: (2023)
Similar Items
-
Randomized Gradient Subspaces for Efficient Large Language Model Training
by: Rajabi, Sahar, et al.
Published: (2025) -
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
by: Hossain, Saad, et al.
Published: (2025) -
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025) -
SubTrack++ : Gradient Subspace Tracking for Scalable LLM Training
by: Rajabi, Sahar, et al.
Published: (2025) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026)