TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hossain, Saad, Tseng, Tom, Pandey, Punya Syon, Vajpayee, Samanvay, Kowal, Matthew, Nonta, Nayeema, Simko, Samuel, Casper, Stephen, Jin, Zhijing, Pelrine, Kellin, Rambhatla, Sirisha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Randomized Gradient Subspaces for Efficient Large Language Model Training
von: Rajabi, Sahar, et al.
Veröffentlicht: (2025)
von: Rajabi, Sahar, et al.
Veröffentlicht: (2025)
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
von: Hossain, Saad, et al.
Veröffentlicht: (2025)
von: Hossain, Saad, et al.
Veröffentlicht: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
SubTrack++ : Gradient Subspace Tracking for Scalable LLM Training
von: Rajabi, Sahar, et al.
Veröffentlicht: (2025)
von: Rajabi, Sahar, et al.
Veröffentlicht: (2025)
BinaryPPO: Efficient Policy Optimization for Binary Classification
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2026)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2026)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
von: Harrasse, Abir, et al.
Veröffentlicht: (2025)
von: Harrasse, Abir, et al.
Veröffentlicht: (2025)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
von: Vennemeyer, Daniel, et al.
Veröffentlicht: (2026)
von: Vennemeyer, Daniel, et al.
Veröffentlicht: (2026)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
von: Murphy, Brendan, et al.
Veröffentlicht: (2025)
von: Murphy, Brendan, et al.
Veröffentlicht: (2025)
International. Tampering with Jerusalem
Veröffentlicht: (1998)
Veröffentlicht: (1998)
Can Go AIs be adversarially robust?
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
Anti-Tamper Radio meets Reconfigurable Intelligent Surface for System-Level Tamper Detection
von: Tabar, Maryam Shaygan, et al.
Veröffentlicht: (2025)
von: Tabar, Maryam Shaygan, et al.
Veröffentlicht: (2025)
BioTamperNet: Affinity-Guided State-Space Model Detecting Tampered Biomedical Images
von: Nandi, Soumyaroop, et al.
Veröffentlicht: (2026)
von: Nandi, Soumyaroop, et al.
Veröffentlicht: (2026)
Tamper-Proofing with Self-Modifying Code
von: Morse, Gregory, et al.
Veröffentlicht: (2026)
von: Morse, Gregory, et al.
Veröffentlicht: (2026)
Towards Universal Quantum Tamper Detection
von: Broadbent, Anne, et al.
Veröffentlicht: (2025)
von: Broadbent, Anne, et al.
Veröffentlicht: (2025)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
von: Ortu, Francesco, et al.
Veröffentlicht: (2026)
von: Ortu, Francesco, et al.
Veröffentlicht: (2026)
LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing
von: Soni, Achint, et al.
Veröffentlicht: (2025)
von: Soni, Achint, et al.
Veröffentlicht: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
von: Che, Zora, et al.
Veröffentlicht: (2025)
von: Che, Zora, et al.
Veröffentlicht: (2025)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
Tamper-Resistant Safeguards for Open-Weight LLMs
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
A Technique for the Detection of PDF Tampering or Forgery
von: Grobler, Gabriel, et al.
Veröffentlicht: (2025)
von: Grobler, Gabriel, et al.
Veröffentlicht: (2025)
Evidence Tampering and Chain of Custody in Layered Attestations
von: Kretz, Ian D., et al.
Veröffentlicht: (2024)
von: Kretz, Ian D., et al.
Veröffentlicht: (2024)
Emergent Persuasion: Will LLMs Persuade Without Being Prompted?
von: Chang, Vincent, et al.
Veröffentlicht: (2025)
von: Chang, Vincent, et al.
Veröffentlicht: (2025)
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025)
Anti-Tamper Protection for Unauthorized Individual Image Generation
von: Li, Zelin, et al.
Veröffentlicht: (2025)
von: Li, Zelin, et al.
Veröffentlicht: (2025)
Research about the Ability of LLM in the Tamper-Detection Area
von: Yang, Xinyu, et al.
Veröffentlicht: (2024)
von: Yang, Xinyu, et al.
Veröffentlicht: (2024)
Information Theoretic Analysis of PUF-Based Tamper Protection
von: Maringer, Georg, et al.
Veröffentlicht: (2025)
von: Maringer, Georg, et al.
Veröffentlicht: (2025)
TextSleuth: Towards Explainable Tampered Text Detection
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
Poster: Camera Tampering Detection for Outdoor IoT Systems
von: Attarha, Shadi, et al.
Veröffentlicht: (2026)
von: Attarha, Shadi, et al.
Veröffentlicht: (2026)
On Split-State Quantum Tamper Detection and Non-Malleability
von: Bergamaschi, Thiago, et al.
Veröffentlicht: (2023)
von: Bergamaschi, Thiago, et al.
Veröffentlicht: (2023)
Tamper-evident Image using JPEG Fixed Points
von: Si, Zhaofeng, et al.
Veröffentlicht: (2025)
von: Si, Zhaofeng, et al.
Veröffentlicht: (2025)
Management of Innovation in Academia: A Case Study in Tampere
von: Tuomo Heinonen
Veröffentlicht: (2015)
von: Tuomo Heinonen
Veröffentlicht: (2015)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
von: Ceraolo, Roberto, et al.
Veröffentlicht: (2024)
von: Ceraolo, Roberto, et al.
Veröffentlicht: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
Recursive Binary Identification with Differential Privacy and Data Tampering Attacks
von: Wang, Jimin, et al.
Veröffentlicht: (2026)
von: Wang, Jimin, et al.
Veröffentlicht: (2026)
Detection of Physiological Data Tampering Attacks with Quantum Machine Learning
von: Onim, Md. Saif Hassan, et al.
Veröffentlicht: (2025)
von: Onim, Md. Saif Hassan, et al.
Veröffentlicht: (2025)
How to Tamper with a Parliament: Strategic Campaigns in Apportionment Elections
von: Bredereck, Robert, et al.
Veröffentlicht: (2026)
von: Bredereck, Robert, et al.
Veröffentlicht: (2026)
UVL2: A Unified Framework for Video Tampering Localization
von: Pei, Pengfei
Veröffentlicht: (2023)
von: Pei, Pengfei
Veröffentlicht: (2023)
Ähnliche Einträge
-
Randomized Gradient Subspaces for Efficient Large Language Model Training
von: Rajabi, Sahar, et al.
Veröffentlicht: (2025) -
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
von: Hossain, Saad, et al.
Veröffentlicht: (2025) -
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025) -
SubTrack++ : Gradient Subspace Tracking for Scalable LLM Training
von: Rajabi, Sahar, et al.
Veröffentlicht: (2025) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2026)