AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Sanyal, Debdeep, Ray, Manodeep, Mandal, Murari |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Realignment Problem: When Right becomes Wrong in LLMs
by: Sharma, Aakash Sen, et al.
Published: (2025)
by: Sharma, Aakash Sen, et al.
Published: (2025)
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Policy Optimization Prefers The Path of Least Resistance
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Confidence is Not Competence
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
by: Sharma, Aakash Sen, et al.
Published: (2025)
by: Sharma, Aakash Sen, et al.
Published: (2025)
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
by: Patil, Parth, et al.
Published: (2026)
by: Patil, Parth, et al.
Published: (2026)
NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic Data
by: Maiti, Agniva, et al.
Published: (2025)
by: Maiti, Agniva, et al.
Published: (2025)
time2time: Causal Intervention in Hidden States to Simulate Rare Events in Time Series Foundation Models
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Purple-teaming LLMs with Adversarial Defender Training
by: Zhou, Jingyan, et al.
Published: (2024)
by: Zhou, Jingyan, et al.
Published: (2024)
LLM-as-a-Judge for Time Series Explanations
by: Sivalingam, Preetham, et al.
Published: (2026)
by: Sivalingam, Preetham, et al.
Published: (2026)
Measuring Representation Robustness in Large Language Models for Geometry
by: Jawandhia, Vedant, et al.
Published: (2026)
by: Jawandhia, Vedant, et al.
Published: (2026)
OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
by: Talokar, Nivya, et al.
Published: (2026)
by: Talokar, Nivya, et al.
Published: (2026)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
by: Agarwal, Parth, et al.
Published: (2025)
by: Agarwal, Parth, et al.
Published: (2025)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
by: Jang, Eyon, et al.
Published: (2026)
by: Jang, Eyon, et al.
Published: (2026)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024)
by: Denison, Carson, et al.
Published: (2024)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
by: Garg, Madhav Krishan, et al.
Published: (2025)
by: Garg, Madhav Krishan, et al.
Published: (2025)
Evaluating LLMs and Pre-trained Models for Text Summarization Across Diverse Datasets
by: Rehman, Tohida, et al.
Published: (2025)
by: Rehman, Tohida, et al.
Published: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
Predicting Anti-microbial Resistance using Large Language Models
by: Yoo, Hyunwoo, et al.
Published: (2024)
by: Yoo, Hyunwoo, et al.
Published: (2024)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
by: Obadinma, Stephen, et al.
Published: (2025)
by: Obadinma, Stephen, et al.
Published: (2025)
Fast Adversarial Training against Textual Adversarial Attacks
by: Yang, Yichen, et al.
Published: (2024)
by: Yang, Yichen, et al.
Published: (2024)
"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents
by: Roy, Soham, et al.
Published: (2026)
by: Roy, Soham, et al.
Published: (2026)
Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
by: Dutta, Aditi, et al.
Published: (2025)
by: Dutta, Aditi, et al.
Published: (2025)
UniBERT: Adversarial Training for Language-Universal Representations
by: Avram, Andrei-Marius, et al.
Published: (2025)
by: Avram, Andrei-Marius, et al.
Published: (2025)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
by: Oepen, Stephan, et al.
Published: (2025)
by: Oepen, Stephan, et al.
Published: (2025)
Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
by: Cui, Sijia, et al.
Published: (2025)
by: Cui, Sijia, et al.
Published: (2025)
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
by: Chen, Sijia, et al.
Published: (2025)
by: Chen, Sijia, et al.
Published: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
AREG: Adversarial Resource Extraction Game for Evaluating Persuasion and Resistance in Large Language Models
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
AdvSumm: Adversarial Training for Bias Mitigation in Text Summarization
by: Gupta, Mukur, et al.
Published: (2025)
by: Gupta, Mukur, et al.
Published: (2025)
Utilizing Local Hierarchy with Adversarial Training for Hierarchical Text Classification
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
How Green are Neural Language Models? Analyzing Energy Consumption in Text Summarization Fine-tuning
by: Rehman, Tohida, et al.
Published: (2025)
by: Rehman, Tohida, et al.
Published: (2025)
TaxoAlign: Scholarly Taxonomy Generation Using Language Models
by: Lahiri, Avishek, et al.
Published: (2025)
by: Lahiri, Avishek, et al.
Published: (2025)
GINopic: Topic Modeling with Graph Isomorphism Network
by: Adhya, Suman, et al.
Published: (2024)
by: Adhya, Suman, et al.
Published: (2024)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
by: Gan, Esther, et al.
Published: (2024)
by: Gan, Esther, et al.
Published: (2024)
Similar Items
-
The Realignment Problem: When Right becomes Wrong in LLMs
by: Sharma, Aakash Sen, et al.
Published: (2025) -
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025) -
Policy Optimization Prefers The Path of Least Resistance
by: Sanyal, Debdeep, et al.
Published: (2025) -
Confidence is Not Competence
by: Sanyal, Debdeep, et al.
Published: (2025) -
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)