PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Seddik, Issam, Souihi, Sami, Tamaazousti, Mohamed, Piergiovanni, Sara Tucci |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
by: Cotti, Luca, et al.
Published: (2025)
by: Cotti, Luca, et al.
Published: (2025)
Unlearning at Scale: Implementing the Right to be Forgotten in Large Language Models
by: X, Abdullah
Published: (2025)
by: X, Abdullah
Published: (2025)
Lightweight LLMs for Network Attack Detection in IoT Networks
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026)
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026)
by: Leonesi, Matteo, et al.
Published: (2026)
Defending against Backdoor Attacks via Module Switching
by: Li, Weijun, et al.
Published: (2025)
by: Li, Weijun, et al.
Published: (2025)
POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models
by: Shao, Yangguang, et al.
Published: (2025)
by: Shao, Yangguang, et al.
Published: (2025)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
by: Lelle, Travis
Published: (2026)
by: Lelle, Travis
Published: (2026)
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction
by: Garcia, Gabriel
Published: (2026)
by: Garcia, Gabriel
Published: (2026)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
by: Wei, Xingyuan, et al.
Published: (2024)
by: Wei, Xingyuan, et al.
Published: (2024)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
DWFS-Obfuscation: Dynamic Weighted Feature Selection for Robust Malware Familial Classification under Obfuscation
by: Wei, Xingyuan, et al.
Published: (2025)
by: Wei, Xingyuan, et al.
Published: (2025)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
by: Gowda, Ishrith
Published: (2026)
by: Gowda, Ishrith
Published: (2026)
QuAnTS: Question Answering on Time Series
by: Divo, Felix, et al.
Published: (2025)
by: Divo, Felix, et al.
Published: (2025)
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
by: Kurian, Ashley, et al.
Published: (2025)
by: Kurian, Ashley, et al.
Published: (2025)
Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models
by: Del Rosario, Ron F.
Published: (2025)
by: Del Rosario, Ron F.
Published: (2025)
Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design
by: Charles, Richard M., et al.
Published: (2025)
by: Charles, Richard M., et al.
Published: (2025)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
Reducing Information Overload: Because Even Security Experts Need to Blink
by: Kuehn, Philipp, et al.
Published: (2022)
by: Kuehn, Philipp, et al.
Published: (2022)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
by: DeLeeuw, Caleb
Published: (2026)
by: DeLeeuw, Caleb
Published: (2026)
Attacking interpretable NLP systems
by: Abdukhamidov, Eldor, et al.
Published: (2025)
by: Abdukhamidov, Eldor, et al.
Published: (2025)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
by: Bercovich, Ivan, et al.
Published: (2026)
by: Bercovich, Ivan, et al.
Published: (2026)
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
by: Bommarito II, Michael J.
Published: (2025)
by: Bommarito II, Michael J.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
ConfusionPrompt: Practical Private Inference for Online Large Language Models
by: Mai, Peihua, et al.
Published: (2023)
by: Mai, Peihua, et al.
Published: (2023)
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
by: Verma, Apurv, et al.
Published: (2024)
by: Verma, Apurv, et al.
Published: (2024)
SecEmb: Sparsity-Aware Secure Federated Learning of On-Device Recommender System with Large Embedding
by: Mai, Peihua, et al.
Published: (2025)
by: Mai, Peihua, et al.
Published: (2025)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
by: Chen, Xinjie, et al.
Published: (2026)
by: Chen, Xinjie, et al.
Published: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
by: Schesch, Benedikt, et al.
Published: (2026)
by: Schesch, Benedikt, et al.
Published: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)
by: Kiashemshaki, Kiana, et al.
Published: (2025)
Mitigating the Impact of Malware Evolution on API Sequence-based Windows Malware Detector
by: Wei, Xingyuan, et al.
Published: (2024)
by: Wei, Xingyuan, et al.
Published: (2024)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025)
by: Palit, Sayon, et al.
Published: (2025)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
by: Zhang, Peiyan, et al.
Published: (2025)
by: Zhang, Peiyan, et al.
Published: (2025)
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
sudoLLM: On Multi-role Alignment of Language Models
by: Saha, Soumadeep, et al.
Published: (2025)
by: Saha, Soumadeep, et al.
Published: (2025)
Split-and-Denoise: Protect large language model inference with local differential privacy
by: Mai, Peihua, et al.
Published: (2023)
by: Mai, Peihua, et al.
Published: (2023)
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
Similar Items
-
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
by: Cotti, Luca, et al.
Published: (2025) -
Unlearning at Scale: Implementing the Right to be Forgotten in Large Language Models
by: X, Abdullah
Published: (2025) -
Lightweight LLMs for Network Attack Detection in IoT Networks
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026) -
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026) -
Defending against Backdoor Attacks via Module Switching
by: Li, Weijun, et al.
Published: (2025)