TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Yupei, Wang, Yanting, Jia, Jinyuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
por: Faruque, Md Omar, et al.
Publicado: (2024)
por: Faruque, Md Omar, et al.
Publicado: (2024)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration
por: Das, Debeshee, et al.
Publicado: (2026)
por: Das, Debeshee, et al.
Publicado: (2026)
TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion
por: Sreekumar, Saideep, et al.
Publicado: (2026)
por: Sreekumar, Saideep, et al.
Publicado: (2026)
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
por: Wang, Hongtao, et al.
Publicado: (2026)
por: Wang, Hongtao, et al.
Publicado: (2026)
MergeGuard: Efficient Thwarting of Trojan Attacks in Machine Learning Models
por: Shabgahi, Soheil Zibakhsh, et al.
Publicado: (2025)
por: Shabgahi, Soheil Zibakhsh, et al.
Publicado: (2025)
Detecting Trojaned DNNs via Spectral Regression Analysis
por: Pasini, Samuele, et al.
Publicado: (2026)
por: Pasini, Samuele, et al.
Publicado: (2026)
Uncertainty-Aware Hardware Trojan Detection Using Multimodal Deep Learning
por: Vishwakarma, Rahul, et al.
Publicado: (2024)
por: Vishwakarma, Rahul, et al.
Publicado: (2024)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
por: Feng, Yunhao, et al.
Publicado: (2026)
por: Feng, Yunhao, et al.
Publicado: (2026)
PromptLocate: Localizing Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data
por: Croce, Nicola, et al.
Publicado: (2025)
por: Croce, Nicola, et al.
Publicado: (2025)
Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
por: Liu, Fazhong, et al.
Publicado: (2026)
por: Liu, Fazhong, et al.
Publicado: (2026)
SAND: A Self-supervised and Adaptive NAS-Driven Framework for Hardware Trojan Detection
por: Pan, Zhixin, et al.
Publicado: (2025)
por: Pan, Zhixin, et al.
Publicado: (2025)
Trojans in Artificial Intelligence (TrojAI) Final Report
por: Reese, Kristopher W., et al.
Publicado: (2026)
por: Reese, Kristopher W., et al.
Publicado: (2026)
An AI Architecture with the Capability to Classify and Explain Hardware Trojans
por: Whitten, Paul, et al.
Publicado: (2024)
por: Whitten, Paul, et al.
Publicado: (2024)
Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging
por: Latibari, Banafsheh Saber, et al.
Publicado: (2025)
por: Latibari, Banafsheh Saber, et al.
Publicado: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
CatchBackdoor: Backdoor Detection via Critical Trojan Neural Path Fuzzing
por: Jin, Haibo, et al.
Publicado: (2021)
por: Jin, Haibo, et al.
Publicado: (2021)
Quantum Properties Trojans (QuPTs) for Attacking Quantum Neural Networks
por: Bhowmik, Sounak, et al.
Publicado: (2025)
por: Bhowmik, Sounak, et al.
Publicado: (2025)
SPICED: Syntactical Bug and Trojan Pattern Identification in A/MS Circuits using LLM-Enhanced Detection
por: Chaudhuri, Jayeeta, et al.
Publicado: (2024)
por: Chaudhuri, Jayeeta, et al.
Publicado: (2024)
TracLLM: A Generic Framework for Attributing Long Context LLMs
por: Wang, Yanting, et al.
Publicado: (2025)
por: Wang, Yanting, et al.
Publicado: (2025)
Neural Trojans
por: Liu, Yuntao, et al.
Publicado: (2017)
por: Liu, Yuntao, et al.
Publicado: (2017)
Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models
por: Wirth, Manuel
Publicado: (2026)
por: Wirth, Manuel
Publicado: (2026)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
por: Sel, Bilgehan, et al.
Publicado: (2026)
por: Sel, Bilgehan, et al.
Publicado: (2026)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
por: Wang, Haoran, et al.
Publicado: (2023)
por: Wang, Haoran, et al.
Publicado: (2023)
SENTAUR: Security EnhaNced Trojan Assessment Using LLMs Against Undesirable Revisions
por: Bhandari, Jitendra, et al.
Publicado: (2024)
por: Bhandari, Jitendra, et al.
Publicado: (2024)
TrojanLoC: LLM-based Framework for RTL Trojan Localization
por: Xiao, Weihua, et al.
Publicado: (2025)
por: Xiao, Weihua, et al.
Publicado: (2025)
An Experimental Study of Trojan Vulnerabilities in UAV Autonomous Landing
por: Ahmari, Reza, et al.
Publicado: (2025)
por: Ahmari, Reza, et al.
Publicado: (2025)
TrojanTime: Backdoor Attacks on Time Series Classification
por: Dong, Chang, et al.
Publicado: (2025)
por: Dong, Chang, et al.
Publicado: (2025)
Solving Trojan Detection Competitions with Linear Weight Classification
por: Huster, Todd, et al.
Publicado: (2024)
por: Huster, Todd, et al.
Publicado: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
por: Liu, Yupei, et al.
Publicado: (2023)
por: Liu, Yupei, et al.
Publicado: (2023)
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
por: Chen, Junxi, et al.
Publicado: (2025)
por: Chen, Junxi, et al.
Publicado: (2025)
Protecting Model Adaptation from Trojans in the Unlabeled Data
por: Sheng, Lijun, et al.
Publicado: (2024)
por: Sheng, Lijun, et al.
Publicado: (2024)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
por: Zou, Wei, et al.
Publicado: (2025)
por: Zou, Wei, et al.
Publicado: (2025)
Trojan Playground: A Reinforcement Learning Framework for Hardware Trojan Insertion and Detection
por: Sarihi, Amin, et al.
Publicado: (2023)
por: Sarihi, Amin, et al.
Publicado: (2023)
HeisenTrojans: They Are Not There Until They Are Triggered
por: Mavurapu, Akshita Reddy, et al.
Publicado: (2023)
por: Mavurapu, Akshita Reddy, et al.
Publicado: (2023)
EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace Method
por: Wang, Yanting, et al.
Publicado: (2026)
por: Wang, Yanting, et al.
Publicado: (2026)
PIArena: A Platform for Prompt Injection Evaluation
por: Geng, Runpeng, et al.
Publicado: (2026)
por: Geng, Runpeng, et al.
Publicado: (2026)
Selective Amnesia: On Efficient, High-Fidelity and Blind Suppression of Backdoor Effects in Trojaned Machine Learning Models
por: Zhu, Rui, et al.
Publicado: (2022)
por: Zhu, Rui, et al.
Publicado: (2022)
Ejemplares similares
-
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
por: Faruque, Md Omar, et al.
Publicado: (2024) -
SecInfer: Preventing Prompt Injection via Inference-time Scaling
por: Liu, Yupei, et al.
Publicado: (2025) -
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
por: Liu, Yupei, et al.
Publicado: (2025) -
Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration
por: Das, Debeshee, et al.
Publicado: (2026) -
TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion
por: Sreekumar, Saideep, et al.
Publicado: (2026)