Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | Maloyan, Narek, Verma, Ekansh, Nutfullin, Bulat, Ashinov, Bislan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
by: Maloyan, Narek, et al.
Published: (2025)
by: Maloyan, Narek, et al.
Published: (2025)
Prompt Injection Attacks in Defended Systems
by: Khomsky, Daniil, et al.
Published: (2024)
by: Khomsky, Daniil, et al.
Published: (2024)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
by: Maloyan, Narek, et al.
Published: (2025)
by: Maloyan, Narek, et al.
Published: (2025)
Trojan Detection Through Pattern Recognition for Large Language Models
by: Bhasin, Vedant, et al.
Published: (2025)
by: Bhasin, Vedant, et al.
Published: (2025)
Solving Trojan Detection Competitions with Linear Weight Classification
by: Huster, Todd, et al.
Published: (2024)
by: Huster, Todd, et al.
Published: (2024)
PETA: Parameter-Efficient Trojan Attacks
by: Hong, Lauren, et al.
Published: (2023)
by: Hong, Lauren, et al.
Published: (2023)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
Analyzing Multi-Head Attention on Trojan BERT Models
by: Wang, Jingwei
Published: (2024)
by: Wang, Jingwei
Published: (2024)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
by: Meier, Dominik, et al.
Published: (2025)
by: Meier, Dominik, et al.
Published: (2025)
Uncertainty-Aware Evaluation for Vision-Language Models
by: Kostumov, Vasily, et al.
Published: (2024)
by: Kostumov, Vasily, et al.
Published: (2024)
Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents
by: Maloyan, Narek, et al.
Published: (2026)
by: Maloyan, Narek, et al.
Published: (2026)
Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents
by: Maloyan, Narek, et al.
Published: (2026)
by: Maloyan, Narek, et al.
Published: (2026)
Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems
by: Maloyan, Narek, et al.
Published: (2026)
by: Maloyan, Narek, et al.
Published: (2026)
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
by: Verma, Vivek, et al.
Published: (2023)
by: Verma, Vivek, et al.
Published: (2023)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
by: Sel, Bilgehan, et al.
Published: (2026)
by: Sel, Bilgehan, et al.
Published: (2026)
Large Language Models for Mathematical Reasoning: Progresses and Challenges
by: Ahn, Janice, et al.
Published: (2024)
by: Ahn, Janice, et al.
Published: (2024)
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
by: Faruque, Md Omar, et al.
Published: (2024)
by: Faruque, Md Omar, et al.
Published: (2024)
Trojan Playground: A Reinforcement Learning Framework for Hardware Trojan Insertion and Detection
by: Sarihi, Amin, et al.
Published: (2023)
by: Sarihi, Amin, et al.
Published: (2023)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
by: Liu, Yupei, et al.
Published: (2025)
by: Liu, Yupei, et al.
Published: (2025)
SSL-Cleanse: Trojan Detection and Mitigation in Self-Supervised Learning
by: Zheng, Mengxin, et al.
Published: (2023)
by: Zheng, Mengxin, et al.
Published: (2023)
On Trojan Signatures in Large Language Models of Code
by: Hussain, Aftab, et al.
Published: (2024)
by: Hussain, Aftab, et al.
Published: (2024)
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
by: Tao, Xijia, et al.
Published: (2024)
by: Tao, Xijia, et al.
Published: (2024)
Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
by: Arcan, Mihael, et al.
Published: (2025)
by: Arcan, Mihael, et al.
Published: (2025)
Detection Avoidance Techniques for Large Language Models
by: Schneider, Sinclair, et al.
Published: (2025)
by: Schneider, Sinclair, et al.
Published: (2025)
Detecting Stylistic Fingerprints of Large Language Models
by: Bitton, Yehonatan, et al.
Published: (2025)
by: Bitton, Yehonatan, et al.
Published: (2025)
Large Language Models in the Abuse Detection Pipeline
by: Kath, Suraj, et al.
Published: (2026)
by: Kath, Suraj, et al.
Published: (2026)
The Philosopher's Stone: Trojaning Plugins of Large Language Models
by: Dong, Tian, et al.
Published: (2023)
by: Dong, Tian, et al.
Published: (2023)
CatchBackdoor: Backdoor Detection via Critical Trojan Neural Path Fuzzing
by: Jin, Haibo, et al.
Published: (2021)
by: Jin, Haibo, et al.
Published: (2021)
The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance
by: Maynard, Andrew D.
Published: (2026)
by: Maynard, Andrew D.
Published: (2026)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?
by: Gu, Zhouhong, et al.
Published: (2024)
by: Gu, Zhouhong, et al.
Published: (2024)
Large Language Models for Sentiment Analysis to Detect Social Challenges: A Use Case with South African Languages
by: Mabokela, Koena Ronny, et al.
Published: (2025)
by: Mabokela, Koena Ronny, et al.
Published: (2025)
Contextual Compression in Retrieval-Augmented Generation for Large Language Models: A Survey
by: Verma, Sourav
Published: (2024)
by: Verma, Sourav
Published: (2024)
Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course
by: Chiang, Cheng-Han, et al.
Published: (2024)
by: Chiang, Cheng-Han, et al.
Published: (2024)
Large Language Models for Toxic Language Detection in Low-Resource Balkan Languages
by: Muminovic, Amel, et al.
Published: (2025)
by: Muminovic, Amel, et al.
Published: (2025)
Evaluating Large Language Models for Detecting Antisemitism
by: Patel, Jay, et al.
Published: (2025)
by: Patel, Jay, et al.
Published: (2025)
Detecting Reference Errors in Scientific Literature with Large Language Models
by: Zhang, Tianmai M., et al.
Published: (2024)
by: Zhang, Tianmai M., et al.
Published: (2024)
Self-training Large Language Models through Knowledge Detection
by: Yeo, Wei Jie, et al.
Published: (2024)
by: Yeo, Wei Jie, et al.
Published: (2024)
Similar Items
-
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
by: Maloyan, Narek, et al.
Published: (2025) -
Prompt Injection Attacks in Defended Systems
by: Khomsky, Daniil, et al.
Published: (2024) -
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
by: Maloyan, Narek, et al.
Published: (2025) -
Trojan Detection Through Pattern Recognition for Large Language Models
by: Bhasin, Vedant, et al.
Published: (2025) -
Solving Trojan Detection Competitions with Linear Weight Classification
by: Huster, Todd, et al.
Published: (2024)