Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Maloyan, Narek, Verma, Ekansh, Nutfullin, Bulat, Ashinov, Bislan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
par: Maloyan, Narek, et autres
Publié: (2025)
par: Maloyan, Narek, et autres
Publié: (2025)
Prompt Injection Attacks in Defended Systems
par: Khomsky, Daniil, et autres
Publié: (2024)
par: Khomsky, Daniil, et autres
Publié: (2024)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
par: Maloyan, Narek, et autres
Publié: (2025)
par: Maloyan, Narek, et autres
Publié: (2025)
Trojan Detection Through Pattern Recognition for Large Language Models
par: Bhasin, Vedant, et autres
Publié: (2025)
par: Bhasin, Vedant, et autres
Publié: (2025)
Solving Trojan Detection Competitions with Linear Weight Classification
par: Huster, Todd, et autres
Publié: (2024)
par: Huster, Todd, et autres
Publié: (2024)
PETA: Parameter-Efficient Trojan Attacks
par: Hong, Lauren, et autres
Publié: (2023)
par: Hong, Lauren, et autres
Publié: (2023)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
par: Cheng, Pengzhou, et autres
Publié: (2024)
par: Cheng, Pengzhou, et autres
Publié: (2024)
Analyzing Multi-Head Attention on Trojan BERT Models
par: Wang, Jingwei
Publié: (2024)
par: Wang, Jingwei
Publié: (2024)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
par: Wang, Haoran, et autres
Publié: (2023)
par: Wang, Haoran, et autres
Publié: (2023)
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
par: Meier, Dominik, et autres
Publié: (2025)
par: Meier, Dominik, et autres
Publié: (2025)
Uncertainty-Aware Evaluation for Vision-Language Models
par: Kostumov, Vasily, et autres
Publié: (2024)
par: Kostumov, Vasily, et autres
Publié: (2024)
Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents
par: Maloyan, Narek, et autres
Publié: (2026)
par: Maloyan, Narek, et autres
Publié: (2026)
Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents
par: Maloyan, Narek, et autres
Publié: (2026)
par: Maloyan, Narek, et autres
Publié: (2026)
Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems
par: Maloyan, Narek, et autres
Publié: (2026)
par: Maloyan, Narek, et autres
Publié: (2026)
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
par: Verma, Vivek, et autres
Publié: (2023)
par: Verma, Vivek, et autres
Publié: (2023)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
par: Hernandez, Adriano
Publié: (2024)
par: Hernandez, Adriano
Publié: (2024)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
par: Sel, Bilgehan, et autres
Publié: (2026)
par: Sel, Bilgehan, et autres
Publié: (2026)
Large Language Models for Mathematical Reasoning: Progresses and Challenges
par: Ahn, Janice, et autres
Publié: (2024)
par: Ahn, Janice, et autres
Publié: (2024)
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
par: Faruque, Md Omar, et autres
Publié: (2024)
par: Faruque, Md Omar, et autres
Publié: (2024)
Trojan Playground: A Reinforcement Learning Framework for Hardware Trojan Insertion and Detection
par: Sarihi, Amin, et autres
Publié: (2023)
par: Sarihi, Amin, et autres
Publié: (2023)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
par: Liu, Yupei, et autres
Publié: (2025)
par: Liu, Yupei, et autres
Publié: (2025)
SSL-Cleanse: Trojan Detection and Mitigation in Self-Supervised Learning
par: Zheng, Mengxin, et autres
Publié: (2023)
par: Zheng, Mengxin, et autres
Publié: (2023)
On Trojan Signatures in Large Language Models of Code
par: Hussain, Aftab, et autres
Publié: (2024)
par: Hussain, Aftab, et autres
Publié: (2024)
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
par: Tao, Xijia, et autres
Publié: (2024)
par: Tao, Xijia, et autres
Publié: (2024)
Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
par: Arcan, Mihael, et autres
Publié: (2025)
par: Arcan, Mihael, et autres
Publié: (2025)
Detection Avoidance Techniques for Large Language Models
par: Schneider, Sinclair, et autres
Publié: (2025)
par: Schneider, Sinclair, et autres
Publié: (2025)
Detecting Stylistic Fingerprints of Large Language Models
par: Bitton, Yehonatan, et autres
Publié: (2025)
par: Bitton, Yehonatan, et autres
Publié: (2025)
Large Language Models in the Abuse Detection Pipeline
par: Kath, Suraj, et autres
Publié: (2026)
par: Kath, Suraj, et autres
Publié: (2026)
The Philosopher's Stone: Trojaning Plugins of Large Language Models
par: Dong, Tian, et autres
Publié: (2023)
par: Dong, Tian, et autres
Publié: (2023)
CatchBackdoor: Backdoor Detection via Critical Trojan Neural Path Fuzzing
par: Jin, Haibo, et autres
Publié: (2021)
par: Jin, Haibo, et autres
Publié: (2021)
The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance
par: Maynard, Andrew D.
Publié: (2026)
par: Maynard, Andrew D.
Publié: (2026)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
par: Samuel, Vinay, et autres
Publié: (2024)
par: Samuel, Vinay, et autres
Publié: (2024)
DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?
par: Gu, Zhouhong, et autres
Publié: (2024)
par: Gu, Zhouhong, et autres
Publié: (2024)
Large Language Models for Sentiment Analysis to Detect Social Challenges: A Use Case with South African Languages
par: Mabokela, Koena Ronny, et autres
Publié: (2025)
par: Mabokela, Koena Ronny, et autres
Publié: (2025)
Contextual Compression in Retrieval-Augmented Generation for Large Language Models: A Survey
par: Verma, Sourav
Publié: (2024)
par: Verma, Sourav
Publié: (2024)
Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course
par: Chiang, Cheng-Han, et autres
Publié: (2024)
par: Chiang, Cheng-Han, et autres
Publié: (2024)
Large Language Models for Toxic Language Detection in Low-Resource Balkan Languages
par: Muminovic, Amel, et autres
Publié: (2025)
par: Muminovic, Amel, et autres
Publié: (2025)
Evaluating Large Language Models for Detecting Antisemitism
par: Patel, Jay, et autres
Publié: (2025)
par: Patel, Jay, et autres
Publié: (2025)
Detecting Reference Errors in Scientific Literature with Large Language Models
par: Zhang, Tianmai M., et autres
Publié: (2024)
par: Zhang, Tianmai M., et autres
Publié: (2024)
Self-training Large Language Models through Knowledge Detection
par: Yeo, Wei Jie, et autres
Publié: (2024)
par: Yeo, Wei Jie, et autres
Publié: (2024)
Documents similaires
-
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
par: Maloyan, Narek, et autres
Publié: (2025) -
Prompt Injection Attacks in Defended Systems
par: Khomsky, Daniil, et autres
Publié: (2024) -
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
par: Maloyan, Narek, et autres
Publié: (2025) -
Trojan Detection Through Pattern Recognition for Large Language Models
par: Bhasin, Vedant, et autres
Publié: (2025) -
Solving Trojan Detection Competitions with Linear Weight Classification
par: Huster, Todd, et autres
Publié: (2024)