Trojans in Artificial Intelligence (TrojAI) Final Report
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914358260400128 |
|---|---|
| author | Reese, Kristopher W. Kulp-McDowall, Taylor Majurski, Michael Blattner, Tim Juba, Derek Bajcsy, Peter Cardone, Antonio Dessauw, Philippe Dima, Alden Kearsley, Anthony J. Kleczynski, Melinda Vasanth, Joel Keyrouz, Walid Ashcraft, Chace Fendley, Neil Staley, Ted Stout, Trevor Carney, Josh Canal, Greg Redman, Will Schmidt, Aurora Hickert, Cameron Paul, William Markowitz, Jared Drenkow, Nathan Shriver, David Connor, Marissa Grimes, Keltin Christiani, Marco Moore, Hayden Widjaja, Jordan Gabert, Kasimir Balakrishnan, Uma Gundimada, Satyanadh Jacobellis, John Lakkur, Sandya Leung, Vitus Roose, Jon Battaglino, Casey Koushanfar, Farinaz Fields, Greg Gu, Xihe Jandali, Yaman Zhang, Xinqiao Javidi, Tara Vartak, Akash Oates, Tim Erichson, Ben Mahoney, Michael Izmailov, Rauf Zhang, Xiangyu Shen, Guangyu Cheng, Siyuan Ma, Shiqing Wang, XiaoFeng Tang, Haixu Tang, Di Chen, Xiaoyi Wang, Zihao Zhu, Rui Jha, Susmit Lin, Xiao Acharya, Manoj Zhou, Weichao Fu, Feisi Kiourti, Panagiota Wang, Chenyu Guo, Zijian Ahmad, H M Sabbir Li, Wenchao Chen, Chao |
| author_facet | Reese, Kristopher W. Kulp-McDowall, Taylor Majurski, Michael Blattner, Tim Juba, Derek Bajcsy, Peter Cardone, Antonio Dessauw, Philippe Dima, Alden Kearsley, Anthony J. Kleczynski, Melinda Vasanth, Joel Keyrouz, Walid Ashcraft, Chace Fendley, Neil Staley, Ted Stout, Trevor Carney, Josh Canal, Greg Redman, Will Schmidt, Aurora Hickert, Cameron Paul, William Markowitz, Jared Drenkow, Nathan Shriver, David Connor, Marissa Grimes, Keltin Christiani, Marco Moore, Hayden Widjaja, Jordan Gabert, Kasimir Balakrishnan, Uma Gundimada, Satyanadh Jacobellis, John Lakkur, Sandya Leung, Vitus Roose, Jon Battaglino, Casey Koushanfar, Farinaz Fields, Greg Gu, Xihe Jandali, Yaman Zhang, Xinqiao Javidi, Tara Vartak, Akash Oates, Tim Erichson, Ben Mahoney, Michael Izmailov, Rauf Zhang, Xiangyu Shen, Guangyu Cheng, Siyuan Ma, Shiqing Wang, XiaoFeng Tang, Haixu Tang, Di Chen, Xiaoyi Wang, Zihao Zhu, Rui Jha, Susmit Lin, Xiao Acharya, Manoj Zhou, Weichao Fu, Feisi Kiourti, Panagiota Wang, Chenyu Guo, Zijian Ahmad, H M Sabbir Li, Wenchao Chen, Chao |
| contents | The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI Trojans. These AI trojans are malicious, hidden backdoors intentionally embedded within an AI model that can cause a system to fail in unexpected ways, or allow a malicious actor to hijack the AI model at will. This multi-year initiative helped to map out the complex nature of the threat, pioneered foundational detection methods, and identified unsolved challenges that require ongoing attention by the burgeoning AI security field. This report synthesizes the program's key findings, including methodologies for detection through weight analysis and trigger inversion, as well as approaches for mitigating Trojan risks in deployed models. Comprehensive test and evaluation results highlight detector performance, sensitivity, and the prevalence of "natural" Trojans. The report concludes with lessons learned and recommendations for advancing AI security research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_07152 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Trojans in Artificial Intelligence (TrojAI) Final Report Reese, Kristopher W. Kulp-McDowall, Taylor Majurski, Michael Blattner, Tim Juba, Derek Bajcsy, Peter Cardone, Antonio Dessauw, Philippe Dima, Alden Kearsley, Anthony J. Kleczynski, Melinda Vasanth, Joel Keyrouz, Walid Ashcraft, Chace Fendley, Neil Staley, Ted Stout, Trevor Carney, Josh Canal, Greg Redman, Will Schmidt, Aurora Hickert, Cameron Paul, William Markowitz, Jared Drenkow, Nathan Shriver, David Connor, Marissa Grimes, Keltin Christiani, Marco Moore, Hayden Widjaja, Jordan Gabert, Kasimir Balakrishnan, Uma Gundimada, Satyanadh Jacobellis, John Lakkur, Sandya Leung, Vitus Roose, Jon Battaglino, Casey Koushanfar, Farinaz Fields, Greg Gu, Xihe Jandali, Yaman Zhang, Xinqiao Javidi, Tara Vartak, Akash Oates, Tim Erichson, Ben Mahoney, Michael Izmailov, Rauf Zhang, Xiangyu Shen, Guangyu Cheng, Siyuan Ma, Shiqing Wang, XiaoFeng Tang, Haixu Tang, Di Chen, Xiaoyi Wang, Zihao Zhu, Rui Jha, Susmit Lin, Xiao Acharya, Manoj Zhou, Weichao Fu, Feisi Kiourti, Panagiota Wang, Chenyu Guo, Zijian Ahmad, H M Sabbir Li, Wenchao Chen, Chao Cryptography and Security Artificial Intelligence Machine Learning The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI Trojans. These AI trojans are malicious, hidden backdoors intentionally embedded within an AI model that can cause a system to fail in unexpected ways, or allow a malicious actor to hijack the AI model at will. This multi-year initiative helped to map out the complex nature of the threat, pioneered foundational detection methods, and identified unsolved challenges that require ongoing attention by the burgeoning AI security field. This report synthesizes the program's key findings, including methodologies for detection through weight analysis and trigger inversion, as well as approaches for mitigating Trojan risks in deployed models. Comprehensive test and evaluation results highlight detector performance, sensitivity, and the prevalence of "natural" Trojans. The report concludes with lessons learned and recommendations for advancing AI security research. |
| title | Trojans in Artificial Intelligence (TrojAI) Final Report |
| topic | Cryptography and Security Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2602.07152 |