If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
Fuente:
arXiv
Saved in:
| Main Author: | Hernandez, Adriano |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
by: Yu, Zhiyuan, et al.
Published: (2024)
by: Yu, Zhiyuan, et al.
Published: (2024)
Ignore Me But Don't Replace Me: Utilizing Non-Linguistic Elements for Pretraining on the Cybersecurity Domain
by: Jang, Eugene, et al.
Published: (2024)
by: Jang, Eugene, et al.
Published: (2024)
Data Reconstruction: When You See It and When You Don't
by: Cohen, Edith, et al.
Published: (2024)
by: Cohen, Edith, et al.
Published: (2024)
Don't Walk the Line: Boundary Guidance for Filtered Generation
by: Ball, Sarah, et al.
Published: (2025)
by: Ball, Sarah, et al.
Published: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
by: Khan, Imran
Published: (2025)
by: Khan, Imran
Published: (2025)
Don't Forget Too Much: Towards Machine Unlearning on Feature Level
by: Xu, Heng, et al.
Published: (2024)
by: Xu, Heng, et al.
Published: (2024)
Reformulation is All You Need: Addressing Malicious Text Features in DNNs
by: Jiang, Yi, et al.
Published: (2025)
by: Jiang, Yi, et al.
Published: (2025)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
Larger-scale Nakamoto-style Blockchains Don't Necessarily Offer Better Security
by: Albrecht, Jannik, et al.
Published: (2024)
by: Albrecht, Jannik, et al.
Published: (2024)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
by: Yoon, Do-hyeon, et al.
Published: (2025)
by: Yoon, Do-hyeon, et al.
Published: (2025)
Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
by: Zhang, Yilin, et al.
Published: (2026)
by: Zhang, Yilin, et al.
Published: (2026)
Every Breath You Don't Take: Deepfake Speech Detection Using Breath
by: Layton, Seth, et al.
Published: (2024)
by: Layton, Seth, et al.
Published: (2024)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
by: García-Carrasco, Jorge, et al.
Published: (2024)
by: García-Carrasco, Jorge, et al.
Published: (2024)
Don't Let MEV Slip: The Costs of Swapping on the Uniswap Protocol
by: Adams, Austin, et al.
Published: (2023)
by: Adams, Austin, et al.
Published: (2023)
Are You Using Reliable Graph Prompts? Trojan Prompt Attacks on Graph Neural Networks
by: Lin, Minhua, et al.
Published: (2024)
by: Lin, Minhua, et al.
Published: (2024)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
by: Cao, Mingyu, et al.
Published: (2024)
by: Cao, Mingyu, et al.
Published: (2024)
Conti Inc.: Understanding the Internal Discussions of a large Ransomware-as-a-Service Operator with Machine Learning
by: Ruellan, Estelle, et al.
Published: (2023)
by: Ruellan, Estelle, et al.
Published: (2023)
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
by: Fan, Mingyuan, et al.
Published: (2026)
by: Fan, Mingyuan, et al.
Published: (2026)
Don't believe everything you read: Understanding and Measuring MCP Behavior under Misleading Tool Descriptions
by: Li, Zhihao, et al.
Published: (2026)
by: Li, Zhihao, et al.
Published: (2026)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
Don't Hash Me Like That: Exposing and Mitigating Hash-Induced Unfairness in Local Differential Privacy
by: Balioglu, Berkay Kemal, et al.
Published: (2025)
by: Balioglu, Berkay Kemal, et al.
Published: (2025)
I Don't Know You, But I Can Catch You: Real-Time Defense against Diverse Adversarial Patches for Object Detectors
by: Lin, Zijin, et al.
Published: (2024)
by: Lin, Zijin, et al.
Published: (2024)
Solving Trojan Detection Competitions with Linear Weight Classification
by: Huster, Todd, et al.
Published: (2024)
by: Huster, Todd, et al.
Published: (2024)
I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
by: Cohen, Roi, et al.
Published: (2024)
by: Cohen, Roi, et al.
Published: (2024)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
by: Sahabandu, Dinuka, et al.
Published: (2024)
by: Sahabandu, Dinuka, et al.
Published: (2024)
Using Hallucinations to Bypass GPT4's Filter
by: Lemkin, Benjamin
Published: (2024)
by: Lemkin, Benjamin
Published: (2024)
"I Don't Use AI for Everything": Exploring Utility, Attitude, and Responsibility of AI-empowered Tools in Software Development
by: Pan, Shidong, et al.
Published: (2024)
by: Pan, Shidong, et al.
Published: (2024)
SecEncoder: Logs are All You Need in Security
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
Don't Let the Claw Grip Your Hand: A Security Analysis and Defense Framework for OpenClaw
by: Shan, Zhengyang, et al.
Published: (2026)
by: Shan, Zhengyang, et al.
Published: (2026)
"Explain, Don't Just Warn!" -- A Real-Time Framework for Generating Phishing Warnings with Contextual Cues
by: Roy, Sayak Saha, et al.
Published: (2025)
by: Roy, Sayak Saha, et al.
Published: (2025)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Trojan Cleansing with Neural Collapse
by: Gu, Xihe, et al.
Published: (2024)
by: Gu, Xihe, et al.
Published: (2024)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
by: Johnson, Daniel D., et al.
Published: (2024)
by: Johnson, Daniel D., et al.
Published: (2024)
Please Don't Kill My Vibe: Empowering Agents with Data Flow Control
by: Summers, Charlie, et al.
Published: (2025)
by: Summers, Charlie, et al.
Published: (2025)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
by: Roy, Debjyoti Saha, et al.
Published: (2024)
by: Roy, Debjyoti Saha, et al.
Published: (2024)
Similar Items
-
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025) -
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
by: Yu, Zhiyuan, et al.
Published: (2024) -
Ignore Me But Don't Replace Me: Utilizing Non-Linguistic Elements for Pretraining on the Cybersecurity Domain
by: Jang, Eugene, et al.
Published: (2024) -
Data Reconstruction: When You See It and When You Don't
by: Cohen, Edith, et al.
Published: (2024) -
Don't Walk the Line: Boundary Guidance for Filtered Generation
by: Ball, Sarah, et al.
Published: (2025)