Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choudhary, Sarthak, Patlan, Atharv Singh, Palumbo, Nils, Hooda, Ashish, Fawaz, Kassem, Jha, Somesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025)
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025)
How Not to Detect Prompt Injections with an LLM
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025)
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025)
PolicyLR: A Logic Representation For Privacy Policies
von: Hooda, Ashish, et al.
Veröffentlicht: (2024)
von: Hooda, Ashish, et al.
Veröffentlicht: (2024)
Dependency-Aware Privacy for Multi-turn Agents
von: Anshumaan, Divyam, et al.
Veröffentlicht: (2026)
von: Anshumaan, Divyam, et al.
Veröffentlicht: (2026)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
Formal Policy Enforcement for Real-World Agentic Systems
von: Palumbo, Nils, et al.
Veröffentlicht: (2026)
von: Palumbo, Nils, et al.
Veröffentlicht: (2026)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
von: Wang, Zi, et al.
Veröffentlicht: (2024)
von: Wang, Zi, et al.
Veröffentlicht: (2024)
Agent Security is a Systems Problem
von: Christodorescu, Mihai, et al.
Veröffentlicht: (2026)
von: Christodorescu, Mihai, et al.
Veröffentlicht: (2026)
What Really is a Member? Discrediting Membership Inference via Poisoning
von: Mangaokar, Neal, et al.
Veröffentlicht: (2025)
von: Mangaokar, Neal, et al.
Veröffentlicht: (2025)
Context manipulation attacks : Web agents are susceptible to corrupted memory
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
Systems Security Foundations for Agentic Computing
von: Christodorescu, Mihai, et al.
Veröffentlicht: (2025)
von: Christodorescu, Mihai, et al.
Veröffentlicht: (2025)
Attacking Byzantine Robust Aggregation in High Dimensions
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2023)
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2023)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
Planting Undetectable Backdoors in Machine Learning Models
von: Goldwasser, Shafi, et al.
Veröffentlicht: (2022)
von: Goldwasser, Shafi, et al.
Veröffentlicht: (2022)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
von: Gong, Chen, et al.
Veröffentlicht: (2022)
von: Gong, Chen, et al.
Veröffentlicht: (2022)
Towards Backdoor Stealthiness in Model Parameter Space
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2025)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
von: Ramesh, Guruprasad Viswanathan, et al.
Veröffentlicht: (2026)
von: Ramesh, Guruprasad Viswanathan, et al.
Veröffentlicht: (2026)
Byzantine-Robust Federated Learning: An Overview With Focus on Developing Sybil-based Attacks to Backdoor Augmented Secure Aggregation Protocols
von: Deshmukh, Atharv
Veröffentlicht: (2024)
von: Deshmukh, Atharv
Veröffentlicht: (2024)
Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models
von: Kalavasis, Alkis, et al.
Veröffentlicht: (2024)
von: Kalavasis, Alkis, et al.
Veröffentlicht: (2024)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
von: Chen, Depeng, et al.
Veröffentlicht: (2024)
von: Chen, Depeng, et al.
Veröffentlicht: (2024)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
von: Li, Zi, et al.
Veröffentlicht: (2026)
von: Li, Zi, et al.
Veröffentlicht: (2026)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2024)
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2024)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
von: Gao, Yue, et al.
Veröffentlicht: (2023)
von: Gao, Yue, et al.
Veröffentlicht: (2023)
Prediction with Expert Advice under Local Differential Privacy
von: Jacobsen, Ben, et al.
Veröffentlicht: (2025)
von: Jacobsen, Ben, et al.
Veröffentlicht: (2025)
Private Continual Counting of Unbounded Streams
von: Jacobsen, Ben, et al.
Veröffentlicht: (2025)
von: Jacobsen, Ben, et al.
Veröffentlicht: (2025)
GPM: The Gaussian Pancake Mechanism for Planting Undetectable Backdoors in Differential Privacy
von: Sun, Haochen, et al.
Veröffentlicht: (2025)
von: Sun, Haochen, et al.
Veröffentlicht: (2025)
Backdooring Bias in Large Language Models
von: Das, Anudeep, et al.
Veröffentlicht: (2026)
von: Das, Anudeep, et al.
Veröffentlicht: (2026)
"Impressively Scary:" Exploring User Perceptions and Reactions to Unraveling Machine Learning Models in Social Media Applications
von: West, Jack, et al.
Veröffentlicht: (2025)
von: West, Jack, et al.
Veröffentlicht: (2025)
An Undetectable Watermark for Generative Image Models
von: Gunn, Sam, et al.
Veröffentlicht: (2024)
von: Gunn, Sam, et al.
Veröffentlicht: (2024)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
von: Hooda, Ashish, et al.
Veröffentlicht: (2024)
von: Hooda, Ashish, et al.
Veröffentlicht: (2024)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
Auto-SPT: Automating Semantic Preserving Transformations for Code
von: Hooda, Ashish, et al.
Veröffentlicht: (2025)
von: Hooda, Ashish, et al.
Veröffentlicht: (2025)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2025)
von: Yu, Miao, et al.
Veröffentlicht: (2025)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
von: You, Ziyang, et al.
Veröffentlicht: (2026)
von: You, Ziyang, et al.
Veröffentlicht: (2026)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
von: Hartman, Max, et al.
Veröffentlicht: (2026)
von: Hartman, Max, et al.
Veröffentlicht: (2026)
Hide and Seek: Fingerprinting Large Language Models with Evolutionary Learning
von: Iourovitski, Dmitri, et al.
Veröffentlicht: (2024)
von: Iourovitski, Dmitri, et al.
Veröffentlicht: (2024)
Efficient Malware Detection with Optimized Learning on High-Dimensional Features
von: Choudhary, Aditya, et al.
Veröffentlicht: (2025)
von: Choudhary, Aditya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025) -
How Not to Detect Prompt Injections with an LLM
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025) -
PolicyLR: A Logic Representation For Privacy Policies
von: Hooda, Ashish, et al.
Veröffentlicht: (2024) -
Dependency-Aware Privacy for Multi-turn Agents
von: Anshumaan, Divyam, et al.
Veröffentlicht: (2026) -
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)