Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
Fuente:
arXiv
Salvato in:
| Autori principali: | Choudhary, Sarthak, Patlan, Atharv Singh, Palumbo, Nils, Hooda, Ashish, Fawaz, Kassem, Jha, Somesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
How Not to Detect Prompt Injections with an LLM
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
PolicyLR: A Logic Representation For Privacy Policies
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
Dependency-Aware Privacy for Multi-turn Agents
di: Anshumaan, Divyam, et al.
Pubblicazione: (2026)
di: Anshumaan, Divyam, et al.
Pubblicazione: (2026)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
di: Mangaokar, Neal, et al.
Pubblicazione: (2024)
di: Mangaokar, Neal, et al.
Pubblicazione: (2024)
Formal Policy Enforcement for Real-World Agentic Systems
di: Palumbo, Nils, et al.
Pubblicazione: (2026)
di: Palumbo, Nils, et al.
Pubblicazione: (2026)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
di: Wang, Zi, et al.
Pubblicazione: (2024)
di: Wang, Zi, et al.
Pubblicazione: (2024)
Agent Security is a Systems Problem
di: Christodorescu, Mihai, et al.
Pubblicazione: (2026)
di: Christodorescu, Mihai, et al.
Pubblicazione: (2026)
What Really is a Member? Discrediting Membership Inference via Poisoning
di: Mangaokar, Neal, et al.
Pubblicazione: (2025)
di: Mangaokar, Neal, et al.
Pubblicazione: (2025)
Context manipulation attacks : Web agents are susceptible to corrupted memory
di: Patlan, Atharv Singh, et al.
Pubblicazione: (2025)
di: Patlan, Atharv Singh, et al.
Pubblicazione: (2025)
Systems Security Foundations for Agentic Computing
di: Christodorescu, Mihai, et al.
Pubblicazione: (2025)
di: Christodorescu, Mihai, et al.
Pubblicazione: (2025)
Attacking Byzantine Robust Aggregation in High Dimensions
di: Choudhary, Sarthak, et al.
Pubblicazione: (2023)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2023)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
di: Patlan, Atharv Singh, et al.
Pubblicazione: (2025)
di: Patlan, Atharv Singh, et al.
Pubblicazione: (2025)
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
di: Patlan, Atharv Singh, et al.
Pubblicazione: (2025)
di: Patlan, Atharv Singh, et al.
Pubblicazione: (2025)
Planting Undetectable Backdoors in Machine Learning Models
di: Goldwasser, Shafi, et al.
Pubblicazione: (2022)
di: Goldwasser, Shafi, et al.
Pubblicazione: (2022)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
di: Gong, Chen, et al.
Pubblicazione: (2022)
di: Gong, Chen, et al.
Pubblicazione: (2022)
Towards Backdoor Stealthiness in Model Parameter Space
di: Xu, Xiaoyun, et al.
Pubblicazione: (2025)
di: Xu, Xiaoyun, et al.
Pubblicazione: (2025)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
di: Ramesh, Guruprasad Viswanathan, et al.
Pubblicazione: (2026)
di: Ramesh, Guruprasad Viswanathan, et al.
Pubblicazione: (2026)
Byzantine-Robust Federated Learning: An Overview With Focus on Developing Sybil-based Attacks to Backdoor Augmented Secure Aggregation Protocols
di: Deshmukh, Atharv
Pubblicazione: (2024)
di: Deshmukh, Atharv
Pubblicazione: (2024)
Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models
di: Kalavasis, Alkis, et al.
Pubblicazione: (2024)
di: Kalavasis, Alkis, et al.
Pubblicazione: (2024)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
di: Chen, Depeng, et al.
Pubblicazione: (2024)
di: Chen, Depeng, et al.
Pubblicazione: (2024)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
di: Li, Zi, et al.
Pubblicazione: (2026)
di: Li, Zi, et al.
Pubblicazione: (2026)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
di: Pal, Soumyadeep, et al.
Pubblicazione: (2024)
di: Pal, Soumyadeep, et al.
Pubblicazione: (2024)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
di: Gao, Yue, et al.
Pubblicazione: (2023)
di: Gao, Yue, et al.
Pubblicazione: (2023)
Prediction with Expert Advice under Local Differential Privacy
di: Jacobsen, Ben, et al.
Pubblicazione: (2025)
di: Jacobsen, Ben, et al.
Pubblicazione: (2025)
Private Continual Counting of Unbounded Streams
di: Jacobsen, Ben, et al.
Pubblicazione: (2025)
di: Jacobsen, Ben, et al.
Pubblicazione: (2025)
GPM: The Gaussian Pancake Mechanism for Planting Undetectable Backdoors in Differential Privacy
di: Sun, Haochen, et al.
Pubblicazione: (2025)
di: Sun, Haochen, et al.
Pubblicazione: (2025)
Backdooring Bias in Large Language Models
di: Das, Anudeep, et al.
Pubblicazione: (2026)
di: Das, Anudeep, et al.
Pubblicazione: (2026)
"Impressively Scary:" Exploring User Perceptions and Reactions to Unraveling Machine Learning Models in Social Media Applications
di: West, Jack, et al.
Pubblicazione: (2025)
di: West, Jack, et al.
Pubblicazione: (2025)
An Undetectable Watermark for Generative Image Models
di: Gunn, Sam, et al.
Pubblicazione: (2024)
di: Gunn, Sam, et al.
Pubblicazione: (2024)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
di: Wang, Chao, et al.
Pubblicazione: (2026)
di: Wang, Chao, et al.
Pubblicazione: (2026)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
di: Ding, Zikang, et al.
Pubblicazione: (2026)
di: Ding, Zikang, et al.
Pubblicazione: (2026)
Auto-SPT: Automating Semantic Preserving Transformations for Code
di: Hooda, Ashish, et al.
Pubblicazione: (2025)
di: Hooda, Ashish, et al.
Pubblicazione: (2025)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
di: Yu, Miao, et al.
Pubblicazione: (2025)
di: Yu, Miao, et al.
Pubblicazione: (2025)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
di: You, Ziyang, et al.
Pubblicazione: (2026)
di: You, Ziyang, et al.
Pubblicazione: (2026)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
di: Hartman, Max, et al.
Pubblicazione: (2026)
di: Hartman, Max, et al.
Pubblicazione: (2026)
Hide and Seek: Fingerprinting Large Language Models with Evolutionary Learning
di: Iourovitski, Dmitri, et al.
Pubblicazione: (2024)
di: Iourovitski, Dmitri, et al.
Pubblicazione: (2024)
Efficient Malware Detection with Optimized Learning on High-Dimensional Features
di: Choudhary, Aditya, et al.
Pubblicazione: (2025)
di: Choudhary, Aditya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025) -
How Not to Detect Prompt Injections with an LLM
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025) -
PolicyLR: A Logic Representation For Privacy Policies
di: Hooda, Ashish, et al.
Pubblicazione: (2024) -
Dependency-Aware Privacy for Multi-turn Agents
di: Anshumaan, Divyam, et al.
Pubblicazione: (2026) -
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
di: Mangaokar, Neal, et al.
Pubblicazione: (2024)