Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kalavasis, Alkis, Karbasi, Amin, Oikonomou, Argyris, Sotiraki, Katerina, Velegkas, Grigoris, Zampetakis, Manolis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Computational Landscape of Replicable Learning
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
Optimal Learners for Realizable Regression: PAC Learning and Online Learning
by: Attias, Idan, et al.
Published: (2023)
by: Attias, Idan, et al.
Published: (2023)
Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks
by: Eggen, Marte, et al.
Published: (2026)
by: Eggen, Marte, et al.
Published: (2026)
Planting Undetectable Backdoors in Machine Learning Models
by: Goldwasser, Shafi, et al.
Published: (2022)
by: Goldwasser, Shafi, et al.
Published: (2022)
Replicable Learning of Large-Margin Halfspaces
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Private Statistical Estimation via Truncation
by: Zampetakis, Manolis, et al.
Published: (2025)
by: Zampetakis, Manolis, et al.
Published: (2025)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
by: Choudhary, Sarthak, et al.
Published: (2026)
by: Choudhary, Sarthak, et al.
Published: (2026)
Smaller Confidence Intervals From IPW Estimators via Data-Dependent Coarsening
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
Transfer Learning Beyond Bounded Density Ratios
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
Prompt Obfuscation for Large Language Models
by: Pape, David, et al.
Published: (2024)
by: Pape, David, et al.
Published: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
by: An, Shengwei, et al.
Published: (2023)
by: An, Shengwei, et al.
Published: (2023)
Cryptographic Backdoor for Neural Networks: Boon and Bane
by: Ngo, Anh Tu, et al.
Published: (2025)
by: Ngo, Anh Tu, et al.
Published: (2025)
Command-line Obfuscation Detection using Small Language Models
by: Outrata, Vojtech, et al.
Published: (2024)
by: Outrata, Vojtech, et al.
Published: (2024)
Prediction-Augmented Trees for Reliable Statistical Inference
by: Kher, Vikram, et al.
Published: (2025)
by: Kher, Vikram, et al.
Published: (2025)
On the Learning Curves of Revenue Maximization
by: Hanneke, Steve, et al.
Published: (2026)
by: Hanneke, Steve, et al.
Published: (2026)
Defending against Backdoor Attack on Deep Neural Networks
by: Cheng, Hao, et al.
Published: (2020)
by: Cheng, Hao, et al.
Published: (2020)
What Makes Treatment Effects Identifiable? Characterizations and Estimators Beyond Unconfoundedness
by: Cai, Yang, et al.
Published: (2025)
by: Cai, Yang, et al.
Published: (2025)
PrivDiffuser: Privacy-Guided Diffusion Model for Data Obfuscation in Sensor Networks
by: Yang, Xin, et al.
Published: (2024)
by: Yang, Xin, et al.
Published: (2024)
DMGNN: Detecting and Mitigating Backdoor Attacks in Graph Neural Networks
by: Sui, Hao, et al.
Published: (2024)
by: Sui, Hao, et al.
Published: (2024)
Adaptive Backdoor Attacks with Reasonable Constraints on Graph Neural Networks
by: Dong, Xuewen, et al.
Published: (2025)
by: Dong, Xuewen, et al.
Published: (2025)
Backdooring Masked Diffusion Language Models
by: Cao, Daniel Yiming, et al.
Published: (2026)
by: Cao, Daniel Yiming, et al.
Published: (2026)
Amalgam: A Framework for Obfuscated Neural Network Training on the Cloud
by: Taki, Sifat Ut, et al.
Published: (2024)
by: Taki, Sifat Ut, et al.
Published: (2024)
On Union-Closedness of Language Generation
by: Hanneke, Steve, et al.
Published: (2025)
by: Hanneke, Steve, et al.
Published: (2025)
BlockDoor: Blocking Backdoor Based Watermarks in Deep Neural Networks
by: Puah, Yi Hao, et al.
Published: (2024)
by: Puah, Yi Hao, et al.
Published: (2024)
URVFL: Undetectable Data Reconstruction Attack on Vertical Federated Learning
by: Yao, Duanyi, et al.
Published: (2024)
by: Yao, Duanyi, et al.
Published: (2024)
Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
by: Gosch, Lukas, et al.
Published: (2024)
by: Gosch, Lukas, et al.
Published: (2024)
An Undetectable Watermark for Generative Image Models
by: Gunn, Sam, et al.
Published: (2024)
by: Gunn, Sam, et al.
Published: (2024)
Learning Mixture Models via Efficient High-dimensional Sparse Fourier Transforms
by: Kalavasis, Alkis, et al.
Published: (2026)
by: Kalavasis, Alkis, et al.
Published: (2026)
(Im)possibility of Automated Hallucination Detection in Large Language Models
by: Karbasi, Amin, et al.
Published: (2025)
by: Karbasi, Amin, et al.
Published: (2025)
BeniFul: Backdoor Defense via Middle Feature Analysis for Deep Neural Networks
by: Li, Xinfu, et al.
Published: (2024)
by: Li, Xinfu, et al.
Published: (2024)
Good-Enough LLM Obfuscation (GELO)
by: Belikov, Anatoly, et al.
Published: (2026)
by: Belikov, Anatoly, et al.
Published: (2026)
Unelicitable Backdoors in Language Models via Cryptographic Transformer Circuits
by: Draguns, Andis, et al.
Published: (2024)
by: Draguns, Andis, et al.
Published: (2024)
Imperio: Language-Guided Backdoor Attacks for Arbitrary Model Control
by: Chow, Ka-Ho, et al.
Published: (2024)
by: Chow, Ka-Ho, et al.
Published: (2024)
Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models
by: Wan, Guangnian, et al.
Published: (2026)
by: Wan, Guangnian, et al.
Published: (2026)
What is Learnable in Valiant's Theory of the Learnable?
by: Hanneke, Steve, et al.
Published: (2026)
by: Hanneke, Steve, et al.
Published: (2026)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
by: Cui, Xinyue, et al.
Published: (2025)
by: Cui, Xinyue, et al.
Published: (2025)
Similar Items
-
On the Computational Landscape of Replicable Learning
by: Kalavasis, Alkis, et al.
Published: (2024) -
Optimal Learners for Realizable Regression: PAC Learning and Online Learning
by: Attias, Idan, et al.
Published: (2023) -
Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks
by: Eggen, Marte, et al.
Published: (2026) -
Planting Undetectable Backdoors in Machine Learning Models
by: Goldwasser, Shafi, et al.
Published: (2022) -
Replicable Learning of Large-Margin Halfspaces
by: Kalavasis, Alkis, et al.
Published: (2024)