DOMBA: Double Model Balancing for Access-Controlled Language Models via Minimum-Bounded Aggregation
Fuente:
arXiv
Saved in:
| Main Authors: | Segal, Tom, Shabtai, Asaf, Elovici, Yuval |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
by: Abaev, Nadya, et al.
Published: (2026)
by: Abaev, Nadya, et al.
Published: (2026)
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
by: Schwartz, Yuval, et al.
Published: (2024)
by: Schwartz, Yuval, et al.
Published: (2024)
Tag&Tab: Pretraining Data Detection in Large Language Models Using Keyword-Based Membership Inference Attack
by: Antebi, Sagiv, et al.
Published: (2025)
by: Antebi, Sagiv, et al.
Published: (2025)
MIA-EPT: Membership Inference Attack via Error Prediction for Tabular Data
by: German, Eyal, et al.
Published: (2025)
by: German, Eyal, et al.
Published: (2025)
GenKubeSec: LLM-Based Kubernetes Misconfiguration Detection, Localization, Reasoning, and Remediation
by: Malul, Ehud, et al.
Published: (2024)
by: Malul, Ehud, et al.
Published: (2024)
RAPID: Robust APT Detection and Investigation Using Context-Aware Deep Learning
by: Amaru, Yonatan, et al.
Published: (2024)
by: Amaru, Yonatan, et al.
Published: (2024)
Real-World Adversarial Attacks on RF-Based Drone Detectors
by: Gazit, Omer, et al.
Published: (2025)
by: Gazit, Omer, et al.
Published: (2025)
QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers
by: Baras, Amit, et al.
Published: (2023)
by: Baras, Amit, et al.
Published: (2023)
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
by: German, Eyal, et al.
Published: (2025)
by: German, Eyal, et al.
Published: (2025)
From Tool Orchestration to Code Execution: A Study of MCP Design Choices
by: Felendler, Yuval, et al.
Published: (2026)
by: Felendler, Yuval, et al.
Published: (2026)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
by: Noah, Amit Finkman, et al.
Published: (2024)
by: Noah, Amit Finkman, et al.
Published: (2024)
RuleGenie: SIEM Detection Rule Set Optimization
by: Shukla, Akansha, et al.
Published: (2025)
by: Shukla, Akansha, et al.
Published: (2025)
LISAA: A Framework for Large Language Model Information Security Awareness Assessment
by: Cohen, Ofir, et al.
Published: (2024)
by: Cohen, Ofir, et al.
Published: (2024)
Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs
by: German, Eyal, et al.
Published: (2025)
by: German, Eyal, et al.
Published: (2025)
Rogue Cell: Adversarial Attack and Defense in Untrusted O-RAN Setup Exploiting the Traffic Steering xApp
by: Aizikovich, Eran, et al.
Published: (2025)
by: Aizikovich, Eran, et al.
Published: (2025)
Detection of Compromised Functions in a Serverless Cloud Environment
by: Lavi, Danielle, et al.
Published: (2024)
by: Lavi, Danielle, et al.
Published: (2024)
DeSparsify: Adversarial Attack Against Token Sparsification Mechanisms in Vision Transformers
by: Yehezkel, Oryan, et al.
Published: (2024)
by: Yehezkel, Oryan, et al.
Published: (2024)
GPT in Sheep's Clothing: The Risk of Customized GPTs
by: Antebi, Sagiv, et al.
Published: (2024)
by: Antebi, Sagiv, et al.
Published: (2024)
KubeGuard: LLM-Assisted Kubernetes Hardening via Configuration Files and Runtime Logs Analysis
by: Cohen, Omri Sgan, et al.
Published: (2025)
by: Cohen, Omri Sgan, et al.
Published: (2025)
Watermarking Makes Language Models Radioactive
by: Sander, Tom, et al.
Published: (2024)
by: Sander, Tom, et al.
Published: (2024)
SecMate: Multi-Agent Adaptive Cybersecurity Troubleshooting with Tri-Context Personalization
by: Meidan, Yair, et al.
Published: (2026)
by: Meidan, Yair, et al.
Published: (2026)
Detecting Training Data of Large Language Models via Expectation Maximization
by: Kim, Gyuwan, et al.
Published: (2024)
by: Kim, Gyuwan, et al.
Published: (2024)
Extracting Training Data from Diffusion Language Models via Infilling
by: Wang, Yihan, et al.
Published: (2026)
by: Wang, Yihan, et al.
Published: (2026)
ATAG: AI-Agent Application Threat Assessment with Attack Graphs
by: Gandhi, Parth Atulbhai, et al.
Published: (2025)
by: Gandhi, Parth Atulbhai, et al.
Published: (2025)
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
by: Wang, Guanchu, et al.
Published: (2024)
by: Wang, Guanchu, et al.
Published: (2024)
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Instructional Fingerprinting of Large Language Models
by: Xu, Jiashu, et al.
Published: (2024)
by: Xu, Jiashu, et al.
Published: (2024)
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026)
by: Subramani, Nishant, et al.
Published: (2026)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
by: Boreiko, Valentyn, et al.
Published: (2024)
by: Boreiko, Valentyn, et al.
Published: (2024)
MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
by: Tong, Xin, et al.
Published: (2025)
by: Tong, Xin, et al.
Published: (2025)
Rethinking How to Evaluate Language Model Jailbreak
by: Cai, Hongyu, et al.
Published: (2024)
by: Cai, Hongyu, et al.
Published: (2024)
Large Language Models in Cybersecurity: State-of-the-Art
by: Motlagh, Farzad Nourmohammadzadeh, et al.
Published: (2024)
by: Motlagh, Farzad Nourmohammadzadeh, et al.
Published: (2024)
Duwak: Dual Watermarks in Large Language Models
by: Zhu, Chaoyi, et al.
Published: (2024)
by: Zhu, Chaoyi, et al.
Published: (2024)
Jailbreaking Large Language Models with Symbolic Mathematics
by: Bethany, Emet, et al.
Published: (2024)
by: Bethany, Emet, et al.
Published: (2024)
Learnable Privacy Neurons Localization in Language Models
by: Chen, Ruizhe, et al.
Published: (2024)
by: Chen, Ruizhe, et al.
Published: (2024)
On Adversarial Robustness of Language Models in Transfer Learning
by: Turbal, Bohdan, et al.
Published: (2024)
by: Turbal, Bohdan, et al.
Published: (2024)
Directional Embedding Smoothing for Robust Vision Language Models
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
The Resurgence of GCG Adversarial Attacks on Large Language Models
by: Tan, Yuting, et al.
Published: (2025)
by: Tan, Yuting, et al.
Published: (2025)
Similar Items
-
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
by: Abaev, Nadya, et al.
Published: (2026) -
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
by: Schwartz, Yuval, et al.
Published: (2024) -
Tag&Tab: Pretraining Data Detection in Large Language Models Using Keyword-Based Membership Inference Attack
by: Antebi, Sagiv, et al.
Published: (2025) -
MIA-EPT: Membership Inference Attack via Error Prediction for Tabular Data
by: German, Eyal, et al.
Published: (2025) -
GenKubeSec: LLM-Based Kubernetes Misconfiguration Detection, Localization, Reasoning, and Remediation
by: Malul, Ehud, et al.
Published: (2024)