SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Shaaban, Mohamed, Elmahallawy, Mohamed |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Secure Aggregation Falls Short: Achieving Long-Term Privacy in Asynchronous Federated Learning for LEO Satellite Networks
por: Elmahallawy, Mohamed, et al.
Publicado: (2025)
por: Elmahallawy, Mohamed, et al.
Publicado: (2025)
Secure and Privacy-Preserving Federated Learning for Next-Generation Underground Mine Safety
por: Elmahallawy, Mohamed, et al.
Publicado: (2025)
por: Elmahallawy, Mohamed, et al.
Publicado: (2025)
Decentralized Trust for Space AI: Blockchain-Based Federated Learning Across Multi-Vendor LEO Satellite Networks
por: Elmahallawy, Mohamed, et al.
Publicado: (2025)
por: Elmahallawy, Mohamed, et al.
Publicado: (2025)
CAPID: Context-Aware PII Detection for Question-Answering Systems
por: Ponomarenko, Mariia, et al.
Publicado: (2026)
por: Ponomarenko, Mariia, et al.
Publicado: (2026)
Detecting Untargeted Attacks and Mitigating Unreliable Updates in Federated Learning for Underground Mining Operations
por: Rahman, Md Sazedur, et al.
Publicado: (2025)
por: Rahman, Md Sazedur, et al.
Publicado: (2025)
Safely Learning with Private Data: A Federated Learning Framework for Large Language Model
por: Zheng, JiaYing, et al.
Publicado: (2024)
por: Zheng, JiaYing, et al.
Publicado: (2024)
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
por: Hughes, Anthony, et al.
Publicado: (2025)
por: Hughes, Anthony, et al.
Publicado: (2025)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
por: Nakka, Krishna Kanth, et al.
Publicado: (2024)
por: Nakka, Krishna Kanth, et al.
Publicado: (2024)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
por: Xie, Yueqi, et al.
Publicado: (2024)
por: Xie, Yueqi, et al.
Publicado: (2024)
PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization
por: Liu, Mingshuo, et al.
Publicado: (2026)
por: Liu, Mingshuo, et al.
Publicado: (2026)
PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
por: Nakka, Krishna Kanth, et al.
Publicado: (2025)
por: Nakka, Krishna Kanth, et al.
Publicado: (2025)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
por: Sun, Bowen, et al.
Publicado: (2026)
por: Sun, Bowen, et al.
Publicado: (2026)
Securing Genomic Data Against Inference Attacks in Federated Learning Environments
por: Pathade, Chetan, et al.
Publicado: (2025)
por: Pathade, Chetan, et al.
Publicado: (2025)
"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents
por: Roy, Soham, et al.
Publicado: (2026)
por: Roy, Soham, et al.
Publicado: (2026)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
por: Wang, Jiawen, et al.
Publicado: (2025)
por: Wang, Jiawen, et al.
Publicado: (2025)
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
por: Alhazbi, Saeif, et al.
Publicado: (2025)
por: Alhazbi, Saeif, et al.
Publicado: (2025)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
por: Zhou, Yukai, et al.
Publicado: (2025)
por: Zhou, Yukai, et al.
Publicado: (2025)
Token-level Data Selection for Safe LLM Fine-tuning
por: Li, Yanping, et al.
Publicado: (2026)
por: Li, Yanping, et al.
Publicado: (2026)
GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
por: Wu, Lichao, et al.
Publicado: (2025)
por: Wu, Lichao, et al.
Publicado: (2025)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
por: Kim, Jinhwa, et al.
Publicado: (2025)
por: Kim, Jinhwa, et al.
Publicado: (2025)
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
por: Shen, Hao, et al.
Publicado: (2025)
por: Shen, Hao, et al.
Publicado: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
por: Li, Yucheng, et al.
Publicado: (2025)
por: Li, Yucheng, et al.
Publicado: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
por: Huang, Caishuang, et al.
Publicado: (2024)
por: Huang, Caishuang, et al.
Publicado: (2024)
Fingerprinting LLMs via Prompt Injection
por: Hu, Yuepeng, et al.
Publicado: (2025)
por: Hu, Yuepeng, et al.
Publicado: (2025)
Efficient Provably Secure Linguistic Steganography via Range Coding
por: Yan, Ruiyi, et al.
Publicado: (2026)
por: Yan, Ruiyi, et al.
Publicado: (2026)
Hidden Data Privacy Breaches in Federated Learning
por: Gong, Xueluan, et al.
Publicado: (2024)
por: Gong, Xueluan, et al.
Publicado: (2024)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
por: Zhu, He, et al.
Publicado: (2026)
por: Zhu, He, et al.
Publicado: (2026)
Lightweight Yet Secure: Secure Scripting Language Generation via Lightweight LLMs
por: Zhang, Keyang, et al.
Publicado: (2026)
por: Zhang, Keyang, et al.
Publicado: (2026)
The Ethics of Interaction: Mitigating Security Threats in LLMs
por: Kumar, Ashutosh, et al.
Publicado: (2024)
por: Kumar, Ashutosh, et al.
Publicado: (2024)
Protecting Privacy in Classifiers by Token Manipulation
por: Harel, Re'em, et al.
Publicado: (2024)
por: Harel, Re'em, et al.
Publicado: (2024)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
por: Liu, Yepeng, et al.
Publicado: (2025)
por: Liu, Yepeng, et al.
Publicado: (2025)
Token-Level Privacy in Large Language Models
por: Harel, Re'em, et al.
Publicado: (2025)
por: Harel, Re'em, et al.
Publicado: (2025)
Privacy-Preserving Transformers: SwiftKey's Differential Privacy Implementation
por: Abouelenin, Abdelrahman, et al.
Publicado: (2025)
por: Abouelenin, Abdelrahman, et al.
Publicado: (2025)
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
por: Salahuddin, Salahuddin, et al.
Publicado: (2025)
por: Salahuddin, Salahuddin, et al.
Publicado: (2025)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
por: Afane, Mohamed, et al.
Publicado: (2025)
por: Afane, Mohamed, et al.
Publicado: (2025)
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
por: Xu, Ning, et al.
Publicado: (2025)
por: Xu, Ning, et al.
Publicado: (2025)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
por: Gu, Tianle, et al.
Publicado: (2025)
por: Gu, Tianle, et al.
Publicado: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
por: Wang, Yanbo, et al.
Publicado: (2026)
por: Wang, Yanbo, et al.
Publicado: (2026)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
por: Xu, Xiangzhe, et al.
Publicado: (2024)
por: Xu, Xiangzhe, et al.
Publicado: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
Ejemplares similares
-
When Secure Aggregation Falls Short: Achieving Long-Term Privacy in Asynchronous Federated Learning for LEO Satellite Networks
por: Elmahallawy, Mohamed, et al.
Publicado: (2025) -
Secure and Privacy-Preserving Federated Learning for Next-Generation Underground Mine Safety
por: Elmahallawy, Mohamed, et al.
Publicado: (2025) -
Decentralized Trust for Space AI: Blockchain-Based Federated Learning Across Multi-Vendor LEO Satellite Networks
por: Elmahallawy, Mohamed, et al.
Publicado: (2025) -
CAPID: Context-Aware PII Detection for Question-Answering Systems
por: Ponomarenko, Mariia, et al.
Publicado: (2026) -
Detecting Untargeted Attacks and Mitigating Unreliable Updates in Federated Learning for Underground Mining Operations
por: Rahman, Md Sazedur, et al.
Publicado: (2025)