AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
Fuente:
arXiv
Saved in:
| Main Author: | Gaikwad, Madhava |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AVEC: Bootstrapping Privacy for Local LLMs
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
CAMP: Cumulative Agentic Masking and Pruning for Privacy Protection in Multi-Turn LLM Conversations
by: Panjwani, Aman
Published: (2026)
by: Panjwani, Aman
Published: (2026)
Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
by: Tsai, Elisa, et al.
Published: (2025)
by: Tsai, Elisa, et al.
Published: (2025)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026)
by: Leonesi, Matteo, et al.
Published: (2026)
Context-aware Privacy Bounds for Linear Queries
by: Zhao, Heng, et al.
Published: (2026)
by: Zhao, Heng, et al.
Published: (2026)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)
by: Wang, Yanshu, et al.
Published: (2026)
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
Powerful Training-Free Membership Inference Against Autoregressive Language Models
by: Ilić, David, et al.
Published: (2026)
by: Ilić, David, et al.
Published: (2026)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
by: Yagoubi, Faouzi El, et al.
Published: (2026)
by: Yagoubi, Faouzi El, et al.
Published: (2026)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
by: DeLeeuw, Caleb
Published: (2026)
by: DeLeeuw, Caleb
Published: (2026)
KidsNanny: A Two-Stage Multimodal Content Moderation Pipeline Integrating Visual Classification, Object Detection, OCR, and Contextual Reasoning for Child Safety
by: Panchal, Viraj, et al.
Published: (2026)
by: Panchal, Viraj, et al.
Published: (2026)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
by: Bercovich, Ivan, et al.
Published: (2026)
by: Bercovich, Ivan, et al.
Published: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
by: Rosenblatt, Lucas, et al.
Published: (2026)
by: Rosenblatt, Lucas, et al.
Published: (2026)
Watermarking for AI Content Detection: A Review on Text, Visual, and Audio Modalities
by: Cao, Lele
Published: (2025)
by: Cao, Lele
Published: (2025)
Whisper Leak: a side-channel attack on Large Language Models
by: McDonald, Geoff, et al.
Published: (2025)
by: McDonald, Geoff, et al.
Published: (2025)
Reasoning-Enhanced Rare-Event Prediction with Balanced Outcome Correction
by: Bulgakov, Vitaly, et al.
Published: (2026)
by: Bulgakov, Vitaly, et al.
Published: (2026)
Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
by: Arora, Sunil, et al.
Published: (2025)
by: Arora, Sunil, et al.
Published: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
by: Zhao, Lepeng, et al.
Published: (2026)
by: Zhao, Lepeng, et al.
Published: (2026)
Towards Modeling Cybersecurity Behavior of Humans in Organizations
by: Kürtz, Klaas Ole
Published: (2026)
by: Kürtz, Klaas Ole
Published: (2026)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
by: Hossain, Ariyan, et al.
Published: (2025)
by: Hossain, Ariyan, et al.
Published: (2025)
CEKER: A Generalizable LLM Framework for Literature Analysis with a Case Study in Unikernel Security
by: Wollman, Alex, et al.
Published: (2024)
by: Wollman, Alex, et al.
Published: (2024)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026)
by: Othman, Refat
Published: (2026)
Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
by: Young, Richard
Published: (2025)
by: Young, Richard
Published: (2025)
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
by: Hartmann, David, et al.
Published: (2026)
by: Hartmann, David, et al.
Published: (2026)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
When Large Language Models are More PersuasiveThan Incentivized Humans, and Why
by: Schoenegger, Philipp, et al.
Published: (2025)
by: Schoenegger, Philipp, et al.
Published: (2025)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
by: Zmanovskii, Nikita
Published: (2025)
by: Zmanovskii, Nikita
Published: (2025)
Deterministic Fuzzy Triage for Legal Compliance Classification and Evidence Retrieval
by: Atri, Rian
Published: (2026)
by: Atri, Rian
Published: (2026)
Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery
by: Junyao, Huang, et al.
Published: (2025)
by: Junyao, Huang, et al.
Published: (2025)
From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection
by: Wang, Mo, et al.
Published: (2026)
by: Wang, Mo, et al.
Published: (2026)
The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
by: Fenech-Borg, Emanuel Z., et al.
Published: (2025)
by: Fenech-Borg, Emanuel Z., et al.
Published: (2025)
AI-Powered Citation Auditing: A Zero-Assumption Protocol for Systematic Reference Verification in Academic Research
by: van Rensburg, L. J. Janse
Published: (2025)
by: van Rensburg, L. J. Janse
Published: (2025)
Scalable and Ethical Insider Threat Detection through Data Synthesis and Analysis by LLMs
by: Gelman, Haywood, et al.
Published: (2025)
by: Gelman, Haywood, et al.
Published: (2025)
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
by: Chen, Renmiao, et al.
Published: (2025)
by: Chen, Renmiao, et al.
Published: (2025)
Benchmarking Deception Probes via Black-to-White Performance Boosts
by: Parrack, Avi, et al.
Published: (2025)
by: Parrack, Avi, et al.
Published: (2025)
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
by: Zhu, Pengyun, et al.
Published: (2026)
by: Zhu, Pengyun, et al.
Published: (2026)
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
by: Dékány, Csaba, et al.
Published: (2025)
by: Dékány, Csaba, et al.
Published: (2025)
Similar Items
-
AVEC: Bootstrapping Privacy for Local LLMs
by: Gaikwad, Madhava
Published: (2025) -
CAMP: Cumulative Agentic Masking and Pruning for Privacy Protection in Multi-Turn LLM Conversations
by: Panjwani, Aman
Published: (2026) -
Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
by: Tsai, Elisa, et al.
Published: (2025) -
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026) -
Context-aware Privacy Bounds for Linear Queries
by: Zhao, Heng, et al.
Published: (2026)