Boundary-targeted Membership Inference Attacks on Safety Classifiers
Fuente:
arXiv
Guardado en:
| Autores principales: | Hughes, Anthony, Goldberg, Alexander, Jha, Prince, Perer, Adam, Aletras, Nikolaos, Mireshghallah, Niloofar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
por: Naseh, Ali, et al.
Publicado: (2025)
por: Naseh, Ali, et al.
Publicado: (2025)
Position: Privacy Is Not Just Memorization!
por: Mireshghallah, Niloofar, et al.
Publicado: (2025)
por: Mireshghallah, Niloofar, et al.
Publicado: (2025)
Incorporating Attribution Importance for Improving Faithfulness Metrics
por: Zhao, Zhixue, et al.
Publicado: (2023)
por: Zhao, Zhixue, et al.
Publicado: (2023)
The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
por: Makroo, Owais, et al.
Publicado: (2025)
por: Makroo, Owais, et al.
Publicado: (2025)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
por: Xue, Huiyin, et al.
Publicado: (2025)
por: Xue, Huiyin, et al.
Publicado: (2025)
Where does output diversity collapse in post-training?
por: Karouzos, Constantinos, et al.
Publicado: (2026)
por: Karouzos, Constantinos, et al.
Publicado: (2026)
An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift
por: Karouzos, Constantinos, et al.
Publicado: (2026)
por: Karouzos, Constantinos, et al.
Publicado: (2026)
We Need to Talk About Classification Evaluation Metrics in NLP
por: Vickers, Peter, et al.
Publicado: (2024)
por: Vickers, Peter, et al.
Publicado: (2024)
Membership Inference Attacks and Privacy in Topic Modeling
por: Manzonelli, Nico, et al.
Publicado: (2024)
por: Manzonelli, Nico, et al.
Publicado: (2024)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
por: Mireshghallah, Niloofar, et al.
Publicado: (2023)
por: Mireshghallah, Niloofar, et al.
Publicado: (2023)
How Private are Language Models in Abstractive Summarization?
por: Hughes, Anthony, et al.
Publicado: (2024)
por: Hughes, Anthony, et al.
Publicado: (2024)
Do Membership Inference Attacks Work on Large Language Models?
por: Duan, Michael, et al.
Publicado: (2024)
por: Duan, Michael, et al.
Publicado: (2024)
Differentially Private Learning Needs Better Model Initialization and Self-Distillation
por: Ngong, Ivoline C., et al.
Publicado: (2024)
por: Ngong, Ivoline C., et al.
Publicado: (2024)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
por: Das, Debeshee, et al.
Publicado: (2024)
por: Das, Debeshee, et al.
Publicado: (2024)
Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks
por: Villegas, Danae Sánchez, et al.
Publicado: (2023)
por: Villegas, Danae Sánchez, et al.
Publicado: (2023)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
por: Meeus, Matthieu, et al.
Publicado: (2024)
por: Meeus, Matthieu, et al.
Publicado: (2024)
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
por: Nguyen, Khanh, et al.
Publicado: (2025)
por: Nguyen, Khanh, et al.
Publicado: (2025)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
por: Puerto, Haritz, et al.
Publicado: (2024)
por: Puerto, Haritz, et al.
Publicado: (2024)
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
por: Tran, Toan, et al.
Publicado: (2025)
por: Tran, Toan, et al.
Publicado: (2025)
Blackbox Model Provenance via Palimpsestic Membership Inference
por: Kuditipudi, Rohith, et al.
Publicado: (2025)
por: Kuditipudi, Rohith, et al.
Publicado: (2025)
Membership Inference Attacks on LLM-based Recommender Systems
por: He, Jiajie, et al.
Publicado: (2025)
por: He, Jiajie, et al.
Publicado: (2025)
CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation
por: Chen, Tong, et al.
Publicado: (2024)
por: Chen, Tong, et al.
Publicado: (2024)
Membership Inference Attacks against Large Vision-Language Models
por: Li, Zhan, et al.
Publicado: (2024)
por: Li, Zhan, et al.
Publicado: (2024)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
por: Fu, Wenjie, et al.
Publicado: (2023)
por: Fu, Wenjie, et al.
Publicado: (2023)
Context-Aware Membership Inference Attacks against Pre-trained Large Language Models
por: Chang, Hongyan, et al.
Publicado: (2024)
por: Chang, Hongyan, et al.
Publicado: (2024)
ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods
por: Xie, Roy, et al.
Publicado: (2024)
por: Xie, Roy, et al.
Publicado: (2024)
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
por: Villegas, Danae Sánchez, et al.
Publicado: (2026)
por: Villegas, Danae Sánchez, et al.
Publicado: (2026)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
por: Hallinan, Skyler, et al.
Publicado: (2025)
por: Hallinan, Skyler, et al.
Publicado: (2025)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
por: Chen, Tong, et al.
Publicado: (2025)
por: Chen, Tong, et al.
Publicado: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
por: Schaeffer, Rylan, et al.
Publicado: (2026)
por: Schaeffer, Rylan, et al.
Publicado: (2026)
Aviation Safety Enhancement via NLP & Deep Learning: Classifying Flight Phases in ATSB Safety Reports
por: Nanyonga, Aziida, et al.
Publicado: (2025)
por: Nanyonga, Aziida, et al.
Publicado: (2025)
Self-Comparison for Dataset-Level Membership Inference in Large (Vision-)Language Models
por: Ren, Jie, et al.
Publicado: (2024)
por: Ren, Jie, et al.
Publicado: (2024)
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
por: Xin, Rui, et al.
Publicado: (2025)
por: Xin, Rui, et al.
Publicado: (2025)
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
por: Hughes, Anthony, et al.
Publicado: (2025)
por: Hughes, Anthony, et al.
Publicado: (2025)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
por: Sahili, Ali Al, et al.
Publicado: (2025)
por: Sahili, Ali Al, et al.
Publicado: (2025)
Reversible Jump Attack to Textual Classifiers with Modification Reduction
por: Ni, Mingze, et al.
Publicado: (2024)
por: Ni, Mingze, et al.
Publicado: (2024)
Characterizing Prompt Compression Methods for Long Context Inference
por: Jha, Siddharth, et al.
Publicado: (2024)
por: Jha, Siddharth, et al.
Publicado: (2024)
Natural Language Processing and Deep Learning Models to Classify Phase of Flight in Aviation Safety Occurrences
por: Nanyonga, Aziida, et al.
Publicado: (2025)
por: Nanyonga, Aziida, et al.
Publicado: (2025)
Scalable Bayesian Low-Rank Adaptation of Large Language Models via Stochastic Variational Subspace Inference
por: Samplawski, Colin, et al.
Publicado: (2025)
por: Samplawski, Colin, et al.
Publicado: (2025)
Ejemplares similares
-
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
por: Naseh, Ali, et al.
Publicado: (2025) -
Position: Privacy Is Not Just Memorization!
por: Mireshghallah, Niloofar, et al.
Publicado: (2025) -
Incorporating Attribution Importance for Improving Faithfulness Metrics
por: Zhao, Zhixue, et al.
Publicado: (2023) -
The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
por: Makroo, Owais, et al.
Publicado: (2025) -
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
por: Xue, Huiyin, et al.
Publicado: (2025)