RAT-Bench: A Comprehensive Benchmark for Text Anonymization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Krčo, Nataša, Yao, Zexi, Meeus, Matthieu, de Montjoye, Yves-Alexandre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
von: Yao, Zexi, et al.
Veröffentlicht: (2025)
von: Yao, Zexi, et al.
Veröffentlicht: (2025)
Counterfactual Influence as a Distributional Quantity
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
Lost in the Averages: A New Specific Setup to Evaluate Membership Inference Attacks Against Machine Learning Models
von: Krčo, Nataša, et al.
Veröffentlicht: (2024)
von: Krčo, Nataša, et al.
Veröffentlicht: (2024)
Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)
The Tail Tells All: Estimating Model-Level Membership Inference Vulnerability Without Reference Models
von: Dodd, Euodia, et al.
Veröffentlicht: (2025)
von: Dodd, Euodia, et al.
Veröffentlicht: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2024)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2024)
Copyright Traps for Large Language Models
von: Meeus, Matthieu, et al.
Veröffentlicht: (2024)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2024)
Synthetic is all you need: removing the auxiliary data assumption for membership inference attacks against synthetic data
von: Guépin, Florent, et al.
Veröffentlicht: (2023)
von: Guépin, Florent, et al.
Veröffentlicht: (2023)
Achilles' Heels: Vulnerable Record Identification in Synthetic Data Publishing
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
von: Li, Lijun, et al.
Veröffentlicht: (2024)
von: Li, Lijun, et al.
Veröffentlicht: (2024)
IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization
von: Frikha, Ahmed, et al.
Veröffentlicht: (2024)
von: Frikha, Ahmed, et al.
Veröffentlicht: (2024)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
TaeBench: Improving Quality of Toxic Adversarial Examples
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
Unlocking the Potential of Large Language Models for Clinical Text Anonymization: A Comparative Study
von: Pissarra, David, et al.
Veröffentlicht: (2024)
von: Pissarra, David, et al.
Veröffentlicht: (2024)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
Exploring the limits of strong membership inference attacks on large language models
von: Hayes, Jamie, et al.
Veröffentlicht: (2025)
von: Hayes, Jamie, et al.
Veröffentlicht: (2025)
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
von: Zhang, Andy K., et al.
Veröffentlicht: (2025)
von: Zhang, Andy K., et al.
Veröffentlicht: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
Downstream Trade-offs of a Family of Text Watermarks
von: Ajith, Anirudh, et al.
Veröffentlicht: (2023)
von: Ajith, Anirudh, et al.
Veröffentlicht: (2023)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
von: Wang, Kun, et al.
Veröffentlicht: (2025)
von: Wang, Kun, et al.
Veröffentlicht: (2025)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms
von: Azarafrooz, Ari
Veröffentlicht: (2026)
von: Azarafrooz, Ari
Veröffentlicht: (2026)
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
von: Shahariar, G M, et al.
Veröffentlicht: (2024)
von: Shahariar, G M, et al.
Veröffentlicht: (2024)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
Adversarial Text Purification: A Large Language Model Approach for Defense
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
von: Liu, Ken Ziyu, et al.
Veröffentlicht: (2025)
von: Liu, Ken Ziyu, et al.
Veröffentlicht: (2025)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
von: Li, Yuexin, et al.
Veröffentlicht: (2026)
von: Li, Yuexin, et al.
Veröffentlicht: (2026)
Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
DeSIA: Attribute Inference Attacks Against Limited Fixed Aggregate Statistics
von: Mao, Yifeng, et al.
Veröffentlicht: (2025)
von: Mao, Yifeng, et al.
Veröffentlicht: (2025)
RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors
von: Bai, Fengshuo, et al.
Veröffentlicht: (2024)
von: Bai, Fengshuo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
von: Yao, Zexi, et al.
Veröffentlicht: (2025) -
Counterfactual Influence as a Distributional Quantity
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025) -
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025) -
Lost in the Averages: A New Specific Setup to Evaluate Membership Inference Attacks Against Machine Learning Models
von: Krčo, Nataša, et al.
Veröffentlicht: (2024) -
Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)