How reparametrization trick broke differentially-private text representation learning
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Habernal, Ivan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?
von: Weiss, Christopher, et al.
Veröffentlicht: (2023)
von: Weiss, Christopher, et al.
Veröffentlicht: (2023)
DP-BART for Privatized Text Rewriting under Local Differential Privacy
von: Igamberdiev, Timour, et al.
Veröffentlicht: (2023)
von: Igamberdiev, Timour, et al.
Veröffentlicht: (2023)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
von: Wang, Jiawen, et al.
Veröffentlicht: (2025)
von: Wang, Jiawen, et al.
Veröffentlicht: (2025)
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024)
Efficient derandomization of differentially private counting queries
von: Ghentiyala, Surendra
Veröffentlicht: (2025)
von: Ghentiyala, Surendra
Veröffentlicht: (2025)
Differentially-private text generation degrades output language quality
von: Çano, Erion, et al.
Veröffentlicht: (2025)
von: Çano, Erion, et al.
Veröffentlicht: (2025)
Private prediction for large-scale synthetic text generation
von: Amin, Kareem, et al.
Veröffentlicht: (2024)
von: Amin, Kareem, et al.
Veröffentlicht: (2024)
How Good is Post-Hoc Watermarking With Language Model Rephrasing?
von: Fernandez, Pierre, et al.
Veröffentlicht: (2025)
von: Fernandez, Pierre, et al.
Veröffentlicht: (2025)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
LLMs can hide text in other text of the same length
von: Norelli, Antonio, et al.
Veröffentlicht: (2025)
von: Norelli, Antonio, et al.
Veröffentlicht: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
von: Zhang, Rui, et al.
Veröffentlicht: (2026)
von: Zhang, Rui, et al.
Veröffentlicht: (2026)
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
von: Banerjee, Somnath, et al.
Veröffentlicht: (2024)
von: Banerjee, Somnath, et al.
Veröffentlicht: (2024)
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
von: Liang, Zi, et al.
Veröffentlicht: (2025)
von: Liang, Zi, et al.
Veröffentlicht: (2025)
What are the attackers doing now? Automating cyber threat intelligence extraction from text on pace with the changing threat landscape: A survey
von: Rahman, Md Rayhanur, et al.
Veröffentlicht: (2021)
von: Rahman, Md Rayhanur, et al.
Veröffentlicht: (2021)
How Susceptible are Large Language Models to Ideological Manipulation?
von: Chen, Kai, et al.
Veröffentlicht: (2024)
von: Chen, Kai, et al.
Veröffentlicht: (2024)
How Vulnerable Are Edge LLMs?
von: Ding, Ao, et al.
Veröffentlicht: (2026)
von: Ding, Ao, et al.
Veröffentlicht: (2026)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
Out of the Cage: How Stochastic Parrots Win in Cyber Security Environments
von: Rigaki, Maria, et al.
Veröffentlicht: (2023)
von: Rigaki, Maria, et al.
Veröffentlicht: (2023)
FLAME: Flexible LLM-Assisted Moderation Engine
von: Bakulin, Ivan, et al.
Veröffentlicht: (2025)
von: Bakulin, Ivan, et al.
Veröffentlicht: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2024)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2024)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
von: Nawal, Aditya, et al.
Veröffentlicht: (2026)
von: Nawal, Aditya, et al.
Veröffentlicht: (2026)
Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage
von: Hoque, Shahinul, et al.
Veröffentlicht: (2026)
von: Hoque, Shahinul, et al.
Veröffentlicht: (2026)
Harnessing large-language models to generate private synthetic text
von: Kurakin, Alexey, et al.
Veröffentlicht: (2023)
von: Kurakin, Alexey, et al.
Veröffentlicht: (2023)
Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
von: He, Haorui, et al.
Veröffentlicht: (2025)
von: He, Haorui, et al.
Veröffentlicht: (2025)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research?
von: Chen, Xiangsen, et al.
Veröffentlicht: (2026)
von: Chen, Xiangsen, et al.
Veröffentlicht: (2026)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
von: Dou, Yipu, et al.
Veröffentlicht: (2026)
von: Dou, Yipu, et al.
Veröffentlicht: (2026)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
ChatNVD: Advancing Cybersecurity Vulnerability Assessment with Large Language Models
von: Chopra, Shivansh, et al.
Veröffentlicht: (2024)
von: Chopra, Shivansh, et al.
Veröffentlicht: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles
von: Siyan, Li, et al.
Veröffentlicht: (2024)
von: Siyan, Li, et al.
Veröffentlicht: (2024)
A Character-based Diffusion Embedding Algorithm for Enhancing the Generation Quality of Generative Linguistic Steganographic Texts
von: Chen, Yingquan, et al.
Veröffentlicht: (2025)
von: Chen, Yingquan, et al.
Veröffentlicht: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
von: Pathade, Chetan
Veröffentlicht: (2025)
von: Pathade, Chetan
Veröffentlicht: (2025)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
Rethinking Backdoor Detection Evaluation for Language Models
von: Yan, Jun, et al.
Veröffentlicht: (2024)
von: Yan, Jun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?
von: Weiss, Christopher, et al.
Veröffentlicht: (2023) -
DP-BART for Privatized Text Rewriting under Local Differential Privacy
von: Igamberdiev, Timour, et al.
Veröffentlicht: (2023) -
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
von: Wang, Jiawen, et al.
Veröffentlicht: (2025) -
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024) -
Efficient derandomization of differentially private counting queries
von: Ghentiyala, Surendra
Veröffentlicht: (2025)