Truthful Text Sanitization Guided by Inference Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pilán, Ildikó, Manzanares-Salor, Benet, Sánchez, David, Lison, Pierre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Conversational Feedback in Scripted versus Spontaneous Dialogues: A Comparative Analysis
von: Pilán, Ildikó, et al.
Veröffentlicht: (2023)
von: Pilán, Ildikó, et al.
Veröffentlicht: (2023)
Protecting De-identified Documents from Search-based Linkage Attacks
von: Lison, Pierre, et al.
Veröffentlicht: (2025)
von: Lison, Pierre, et al.
Veröffentlicht: (2025)
Stronger Re-identification Attacks through Reasoning and Aggregation
von: Charpentier, Lucas Georges Gabriel, et al.
Veröffentlicht: (2025)
von: Charpentier, Lucas Georges Gabriel, et al.
Veröffentlicht: (2025)
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
von: Kennington, Casey, et al.
Veröffentlicht: (2025)
von: Kennington, Casey, et al.
Veröffentlicht: (2025)
Re-identification of De-identified Documents with Autoregressive Infilling
von: Charpentier, Lucas Georges Gabriel, et al.
Veröffentlicht: (2025)
von: Charpentier, Lucas Georges Gabriel, et al.
Veröffentlicht: (2025)
Enhancing Naturalness in LLM-Generated Utterances through Disfluency Insertion
von: Hassan, Syed Zohaib, et al.
Veröffentlicht: (2024)
von: Hassan, Syed Zohaib, et al.
Veröffentlicht: (2024)
Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
von: Kåsene, Vebjørn Haug, et al.
Veröffentlicht: (2025)
von: Kåsene, Vebjørn Haug, et al.
Veröffentlicht: (2025)
Machine Text Detectors are Membership Inference Attacks
von: Koike, Ryuto, et al.
Veröffentlicht: (2025)
von: Koike, Ryuto, et al.
Veröffentlicht: (2025)
Evaluating LLMs on Generating Age-Appropriate Child-Like Conversations
von: Hassan, Syed Zohaib, et al.
Veröffentlicht: (2025)
von: Hassan, Syed Zohaib, et al.
Veröffentlicht: (2025)
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
von: Qiu, Huming, et al.
Veröffentlicht: (2024)
von: Qiu, Huming, et al.
Veröffentlicht: (2024)
Knowledge Sanitization of Large Language Models
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2023)
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2023)
Leverage Unlearning to Sanitize LLMs
von: Boutet, Antoine, et al.
Veröffentlicht: (2025)
von: Boutet, Antoine, et al.
Veröffentlicht: (2025)
The Empirical Impact of Data Sanitization on Language Models
von: Pal, Anwesan, et al.
Veröffentlicht: (2024)
von: Pal, Anwesan, et al.
Veröffentlicht: (2024)
Unmasking Database Vulnerabilities: Zero-Knowledge Schema Inference Attacks in Text-to-SQL Systems
von: Klisura, Đorđe, et al.
Veröffentlicht: (2024)
von: Klisura, Đorđe, et al.
Veröffentlicht: (2024)
A Systematic Approach to Predict the Impact of Cybersecurity Vulnerabilities Using LLMs
von: Høst, Anders Mølmen, et al.
Veröffentlicht: (2025)
von: Høst, Anders Mølmen, et al.
Veröffentlicht: (2025)
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
von: Hoscilowicz, Jakub, et al.
Veröffentlicht: (2024)
von: Hoscilowicz, Jakub, et al.
Veröffentlicht: (2024)
Here's a Free Lunch: Sanitizing Backdoored Models with Model Merge
von: Arora, Ansh, et al.
Veröffentlicht: (2024)
von: Arora, Ansh, et al.
Veröffentlicht: (2024)
Truth Neurons
von: Li, Haohang, et al.
Veröffentlicht: (2025)
von: Li, Haohang, et al.
Veröffentlicht: (2025)
Do Membership Inference Attacks Work on Large Language Models?
von: Duan, Michael, et al.
Veröffentlicht: (2024)
von: Duan, Michael, et al.
Veröffentlicht: (2024)
Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks
von: Zhou, Yue, et al.
Veröffentlicht: (2024)
von: Zhou, Yue, et al.
Veröffentlicht: (2024)
Where-to-Unmask: Ground-Truth-Guided Unmasking Order Learning for Masked Diffusion Language Models
von: Asano, Hikaru, et al.
Veröffentlicht: (2026)
von: Asano, Hikaru, et al.
Veröffentlicht: (2026)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs
von: Yaldiz, Duygu Nur, et al.
Veröffentlicht: (2025)
von: Yaldiz, Duygu Nur, et al.
Veröffentlicht: (2025)
Truth Knows No Language: Evaluating Truthfulness Beyond English
von: Figueras, Blanca Calvo, et al.
Veröffentlicht: (2025)
von: Figueras, Blanca Calvo, et al.
Veröffentlicht: (2025)
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
TruthStance: An Annotated Dataset of Conversations on Truth Social
von: Ameen, Fathima, et al.
Veröffentlicht: (2026)
von: Ameen, Fathima, et al.
Veröffentlicht: (2026)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
The Geometries of Truth Are Orthogonal Across Tasks
von: Azizian, Waiss, et al.
Veröffentlicht: (2025)
von: Azizian, Waiss, et al.
Veröffentlicht: (2025)
NatLogAttack: A Framework for Attacking Natural Language Inference Models with Natural Logic
von: Zheng, Zi'ou, et al.
Veröffentlicht: (2023)
von: Zheng, Zi'ou, et al.
Veröffentlicht: (2023)
PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization
von: Liu, Mingshuo, et al.
Veröffentlicht: (2026)
von: Liu, Mingshuo, et al.
Veröffentlicht: (2026)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
von: Fu, Yao, et al.
Veröffentlicht: (2025)
von: Fu, Yao, et al.
Veröffentlicht: (2025)
The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
von: Makroo, Owais, et al.
Veröffentlicht: (2025)
von: Makroo, Owais, et al.
Veröffentlicht: (2025)
PivotAttack: Rethinking the Search Trajectory in Hard-Label Text Attacks via Pivot Words
von: Liang, Yuzhi, et al.
Veröffentlicht: (2026)
von: Liang, Yuzhi, et al.
Veröffentlicht: (2026)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
Understanding the Effects of Iterative Prompting on Truthfulness
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2024)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2024)
On the Universal Truthfulness Hyperplane Inside LLMs
von: Liu, Junteng, et al.
Veröffentlicht: (2024)
von: Liu, Junteng, et al.
Veröffentlicht: (2024)
Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
von: Bao, Yuntai, et al.
Veröffentlicht: (2025)
von: Bao, Yuntai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Conversational Feedback in Scripted versus Spontaneous Dialogues: A Comparative Analysis
von: Pilán, Ildikó, et al.
Veröffentlicht: (2023) -
Protecting De-identified Documents from Search-based Linkage Attacks
von: Lison, Pierre, et al.
Veröffentlicht: (2025) -
Stronger Re-identification Attacks through Reasoning and Aggregation
von: Charpentier, Lucas Georges Gabriel, et al.
Veröffentlicht: (2025) -
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
von: Kennington, Casey, et al.
Veröffentlicht: (2025) -
Re-identification of De-identified Documents with Autoregressive Infilling
von: Charpentier, Lucas Georges Gabriel, et al.
Veröffentlicht: (2025)