When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Xinyue, Shen, Yun, Backes, Michael, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Voice Jailbreak Attacks Against GPT-4o
by: Shen, Xinyue, et al.
Published: (2024)
by: Shen, Xinyue, et al.
Published: (2024)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
MGTBench: Benchmarking Machine-Generated Text Detection
by: He, Xinlei, et al.
Published: (2023)
by: He, Xinlei, et al.
Published: (2023)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
The Challenge of Identifying the Origin of Black-Box Large Language Models
by: Yang, Ziqing, et al.
Published: (2025)
by: Yang, Ziqing, et al.
Published: (2025)
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
by: Wu, Yixin, et al.
Published: (2025)
by: Wu, Yixin, et al.
Published: (2025)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
Instruction Backdoor Attacks Against Customized LLMs
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
by: Wen, Rui, et al.
Published: (2024)
by: Wen, Rui, et al.
Published: (2024)
Vera Verto: Multimodal Hijacking Attack
by: Zhang, Minxing, et al.
Published: (2024)
by: Zhang, Minxing, et al.
Published: (2024)
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
by: Wen, Rui, et al.
Published: (2025)
by: Wen, Rui, et al.
Published: (2025)
Transferable Availability Poisoning Attacks
by: Liu, Yiyong, et al.
Published: (2023)
by: Liu, Yiyong, et al.
Published: (2023)
Excessive Reasoning Attack on Reasoning LLMs
by: Si, Wai Man, et al.
Published: (2025)
by: Si, Wai Man, et al.
Published: (2025)
When Evaluation Becomes a Side Channel: Regime Leakage and Structural Mitigations for Alignment Assessment
by: Santos-Grueiro, Igor
Published: (2026)
by: Santos-Grueiro, Igor
Published: (2026)
Location Leakage in Federated Signal Maps
by: Bakopoulou, Evita, et al.
Published: (2021)
by: Bakopoulou, Evita, et al.
Published: (2021)
Provably Cost-Sensitive Adversarial Defense via Randomized Smoothing
by: Xin, Yuan, et al.
Published: (2023)
by: Xin, Yuan, et al.
Published: (2023)
On the Abuse and Detection of Polyglot Files
by: Koch, Luke, et al.
Published: (2024)
by: Koch, Luke, et al.
Published: (2024)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
by: Yang, Ziqing, et al.
Published: (2024)
by: Yang, Ziqing, et al.
Published: (2024)
Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
IstGPT: LLM-based Anomaly Detection for Spatial-Temporal Graph in Industrial Systems
by: Zhang, Yuchen, et al.
Published: (2026)
by: Zhang, Yuchen, et al.
Published: (2026)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
by: Chen, Zeyuan, et al.
Published: (2026)
by: Chen, Zeyuan, et al.
Published: (2026)
On the Role of Similarity in Detecting Masquerading Files
by: Oliver, Jonathan, et al.
Published: (2024)
by: Oliver, Jonathan, et al.
Published: (2024)
Understanding Deep Gradient Leakage via Inversion Influence Functions
by: Zhang, Haobo, et al.
Published: (2023)
by: Zhang, Haobo, et al.
Published: (2023)
Link Stealing Attacks Against Inductive Graph Neural Networks
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
by: Akkus, Atilla, et al.
Published: (2024)
by: Akkus, Atilla, et al.
Published: (2024)
Learning to Localize Leakage of Cryptographic Sensitive Variables
by: Gammell, Jimmy, et al.
Published: (2025)
by: Gammell, Jimmy, et al.
Published: (2025)
BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning
by: Wu, Baoyuan, et al.
Published: (2024)
by: Wu, Baoyuan, et al.
Published: (2024)
Analysis of Privacy Leakage in Federated Large Language Models
by: Vu, Minh N., et al.
Published: (2024)
by: Vu, Minh N., et al.
Published: (2024)
Information Leakage from Embedding in Large Language Models
by: Wan, Zhipeng, et al.
Published: (2024)
by: Wan, Zhipeng, et al.
Published: (2024)
Building Gradient Bridges: Label Leakage from Restricted Gradient Sharing in Federated Learning
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
by: Schmotz, David, et al.
Published: (2026)
by: Schmotz, David, et al.
Published: (2026)
Zero-Knowledge Proof-based Verifiable Decentralized Machine Learning in Communication Network: A Comprehensive Survey
by: Xing, Zhibo, et al.
Published: (2023)
by: Xing, Zhibo, et al.
Published: (2023)
Discovering Universal Activation Directions for PII Leakage in Language Models
by: Marchyok, Leo, et al.
Published: (2026)
by: Marchyok, Leo, et al.
Published: (2026)
Information Leakage from Data Updates in Machine Learning Models
by: Hui, Tian, et al.
Published: (2023)
by: Hui, Tian, et al.
Published: (2023)
Similar Items
-
Voice Jailbreak Attacks Against GPT-4o
by: Shen, Xinyue, et al.
Published: (2024) -
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
by: Shen, Xinyue, et al.
Published: (2023) -
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023) -
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024) -
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
by: Chu, Junjie, et al.
Published: (2024)