PANORAMA: A synthetic PII-laced dataset for studying sensitive data memorization in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Selvam, Sriram, Ghosh, Anneswa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
Scalable multilingual PII annotation for responsible AI in LLMs
von: Meena, Bharti, et al.
Veröffentlicht: (2025)
von: Meena, Bharti, et al.
Veröffentlicht: (2025)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework
von: Luo, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Luo, Xiaoyu, et al.
Veröffentlicht: (2026)
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
von: Zaratiana, Urchade, et al.
Veröffentlicht: (2026)
von: Zaratiana, Urchade, et al.
Veröffentlicht: (2026)
Disentangling generalization and memorization in large language models using chess
von: Pleiss, Leonard S., et al.
Veröffentlicht: (2026)
von: Pleiss, Leonard S., et al.
Veröffentlicht: (2026)
PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization
von: Liu, Mingshuo, et al.
Veröffentlicht: (2026)
von: Liu, Mingshuo, et al.
Veröffentlicht: (2026)
Fine-Tuning Over Architectural Complexity: Broad-Coverage PII Detection on PIIBench with DeBERTa
von: Jha, Pritesh
Veröffentlicht: (2026)
von: Jha, Pritesh
Veröffentlicht: (2026)
Locale-Conditioned Few-Shot Prompting Mitigates Demonstration Regurgitation in On-Device PII Substitution with Small Language Models
von: Sadani, Anuj, et al.
Veröffentlicht: (2026)
von: Sadani, Anuj, et al.
Veröffentlicht: (2026)
Deep sequence models tend to memorize geometrically; it is unclear why
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025)
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025)
FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
von: Shen, Hao, et al.
Veröffentlicht: (2025)
von: Shen, Hao, et al.
Veröffentlicht: (2025)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2023)
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2023)
Mapping Clinical Doubt: Locating Linguistic Uncertainty in LLMs
von: Sridhar, Srivarshinee, et al.
Veröffentlicht: (2025)
von: Sridhar, Srivarshinee, et al.
Veröffentlicht: (2025)
PromptMind Team at EHRSQL-2024: Improving Reliability of SQL Generation using Ensemble LLMs
von: Gundabathula, Satya K, et al.
Veröffentlicht: (2024)
von: Gundabathula, Satya K, et al.
Veröffentlicht: (2024)
LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
von: Zolnour, Ali, et al.
Veröffentlicht: (2025)
von: Zolnour, Ali, et al.
Veröffentlicht: (2025)
Code-Mixer Ya Nahi: Novel Approaches to Measuring Multilingual LLMs' Code-Mixing Capabilities
von: Gupta, Ayushman, et al.
Veröffentlicht: (2024)
von: Gupta, Ayushman, et al.
Veröffentlicht: (2024)
Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows
von: Sharan, Aditya, et al.
Veröffentlicht: (2026)
von: Sharan, Aditya, et al.
Veröffentlicht: (2026)
Can LLMs Augment Low-Resource Reading Comprehension Datasets? Opportunities and Challenges
von: Samuel, Vinay, et al.
Veröffentlicht: (2023)
von: Samuel, Vinay, et al.
Veröffentlicht: (2023)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
Mined Prompting and Metadata-Guided Generation for Wound Care Visual Question Answering
von: Durgapraveen, Bavana, et al.
Veröffentlicht: (2025)
von: Durgapraveen, Bavana, et al.
Veröffentlicht: (2025)
Entity-Augmented Neuroscience Knowledge Retrieval Using Ontology and Semantic Understanding Capability of LLM
von: Ta, Pralaypati, et al.
Veröffentlicht: (2025)
von: Ta, Pralaypati, et al.
Veröffentlicht: (2025)
Large Language Models (LLMs) for Source Code Analysis: applications, models and datasets
von: Jelodar, Hamed, et al.
Veröffentlicht: (2025)
von: Jelodar, Hamed, et al.
Veröffentlicht: (2025)
Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis
von: Yadav, Anushka, et al.
Veröffentlicht: (2025)
von: Yadav, Anushka, et al.
Veröffentlicht: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
Tagarela - A Portuguese speech dataset from podcasts
von: de Oliveira, Frederico Santos, et al.
Veröffentlicht: (2026)
von: de Oliveira, Frederico Santos, et al.
Veröffentlicht: (2026)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
von: Mohseni, Seyedreza, et al.
Veröffentlicht: (2024)
von: Mohseni, Seyedreza, et al.
Veröffentlicht: (2024)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Applicability of Large Language Models and Generative Models for Legal Case Judgement Summarization
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
Experiments with truth using Machine Learning: Spectral analysis and explainable classification of synthetic, false, and genuine information
von: Pendyala, Vishnu S., et al.
Veröffentlicht: (2024)
von: Pendyala, Vishnu S., et al.
Veröffentlicht: (2024)
Pre-training LLMs using human-like development data corpus
von: Bhardwaj, Khushi, et al.
Veröffentlicht: (2023)
von: Bhardwaj, Khushi, et al.
Veröffentlicht: (2023)
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
von: Juneja, Gurusha, et al.
Veröffentlicht: (2025)
von: Juneja, Gurusha, et al.
Veröffentlicht: (2025)
VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation
von: Pan, Jingheng, et al.
Veröffentlicht: (2026)
von: Pan, Jingheng, et al.
Veröffentlicht: (2026)
Plancraft: an evaluation dataset for planning with LLM agents
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility
von: Shahariar, G M, et al.
Veröffentlicht: (2026)
von: Shahariar, G M, et al.
Veröffentlicht: (2026)
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
von: Burns, Thomas F, et al.
Veröffentlicht: (2025)
von: Burns, Thomas F, et al.
Veröffentlicht: (2025)
Hey, wait a minute: on at-issue sensitivity in Language Models
von: Kim, Sanghee J., et al.
Veröffentlicht: (2025)
von: Kim, Sanghee J., et al.
Veröffentlicht: (2025)
CrisiText: A dataset of warning messages for LLM training in emergency communication
von: Gonella, Giacomo, et al.
Veröffentlicht: (2025)
von: Gonella, Giacomo, et al.
Veröffentlicht: (2025)
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
von: Dong, Mingwen, et al.
Veröffentlicht: (2024)
von: Dong, Mingwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024) -
Scalable multilingual PII annotation for responsible AI in LLMs
von: Meena, Bharti, et al.
Veröffentlicht: (2025) -
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024) -
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework
von: Luo, Xiaoyu, et al.
Veröffentlicht: (2026) -
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
von: Zaratiana, Urchade, et al.
Veröffentlicht: (2026)