CAPID: Context-Aware PII Detection for Question-Answering Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ponomarenko, Mariia, Abedini, Sepideh, Shafieinejad, Masoumeh, Emerson, D. B., Mohapatra, Shubhankar, He, Xi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908826565869568
author Ponomarenko, Mariia
Abedini, Sepideh
Shafieinejad, Masoumeh
Emerson, D. B.
Mohapatra, Shubhankar
He, Xi
author_facet Ponomarenko, Mariia
Abedini, Sepideh
Shafieinejad, Masoumeh
Emerson, D. B.
Mohapatra, Shubhankar
He, Xi
contents Detecting personally identifiable information (PII) in user queries is critical for ensuring privacy in question-answering systems. Current approaches mainly redact all PII, disregarding the fact that some of them may be contextually relevant to the user's question, resulting in a degradation of response quality. Large language models (LLMs) might be able to help determine which PII are relevant, but due to their closed source nature and lack of privacy guarantees, they are unsuitable for sensitive data processing. To achieve privacy-preserving PII detection, we propose CAPID, a practical approach that fine-tunes a locally owned small language model (SLM) that filters sensitive information before it is passed to LLMs for QA. However, existing datasets do not capture the context-dependent relevance of PII needed to train such a model effectively. To fill this gap, we propose a synthetic data generation pipeline that leverages LLMs to produce a diverse, domain-rich dataset spanning multiple PII types and relevance levels. Using this dataset, we fine-tune an SLM to detect PII spans, classify their types, and estimate contextual relevance. Our experiments show that relevance-aware PII detection with a fine-tuned SLM substantially outperforms existing baselines in span, relevance and type accuracy while preserving significantly higher downstream utility under anonymization.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10074
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CAPID: Context-Aware PII Detection for Question-Answering Systems
Ponomarenko, Mariia
Abedini, Sepideh
Shafieinejad, Masoumeh
Emerson, D. B.
Mohapatra, Shubhankar
He, Xi
Cryptography and Security
Computation and Language
Detecting personally identifiable information (PII) in user queries is critical for ensuring privacy in question-answering systems. Current approaches mainly redact all PII, disregarding the fact that some of them may be contextually relevant to the user's question, resulting in a degradation of response quality. Large language models (LLMs) might be able to help determine which PII are relevant, but due to their closed source nature and lack of privacy guarantees, they are unsuitable for sensitive data processing. To achieve privacy-preserving PII detection, we propose CAPID, a practical approach that fine-tunes a locally owned small language model (SLM) that filters sensitive information before it is passed to LLMs for QA. However, existing datasets do not capture the context-dependent relevance of PII needed to train such a model effectively. To fill this gap, we propose a synthetic data generation pipeline that leverages LLMs to produce a diverse, domain-rich dataset spanning multiple PII types and relevance levels. Using this dataset, we fine-tune an SLM to detect PII spans, classify their types, and estimate contextual relevance. Our experiments show that relevance-aware PII detection with a fine-tuned SLM substantially outperforms existing baselines in span, relevance and type accuracy while preserving significantly higher downstream utility under anonymization.
title CAPID: Context-Aware PII Detection for Question-Answering Systems
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2602.10074