PIIBench: A Unified Multi-Source Benchmark Corpus for Personally Identifiable Information Detection
Fuente:
arXiv
Saved in:
| Main Author: | Jha, Pritesh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-Tuning Over Architectural Complexity: Broad-Coverage PII Detection on PIIBench with DeBERTa
by: Jha, Pritesh
Published: (2026)
by: Jha, Pritesh
Published: (2026)
RaV-IDP: A Reconstruction-as-Validation Framework for Faithful Intelligent Document Processing
by: Jha, Pritesh
Published: (2026)
by: Jha, Pritesh
Published: (2026)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
by: Lu, Zhiyuan, et al.
Published: (2026)
by: Lu, Zhiyuan, et al.
Published: (2026)
Enhancing the De-identification of Personally Identifiable Information in Educational Data
by: Ji, Zilyu, et al.
Published: (2025)
by: Ji, Zilyu, et al.
Published: (2025)
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
by: Zaratiana, Urchade, et al.
Published: (2026)
by: Zaratiana, Urchade, et al.
Published: (2026)
EuskañolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching
by: Heredia, Maite, et al.
Published: (2025)
by: Heredia, Maite, et al.
Published: (2025)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models
by: Singh, Abhinav Kumar, et al.
Published: (2026)
by: Singh, Abhinav Kumar, et al.
Published: (2026)
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings
by: Čechovič, Marko, et al.
Published: (2025)
by: Čechovič, Marko, et al.
Published: (2025)
UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition
by: Hamad, Nagham, et al.
Published: (2025)
by: Hamad, Nagham, et al.
Published: (2025)
GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization
by: Ye, Yangfan, et al.
Published: (2024)
by: Ye, Yangfan, et al.
Published: (2024)
Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
by: Allbert, Rumi, et al.
Published: (2024)
by: Allbert, Rumi, et al.
Published: (2024)
Identifying Multiple Personalities in Large Language Models with External Evaluation
by: Song, Xiaoyang, et al.
Published: (2024)
by: Song, Xiaoyang, et al.
Published: (2024)
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework
by: Luo, Xiaoyu, et al.
Published: (2026)
by: Luo, Xiaoyu, et al.
Published: (2026)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
by: Kim, Joongwon, et al.
Published: (2024)
by: Kim, Joongwon, et al.
Published: (2024)
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
by: Hu, Yulin, et al.
Published: (2026)
by: Hu, Yulin, et al.
Published: (2026)
MCFEND: A Multi-source Benchmark Dataset for Chinese Fake News Detection
by: Li, Yupeng, et al.
Published: (2024)
by: Li, Yupeng, et al.
Published: (2024)
Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark
by: Jiang, Feng, et al.
Published: (2023)
by: Jiang, Feng, et al.
Published: (2023)
Benchmarking and Improving LLM Robustness for Personalized Generation
by: Okite, Chimaobi, et al.
Published: (2025)
by: Okite, Chimaobi, et al.
Published: (2025)
PrefDisco: Benchmarking Proactive Personalized Reasoning
by: Li, Shuyue Stella, et al.
Published: (2025)
by: Li, Shuyue Stella, et al.
Published: (2025)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
by: Huang, Xu, et al.
Published: (2025)
by: Huang, Xu, et al.
Published: (2025)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information
by: Li, Yingya, et al.
Published: (2024)
by: Li, Yingya, et al.
Published: (2024)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
by: Jin, Jiho, et al.
Published: (2026)
by: Jin, Jiho, et al.
Published: (2026)
Tell me what I need to know: Exploring LLM-based (Personalized) Abstractive Multi-Source Meeting Summarization
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
EAG: Extract and Generate Multi-way Aligned Corpus for Complete Multi-lingual Neural Machine Translation
by: Xu, Yulin, et al.
Published: (2022)
by: Xu, Yulin, et al.
Published: (2022)
The Energy of Falsehood: Detecting Hallucinations via Diffusion Model Likelihoods
by: Gautam, Arpit Singh, et al.
Published: (2026)
by: Gautam, Arpit Singh, et al.
Published: (2026)
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
by: Sun, Xiongtao, et al.
Published: (2024)
by: Sun, Xiongtao, et al.
Published: (2024)
Personalized Turn-Level User Conversation Satisfaction Benchmark
by: Wang, Zhefan, et al.
Published: (2026)
by: Wang, Zhefan, et al.
Published: (2026)
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages
by: Isbarov, Jafar, et al.
Published: (2025)
by: Isbarov, Jafar, et al.
Published: (2025)
The PLLuM Instruction Corpus
by: Pęzik, Piotr, et al.
Published: (2025)
by: Pęzik, Piotr, et al.
Published: (2025)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
by: Emery, Deanna, et al.
Published: (2025)
by: Emery, Deanna, et al.
Published: (2025)
Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
LLM-Based Section Identifiers Excel on Open Source but Stumble in Real World Applications
by: Krishnamoorthy, Saranya, et al.
Published: (2024)
by: Krishnamoorthy, Saranya, et al.
Published: (2024)
InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
by: Du, Yaxin, et al.
Published: (2025)
by: Du, Yaxin, et al.
Published: (2025)
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
by: Dong, Yijiang River, et al.
Published: (2025)
by: Dong, Yijiang River, et al.
Published: (2025)
PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems
by: Yu, Jiongchi, et al.
Published: (2026)
by: Yu, Jiongchi, et al.
Published: (2026)
Similar Items
-
Fine-Tuning Over Architectural Complexity: Broad-Coverage PII Detection on PIIBench with DeBERTa
by: Jha, Pritesh
Published: (2026) -
RaV-IDP: A Reconstruction-as-Validation Framework for Faithful Intelligent Document Processing
by: Jha, Pritesh
Published: (2026) -
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
by: Lu, Zhiyuan, et al.
Published: (2026) -
Enhancing the De-identification of Personally Identifiable Information in Educational Data
by: Ji, Zilyu, et al.
Published: (2025) -
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
by: Zaratiana, Urchade, et al.
Published: (2026)