Detecting Non-Membership in LLM Training Data via Rank Correlations
Fuente:
arXiv
Saved in:
| Main Authors: | Shetty, Pranav, Haque, Mirazul, Ma, Zhiqiang, Liu, Xiaomo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
by: Shetty, Pranav, et al.
Published: (2025)
by: Shetty, Pranav, et al.
Published: (2025)
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
by: Sibue, Mathieu, et al.
Published: (2026)
by: Sibue, Mathieu, et al.
Published: (2026)
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
by: S, Santosh T. Y. S., et al.
Published: (2025)
by: S, Santosh T. Y. S., et al.
Published: (2025)
DocLLM: A layout-aware generative language model for multimodal document understanding
by: Wang, Dongsheng, et al.
Published: (2023)
by: Wang, Dongsheng, et al.
Published: (2023)
Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation
by: Li, Xianzhi, et al.
Published: (2024)
by: Li, Xianzhi, et al.
Published: (2024)
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
by: German, Eyal, et al.
Published: (2025)
by: German, Eyal, et al.
Published: (2025)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
by: Zheng, Jiangrui, et al.
Published: (2023)
by: Zheng, Jiangrui, et al.
Published: (2023)
ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph Reconstruction
by: Yu, Yan, et al.
Published: (2025)
by: Yu, Yan, et al.
Published: (2025)
SlimPajama-DC: Understanding Data Combinations for LLM Training
by: Shen, Zhiqiang, et al.
Published: (2023)
by: Shen, Zhiqiang, et al.
Published: (2023)
Who Wrote the Book? Detecting and Attributing LLM Ghostwriters
by: Shetty, Anudeex, et al.
Published: (2026)
by: Shetty, Anudeex, et al.
Published: (2026)
Sequences of Logits Reveal the Low Rank Structure of Language Models
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
R.R.: Unveiling LLM Training Privacy through Recollection and Ranking
by: Meng, Wenlong, et al.
Published: (2025)
by: Meng, Wenlong, et al.
Published: (2025)
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
by: Yang, Shidong, et al.
Published: (2026)
by: Yang, Shidong, et al.
Published: (2026)
LLM-Assisted Question-Answering on Technical Documents Using Structured Data-Aware Retrieval Augmented Generation
by: Sobhan, Shadman, et al.
Published: (2025)
by: Sobhan, Shadman, et al.
Published: (2025)
Human Texts Are Outliers: Detecting LLM-generated Texts via Out-of-distribution Detection
by: Zeng, Cong, et al.
Published: (2025)
by: Zeng, Cong, et al.
Published: (2025)
Translating between SQL Dialects for Cloud Migration
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
Accelerating materials discovery for polymer solar cells: Data-driven insights enabled by natural language processing
by: Shetty, Pranav, et al.
Published: (2024)
by: Shetty, Pranav, et al.
Published: (2024)
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
by: Tran, Toan, et al.
Published: (2025)
by: Tran, Toan, et al.
Published: (2025)
CodeMirage: Hallucinations in Code Generated by Large Language Models
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation
by: Ma, Liqun, et al.
Published: (2024)
by: Ma, Liqun, et al.
Published: (2024)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
by: Chen, Zixun, et al.
Published: (2025)
by: Chen, Zixun, et al.
Published: (2025)
SpecDetect: Simple, Fast, and Training-Free Detection of LLM-Generated Text via Spectral Analysis
by: Luo, Haitong, et al.
Published: (2025)
by: Luo, Haitong, et al.
Published: (2025)
Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection
by: Li, Mingzhe, et al.
Published: (2026)
by: Li, Mingzhe, et al.
Published: (2026)
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
by: Haque, Mirazul, et al.
Published: (2026)
by: Haque, Mirazul, et al.
Published: (2026)
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Efficient Code LLM Training via Distribution-Consistent and Diversity-Aware Data Selection
by: Lyu, Weijie, et al.
Published: (2025)
by: Lyu, Weijie, et al.
Published: (2025)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
by: Shetty, Anudeex, et al.
Published: (2026)
by: Shetty, Anudeex, et al.
Published: (2026)
Transferable and Efficient Non-Factual Content Detection via Probe Training with Offline Consistency Checking
by: Zhang, Xiaokang, et al.
Published: (2024)
by: Zhang, Xiaokang, et al.
Published: (2024)
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
by: Li, Wenkai, et al.
Published: (2024)
by: Li, Wenkai, et al.
Published: (2024)
Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
Timber: Training-free Instruct Model Refining with Base via Effective Rank
by: Wu, Taiqiang, et al.
Published: (2025)
by: Wu, Taiqiang, et al.
Published: (2025)
Curriculum-style Data Augmentation for LLM-based Metaphor Detection
by: Jia, Kaidi, et al.
Published: (2024)
by: Jia, Kaidi, et al.
Published: (2024)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
by: Liu, Yule, et al.
Published: (2025)
by: Liu, Yule, et al.
Published: (2025)
Detecting RLVR Training Data via Structural Convergence of Reasoning
by: Zhang, Hongbo, et al.
Published: (2026)
by: Zhang, Hongbo, et al.
Published: (2026)
Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
by: Fan, Qihui, et al.
Published: (2026)
by: Fan, Qihui, et al.
Published: (2026)
Where is this coming from? Making groundedness count in the evaluation of Document VQA models
by: Nourbakhsh, Armineh, et al.
Published: (2025)
by: Nourbakhsh, Armineh, et al.
Published: (2025)
Similar Items
-
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
by: Shetty, Pranav, et al.
Published: (2025) -
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
by: Zmigrod, Ran, et al.
Published: (2024) -
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
by: Sibue, Mathieu, et al.
Published: (2026) -
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
by: S, Santosh T. Y. S., et al.
Published: (2025) -
DocLLM: A layout-aware generative language model for multimodal document understanding
by: Wang, Dongsheng, et al.
Published: (2023)