Automated Detection of Pre-training Text in Black-box LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Ruihan, Shang, Yu-Ming, Peng, Jiankun, Luo, Wei, Wang, Yazhe, Zhang, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
TransGPT: Multi-modal Generative Pre-trained Transformer for Transportation
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
STAGE: Simplified Text-Attributed Graph Embeddings Using Pre-trained LLMs
by: Zolnai-Lucas, Aaron, et al.
Published: (2024)
by: Zolnai-Lucas, Aaron, et al.
Published: (2024)
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
by: Wang, Yubo, et al.
Published: (2023)
by: Wang, Yubo, et al.
Published: (2023)
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025)
by: Ren, Yanwei, et al.
Published: (2025)
Pre-trained Large Language Models for Financial Sentiment Analysis
by: Luo, Wei, et al.
Published: (2024)
by: Luo, Wei, et al.
Published: (2024)
Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
by: He, Yutong, et al.
Published: (2024)
by: He, Yutong, et al.
Published: (2024)
Black-box Prompt Tuning with Subspace Learning
by: Zheng, Yuanhang, et al.
Published: (2023)
by: Zheng, Yuanhang, et al.
Published: (2023)
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
SAMGPT: Text-free Graph Foundation Model for Multi-domain Pre-training and Cross-domain Adaptation
by: Yu, Xingtong, et al.
Published: (2025)
by: Yu, Xingtong, et al.
Published: (2025)
Boosting Explainability through Selective Rationalization in Pre-trained Language Models
by: Yuan, Libing, et al.
Published: (2025)
by: Yuan, Libing, et al.
Published: (2025)
Pre-training LLMs using human-like development data corpus
by: Bhardwaj, Khushi, et al.
Published: (2023)
by: Bhardwaj, Khushi, et al.
Published: (2023)
Probing Language Models for Pre-training Data Detection
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
Scaling LLM Pre-training with Vocabulary Curriculum
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
Atomic Calibration of LLMs in Long-Form Generations
by: Zhang, Caiqi, et al.
Published: (2024)
by: Zhang, Caiqi, et al.
Published: (2024)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
by: Bhattacharjee, Amrita, et al.
Published: (2023)
by: Bhattacharjee, Amrita, et al.
Published: (2023)
DocMamba: Efficient Document Pre-training with State Space Model
by: Hu, Pengfei, et al.
Published: (2024)
by: Hu, Pengfei, et al.
Published: (2024)
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
by: Taghian, Mehran, et al.
Published: (2026)
by: Taghian, Mehran, et al.
Published: (2026)
Relational Prompt-based Pre-trained Language Models for Social Event Detection
by: Li, Pu, et al.
Published: (2024)
by: Li, Pu, et al.
Published: (2024)
DPIC: Decoupling Prompt and Intrinsic Characteristics for LLM Generated Text Detection
by: Yu, Xiao, et al.
Published: (2023)
by: Yu, Xiao, et al.
Published: (2023)
CDGP: Automatic Cloze Distractor Generation based on Pre-trained Language Model
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
by: Peng, Jiahui, et al.
Published: (2025)
by: Peng, Jiahui, et al.
Published: (2025)
Pre-training Distillation for Large Language Models: A Design Space Exploration
by: Peng, Hao, et al.
Published: (2024)
by: Peng, Hao, et al.
Published: (2024)
LangCell: Language-Cell Pre-training for Cell Identity Understanding
by: Zhao, Suyuan, et al.
Published: (2024)
by: Zhao, Suyuan, et al.
Published: (2024)
DataMan: Data Manager for Pre-training Large Language Models
by: Peng, Ru, et al.
Published: (2025)
by: Peng, Ru, et al.
Published: (2025)
Synergistic Anchored Contrastive Pre-training for Few-Shot Relation Extraction
by: Luo, Da, et al.
Published: (2023)
by: Luo, Da, et al.
Published: (2023)
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
by: Luo, Kangcheng, et al.
Published: (2025)
by: Luo, Kangcheng, et al.
Published: (2025)
PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
by: Luo, Ruilin, et al.
Published: (2024)
by: Luo, Ruilin, et al.
Published: (2024)
DataVisT5: A Pre-trained Language Model for Jointly Understanding Text and Data Visualization
by: Wan, Zhuoyue, et al.
Published: (2024)
by: Wan, Zhuoyue, et al.
Published: (2024)
ExpNote: Black-box Large Language Models are Better Task Solvers with Experience Notebook
by: Sun, Wangtao, et al.
Published: (2023)
by: Sun, Wangtao, et al.
Published: (2023)
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
by: Witold, Waligóra
Published: (2024)
by: Witold, Waligóra
Published: (2024)
Multi-domain Knowledge Graph Collaborative Pre-training and Prompt Tuning for Diverse Downstream Tasks
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Epistemological Bias As a Means for the Automated Detection of Injustices in Text
by: Andrews, Kenya, et al.
Published: (2024)
by: Andrews, Kenya, et al.
Published: (2024)
Detection Method for Prompt Injection by Integrating Pre-trained Model and Heuristic Feature Engineering
by: Ji, Yi, et al.
Published: (2025)
by: Ji, Yi, et al.
Published: (2025)
Blacks is to Anger as Whites is to Joy? Understanding Latent Affective Bias in Large Pre-trained Neural Language Models
by: Kadan, Anoop, et al.
Published: (2023)
by: Kadan, Anoop, et al.
Published: (2023)
Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning
by: Zhou, Xiaotian, et al.
Published: (2026)
by: Zhou, Xiaotian, et al.
Published: (2026)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
by: Agrawal, Tanmay
Published: (2025)
by: Agrawal, Tanmay
Published: (2025)
Similar Items
-
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025) -
TransGPT: Multi-modal Generative Pre-trained Transformer for Transportation
by: Wang, Peng, et al.
Published: (2024) -
STAGE: Simplified Text-Attributed Graph Embeddings Using Pre-trained LLMs
by: Zolnai-Lucas, Aaron, et al.
Published: (2024) -
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
by: Wang, Yubo, et al.
Published: (2023) -
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025)