Spiking the training data to correct for test set contamination
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Johnny Tian-Zheng, Li, Jerry, Godbole, Ameya, Jia, Robin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
von: Godbole, Ameya, et al.
Veröffentlicht: (2025)
von: Godbole, Ameya, et al.
Veröffentlicht: (2025)
Hubble: a Model Suite to Advance the Study of LLM Memorization
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2025)
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2025)
Proving membership in LLM pretraining data via data watermarks
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2024)
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2024)
Interrogating LLM design under a fair learning doctrine
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2025)
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2025)
MuPlon: Multi-Path Causal Optimization for Claim Verification through Controlling Confounding
von: Guo, Hanghui, et al.
Veröffentlicht: (2025)
von: Guo, Hanghui, et al.
Veröffentlicht: (2025)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
von: Cui, Xinyue, et al.
Veröffentlicht: (2025)
von: Cui, Xinyue, et al.
Veröffentlicht: (2025)
Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models
von: Egami, Naoki, et al.
Veröffentlicht: (2023)
von: Egami, Naoki, et al.
Veröffentlicht: (2023)
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation
von: Arias, Esteban Garces, et al.
Veröffentlicht: (2024)
von: Arias, Esteban Garces, et al.
Veröffentlicht: (2024)
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
von: Hu, Zizhao, et al.
Veröffentlicht: (2026)
von: Hu, Zizhao, et al.
Veröffentlicht: (2026)
Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
LLMs as Implicit Imputers: Uncertainty Should Scale with Missing Information
von: van Buuren, Stef
Veröffentlicht: (2026)
von: van Buuren, Stef
Veröffentlicht: (2026)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
LIDS: LLM Summary Inference Under the Layered Lens
von: Park, Dylan, et al.
Veröffentlicht: (2026)
von: Park, Dylan, et al.
Veröffentlicht: (2026)
Learning Dynamic Representations and Policies from Multimodal Clinical Time-Series with Informative Missingness
von: Liang, Zihan, et al.
Veröffentlicht: (2026)
von: Liang, Zihan, et al.
Veröffentlicht: (2026)
Omitted Variable Bias in Language Models Under Distribution Shift
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingness
von: Liang, Zihan, et al.
Veröffentlicht: (2025)
von: Liang, Zihan, et al.
Veröffentlicht: (2025)
Causal Graph Discovery with Retrieval-Augmented Generation based Large Language Models
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2024)
Optimizing Language Models for Human Preferences is a Causal Inference Problem
von: Lin, Victoria, et al.
Veröffentlicht: (2024)
von: Lin, Victoria, et al.
Veröffentlicht: (2024)
How to Evaluate Entity Resolution Systems: An Entity-Centric Framework with Application to Inventor Name Disambiguation
von: Binette, Olivier, et al.
Veröffentlicht: (2024)
von: Binette, Olivier, et al.
Veröffentlicht: (2024)
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
von: Tu, Ruibo, et al.
Veröffentlicht: (2024)
von: Tu, Ruibo, et al.
Veröffentlicht: (2024)
TWIN-GPT: Digital Twins for Clinical Trials via Large Language Model
von: Wang, Yue, et al.
Veröffentlicht: (2024)
von: Wang, Yue, et al.
Veröffentlicht: (2024)
Discovering influential text using convolutional neural networks
von: Ayers, Megan, et al.
Veröffentlicht: (2024)
von: Ayers, Megan, et al.
Veröffentlicht: (2024)
Zipf Distributions from Two-Stage Symbolic Processes: Stability Under Stochastic Lexical Filtering
von: Berman, Vladimir
Veröffentlicht: (2025)
von: Berman, Vladimir
Veröffentlicht: (2025)
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
von: Kern, Christoph, et al.
Veröffentlicht: (2023)
von: Kern, Christoph, et al.
Veröffentlicht: (2023)
End-To-End Causal Effect Estimation from Unstructured Natural Language Data
von: Dhawan, Nikita, et al.
Veröffentlicht: (2024)
von: Dhawan, Nikita, et al.
Veröffentlicht: (2024)
Proximal Causal Inference With Text Data
von: Chen, Jacob M., et al.
Veröffentlicht: (2024)
von: Chen, Jacob M., et al.
Veröffentlicht: (2024)
Causal Inference on Outcomes Learned from Text
von: Modarressi, Iman, et al.
Veröffentlicht: (2025)
von: Modarressi, Iman, et al.
Veröffentlicht: (2025)
A Design-based Solution for Causal Inference with Text: Can a Language Model Be Too Large?
von: Tierney, Graham, et al.
Veröffentlicht: (2025)
von: Tierney, Graham, et al.
Veröffentlicht: (2025)
(Mis)Fitting: A Survey of Scaling Laws
von: Li, Margaret, et al.
Veröffentlicht: (2025)
von: Li, Margaret, et al.
Veröffentlicht: (2025)
Debiasing Watermarks for Large Language Models via Maximal Coupling
von: Xie, Yangxinyu, et al.
Veröffentlicht: (2024)
von: Xie, Yangxinyu, et al.
Veröffentlicht: (2024)
Robust Detection of Watermarks for Large Language Models Under Human Edits
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
von: Chew, Robert, et al.
Veröffentlicht: (2026)
von: Chew, Robert, et al.
Veröffentlicht: (2026)
Text Rationalization for Robust Causal Effect Estimation
von: Zhang, Lijinghua, et al.
Veröffentlicht: (2025)
von: Zhang, Lijinghua, et al.
Veröffentlicht: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
RCT Rejection Sampling for Causal Estimation Evaluation
von: Keith, Katherine A., et al.
Veröffentlicht: (2023)
von: Keith, Katherine A., et al.
Veröffentlicht: (2023)
Evaluating Interventional Reasoning Capabilities of Large Language Models
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
The Leaderboard Illusion
von: Singh, Shivalika, et al.
Veröffentlicht: (2025)
von: Singh, Shivalika, et al.
Veröffentlicht: (2025)
Majority of the Bests: Improving Best-of-N via Bootstrapping
von: Rakhsha, Amin, et al.
Veröffentlicht: (2025)
von: Rakhsha, Amin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
von: Godbole, Ameya, et al.
Veröffentlicht: (2025) -
Hubble: a Model Suite to Advance the Study of LLM Memorization
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2025) -
Proving membership in LLM pretraining data via data watermarks
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2024) -
Interrogating LLM design under a fair learning doctrine
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2025) -
MuPlon: Multi-Path Causal Optimization for Claim Verification through Controlling Confounding
von: Guo, Hanghui, et al.
Veröffentlicht: (2025)