All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zeyu, Chen, Ryan, Stadie, Bradly C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting
von: Zhang, Zeyu, et al.
Veröffentlicht: (2026)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2026)
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
von: Zhang, Lunjun, et al.
Veröffentlicht: (2026)
von: Zhang, Lunjun, et al.
Veröffentlicht: (2026)
The Gradient of Algebraic Model Counting
von: Maene, Jaron, et al.
Veröffentlicht: (2025)
von: Maene, Jaron, et al.
Veröffentlicht: (2025)
CountsDiff: A Diffusion Model on the Natural Numbers for Generation and Imputation of Count-Based Data
von: Soatto, Renzo G., et al.
Veröffentlicht: (2026)
von: Soatto, Renzo G., et al.
Veröffentlicht: (2026)
TACNET: Temporal Audio Source Counting Network
von: Ahmadnejad, Amirreza, et al.
Veröffentlicht: (2023)
von: Ahmadnejad, Amirreza, et al.
Veröffentlicht: (2023)
RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation
von: Kattamuri, Ashish, et al.
Veröffentlicht: (2025)
von: Kattamuri, Ashish, et al.
Veröffentlicht: (2025)
Counting Still Counts: Understanding Neural Complex Query Answering Through Query Relaxation
von: Brunink, Yannick, et al.
Veröffentlicht: (2025)
von: Brunink, Yannick, et al.
Veröffentlicht: (2025)
Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity
von: Li, Bojie
Veröffentlicht: (2026)
von: Li, Bojie
Veröffentlicht: (2026)
State Contamination in Memory-Augmented LLM Agents
von: Wang, Yian, et al.
Veröffentlicht: (2026)
von: Wang, Yian, et al.
Veröffentlicht: (2026)
Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models
von: Li, Weixian Waylon, et al.
Veröffentlicht: (2026)
von: Li, Weixian Waylon, et al.
Veröffentlicht: (2026)
Counting Hours, Counting Losses: The Toll of Unpredictable Work Schedules on Financial Security
von: Nokhiz, Pegah, et al.
Veröffentlicht: (2025)
von: Nokhiz, Pegah, et al.
Veröffentlicht: (2025)
Wonderful Team: Zero-Shot Physical Task Planning with Visual LLMs
von: Wang, Zidan, et al.
Veröffentlicht: (2024)
von: Wang, Zidan, et al.
Veröffentlicht: (2024)
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
von: Jia, Furong, et al.
Veröffentlicht: (2025)
von: Jia, Furong, et al.
Veröffentlicht: (2025)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
von: Mestha, Harshvardhan, et al.
Veröffentlicht: (2024)
von: Mestha, Harshvardhan, et al.
Veröffentlicht: (2024)
Do Attention Heads Compete or Cooperate during Counting?
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025)
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025)
Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor Decomposition
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
Explainable Fuzzy GNNs for Leak Detection in Water Distribution Networks
von: Khaled, Qusai, et al.
Veröffentlicht: (2026)
von: Khaled, Qusai, et al.
Veröffentlicht: (2026)
Online Detection of Anomalies in Temporal Knowledge Graphs with Interpretability
von: Zhang, Jiasheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiasheng, et al.
Veröffentlicht: (2024)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
von: Dwibedi, Debidatta, et al.
Veröffentlicht: (2024)
von: Dwibedi, Debidatta, et al.
Veröffentlicht: (2024)
VDSC: Enhancing Exploration Timing with Value Discrepancy and State Counts
von: Captari, Marius, et al.
Veröffentlicht: (2024)
von: Captari, Marius, et al.
Veröffentlicht: (2024)
Grid-Mapping Pseudo-Count Constraint for Offline Reinforcement Learning
von: Shen, Yi, et al.
Veröffentlicht: (2024)
von: Shen, Yi, et al.
Veröffentlicht: (2024)
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts
von: Chung, Youngseog, et al.
Veröffentlicht: (2024)
von: Chung, Youngseog, et al.
Veröffentlicht: (2024)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
von: Wang, Xinglin, et al.
Veröffentlicht: (2025)
von: Wang, Xinglin, et al.
Veröffentlicht: (2025)
When Can Transformers Count to n?
von: Yehudai, Gilad, et al.
Veröffentlicht: (2024)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2024)
When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation
von: Ward, Joshua, et al.
Veröffentlicht: (2025)
von: Ward, Joshua, et al.
Veröffentlicht: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
Contextual Counting: A Mechanistic Study of Transformers on a Quantitative Task
von: Golkar, Siavash, et al.
Veröffentlicht: (2024)
von: Golkar, Siavash, et al.
Veröffentlicht: (2024)
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
von: Dwibedi, Debidatta, et al.
Veröffentlicht: (2024)
von: Dwibedi, Debidatta, et al.
Veröffentlicht: (2024)
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
von: Ling Team, et al.
Veröffentlicht: (2025)
von: Ling Team, et al.
Veröffentlicht: (2025)
Anomaly Detection with Adaptive and Aggressive Rejection for Contaminated Training Data
von: Lee, Jungi, et al.
Veröffentlicht: (2025)
von: Lee, Jungi, et al.
Veröffentlicht: (2025)
Structured Temporal Causality for Interpretable Multivariate Time Series Anomaly Detection
von: Cho, Dongchan, et al.
Veröffentlicht: (2025)
von: Cho, Dongchan, et al.
Veröffentlicht: (2025)
Every Character Counts: From Vulnerability to Defense in Phishing Detection
von: Chiper, Maria, et al.
Veröffentlicht: (2025)
von: Chiper, Maria, et al.
Veröffentlicht: (2025)
All AI Models are Wrong, but Some are Optimal
von: Anand, Akhil S, et al.
Veröffentlicht: (2025)
von: Anand, Akhil S, et al.
Veröffentlicht: (2025)
Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS
von: DeLorenzo, Matthew, et al.
Veröffentlicht: (2024)
von: DeLorenzo, Matthew, et al.
Veröffentlicht: (2024)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
More Agents Is All You Need
von: Li, Junyou, et al.
Veröffentlicht: (2024)
von: Li, Junyou, et al.
Veröffentlicht: (2024)
D2 Actor Critic: Diffusion Actor Meets Distributional Critic
von: Zhang, Lunjun, et al.
Veröffentlicht: (2025)
von: Zhang, Lunjun, et al.
Veröffentlicht: (2025)
Parallel Sampling via Counting
von: Anari, Nima, et al.
Veröffentlicht: (2024)
von: Anari, Nima, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting
von: Zhang, Zeyu, et al.
Veröffentlicht: (2026) -
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
von: Zhang, Lunjun, et al.
Veröffentlicht: (2026) -
The Gradient of Algebraic Model Counting
von: Maene, Jaron, et al.
Veröffentlicht: (2025) -
CountsDiff: A Diffusion Model on the Natural Numbers for Generation and Imputation of Count-Based Data
von: Soatto, Renzo G., et al.
Veröffentlicht: (2026) -
TACNET: Temporal Audio Source Counting Network
von: Ahmadnejad, Amirreza, et al.
Veröffentlicht: (2023)