On The Fragility of Benchmark Contamination Detection in Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Han, Li, Haoyu, Ko, Brian, Zhang, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Benchmark Datasets Should Be Contamination-Resistant
by: Al-Lawati, Ali, et al.
Published: (2026)
by: Al-Lawati, Ali, et al.
Published: (2026)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
OCGEC: One-class Graph Embedding Classification for DNN Backdoor Detection
by: Jiang, Haoyu, et al.
Published: (2023)
by: Jiang, Haoyu, et al.
Published: (2023)
Beyond Detection: A Comprehensive Benchmark and Study on Representation Learning for Fine-Grained Webshell Family Classification
by: Han, Feijiang
Published: (2025)
by: Han, Feijiang
Published: (2025)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Conditional Adversarial Fragility in Financial Machine Learning under Macroeconomic Stress
by: Baviskar, Samruddhi
Published: (2025)
by: Baviskar, Samruddhi
Published: (2025)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
by: Hung, Kuo-Han, et al.
Published: (2024)
by: Hung, Kuo-Han, et al.
Published: (2024)
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
Magnitude-based Neuron Pruning for Backdoor Defens
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
A Customer Level Fraudulent Activity Detection Benchmark for Enhancing Machine Learning Model Research and Evaluation
by: Jing, Phoebe, et al.
Published: (2024)
by: Jing, Phoebe, et al.
Published: (2024)
Large Language Models in Wireless Application Design: In-Context Learning-enhanced Automatic Network Intrusion Detection
by: Zhang, Han, et al.
Published: (2024)
by: Zhang, Han, et al.
Published: (2024)
A Semantic and Clean-label Backdoor Attack against Graph Convolutional Networks
by: Dai, Jiazhu, et al.
Published: (2025)
by: Dai, Jiazhu, et al.
Published: (2025)
Effective backdoor attack on graph neural networks in link prediction tasks
by: Dai, Jiazhu, et al.
Published: (2024)
by: Dai, Jiazhu, et al.
Published: (2024)
KnowGraph: Knowledge-Enabled Anomaly Detection via Logical Reasoning on Graph Data
by: Zhou, Andy, et al.
Published: (2024)
by: Zhou, Andy, et al.
Published: (2024)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
by: Poppi, Samuele, et al.
Published: (2024)
by: Poppi, Samuele, et al.
Published: (2024)
Detecting Benchmark Contamination Through Watermarking
by: Sander, Tom, et al.
Published: (2025)
by: Sander, Tom, et al.
Published: (2025)
Membership Inference Attack with Partial Features
by: Wang, Xurun, et al.
Published: (2025)
by: Wang, Xurun, et al.
Published: (2025)
Exploring Robust Intrusion Detection: A Benchmark Study of Feature Transferability in IoT Botnet Attack Detection
by: Guerra-Manzanares, Alejandro, et al.
Published: (2026)
by: Guerra-Manzanares, Alejandro, et al.
Published: (2026)
Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models
by: Wang, Jeffrey G., et al.
Published: (2024)
by: Wang, Jeffrey G., et al.
Published: (2024)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
by: Li, Fengpeng, et al.
Published: (2026)
by: Li, Fengpeng, et al.
Published: (2026)
How Does a Deep Learning Model Architecture Impact Its Privacy? A Comprehensive Study of Privacy Attacks on CNNs and Transformers
by: Zhang, Guangsheng, et al.
Published: (2022)
by: Zhang, Guangsheng, et al.
Published: (2022)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
A Generative Approach to Surrogate-based Black-box Attacks
by: Moraffah, Raha, et al.
Published: (2024)
by: Moraffah, Raha, et al.
Published: (2024)
Accelerating IoV Intrusion Detection: Benchmarking GPU-Accelerated vs CPU-Based ML Libraries
by: Çolhak, Furkan, et al.
Published: (2025)
by: Çolhak, Furkan, et al.
Published: (2025)
Explainable AI for Comparative Analysis of Intrusion Detection Models
by: Corea, Pap M., et al.
Published: (2024)
by: Corea, Pap M., et al.
Published: (2024)
Data Valuation and Detections in Federated Learning
by: Li, Wenqian, et al.
Published: (2023)
by: Li, Wenqian, et al.
Published: (2023)
VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
by: Yu, Simon, et al.
Published: (2025)
by: Yu, Simon, et al.
Published: (2025)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023)
by: Golchin, Shahriar, et al.
Published: (2023)
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models
by: Krishna, Arjun, et al.
Published: (2025)
by: Krishna, Arjun, et al.
Published: (2025)
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
by: Guo, Yihao, et al.
Published: (2025)
by: Guo, Yihao, et al.
Published: (2025)
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
by: Wang, Ziyao, et al.
Published: (2025)
by: Wang, Ziyao, et al.
Published: (2025)
LAN: Learning Adaptive Neighbors for Real-Time Insider Threat Detection
by: Cai, Xiangrui, et al.
Published: (2024)
by: Cai, Xiangrui, et al.
Published: (2024)
PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
Finetuning Large Language Models for Vulnerability Detection
by: Shestov, Alexey, et al.
Published: (2024)
by: Shestov, Alexey, et al.
Published: (2024)
Similar Items
-
LLM Benchmark Datasets Should Be Contamination-Resistant
by: Al-Lawati, Ali, et al.
Published: (2026) -
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026) -
OCGEC: One-class Graph Embedding Classification for DNN Backdoor Detection
by: Jiang, Haoyu, et al.
Published: (2023) -
Beyond Detection: A Comprehensive Benchmark and Study on Representation Learning for Fine-Grained Webshell Family Classification
by: Han, Feijiang
Published: (2025) -
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)