DyePack: Provably Flagging Test Set Contamination in LLMs Using Backdoors
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Yize, Wang, Wenxiao, Moayeri, Mazda, Feizi, Soheil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
by: Hosseini, Parsa, et al.
Published: (2025)
by: Hosseini, Parsa, et al.
Published: (2025)
PRIME: Prioritizing Interpretability in Failure Mode Extraction
by: Rezaei, Keivan, et al.
Published: (2023)
by: Rezaei, Keivan, et al.
Published: (2023)
Tool Preferences in Agentic LLMs are Unreliable
by: Faghih, Kazem, et al.
Published: (2025)
by: Faghih, Kazem, et al.
Published: (2025)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
by: Wang, Wenxiao, et al.
Published: (2025)
by: Wang, Wenxiao, et al.
Published: (2025)
Maestro: Joint Graph & Config Optimization for Reliable AI Agents
by: Wang, Wenxiao, et al.
Published: (2025)
by: Wang, Wenxiao, et al.
Published: (2025)
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
by: Chegini, Atoosa, et al.
Published: (2026)
by: Chegini, Atoosa, et al.
Published: (2026)
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024)
by: Zarei, Arman, et al.
Published: (2024)
Rethinking Artistic Copyright Infringements in the Era of Text-to-Image Generative Models
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Can AI-Generated Text be Reliably Detected?
by: Sadasivan, Vinu Sankar, et al.
Published: (2023)
by: Sadasivan, Vinu Sankar, et al.
Published: (2023)
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
by: Saha, Shoumik, et al.
Published: (2025)
by: Saha, Shoumik, et al.
Published: (2025)
SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents
by: Saberi, Mehrdad, et al.
Published: (2026)
by: Saberi, Mehrdad, et al.
Published: (2026)
Endor: Hardware-Friendly Sparse Format for Offloaded LLM Inference
by: Joo, Donghyeon, et al.
Published: (2024)
by: Joo, Donghyeon, et al.
Published: (2024)
Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
by: Pournemat, Mobina, et al.
Published: (2025)
by: Pournemat, Mobina, et al.
Published: (2025)
Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
by: Schaeffer, Rylan, et al.
Published: (2026)
by: Schaeffer, Rylan, et al.
Published: (2026)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025)
by: Balasubramanian, Sriram, et al.
Published: (2025)
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
by: Tseng, Yuan, et al.
Published: (2025)
by: Tseng, Yuan, et al.
Published: (2025)
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
DCR: Quantifying Data Contamination in LLMs Evaluation
by: Xu, Cheng, et al.
Published: (2025)
by: Xu, Cheng, et al.
Published: (2025)
RESTOR: Knowledge Recovery in Machine Unlearning
by: Rezaei, Keivan, et al.
Published: (2024)
by: Rezaei, Keivan, et al.
Published: (2024)
Early Stopping for Large Reasoning Models via Confidence Dynamics
by: Hosseini, Parsa, et al.
Published: (2026)
by: Hosseini, Parsa, et al.
Published: (2026)
On Mechanistic Circuits for Extractive Question-Answering
by: Basu, Samyadeep, et al.
Published: (2025)
by: Basu, Samyadeep, et al.
Published: (2025)
Benchmarking ChatGPT and DeepSeek in April 2025: A Novel Dual Perspective Sentiment Analysis Using Lexicon-Based and Deep Learning Approaches
by: Alhusseini, Maryam Mahdi, et al.
Published: (2025)
by: Alhusseini, Maryam Mahdi, et al.
Published: (2025)
Pruning Strategies for Backdoor Defense in LLMs
by: Chapagain, Santosh, et al.
Published: (2025)
by: Chapagain, Santosh, et al.
Published: (2025)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
by: Wen, Rui, et al.
Published: (2026)
by: Wen, Rui, et al.
Published: (2026)
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
by: Liu, Yize, et al.
Published: (2025)
by: Liu, Yize, et al.
Published: (2025)
Quantifying Data Contamination in Psychometric Evaluations of LLMs
by: Han, Jongwook, et al.
Published: (2025)
by: Han, Jongwook, et al.
Published: (2025)
Fast Adversarial Attacks on Language Models In One GPU Minute
by: Sadasivan, Vinu Sankar, et al.
Published: (2024)
by: Sadasivan, Vinu Sankar, et al.
Published: (2024)
MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication
by: Sambara, Sraavya, et al.
Published: (2026)
by: Sambara, Sraavya, et al.
Published: (2026)
Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025)
by: Balasubramanian, Sriram, et al.
Published: (2025)
Task-Aware Calibration: Provably Optimal Decoding in LLMs
by: Tomov, Tim, et al.
Published: (2026)
by: Tomov, Tim, et al.
Published: (2026)
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
by: Gao, Lang, et al.
Published: (2025)
by: Gao, Lang, et al.
Published: (2025)
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
by: Zhao, Yize, et al.
Published: (2025)
by: Zhao, Yize, et al.
Published: (2025)
Explain the Flag: Contextualizing Hate Speech Beyond Censorship
by: Liartis, Jason, et al.
Published: (2026)
by: Liartis, Jason, et al.
Published: (2026)
Explore the Reasoning Capability of LLMs in the Chess Testbed
by: Wang, Shu, et al.
Published: (2024)
by: Wang, Shu, et al.
Published: (2024)
Threshold Filtering Packing for Supervised Fine-Tuning: Training Related Samples within Packs
by: Dong, Jiancheng, et al.
Published: (2024)
by: Dong, Jiancheng, et al.
Published: (2024)
Similar Items
-
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
by: Hosseini, Parsa, et al.
Published: (2025) -
PRIME: Prioritizing Interpretability in Failure Mode Extraction
by: Rezaei, Keivan, et al.
Published: (2023) -
Tool Preferences in Agentic LLMs are Unreliable
by: Faghih, Kazem, et al.
Published: (2025) -
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
by: Wang, Wenxiao, et al.
Published: (2025) -
Maestro: Joint Graph & Config Optimization for Reliable AI Agents
by: Wang, Wenxiao, et al.
Published: (2025)