Test of Time: Rethinking Temporal Signal of Benchmark Contamination
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Terry Jingchen, Dev, Gopal, Wang, Ning, Obreiter, Max, Pandey, Punya Syon, Samway, Keenan, Jiang, Wenyuan, Huang, Yinya, Schölkopf, Bernhard, Sachan, Mrinmaya, Jin, Zhijing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
por: Ortu, Francesco, et al.
Publicado: (2026)
por: Ortu, Francesco, et al.
Publicado: (2026)
BinaryPPO: Efficient Policy Optimization for Binary Classification
por: Pandey, Punya Syon, et al.
Publicado: (2026)
por: Pandey, Punya Syon, et al.
Publicado: (2026)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
por: Harrasse, Abir, et al.
Publicado: (2025)
por: Harrasse, Abir, et al.
Publicado: (2025)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
por: Ceraolo, Roberto, et al.
Publicado: (2024)
por: Ceraolo, Roberto, et al.
Publicado: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
por: Simko, Samuel, et al.
Publicado: (2025)
por: Simko, Samuel, et al.
Publicado: (2025)
Uncovering Hidden Correctness in LLM Causal Reasoning via Symbolic Verification
por: He, Paul, et al.
Publicado: (2026)
por: He, Paul, et al.
Publicado: (2026)
Are Language Models Consequentialist or Deontological Moral Reasoners?
por: Samway, Keenan, et al.
Publicado: (2025)
por: Samway, Keenan, et al.
Publicado: (2025)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
por: Draye, Florent, et al.
Publicado: (2026)
por: Draye, Florent, et al.
Publicado: (2026)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
Lean Meets Theoretical Computer Science: Scalable Synthesis of Theorem Proving Challenges in Formal-Informal Pairs
por: Zhang, Terry Jingchen, et al.
Publicado: (2025)
por: Zhang, Terry Jingchen, et al.
Publicado: (2025)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
por: Piatti, Giorgio, et al.
Publicado: (2024)
por: Piatti, Giorgio, et al.
Publicado: (2024)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
por: Jenny, David F., et al.
Publicado: (2023)
por: Jenny, David F., et al.
Publicado: (2023)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
por: Liu, Xinge, et al.
Publicado: (2026)
por: Liu, Xinge, et al.
Publicado: (2026)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
por: Ortu, Francesco, et al.
Publicado: (2024)
por: Ortu, Francesco, et al.
Publicado: (2024)
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
por: Lyu, Zhiheng, et al.
Publicado: (2024)
por: Lyu, Zhiheng, et al.
Publicado: (2024)
When Do Language Models Endorse Limitations on Human Rights Principles?
por: Samway, Keenan, et al.
Publicado: (2026)
por: Samway, Keenan, et al.
Publicado: (2026)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
por: Piedrahita, David Guzman, et al.
Publicado: (2025)
por: Piedrahita, David Guzman, et al.
Publicado: (2025)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
CausalCite: A Causal Formulation of Paper Citations
por: Kumar, Ishan, et al.
Publicado: (2023)
por: Kumar, Ishan, et al.
Publicado: (2023)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
por: Opedal, Andreas, et al.
Publicado: (2024)
por: Opedal, Andreas, et al.
Publicado: (2024)
Can Large Language Models Infer Causation from Correlation?
por: Jin, Zhijing, et al.
Publicado: (2023)
por: Jin, Zhijing, et al.
Publicado: (2023)
Implicit Personalization in Language Models: A Systematic Study
por: Jin, Zhijing, et al.
Publicado: (2024)
por: Jin, Zhijing, et al.
Publicado: (2024)
Can Theoretical Physics Research Benefit from Language Agents?
por: Lu, Sirui, et al.
Publicado: (2025)
por: Lu, Sirui, et al.
Publicado: (2025)
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning
por: Cui, Shaobo, et al.
Publicado: (2024)
por: Cui, Shaobo, et al.
Publicado: (2024)
Educators' Perceptions of Large Language Models as Tutors: Comparing Human and AI Tutors in a Blind Text-only Setting
por: Chowdhury, Sankalan Pal, et al.
Publicado: (2025)
por: Chowdhury, Sankalan Pal, et al.
Publicado: (2025)
Learning to Reason Efficiently with A* Post-Training
por: Opedal, Andreas, et al.
Publicado: (2026)
por: Opedal, Andreas, et al.
Publicado: (2026)
Fluid Representations in Reasoning Models
por: Kharlapenko, Dmitrii, et al.
Publicado: (2026)
por: Kharlapenko, Dmitrii, et al.
Publicado: (2026)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
por: Hossain, Saad, et al.
Publicado: (2026)
por: Hossain, Saad, et al.
Publicado: (2026)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
por: Vennemeyer, Daniel, et al.
Publicado: (2026)
por: Vennemeyer, Daniel, et al.
Publicado: (2026)
Causality can systematically address the monsters under the bench(marks)
por: Leeb, Felix, et al.
Publicado: (2025)
por: Leeb, Felix, et al.
Publicado: (2025)
Causal Responsibility Attribution for Human-AI Collaboration
por: Qi, Yahang, et al.
Publicado: (2024)
por: Qi, Yahang, et al.
Publicado: (2024)
SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
por: Xiang, Kun, et al.
Publicado: (2025)
por: Xiang, Kun, et al.
Publicado: (2025)
CLadder: Assessing Causal Reasoning in Language Models
por: Jin, Zhijing, et al.
Publicado: (2023)
por: Jin, Zhijing, et al.
Publicado: (2023)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
por: Cui, Peng, et al.
Publicado: (2025)
por: Cui, Peng, et al.
Publicado: (2025)
Language Model Alignment in Multilingual Trolley Problems
por: Jin, Zhijing, et al.
Publicado: (2024)
por: Jin, Zhijing, et al.
Publicado: (2024)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
por: Kassem, Aly M., et al.
Publicado: (2025)
por: Kassem, Aly M., et al.
Publicado: (2025)
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
por: Lalwani, Abhinav, et al.
Publicado: (2024)
por: Lalwani, Abhinav, et al.
Publicado: (2024)
Are Language Models Efficient Reasoners? A Perspective from Logic Programming
por: Opedal, Andreas, et al.
Publicado: (2025)
por: Opedal, Andreas, et al.
Publicado: (2025)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
por: Opedal, Andreas, et al.
Publicado: (2024)
por: Opedal, Andreas, et al.
Publicado: (2024)
Ejemplares similares
-
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
por: Ortu, Francesco, et al.
Publicado: (2026) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
por: Pandey, Punya Syon, et al.
Publicado: (2026) -
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
por: Harrasse, Abir, et al.
Publicado: (2025) -
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
por: Ceraolo, Roberto, et al.
Publicado: (2024) -
Improving Large Language Model Safety with Contrastive Representation Learning
por: Simko, Samuel, et al.
Publicado: (2025)