Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Linbo, Zhao, Jinman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training on the Benchmark Is Not All You Need
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Contrast Is All You Need
by: Kilic, Burak, et al.
Published: (2023)
by: Kilic, Burak, et al.
Published: (2023)
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Rho-1: Not All Tokens Are What You Need
by: Lin, Zhenghao, et al.
Published: (2024)
by: Lin, Zhenghao, et al.
Published: (2024)
More Agents Is All You Need
by: Li, Junyou, et al.
Published: (2024)
by: Li, Junyou, et al.
Published: (2024)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
by: Wu, Zongqian, et al.
Published: (2025)
by: Wu, Zongqian, et al.
Published: (2025)
Tensor Product Attention Is All You Need
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
Sparsity May Be All You Need: Sparse Random Parameter Adaptation
by: Rios, Jesus, et al.
Published: (2025)
by: Rios, Jesus, et al.
Published: (2025)
Self-Verification is All You Need To Pass The Japanese Bar Examination
by: Shin, Andrew
Published: (2026)
by: Shin, Andrew
Published: (2026)
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
by: Bai, Yuelin, et al.
Published: (2024)
by: Bai, Yuelin, et al.
Published: (2024)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
by: Lu, Xiaoding, et al.
Published: (2024)
by: Lu, Xiaoding, et al.
Published: (2024)
Rethinking Data Selection at Scale: Random Selection is Almost All You Need
by: Xia, Tingyu, et al.
Published: (2024)
by: Xia, Tingyu, et al.
Published: (2024)
Communication is All You Need: Persuasion Dataset Construction via Multi-LLM Communication
by: Ma, Weicheng, et al.
Published: (2025)
by: Ma, Weicheng, et al.
Published: (2025)
Towards Supporting Legal Argumentation with NLP: Is More Data Really All You Need?
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
by: Dhaliwal, Mehak, et al.
Published: (2026)
by: Dhaliwal, Mehak, et al.
Published: (2026)
Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning
by: Som, Anirudh, et al.
Published: (2023)
by: Som, Anirudh, et al.
Published: (2023)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
by: Li, Chuhan, et al.
Published: (2024)
by: Li, Chuhan, et al.
Published: (2024)
From Biased Chatbots to Biased Agents: Examining Role Assignment Effects on LLM Agent Robustness
by: Cao, Linbo, et al.
Published: (2026)
by: Cao, Linbo, et al.
Published: (2026)
Synthetic Data RL: Task Definition Is All You Need
by: Guo, Yiduo, et al.
Published: (2025)
by: Guo, Yiduo, et al.
Published: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
by: Nguyen-Tri, Quan, et al.
Published: (2025)
by: Nguyen-Tri, Quan, et al.
Published: (2025)
Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization
by: Dong, Zhijin
Published: (2025)
by: Dong, Zhijin
Published: (2025)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
by: Li, Ruanjun, et al.
Published: (2025)
by: Li, Ruanjun, et al.
Published: (2025)
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
SecEncoder: Logs are All You Need in Security
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
by: Li, Mengjie, et al.
Published: (2025)
by: Li, Mengjie, et al.
Published: (2025)
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
by: Bsharat, Sondos Mahmoud, et al.
Published: (2023)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2023)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
by: Steinmetz, Cody, et al.
Published: (2025)
by: Steinmetz, Cody, et al.
Published: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
by: Liu, Yiyang, et al.
Published: (2025)
by: Liu, Yiyang, et al.
Published: (2025)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
by: Gomaa, Eyad, et al.
Published: (2024)
by: Gomaa, Eyad, et al.
Published: (2024)
ViC: Virtual Compiler Is All You Need For Assembly Code Search
by: Gao, Zeyu, et al.
Published: (2024)
by: Gao, Zeyu, et al.
Published: (2024)
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
by: Tan, Yingshui, et al.
Published: (2024)
by: Tan, Yingshui, et al.
Published: (2024)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
by: Bai, Yuqi, et al.
Published: (2025)
by: Bai, Yuqi, et al.
Published: (2025)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
by: Lu, Jinghui, et al.
Published: (2025)
by: Lu, Jinghui, et al.
Published: (2025)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
by: Gan, Chunjing, et al.
Published: (2024)
by: Gan, Chunjing, et al.
Published: (2024)
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
by: Huang, Tiancheng, et al.
Published: (2025)
by: Huang, Tiancheng, et al.
Published: (2025)
SwaQuAD-24: QA Benchmark Dataset in Swahili
by: Kondoro, Alfred Malengo
Published: (2024)
by: Kondoro, Alfred Malengo
Published: (2024)
Similar Items
-
Training on the Benchmark Is Not All You Need
by: Ni, Shiwen, et al.
Published: (2024) -
Contrast Is All You Need
by: Kilic, Burak, et al.
Published: (2023) -
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025) -
Not All Documents Are What You Need for Extracting Instruction Tuning Data
by: Zhang, Chi, et al.
Published: (2025) -
Rho-1: Not All Tokens Are What You Need
by: Lin, Zhenghao, et al.
Published: (2024)