MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qihao, Huang, Yangyu, Lv, Tengchao, Cui, Lei, Sun, Qinzheng, Mao, Shaoguang, Zhang, Xin, Xin, Ying, Yin, Qiufeng, Li, Scarlett, Wei, Furu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problems
by: Li, Zongqian, et al.
Published: (2026)
by: Li, Zongqian, et al.
Published: (2026)
RedStone: Curating General, Code, Math, and QA Data for Large Language Models
by: Chang, Yaoyao, et al.
Published: (2024)
by: Chang, Yaoyao, et al.
Published: (2024)
Data Efficacy for Language Model Training
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
by: Huang, Yangyu, et al.
Published: (2025)
by: Huang, Yangyu, et al.
Published: (2025)
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
by: Mayr, Martin, et al.
Published: (2026)
by: Mayr, Martin, et al.
Published: (2026)
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
by: Pfister, Rolf, et al.
Published: (2025)
by: Pfister, Rolf, et al.
Published: (2025)
An on-demand resource allocation algorithm for a quantum network hub and its performance analysis
by: Gauthier, Scarlett, et al.
Published: (2024)
by: Gauthier, Scarlett, et al.
Published: (2024)
PIM or CXL-PIM? Understanding Architectural Trade-offs Through Large-Scale Benchmarking
by: Lee, I-Ting, et al.
Published: (2025)
by: Lee, I-Ting, et al.
Published: (2025)
Energy Efficiency Analysis of Active RIS-enhanced Wireless Network under Power-Sum Constraint
by: Xin, Jingdie, et al.
Published: (2025)
by: Xin, Jingdie, et al.
Published: (2025)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs
by: Chen, Mingkai, et al.
Published: (2024)
by: Chen, Mingkai, et al.
Published: (2024)
Systematic Performance Evaluation Framework for LEO Mega-Constellation Satellite Networks
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
Green AI: Exploring Carbon Footprints, Mitigation Strategies, and Trade Offs in Large Language Model Training
by: Liu, Vivian, et al.
Published: (2024)
by: Liu, Vivian, et al.
Published: (2024)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
by: Shen, Siyuan, et al.
Published: (2025)
by: Shen, Siyuan, et al.
Published: (2025)
Wasure: A Modular Toolkit for Comprehensive WebAssembly Benchmarking
by: Carissimi, Riccardo, et al.
Published: (2026)
by: Carissimi, Riccardo, et al.
Published: (2026)
A Continuous Benchmarking Infrastructure for High-Performance Computing Applications
by: Alt, Christoph, et al.
Published: (2024)
by: Alt, Christoph, et al.
Published: (2024)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
by: Zhou, Fang, et al.
Published: (2026)
by: Zhou, Fang, et al.
Published: (2026)
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
by: Salaria, Shweta, et al.
Published: (2025)
by: Salaria, Shweta, et al.
Published: (2025)
WritePolicyBench: Benchmarking Memory Write Policies under Byte Budgets
by: Cham, Edgard El
Published: (2026)
by: Cham, Edgard El
Published: (2026)
Ecoscape: Fault Tolerance Benchmark for Adaptive Remediation Strategies in Real-Time Edge ML
by: Reiter, Hendrik, et al.
Published: (2025)
by: Reiter, Hendrik, et al.
Published: (2025)
End-to-End Throughput Benchmarking of Portable Deterministic CNN-Based Signal Processing Pipelines
by: Boerkamp, Christiaan, et al.
Published: (2026)
by: Boerkamp, Christiaan, et al.
Published: (2026)
Energy-Efficient Software Development: A Multi-dimensional Empirical Analysis of Stack Overflow
by: Jin, Bihui, et al.
Published: (2024)
by: Jin, Bihui, et al.
Published: (2024)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026)
by: Ding, Jiabiao, et al.
Published: (2026)
Benchmark-based Study of CPU/GPU Power-Related Features through JAX and TensorFlow
by: Tchakoute, Roblex Nana, et al.
Published: (2025)
by: Tchakoute, Roblex Nana, et al.
Published: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
by: Zhao, Yanbo, et al.
Published: (2025)
by: Zhao, Yanbo, et al.
Published: (2025)
XRFlux: Virtual Reality Benchmark for Edge Caching Systems
by: Alfares, Nader, et al.
Published: (2024)
by: Alfares, Nader, et al.
Published: (2024)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
Heterogeneous Memory Benchmarking Toolkit
by: Ghaemi, Golsana, et al.
Published: (2025)
by: Ghaemi, Golsana, et al.
Published: (2025)
Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs
by: Yu, Dingcui, et al.
Published: (2025)
by: Yu, Dingcui, et al.
Published: (2025)
Multi-Strided Access Patterns to Boost Hardware Prefetching
by: Blom, Miguel O., et al.
Published: (2024)
by: Blom, Miguel O., et al.
Published: (2024)
Introducing the Arm-membench Throughput Benchmark
by: Burth, Cyrill, et al.
Published: (2025)
by: Burth, Cyrill, et al.
Published: (2025)
Benchmarking GPUs on SVBRDF Extractor Model
by: Kandel, Narayan, et al.
Published: (2023)
by: Kandel, Narayan, et al.
Published: (2023)
Towards Multi-dimensional Elasticity for Pervasive Stream Processing Services
by: Sedlak, Boris, et al.
Published: (2025)
by: Sedlak, Boris, et al.
Published: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
Efficient Hybrid Amplitude-Phase Quantization for Multi-Antenna Relay System
by: Kim, Changdae, et al.
Published: (2025)
by: Kim, Changdae, et al.
Published: (2025)
ppOpen-AT: A Directive-base Auto-tuning Language
by: Katagiri, Takahiro
Published: (2024)
by: Katagiri, Takahiro
Published: (2024)
Efficiently Ranking Software Variants with Minimal Benchmarks
by: Matricon, Théo, et al.
Published: (2025)
by: Matricon, Théo, et al.
Published: (2025)
Spatiotemporal Non-Uniformity-Aware Online Task Scheduling in Collaborative Edge Computing for Industrial Internet of Things
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Similar Items
-
Scaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problems
by: Li, Zongqian, et al.
Published: (2026) -
RedStone: Curating General, Code, Math, and QA Data for Large Language Models
by: Chang, Yaoyao, et al.
Published: (2024) -
Data Efficacy for Language Model Training
by: Dai, Yalun, et al.
Published: (2025) -
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
by: Huang, Yangyu, et al.
Published: (2025) -
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
by: Mayr, Martin, et al.
Published: (2026)