OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Mengzhang, Gao, Xin, Li, Yu, Lin, Honglin, Liu, Zheng, Pan, Zhuoshi, Pei, Qizhi, Shang, Xiaoran, Sun, Mengyuan, Tang, Zinan, Wang, Xiaoyang, Zhong, Zhanping, Zhu, Yun, Lin, Dahua, He, Conghui, Wu, Lijun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer
von: Lin, Honglin, et al.
Veröffentlicht: (2025)
von: Lin, Honglin, et al.
Veröffentlicht: (2025)
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
von: Liu, Zheng, et al.
Veröffentlicht: (2026)
von: Liu, Zheng, et al.
Veröffentlicht: (2026)
Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
von: Tang, Zinan, et al.
Veröffentlicht: (2025)
von: Tang, Zinan, et al.
Veröffentlicht: (2025)
Unlocking Data Value in Finance: A Study on Distillation and Difficulty-Aware Training
von: Cao, Chuxue, et al.
Veröffentlicht: (2026)
von: Cao, Chuxue, et al.
Veröffentlicht: (2026)
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
von: Lin, Honglin, et al.
Veröffentlicht: (2026)
von: Lin, Honglin, et al.
Veröffentlicht: (2026)
Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
von: Lin, Honglin, et al.
Veröffentlicht: (2025)
von: Lin, Honglin, et al.
Veröffentlicht: (2025)
LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2025)
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2025)
MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
von: Lin, Honglin, et al.
Veröffentlicht: (2026)
von: Lin, Honglin, et al.
Veröffentlicht: (2026)
IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment
von: Ming, Chenlin, et al.
Veröffentlicht: (2025)
von: Ming, Chenlin, et al.
Veröffentlicht: (2025)
A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2025)
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2025)
OpenDataLab: Empowering General Artificial Intelligence with Open Datasets
von: He, Conghui, et al.
Veröffentlicht: (2024)
von: He, Conghui, et al.
Veröffentlicht: (2024)
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
MLIP Arena: Advancing Fairness and Transparency in Machine Learning Interatomic Potentials via an Open, Accessible Benchmark Platform
von: Chiang, Yuan, et al.
Veröffentlicht: (2025)
von: Chiang, Yuan, et al.
Veröffentlicht: (2025)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
von: Zhang, Borong, et al.
Veröffentlicht: (2025)
von: Zhang, Borong, et al.
Veröffentlicht: (2025)
Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
von: Luo, Haipeng, et al.
Veröffentlicht: (2024)
von: Luo, Haipeng, et al.
Veröffentlicht: (2024)
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences
von: Lin, Fangzhou, et al.
Veröffentlicht: (2026)
von: Lin, Fangzhou, et al.
Veröffentlicht: (2026)
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
GenAI Arena: An Open Evaluation Platform for Generative Models
von: Jiang, Dongfu, et al.
Veröffentlicht: (2024)
von: Jiang, Dongfu, et al.
Veröffentlicht: (2024)
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence
von: Liu, Peigen, et al.
Veröffentlicht: (2026)
von: Liu, Peigen, et al.
Veröffentlicht: (2026)
CalArena: A Large-Scale Post-Hoc Calibration Benchmark
von: Berta, Eugène, et al.
Veröffentlicht: (2026)
von: Berta, Eugène, et al.
Veröffentlicht: (2026)
3D Arena: An Open Platform for Generative 3D Evaluation
von: Ebert, Dylan
Veröffentlicht: (2025)
von: Ebert, Dylan
Veröffentlicht: (2025)
JIR-Arena: The First Benchmark Dataset for Just-in-time Information Recommendation
von: Yang, Ke, et al.
Veröffentlicht: (2025)
von: Yang, Ke, et al.
Veröffentlicht: (2025)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
von: Huang, Zhen, et al.
Veröffentlicht: (2024)
von: Huang, Zhen, et al.
Veröffentlicht: (2024)
Arenas
Veröffentlicht: (2010)
Veröffentlicht: (2010)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models
von: Jia, Pengyue, et al.
Veröffentlicht: (2025)
von: Jia, Pengyue, et al.
Veröffentlicht: (2025)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
von: Gao, Xin, et al.
Veröffentlicht: (2025) -
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
von: Li, Yu, et al.
Veröffentlicht: (2026) -
MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer
von: Lin, Honglin, et al.
Veröffentlicht: (2025) -
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
von: Liu, Zheng, et al.
Veröffentlicht: (2026) -
Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
von: Tang, Zinan, et al.
Veröffentlicht: (2025)