One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shen, Chengyu, Hou, Yanheng, Pan, Minghui, He, Runming, Wong, Zhen Hao, Qiang, Meiyi, Liu, Zhou, Liang, Hao, Lai, Peichao, Sheng, Zeang, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
par: Feng, Hengyi, et autres
Publié: (2025)
par: Feng, Hengyi, et autres
Publié: (2025)
Let's Verify Math Questions Step by Step
par: Shen, Chengyu, et autres
Publié: (2025)
par: Shen, Chengyu, et autres
Publié: (2025)
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
par: Ma, Lu, et autres
Publié: (2025)
par: Ma, Lu, et autres
Publié: (2025)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
par: Feng, Hengyi, et autres
Publié: (2026)
par: Feng, Hengyi, et autres
Publié: (2026)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
par: Wong, Zhen Hao, et autres
Publié: (2025)
par: Wong, Zhen Hao, et autres
Publié: (2025)
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
par: Wong, Zhen Hao, et autres
Publié: (2025)
par: Wong, Zhen Hao, et autres
Publié: (2025)
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
par: Liang, Hao, et autres
Publié: (2025)
par: Liang, Hao, et autres
Publié: (2025)
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
par: Liang, Hao, et autres
Publié: (2026)
par: Liang, Hao, et autres
Publié: (2026)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
par: You, Qijie, et autres
Publié: (2026)
par: You, Qijie, et autres
Publié: (2026)
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
par: Guo, Tianyu, et autres
Publié: (2025)
par: Guo, Tianyu, et autres
Publié: (2025)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
par: Xu, Weiwei, et autres
Publié: (2024)
par: Xu, Weiwei, et autres
Publié: (2024)
Can LLMs be Good Graph Judge for Knowledge Graph Construction?
par: Huang, Haoyu, et autres
Publié: (2024)
par: Huang, Haoyu, et autres
Publié: (2024)
A2Eval: Agentic and Automated Evaluation for Embodied Brain
par: Zhang, Shuai, et autres
Publié: (2026)
par: Zhang, Shuai, et autres
Publié: (2026)
Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning
par: Lai, Peichao, et autres
Publié: (2024)
par: Lai, Peichao, et autres
Publié: (2024)
VeriMind: Agentic LLM for Automated Verilog Generation with a Novel Evaluation Metric
par: Nadimi, Bardia, et autres
Publié: (2025)
par: Nadimi, Bardia, et autres
Publié: (2025)
Acceleration Algorithms in GNNs: A Survey
par: Ma, Lu, et autres
Publié: (2024)
par: Ma, Lu, et autres
Publié: (2024)
DARO: Difficulty-Aware Reweighting Policy Optimization
par: Zhou, Jingyu, et autres
Publié: (2025)
par: Zhou, Jingyu, et autres
Publié: (2025)
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
par: Wang, Yibo, et autres
Publié: (2026)
par: Wang, Yibo, et autres
Publié: (2026)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
par: Liang, Hao, et autres
Publié: (2025)
par: Liang, Hao, et autres
Publié: (2025)
RepEval: Effective Text Evaluation with LLM Representation
par: Sheng, Shuqian, et autres
Publié: (2024)
par: Sheng, Shuqian, et autres
Publié: (2024)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
par: Liang, Hao, et autres
Publié: (2026)
par: Liang, Hao, et autres
Publié: (2026)
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
par: Liang, Hao, et autres
Publié: (2026)
par: Liang, Hao, et autres
Publié: (2026)
Towards Scalable and Deep Graph Neural Networks via Noise Masking
par: Liang, Yuxuan, et autres
Publié: (2024)
par: Liang, Yuxuan, et autres
Publié: (2024)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
par: Cai, Qifeng, et autres
Publié: (2025)
par: Cai, Qifeng, et autres
Publié: (2025)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
par: Wei, Tianjun, et autres
Publié: (2025)
par: Wei, Tianjun, et autres
Publié: (2025)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
par: Liang, Chumeng, et autres
Publié: (2025)
par: Liang, Chumeng, et autres
Publié: (2025)
On Randomness in Agentic Evals
par: Bjarnason, Bjarni Haukur, et autres
Publié: (2026)
par: Bjarnason, Bjarni Haukur, et autres
Publié: (2026)
An Agentic System for Rare Disease Diagnosis with Traceable Reasoning
par: Zhao, Weike, et autres
Publié: (2025)
par: Zhao, Weike, et autres
Publié: (2025)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
par: Yehudai, Asaf, et autres
Publié: (2026)
par: Yehudai, Asaf, et autres
Publié: (2026)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
par: Wu, JiaRu, et autres
Publié: (2025)
par: Wu, JiaRu, et autres
Publié: (2025)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
par: Wang, Yixu, et autres
Publié: (2025)
par: Wang, Yixu, et autres
Publié: (2025)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
par: Li, Zhen, et autres
Publié: (2025)
par: Li, Zhen, et autres
Publié: (2025)
AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents
par: Liang, Yuanzhi, et autres
Publié: (2024)
par: Liang, Yuanzhi, et autres
Publié: (2024)
Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction
par: Bi, Shuzhen, et autres
Publié: (2026)
par: Bi, Shuzhen, et autres
Publié: (2026)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
par: Liang, Hao, et autres
Publié: (2024)
par: Liang, Hao, et autres
Publié: (2024)
SycEval: Evaluating LLM Sycophancy
par: Fanous, Aaron, et autres
Publié: (2025)
par: Fanous, Aaron, et autres
Publié: (2025)
Existence and nonrelativistic limit of ground states to nonlinear Dirac equation
par: Chen, Pan, et autres
Publié: (2025)
par: Chen, Pan, et autres
Publié: (2025)
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
par: Chen, Yongrui, et autres
Publié: (2025)
par: Chen, Yongrui, et autres
Publié: (2025)
DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review
par: Weng, Yixuan, et autres
Publié: (2026)
par: Weng, Yixuan, et autres
Publié: (2026)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
par: He, Yuanqin, et autres
Publié: (2024)
par: He, Yuanqin, et autres
Publié: (2024)
Documents similaires
-
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
par: Feng, Hengyi, et autres
Publié: (2025) -
Let's Verify Math Questions Step by Step
par: Shen, Chengyu, et autres
Publié: (2025) -
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
par: Ma, Lu, et autres
Publié: (2025) -
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
par: Feng, Hengyi, et autres
Publié: (2026) -
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
par: Wong, Zhen Hao, et autres
Publié: (2025)