Gespeichert in:
| Hauptverfasser: | Yang, Xinwei, Liu, Zhaofeng, Huang, Chen, Zhang, Jiashuai, Zhang, Tong, Zhang, Yifan, Lei, Wenqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2505.16667 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
von: Guo, Ruiling, et al.
Veröffentlicht: (2025)
von: Guo, Ruiling, et al.
Veröffentlicht: (2025)
Cross-Space Adaptive Filter: Integrating Graph Topology and Node Attributes for Alleviating the Over-smoothing Problem
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
Towards Better Generalization via Distributional Input Projection Network
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
von: Chen, Weiyi, et al.
Veröffentlicht: (2026)
von: Chen, Weiyi, et al.
Veröffentlicht: (2026)
E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)
CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
Benchmarking LLM Tool-Use in the Wild
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
Evaluation of Retrieval-Augmented Generation: A Survey
von: Yu, Hao, et al.
Veröffentlicht: (2024)
von: Yu, Hao, et al.
Veröffentlicht: (2024)
Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark
von: Zhang, Longgang, et al.
Veröffentlicht: (2026)
von: Zhang, Longgang, et al.
Veröffentlicht: (2026)
The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?
von: Xun, Yuan, et al.
Veröffentlicht: (2025)
von: Xun, Yuan, et al.
Veröffentlicht: (2025)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
von: Zhou, Yifan, et al.
Veröffentlicht: (2026)
von: Zhou, Yifan, et al.
Veröffentlicht: (2026)
BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD
von: Zhang, Haozhe, et al.
Veröffentlicht: (2026)
von: Zhang, Haozhe, et al.
Veröffentlicht: (2026)
Advancing ESG Intelligence: An Expert-level Agent and Comprehensive Benchmark for Sustainable Finance
von: Zhao, Yilei, et al.
Veröffentlicht: (2026)
von: Zhao, Yilei, et al.
Veröffentlicht: (2026)
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Competition-Level Problems are Effective LLM Evaluators
von: Huang, Yiming, et al.
Veröffentlicht: (2023)
von: Huang, Yiming, et al.
Veröffentlicht: (2023)
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution
von: Li, Han, et al.
Veröffentlicht: (2026)
von: Li, Han, et al.
Veröffentlicht: (2026)
CodeApex: A Bilingual Programming Evaluation Benchmark for Large Language Models
von: Fu, Lingyue, et al.
Veröffentlicht: (2023)
von: Fu, Lingyue, et al.
Veröffentlicht: (2023)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
OpenGLT: A Comprehensive Benchmark of Graph Neural Networks for Graph-Level Tasks
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering
von: Zhang, Weikang, et al.
Veröffentlicht: (2026)
von: Zhang, Weikang, et al.
Veröffentlicht: (2026)
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
Generalizable Agent Modeling for Agent Collaboration-Competition Adaptation with Multi-Retrieval and Dynamic Generation
von: Wang, Chenxu, et al.
Veröffentlicht: (2025)
von: Wang, Chenxu, et al.
Veröffentlicht: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
Cooperative-Competitive Team Play of Real-World Craft Robots
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning
von: Zhang, Min, et al.
Veröffentlicht: (2024)
von: Zhang, Min, et al.
Veröffentlicht: (2024)
OCDB: Revisiting Causal Discovery with a Comprehensive Benchmark and Evaluation Framework
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
Human-centered In-building Embodied Delivery Benchmark
von: Xu, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoqun, et al.
Veröffentlicht: (2024)
TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs
von: Li, Zhuofeng, et al.
Veröffentlicht: (2024)
von: Li, Zhuofeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
von: Guo, Ruiling, et al.
Veröffentlicht: (2025) -
Cross-Space Adaptive Filter: Integrating Graph Topology and Node Attributes for Alleviating the Over-smoothing Problem
von: Huang, Chen, et al.
Veröffentlicht: (2024) -
Towards Better Generalization via Distributional Input Projection Network
von: Hao, Yifan, et al.
Veröffentlicht: (2025) -
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
von: Chen, Weiyi, et al.
Veröffentlicht: (2026) -
E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)