HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Fengji, Wu, Linquan, Bai, Huiyu, Lin, Guancheng, Li, Xiao, Yu, Xiao, Wang, Yue, Chen, Bei, Keung, Jacky |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
von: Wu, Linquan, et al.
Veröffentlicht: (2026)
von: Wu, Linquan, et al.
Veröffentlicht: (2026)
Qiskit HumanEval: An Evaluation Benchmark For Quantum Code Generative Models
von: Vishwakarma, Sanjay, et al.
Veröffentlicht: (2024)
von: Vishwakarma, Sanjay, et al.
Veröffentlicht: (2024)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
HumanEval on Latest GPT Models -- 2024
von: Li, Daniel, et al.
Veröffentlicht: (2024)
von: Li, Daniel, et al.
Veröffentlicht: (2024)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
A$^2$Search: Ambiguity-Aware Question Answering with Reinforcement Learning
von: Zhang, Fengji, et al.
Veröffentlicht: (2025)
von: Zhang, Fengji, et al.
Veröffentlicht: (2025)
Data Preparation for Deep Learning based Code Smell Detection: A Systematic Literature Review
von: Zhang, Fengji, et al.
Veröffentlicht: (2024)
von: Zhang, Fengji, et al.
Veröffentlicht: (2024)
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
von: Ashrafi, Nazmus
Veröffentlicht: (2026)
von: Ashrafi, Nazmus
Veröffentlicht: (2026)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
von: Bradbury, Jeremy S., et al.
Veröffentlicht: (2024)
von: Bradbury, Jeremy S., et al.
Veröffentlicht: (2024)
Lightweight Model Editing for LLMs to Correct Deprecated API Recommendations
von: Lin, Guancheng, et al.
Veröffentlicht: (2025)
von: Lin, Guancheng, et al.
Veröffentlicht: (2025)
R2ComSync: Improving Code-Comment Synchronization with In-Context Learning and Reranking
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
Reactor Mk.1 performances: MMLU, HumanEval and BBH test results
von: Dunham, TJ, et al.
Veröffentlicht: (2024)
von: Dunham, TJ, et al.
Veröffentlicht: (2024)
ClassEval-T: Evaluating Large Language Models in Class-Level Code Translation
von: Xue, Pengyu, et al.
Veröffentlicht: (2024)
von: Xue, Pengyu, et al.
Veröffentlicht: (2024)
Towards Engineering Multi-Agent LLMs: A Protocol-Driven Approach
von: Mao, Zhenyu, et al.
Veröffentlicht: (2025)
von: Mao, Zhenyu, et al.
Veröffentlicht: (2025)
Fight Fire with Fire: How Much Can We Trust ChatGPT on Source Code-Related Tasks?
von: Yu, Xiao, et al.
Veröffentlicht: (2024)
von: Yu, Xiao, et al.
Veröffentlicht: (2024)
FaceSleuth-R: Adaptive Orientation-Aware Attention for Robust Micro-Expression Recognition
von: Wu, Linquan, et al.
Veröffentlicht: (2025)
von: Wu, Linquan, et al.
Veröffentlicht: (2025)
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
von: Luo, Hanjun, et al.
Veröffentlicht: (2025)
von: Luo, Hanjun, et al.
Veröffentlicht: (2025)
SysTradeBench: An Iterative Build-Test-Patch Benchmark for Strategy-to-Code Trading Systems with Drift-Aware Diagnostics
von: Cao, Yuchen, et al.
Veröffentlicht: (2026)
von: Cao, Yuchen, et al.
Veröffentlicht: (2026)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
von: Feng, Jia, et al.
Veröffentlicht: (2024)
von: Feng, Jia, et al.
Veröffentlicht: (2024)
Chart2Code-MoLA: Efficient Multi-Modal Code Generation via Adaptive Expert Routing
von: Wang, Yifei, et al.
Veröffentlicht: (2025)
von: Wang, Yifei, et al.
Veröffentlicht: (2025)
Advancing Autonomous Driving System Testing: Demands, Challenges, and Future Directions
von: Liao, Yihan, et al.
Veröffentlicht: (2025)
von: Liao, Yihan, et al.
Veröffentlicht: (2025)
HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2025)
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2025)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
von: Wu, Jie JW, et al.
Veröffentlicht: (2024)
von: Wu, Jie JW, et al.
Veröffentlicht: (2024)
Hybrid Privacy Policy-Code Consistency Check using Knowledge Graphs and LLMs
von: Mao, Zhenyu, et al.
Veröffentlicht: (2025)
von: Mao, Zhenyu, et al.
Veröffentlicht: (2025)
R2Code: A Self-Reflective LLM Framework for Requirements-to-Code Traceability
von: Wang, Yifei, et al.
Veröffentlicht: (2026)
von: Wang, Yifei, et al.
Veröffentlicht: (2026)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
von: Liang, Chumeng, et al.
Veröffentlicht: (2025)
von: Liang, Chumeng, et al.
Veröffentlicht: (2025)
Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face
von: Liu, Yujian, et al.
Veröffentlicht: (2026)
von: Liu, Yujian, et al.
Veröffentlicht: (2026)
OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning
von: Lu, Taiting, et al.
Veröffentlicht: (2026)
von: Lu, Taiting, et al.
Veröffentlicht: (2026)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
von: Wu, Qiming, et al.
Veröffentlicht: (2024)
von: Wu, Qiming, et al.
Veröffentlicht: (2024)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
von: Ye, Zhifan, et al.
Veröffentlicht: (2025)
von: Ye, Zhifan, et al.
Veröffentlicht: (2025)
Delving into Parameter-Efficient Fine-Tuning in Code Change Learning: An Empirical Study
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
Infinitely many associated primes of local cohomology modules of ramified regular local rings
von: Ma, Linquan
Veröffentlicht: (2026)
von: Ma, Linquan
Veröffentlicht: (2026)
DiagramNet: An End-to-End Recognition Framework and Dataset for Non-Standard System-Level Diagrams
von: Lou, Jincheng, et al.
Veröffentlicht: (2026)
von: Lou, Jincheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
von: Wu, Linquan, et al.
Veröffentlicht: (2026) -
Qiskit HumanEval: An Evaluation Benchmark For Quantum Code Generative Models
von: Vishwakarma, Sanjay, et al.
Veröffentlicht: (2024) -
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023) -
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
von: Zhang, Shudan, et al.
Veröffentlicht: (2024) -
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)