ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chenchen, Li, Yuhang, Xu, Can, Liu, Jiaheng, Liu, Ao, Zhou, Changzhi, Deng, Ken, Wu, Dengpeng, Huang, Guanhua, Li, Kejiao, Yi, Qi, Xiong, Ruibin, Hu, Shihui, Zhang, Yue, Jiang, Yuhao, Xu, Zenan, Zhang, Yuanxing, Zhou, Wiggin, Zhou, Chayse, Lian, Fengzong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
by: Li, Yuhang, et al.
Published: (2025)
by: Li, Yuhang, et al.
Published: (2025)
Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework
by: Xu, Zenan, et al.
Published: (2025)
by: Xu, Zenan, et al.
Published: (2025)
Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
by: Gong, Xue, et al.
Published: (2026)
by: Gong, Xue, et al.
Published: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
by: Chou, Jason, et al.
Published: (2025)
by: Chou, Jason, et al.
Published: (2025)
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
by: Huang, Guanhua, et al.
Published: (2025)
by: Huang, Guanhua, et al.
Published: (2025)
Adaptive Deep Reasoning: Triggering Deep Thinking When Needed
by: Wang, Yunhao, et al.
Published: (2025)
by: Wang, Yunhao, et al.
Published: (2025)
Vibe AIGC: A New Paradigm for Content Generation via Agentic Orchestration
by: Liu, Jiaheng, et al.
Published: (2026)
by: Liu, Jiaheng, et al.
Published: (2026)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
by: Tang, Yuqi, et al.
Published: (2026)
by: Tang, Yuqi, et al.
Published: (2026)
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
Tailored Regulation of Graphite Microcrystals via Tandem Catalytic Carbonization for Enhanced Electrochemical Performance of Hard Carbon in the Low‐Voltage Plateau
by: Limin Zhou, et al.
Published: (2024)
by: Limin Zhou, et al.
Published: (2024)
Novel Translations
by: Wiggin, Bethany
Published: (2023)
by: Wiggin, Bethany
Published: (2023)
Enduring colonial legacies in Philadelphia
by: Bethany Wiggin
Published: (2024)
by: Bethany Wiggin
Published: (2024)
VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining
by: Zhu, Xuanyu, et al.
Published: (2026)
by: Zhu, Xuanyu, et al.
Published: (2026)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
by: Wu, Yanan, et al.
Published: (2024)
by: Wu, Yanan, et al.
Published: (2024)
CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
by: Zhang, Alexander, et al.
Published: (2025)
by: Zhang, Alexander, et al.
Published: (2025)
Rotor-Failure-Aware Quadrotors Flight in Unknown Environments
by: Zhou, Xiaobin, et al.
Published: (2025)
by: Zhou, Xiaobin, et al.
Published: (2025)
Fractional dual‐phase‐lag thermoelastic diffusion based on nonlocal elasticity and laser‐ablated impact response for 1D sandwich metallic composites
by: Jiaheng Liu, et al.
Published: (2026)
by: Jiaheng Liu, et al.
Published: (2026)
Reinforcement Learning on Pre-Training Data
by: Li, Siheng, et al.
Published: (2025)
by: Li, Siheng, et al.
Published: (2025)
Task Abstention for Large Language Models in Code Generation
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
CWSSNet: Hyperspectral Image Classification Enhanced by Wavelet Domain Convolution
by: Tong, Yulin, et al.
Published: (2025)
by: Tong, Yulin, et al.
Published: (2025)
Learning Top-k Subtask Planning Tree based on Discriminative Representation Pre-training for Decision Making
by: Ruan, Jingqing, et al.
Published: (2023)
by: Ruan, Jingqing, et al.
Published: (2023)
Fair Conformal Classification via Learning Representation-Based Groups
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
New Hampshire: The Automated Information System.
by: Wiggin, Kendall F.
Published: (1996)
by: Wiggin, Kendall F.
Published: (1996)
Gateway 2000: A Strategic Plan for the New Hampshire Automated Information System.
by: Wiggin, Kendall F.
Published: (1993)
by: Wiggin, Kendall F.
Published: (1993)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024)
by: Ouyang, Kun, et al.
Published: (2024)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
by: Li, Youquan, et al.
Published: (2024)
by: Li, Youquan, et al.
Published: (2024)
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
by: Zhai, Zenan, et al.
Published: (2025)
by: Zhai, Zenan, et al.
Published: (2025)
An integrated bioinformatics analysis reveals IRF8 as a critical biomarker for immune infiltration in atherosclerosis advance
by: Donglai Zhou, et al.
Published: (2024)
by: Donglai Zhou, et al.
Published: (2024)
PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning
by: Zhang, Luan, et al.
Published: (2026)
by: Zhang, Luan, et al.
Published: (2026)
Granular parakeratosis
by: Shihui Zhou, et al.
Published: (2025)
by: Shihui Zhou, et al.
Published: (2025)
Lease Accounting Standard (ASC 842), Temporary Book‐Tax Differences, and Capital Market Uncertainty
by: Yan Zhou, et al.
Published: (2025)
by: Yan Zhou, et al.
Published: (2025)
How Does a Firm's Information Environment Influence CEO Compensation?
by: Shihui Fan, et al.
Published: (2024)
by: Shihui Fan, et al.
Published: (2024)
Second-order monotonicity conditions and mean field games with volatility control
by: Mou, Chenchen, et al.
Published: (2025)
by: Mou, Chenchen, et al.
Published: (2025)
MIO: A Foundation Model on Multimodal Tokens
by: Wang, Zekun, et al.
Published: (2024)
by: Wang, Zekun, et al.
Published: (2024)
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
by: Xu, Pengju, et al.
Published: (2025)
by: Xu, Pengju, et al.
Published: (2025)
Frequency response of the nonclassicality and its correspondence to the classical dynamics
by: Shihui Zhang
Published: (2016)
by: Shihui Zhang
Published: (2016)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
Generative Model Watermarking Suppressing High-Frequency Artifacts
by: Zhang, Li, et al.
Published: (2023)
by: Zhang, Li, et al.
Published: (2023)
The Art of Efficient Reasoning: Data, Reward, and Optimization
by: Wu, Taiqiang, et al.
Published: (2026)
by: Wu, Taiqiang, et al.
Published: (2026)
Similar Items
-
ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
by: Li, Yuhang, et al.
Published: (2025) -
Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework
by: Xu, Zenan, et al.
Published: (2025) -
Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
by: Gong, Xue, et al.
Published: (2026) -
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
by: Chou, Jason, et al.
Published: (2025) -
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
by: Huang, Guanhua, et al.
Published: (2025)