ManiBench: A Benchmark for Testing Visual-Logic Drift and Syntactic Hallucinations in Manim Code Generation
Fuente:
arXiv
Saved in:
| Main Author: | Oli, Nabin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manimator: Transforming Research Papers into Visual Explanations
by: P, Samarth, et al.
Published: (2025)
by: P, Samarth, et al.
Published: (2025)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Training and Agentic Inference Strategies for LLM-based Manim Animation Generation
by: Silva, Ravidu Suien Rammuni, et al.
Published: (2026)
by: Silva, Ravidu Suien Rammuni, et al.
Published: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
by: Xu, Qiao, et al.
Published: (2026)
by: Xu, Qiao, et al.
Published: (2026)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
by: Lee, Yunseo, et al.
Published: (2025)
by: Lee, Yunseo, et al.
Published: (2025)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
by: Bao, Forrest Sheng, et al.
Published: (2024)
by: Bao, Forrest Sheng, et al.
Published: (2024)
BenchScope: How Many Independent Signals Does Your Benchmark Provide?
by: Sha, Tommy, et al.
Published: (2026)
by: Sha, Tommy, et al.
Published: (2026)
ChaosBench-Logic: A Benchmark for Logical and Symbolic Reasoning on Chaotic Dynamical Systems
by: Thomas, Noel
Published: (2026)
by: Thomas, Noel
Published: (2026)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
by: Gioacchini, Luca, et al.
Published: (2024)
by: Gioacchini, Luca, et al.
Published: (2024)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
by: Sun, Siqi, et al.
Published: (2026)
by: Sun, Siqi, et al.
Published: (2026)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
by: Chen, Kesheng, et al.
Published: (2026)
by: Chen, Kesheng, et al.
Published: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
by: Ivanov, Maksim, et al.
Published: (2026)
by: Ivanov, Maksim, et al.
Published: (2026)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
by: Cui, Yi
Published: (2025)
by: Cui, Yi
Published: (2025)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation
by: He, Yibo, et al.
Published: (2025)
by: He, Yibo, et al.
Published: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models
by: Zuo, Kaiwen, et al.
Published: (2024)
by: Zuo, Kaiwen, et al.
Published: (2024)
Web-Bench: A LLM Code Benchmark Based on Web Standards and Frameworks
by: Xu, Kai, et al.
Published: (2025)
by: Xu, Kai, et al.
Published: (2025)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
by: Zhuo, Terry Yue, et al.
Published: (2024)
by: Zhuo, Terry Yue, et al.
Published: (2024)
SWE Context Bench: A Benchmark for Context Learning in Coding
by: Zhu, Jiayuan, et al.
Published: (2026)
by: Zhu, Jiayuan, et al.
Published: (2026)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
by: Li, Xinghang, et al.
Published: (2025)
by: Li, Xinghang, et al.
Published: (2025)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
by: Padwal, Vedant
Published: (2026)
by: Padwal, Vedant
Published: (2026)
Benchmarking Generative Models on Computational Thinking Tests in Elementary Visual Programming
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models
by: Wang, JiYang, et al.
Published: (2026)
by: Wang, JiYang, et al.
Published: (2026)
Code Hallucination
by: Rahman, Mirza Masfiqur, et al.
Published: (2024)
by: Rahman, Mirza Masfiqur, et al.
Published: (2024)
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
by: Huang, Jiahao, et al.
Published: (2026)
by: Huang, Jiahao, et al.
Published: (2026)
SynBench: A Benchmark for Differentially Private Text Generation
by: Sun, Yidan, et al.
Published: (2025)
by: Sun, Yidan, et al.
Published: (2025)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026)
by: Zhou, Qixing, et al.
Published: (2026)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
by: Hua, Tianyu, et al.
Published: (2025)
by: Hua, Tianyu, et al.
Published: (2025)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
by: Jing, Lucas, et al.
Published: (2026)
by: Jing, Lucas, et al.
Published: (2026)
MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
Hallucination Benchmark in Medical Visual Question Answering
by: Wu, Jinge, et al.
Published: (2024)
by: Wu, Jinge, et al.
Published: (2024)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
Similar Items
-
Manimator: Transforming Research Papers into Visual Explanations
by: P, Samarth, et al.
Published: (2025) -
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
by: Jiang, Nan, et al.
Published: (2024) -
Training and Agentic Inference Strategies for LLM-based Manim Animation Generation
by: Silva, Ravidu Suien Rammuni, et al.
Published: (2026) -
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
by: Xu, Qiao, et al.
Published: (2026) -
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
by: Lee, Yunseo, et al.
Published: (2025)