MolViBench: Evaluating LLMs on Molecular Vibe Coding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiatong, Ren, Yuxuan, Wang, Weida, Zheng, Changmeng, Wei, Xiao-yong, Li, Qing, Bian, Yatao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
by: Li, Jiatong, et al.
Published: (2025)
by: Li, Jiatong, et al.
Published: (2025)
Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal Reasoning
by: Liang, Dayong, et al.
Published: (2025)
by: Liang, Dayong, et al.
Published: (2025)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
by: Shi, Weikang, et al.
Published: (2026)
by: Shi, Weikang, et al.
Published: (2026)
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
by: Zhao, Guojiang, et al.
Published: (2025)
by: Zhao, Guojiang, et al.
Published: (2025)
MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts
by: Li, Jiatong, et al.
Published: (2024)
by: Li, Jiatong, et al.
Published: (2024)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
by: Tran, Hung, et al.
Published: (2026)
by: Tran, Hung, et al.
Published: (2026)
Membership Inference on LLMs in the Wild
by: Yi, Jiatong, et al.
Published: (2026)
by: Yi, Jiatong, et al.
Published: (2026)
Vibe Coding, Interface Flattening
by: Jin, Hongrui
Published: (2025)
by: Jin, Hongrui
Published: (2025)
GLM-5: from Vibe Coding to Agentic Engineering
by: GLM-5-Team, et al.
Published: (2026)
by: GLM-5-Team, et al.
Published: (2026)
A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning
by: Zheng, Changmeng, et al.
Published: (2024)
by: Zheng, Changmeng, et al.
Published: (2024)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
by: Zhao, Songwen, et al.
Published: (2025)
by: Zhao, Songwen, et al.
Published: (2025)
Vibe Checker: Aligning Code Evaluation with Human Preference
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
by: Yang, Yilun, et al.
Published: (2025)
by: Yang, Yilun, et al.
Published: (2025)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
LongCodeBench: Evaluating Coding LLMs at 1M Context Windows
by: Rando, Stefano, et al.
Published: (2025)
by: Rando, Stefano, et al.
Published: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
by: Huang, Jen-tse, et al.
Published: (2023)
by: Huang, Jen-tse, et al.
Published: (2023)
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
by: Inc, Xiaohongshu
Published: (2026)
by: Inc, Xiaohongshu
Published: (2026)
MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation
by: Cai, Feiyang, et al.
Published: (2025)
by: Cai, Feiyang, et al.
Published: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
by: Zhai, Zenan, et al.
Published: (2025)
by: Zhai, Zenan, et al.
Published: (2025)
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
by: Yu, Boxi, et al.
Published: (2025)
by: Yu, Boxi, et al.
Published: (2025)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024)
by: Zong, Qing, et al.
Published: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
From Code-Centric to Concept-Centric: Teaching NLP with LLM-Assisted "Vibe Coding"
by: Al-Khalifa, Hend
Published: (2026)
by: Al-Khalifa, Hend
Published: (2026)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
by: Li, Hanyu, et al.
Published: (2025)
by: Li, Hanyu, et al.
Published: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
by: Chen, Zaoyu, et al.
Published: (2026)
by: Chen, Zaoyu, et al.
Published: (2026)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
by: Xu, Shihao, et al.
Published: (2026)
by: Xu, Shihao, et al.
Published: (2026)
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
by: Ge, Wentao, et al.
Published: (2023)
by: Ge, Wentao, et al.
Published: (2023)
Conveying Imagistic Thinking in Traditional Chinese Medicine Translation: A Prompt Engineering and LLM-Based Evaluation Framework
by: Han, Jiatong
Published: (2025)
by: Han, Jiatong
Published: (2025)
GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
by: Liu, Pengfei, et al.
Published: (2023)
by: Liu, Pengfei, et al.
Published: (2023)
MolSight: Molecular Property Prediction with Images
by: Baranwal, Aaditya, et al.
Published: (2026)
by: Baranwal, Aaditya, et al.
Published: (2026)
MolMetaLM: a Physicochemical Knowledge-Guided Molecular Meta Language Model
by: Wu, Yifan, et al.
Published: (2024)
by: Wu, Yifan, et al.
Published: (2024)
CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing
by: Hu, Yucheng, et al.
Published: (2026)
by: Hu, Yucheng, et al.
Published: (2026)
VibeVoice Technical Report
by: Peng, Zhiliang, et al.
Published: (2025)
by: Peng, Zhiliang, et al.
Published: (2025)
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
by: Liu, Zhiyuan, et al.
Published: (2023)
by: Liu, Zhiyuan, et al.
Published: (2023)
Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Similar Items
-
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
by: Li, Jiatong, et al.
Published: (2025) -
Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal Reasoning
by: Liang, Dayong, et al.
Published: (2025) -
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025) -
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
by: Shi, Weikang, et al.
Published: (2026) -
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
by: Zhao, Guojiang, et al.
Published: (2025)