MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Shbita, Basel, Ahmed, Farhan, DeLuca, Chad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLMON: An LLM-native Markup Language to Leverage Structure and Semantics at the LLM Interface
di: Hind, Michael, et al.
Pubblicazione: (2026)
di: Hind, Michael, et al.
Pubblicazione: (2026)
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
di: Daghighfarsoodeh, Alireza, et al.
Pubblicazione: (2025)
di: Daghighfarsoodeh, Alireza, et al.
Pubblicazione: (2025)
Benchmarking Large Language Models with Integer Sequence Generation Tasks
di: O'Malley, Daniel, et al.
Pubblicazione: (2024)
di: O'Malley, Daniel, et al.
Pubblicazione: (2024)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)
Runtime-Structured Task Decomposition for Agentic Coding Systems
di: Asthana, Shubhi, et al.
Pubblicazione: (2026)
di: Asthana, Shubhi, et al.
Pubblicazione: (2026)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
di: Zhou, Xingjian, et al.
Pubblicazione: (2024)
di: Zhou, Xingjian, et al.
Pubblicazione: (2024)
RESTestBench: A Benchmark for Evaluating the Effectiveness of LLM-Generated REST API Test Cases from NL Requirements
di: Kogler, Leon, et al.
Pubblicazione: (2026)
di: Kogler, Leon, et al.
Pubblicazione: (2026)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
di: Cao, Jialun, et al.
Pubblicazione: (2024)
di: Cao, Jialun, et al.
Pubblicazione: (2024)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
GeoSQL-Eval: First Evaluation of LLMs on PostGIS-Based NL2GeoSQL Queries
di: Hou, Shuyang, et al.
Pubblicazione: (2025)
di: Hou, Shuyang, et al.
Pubblicazione: (2025)
ScenicNL: Generating Probabilistic Scenario Programs from Natural Language
di: Elmaaroufi, Karim, et al.
Pubblicazione: (2024)
di: Elmaaroufi, Karim, et al.
Pubblicazione: (2024)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
di: Li, Yuangang, et al.
Pubblicazione: (2026)
di: Li, Yuangang, et al.
Pubblicazione: (2026)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025)
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025)
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
di: Sharma, Aman, et al.
Pubblicazione: (2026)
di: Sharma, Aman, et al.
Pubblicazione: (2026)
SQLong: Enhanced NL2SQL for Longer Contexts with LLMs
di: Nguyen, Dai Quoc, et al.
Pubblicazione: (2025)
di: Nguyen, Dai Quoc, et al.
Pubblicazione: (2025)
DeepSeq: High-Throughput Single-Cell RNA Sequencing Data Labeling via Web Search-Augmented Agentic Generative AI Foundation Models
di: Dajani, Saleem A. Al, et al.
Pubblicazione: (2025)
di: Dajani, Saleem A. Al, et al.
Pubblicazione: (2025)
DafnyBench: A Benchmark for Formal Software Verification
di: Loughridge, Chloe, et al.
Pubblicazione: (2024)
di: Loughridge, Chloe, et al.
Pubblicazione: (2024)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
di: Wang, Lilin, et al.
Pubblicazione: (2025)
di: Wang, Lilin, et al.
Pubblicazione: (2025)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
di: Xie, Zichen, et al.
Pubblicazione: (2026)
di: Xie, Zichen, et al.
Pubblicazione: (2026)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
di: Zhao, Zhimin, et al.
Pubblicazione: (2026)
di: Zhao, Zhimin, et al.
Pubblicazione: (2026)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
di: Galimzyanov, Timur, et al.
Pubblicazione: (2024)
di: Galimzyanov, Timur, et al.
Pubblicazione: (2024)
SWE-Bench-CL: Continual Learning for Coding Agents
di: Joshi, Thomas, et al.
Pubblicazione: (2025)
di: Joshi, Thomas, et al.
Pubblicazione: (2025)
WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks
di: Wornow, Michael, et al.
Pubblicazione: (2024)
di: Wornow, Michael, et al.
Pubblicazione: (2024)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
di: Xie, Zichen, et al.
Pubblicazione: (2026)
di: Xie, Zichen, et al.
Pubblicazione: (2026)
CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
di: Xu, Jiacheng, et al.
Pubblicazione: (2025)
di: Xu, Jiacheng, et al.
Pubblicazione: (2025)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
di: Diehl, Patrick, et al.
Pubblicazione: (2025)
di: Diehl, Patrick, et al.
Pubblicazione: (2025)
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
di: Bogavelli, Tara, et al.
Pubblicazione: (2026)
di: Bogavelli, Tara, et al.
Pubblicazione: (2026)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
di: Majdinasab, Vahid, et al.
Pubblicazione: (2025)
di: Majdinasab, Vahid, et al.
Pubblicazione: (2025)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
di: Arora, Avi, et al.
Pubblicazione: (2025)
di: Arora, Avi, et al.
Pubblicazione: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
di: Mündler, Niels, et al.
Pubblicazione: (2024)
di: Mündler, Niels, et al.
Pubblicazione: (2024)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
di: Rank, Ben, et al.
Pubblicazione: (2026)
di: Rank, Ben, et al.
Pubblicazione: (2026)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
di: Zheng, Qinkai, et al.
Pubblicazione: (2023)
di: Zheng, Qinkai, et al.
Pubblicazione: (2023)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
di: Guo, Jiawei, et al.
Pubblicazione: (2024)
di: Guo, Jiawei, et al.
Pubblicazione: (2024)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
di: Wu, Yixi, et al.
Pubblicazione: (2024)
di: Wu, Yixi, et al.
Pubblicazione: (2024)
Themisto: Jupyter-Based Runtime Benchmark
di: Grotov, Konstantin, et al.
Pubblicazione: (2025)
di: Grotov, Konstantin, et al.
Pubblicazione: (2025)
Are Large Language Models Memorizing Bug Benchmarks?
di: Ramos, Daniel, et al.
Pubblicazione: (2024)
di: Ramos, Daniel, et al.
Pubblicazione: (2024)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
di: Liu, Hugh Xuechen, et al.
Pubblicazione: (2026)
di: Liu, Hugh Xuechen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
LLMON: An LLM-native Markup Language to Leverage Structure and Semantics at the LLM Interface
di: Hind, Michael, et al.
Pubblicazione: (2026) -
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
di: Asthana, Shubhi, et al.
Pubblicazione: (2025) -
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
di: Daghighfarsoodeh, Alireza, et al.
Pubblicazione: (2025) -
Benchmarking Large Language Models with Integer Sequence Generation Tasks
di: O'Malley, Daniel, et al.
Pubblicazione: (2024) -
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)