Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Hugh Xuechen, Tatar, Kıvanç |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
di: Li, Xin-Ye, et al.
Pubblicazione: (2026)
di: Li, Xin-Ye, et al.
Pubblicazione: (2026)
Grounding Machine Creativity in Game Design Knowledge Representations: Empirical Probing of LLM-Based Executable Synthesis of Goal Playable Patterns under Structural Constraints
di: Liu, Hugh Xuechen, et al.
Pubblicazione: (2026)
di: Liu, Hugh Xuechen, et al.
Pubblicazione: (2026)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
di: Jiang, Shan, et al.
Pubblicazione: (2026)
di: Jiang, Shan, et al.
Pubblicazione: (2026)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025)
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
di: Gu, Alex, et al.
Pubblicazione: (2024)
di: Gu, Alex, et al.
Pubblicazione: (2024)
Automatic Generation of Executable BPMN Models from Medical Guidelines
di: Sekar, Praveen Kumar Menaka, et al.
Pubblicazione: (2026)
di: Sekar, Praveen Kumar Menaka, et al.
Pubblicazione: (2026)
VecTrans: Enhancing Compiler Auto-Vectorization through LLM-Assisted Code Transformations
di: Zheng, Zhongchun, et al.
Pubblicazione: (2025)
di: Zheng, Zhongchun, et al.
Pubblicazione: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
di: Xu, WeiZhe, et al.
Pubblicazione: (2026)
di: Xu, WeiZhe, et al.
Pubblicazione: (2026)
CAPE: Capability Achievement via Policy Execution
di: Ball, David
Pubblicazione: (2025)
di: Ball, David
Pubblicazione: (2025)
How Robustly do LLMs Understand Execution Semantics?
di: Spiess, Claudio, et al.
Pubblicazione: (2026)
di: Spiess, Claudio, et al.
Pubblicazione: (2026)
Feedback Over Form: Why Execution Feedback Matters More Than Pipeline Topology in 1-3B Code Generation
di: McAndrews, Charles Junichi
Pubblicazione: (2026)
di: McAndrews, Charles Junichi
Pubblicazione: (2026)
A Stochastic Differential Equation Framework for Multi-Objective LLM Interactions: Dynamical Systems Analysis with Code Generation Applications
di: Shukla, Shivani, et al.
Pubblicazione: (2025)
di: Shukla, Shivani, et al.
Pubblicazione: (2025)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
di: Li, Yuangang, et al.
Pubblicazione: (2026)
di: Li, Yuangang, et al.
Pubblicazione: (2026)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
di: Chen, Simin, et al.
Pubblicazione: (2025)
di: Chen, Simin, et al.
Pubblicazione: (2025)
Understanding LLM-Driven Test Oracle Generation
di: Bodicoat, Adam, et al.
Pubblicazione: (2026)
di: Bodicoat, Adam, et al.
Pubblicazione: (2026)
Mutation-Guided LLM-based Test Generation at Meta
di: Foster, Christopher, et al.
Pubblicazione: (2025)
di: Foster, Christopher, et al.
Pubblicazione: (2025)
Engineering LLM Powered Multi-agent Framework for Autonomous CloudOps
di: Parthasarathy, Kannan, et al.
Pubblicazione: (2025)
di: Parthasarathy, Kannan, et al.
Pubblicazione: (2025)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
di: Manglik, Akshay, et al.
Pubblicazione: (2026)
di: Manglik, Akshay, et al.
Pubblicazione: (2026)
Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution
di: Yu, Kai, et al.
Pubblicazione: (2026)
di: Yu, Kai, et al.
Pubblicazione: (2026)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
di: Diehl, Patrick, et al.
Pubblicazione: (2025)
di: Diehl, Patrick, et al.
Pubblicazione: (2025)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
di: Samsonau, Sergey V.
Pubblicazione: (2026)
di: Samsonau, Sergey V.
Pubblicazione: (2026)
How Robust are LLM-Generated Library Imports? An Empirical Study using Stack Overflow
di: Latendresse, Jasmine, et al.
Pubblicazione: (2025)
di: Latendresse, Jasmine, et al.
Pubblicazione: (2025)
Automated Unity Game Template Generation from GDDs via NLP and Multi-Modal LLMs
di: Hassan, Amna
Pubblicazione: (2025)
di: Hassan, Amna
Pubblicazione: (2025)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
di: Shah, Mehil B, et al.
Pubblicazione: (2025)
di: Shah, Mehil B, et al.
Pubblicazione: (2025)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
di: Shbita, Basel, et al.
Pubblicazione: (2025)
di: Shbita, Basel, et al.
Pubblicazione: (2025)
Generative AI to Generate Test Data Generators
di: Baudry, Benoit, et al.
Pubblicazione: (2024)
di: Baudry, Benoit, et al.
Pubblicazione: (2024)
Automatic Programming: Large Language Models and Beyond
di: Lyu, Michael R., et al.
Pubblicazione: (2024)
di: Lyu, Michael R., et al.
Pubblicazione: (2024)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
di: Wu, Yixi, et al.
Pubblicazione: (2024)
di: Wu, Yixi, et al.
Pubblicazione: (2024)
Automatic Detection of LLM-Generated Code: A Comparative Case Study of Contemporary Models Across Function and Class Granularities
di: Rahman, Musfiqur, et al.
Pubblicazione: (2024)
di: Rahman, Musfiqur, et al.
Pubblicazione: (2024)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
di: Rank, Ben, et al.
Pubblicazione: (2026)
di: Rank, Ben, et al.
Pubblicazione: (2026)
JARVIS: A Multi-Agent Code Assistant for High-Quality EDA Script Generation
di: Pasandi, Ghasem, et al.
Pubblicazione: (2025)
di: Pasandi, Ghasem, et al.
Pubblicazione: (2025)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
di: Bisztray, Tamas, et al.
Pubblicazione: (2025)
di: Bisztray, Tamas, et al.
Pubblicazione: (2025)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
di: Huang, Jiawei, et al.
Pubblicazione: (2026)
di: Huang, Jiawei, et al.
Pubblicazione: (2026)
The Dual-State Architecture for Reliable LLM Agents
di: Thompson, Matthew
Pubblicazione: (2025)
di: Thompson, Matthew
Pubblicazione: (2025)
PCodeTrans: Translate Decompiled Pseudocode to Compilable and Executable Equivalent
di: Cui, Yuxin, et al.
Pubblicazione: (2026)
di: Cui, Yuxin, et al.
Pubblicazione: (2026)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
di: Zhao, Zhimin, et al.
Pubblicazione: (2026)
di: Zhao, Zhimin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
di: Li, Xin-Ye, et al.
Pubblicazione: (2026) -
Grounding Machine Creativity in Game Design Knowledge Representations: Empirical Probing of LLM-Based Executable Synthesis of Goal Playable Patterns under Structural Constraints
di: Liu, Hugh Xuechen, et al.
Pubblicazione: (2026) -
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
di: Jiang, Shan, et al.
Pubblicazione: (2026) -
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025) -
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
di: Gu, Alex, et al.
Pubblicazione: (2024)