GameDevBench: Evaluating Agentic Capabilities Through Game Development
Fuente:
arXiv
Guardado en:
| Autores principales: | Chi, Wayne, Fang, Yixiong, Yayavaram, Arnav, Yayavaram, Siddharth, Karten, Seth, Wei, Qiuhong Anna, Chen, Runkun, Wang, Alexander, Chen, Valerie, Talwalkar, Ameet, Donahue, Chris |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
por: Yayavaram, Arnav, et al.
Publicado: (2025)
por: Yayavaram, Arnav, et al.
Publicado: (2025)
The Impact of Element Ordering on LM Agent Performance
por: Chi, Wayne, et al.
Publicado: (2024)
por: Chi, Wayne, et al.
Publicado: (2024)
Comparing Developer and LLM Biases in Code Evaluation
por: Mittal, Aditya, et al.
Publicado: (2026)
por: Mittal, Aditya, et al.
Publicado: (2026)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
por: Chi, Wayne, et al.
Publicado: (2025)
por: Chi, Wayne, et al.
Publicado: (2025)
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
por: Chi, Wayne, et al.
Publicado: (2025)
por: Chi, Wayne, et al.
Publicado: (2025)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
por: Chen, Valerie, et al.
Publicado: (2025)
por: Chen, Valerie, et al.
Publicado: (2025)
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
por: Yu, Rebecca, et al.
Publicado: (2025)
por: Yu, Rebecca, et al.
Publicado: (2025)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
por: Chen, Valerie, et al.
Publicado: (2026)
por: Chen, Valerie, et al.
Publicado: (2026)
How Gaming Could Improve Information Literacy
por: Doshi, Ameet
Publicado: (2006)
por: Doshi, Ameet
Publicado: (2006)
Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
por: Karten, Seth, et al.
Publicado: (2026)
por: Karten, Seth, et al.
Publicado: (2026)
Do LLMs exhibit human-like response biases? A case study in survey design
por: Tjuatja, Lindia, et al.
Publicado: (2023)
por: Tjuatja, Lindia, et al.
Publicado: (2023)
GameDevDojo -- An Educational Game for Teaching Game Development Concepts
por: Holly, Michael, et al.
Publicado: (2024)
por: Holly, Michael, et al.
Publicado: (2024)
GameTileNet: A Semantic Dataset for Low-Resolution Game Art in Procedural Content Generation
por: Chen, Yi-Chun, et al.
Publicado: (2025)
por: Chen, Yi-Chun, et al.
Publicado: (2025)
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
por: Shen, Junhong, et al.
Publicado: (2024)
por: Shen, Junhong, et al.
Publicado: (2024)
Need Help? Designing Proactive AI Assistants for Programming
por: Chen, Valerie, et al.
Publicado: (2024)
por: Chen, Valerie, et al.
Publicado: (2024)
CodingGenie: A Proactive LLM-Powered Programming Assistant
por: Zhao, Sebastian, et al.
Publicado: (2025)
por: Zhao, Sebastian, et al.
Publicado: (2025)
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
por: Pan, Jane, et al.
Publicado: (2025)
por: Pan, Jane, et al.
Publicado: (2025)
Automatic Generation of High-Performance RL Environments
por: Karten, Seth, et al.
Publicado: (2026)
por: Karten, Seth, et al.
Publicado: (2026)
PokéChamp: an Expert-level Minimax Language Agent
por: Karten, Seth, et al.
Publicado: (2025)
por: Karten, Seth, et al.
Publicado: (2025)
Audio Language Model for Deepfake Detection Grounded in Acoustic Chain-of-Thought
por: Chen, Runkun, et al.
Publicado: (2026)
por: Chen, Runkun, et al.
Publicado: (2026)
FightLadder: A Benchmark for Competitive Multi-Agent Reinforcement Learning
por: Li, Wenzhe, et al.
Publicado: (2024)
por: Li, Wenzhe, et al.
Publicado: (2024)
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
por: Chen, Yi-Chun, et al.
Publicado: (2025)
por: Chen, Yi-Chun, et al.
Publicado: (2025)
The Censorship "Game."
por: Mahood, Wayne
Publicado: (1983)
por: Mahood, Wayne
Publicado: (1983)
Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
por: Shi, Chengshuai, et al.
Publicado: (2026)
por: Shi, Chengshuai, et al.
Publicado: (2026)
Agreement-Based Cascading for Efficient Inference
por: Kolawole, Steven, et al.
Publicado: (2024)
por: Kolawole, Steven, et al.
Publicado: (2024)
Learning Personalized Decision Support Policies
por: Bhatt, Umang, et al.
Publicado: (2023)
por: Bhatt, Umang, et al.
Publicado: (2023)
Provably tuning the ElasticNet across instances
por: Balcan, Maria-Florina, et al.
Publicado: (2022)
por: Balcan, Maria-Florina, et al.
Publicado: (2022)
Learning to Relax: Setting Solver Parameters Across a Sequence of Linear System Instances
por: Khodak, Mikhail, et al.
Publicado: (2023)
por: Khodak, Mikhail, et al.
Publicado: (2023)
Where Does My Model Underperform? A Human Evaluation of Slice Discovery Algorithms
por: Johnson, Nari, et al.
Publicado: (2023)
por: Johnson, Nari, et al.
Publicado: (2023)
OpenGame: Open Agentic Coding for Games
por: Jiang, Yilei, et al.
Publicado: (2026)
por: Jiang, Yilei, et al.
Publicado: (2026)
RECODE: Reasoning Through Code Generation for Visual Question Answering
por: Shen, Junhong, et al.
Publicado: (2025)
por: Shen, Junhong, et al.
Publicado: (2025)
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
por: Li, Sijie, et al.
Publicado: (2025)
por: Li, Sijie, et al.
Publicado: (2025)
FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
por: Feng, Shengyu, et al.
Publicado: (2025)
por: Feng, Shengyu, et al.
Publicado: (2025)
Phase Transitions in Turnpike Theory For Mean-Field Games
por: Karuturi, Siddharth
Publicado: (2026)
por: Karuturi, Siddharth
Publicado: (2026)
Multi-Scale Feature Fusion Quantum Depthwise Convolutional Neural Networks for Text Classification
por: Chen, Yixiong, et al.
Publicado: (2024)
por: Chen, Yixiong, et al.
Publicado: (2024)
Addressing bias in Recommender Systems: A Case Study on Data Debiasing Techniques in Mobile Games
por: Wang, Yixiong, et al.
Publicado: (2024)
por: Wang, Yixiong, et al.
Publicado: (2024)
Rational Capability in Concurrent Games
por: Li, Yinfeng, et al.
Publicado: (2025)
por: Li, Yinfeng, et al.
Publicado: (2025)
LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
por: Karten, Seth, et al.
Publicado: (2025)
por: Karten, Seth, et al.
Publicado: (2025)
Impact of Decentralized Learning on Player Utilities in Stackelberg Games
por: Donahue, Kate, et al.
Publicado: (2024)
por: Donahue, Kate, et al.
Publicado: (2024)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
por: Lin, Jingru, et al.
Publicado: (2025)
por: Lin, Jingru, et al.
Publicado: (2025)
Ejemplares similares
-
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
por: Yayavaram, Arnav, et al.
Publicado: (2025) -
The Impact of Element Ordering on LM Agent Performance
por: Chi, Wayne, et al.
Publicado: (2024) -
Comparing Developer and LLM Biases in Code Evaluation
por: Mittal, Aditya, et al.
Publicado: (2026) -
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
por: Chi, Wayne, et al.
Publicado: (2025) -
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
por: Chi, Wayne, et al.
Publicado: (2025)