AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations
Fuente:
arXiv
Salvato in:
| Autore principale: | Ovezmyradov, Berdymyrat |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models
di: Tan, Chee Wei, et al.
Pubblicazione: (2026)
di: Tan, Chee Wei, et al.
Pubblicazione: (2026)
Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
di: Pant, Aakash, et al.
Pubblicazione: (2026)
di: Pant, Aakash, et al.
Pubblicazione: (2026)
Moonshine: Distilling Game Content Generators into Steerable Generative Models
di: Nie, Yuhe, et al.
Pubblicazione: (2024)
di: Nie, Yuhe, et al.
Pubblicazione: (2024)
Large Language Models for Water Distribution Systems Modeling and Decision-Making
di: Goldshtein, Yinon, et al.
Pubblicazione: (2025)
di: Goldshtein, Yinon, et al.
Pubblicazione: (2025)
A Prompt Engineering Approach and a Knowledge Graph based Framework for Tackling Legal Implications of Large Language Model Answers
di: Hannah, George, et al.
Pubblicazione: (2024)
di: Hannah, George, et al.
Pubblicazione: (2024)
Causal Reinforcement Learning for Complex Card Games: A Magic The Gathering Benchmark
di: Cunha, Cristiano da Costa, et al.
Pubblicazione: (2026)
di: Cunha, Cristiano da Costa, et al.
Pubblicazione: (2026)
Scaling Laws for State Dynamics in Large Language Models
di: Li, Jacob X, et al.
Pubblicazione: (2025)
di: Li, Jacob X, et al.
Pubblicazione: (2025)
Playing Games with My Heart: An Evaluation of AI Companion Apps
di: Rauh, Maribeth, et al.
Pubblicazione: (2026)
di: Rauh, Maribeth, et al.
Pubblicazione: (2026)
Collaborative AI Enhances Image Understanding in Materials Science
di: Yin, Ruoyan Avery, et al.
Pubblicazione: (2025)
di: Yin, Ruoyan Avery, et al.
Pubblicazione: (2025)
Large Language Models for Judicial Entity Extraction: A Comparative Study
di: Hussain, Atin Sakkeer, et al.
Pubblicazione: (2024)
di: Hussain, Atin Sakkeer, et al.
Pubblicazione: (2024)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
Combining Insights From Multiple Large Language Models Improves Diagnostic Accuracy
di: Barabucci, Gioele, et al.
Pubblicazione: (2024)
di: Barabucci, Gioele, et al.
Pubblicazione: (2024)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
di: Wu, Dekun, et al.
Pubblicazione: (2023)
di: Wu, Dekun, et al.
Pubblicazione: (2023)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
di: Hong, Yoosung
Pubblicazione: (2026)
di: Hong, Yoosung
Pubblicazione: (2026)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
di: Karpurapu, Shanthi, et al.
Pubblicazione: (2024)
di: Karpurapu, Shanthi, et al.
Pubblicazione: (2024)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
di: Ghandi, Taraneh, et al.
Pubblicazione: (2026)
di: Ghandi, Taraneh, et al.
Pubblicazione: (2026)
Open-TI: Open Traffic Intelligence with Augmented Language Model
di: Da, Longchao, et al.
Pubblicazione: (2023)
di: Da, Longchao, et al.
Pubblicazione: (2023)
It's not the Language Model, it's the Tool: Deterministic Mediation for Scientific Workflows
di: Adamidis, Marios, et al.
Pubblicazione: (2026)
di: Adamidis, Marios, et al.
Pubblicazione: (2026)
STLLM-DF: A Spatial-Temporal Large Language Model with Diffusion for Enhanced Multi-Mode Traffic System Forecasting
di: Shao, Zhiqi, et al.
Pubblicazione: (2024)
di: Shao, Zhiqi, et al.
Pubblicazione: (2024)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
di: Hasan, Mohammed Rakibul
Pubblicazione: (2026)
di: Hasan, Mohammed Rakibul
Pubblicazione: (2026)
CAD-Prompted Generative Models: A Pathway to Feasible and Novel Engineering Designs
di: Chong, Leah, et al.
Pubblicazione: (2024)
di: Chong, Leah, et al.
Pubblicazione: (2024)
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
di: Dong, Jia-Kai, et al.
Pubblicazione: (2025)
di: Dong, Jia-Kai, et al.
Pubblicazione: (2025)
Yahtzee: Reinforcement Learning Techniques for Stochastic Combinatorial Games
di: Pape, Nicholas A.
Pubblicazione: (2025)
di: Pape, Nicholas A.
Pubblicazione: (2025)
Design Structure Matrix Modularization with Large Language Models
di: Jiang, Shuo, et al.
Pubblicazione: (2026)
di: Jiang, Shuo, et al.
Pubblicazione: (2026)
Enhancing Multi-Criteria Decision Analysis with AI: Integrating Analytic Hierarchy Process and GPT-4 for Automated Decision Support
di: Svoboda, Igor, et al.
Pubblicazione: (2024)
di: Svoboda, Igor, et al.
Pubblicazione: (2024)
Agent-Aided Design for Dynamic CAD Models
di: Adler, Mitch, et al.
Pubblicazione: (2026)
di: Adler, Mitch, et al.
Pubblicazione: (2026)
Large Language Models for Combinatorial Optimization of Design Structure Matrix
di: Jiang, Shuo, et al.
Pubblicazione: (2025)
di: Jiang, Shuo, et al.
Pubblicazione: (2025)
Design Process of a Self Adaptive Smart Serious Games Ecosystem
di: Tao, X., et al.
Pubblicazione: (2025)
di: Tao, X., et al.
Pubblicazione: (2025)
Robust Agent Compensation (RAC): Teaching AI Agents to Compensate
di: Perera, Srinath, et al.
Pubblicazione: (2026)
di: Perera, Srinath, et al.
Pubblicazione: (2026)
ASD-Bench: A Four-Axis Comprehensive Benchmark of AI Models for Autism Spectrum Disorder
di: Singh, Shubhankit, et al.
Pubblicazione: (2026)
di: Singh, Shubhankit, et al.
Pubblicazione: (2026)
CellTypeAgent: Trustworthy cell type annotation with Large Language Models
di: Chen, Jiawen, et al.
Pubblicazione: (2025)
di: Chen, Jiawen, et al.
Pubblicazione: (2025)
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
di: Hooshyar, Danial, et al.
Pubblicazione: (2023)
di: Hooshyar, Danial, et al.
Pubblicazione: (2023)
Reinforcement of Explainability of ChatGPT Prompts by Embedding Breast Cancer Self-Screening Rules into AI Responses
di: Khan, Yousef, et al.
Pubblicazione: (2024)
di: Khan, Yousef, et al.
Pubblicazione: (2024)
Understanding Clinical Decision-Making in Traditional East Asian Medicine through Dimensionality Reduction: An Empirical Investigation
di: Bae, Hyojin, et al.
Pubblicazione: (2024)
di: Bae, Hyojin, et al.
Pubblicazione: (2024)
BLT: Can Large Language Models Handle Basic Legal Text?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
How to Evaluate Medical AI
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025)
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025)
TelePlanNet: An AI-Driven Framework for Efficient Telecom Network Planning
di: Deng, Zongyuan, et al.
Pubblicazione: (2025)
di: Deng, Zongyuan, et al.
Pubblicazione: (2025)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
di: Karki, Siddhant, et al.
Pubblicazione: (2025)
di: Karki, Siddhant, et al.
Pubblicazione: (2025)
Inference-Time Intervention in Large Language Models for Reliable Requirement Verification
di: Darm, Paul, et al.
Pubblicazione: (2025)
di: Darm, Paul, et al.
Pubblicazione: (2025)
Smart Recommendations for Renting Bikes in Bike Sharing Systems
di: Billhardt, Holger, et al.
Pubblicazione: (2024)
di: Billhardt, Holger, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models
di: Tan, Chee Wei, et al.
Pubblicazione: (2026) -
Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
di: Pant, Aakash, et al.
Pubblicazione: (2026) -
Moonshine: Distilling Game Content Generators into Steerable Generative Models
di: Nie, Yuhe, et al.
Pubblicazione: (2024) -
Large Language Models for Water Distribution Systems Modeling and Decision-Making
di: Goldshtein, Yinon, et al.
Pubblicazione: (2025) -
A Prompt Engineering Approach and a Knowledge Graph based Framework for Tackling Legal Implications of Large Language Model Answers
di: Hannah, George, et al.
Pubblicazione: (2024)