AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Ovezmyradov, Berdymyrat |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models
par: Tan, Chee Wei, et autres
Publié: (2026)
par: Tan, Chee Wei, et autres
Publié: (2026)
Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
par: Pant, Aakash, et autres
Publié: (2026)
par: Pant, Aakash, et autres
Publié: (2026)
Moonshine: Distilling Game Content Generators into Steerable Generative Models
par: Nie, Yuhe, et autres
Publié: (2024)
par: Nie, Yuhe, et autres
Publié: (2024)
Large Language Models for Water Distribution Systems Modeling and Decision-Making
par: Goldshtein, Yinon, et autres
Publié: (2025)
par: Goldshtein, Yinon, et autres
Publié: (2025)
A Prompt Engineering Approach and a Knowledge Graph based Framework for Tackling Legal Implications of Large Language Model Answers
par: Hannah, George, et autres
Publié: (2024)
par: Hannah, George, et autres
Publié: (2024)
Causal Reinforcement Learning for Complex Card Games: A Magic The Gathering Benchmark
par: Cunha, Cristiano da Costa, et autres
Publié: (2026)
par: Cunha, Cristiano da Costa, et autres
Publié: (2026)
Scaling Laws for State Dynamics in Large Language Models
par: Li, Jacob X, et autres
Publié: (2025)
par: Li, Jacob X, et autres
Publié: (2025)
Playing Games with My Heart: An Evaluation of AI Companion Apps
par: Rauh, Maribeth, et autres
Publié: (2026)
par: Rauh, Maribeth, et autres
Publié: (2026)
Collaborative AI Enhances Image Understanding in Materials Science
par: Yin, Ruoyan Avery, et autres
Publié: (2025)
par: Yin, Ruoyan Avery, et autres
Publié: (2025)
Large Language Models for Judicial Entity Extraction: A Comparative Study
par: Hussain, Atin Sakkeer, et autres
Publié: (2024)
par: Hussain, Atin Sakkeer, et autres
Publié: (2024)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
par: Zhu, Yuxuan, et autres
Publié: (2025)
par: Zhu, Yuxuan, et autres
Publié: (2025)
Combining Insights From Multiple Large Language Models Improves Diagnostic Accuracy
par: Barabucci, Gioele, et autres
Publié: (2024)
par: Barabucci, Gioele, et autres
Publié: (2024)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
par: Wu, Dekun, et autres
Publié: (2023)
par: Wu, Dekun, et autres
Publié: (2023)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
par: Hong, Yoosung
Publié: (2026)
par: Hong, Yoosung
Publié: (2026)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
par: Karpurapu, Shanthi, et autres
Publié: (2024)
par: Karpurapu, Shanthi, et autres
Publié: (2024)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
par: Ghandi, Taraneh, et autres
Publié: (2026)
par: Ghandi, Taraneh, et autres
Publié: (2026)
Open-TI: Open Traffic Intelligence with Augmented Language Model
par: Da, Longchao, et autres
Publié: (2023)
par: Da, Longchao, et autres
Publié: (2023)
It's not the Language Model, it's the Tool: Deterministic Mediation for Scientific Workflows
par: Adamidis, Marios, et autres
Publié: (2026)
par: Adamidis, Marios, et autres
Publié: (2026)
STLLM-DF: A Spatial-Temporal Large Language Model with Diffusion for Enhanced Multi-Mode Traffic System Forecasting
par: Shao, Zhiqi, et autres
Publié: (2024)
par: Shao, Zhiqi, et autres
Publié: (2024)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
par: Hasan, Mohammed Rakibul
Publié: (2026)
par: Hasan, Mohammed Rakibul
Publié: (2026)
CAD-Prompted Generative Models: A Pathway to Feasible and Novel Engineering Designs
par: Chong, Leah, et autres
Publié: (2024)
par: Chong, Leah, et autres
Publié: (2024)
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
par: Dong, Jia-Kai, et autres
Publié: (2025)
par: Dong, Jia-Kai, et autres
Publié: (2025)
Yahtzee: Reinforcement Learning Techniques for Stochastic Combinatorial Games
par: Pape, Nicholas A.
Publié: (2025)
par: Pape, Nicholas A.
Publié: (2025)
Design Structure Matrix Modularization with Large Language Models
par: Jiang, Shuo, et autres
Publié: (2026)
par: Jiang, Shuo, et autres
Publié: (2026)
Enhancing Multi-Criteria Decision Analysis with AI: Integrating Analytic Hierarchy Process and GPT-4 for Automated Decision Support
par: Svoboda, Igor, et autres
Publié: (2024)
par: Svoboda, Igor, et autres
Publié: (2024)
Agent-Aided Design for Dynamic CAD Models
par: Adler, Mitch, et autres
Publié: (2026)
par: Adler, Mitch, et autres
Publié: (2026)
Large Language Models for Combinatorial Optimization of Design Structure Matrix
par: Jiang, Shuo, et autres
Publié: (2025)
par: Jiang, Shuo, et autres
Publié: (2025)
Design Process of a Self Adaptive Smart Serious Games Ecosystem
par: Tao, X., et autres
Publié: (2025)
par: Tao, X., et autres
Publié: (2025)
Robust Agent Compensation (RAC): Teaching AI Agents to Compensate
par: Perera, Srinath, et autres
Publié: (2026)
par: Perera, Srinath, et autres
Publié: (2026)
ASD-Bench: A Four-Axis Comprehensive Benchmark of AI Models for Autism Spectrum Disorder
par: Singh, Shubhankit, et autres
Publié: (2026)
par: Singh, Shubhankit, et autres
Publié: (2026)
CellTypeAgent: Trustworthy cell type annotation with Large Language Models
par: Chen, Jiawen, et autres
Publié: (2025)
par: Chen, Jiawen, et autres
Publié: (2025)
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
par: Hooshyar, Danial, et autres
Publié: (2023)
par: Hooshyar, Danial, et autres
Publié: (2023)
Reinforcement of Explainability of ChatGPT Prompts by Embedding Breast Cancer Self-Screening Rules into AI Responses
par: Khan, Yousef, et autres
Publié: (2024)
par: Khan, Yousef, et autres
Publié: (2024)
Understanding Clinical Decision-Making in Traditional East Asian Medicine through Dimensionality Reduction: An Empirical Investigation
par: Bae, Hyojin, et autres
Publié: (2024)
par: Bae, Hyojin, et autres
Publié: (2024)
BLT: Can Large Language Models Handle Basic Legal Text?
par: Blair-Stanek, Andrew, et autres
Publié: (2023)
par: Blair-Stanek, Andrew, et autres
Publié: (2023)
How to Evaluate Medical AI
par: Kopanichuk, Ilia, et autres
Publié: (2025)
par: Kopanichuk, Ilia, et autres
Publié: (2025)
TelePlanNet: An AI-Driven Framework for Efficient Telecom Network Planning
par: Deng, Zongyuan, et autres
Publié: (2025)
par: Deng, Zongyuan, et autres
Publié: (2025)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
par: Karki, Siddhant, et autres
Publié: (2025)
par: Karki, Siddhant, et autres
Publié: (2025)
Inference-Time Intervention in Large Language Models for Reliable Requirement Verification
par: Darm, Paul, et autres
Publié: (2025)
par: Darm, Paul, et autres
Publié: (2025)
Smart Recommendations for Renting Bikes in Bike Sharing Systems
par: Billhardt, Holger, et autres
Publié: (2024)
par: Billhardt, Holger, et autres
Publié: (2024)
Documents similaires
-
Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models
par: Tan, Chee Wei, et autres
Publié: (2026) -
Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
par: Pant, Aakash, et autres
Publié: (2026) -
Moonshine: Distilling Game Content Generators into Steerable Generative Models
par: Nie, Yuhe, et autres
Publié: (2024) -
Large Language Models for Water Distribution Systems Modeling and Decision-Making
par: Goldshtein, Yinon, et autres
Publié: (2025) -
A Prompt Engineering Approach and a Knowledge Graph based Framework for Tackling Legal Implications of Large Language Model Answers
par: Hannah, George, et autres
Publié: (2024)