Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
Fuente:
arXiv
Saved in:
| Main Authors: | Isaza-Giraldo, Andrés, Bala, Paulo, Pereira, Lucas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025)
by: McCullum, Lucas, et al.
Published: (2025)
Flowco: Rethinking Data Analysis in the Age of LLMs
by: Freund, Stephen N., et al.
Published: (2025)
by: Freund, Stephen N., et al.
Published: (2025)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Evaluating LLMs as Human Surrogates in Controlled Experiments
by: Hoq, Adnan, et al.
Published: (2026)
by: Hoq, Adnan, et al.
Published: (2026)
Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf Game
by: Du, Silin, et al.
Published: (2024)
by: Du, Silin, et al.
Published: (2024)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
by: Liu, Hongtao, et al.
Published: (2025)
by: Liu, Hongtao, et al.
Published: (2025)
Hacc-Man: An Arcade Game for Jailbreaking LLMs
by: Valentim, Matheus, et al.
Published: (2024)
by: Valentim, Matheus, et al.
Published: (2024)
Game Plot Design with an LLM-powered Assistant: An Empirical Study with Game Designers
by: Alavi, Seyed Hossein, et al.
Published: (2024)
by: Alavi, Seyed Hossein, et al.
Published: (2024)
Game Development as Human-LLM Interaction
by: Hong, Jiale, et al.
Published: (2024)
by: Hong, Jiale, et al.
Published: (2024)
Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases
by: Ford, Casey, et al.
Published: (2026)
by: Ford, Casey, et al.
Published: (2026)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
Large Language Models and Games: A Survey and Roadmap
by: Gallotta, Roberto, et al.
Published: (2024)
by: Gallotta, Roberto, et al.
Published: (2024)
Playing 20 Question Game with Policy-Based Reinforcement Learning
by: Hu, Huang, et al.
Published: (2018)
by: Hu, Huang, et al.
Published: (2018)
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
by: Saakyan, Arkadiy, et al.
Published: (2025)
by: Saakyan, Arkadiy, et al.
Published: (2025)
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
Automatic Bug Detection in LLM-Powered Text-Based Games Using LLMs
by: Jin, Claire, et al.
Published: (2024)
by: Jin, Claire, et al.
Published: (2024)
HICode: Hierarchical Inductive Coding with LLMs
by: Zhong, Mian, et al.
Published: (2025)
by: Zhong, Mian, et al.
Published: (2025)
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
by: Li, Shuqing, et al.
Published: (2025)
by: Li, Shuqing, et al.
Published: (2025)
Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game
by: Ye, Rong, et al.
Published: (2025)
by: Ye, Rong, et al.
Published: (2025)
The BS-meter: A ChatGPT-Trained Instrument to Detect Sloppy Language-Games
by: Trevisan, Alessandro, et al.
Published: (2024)
by: Trevisan, Alessandro, et al.
Published: (2024)
DreamGarden: A Designer Assistant for Growing Games from a Single Prompt
by: Earle, Sam, et al.
Published: (2024)
by: Earle, Sam, et al.
Published: (2024)
Offscript: Automated Auditing of Instruction Adherence in LLMs
by: Clark, Nicholas, et al.
Published: (2025)
by: Clark, Nicholas, et al.
Published: (2025)
Can LLMs Generate Visualizations with Dataless Prompts?
by: Coelho, Darius, et al.
Published: (2024)
by: Coelho, Darius, et al.
Published: (2024)
Aligning LLMs with Individual Preferences via Interaction
by: Wu, Shujin, et al.
Published: (2024)
by: Wu, Shujin, et al.
Published: (2024)
Are Today's LLMs Ready to Explain Well-Being Concepts?
by: Jiang, Bohan, et al.
Published: (2025)
by: Jiang, Bohan, et al.
Published: (2025)
Clinical knowledge in LLMs does not translate to human interactions
by: Bean, Andrew M., et al.
Published: (2025)
by: Bean, Andrew M., et al.
Published: (2025)
AutoLife: Automatic Life Journaling with Smartphones and LLMs
by: Xu, Huatao, et al.
Published: (2024)
by: Xu, Huatao, et al.
Published: (2024)
HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?
by: Ji, Sijie, et al.
Published: (2024)
by: Ji, Sijie, et al.
Published: (2024)
Direct Advantage Regression: Aligning LLMs with Online AI Reward
by: He, Li, et al.
Published: (2025)
by: He, Li, et al.
Published: (2025)
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
by: Zhu, Shaojie, et al.
Published: (2023)
by: Zhu, Shaojie, et al.
Published: (2023)
AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers
by: Wuttke, Alexander, et al.
Published: (2024)
by: Wuttke, Alexander, et al.
Published: (2024)
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024)
by: Shaikh, Ammar, et al.
Published: (2024)
Copiloting Diagnosis of Autism in Real Clinical Scenarios via LLMs
by: Jiang, Yi, et al.
Published: (2024)
by: Jiang, Yi, et al.
Published: (2024)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
by: Park, Jung In, et al.
Published: (2024)
by: Park, Jung In, et al.
Published: (2024)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
RAG-based EEG-to-Text Translation Using Deep Learning and LLMs
by: Collautti, Enrico, et al.
Published: (2026)
by: Collautti, Enrico, et al.
Published: (2026)
Similar Items
-
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025) -
Flowco: Rethinking Data Analysis in the Age of LLMs
by: Freund, Stephen N., et al.
Published: (2025) -
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
by: Yadav, Neemesh, et al.
Published: (2025) -
Evaluating LLMs as Human Surrogates in Controlled Experiments
by: Hoq, Adnan, et al.
Published: (2026) -
Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf Game
by: Du, Silin, et al.
Published: (2024)