StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
Fuente:
arXiv
Guardado en:
| Autores principales: | Tan, Weihao, Jiang, Changjiu, Duan, Yu, Lei, Mingcong, Li, Jiageng, Hong, Yitian, Wang, Xinrun, An, Bo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter
por: Park, Junyeong, et al.
Publicado: (2025)
por: Park, Junyeong, et al.
Publicado: (2025)
JaxLife: An Open-Ended Agentic Simulator
por: Lu, Chris, et al.
Publicado: (2024)
por: Lu, Chris, et al.
Publicado: (2024)
Democratizing Game Modding with GenAI: A Case Study of StarCharM, a Stardew Valley Character Maker
por: Miralvand, Hamid Zand, et al.
Publicado: (2025)
por: Miralvand, Hamid Zand, et al.
Publicado: (2025)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
por: Tan, Weihao, et al.
Publicado: (2024)
por: Tan, Weihao, et al.
Publicado: (2024)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
por: Xie, Tianbao, et al.
Publicado: (2024)
por: Xie, Tianbao, et al.
Publicado: (2024)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
por: Li, Hengzhi, et al.
Publicado: (2025)
por: Li, Hengzhi, et al.
Publicado: (2025)
KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks
por: Duan, Zhangqi, et al.
Publicado: (2026)
por: Duan, Zhangqi, et al.
Publicado: (2026)
MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
por: Shaar, Shaden, et al.
Publicado: (2026)
por: Shaar, Shaden, et al.
Publicado: (2026)
CrafText Benchmark: Advancing Instruction Following in Complex Multimodal Open-Ended World
por: Volovikova, Zoya, et al.
Publicado: (2025)
por: Volovikova, Zoya, et al.
Publicado: (2025)
Grading Open‐Ended Questions Using LLMs and RAG
por: Jacobo Farray Rodríguez, et al.
Publicado: (2025)
por: Jacobo Farray Rodríguez, et al.
Publicado: (2025)
OpenEP: Open-Ended Future Event Prediction
por: Guan, Yong, et al.
Publicado: (2024)
por: Guan, Yong, et al.
Publicado: (2024)
Open-Ended Video Game Glitch Detection with Agentic Reasoning and Temporal Grounding
por: Zheng, Muyang, et al.
Publicado: (2026)
por: Zheng, Muyang, et al.
Publicado: (2026)
Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
por: Mannekote, Amogh, et al.
Publicado: (2024)
por: Mannekote, Amogh, et al.
Publicado: (2024)
How Intrinsic Motivation Underlies Embodied Open-Ended Behavior
por: Moreno-Bote, Rubén, et al.
Publicado: (2026)
por: Moreno-Bote, Rubén, et al.
Publicado: (2026)
Generating Planning Feedback for Open-Ended Programming Exercises with LLMs
por: Demirtaş, Mehmet Arif, et al.
Publicado: (2025)
por: Demirtaş, Mehmet Arif, et al.
Publicado: (2025)
Dojo: A Differentiable Physics Engine for Robotics
por: Howell, Taylor A., et al.
Publicado: (2022)
por: Howell, Taylor A., et al.
Publicado: (2022)
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
por: Xu, Xinrun, et al.
Publicado: (2025)
por: Xu, Xinrun, et al.
Publicado: (2025)
Quality Control in Open-Ended Crowdsourcing: A Survey
por: Chai, Lei, et al.
Publicado: (2024)
por: Chai, Lei, et al.
Publicado: (2024)
Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
por: Bo, Weihao, et al.
Publicado: (2025)
por: Bo, Weihao, et al.
Publicado: (2025)
On Creativity and Open-Endedness
por: Soros, L. B., et al.
Publicado: (2024)
por: Soros, L. B., et al.
Publicado: (2024)
The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation
por: Carlsson, Fredrik, et al.
Publicado: (2024)
por: Carlsson, Fredrik, et al.
Publicado: (2024)
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
por: Matthews, Michael, et al.
Publicado: (2024)
por: Matthews, Michael, et al.
Publicado: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning
por: Liu, Wanlong, et al.
Publicado: (2026)
por: Liu, Wanlong, et al.
Publicado: (2026)
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
por: Ma, Bolei, et al.
Publicado: (2025)
por: Ma, Bolei, et al.
Publicado: (2025)
LeakDojo: Decoding the Leakage Threats of RAG Systems
por: Zhang, Maosen, et al.
Publicado: (2026)
por: Zhang, Maosen, et al.
Publicado: (2026)
Enabling Adaptive Agent Training in Open-Ended Simulators by Targeting Diversity
por: Costales, Robby, et al.
Publicado: (2024)
por: Costales, Robby, et al.
Publicado: (2024)
ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents
por: Shao, Jie-Jing, et al.
Publicado: (2024)
por: Shao, Jie-Jing, et al.
Publicado: (2024)
MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
por: Qiang, Rushi, et al.
Publicado: (2025)
por: Qiang, Rushi, et al.
Publicado: (2025)
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
por: Jeune, Pierre Le, et al.
Publicado: (2026)
por: Jeune, Pierre Le, et al.
Publicado: (2026)
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
por: Zhao, Wei, et al.
Publicado: (2025)
por: Zhao, Wei, et al.
Publicado: (2025)
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
por: Yang, Saelyne, et al.
Publicado: (2026)
por: Yang, Saelyne, et al.
Publicado: (2026)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
por: Demchak, Nathaniel, et al.
Publicado: (2024)
por: Demchak, Nathaniel, et al.
Publicado: (2024)
Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
por: Bhatti, Hunzalah Hassan, et al.
Publicado: (2025)
por: Bhatti, Hunzalah Hassan, et al.
Publicado: (2025)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
por: Ivanov, Andrei, et al.
Publicado: (2025)
por: Ivanov, Andrei, et al.
Publicado: (2025)
BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds
por: Wang, Huayi, et al.
Publicado: (2025)
por: Wang, Huayi, et al.
Publicado: (2025)
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
por: Zhuo, Terry Yue, et al.
Publicado: (2025)
por: Zhuo, Terry Yue, et al.
Publicado: (2025)
Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams
por: Liu, Zewen, et al.
Publicado: (2026)
por: Liu, Zewen, et al.
Publicado: (2026)
O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL
por: Yao, Yi, et al.
Publicado: (2026)
por: Yao, Yi, et al.
Publicado: (2026)
CausalEvolve: Towards Open-Ended Discovery with Causal Scratchpad
por: Chen, Yongqiang, et al.
Publicado: (2026)
por: Chen, Yongqiang, et al.
Publicado: (2026)
Ejemplares similares
-
CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter
por: Park, Junyeong, et al.
Publicado: (2025) -
JaxLife: An Open-Ended Agentic Simulator
por: Lu, Chris, et al.
Publicado: (2024) -
Democratizing Game Modding with GenAI: A Case Study of StarCharM, a Stardew Valley Character Maker
por: Miralvand, Hamid Zand, et al.
Publicado: (2025) -
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
por: Tan, Weihao, et al.
Publicado: (2024) -
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
por: Xie, Tianbao, et al.
Publicado: (2024)