SmartPlay: A Benchmark for LLMs as Intelligent Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Yue, Tang, Xuan, Mitchell, Tom M., Li, Yuanzhi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals
por: Wu, Yue, et al.
Publicado: (2023)
por: Wu, Yue, et al.
Publicado: (2023)
AgentKit: Structured LLM Reasoning with Dynamic Graphs
por: Wu, Yue, et al.
Publicado: (2024)
por: Wu, Yue, et al.
Publicado: (2024)
Learning to play: A Multimodal Agent for 3D Game-Play
por: Yue, Yuguang, et al.
Publicado: (2025)
por: Yue, Yuguang, et al.
Publicado: (2025)
Power Plays: Unleashing Machine Learning Magic in Smart Grids
por: Rashid, Abdur, et al.
Publicado: (2024)
por: Rashid, Abdur, et al.
Publicado: (2024)
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
por: Wang, Sai, et al.
Publicado: (2025)
por: Wang, Sai, et al.
Publicado: (2025)
Bidirectional Distillation: A Mixed-Play Framework for Multi-Agent Generalizable Behaviors
por: Feng, Lang, et al.
Publicado: (2025)
por: Feng, Lang, et al.
Publicado: (2025)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
por: Li, Manling, et al.
Publicado: (2024)
por: Li, Manling, et al.
Publicado: (2024)
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
por: De Brouwer, Edward, et al.
Publicado: (2026)
por: De Brouwer, Edward, et al.
Publicado: (2026)
2026 Roadmap on Artificial Intelligence and Machine Learning for Smart Manufacturing
por: Lee, Jay, et al.
Publicado: (2026)
por: Lee, Jay, et al.
Publicado: (2026)
Learning Game-Playing Agents with Generative Code Optimization
por: Kuang, Zhiyi, et al.
Publicado: (2025)
por: Kuang, Zhiyi, et al.
Publicado: (2025)
Self-Improving AI Agents through Self-Play
por: Chojecki, Przemyslaw
Publicado: (2025)
por: Chojecki, Przemyslaw
Publicado: (2025)
SmartBench: Evaluating LLMs in Smart Homes with Anomalous Device States and Behavioral Contexts
por: Zou, Qingsong, et al.
Publicado: (2026)
por: Zou, Qingsong, et al.
Publicado: (2026)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
por: Davoodi, Arash Gholami, et al.
Publicado: (2024)
por: Davoodi, Arash Gholami, et al.
Publicado: (2024)
Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
por: Wu, Skyler, et al.
Publicado: (2025)
por: Wu, Skyler, et al.
Publicado: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
por: Zhang, Xiyuan, et al.
Publicado: (2026)
por: Zhang, Xiyuan, et al.
Publicado: (2026)
Read to Play (R2-Play): Decision Transformer with Multimodal Game Instruction
por: Jin, Yonggang, et al.
Publicado: (2024)
por: Jin, Yonggang, et al.
Publicado: (2024)
CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
por: Zhu, Jiace, et al.
Publicado: (2026)
por: Zhu, Jiace, et al.
Publicado: (2026)
Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP
por: Chen, Zixiang, et al.
Publicado: (2023)
por: Chen, Zixiang, et al.
Publicado: (2023)
Physics of Language Models: Part 1, Learning Hierarchical Language Structures
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2023)
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2023)
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2023)
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2023)
Physics of Language Models: Part 3.2, Knowledge Manipulation
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2023)
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2023)
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2024)
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2024)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
por: Chen, Zhaorun, et al.
Publicado: (2025)
por: Chen, Zhaorun, et al.
Publicado: (2025)
TextAtari: 100K Frames Game Playing with Language Agents
por: Li, Wenhao, et al.
Publicado: (2025)
por: Li, Wenhao, et al.
Publicado: (2025)
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
por: Xu, Zelai, et al.
Publicado: (2023)
por: Xu, Zelai, et al.
Publicado: (2023)
DataSciBench: An LLM Agent Benchmark for Data Science
por: Zhang, Dan, et al.
Publicado: (2025)
por: Zhang, Dan, et al.
Publicado: (2025)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
por: Tang, Xuan, et al.
Publicado: (2025)
por: Tang, Xuan, et al.
Publicado: (2025)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
por: Shen, Yuanzhe, et al.
Publicado: (2026)
por: Shen, Yuanzhe, et al.
Publicado: (2026)
Intelligence as Trajectory-Dominant Pareto Optimization
por: Khanh, Truong Xuan, et al.
Publicado: (2026)
por: Khanh, Truong Xuan, et al.
Publicado: (2026)
EVGeoQA: Benchmarking LLMs on Dynamic, Multi-Objective Geo-Spatial Exploration
por: Wu, Jianfei, et al.
Publicado: (2026)
por: Wu, Jianfei, et al.
Publicado: (2026)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
por: Huang, Kaixuan, et al.
Publicado: (2025)
por: Huang, Kaixuan, et al.
Publicado: (2025)
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
por: Rawles, Christopher, et al.
Publicado: (2024)
por: Rawles, Christopher, et al.
Publicado: (2024)
BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction
por: Zhou, Zikang, et al.
Publicado: (2024)
por: Zhou, Zikang, et al.
Publicado: (2024)
Selective Prompting Tuning for Personalized Conversations with LLMs
por: Huang, Qiushi, et al.
Publicado: (2024)
por: Huang, Qiushi, et al.
Publicado: (2024)
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts
por: Chung, Youngseog, et al.
Publicado: (2024)
por: Chung, Youngseog, et al.
Publicado: (2024)
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
por: He, Yu, et al.
Publicado: (2025)
por: He, Yu, et al.
Publicado: (2025)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
por: Skelic, Lejla, et al.
Publicado: (2025)
por: Skelic, Lejla, et al.
Publicado: (2025)
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
por: Choudhury, Deepro, et al.
Publicado: (2025)
por: Choudhury, Deepro, et al.
Publicado: (2025)
Self-Play Preference Optimization for Language Model Alignment
por: Wu, Yue, et al.
Publicado: (2024)
por: Wu, Yue, et al.
Publicado: (2024)
Ejemplares similares
-
Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals
por: Wu, Yue, et al.
Publicado: (2023) -
AgentKit: Structured LLM Reasoning with Dynamic Graphs
por: Wu, Yue, et al.
Publicado: (2024) -
Learning to play: A Multimodal Agent for 3D Game-Play
por: Yue, Yuguang, et al.
Publicado: (2025) -
Power Plays: Unleashing Machine Learning Magic in Smart Grids
por: Rashid, Abdur, et al.
Publicado: (2024) -
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
por: Wang, Sai, et al.
Publicado: (2025)