SmartPlay: A Benchmark for LLMs as Intelligent Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Yue, Tang, Xuan, Mitchell, Tom M., Li, Yuanzhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals
von: Wu, Yue, et al.
Veröffentlicht: (2023)
von: Wu, Yue, et al.
Veröffentlicht: (2023)
AgentKit: Structured LLM Reasoning with Dynamic Graphs
von: Wu, Yue, et al.
Veröffentlicht: (2024)
von: Wu, Yue, et al.
Veröffentlicht: (2024)
Learning to play: A Multimodal Agent for 3D Game-Play
von: Yue, Yuguang, et al.
Veröffentlicht: (2025)
von: Yue, Yuguang, et al.
Veröffentlicht: (2025)
Power Plays: Unleashing Machine Learning Magic in Smart Grids
von: Rashid, Abdur, et al.
Veröffentlicht: (2024)
von: Rashid, Abdur, et al.
Veröffentlicht: (2024)
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
von: Wang, Sai, et al.
Veröffentlicht: (2025)
von: Wang, Sai, et al.
Veröffentlicht: (2025)
Bidirectional Distillation: A Mixed-Play Framework for Multi-Agent Generalizable Behaviors
von: Feng, Lang, et al.
Veröffentlicht: (2025)
von: Feng, Lang, et al.
Veröffentlicht: (2025)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
von: Li, Manling, et al.
Veröffentlicht: (2024)
von: Li, Manling, et al.
Veröffentlicht: (2024)
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
von: De Brouwer, Edward, et al.
Veröffentlicht: (2026)
von: De Brouwer, Edward, et al.
Veröffentlicht: (2026)
2026 Roadmap on Artificial Intelligence and Machine Learning for Smart Manufacturing
von: Lee, Jay, et al.
Veröffentlicht: (2026)
von: Lee, Jay, et al.
Veröffentlicht: (2026)
Learning Game-Playing Agents with Generative Code Optimization
von: Kuang, Zhiyi, et al.
Veröffentlicht: (2025)
von: Kuang, Zhiyi, et al.
Veröffentlicht: (2025)
Self-Improving AI Agents through Self-Play
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
SmartBench: Evaluating LLMs in Smart Homes with Anomalous Device States and Behavioral Contexts
von: Zou, Qingsong, et al.
Veröffentlicht: (2026)
von: Zou, Qingsong, et al.
Veröffentlicht: (2026)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2024)
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2024)
Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
von: Wu, Skyler, et al.
Veröffentlicht: (2025)
von: Wu, Skyler, et al.
Veröffentlicht: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
von: Mohammadi, Mahmoud, et al.
Veröffentlicht: (2025)
von: Mohammadi, Mahmoud, et al.
Veröffentlicht: (2025)
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
von: Zhang, Xiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Xiyuan, et al.
Veröffentlicht: (2026)
Read to Play (R2-Play): Decision Transformer with Multimodal Game Instruction
von: Jin, Yonggang, et al.
Veröffentlicht: (2024)
von: Jin, Yonggang, et al.
Veröffentlicht: (2024)
CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
von: Zhu, Jiace, et al.
Veröffentlicht: (2026)
von: Zhu, Jiace, et al.
Veröffentlicht: (2026)
Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP
von: Chen, Zixiang, et al.
Veröffentlicht: (2023)
von: Chen, Zixiang, et al.
Veröffentlicht: (2023)
Physics of Language Models: Part 1, Learning Hierarchical Language Structures
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
Physics of Language Models: Part 3.2, Knowledge Manipulation
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
von: Chen, Zhaorun, et al.
Veröffentlicht: (2025)
von: Chen, Zhaorun, et al.
Veröffentlicht: (2025)
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024)
TextAtari: 100K Frames Game Playing with Language Agents
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
von: Xu, Zelai, et al.
Veröffentlicht: (2023)
von: Xu, Zelai, et al.
Veröffentlicht: (2023)
DataSciBench: An LLM Agent Benchmark for Data Science
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
Intelligence as Trajectory-Dominant Pareto Optimization
von: Khanh, Truong Xuan, et al.
Veröffentlicht: (2026)
von: Khanh, Truong Xuan, et al.
Veröffentlicht: (2026)
EVGeoQA: Benchmarking LLMs on Dynamic, Multi-Objective Geo-Spatial Exploration
von: Wu, Jianfei, et al.
Veröffentlicht: (2026)
von: Wu, Jianfei, et al.
Veröffentlicht: (2026)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
von: Huang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2025)
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
von: Rawles, Christopher, et al.
Veröffentlicht: (2024)
von: Rawles, Christopher, et al.
Veröffentlicht: (2024)
BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction
von: Zhou, Zikang, et al.
Veröffentlicht: (2024)
von: Zhou, Zikang, et al.
Veröffentlicht: (2024)
Selective Prompting Tuning for Personalized Conversations with LLMs
von: Huang, Qiushi, et al.
Veröffentlicht: (2024)
von: Huang, Qiushi, et al.
Veröffentlicht: (2024)
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts
von: Chung, Youngseog, et al.
Veröffentlicht: (2024)
von: Chung, Youngseog, et al.
Veröffentlicht: (2024)
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
von: Skelic, Lejla, et al.
Veröffentlicht: (2025)
von: Skelic, Lejla, et al.
Veröffentlicht: (2025)
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
von: Choudhury, Deepro, et al.
Veröffentlicht: (2025)
von: Choudhury, Deepro, et al.
Veröffentlicht: (2025)
TestAgent: An Adaptive and Intelligent Expert for Human Assessment
von: Yu, Junhao, et al.
Veröffentlicht: (2025)
von: Yu, Junhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals
von: Wu, Yue, et al.
Veröffentlicht: (2023) -
AgentKit: Structured LLM Reasoning with Dynamic Graphs
von: Wu, Yue, et al.
Veröffentlicht: (2024) -
Learning to play: A Multimodal Agent for 3D Game-Play
von: Yue, Yuguang, et al.
Veröffentlicht: (2025) -
Power Plays: Unleashing Machine Learning Magic in Smart Grids
von: Rashid, Abdur, et al.
Veröffentlicht: (2024) -
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
von: Wang, Sai, et al.
Veröffentlicht: (2025)