PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?
Fuente:
arXiv
Saved in:
| Main Authors: | Hua, Dongdong, Sun, Yifei, Huang, Renhong, Gao, Feng, Wang, Chunping, Yang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAT: RunAnyThing via Fully Automated Environment Configuration
by: Huang, Renhong, et al.
Published: (2026)
by: Huang, Renhong, et al.
Published: (2026)
From LLM-Driven Trading Card Generation to Procedural Relatedness: A Pokémon Case Study
by: Pfau, Johannes, et al.
Published: (2026)
by: Pfau, Johannes, et al.
Published: (2026)
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
by: Zheng, Qiaoyuan, et al.
Published: (2026)
by: Zheng, Qiaoyuan, et al.
Published: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémon
by: Angliss, Cameron, et al.
Published: (2025)
by: Angliss, Cameron, et al.
Published: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
$C^3$-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
by: Yu, Peijie, et al.
Published: (2025)
by: Yu, Peijie, et al.
Published: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models
by: Hu, Sihao, et al.
Published: (2024)
by: Hu, Sihao, et al.
Published: (2024)
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
by: Yu, Peijie, et al.
Published: (2025)
by: Yu, Peijie, et al.
Published: (2025)
Learning to Beat ByteRL: Exploitability of Collectible Card Game Agents
by: Haluska, Radovan, et al.
Published: (2024)
by: Haluska, Radovan, et al.
Published: (2024)
RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy Optimization
by: Zhang, Siwei, et al.
Published: (2026)
by: Zhang, Siwei, et al.
Published: (2026)
MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy Optimization
by: Huang, Renhong, et al.
Published: (2026)
by: Huang, Renhong, et al.
Published: (2026)
Experience Transfer for Multimodal LLM Agents in Minecraft Game
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
A Multi-Agent Pokemon Tournament for Evaluating Strategic Reasoning of Large Language Models
by: Yashwanth, Tadisetty Sai, et al.
Published: (2025)
by: Yashwanth, Tadisetty Sai, et al.
Published: (2025)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
by: Wu, Yunze, et al.
Published: (2025)
by: Wu, Yunze, et al.
Published: (2025)
TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
by: Yan, Lewen, et al.
Published: (2025)
by: Yan, Lewen, et al.
Published: (2025)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
by: Lu, Jiaxuan, et al.
Published: (2026)
by: Lu, Jiaxuan, et al.
Published: (2026)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026)
by: Rank, Ben, et al.
Published: (2026)
LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents
by: Xiao, Chang, et al.
Published: (2024)
by: Xiao, Chang, et al.
Published: (2024)
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
by: Jain, Daksh, et al.
Published: (2025)
by: Jain, Daksh, et al.
Published: (2025)
CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
ContractBench: Can LLM Agents Preserve Observation Contracts?
by: Wang, Jicheng, et al.
Published: (2026)
by: Wang, Jicheng, et al.
Published: (2026)
SkillMaster: Toward Autonomous Skill Mastery in LLM Agents
by: Yang, Min, et al.
Published: (2026)
by: Yang, Min, et al.
Published: (2026)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
by: Zhou, Yifan, et al.
Published: (2026)
by: Zhou, Yifan, et al.
Published: (2026)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
by: Liu, Wenrui, et al.
Published: (2025)
by: Liu, Wenrui, et al.
Published: (2025)
TradingAgents: Multi-Agents LLM Financial Trading Framework
by: Xiao, Yijia, et al.
Published: (2024)
by: Xiao, Yijia, et al.
Published: (2024)
UrzaGPT: LoRA-Tuned Large Language Models for Card Selection in Collectible Card Games
by: Bertram, Timo
Published: (2025)
by: Bertram, Timo
Published: (2025)
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
by: Wang, Qiyao, et al.
Published: (2026)
by: Wang, Qiyao, et al.
Published: (2026)
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
by: Jiang, Chenxi, et al.
Published: (2025)
by: Jiang, Chenxi, et al.
Published: (2025)
DeliveryBench: Can Agents Earn Profit in Real World?
by: Mao, Lingjun, et al.
Published: (2025)
by: Mao, Lingjun, et al.
Published: (2025)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
Can LLM Agents Sustain Long-Horizon Organizational Dynamics?
by: Zhu, Xuancheng, et al.
Published: (2026)
by: Zhu, Xuancheng, et al.
Published: (2026)
A Taxonomy of Collectible Card Games from a Game-Playing AI Perspective
by: Vieira, Ronaldo e Silva, et al.
Published: (2024)
by: Vieira, Ronaldo e Silva, et al.
Published: (2024)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
by: He, Jiawei, et al.
Published: (2026)
by: He, Jiawei, et al.
Published: (2026)
FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents
by: Li, Qizheng, et al.
Published: (2026)
by: Li, Qizheng, et al.
Published: (2026)
A Survey on LLM-powered Agents for Recommender Systems
by: Peng, Qiyao, et al.
Published: (2025)
by: Peng, Qiyao, et al.
Published: (2025)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
Similar Items
-
RAT: RunAnyThing via Fully Automated Environment Configuration
by: Huang, Renhong, et al.
Published: (2026) -
From LLM-Driven Trading Card Generation to Procedural Relatedness: A Pokémon Case Study
by: Pfau, Johannes, et al.
Published: (2026) -
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
by: Zheng, Qiaoyuan, et al.
Published: (2026) -
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026) -
VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémon
by: Angliss, Cameron, et al.
Published: (2025)