The Othello AI Arena: Evaluating Intelligent Systems Through Limited-Time Adaptation to Unseen Boards
Fuente:
arXiv
Saved in:
| Main Author: | Kim, Sundong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
System 2 Reasoning for Human-AI Alignment: Generality and Adaptivity via ARC-AGI
by: Kim, Sejin, et al.
Published: (2024)
by: Kim, Sejin, et al.
Published: (2024)
Othello is Solved
by: Takizawa, Hiroki
Published: (2023)
by: Takizawa, Hiroki
Published: (2023)
GIFARC: Synthetic Dataset for Leveraging Human-Intuitive Analogies to Elevate AI Reasoning
by: Sim, Woochang, et al.
Published: (2025)
by: Sim, Woochang, et al.
Published: (2025)
Addressing and Visualizing Misalignments in Human Task-Solving Trajectories
by: Kim, Sejin, et al.
Published: (2024)
by: Kim, Sejin, et al.
Published: (2024)
What if Othello-Playing Language Models Could See?
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
Causal-Paced Deep Reinforcement Learning
by: Cho, Geonwoo, et al.
Published: (2025)
by: Cho, Geonwoo, et al.
Published: (2025)
Abductive Symbolic Solver on Abstraction and Reasoning Corpus
by: Lim, Mintaek, et al.
Published: (2024)
by: Lim, Mintaek, et al.
Published: (2024)
Enhancing Analogical Reasoning in the Abstraction and Reasoning Corpus via Model-Based RL
by: Lee, Jihwan, et al.
Published: (2024)
by: Lee, Jihwan, et al.
Published: (2024)
Automatically Finding Rule-Based Neurons in OthelloGPT
by: Singh, Aditya, et al.
Published: (2025)
by: Singh, Aditya, et al.
Published: (2025)
Can Large Language Models Develop Gambling Addiction?
by: Lee, Seungpil, et al.
Published: (2025)
by: Lee, Seungpil, et al.
Published: (2025)
TextME: Bridging Unseen Modalities Through Text Descriptions
by: Hong, Soyeon, et al.
Published: (2026)
by: Hong, Soyeon, et al.
Published: (2026)
ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
by: Kim, Sejin, et al.
Published: (2025)
by: Kim, Sejin, et al.
Published: (2025)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
TRACED: Transition-aware Regret Approximation with Co-learnability for Environment Design
by: Cho, Geonwoo, et al.
Published: (2025)
by: Cho, Geonwoo, et al.
Published: (2025)
Music Arena: Live Evaluation for Text-to-Music
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
Unseen Fake News Detection Through Casual Debiasing
by: Gong, Shuzhi, et al.
Published: (2025)
by: Gong, Shuzhi, et al.
Published: (2025)
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
by: Lee, Kwanhee, et al.
Published: (2025)
by: Lee, Kwanhee, et al.
Published: (2025)
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023)
by: Wu, Ti-Rong, et al.
Published: (2023)
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
AMPED: Adaptive Multi-objective Projection for balancing Exploration and skill Diversification
by: Cho, Geonwoo, et al.
Published: (2025)
by: Cho, Geonwoo, et al.
Published: (2025)
SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
by: Li, Changyi, et al.
Published: (2026)
by: Li, Changyi, et al.
Published: (2026)
Development, Evaluation, and Deployment of a Multi-Agent System for Thoracic Tumor Board
by: Ellis-Caleo, Tim, et al.
Published: (2026)
by: Ellis-Caleo, Tim, et al.
Published: (2026)
A Methodology for Thermal Limit Bias Predictability Through Artificial Intelligence
by: Tunga, Anirudh, et al.
Published: (2026)
by: Tunga, Anirudh, et al.
Published: (2026)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
by: Goffinet, Etienne, et al.
Published: (2025)
by: Goffinet, Etienne, et al.
Published: (2025)
mOthello: When Do Cross-Lingual Representation Alignment and Cross-Lingual Transfer Emerge in Multilingual Models?
by: Hua, Tianze, et al.
Published: (2024)
by: Hua, Tianze, et al.
Published: (2024)
ARCLE: The Abstraction and Reasoning Corpus Learning Environment for Reinforcement Learning
by: Lee, Hosung, et al.
Published: (2024)
by: Lee, Hosung, et al.
Published: (2024)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
by: Kim, Yunho, et al.
Published: (2024)
by: Kim, Yunho, et al.
Published: (2024)
ATLANTIS: AI-driven Threat Localization, Analysis, and Triage Intelligence System
by: Kim, Taesoo, et al.
Published: (2025)
by: Kim, Taesoo, et al.
Published: (2025)
Evaluating Artificial Intelligence Through a Christian Understanding of Human Flourishing
by: Skytland, Nicholas, et al.
Published: (2026)
by: Skytland, Nicholas, et al.
Published: (2026)
RecSys Arena: Pair-wise Recommender System Evaluation with Large Language Models
by: Wu, Zhuo, et al.
Published: (2024)
by: Wu, Zhuo, et al.
Published: (2024)
Accelerating AI Development with Cyber Arenas
by: Cashman, William, et al.
Published: (2025)
by: Cashman, William, et al.
Published: (2025)
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence
by: Liu, Peigen, et al.
Published: (2026)
by: Liu, Peigen, et al.
Published: (2026)
Amplifying Human Creativity and Problem Solving with AI Through Generative Collective Intelligence
by: Kehler, Thomas P., et al.
Published: (2025)
by: Kehler, Thomas P., et al.
Published: (2025)
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
by: Bonatti, Rogerio, et al.
Published: (2024)
by: Bonatti, Rogerio, et al.
Published: (2024)
Generating Unseen Code Tests In Infinitum
by: Zalmanovici, Marcel, et al.
Published: (2024)
by: Zalmanovici, Marcel, et al.
Published: (2024)
InterPol: De-anonymizing LM Arena via Interpolated Preference Learning
by: Cho, Minsung, et al.
Published: (2026)
by: Cho, Minsung, et al.
Published: (2026)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
by: Wang, Daoyu, et al.
Published: (2025)
by: Wang, Daoyu, et al.
Published: (2025)
Similar Items
-
System 2 Reasoning for Human-AI Alignment: Generality and Adaptivity via ARC-AGI
by: Kim, Sejin, et al.
Published: (2024) -
Othello is Solved
by: Takizawa, Hiroki
Published: (2023) -
GIFARC: Synthetic Dataset for Leveraging Human-Intuitive Analogies to Elevate AI Reasoning
by: Sim, Woochang, et al.
Published: (2025) -
Addressing and Visualizing Misalignments in Human Task-Solving Trajectories
by: Kim, Sejin, et al.
Published: (2024) -
What if Othello-Playing Language Models Could See?
by: Chen, Xinyi, et al.
Published: (2025)