Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Dongmin, Kim, Minkyu, Choi, Beongjun, Kim, Junhyuck, Lee, Keon, Lee, Jonghyun, Park, Inkyu, Lee, Byeong-Uk, Hwang, Jaeyoung, Ahn, Jaewoo, Mahabaleshwarkar, Ameya S., Kartal, Bilal, Biswas, Pritam, Suhara, Yoshi, Lee, Kangwook, Cho, Jaewoong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
by: Park, Jaehyun, et al.
Published: (2026)
by: Park, Jaehyun, et al.
Published: (2026)
When2Call: When (not) to Call Tools
by: Ross, Hayley, et al.
Published: (2025)
by: Ross, Hayley, et al.
Published: (2025)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Raon-Speech Technical Report
by: Kim, Beomsoo, et al.
Published: (2026)
by: Kim, Beomsoo, et al.
Published: (2026)
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
by: Kim, Semin, et al.
Published: (2026)
by: Kim, Semin, et al.
Published: (2026)
Image Clustering Conditioned on Text Criteria
by: Kwon, Sehyun, et al.
Published: (2023)
by: Kwon, Sehyun, et al.
Published: (2023)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
by: Kim, Haechan, et al.
Published: (2026)
by: Kim, Haechan, et al.
Published: (2026)
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
by: Kim, Junhyuck, et al.
Published: (2024)
by: Kim, Junhyuck, et al.
Published: (2024)
Test-time Alignment of Diffusion Models without Reward Over-optimization
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
SAiD: Speech-driven Blendshape Facial Animation with Diffusion
by: Park, Inkyu, et al.
Published: (2023)
by: Park, Inkyu, et al.
Published: (2023)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
by: Kim, Jaehyeon, et al.
Published: (2024)
by: Kim, Jaehyeon, et al.
Published: (2024)
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
by: Kim, Jaehyeon, et al.
Published: (2024)
by: Kim, Jaehyeon, et al.
Published: (2024)
Normalized Convolutional Neural Network
by: Kim, Dongsuk, et al.
Published: (2020)
by: Kim, Dongsuk, et al.
Published: (2020)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
by: Park, Jongho, et al.
Published: (2024)
by: Park, Jongho, et al.
Published: (2024)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
by: Lee, Sanghyun, et al.
Published: (2025)
by: Lee, Sanghyun, et al.
Published: (2025)
Impact of Regularization on Calibration and Robustness: from the Representation Space Perspective
by: Park, Jonghyun, et al.
Published: (2024)
by: Park, Jonghyun, et al.
Published: (2024)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023)
by: Yang, Seongjun, et al.
Published: (2023)
A Delayed Acceptance Auxiliary Variable MCMC for Spatial Models with Intractable Likelihood Function
by: Lee, Jong Hyeon, et al.
Published: (2025)
by: Lee, Jong Hyeon, et al.
Published: (2025)
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks
by: Heo, Chiyeong, et al.
Published: (2026)
by: Heo, Chiyeong, et al.
Published: (2026)
Learning to Control Camera Exposure via Reinforcement Learning
by: Lee, Kyunghyun, et al.
Published: (2024)
by: Lee, Kyunghyun, et al.
Published: (2024)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
by: Lee, Sanghyun, et al.
Published: (2025)
by: Lee, Sanghyun, et al.
Published: (2025)
VR-Pipe: Streamlining Hardware Graphics Pipeline for Volume Rendering
by: Lee, Junseo, et al.
Published: (2025)
by: Lee, Junseo, et al.
Published: (2025)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
by: Lee, Keon, et al.
Published: (2024)
by: Lee, Keon, et al.
Published: (2024)
Bioelectrosynthesis of Signaling Molecules for Selective Modulation of Cell Signaling
by: Myeongeun Lee, et al.
Published: (2025)
by: Myeongeun Lee, et al.
Published: (2025)
Challenges and Lessons from MIDOG 2025: A Two-Stage Approach to Domain-Robust Mitotic Figure Detection
by: Song, Euiseop, et al.
Published: (2025)
by: Song, Euiseop, et al.
Published: (2025)
Source Identification in Abstractive Summarization
by: Suhara, Yoshi, et al.
Published: (2024)
by: Suhara, Yoshi, et al.
Published: (2024)
A Stein Gradient Descent Approach for Doubly Intractable Distributions
by: Lee, Heesang, et al.
Published: (2024)
by: Lee, Heesang, et al.
Published: (2024)
Transparent Networks for Multivariate Time Series
by: Kim, Minkyu, et al.
Published: (2024)
by: Kim, Minkyu, et al.
Published: (2024)
KMTNet View of Blue Large-amplitude Pulsators Toward the Galactic Bulge: I. Discovery of Wide-orbit Companions in OGLE-BLAP-006
by: Kim, Seung-Lee, et al.
Published: (2025)
by: Kim, Seung-Lee, et al.
Published: (2025)
Active Learning for Continual Learning: Keeping the Past Alive in the Present
by: Park, Jaehyun, et al.
Published: (2025)
by: Park, Jaehyun, et al.
Published: (2025)
TRACE: Your Diffusion Model is Secretly an Instance Edge Detector
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model Agents
by: Jeong, Jongwon, et al.
Published: (2026)
by: Jeong, Jongwon, et al.
Published: (2026)
Overlap Weights for Binary Outcomes: A Performance Assessment
by: Seo Young Park, et al.
Published: (2025)
by: Seo Young Park, et al.
Published: (2025)
Adversarial Reinforcement Learning Framework for ESP Cheater Simulation
by: Park, Inkyu, et al.
Published: (2025)
by: Park, Inkyu, et al.
Published: (2025)
How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects
by: Lee, Wonkwang, et al.
Published: (2025)
by: Lee, Wonkwang, et al.
Published: (2025)
Multiplicative Thom-Sebastiani for Bernstein-Sato polynomials
by: Lee, Jonghyun
Published: (2024)
by: Lee, Jonghyun
Published: (2024)
RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen Objects
by: Kim, Jaeguk, et al.
Published: (2025)
by: Kim, Jaeguk, et al.
Published: (2025)
Battery‐Free, Wireless Multi‐Modal Sensor, and Actuator Array System for Pressure Injury Prevention
by: Hyeonseok Han, et al.
Published: (2024)
by: Hyeonseok Han, et al.
Published: (2024)
Similar Items
-
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024) -
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
by: Park, Jaehyun, et al.
Published: (2026) -
When2Call: When (not) to Call Tools
by: Ross, Hayley, et al.
Published: (2025) -
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026) -
Raon-Speech Technical Report
by: Kim, Beomsoo, et al.
Published: (2026)