GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yunzhe, Xu, Runhui, Zheng, Kexin, Zhang, Tianyi, Kogundi, Jayavibhav Niranjan, Hans, Soham, Ustun, Volkan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
Abstracting Geo-specific Terrains to Scale Up Reinforcement Learning
von: Ustun, Volkan, et al.
Veröffentlicht: (2025)
von: Ustun, Volkan, et al.
Veröffentlicht: (2025)
A Data-Driven Discretized CS:GO Simulation Environment to Facilitate Strategic Multi-Agent Planning Research
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026)
von: Cao, Meng, et al.
Veröffentlicht: (2026)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
Spontaneous Theory of Mind for Artificial Intelligence
von: Gurney, Nikolos, et al.
Veröffentlicht: (2024)
von: Gurney, Nikolos, et al.
Veröffentlicht: (2024)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
SyncVIS: Synchronized Video Instance Segmentation
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
von: Krojer, Benno, et al.
Veröffentlicht: (2025)
von: Krojer, Benno, et al.
Veröffentlicht: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
DeVAn: Dense Video Annotation for Video-Language Models
von: Liu, Tingkai, et al.
Veröffentlicht: (2023)
von: Liu, Tingkai, et al.
Veröffentlicht: (2023)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
von: Weck, Benno, et al.
Veröffentlicht: (2026)
von: Weck, Benno, et al.
Veröffentlicht: (2026)
AdsQA: Towards Advertisement Video Understanding
von: Long, Xinwei, et al.
Veröffentlicht: (2025)
von: Long, Xinwei, et al.
Veröffentlicht: (2025)
The Emotional Impact of Game Duration: A Framework for Understanding Player Emotions in Extended Gameplay Sessions
von: Kumar, Anoop, et al.
Veröffentlicht: (2024)
von: Kumar, Anoop, et al.
Veröffentlicht: (2024)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay
von: Lan, Yihuai, et al.
Veröffentlicht: (2023)
von: Lan, Yihuai, et al.
Veröffentlicht: (2023)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
Detecting AI Assistance in Abstract Complex Tasks
von: King, Tyler, et al.
Veröffentlicht: (2025)
von: King, Tyler, et al.
Veröffentlicht: (2025)
Audio-Sync Video Generation with Multi-Stream Temporal Control
von: Weng, Shuchen, et al.
Veröffentlicht: (2025)
von: Weng, Shuchen, et al.
Veröffentlicht: (2025)
Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge
von: Sun, Yuyang, et al.
Veröffentlicht: (2026)
von: Sun, Yuyang, et al.
Veröffentlicht: (2026)
Gameplay Highlights Generation
von: Edithal, Vignesh, et al.
Veröffentlicht: (2025)
von: Edithal, Vignesh, et al.
Veröffentlicht: (2025)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
von: Zhou, Zhiyu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhiyu, et al.
Veröffentlicht: (2026)
SportQA: A Benchmark for Sports Understanding in Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
von: Guo, Xuehang, et al.
Veröffentlicht: (2025)
von: Guo, Xuehang, et al.
Veröffentlicht: (2025)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning
von: Jiang, Zhiheng, et al.
Veröffentlicht: (2026)
von: Jiang, Zhiheng, et al.
Veröffentlicht: (2026)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
Understanding Complexity in VideoQA via Visual Program Generation
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
ChartAB: A Benchmark for Chart Grounding & Dense Alignment
von: Bansal, Aniruddh, et al.
Veröffentlicht: (2025)
von: Bansal, Aniruddh, et al.
Veröffentlicht: (2025)
TimeLogic: A Temporal Logic Benchmark for Video QA
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
von: Kartha, Aaryaman, et al.
Veröffentlicht: (2025)
von: Kartha, Aaryaman, et al.
Veröffentlicht: (2025)
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
von: Liang, Fangzhou, et al.
Veröffentlicht: (2025)
von: Liang, Fangzhou, et al.
Veröffentlicht: (2025)
TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmark
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
CaT-BENCH: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans
von: Lal, Yash Kumar, et al.
Veröffentlicht: (2024)
von: Lal, Yash Kumar, et al.
Veröffentlicht: (2024)
Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image Models
von: Ventura, Mor, et al.
Veröffentlicht: (2023)
von: Ventura, Mor, et al.
Veröffentlicht: (2023)
GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
von: Long, Weicai, et al.
Veröffentlicht: (2026)
von: Long, Weicai, et al.
Veröffentlicht: (2026)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025) -
Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025) -
Abstracting Geo-specific Terrains to Scale Up Reinforcement Learning
von: Ustun, Volkan, et al.
Veröffentlicht: (2025) -
A Data-Driven Discretized CS:GO Simulation Environment to Facilitate Strategic Multi-Agent Planning Research
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025) -
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026)