ING-VP: MLLMs cannot Play Easy Vision-based Games Yet
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Haoran, Guo, Hangyu, Guo, Shuyue, Cao, Meng, Huang, Wenhao, Liu, Jiaheng, Zhang, Ge |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
par: Cao, Meng, et autres
Publié: (2024)
par: Cao, Meng, et autres
Publié: (2024)
LLaVA-OneVision: Easy Visual Task Transfer
par: Li, Bo, et autres
Publié: (2024)
par: Li, Bo, et autres
Publié: (2024)
LongIns: A Challenging Long-context Instruction-based Exam for LLMs
par: Gavin, Shawn, et autres
Publié: (2024)
par: Gavin, Shawn, et autres
Publié: (2024)
CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
par: Peng, Zhongyuan, et autres
Publié: (2025)
par: Peng, Zhongyuan, et autres
Publié: (2025)
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
par: Zheng, Tianyu, et autres
Publié: (2024)
par: Zheng, Tianyu, et autres
Publié: (2024)
Vision Language Models Are Not (Yet) Spelling Correctors
par: Liang, Junhong, et autres
Publié: (2025)
par: Liang, Junhong, et autres
Publié: (2025)
Can MLLMs Understand the Deep Implication Behind Chinese Images?
par: Zhang, Chenhao, et autres
Publié: (2024)
par: Zhang, Chenhao, et autres
Publié: (2024)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
par: Mi, Hongze, et autres
Publié: (2024)
par: Mi, Hongze, et autres
Publié: (2024)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
par: Wang, Xiaoyang, et autres
Publié: (2025)
par: Wang, Xiaoyang, et autres
Publié: (2025)
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
par: Liang, Yiming, et autres
Publié: (2024)
par: Liang, Yiming, et autres
Publié: (2024)
GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models
par: Li, Shilong, et autres
Publié: (2024)
par: Li, Shilong, et autres
Publié: (2024)
Ethical Considerations of Large Language Models in Game Playing
par: Zhang, Qingquan, et autres
Publié: (2025)
par: Zhang, Qingquan, et autres
Publié: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
par: Yu, En, et autres
Publié: (2025)
par: Yu, En, et autres
Publié: (2025)
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
par: Wang, Zekun Moore, et autres
Publié: (2023)
par: Wang, Zekun Moore, et autres
Publié: (2023)
Affordance Benchmark for MLLMs
par: Wang, Junying, et autres
Publié: (2025)
par: Wang, Junying, et autres
Publié: (2025)
MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
par: Wang, Pei, et autres
Publié: (2024)
par: Wang, Pei, et autres
Publié: (2024)
Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty
par: Ling, Zehui, et autres
Publié: (2025)
par: Ling, Zehui, et autres
Publié: (2025)
p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay
par: Zhang, Jun, et autres
Publié: (2024)
par: Zhang, Jun, et autres
Publié: (2024)
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
par: Ling Team, et autres
Publié: (2025)
par: Ling Team, et autres
Publié: (2025)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
par: Li, Shilong, et autres
Publié: (2024)
par: Li, Shilong, et autres
Publié: (2024)
Learning to Play Like Humans: A Framework for LLM Adaptation in Interactive Fiction Games
par: Zhang, Jinming, et autres
Publié: (2025)
par: Zhang, Jinming, et autres
Publié: (2025)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
par: Chen, Qiguang, et autres
Publié: (2026)
par: Chen, Qiguang, et autres
Publié: (2026)
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
par: Yang, Enneng, et autres
Publié: (2024)
par: Yang, Enneng, et autres
Publié: (2024)
TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
par: Li, Yizhi, et autres
Publié: (2025)
par: Li, Yizhi, et autres
Publié: (2025)
TextAtari: 100K Frames Game Playing with Language Agents
par: Li, Wenhao, et autres
Publié: (2025)
par: Li, Wenhao, et autres
Publié: (2025)
Play to Generalize: Learning to Reason Through Game Play
par: Xie, Yunfei, et autres
Publié: (2025)
par: Xie, Yunfei, et autres
Publié: (2025)
A Text-to-Game Engine for UGC-Based Role-Playing Games
par: Zhang, Lei, et autres
Publié: (2024)
par: Zhang, Lei, et autres
Publié: (2024)
Cultivating Game Sense for Yourself: Making VLMs Gaming Experts
par: Lu, Wenxuan, et autres
Publié: (2025)
par: Lu, Wenxuan, et autres
Publié: (2025)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
par: Ye, Hengwei, et autres
Publié: (2026)
par: Ye, Hengwei, et autres
Publié: (2026)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
par: Yao, Huanjin, et autres
Publié: (2025)
par: Yao, Huanjin, et autres
Publié: (2025)
Do MLLMs Really Understand the Charts?
par: Zhang, Xiao, et autres
Publié: (2025)
par: Zhang, Xiao, et autres
Publié: (2025)
Test-Time-Matching: Decouple Personality, Memory, and Linguistic Style in LLM-based Role-Playing Language Agent
par: Zhan, Xiaoyu, et autres
Publié: (2025)
par: Zhan, Xiaoyu, et autres
Publié: (2025)
FCPE: A Fast Context-based Pitch Estimation Model
par: Luo, Yuxin, et autres
Publié: (2025)
par: Luo, Yuxin, et autres
Publié: (2025)
DDK: Distilling Domain Knowledge for Efficient Large Language Models
par: Liu, Jiaheng, et autres
Publié: (2024)
par: Liu, Jiaheng, et autres
Publié: (2024)
Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
par: Zhang, Yikai, et autres
Publié: (2025)
par: Zhang, Yikai, et autres
Publié: (2025)
AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs
par: Zhu, Han, et autres
Publié: (2026)
par: Zhu, Han, et autres
Publié: (2026)
ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
par: He, Liyang, et autres
Publié: (2025)
par: He, Liyang, et autres
Publié: (2025)
MdEval: Massively Multilingual Code Debugging
par: Liu, Shukai, et autres
Publié: (2024)
par: Liu, Shukai, et autres
Publié: (2024)
AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
par: Li, Ziming, et autres
Publié: (2024)
par: Li, Ziming, et autres
Publié: (2024)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
par: Hudi, Frederikus, et autres
Publié: (2025)
par: Hudi, Frederikus, et autres
Publié: (2025)
Documents similaires
-
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
par: Cao, Meng, et autres
Publié: (2024) -
LLaVA-OneVision: Easy Visual Task Transfer
par: Li, Bo, et autres
Publié: (2024) -
LongIns: A Challenging Long-context Instruction-based Exam for LLMs
par: Gavin, Shawn, et autres
Publié: (2024) -
CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
par: Peng, Zhongyuan, et autres
Publié: (2025) -
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
par: Zheng, Tianyu, et autres
Publié: (2024)