See, Symbolize, Act: Grounding VLMs with Spatial Representations for Better Gameplay
Fuente:
arXiv
Saved in:
| Main Authors: | Baghel, Ashish, Chopra, Paras |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Discovering Reinforcement Learning Interfaces with Large Language Models
by: Jaswal, Akshat Singh, et al.
Published: (2026)
by: Jaswal, Akshat Singh, et al.
Published: (2026)
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
by: Singh, Simardeep, et al.
Published: (2026)
by: Singh, Simardeep, et al.
Published: (2026)
Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
by: Trehan, Dhruv, et al.
Published: (2026)
by: Trehan, Dhruv, et al.
Published: (2026)
The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
by: Sharma, Aman, et al.
Published: (2025)
by: Sharma, Aman, et al.
Published: (2025)
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
by: Sharma, Aman, et al.
Published: (2025)
by: Sharma, Aman, et al.
Published: (2025)
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
by: Sharma, Aman, et al.
Published: (2026)
by: Sharma, Aman, et al.
Published: (2026)
Hybrid Neural World Models
by: Lakshmanan, Pranav, et al.
Published: (2026)
by: Lakshmanan, Pranav, et al.
Published: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
AWE: Adaptive Agents for Dynamic Web Penetration Testing
by: Jaswal, Akshat Singh, et al.
Published: (2026)
by: Jaswal, Akshat Singh, et al.
Published: (2026)
Building Interpretable Models for Moral Decision-Making
by: Goel, Mayank, et al.
Published: (2026)
by: Goel, Mayank, et al.
Published: (2026)
METIS: Mentoring Engine for Thoughtful Inquiry & Solutions
by: Kumar, Abhinav Rajeev, et al.
Published: (2026)
by: Kumar, Abhinav Rajeev, et al.
Published: (2026)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Do GFlowNets Transfer? Case Study on the Game of 24/42
by: Gupta, Adesh, et al.
Published: (2025)
by: Gupta, Adesh, et al.
Published: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
by: Zhang, Jianshu, et al.
Published: (2026)
by: Zhang, Jianshu, et al.
Published: (2026)
To See the Unseen: on the Generalization Ability of Transformers in Symbolic Reasoning
by: Lazić, Nevena, et al.
Published: (2026)
by: Lazić, Nevena, et al.
Published: (2026)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
The NES Video-Music Database: A Dataset of Symbolic Video Game Music Paired with Gameplay Videos
by: Cardoso, Igor, et al.
Published: (2024)
by: Cardoso, Igor, et al.
Published: (2024)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
Grounding Methods for Neural-Symbolic AI
by: Ontiveros, Rodrigo Castellano, et al.
Published: (2025)
by: Ontiveros, Rodrigo Castellano, et al.
Published: (2025)
Model-Grounded Symbolic Artificial Intelligence Systems Learning and Reasoning with Model-Grounded Symbolic Artificial Intelligence Systems
by: Chattopadhyay, Aniruddha, et al.
Published: (2025)
by: Chattopadhyay, Aniruddha, et al.
Published: (2025)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
Measuring How (Not Just Whether) VLMs Build Common Ground
by: Imai, Saki, et al.
Published: (2025)
by: Imai, Saki, et al.
Published: (2025)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
by: Liu, Zhining, et al.
Published: (2025)
by: Liu, Zhining, et al.
Published: (2025)
Semi-Strongly solved: a New Definition Leading Computer to Perfect Gameplay
by: Takizawa, Hiroki
Published: (2024)
by: Takizawa, Hiroki
Published: (2024)
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
by: Vaishnav, Mohit, et al.
Published: (2026)
by: Vaishnav, Mohit, et al.
Published: (2026)
pixelLOG: Logging of Online Gameplay for Cognitive Research
by: Lu, Zeyu, et al.
Published: (2026)
by: Lu, Zeyu, et al.
Published: (2026)
Grammar and Gameplay-aligned RL for Game Description Generation with LLMs
by: Tanaka, Tsunehiko, et al.
Published: (2025)
by: Tanaka, Tsunehiko, et al.
Published: (2025)
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
GroundAct: Can LLM Agents Ground Actions in Environmental States?
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
by: Chen, Weixin, et al.
Published: (2026)
by: Chen, Weixin, et al.
Published: (2026)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
by: Wang, Dianyi, et al.
Published: (2025)
by: Wang, Dianyi, et al.
Published: (2025)
A Mechanistic Investigation of Supervised Fine Tuning
by: Chopra, Ruhaan
Published: (2026)
by: Chopra, Ruhaan
Published: (2026)
Don't Look at the Numbers: Visual Anchoring Bias and Layer-wise Representation in VLMs
by: Shalankin, M.
Published: (2026)
by: Shalankin, M.
Published: (2026)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
by: Pan, Zhenyu, et al.
Published: (2025)
by: Pan, Zhenyu, et al.
Published: (2025)
Can Large Language Models Act as Symbolic Reasoners?
by: Sullivan, Rob, et al.
Published: (2024)
by: Sullivan, Rob, et al.
Published: (2024)
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
by: Yu, Songsong, et al.
Published: (2025)
by: Yu, Songsong, et al.
Published: (2025)
The Garden of Forking Paths: Narrative Arc-Conditioned Gameplay Planning
by: Wen, Yunge, et al.
Published: (2026)
by: Wen, Yunge, et al.
Published: (2026)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
by: Liu, Li, et al.
Published: (2024)
by: Liu, Li, et al.
Published: (2024)
Finite Automata Extraction: Low-data World Model Learning as Programs from Gameplay Video
by: Goel, Dave, et al.
Published: (2025)
by: Goel, Dave, et al.
Published: (2025)
Similar Items
-
Discovering Reinforcement Learning Interfaces with Large Language Models
by: Jaswal, Akshat Singh, et al.
Published: (2026) -
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
by: Singh, Simardeep, et al.
Published: (2026) -
Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
by: Trehan, Dhruv, et al.
Published: (2026) -
The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
by: Sharma, Aman, et al.
Published: (2025) -
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
by: Sharma, Aman, et al.
Published: (2025)