VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Yufan, Tertikas, Konstantinos, Maiti, Shalini, Han, Junlin, Zhang, Tong, Süsstrunk, Sabine, Kokkinos, Filippos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
by: Maiti, Shalini, et al.
Published: (2025)
by: Maiti, Shalini, et al.
Published: (2025)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
by: Han, Junlin, et al.
Published: (2025)
by: Han, Junlin, et al.
Published: (2025)
Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
by: Chen, Yule, et al.
Published: (2025)
by: Chen, Yule, et al.
Published: (2025)
VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models
by: Han, Junlin, et al.
Published: (2024)
by: Han, Junlin, et al.
Published: (2024)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image Editing
by: Ren, Yufan, et al.
Published: (2025)
by: Ren, Yufan, et al.
Published: (2025)
Flex3D: Feed-Forward 3D Generation with Flexible Reconstruction Model and Input View Curation
by: Han, Junlin, et al.
Published: (2024)
by: Han, Junlin, et al.
Published: (2024)
AdaNCA: Neural Cellular Automata As Adaptors For More Robust Vision Transformer
by: Xu, Yitao, et al.
Published: (2024)
by: Xu, Yitao, et al.
Published: (2024)
Reasoning or Pattern Matching? Probing Large Vision-Language Models with Visual Puzzles
by: Lymperaiou, Maria, et al.
Published: (2026)
by: Lymperaiou, Maria, et al.
Published: (2026)
Adaptive Multi-step Refinement Network for Robust Point Cloud Registration
by: Chen, Zhi, et al.
Published: (2023)
by: Chen, Zhi, et al.
Published: (2023)
NEMTO: Neural Environment Matting for Novel View and Relighting Synthesis of Transparent Objects
by: Wang, Dongqing, et al.
Published: (2023)
by: Wang, Dongqing, et al.
Published: (2023)
CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
by: Cai, Jie, et al.
Published: (2025)
by: Cai, Jie, et al.
Published: (2025)
DSI2I: Dense Style for Unpaired Image-to-Image Translation
by: Ozaydin, Baran, et al.
Published: (2022)
by: Ozaydin, Baran, et al.
Published: (2022)
InNeRF360: Text-Guided 3D-Consistent Object Inpainting on 360-degree Neural Radiance Fields
by: Wang, Dongqing, et al.
Published: (2023)
by: Wang, Dongqing, et al.
Published: (2023)
PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
by: Shan, Haozhe, et al.
Published: (2026)
by: Shan, Haozhe, et al.
Published: (2026)
Canonical Latent Representations in Conditional Diffusion Models
by: Xu, Yitao, et al.
Published: (2025)
by: Xu, Yitao, et al.
Published: (2025)
Enhancing Frequency Forgery Clues for Diffusion-Generated Image Detection
by: Zhang, Daichi, et al.
Published: (2025)
by: Zhang, Daichi, et al.
Published: (2025)
IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
by: Ji, Yuheng, et al.
Published: (2025)
by: Ji, Yuheng, et al.
Published: (2025)
WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
by: Alper, Morris, et al.
Published: (2025)
by: Alper, Morris, et al.
Published: (2025)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
Emergent Dynamics in Neural Cellular Automata
by: Xu, Yitao, et al.
Published: (2024)
by: Xu, Yitao, et al.
Published: (2024)
2-Shots in the Dark: Low-Light Denoising with Minimal Data Acquisition
by: Lu, Liying, et al.
Published: (2025)
by: Lu, Liying, et al.
Published: (2025)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
by: Zhu, Fengbin, et al.
Published: (2024)
by: Zhu, Fengbin, et al.
Published: (2024)
Efficient Temporally-Aware DeepFake Detection using H.264 Motion Vectors
by: Grönquist, Peter, et al.
Published: (2023)
by: Grönquist, Peter, et al.
Published: (2023)
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
by: Li, Siqi, et al.
Published: (2025)
by: Li, Siqi, et al.
Published: (2025)
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
by: Luo, Zhiming, et al.
Published: (2026)
by: Luo, Zhiming, et al.
Published: (2026)
TempSAL -- Uncovering Temporal Information for Deep Saliency Prediction
by: Aydemir, Bahar, et al.
Published: (2023)
by: Aydemir, Bahar, et al.
Published: (2023)
Data Augmentation via Latent Diffusion for Saliency Prediction
by: Aydemir, Bahar, et al.
Published: (2024)
by: Aydemir, Bahar, et al.
Published: (2024)
AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
by: Bao, Han, et al.
Published: (2024)
by: Bao, Han, et al.
Published: (2024)
Unlocking Comics: The AI4VA Dataset for Visual Understanding
by: Grönquist, Peter, et al.
Published: (2024)
by: Grönquist, Peter, et al.
Published: (2024)
Color Constancy in Hyperspectral Imaging via Reduced Spectral Spaces
by: Vidarsson, G. Dofri, et al.
Published: (2026)
by: Vidarsson, G. Dofri, et al.
Published: (2026)
ChartBench: A Benchmark for Complex Visual Reasoning in Charts
by: Xu, Zhengzhuo, et al.
Published: (2023)
by: Xu, Zhengzhuo, et al.
Published: (2023)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
by: Ghosal, Deepanway, et al.
Published: (2024)
by: Ghosal, Deepanway, et al.
Published: (2024)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
by: Zhu, Yingjie, et al.
Published: (2024)
by: Zhu, Yingjie, et al.
Published: (2024)
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
by: Yuan, Botai, et al.
Published: (2025)
by: Yuan, Botai, et al.
Published: (2025)
Unsupervised 2D-3D lifting of non-rigid objects using local constraints
by: Maiti, Shalini, et al.
Published: (2025)
by: Maiti, Shalini, et al.
Published: (2025)
AgroBench: Vision-Language Model Benchmark in Agriculture
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
Similar Items
-
Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
by: Maiti, Shalini, et al.
Published: (2025) -
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
by: Han, Junlin, et al.
Published: (2025) -
Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
by: Chen, Yule, et al.
Published: (2025) -
VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models
by: Han, Junlin, et al.
Published: (2024) -
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)