Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
Fuente:
arXiv
Guardado en:
| Autores principales: | Xie, Peijin, Sun, Lin, Liu, Bingquan, Wang, Dexin, Zhang, Xiangzheng, Sun, Chengjie, Zhang, Jiajia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
por: Xie, Peijin, et al.
Publicado: (2025)
por: Xie, Peijin, et al.
Publicado: (2025)
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
por: Khalil, Ahmad, et al.
Publicado: (2025)
por: Khalil, Ahmad, et al.
Publicado: (2025)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
por: Zhang, Yi, et al.
Publicado: (2025)
por: Zhang, Yi, et al.
Publicado: (2025)
MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
por: Liu, Xinyu, et al.
Publicado: (2025)
por: Liu, Xinyu, et al.
Publicado: (2025)
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
por: Lu, Zhuqiang, et al.
Publicado: (2024)
por: Lu, Zhuqiang, et al.
Publicado: (2024)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
por: Zhan, Yu-Wei, et al.
Publicado: (2025)
por: Zhan, Yu-Wei, et al.
Publicado: (2025)
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
por: Khalil, Ahmad, et al.
Publicado: (2025)
por: Khalil, Ahmad, et al.
Publicado: (2025)
IM-IAD: Industrial Image Anomaly Detection Benchmark in Manufacturing
por: Xie, Guoyang, et al.
Publicado: (2023)
por: Xie, Guoyang, et al.
Publicado: (2023)
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
por: Qian, Shun, et al.
Publicado: (2024)
por: Qian, Shun, et al.
Publicado: (2024)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
por: Zhou, Zhiyu, et al.
Publicado: (2026)
por: Zhou, Zhiyu, et al.
Publicado: (2026)
Spatial-Aware Latent Initialization for Controllable Image Generation
por: Sun, Wenqiang, et al.
Publicado: (2024)
por: Sun, Wenqiang, et al.
Publicado: (2024)
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
por: Yu, Jiongze, et al.
Publicado: (2026)
por: Yu, Jiongze, et al.
Publicado: (2026)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
por: Bulat, Adrian, et al.
Publicado: (2026)
por: Bulat, Adrian, et al.
Publicado: (2026)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
por: Jin, Yizhang, et al.
Publicado: (2024)
por: Jin, Yizhang, et al.
Publicado: (2024)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
por: Liu, Chonghan, et al.
Publicado: (2025)
por: Liu, Chonghan, et al.
Publicado: (2025)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
por: Wang, Wei, et al.
Publicado: (2026)
por: Wang, Wei, et al.
Publicado: (2026)
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning
por: Wu, Nemin, et al.
Publicado: (2024)
por: Wu, Nemin, et al.
Publicado: (2024)
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
por: Lin, Jingli, et al.
Publicado: (2025)
por: Lin, Jingli, et al.
Publicado: (2025)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
por: Xu, Zelin, et al.
Publicado: (2026)
por: Xu, Zelin, et al.
Publicado: (2026)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
por: Wang, Xingrui, et al.
Publicado: (2025)
por: Wang, Xingrui, et al.
Publicado: (2025)
RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection
por: Wang, Zhuo, et al.
Publicado: (2025)
por: Wang, Zhuo, et al.
Publicado: (2025)
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
por: Wu, Haoning, et al.
Publicado: (2025)
por: Wu, Haoning, et al.
Publicado: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
por: Jia, Mengdi, et al.
Publicado: (2025)
por: Jia, Mengdi, et al.
Publicado: (2025)
RISE-Video: Can Video Generators Decode Implicit World Rules?
por: Liu, Mingxin, et al.
Publicado: (2026)
por: Liu, Mingxin, et al.
Publicado: (2026)
Object-Centric 3D Gaussian Splatting for Strawberry Plant Reconstruction and Phenotyping
por: Li, Jiajia, et al.
Publicado: (2025)
por: Li, Jiajia, et al.
Publicado: (2025)
HGNet: High-Order Spatial Awareness Hypergraph and Multi-Scale Context Attention Network for Colorectal Polyp Detection
por: Liu, Xiaofang, et al.
Publicado: (2025)
por: Liu, Xiaofang, et al.
Publicado: (2025)
OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU
por: Sun, Qi, et al.
Publicado: (2025)
por: Sun, Qi, et al.
Publicado: (2025)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
por: Guo, Zhongbin, et al.
Publicado: (2026)
por: Guo, Zhongbin, et al.
Publicado: (2026)
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
por: Zhou, Weilin, et al.
Publicado: (2026)
por: Zhou, Weilin, et al.
Publicado: (2026)
Benchmarking Pathology Foundation Models for Spatial Domain Understanding
por: Zhao, Bokai, et al.
Publicado: (2026)
por: Zhao, Bokai, et al.
Publicado: (2026)
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
por: Wang, Xinghao, et al.
Publicado: (2026)
por: Wang, Xinghao, et al.
Publicado: (2026)
Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Map
por: Chang, Xinyuan, et al.
Publicado: (2024)
por: Chang, Xinyuan, et al.
Publicado: (2024)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
por: Ropero, Fernando, et al.
Publicado: (2026)
por: Ropero, Fernando, et al.
Publicado: (2026)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
por: Cai, Yuxuan, et al.
Publicado: (2025)
por: Cai, Yuxuan, et al.
Publicado: (2025)
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
por: Wu, Qi, et al.
Publicado: (2025)
por: Wu, Qi, et al.
Publicado: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
por: Tian, Kexin, et al.
Publicado: (2025)
por: Tian, Kexin, et al.
Publicado: (2025)
LoopNav: Benchmarking Spatial Consistency in World Models
por: Lian, Kewei, et al.
Publicado: (2025)
por: Lian, Kewei, et al.
Publicado: (2025)
Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach
por: Wang, Xincheng, et al.
Publicado: (2025)
por: Wang, Xincheng, et al.
Publicado: (2025)
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
por: Liu, Yuhong, et al.
Publicado: (2025)
por: Liu, Yuhong, et al.
Publicado: (2025)
MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
por: Jiang, Xi, et al.
Publicado: (2024)
por: Jiang, Xi, et al.
Publicado: (2024)
Ejemplares similares
-
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
por: Xie, Peijin, et al.
Publicado: (2025) -
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
por: Khalil, Ahmad, et al.
Publicado: (2025) -
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
por: Zhang, Yi, et al.
Publicado: (2025) -
MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
por: Liu, Xinyu, et al.
Publicado: (2025) -
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
por: Lu, Zhuqiang, et al.
Publicado: (2024)