Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
Fuente:
arXiv
Salvato in:
| Autori principali: | Xie, Peijin, Sun, Lin, Liu, Bingquan, Wang, Dexin, Zhang, Xiangzheng, Sun, Chengjie, Zhang, Jiajia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
di: Xie, Peijin, et al.
Pubblicazione: (2025)
di: Xie, Peijin, et al.
Pubblicazione: (2025)
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
di: Khalil, Ahmad, et al.
Pubblicazione: (2025)
di: Khalil, Ahmad, et al.
Pubblicazione: (2025)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
di: Zhang, Yi, et al.
Pubblicazione: (2025)
di: Zhang, Yi, et al.
Pubblicazione: (2025)
MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
di: Lu, Zhuqiang, et al.
Pubblicazione: (2024)
di: Lu, Zhuqiang, et al.
Pubblicazione: (2024)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
di: Zhan, Yu-Wei, et al.
Pubblicazione: (2025)
di: Zhan, Yu-Wei, et al.
Pubblicazione: (2025)
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
di: Khalil, Ahmad, et al.
Pubblicazione: (2025)
di: Khalil, Ahmad, et al.
Pubblicazione: (2025)
IM-IAD: Industrial Image Anomaly Detection Benchmark in Manufacturing
di: Xie, Guoyang, et al.
Pubblicazione: (2023)
di: Xie, Guoyang, et al.
Pubblicazione: (2023)
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
di: Qian, Shun, et al.
Pubblicazione: (2024)
di: Qian, Shun, et al.
Pubblicazione: (2024)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
di: Zhou, Zhiyu, et al.
Pubblicazione: (2026)
di: Zhou, Zhiyu, et al.
Pubblicazione: (2026)
Spatial-Aware Latent Initialization for Controllable Image Generation
di: Sun, Wenqiang, et al.
Pubblicazione: (2024)
di: Sun, Wenqiang, et al.
Pubblicazione: (2024)
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
di: Yu, Jiongze, et al.
Pubblicazione: (2026)
di: Yu, Jiongze, et al.
Pubblicazione: (2026)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
di: Bulat, Adrian, et al.
Pubblicazione: (2026)
di: Bulat, Adrian, et al.
Pubblicazione: (2026)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
di: Jin, Yizhang, et al.
Pubblicazione: (2024)
di: Jin, Yizhang, et al.
Pubblicazione: (2024)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
di: Liu, Chonghan, et al.
Pubblicazione: (2025)
di: Liu, Chonghan, et al.
Pubblicazione: (2025)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
di: Wang, Wei, et al.
Pubblicazione: (2026)
di: Wang, Wei, et al.
Pubblicazione: (2026)
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning
di: Wu, Nemin, et al.
Pubblicazione: (2024)
di: Wu, Nemin, et al.
Pubblicazione: (2024)
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
di: Lin, Jingli, et al.
Pubblicazione: (2025)
di: Lin, Jingli, et al.
Pubblicazione: (2025)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
di: Xu, Zelin, et al.
Pubblicazione: (2026)
di: Xu, Zelin, et al.
Pubblicazione: (2026)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection
di: Wang, Zhuo, et al.
Pubblicazione: (2025)
di: Wang, Zhuo, et al.
Pubblicazione: (2025)
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
di: Wu, Haoning, et al.
Pubblicazione: (2025)
di: Wu, Haoning, et al.
Pubblicazione: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
di: Jia, Mengdi, et al.
Pubblicazione: (2025)
di: Jia, Mengdi, et al.
Pubblicazione: (2025)
RISE-Video: Can Video Generators Decode Implicit World Rules?
di: Liu, Mingxin, et al.
Pubblicazione: (2026)
di: Liu, Mingxin, et al.
Pubblicazione: (2026)
Object-Centric 3D Gaussian Splatting for Strawberry Plant Reconstruction and Phenotyping
di: Li, Jiajia, et al.
Pubblicazione: (2025)
di: Li, Jiajia, et al.
Pubblicazione: (2025)
HGNet: High-Order Spatial Awareness Hypergraph and Multi-Scale Context Attention Network for Colorectal Polyp Detection
di: Liu, Xiaofang, et al.
Pubblicazione: (2025)
di: Liu, Xiaofang, et al.
Pubblicazione: (2025)
OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU
di: Sun, Qi, et al.
Pubblicazione: (2025)
di: Sun, Qi, et al.
Pubblicazione: (2025)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
di: Zhou, Weilin, et al.
Pubblicazione: (2026)
di: Zhou, Weilin, et al.
Pubblicazione: (2026)
Benchmarking Pathology Foundation Models for Spatial Domain Understanding
di: Zhao, Bokai, et al.
Pubblicazione: (2026)
di: Zhao, Bokai, et al.
Pubblicazione: (2026)
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
di: Wang, Xinghao, et al.
Pubblicazione: (2026)
di: Wang, Xinghao, et al.
Pubblicazione: (2026)
Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Map
di: Chang, Xinyuan, et al.
Pubblicazione: (2024)
di: Chang, Xinyuan, et al.
Pubblicazione: (2024)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
di: Ropero, Fernando, et al.
Pubblicazione: (2026)
di: Ropero, Fernando, et al.
Pubblicazione: (2026)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
di: Cai, Yuxuan, et al.
Pubblicazione: (2025)
di: Cai, Yuxuan, et al.
Pubblicazione: (2025)
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
di: Wu, Qi, et al.
Pubblicazione: (2025)
di: Wu, Qi, et al.
Pubblicazione: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
di: Tian, Kexin, et al.
Pubblicazione: (2025)
di: Tian, Kexin, et al.
Pubblicazione: (2025)
LoopNav: Benchmarking Spatial Consistency in World Models
di: Lian, Kewei, et al.
Pubblicazione: (2025)
di: Lian, Kewei, et al.
Pubblicazione: (2025)
Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach
di: Wang, Xincheng, et al.
Pubblicazione: (2025)
di: Wang, Xincheng, et al.
Pubblicazione: (2025)
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
di: Liu, Yuhong, et al.
Pubblicazione: (2025)
di: Liu, Yuhong, et al.
Pubblicazione: (2025)
MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
di: Jiang, Xi, et al.
Pubblicazione: (2024)
di: Jiang, Xi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
di: Xie, Peijin, et al.
Pubblicazione: (2025) -
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
di: Khalil, Ahmad, et al.
Pubblicazione: (2025) -
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
di: Zhang, Yi, et al.
Pubblicazione: (2025) -
MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
di: Liu, Xinyu, et al.
Pubblicazione: (2025) -
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
di: Lu, Zhuqiang, et al.
Pubblicazione: (2024)