InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Deng, Nianchen, Gu, Lixin, Ye, Shenglong, He, Yinan, Chen, Zhe, Li, Songze, Wang, Haomin, Wei, Xingguang, Yang, Tianshuo, Dou, Min, He, Tong, Shao, Wenqi, Zhang, Kaipeng, Wang, Yi, Shi, Botian, Zhang, Yanting, Dai, Jifeng, Qiao, Yu, Zhang, Hongjie, Wang, Wenhai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
por: Wei, Xingguang, et al.
Publicado: (2025)
por: Wei, Xingguang, et al.
Publicado: (2025)
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
por: Wang, Haomin, et al.
Publicado: (2025)
por: Wang, Haomin, et al.
Publicado: (2025)
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
por: Zhu, Jinguo, et al.
Publicado: (2025)
por: Zhu, Jinguo, et al.
Publicado: (2025)
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
por: Wang, Weiyun, et al.
Publicado: (2025)
por: Wang, Weiyun, et al.
Publicado: (2025)
ArchCAD-400K: A Large-Scale CAD drawings Dataset and New Baseline for Panoptic Symbol Spotting
por: Luo, Ruifeng, et al.
Publicado: (2025)
por: Luo, Ruifeng, et al.
Publicado: (2025)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
por: Gao, Zhangwei, et al.
Publicado: (2024)
por: Gao, Zhangwei, et al.
Publicado: (2024)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
por: Wang, Yi, et al.
Publicado: (2024)
por: Wang, Yi, et al.
Publicado: (2024)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
por: Wang, Yi, et al.
Publicado: (2025)
por: Wang, Yi, et al.
Publicado: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
por: Luo, Gen, et al.
Publicado: (2025)
por: Luo, Gen, et al.
Publicado: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
por: Jia, Mengdi, et al.
Publicado: (2025)
por: Jia, Mengdi, et al.
Publicado: (2025)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
por: Chen, Zhe, et al.
Publicado: (2023)
por: Chen, Zhe, et al.
Publicado: (2023)
Docopilot: Improving Multimodal Models for Document-Level Understanding
por: Duan, Yuchen, et al.
Publicado: (2025)
por: Duan, Yuchen, et al.
Publicado: (2025)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
por: Meng, Fanqing, et al.
Publicado: (2024)
por: Meng, Fanqing, et al.
Publicado: (2024)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
por: Zhao, Lirui, et al.
Publicado: (2024)
por: Zhao, Lirui, et al.
Publicado: (2024)
ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality
por: He, Yefei, et al.
Publicado: (2024)
por: He, Yefei, et al.
Publicado: (2024)
MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
por: Meng, Fanqing, et al.
Publicado: (2025)
por: Meng, Fanqing, et al.
Publicado: (2025)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
por: Luo, Gen, et al.
Publicado: (2024)
por: Luo, Gen, et al.
Publicado: (2024)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
por: Chen, Xinyi, et al.
Publicado: (2025)
por: Chen, Xinyi, et al.
Publicado: (2025)
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
por: Dong, Xiaoyi, et al.
Publicado: (2024)
por: Dong, Xiaoyi, et al.
Publicado: (2024)
Scattering and Gathering for Spatially Varying Blurs
por: Chimitt, Nicholas, et al.
Publicado: (2023)
por: Chimitt, Nicholas, et al.
Publicado: (2023)
Needle In A Multimodal Haystack
por: Wang, Weiyun, et al.
Publicado: (2024)
por: Wang, Weiyun, et al.
Publicado: (2024)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
por: Wang, Yi, et al.
Publicado: (2023)
por: Wang, Yi, et al.
Publicado: (2023)
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
por: Wu, Haoning, et al.
Publicado: (2025)
por: Wu, Haoning, et al.
Publicado: (2025)
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
por: Zhang, Pan, et al.
Publicado: (2024)
por: Zhang, Pan, et al.
Publicado: (2024)
Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments
por: Yang, Yang, et al.
Publicado: (2024)
por: Yang, Yang, et al.
Publicado: (2024)
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction
por: Xiong, Kai, et al.
Publicado: (2026)
por: Xiong, Kai, et al.
Publicado: (2026)
Spatial Data and Evaluation Indicators for Eco-Tourism Development Value in Taihang Honggu National Forest Park, China
por: Zhang, Wenqi
Publicado: (2025)
por: Zhang, Wenqi
Publicado: (2025)
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
por: Dong, Xiaoyi, et al.
Publicado: (2024)
por: Dong, Xiaoyi, et al.
Publicado: (2024)
Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
por: Zhu, Zheng, et al.
Publicado: (2024)
por: Zhu, Zheng, et al.
Publicado: (2024)
ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
por: Cui, Long, et al.
Publicado: (2025)
por: Cui, Long, et al.
Publicado: (2025)
ModiGen: A Large Language Model-Based Workflow for Multi-Task Modelica Code Generation
por: Xiang, Jiahui, et al.
Publicado: (2025)
por: Xiang, Jiahui, et al.
Publicado: (2025)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
por: Wang, Weiyun, et al.
Publicado: (2024)
por: Wang, Weiyun, et al.
Publicado: (2024)
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
por: Chen, Zhe, et al.
Publicado: (2024)
por: Chen, Zhe, et al.
Publicado: (2024)
VideoChat: Chat-Centric Video Understanding
por: Li, KunChang, et al.
Publicado: (2023)
por: Li, KunChang, et al.
Publicado: (2023)
Intern-S1: A Scientific Multimodal Foundation Model
por: Bai, Lei, et al.
Publicado: (2025)
por: Bai, Lei, et al.
Publicado: (2025)
Adaptive Blind Super-Resolution Network for Spatial-Specific and Spatial-Agnostic Degradations
por: Wen, Weilei, et al.
Publicado: (2025)
por: Wen, Weilei, et al.
Publicado: (2025)
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
por: Zhang, Pan, et al.
Publicado: (2024)
por: Zhang, Pan, et al.
Publicado: (2024)
Scalable tensor network algorithm for thermal quantum many-body systems in two dimension
por: Zhang, Meng, et al.
Publicado: (2024)
por: Zhang, Meng, et al.
Publicado: (2024)
Knowledge Discovery in Spatial Cartographic Information Retrieval.
por: Yu, Lixin
Publicado: (1999)
por: Yu, Lixin
Publicado: (1999)
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
por: Jing, Linglin, et al.
Publicado: (2025)
por: Jing, Linglin, et al.
Publicado: (2025)
Ejemplares similares
-
Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
por: Wei, Xingguang, et al.
Publicado: (2025) -
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
por: Wang, Haomin, et al.
Publicado: (2025) -
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
por: Zhu, Jinguo, et al.
Publicado: (2025) -
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
por: Wang, Weiyun, et al.
Publicado: (2025) -
ArchCAD-400K: A Large-Scale CAD drawings Dataset and New Baseline for Panoptic Symbol Spotting
por: Luo, Ruifeng, et al.
Publicado: (2025)