Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Diankun, Liu, Fangfu, Hung, Yi-Hsin, Duan, Yueqi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
por: Zhou, Yue, et al.
Publicado: (2025)
por: Zhou, Yue, et al.
Publicado: (2025)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
por: Chen, Yuhao, et al.
Publicado: (2026)
por: Chen, Yuhao, et al.
Publicado: (2026)
Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformers
por: Wang, Binxu, et al.
Publicado: (2026)
por: Wang, Binxu, et al.
Publicado: (2026)
MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
por: Dong, Yi, et al.
Publicado: (2025)
por: Dong, Yi, et al.
Publicado: (2025)
Perceptual Flow Network for Visually Grounded Reasoning
por: Li, Yangfu, et al.
Publicado: (2026)
por: Li, Yangfu, et al.
Publicado: (2026)
ESCAPE: Energy-based Selective Adaptive Correction for Out-of-distribution 3D Human Pose Estimation
por: Bidulka, Luke, et al.
Publicado: (2024)
por: Bidulka, Luke, et al.
Publicado: (2024)
Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
por: Huang, Yuhang, et al.
Publicado: (2025)
por: Huang, Yuhang, et al.
Publicado: (2025)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
por: Liu, Zhi
Publicado: (2026)
por: Liu, Zhi
Publicado: (2026)
Reasoning is a Modality
por: Liu, Zhiguang, et al.
Publicado: (2026)
por: Liu, Zhiguang, et al.
Publicado: (2026)
Spatially Optimized Compact Deep Metric Learning Model for Similarity Search
por: Islam, Md. Farhadul, et al.
Publicado: (2024)
por: Islam, Md. Farhadul, et al.
Publicado: (2024)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
por: Wang, Zhaohui, et al.
Publicado: (2025)
por: Wang, Zhaohui, et al.
Publicado: (2025)
Invariant Representation via Decoupling Style and Spurious Features from Images
por: Li, Ruimeng, et al.
Publicado: (2023)
por: Li, Ruimeng, et al.
Publicado: (2023)
MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding
por: Wang, Chiyue, et al.
Publicado: (2026)
por: Wang, Chiyue, et al.
Publicado: (2026)
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
por: Bu, Weijue, et al.
Publicado: (2025)
por: Bu, Weijue, et al.
Publicado: (2025)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
por: Caselles-Dupré, Hugo, et al.
Publicado: (2026)
por: Caselles-Dupré, Hugo, et al.
Publicado: (2026)
A Hybrid Deep Learning and Model-Checking Framework for Accurate Brain Tumor Detection and Validation
por: Elfatimi, Elhoucine, et al.
Publicado: (2024)
por: Elfatimi, Elhoucine, et al.
Publicado: (2024)
Adversarial Robustness of Deep Learning-Based Thyroid Nodule Segmentation in Ultrasound
por: Dietrich, Nicholas, et al.
Publicado: (2026)
por: Dietrich, Nicholas, et al.
Publicado: (2026)
FlightScope: An Experimental Comparative Review of Aircraft Detection Algorithms in Satellite Imagery
por: Ghazouali, Safouane El, et al.
Publicado: (2024)
por: Ghazouali, Safouane El, et al.
Publicado: (2024)
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
por: Ziakas, Christos, et al.
Publicado: (2025)
por: Ziakas, Christos, et al.
Publicado: (2025)
An Active Inference Model of Covert and Overt Visual Attention
por: Mišić, Tin, et al.
Publicado: (2025)
por: Mišić, Tin, et al.
Publicado: (2025)
Compound and Parallel Modes of Tropical Convolutional Neural Networks
por: Li, Mingbo, et al.
Publicado: (2025)
por: Li, Mingbo, et al.
Publicado: (2025)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
por: Tang, Guowei
Publicado: (2026)
por: Tang, Guowei
Publicado: (2026)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
por: Ji, Binbin, et al.
Publicado: (2025)
por: Ji, Binbin, et al.
Publicado: (2025)
Adapting Multimodal Foundation Models for Few-Shot Learning: A Comprehensive Study on Contrastive Captioners
por: Narasinghe, N. K. B. M. P. K. B., et al.
Publicado: (2025)
por: Narasinghe, N. K. B. M. P. K. B., et al.
Publicado: (2025)
PhysVid: Physics Aware Local Conditioning for Generative Video Models
por: Pathak, Saurabh, et al.
Publicado: (2026)
por: Pathak, Saurabh, et al.
Publicado: (2026)
Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
por: Naeen, Mohammad Ali Etemadi, et al.
Publicado: (2025)
por: Naeen, Mohammad Ali Etemadi, et al.
Publicado: (2025)
Pointing-Guided Target Estimation via Transformer-Based Attention
por: Müller, Luca, et al.
Publicado: (2025)
por: Müller, Luca, et al.
Publicado: (2025)
nuScenes Knowledge Graph -- A comprehensive semantic representation of traffic scenes for trajectory prediction
por: Mlodzian, Leon, et al.
Publicado: (2023)
por: Mlodzian, Leon, et al.
Publicado: (2023)
Detection and Simulation of Urban Heat Islands Using a Fine-Tuned Geospatial Foundation Model
por: Kreismann, David
Publicado: (2025)
por: Kreismann, David
Publicado: (2025)
Advancing Attribution-Based Neural Network Explainability through Relative Absolute Magnitude Layer-Wise Relevance Propagation and Multi-Component Evaluation
por: Vukadin, Davor, et al.
Publicado: (2024)
por: Vukadin, Davor, et al.
Publicado: (2024)
Text-to-Events: Synthetic Event Camera Streams from Conditional Text Input
por: Ott, Joachim, et al.
Publicado: (2024)
por: Ott, Joachim, et al.
Publicado: (2024)
Part-Level 3D Gaussian Vehicle Generation with Joint and Hinge Axis Estimation
por: Qian, Shiyao, et al.
Publicado: (2026)
por: Qian, Shiyao, et al.
Publicado: (2026)
Efficient Attention: Attention with Linear Complexities
por: Shen, Zhuoran, et al.
Publicado: (2018)
por: Shen, Zhuoran, et al.
Publicado: (2018)
NFR: Neural Feature-Guided Non-Rigid Shape Registration
por: Chen, Zhangquan, et al.
Publicado: (2025)
por: Chen, Zhangquan, et al.
Publicado: (2025)
OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
por: Huang, Mark He, et al.
Publicado: (2025)
por: Huang, Mark He, et al.
Publicado: (2025)
Ultra-Low-Latency Spiking Neural Networks with Temporal-Dependent Integrate-and-Fire Neuron Model for Objects Detection
por: Zhang, Chengjun, et al.
Publicado: (2025)
por: Zhang, Chengjun, et al.
Publicado: (2025)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
JotlasNet: Joint Tensor Low-Rank and Attention-based Sparse Unrolling Network for Accelerating Dynamic MRI
por: Zhang, Yinghao, et al.
Publicado: (2025)
por: Zhang, Yinghao, et al.
Publicado: (2025)
Robustness Feature Adapter for Efficient Adversarial Training
por: Wu, Quanwei, et al.
Publicado: (2025)
por: Wu, Quanwei, et al.
Publicado: (2025)
Ejemplares similares
-
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
por: Zhou, Yue, et al.
Publicado: (2025) -
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
por: Chen, Yuhao, et al.
Publicado: (2026) -
Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformers
por: Wang, Binxu, et al.
Publicado: (2026) -
MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
por: Dong, Yi, et al.
Publicado: (2025) -
Perceptual Flow Network for Visually Grounded Reasoning
por: Li, Yangfu, et al.
Publicado: (2026)