Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Xiaoyu, Huang, Wenxuan, Sun, Hao, Fu, Xinyu, Ma, Changfeng, Cao, Shaosheng, Jia, Bohan, Lin, Shaohui, Yin, Zhenfei, Bai, Lei, Ouyang, Wanli, Li, Yuanqi, Guo, Jie, Guo, Yanwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-Time-Matching: Decouple Personality, Memory, and Linguistic Style in LLM-based Role-Playing Language Agent
by: Zhan, Xiaoyu, et al.
Published: (2025)
by: Zhan, Xiaoyu, et al.
Published: (2025)
Semantic Human Mesh Reconstruction with Textures
by: Zhan, Xiaoyu, et al.
Published: (2024)
by: Zhan, Xiaoyu, et al.
Published: (2024)
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2025)
by: Huang, Wenxuan, et al.
Published: (2025)
Interleaving Reasoning for Better Text-to-Image Generation
by: Huang, Wenxuan, et al.
Published: (2025)
by: Huang, Wenxuan, et al.
Published: (2025)
On the Error Analysis of 3D Gaussian Splatting and an Optimal Projection Strategy
by: Huang, Letian, et al.
Published: (2024)
by: Huang, Letian, et al.
Published: (2024)
360-GS: Layout-guided Panoramic Gaussian Splatting For Indoor Roaming
by: Bai, Jiayang, et al.
Published: (2024)
by: Bai, Jiayang, et al.
Published: (2024)
Sparse Point Cloud Patches Rendering via Splitting 2D Gaussians
by: Changfeng, Ma, et al.
Published: (2025)
by: Changfeng, Ma, et al.
Published: (2025)
Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2026)
by: Huang, Wenxuan, et al.
Published: (2026)
Flatten The Complex: Joint B-Rep Generation via Compositional $k$-Cell Particles
by: Lu, Junran, et al.
Published: (2026)
by: Lu, Junran, et al.
Published: (2026)
CompBench: Benchmarking Complex Instruction-guided Image Editing
by: Jia, Bohan, et al.
Published: (2025)
by: Jia, Bohan, et al.
Published: (2025)
UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars
by: Zhan, Xiaoyu, et al.
Published: (2026)
by: Zhan, Xiaoyu, et al.
Published: (2026)
Spectral-GS: Taming 3D Gaussian Splatting with Spectral Entropy
by: Huang, Letian, et al.
Published: (2024)
by: Huang, Letian, et al.
Published: (2024)
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
by: Zeng, Yu, et al.
Published: (2026)
by: Zeng, Yu, et al.
Published: (2026)
GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo Collections
by: Bai, Haiyang, et al.
Published: (2025)
by: Bai, Haiyang, et al.
Published: (2025)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
by: Huang, Wenxuan, et al.
Published: (2024)
by: Huang, Wenxuan, et al.
Published: (2024)
Parameterize Structure with Differentiable Template for 3D Shape Generation
by: Ma, Changfeng, et al.
Published: (2024)
by: Ma, Changfeng, et al.
Published: (2024)
DOGE: Differentiable Bezier Graph Optimization for Road Network Extraction
by: Sun, Jiahui, et al.
Published: (2025)
by: Sun, Jiahui, et al.
Published: (2025)
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
by: Xue, Xiangyuan, et al.
Published: (2026)
by: Xue, Xiangyuan, et al.
Published: (2026)
BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splatting
by: Yu, Jiaxing, et al.
Published: (2026)
by: Yu, Jiaxing, et al.
Published: (2026)
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
by: Xu, Wanghan, et al.
Published: (2026)
by: Xu, Wanghan, et al.
Published: (2026)
FINER++: Building a Family of Variable-periodic Functions for Activating Implicit Neural Representation
by: Zhu, Hao, et al.
Published: (2024)
by: Zhu, Hao, et al.
Published: (2024)
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
by: Liu, Dingning, et al.
Published: (2024)
by: Liu, Dingning, et al.
Published: (2024)
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?
by: Li, Bangyan, et al.
Published: (2025)
by: Li, Bangyan, et al.
Published: (2025)
GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
by: Wang, Yiqi, et al.
Published: (2024)
by: Wang, Yiqi, et al.
Published: (2024)
LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning
by: Yang, Xinran, et al.
Published: (2025)
by: Yang, Xinran, et al.
Published: (2025)
AnchorRoute: Human Motion Synthesis with Interval-Routed Sparse Contro
by: Fang, Pengcheng, et al.
Published: (2026)
by: Fang, Pengcheng, et al.
Published: (2026)
Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services
by: Guo, Hongcheng, et al.
Published: (2025)
by: Guo, Hongcheng, et al.
Published: (2025)
Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model
by: Jiao, Siwen, et al.
Published: (2026)
by: Jiao, Siwen, et al.
Published: (2026)
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
by: Xue, Xiangyuan, et al.
Published: (2025)
by: Xue, Xiangyuan, et al.
Published: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
by: Wang, Peijie, et al.
Published: (2025)
by: Wang, Peijie, et al.
Published: (2025)
TransparentGS: Fast Inverse Rendering of Transparent Objects with Gaussians
by: Huang, Letian, et al.
Published: (2025)
by: Huang, Letian, et al.
Published: (2025)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
by: Chen, Zeren, et al.
Published: (2023)
by: Chen, Zeren, et al.
Published: (2023)
ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability
by: Liu, Wenhan, et al.
Published: (2025)
by: Liu, Wenhan, et al.
Published: (2025)
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Code-driven Number Sequence Calculation: Enhancing the inductive Reasoning Abilities of Large Language Models
by: Chen, Kedi, et al.
Published: (2025)
by: Chen, Kedi, et al.
Published: (2025)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
by: Wei, Jiaqi, et al.
Published: (2026)
by: Wei, Jiaqi, et al.
Published: (2026)
Similar Items
-
Test-Time-Matching: Decouple Personality, Memory, and Linguistic Style in LLM-based Role-Playing Language Agent
by: Zhan, Xiaoyu, et al.
Published: (2025) -
Semantic Human Mesh Reconstruction with Textures
by: Zhan, Xiaoyu, et al.
Published: (2024) -
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2025) -
Interleaving Reasoning for Better Text-to-Image Generation
by: Huang, Wenxuan, et al.
Published: (2025) -
On the Error Analysis of 3D Gaussian Splatting and an Optimal Projection Strategy
by: Huang, Letian, et al.
Published: (2024)