Similar Items
On Computational Limits of FlowAR Models: Expressivity and Efficiency
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
by: Ke, Yekun, et al.
Published: (2025)
by: Ke, Yekun, et al.
Published: (2025)
Automatic Channel Pruning for Multi-Head Attention
by: Lee, Eunho, et al.
Published: (2024)
by: Lee, Eunho, et al.
Published: (2024)
fruit-SALAD: A Style Aligned Artwork Dataset to reveal similarity perception in image embeddings
by: Ohm, Tillmann, et al.
Published: (2024)
by: Ohm, Tillmann, et al.
Published: (2024)
Limits of Spatial Imagery Reasoning in Frontier LLM Models
by: Hayashi, Sergio Y., et al.
Published: (2026)
by: Hayashi, Sergio Y., et al.
Published: (2026)
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026)
by: Liu, Dongrui, et al.
Published: (2026)
OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
by: Thwal, Chu Myaet, et al.
Published: (2024)
by: Thwal, Chu Myaet, et al.
Published: (2024)
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
by: Liu, Zishan, et al.
Published: (2026)
by: Liu, Zishan, et al.
Published: (2026)
Image-Intrinsic Priors for Integrated Circuit Defect Detection and Novel Class Discovery via Self-Supervised Learning
by: Zhao, Botong., et al.
Published: (2025)
by: Zhao, Botong., et al.
Published: (2025)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026)
by: Chen, Siyi, et al.
Published: (2026)
CoV: Chain-of-View Prompting for Spatial Reasoning
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
by: Ni, Zanlin, et al.
Published: (2024)
by: Ni, Zanlin, et al.
Published: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning
by: Xu, Huilin, et al.
Published: (2025)
by: Xu, Huilin, et al.
Published: (2025)
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
by: Hu, Youbing, et al.
Published: (2024)
by: Hu, Youbing, et al.
Published: (2024)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
by: Pan, Zhenyu, et al.
Published: (2025)
by: Pan, Zhenyu, et al.
Published: (2025)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
by: Yu, Hu, et al.
Published: (2024)
by: Yu, Hu, et al.
Published: (2024)
Texture Image Synthesis Using Spatial GAN Based on Vision Transformers
by: Salari, Elahe, et al.
Published: (2025)
by: Salari, Elahe, et al.
Published: (2025)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
by: Wu, Junfei, et al.
Published: (2025)
by: Wu, Junfei, et al.
Published: (2025)
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps
by: Gao, Xiangjun, et al.
Published: (2026)
by: Gao, Xiangjun, et al.
Published: (2026)
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
by: Tang, Jianting, et al.
Published: (2025)
by: Tang, Jianting, et al.
Published: (2025)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
by: Li, Hongxing, et al.
Published: (2025)
by: Li, Hongxing, et al.
Published: (2025)
Algorithm Research of ELMo Word Embedding and Deep Learning Multimodal Transformer in Image Description
by: Cheng, Xiaohan, et al.
Published: (2024)
by: Cheng, Xiaohan, et al.
Published: (2024)
Exploring Non-Local Spatial-Angular Correlations with a Hybrid Mamba-Transformer Framework for Light Field Super-Resolution
by: Liu, Haosong, et al.
Published: (2025)
by: Liu, Haosong, et al.
Published: (2025)
DiMSUM: Diffusion Mamba -- A Scalable and Unified Spatial-Frequency Method for Image Generation
by: Phung, Hao, et al.
Published: (2024)
by: Phung, Hao, et al.
Published: (2024)
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
by: Guo, Xingang, et al.
Published: (2025)
by: Guo, Xingang, et al.
Published: (2025)
SentiFormer: Metadata Enhanced Transformer for Image Sentiment Analysis
by: Feng, Bin, et al.
Published: (2025)
by: Feng, Bin, et al.
Published: (2025)
UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Science
by: Zhang, Jie, et al.
Published: (2026)
by: Zhang, Jie, et al.
Published: (2026)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
by: Duan, Chengqi, et al.
Published: (2025)
by: Duan, Chengqi, et al.
Published: (2025)
Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
VERDI: VLM-Embedded Reasoning for Autonomous Driving
by: Feng, Bowen, et al.
Published: (2025)
by: Feng, Bowen, et al.
Published: (2025)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
Sub-token ViT Embedding via Stochastic Resonance Transformers
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
Latent Bias Alignment for High-Fidelity Diffusion Inversion in Real-World Image Reconstruction and Manipulation
by: Chen, Weiming, et al.
Published: (2026)
by: Chen, Weiming, et al.
Published: (2026)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
by: Feng, Jie, et al.
Published: (2025)
by: Feng, Jie, et al.
Published: (2025)
Make Geometry Matter for Spatial Reasoning
by: Zhang, Shihua, et al.
Published: (2026)
by: Zhang, Shihua, et al.
Published: (2026)
Geometrically-Constrained Agent for Spatial Reasoning
by: Chen, Zeren, et al.
Published: (2025)
by: Chen, Zeren, et al.
Published: (2025)
Similar Items
-
On Computational Limits of FlowAR Models: Expressivity and Efficiency
by: Cao, Yang, et al.
Published: (2025) -
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
by: Ke, Yekun, et al.
Published: (2025) -
Automatic Channel Pruning for Multi-Head Attention
by: Lee, Eunho, et al.
Published: (2024) -
fruit-SALAD: A Style Aligned Artwork Dataset to reveal similarity perception in image embeddings
by: Ohm, Tillmann, et al.
Published: (2024) -
Limits of Spatial Imagery Reasoning in Frontier LLM Models
by: Hayashi, Sergio Y., et al.
Published: (2026)