4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhangquan, Zhang, Manyuan, Yu, Xinlei, An, Xiang, Li, Bo, Xie, Xin, Wang, ZiDong, Sun, Mingze, Chen, Shuang, Li, Hongyu, Hu, Xiaobin, Huang, Ruqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
NFR: Neural Feature-Guided Non-Rigid Shape Registration
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
von: Peng, Hongxing, et al.
Veröffentlicht: (2025)
von: Peng, Hongxing, et al.
Veröffentlicht: (2025)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025)
von: Li, Huibin, et al.
Veröffentlicht: (2025)
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
von: Meng, Zi, et al.
Veröffentlicht: (2026)
von: Meng, Zi, et al.
Veröffentlicht: (2026)
Collaborative AI Enhances Image Understanding in Materials Science
von: Yin, Ruoyan Avery, et al.
Veröffentlicht: (2025)
von: Yin, Ruoyan Avery, et al.
Veröffentlicht: (2025)
Multi-modal Sensor Fusion for Auto Driving Perception: A Survey
von: Huang, Keli, et al.
Veröffentlicht: (2022)
von: Huang, Keli, et al.
Veröffentlicht: (2022)
PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance
von: Satish, Siddarth Nilol Kundur, et al.
Veröffentlicht: (2026)
von: Satish, Siddarth Nilol Kundur, et al.
Veröffentlicht: (2026)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
von: Wu, Kunlin, et al.
Veröffentlicht: (2026)
von: Wu, Kunlin, et al.
Veröffentlicht: (2026)
A Novel Dataset for Flood Detection Robust to Seasonal Changes in Satellite Imagery
von: Jang, Youngsun, et al.
Veröffentlicht: (2025)
von: Jang, Youngsun, et al.
Veröffentlicht: (2025)
BOOST: Out-of-Distribution-Informed Adaptive Sampling for Bias Mitigation in Stylistic Convolutional Neural Networks
von: Vijendran, Mridula, et al.
Veröffentlicht: (2025)
von: Vijendran, Mridula, et al.
Veröffentlicht: (2025)
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
von: Zhu, Chenglin, et al.
Veröffentlicht: (2025)
von: Zhu, Chenglin, et al.
Veröffentlicht: (2025)
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
von: Chen, Yiteng, et al.
Veröffentlicht: (2025)
von: Chen, Yiteng, et al.
Veröffentlicht: (2025)
STimage-1K4M: A histopathology image-gene expression dataset for spatial transcriptomics
von: Chen, Jiawen, et al.
Veröffentlicht: (2024)
von: Chen, Jiawen, et al.
Veröffentlicht: (2024)
Advanced Long-term Earth System Forecasting
von: Wu, Hao, et al.
Veröffentlicht: (2025)
von: Wu, Hao, et al.
Veröffentlicht: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
von: Riva, Paolo, et al.
Veröffentlicht: (2026)
von: Riva, Paolo, et al.
Veröffentlicht: (2026)
Group Multi-View Transformer for 3D Shape Analysis with Spatial Encoding
von: Xu, Lixiang, et al.
Veröffentlicht: (2023)
von: Xu, Lixiang, et al.
Veröffentlicht: (2023)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
von: Li, Yayuan, et al.
Veröffentlicht: (2025)
von: Li, Yayuan, et al.
Veröffentlicht: (2025)
Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
von: Shi, Hongrui, et al.
Veröffentlicht: (2025)
von: Shi, Hongrui, et al.
Veröffentlicht: (2025)
DV-Matcher: Deformation-based Non-Rigid Point Cloud Matching Guided by Pre-trained Visual Features
von: Chen, Zhangquan, et al.
Veröffentlicht: (2024)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2024)
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
von: Xiao, Jiasong, et al.
Veröffentlicht: (2026)
von: Xiao, Jiasong, et al.
Veröffentlicht: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
4D Synchronized Fields: Motion-Language Gaussian Splatting for Temporal Scene Understanding
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2026)
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2026)
Flexible-weighted Chamfer Distance: Enhanced Objective Function for Point Cloud Completion
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
FlightScope: An Experimental Comparative Review of Aircraft Detection Algorithms in Satellite Imagery
von: Ghazouali, Safouane El, et al.
Veröffentlicht: (2024)
von: Ghazouali, Safouane El, et al.
Veröffentlicht: (2024)
Aximorphic Perspective Projection Model for Immersive Imagery
von: Fober, Jakub Maksymilian
Veröffentlicht: (2021)
von: Fober, Jakub Maksymilian
Veröffentlicht: (2021)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
von: Zhang, Sinin, et al.
Veröffentlicht: (2026)
von: Zhang, Sinin, et al.
Veröffentlicht: (2026)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
Web-Scale Collection of Video Data for 4D Animal Reconstruction
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2025)
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
A Simple Baseline for Streaming Video Understanding
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025)
An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
von: Li, Zhi, et al.
Veröffentlicht: (2025)
von: Li, Zhi, et al.
Veröffentlicht: (2025)
SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model
von: Li, Xinqing, et al.
Veröffentlicht: (2025)
von: Li, Xinqing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025) -
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025) -
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026) -
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025) -
NFR: Neural Feature-Guided Non-Rigid Shape Registration
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)