Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Xueying, Li, Wenhao, Qian, Quanhao, Zhao, Deli, Lu, Shijian, Zhang, Gongjie, Xu, Ran |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026)
by: Zhang, Gongjie, et al.
Published: (2026)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
by: Wang, Jiuniu, et al.
Published: (2025)
by: Wang, Jiuniu, et al.
Published: (2025)
Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting
by: Miao, Xingyu, et al.
Published: (2025)
by: Miao, Xingyu, et al.
Published: (2025)
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning
by: Jiang, Xueying, et al.
Published: (2025)
by: Jiang, Xueying, et al.
Published: (2025)
Multimodal 3D Reasoning Segmentation with Complex Scenes
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
Weakly Supervised Monocular 3D Detection with a Single-View Image
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
Modeling Continuous Motion for 3D Point Cloud Object Tracking
by: Luo, Zhipeng, et al.
Published: (2023)
by: Luo, Zhipeng, et al.
Published: (2023)
MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked Autoencoders
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
by: Qian, Quanhao, et al.
Published: (2025)
by: Qian, Quanhao, et al.
Published: (2025)
LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
by: Jin, Sheng, et al.
Published: (2024)
by: Jin, Sheng, et al.
Published: (2024)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
by: Nie, Jiahao, et al.
Published: (2026)
by: Nie, Jiahao, et al.
Published: (2026)
RCDN: Towards Robust Camera-Insensitivity Collaborative Perception via Dynamic Feature-based 3D Neural Modeling
by: Wang, Tianhang, et al.
Published: (2024)
by: Wang, Tianhang, et al.
Published: (2024)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
by: Yin, Weijie, et al.
Published: (2025)
by: Yin, Weijie, et al.
Published: (2025)
MUC: Mixture of Uncalibrated Cameras for Robust 3D Human Body Reconstruction
by: Zhu, Yitao, et al.
Published: (2024)
by: Zhu, Yitao, et al.
Published: (2024)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
by: Nie, Jiahao, et al.
Published: (2024)
by: Nie, Jiahao, et al.
Published: (2024)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
TrafficLoc: Localizing Traffic Surveillance Cameras in 3D Scenes
by: Xia, Yan, et al.
Published: (2024)
by: Xia, Yan, et al.
Published: (2024)
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
FreGS: 3D Gaussian Splatting with Progressive Frequency Regularization
by: Zhang, Jiahui, et al.
Published: (2024)
by: Zhang, Jiahui, et al.
Published: (2024)
EventFlash: Towards Efficient MLLMs for Event-Based Vision
by: Liu, Shaoyu, et al.
Published: (2026)
by: Liu, Shaoyu, et al.
Published: (2026)
L3DR: 3D-aware LiDAR Diffusion and Rectification
by: Liu, Quan, et al.
Published: (2026)
by: Liu, Quan, et al.
Published: (2026)
On Accurate and Robust Estimation of 3D and 2D Circular Center: Method and Application to Camera-Lidar Calibration
by: Jiang, Jiajun, et al.
Published: (2025)
by: Jiang, Jiajun, et al.
Published: (2025)
Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence Mining
by: Nie, Jiahao, et al.
Published: (2024)
by: Nie, Jiahao, et al.
Published: (2024)
RynnEC: Bringing MLLMs into Embodied World
by: Dang, Ronghao, et al.
Published: (2025)
by: Dang, Ronghao, et al.
Published: (2025)
SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation
by: Xu, Muyu, et al.
Published: (2025)
by: Xu, Muyu, et al.
Published: (2025)
Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
by: Tian, Huiyuan, et al.
Published: (2025)
by: Tian, Huiyuan, et al.
Published: (2025)
A Survey of Label-Efficient Deep Learning for 3D Point Clouds
by: Xiao, Aoran, et al.
Published: (2023)
by: Xiao, Aoran, et al.
Published: (2023)
StyleGaussian: Instant 3D Style Transfer with Gaussian Splatting
by: Liu, Kunhao, et al.
Published: (2024)
by: Liu, Kunhao, et al.
Published: (2024)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
SELU: Self-Learning Embodied MLLMs in Unknown Environments
by: Li, Boyu, et al.
Published: (2024)
by: Li, Boyu, et al.
Published: (2024)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
by: Lu, Junxin, et al.
Published: (2026)
by: Lu, Junxin, et al.
Published: (2026)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
by: Jiang, Dongfu, et al.
Published: (2025)
by: Jiang, Dongfu, et al.
Published: (2025)
Similar Items
-
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026) -
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
by: Wang, Jiuniu, et al.
Published: (2025) -
Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting
by: Miao, Xingyu, et al.
Published: (2025) -
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
by: Li, Wenhao, et al.
Published: (2026) -
Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning
by: Jiang, Xueying, et al.
Published: (2025)