Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Haoyuan, Zhou, Yanpeng, Gao, Yufei, Tang, Tao, Han, Jianhua, Yuan, Yujie, Chen, Dave Zhenyu, Bian, Jiawang, Xu, Hang, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
GS-CLIP: Gaussian Splatting for Contrastive Language-Image-3D Pretraining from Real-World Data
by: Li, Haoyuan, et al.
Published: (2024)
by: Li, Haoyuan, et al.
Published: (2024)
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
by: Jeddi, Ahmadreza, et al.
Published: (2026)
by: Jeddi, Ahmadreza, et al.
Published: (2026)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
by: Pan, Zhenyu, et al.
Published: (2025)
by: Pan, Zhenyu, et al.
Published: (2025)
Are VLMs Really Blind
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025)
by: Zhang, Nonghai, et al.
Published: (2025)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
3D Primitives are a Spatial Language for VLMs
by: Liu, Junze, et al.
Published: (2026)
by: Liu, Junze, et al.
Published: (2026)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
by: He, Jixuan, et al.
Published: (2026)
by: He, Jixuan, et al.
Published: (2026)
When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs
by: Dai, Aobotao, et al.
Published: (2025)
by: Dai, Aobotao, et al.
Published: (2025)
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
by: Chen, Jierun, et al.
Published: (2025)
by: Chen, Jierun, et al.
Published: (2025)
Monocular 3D Object Position Estimation with VLMs for Human-Robot Interaction
by: Wahl, Ari, et al.
Published: (2026)
by: Wahl, Ari, et al.
Published: (2026)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
by: Huang, Yiyang, et al.
Published: (2025)
by: Huang, Yiyang, et al.
Published: (2025)
DreamVTON: Customizing 3D Virtual Try-on with Personalized Diffusion Models
by: Xie, Zhenyu, et al.
Published: (2024)
by: Xie, Zhenyu, et al.
Published: (2024)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
by: Gavrikov, Paul, et al.
Published: (2025)
by: Gavrikov, Paul, et al.
Published: (2025)
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
by: Chen, Jingkun, et al.
Published: (2026)
by: Chen, Jingkun, et al.
Published: (2026)
LLaVA$^3$: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs
by: Petit, Doriand, et al.
Published: (2025)
by: Petit, Doriand, et al.
Published: (2025)
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
by: Pang, Zirui, et al.
Published: (2025)
by: Pang, Zirui, et al.
Published: (2025)
From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
by: Cao, Ang, et al.
Published: (2025)
by: Cao, Ang, et al.
Published: (2025)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
by: Wang, Shengao, et al.
Published: (2025)
by: Wang, Shengao, et al.
Published: (2025)
Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT
by: Sun, Peng, et al.
Published: (2026)
by: Sun, Peng, et al.
Published: (2026)
Collaboration: It Really Does Work!
by: Youssef, Jennifer L.
Published: (2005)
by: Youssef, Jennifer L.
Published: (2005)
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving
by: Tang, Zecong, et al.
Published: (2026)
by: Tang, Zecong, et al.
Published: (2026)
Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment
by: Cambrin, Daniele Rege, et al.
Published: (2024)
by: Cambrin, Daniele Rege, et al.
Published: (2024)
Should VLMs be Pre-trained with Image Data?
by: Keh, Sedrick, et al.
Published: (2025)
by: Keh, Sedrick, et al.
Published: (2025)
How Do VLAs Effectively Inherit from VLMs?
by: Zhang, Chuheng, et al.
Published: (2025)
by: Zhang, Chuheng, et al.
Published: (2025)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
by: Ma, Xianzheng, et al.
Published: (2026)
by: Ma, Xianzheng, et al.
Published: (2026)
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud Pretraining
by: Yan, Siming, et al.
Published: (2023)
by: Yan, Siming, et al.
Published: (2023)
AnchorSplat: Feed-Forward 3D Gaussian Splatting with 3D Geometric Priors
by: Zhang, Xiaoxue, et al.
Published: (2026)
by: Zhang, Xiaoxue, et al.
Published: (2026)
3D-GSRD: 3D Molecular Graph Auto-Encoder with Selective Re-mask Decoding
by: Wu, Chang, et al.
Published: (2025)
by: Wu, Chang, et al.
Published: (2025)
Hyper3D: Efficient 3D Representation via Hybrid Triplane and Octree Feature for Enhanced 3D Shape Variational Auto-Encoders
by: Guo, Jingyu, et al.
Published: (2025)
by: Guo, Jingyu, et al.
Published: (2025)
Bringing Your Portrait to 3D Presence
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
by: Li, Tang, et al.
Published: (2024)
by: Li, Tang, et al.
Published: (2024)
Similar Items
-
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
by: Li, Haoyuan, et al.
Published: (2025) -
GS-CLIP: Gaussian Splatting for Contrastive Language-Image-3D Pretraining from Real-World Data
by: Li, Haoyuan, et al.
Published: (2024) -
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
by: Jeddi, Ahmadreza, et al.
Published: (2026) -
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
by: Pan, Zhenyu, et al.
Published: (2025) -
Are VLMs Really Blind
by: Singh, Ayush, et al.
Published: (2024)