3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Hao, Huang, Ting, Zhang, Zeyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
3D CoCa: Contrastive Learners are 3D Captioners
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
di: Chen, Yixiong, et al.
Pubblicazione: (2025)
di: Chen, Yixiong, et al.
Pubblicazione: (2025)
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
di: Patock, Jake R., et al.
Pubblicazione: (2025)
di: Patock, Jake R., et al.
Pubblicazione: (2025)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners
di: Ling, Lu, et al.
Pubblicazione: (2025)
di: Ling, Lu, et al.
Pubblicazione: (2025)
UniMesh: Unifying 3D Mesh Understanding and Generation
di: Huang, Peng, et al.
Pubblicazione: (2026)
di: Huang, Peng, et al.
Pubblicazione: (2026)
DragMesh: Interactive 3D Generation Made Easy
di: Zhang, Tianshan, et al.
Pubblicazione: (2025)
di: Zhang, Tianshan, et al.
Pubblicazione: (2025)
GS-Bias: Global-Spatial Bias Learner for Single-Image Test-Time Adaptation of Vision-Language Models
di: Huang, Zhaohong, et al.
Pubblicazione: (2025)
di: Huang, Zhaohong, et al.
Pubblicazione: (2025)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
di: Li, Peize, et al.
Pubblicazione: (2026)
di: Li, Peize, et al.
Pubblicazione: (2026)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
di: Ma, Wufei, et al.
Pubblicazione: (2025)
di: Ma, Wufei, et al.
Pubblicazione: (2025)
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
di: Gao, Yuanyuan, et al.
Pubblicazione: (2026)
di: Gao, Yuanyuan, et al.
Pubblicazione: (2026)
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
di: Zhang, Nonghai, et al.
Pubblicazione: (2025)
di: Zhang, Nonghai, et al.
Pubblicazione: (2025)
Nav-R1: Reasoning and Navigation in Embodied Scenes
di: Liu, Qingxiang, et al.
Pubblicazione: (2025)
di: Liu, Qingxiang, et al.
Pubblicazione: (2025)
Code2Worlds: Empowering Coding LLMs for 4D World Generation
di: Zhang, Yi, et al.
Pubblicazione: (2026)
di: Zhang, Yi, et al.
Pubblicazione: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
di: Ma, Guoqing, et al.
Pubblicazione: (2026)
di: Ma, Guoqing, et al.
Pubblicazione: (2026)
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
di: Liu, Fangfu, et al.
Pubblicazione: (2026)
di: Liu, Fangfu, et al.
Pubblicazione: (2026)
Test-Time 3D Occupancy Prediction
di: Zhang, Fengyi, et al.
Pubblicazione: (2025)
di: Zhang, Fengyi, et al.
Pubblicazione: (2025)
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing
di: Sun, Xiaokun, et al.
Pubblicazione: (2026)
di: Sun, Xiaokun, et al.
Pubblicazione: (2026)
Light4D: Training-Free Extreme Viewpoint 4D Video Relighting
di: Wu, Zhenghuang, et al.
Pubblicazione: (2026)
di: Wu, Zhenghuang, et al.
Pubblicazione: (2026)
Contrastive Masked Autoencoders are Stronger Vision Learners
di: Huang, Zhicheng, et al.
Pubblicazione: (2022)
di: Huang, Zhicheng, et al.
Pubblicazione: (2022)
4D Contrastive Superflows are Dense 3D Representation Learners
di: Xu, Xiang, et al.
Pubblicazione: (2024)
di: Xu, Xiang, et al.
Pubblicazione: (2024)
Generalizable Sparse-View 3D Reconstruction from Unconstrained Images
di: Gupta, Vinayak, et al.
Pubblicazione: (2026)
di: Gupta, Vinayak, et al.
Pubblicazione: (2026)
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
di: Jin, Li, et al.
Pubblicazione: (2026)
di: Jin, Li, et al.
Pubblicazione: (2026)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
di: Gao, Yipeng, et al.
Pubblicazione: (2023)
di: Gao, Yipeng, et al.
Pubblicazione: (2023)
StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation
di: Ren, Zeyu, et al.
Pubblicazione: (2026)
di: Ren, Zeyu, et al.
Pubblicazione: (2026)
GGHead: Fast and Generalizable 3D Gaussian Heads
di: Kirschstein, Tobias, et al.
Pubblicazione: (2024)
di: Kirschstein, Tobias, et al.
Pubblicazione: (2024)
MultiCo3D: Multi-Label Voxel Contrast for One-Shot Incremental Segmentation of 3D Neuroimages
di: Xu, Hao, et al.
Pubblicazione: (2025)
di: Xu, Hao, et al.
Pubblicazione: (2025)
HuGDiffusion: Generalizable Single-Image Human Rendering via 3D Gaussian Diffusion
di: Tang, Yingzhi, et al.
Pubblicazione: (2025)
di: Tang, Yingzhi, et al.
Pubblicazione: (2025)
CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners
di: Yao, Yunzhi, et al.
Pubblicazione: (2025)
di: Yao, Yunzhi, et al.
Pubblicazione: (2025)
MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss
di: Zhao, Can, et al.
Pubblicazione: (2025)
di: Zhao, Can, et al.
Pubblicazione: (2025)
CaMML: Context-Aware Multimodal Learner for Large Models
di: Chen, Yixin, et al.
Pubblicazione: (2024)
di: Chen, Yixin, et al.
Pubblicazione: (2024)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
di: Wang, Pan, et al.
Pubblicazione: (2025)
di: Wang, Pan, et al.
Pubblicazione: (2025)
SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
di: Fedele, Elisabetta, et al.
Pubblicazione: (2025)
di: Fedele, Elisabetta, et al.
Pubblicazione: (2025)
SegTTA: Training-Free Test-Time Augmentation for Zero-Shot Medical Imaging Segmentation
di: Yao, Yihong, et al.
Pubblicazione: (2026)
di: Yao, Yihong, et al.
Pubblicazione: (2026)
HSG: Hyperbolic Scene Graph
di: Wang, Liyang, et al.
Pubblicazione: (2026)
di: Wang, Liyang, et al.
Pubblicazione: (2026)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
Generalizable Synthetic Image Detection via Language-guided Contrastive Learning
di: Wu, Haiwei, et al.
Pubblicazione: (2023)
di: Wu, Haiwei, et al.
Pubblicazione: (2023)
Reconstructing 4D Spatial Intelligence: A Survey
di: Cao, Yukang, et al.
Pubblicazione: (2025)
di: Cao, Yukang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
3D CoCa: Contrastive Learners are 3D Captioners
di: Huang, Ting, et al.
Pubblicazione: (2025) -
CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
di: Chen, Yixiong, et al.
Pubblicazione: (2025) -
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
di: Patock, Jake R., et al.
Pubblicazione: (2025) -
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025) -
I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners
di: Ling, Lu, et al.
Pubblicazione: (2025)