3D CoCa: Contrastive Learners are 3D Captioners
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Ting, Zhang, Zeyu, Wang, Yemin, Tang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence
by: Tang, Hao, et al.
Published: (2026)
by: Tang, Hao, et al.
Published: (2026)
CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
by: Chen, Yixiong, et al.
Published: (2025)
by: Chen, Yixiong, et al.
Published: (2025)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
by: Patock, Jake R., et al.
Published: (2025)
by: Patock, Jake R., et al.
Published: (2025)
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025)
by: Zhang, Nonghai, et al.
Published: (2025)
UniMesh: Unifying 3D Mesh Understanding and Generation
by: Huang, Peng, et al.
Published: (2026)
by: Huang, Peng, et al.
Published: (2026)
DragMesh: Interactive 3D Generation Made Easy
by: Zhang, Tianshan, et al.
Published: (2025)
by: Zhang, Tianshan, et al.
Published: (2025)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
by: Li, Peize, et al.
Published: (2026)
by: Li, Peize, et al.
Published: (2026)
SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation
by: Ren, Junlong, et al.
Published: (2025)
by: Ren, Junlong, et al.
Published: (2025)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
by: Gao, Yipeng, et al.
Published: (2023)
by: Gao, Yipeng, et al.
Published: (2023)
MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
4D Contrastive Superflows are Dense 3D Representation Learners
by: Xu, Xiang, et al.
Published: (2024)
by: Xu, Xiang, et al.
Published: (2024)
Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective
by: Zhu, Duowang, et al.
Published: (2025)
by: Zhu, Duowang, et al.
Published: (2025)
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
by: Ma, Ziping, et al.
Published: (2024)
by: Ma, Ziping, et al.
Published: (2024)
MultiCo3D: Multi-Label Voxel Contrast for One-Shot Incremental Segmentation of 3D Neuroimages
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
by: Huang, Junming, et al.
Published: (2026)
by: Huang, Junming, et al.
Published: (2026)
Code2Worlds: Empowering Coding LLMs for 4D World Generation
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
Light4D: Training-Free Extreme Viewpoint 4D Video Relighting
by: Wu, Zhenghuang, et al.
Published: (2026)
by: Wu, Zhenghuang, et al.
Published: (2026)
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
by: Jin, Bu, et al.
Published: (2024)
by: Jin, Bu, et al.
Published: (2024)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
by: Tang, Changli, et al.
Published: (2026)
by: Tang, Changli, et al.
Published: (2026)
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
by: Jin, Li, et al.
Published: (2026)
by: Jin, Li, et al.
Published: (2026)
ContrastiveGaussian: High-Fidelity 3D Generation with Contrastive Learning and Gaussian Splatting
by: Liu, Junbang, et al.
Published: (2025)
by: Liu, Junbang, et al.
Published: (2025)
View Selection for 3D Captioning via Diffusion Ranking
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Bi-directional Contextual Attention for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion
by: Wen, Hao, et al.
Published: (2024)
by: Wen, Hao, et al.
Published: (2024)
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
by: Huang, Sheng-Yu, et al.
Published: (2026)
by: Huang, Sheng-Yu, et al.
Published: (2026)
Nav-R1: Reasoning and Navigation in Embodied Scenes
by: Liu, Qingxiang, et al.
Published: (2025)
by: Liu, Qingxiang, et al.
Published: (2025)
DEAP-3DSAM: Decoder Enhanced and Auto Prompt SAM for 3D Medical Image Segmentation
by: Chen, Fangda, et al.
Published: (2025)
by: Chen, Fangda, et al.
Published: (2025)
ExCap3D: Expressive 3D Scene Understanding via Object Captioning with Varying Detail
by: Yeshwanth, Chandan, et al.
Published: (2025)
by: Yeshwanth, Chandan, et al.
Published: (2025)
Tetrahedron Splatting for 3D Generation
by: Gu, Chun, et al.
Published: (2024)
by: Gu, Chun, et al.
Published: (2024)
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning
by: Wang, Xuan, et al.
Published: (2023)
by: Wang, Xuan, et al.
Published: (2023)
Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework
by: Chen, Yi-Ting, et al.
Published: (2025)
by: Chen, Yi-Ting, et al.
Published: (2025)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
Emu3.5: Native Multimodal Models are World Learners
by: Cui, Yufeng, et al.
Published: (2025)
by: Cui, Yufeng, et al.
Published: (2025)
3D-Consistent Human Avatars with Sparse Inputs via Gaussian Splatting and Contrastive Learning
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets
by: Li, Zeyu, et al.
Published: (2024)
by: Li, Zeyu, et al.
Published: (2024)
Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection
by: Yu, Yuyang, et al.
Published: (2025)
by: Yu, Yuyang, et al.
Published: (2025)
Similar Items
-
3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence
by: Tang, Hao, et al.
Published: (2026) -
CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
by: Chen, Yixiong, et al.
Published: (2025) -
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
by: Huang, Ting, et al.
Published: (2025) -
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
by: Patock, Jake R., et al.
Published: (2025) -
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025)