View Selection for 3D Captioning via Diffusion Ranking
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Tiange, Johnson, Justin, Lee, Honglak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Repurposing 2D Diffusion Models for 3D Shape Completion
by: He, Yao, et al.
Published: (2025)
by: He, Yao, et al.
Published: (2025)
View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View Selection
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement
by: Luo, Yihang, et al.
Published: (2024)
by: Luo, Yihang, et al.
Published: (2024)
Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation
by: Xiang, Tiange, et al.
Published: (2025)
by: Xiang, Tiange, et al.
Published: (2025)
Online 3D Gaussian Splatting Modeling with Novel View Selection
by: Lee, Byeonggwon, et al.
Published: (2025)
by: Lee, Byeonggwon, et al.
Published: (2025)
Dual Caption Preference Optimization for Diffusion Models
by: Saeidi, Amir, et al.
Published: (2025)
by: Saeidi, Amir, et al.
Published: (2025)
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Diffusion Models
by: Zhu, Ruishu, et al.
Published: (2025)
by: Zhu, Ruishu, et al.
Published: (2025)
Label-Efficient 3D Brain Segmentation via Complementary 2D Diffusion Models with Orthogonal Views
by: Cho, Jihoon, et al.
Published: (2024)
by: Cho, Jihoon, et al.
Published: (2024)
Bi-directional Contextual Attention for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
by: Qu, Huaizhi, et al.
Published: (2025)
by: Qu, Huaizhi, et al.
Published: (2025)
AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
by: Chen, Yutian, et al.
Published: (2026)
by: Chen, Yutian, et al.
Published: (2026)
Diffusion Low Rank Hybrid Reconstruction for Sparse View Medical Imaging
by: Deng, Zongyin, et al.
Published: (2025)
by: Deng, Zongyin, et al.
Published: (2025)
ViewCraft3D: High-Fidelity and View-Consistent 3D Vector Graphics Synthesis
by: Wang, Chuang, et al.
Published: (2025)
by: Wang, Chuang, et al.
Published: (2025)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
by: Jang, Yunseok, et al.
Published: (2025)
by: Jang, Yunseok, et al.
Published: (2025)
DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis
by: Gu, Yuming, et al.
Published: (2023)
by: Gu, Yuming, et al.
Published: (2023)
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism
by: Wang, Zhen, et al.
Published: (2023)
by: Wang, Zhen, et al.
Published: (2023)
3D CoCa: Contrastive Learners are 3D Captioners
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
MVLight: Relightable Text-to-3D Generation via Light-conditioned Multi-View Diffusion
by: Shim, Dongseok, et al.
Published: (2024)
by: Shim, Dongseok, et al.
Published: (2024)
DMAligner: Enhancing Image Alignment via Diffusion Model Based View Synthesis
by: Luo, Xinglong, et al.
Published: (2026)
by: Luo, Xinglong, et al.
Published: (2026)
CAT3D: Create Anything in 3D with Multi-View Diffusion Models
by: Gao, Ruiqi, et al.
Published: (2024)
by: Gao, Ruiqi, et al.
Published: (2024)
ExCap3D: Expressive 3D Scene Understanding via Object Captioning with Varying Detail
by: Yeshwanth, Chandan, et al.
Published: (2025)
by: Yeshwanth, Chandan, et al.
Published: (2025)
DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions
by: Lin, Xiaoyu, et al.
Published: (2025)
by: Lin, Xiaoyu, et al.
Published: (2025)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
by: Chu, Sanghyeok, et al.
Published: (2026)
by: Chu, Sanghyeok, et al.
Published: (2026)
Consistent-1-to-3: Consistent Image to 3D View Synthesis via Geometry-aware Diffusion Models
by: Ye, Jianglong, et al.
Published: (2023)
by: Ye, Jianglong, et al.
Published: (2023)
Sparse-View 3D Gaussian Splatting in the Wild
by: Park, Wongi, et al.
Published: (2026)
by: Park, Wongi, et al.
Published: (2026)
Towards Scalable Language-Image Pre-training for 3D Medical Imaging
by: Zhao, Chenhui, et al.
Published: (2025)
by: Zhao, Chenhui, et al.
Published: (2025)
SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation
by: Ren, Junlong, et al.
Published: (2025)
by: Ren, Junlong, et al.
Published: (2025)
OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery
by: Liang, Xiucheng, et al.
Published: (2025)
by: Liang, Xiucheng, et al.
Published: (2025)
From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis
by: Huang, Ranran, et al.
Published: (2026)
by: Huang, Ranran, et al.
Published: (2026)
MVPainter: Accurate and Detailed 3D Texture Generation via Multi-View Diffusion with Geometric Control
by: Shao, Mingqi, et al.
Published: (2025)
by: Shao, Mingqi, et al.
Published: (2025)
DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing
by: Ai, Jinxin, et al.
Published: (2026)
by: Ai, Jinxin, et al.
Published: (2026)
Masked Diffusion Captioning for Visual Feature Learning
by: Feng, Chao, et al.
Published: (2025)
by: Feng, Chao, et al.
Published: (2025)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D
by: Cheng, Wei, et al.
Published: (2024)
by: Cheng, Wei, et al.
Published: (2024)
OccFusion: Rendering Occluded Humans with Generative Diffusion Priors
by: Sun, Adam, et al.
Published: (2024)
by: Sun, Adam, et al.
Published: (2024)
View-Consistent Diffusion Representations for 3D-Consistent Video Generation
by: Danier, Duolikun, et al.
Published: (2025)
by: Danier, Duolikun, et al.
Published: (2025)
MVSDet: Multi-View Indoor 3D Object Detection via Efficient Plane Sweeps
by: Xu, Yating, et al.
Published: (2024)
by: Xu, Yating, et al.
Published: (2024)
Similar Items
-
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025) -
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025) -
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024) -
Repurposing 2D Diffusion Models for 3D Shape Completion
by: He, Yao, et al.
Published: (2025) -
View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View Selection
by: Zhang, Qi, et al.
Published: (2024)