GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Tong, Yang, Guandao, Li, Zhibing, Zhang, Kai, Liu, Ziwei, Guibas, Leonidas, Lin, Dahua, Wetzstein, Gordon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
by: Huang, Ian, et al.
Published: (2024)
by: Huang, Ian, et al.
Published: (2024)
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
by: Deng, Youming, et al.
Published: (2025)
by: Deng, Youming, et al.
Published: (2025)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
by: Zhang, Yuhan, et al.
Published: (2025)
by: Zhang, Yuhan, et al.
Published: (2025)
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
by: Ackermann, Jan, et al.
Published: (2025)
by: Ackermann, Jan, et al.
Published: (2025)
Robust Symmetry Detection via Riemannian Langevin Dynamics
by: Je, Jihyeon, et al.
Published: (2024)
by: Je, Jihyeon, et al.
Published: (2024)
InfoGaussian: Structure-Aware Dynamic Gaussians through Lightweight Information Shaping
by: Zhang, Yunchao, et al.
Published: (2024)
by: Zhang, Yunchao, et al.
Published: (2024)
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
by: Zhang, Yuhan, et al.
Published: (2025)
by: Zhang, Yuhan, et al.
Published: (2025)
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
by: Zheng, Yang, et al.
Published: (2025)
by: Zheng, Yang, et al.
Published: (2025)
PhysAvatar: Learning the Physics of Dressed 3D Avatars from Visual Observations
by: Zheng, Yang, et al.
Published: (2024)
by: Zheng, Yang, et al.
Published: (2024)
pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
by: Chen, Hansheng, et al.
Published: (2025)
by: Chen, Hansheng, et al.
Published: (2025)
Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
by: Fang, Ye, et al.
Published: (2024)
by: Fang, Ye, et al.
Published: (2024)
Animal Pose Labeling Using General-Purpose Point Trackers
by: Pan, Zhuoyang, et al.
Published: (2025)
by: Pan, Zhuoyang, et al.
Published: (2025)
4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling
by: Bahmani, Sherwin, et al.
Published: (2023)
by: Bahmani, Sherwin, et al.
Published: (2023)
SS4D: Native 4D Generative Model via Structured Spacetime Latents
by: Li, Zhibing, et al.
Published: (2025)
by: Li, Zhibing, et al.
Published: (2025)
Harnessing GPT-4V(ision) for Insurance: A Preliminary Exploration
by: Lin, Chenwei, et al.
Published: (2024)
by: Lin, Chenwei, et al.
Published: (2024)
Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
by: Yang, Zhengyuan, et al.
Published: (2023)
by: Yang, Zhengyuan, et al.
Published: (2023)
Asymmetric Flow Models
by: Chen, Hansheng, et al.
Published: (2026)
by: Chen, Hansheng, et al.
Published: (2026)
Orthogonal Adaptation for Modular Customization of Diffusion Models
by: Po, Ryan, et al.
Published: (2023)
by: Po, Ryan, et al.
Published: (2023)
Diffusion Self-Distillation for Zero-Shot Customized Image Generation
by: Cai, Shengqu, et al.
Published: (2024)
by: Cai, Shengqu, et al.
Published: (2024)
GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
by: Wake, Naoki, et al.
Published: (2023)
by: Wake, Naoki, et al.
Published: (2023)
Video World Models with Long-term Spatial Memory
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
by: Deng, Boyang, et al.
Published: (2024)
by: Deng, Boyang, et al.
Published: (2024)
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
by: Chen, Hansheng, et al.
Published: (2024)
by: Chen, Hansheng, et al.
Published: (2024)
GPT-4V(ision) Unsuitable for Clinical Care and Education: A Clinician-Evaluated Assessment
by: Senkaiahliyan, Senthujan, et al.
Published: (2023)
by: Senkaiahliyan, Senthujan, et al.
Published: (2023)
AIpparel: A Multimodal Foundation Model for Digital Garments
by: Nakayama, Kiyohiro, et al.
Published: (2024)
by: Nakayama, Kiyohiro, et al.
Published: (2024)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024)
by: Zheng, Boyuan, et al.
Published: (2024)
Neural Control Variates with Automatic Integration
by: Li, Zilu, et al.
Published: (2024)
by: Li, Zilu, et al.
Published: (2024)
GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance Modeling
by: Zheng, Yang, et al.
Published: (2025)
by: Zheng, Yang, et al.
Published: (2025)
3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation
by: Chen, Hansheng, et al.
Published: (2024)
by: Chen, Hansheng, et al.
Published: (2024)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
by: Kuang, Zhengfei, et al.
Published: (2024)
by: Kuang, Zhengfei, et al.
Published: (2024)
HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting
by: Liu, Xian, et al.
Published: (2023)
by: Liu, Xian, et al.
Published: (2023)
How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation
by: Han, Zhongyi, et al.
Published: (2023)
by: Han, Zhongyi, et al.
Published: (2023)
Dynamic Gaussian Marbles for Novel View Synthesis of Casual Monocular Videos
by: Stearns, Colton, et al.
Published: (2024)
by: Stearns, Colton, et al.
Published: (2024)
NeRF Revisited: Fixing Quadrature Instability in Volume Rendering
by: Uy, Mikaela Angelina, et al.
Published: (2023)
by: Uy, Mikaela Angelina, et al.
Published: (2023)
Gaussian Mixture Flow Matching Models
by: Chen, Hansheng, et al.
Published: (2025)
by: Chen, Hansheng, et al.
Published: (2025)
OCH3R: Object-Centric Holistic 3D Reconstruction
by: Du, Yi, et al.
Published: (2026)
by: Du, Yi, et al.
Published: (2026)
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
by: Deng, Boyang, et al.
Published: (2025)
by: Deng, Boyang, et al.
Published: (2025)
Similar Items
-
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
by: Huang, Ian, et al.
Published: (2024) -
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
by: Wu, Tong, et al.
Published: (2024) -
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
by: Deng, Youming, et al.
Published: (2025) -
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
by: Zhang, Yuhan, et al.
Published: (2025) -
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
by: Yang, Shuai, et al.
Published: (2024)