VP-LLM: Text-Driven 3D Volume Completion with Large Language Models through Patchification
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jianmeng, Liu, Yichen, Zhang, Yuyao, Meng, Zeyuan, Tai, Yu-Wing, Tang, Chi-Keung |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
Multimodal Generation of Animatable 3D Human Models with AvatarForge
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
ChatCam: Empowering Camera Control through Conversational AI
by: Liu, Xinhang, et al.
Published: (2024)
by: Liu, Xinhang, et al.
Published: (2024)
ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
SANeRF-HQ: Segment Anything for NeRF in High Quality
by: Liu, Yichen, et al.
Published: (2023)
by: Liu, Yichen, et al.
Published: (2023)
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
by: Huang-Menders, Alexander, et al.
Published: (2025)
by: Huang-Menders, Alexander, et al.
Published: (2025)
FED-NeRF: Achieve High 3D Consistency and Temporal Coherence for Face Video Editing on Dynamic NeRF
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
InceptionHuman: Controllable Prompt-to-NeRF for Photorealistic 3D Human Generation
by: Kao, Shiu-hong, et al.
Published: (2023)
by: Kao, Shiu-hong, et al.
Published: (2023)
Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning
by: Zhao, Yubo, et al.
Published: (2025)
by: Zhao, Yubo, et al.
Published: (2025)
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
C3LLM: Conditional Multimodal Content Generation Using Large Language Models
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
by: Liu, Xinhang, et al.
Published: (2023)
by: Liu, Xinhang, et al.
Published: (2023)
UltraGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
by: Zhang, Yuyao, et al.
Published: (2025)
by: Zhang, Yuyao, et al.
Published: (2025)
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
by: Ao, Yuzhuo, et al.
Published: (2026)
by: Ao, Yuzhuo, et al.
Published: (2026)
Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
by: Kao, Shiu-hong, et al.
Published: (2026)
by: Kao, Shiu-hong, et al.
Published: (2026)
Inpaint4DNeRF: Promptable Spatio-Temporal NeRF Inpainting with Generative Diffusion Models
by: Jiang, Han, et al.
Published: (2023)
by: Jiang, Han, et al.
Published: (2023)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
DragVideo: Interactive Drag-style Video Editing
by: Deng, Yufan, et al.
Published: (2023)
by: Deng, Yufan, et al.
Published: (2023)
UVRM: A Scalable 3D Reconstruction Model from Unposed Videos
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling
by: Mai, Ziyang, et al.
Published: (2026)
by: Mai, Ziyang, et al.
Published: (2026)
Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling
by: Liu, Xinhang, et al.
Published: (2024)
by: Liu, Xinhang, et al.
Published: (2024)
Pyramidal Patchification Flow for Visual Generation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing
by: Zhang, Yuyao, et al.
Published: (2026)
by: Zhang, Yuyao, et al.
Published: (2026)
GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
by: Zhou, Junwei, et al.
Published: (2025)
by: Zhou, Junwei, et al.
Published: (2025)
DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
by: Srivastava, Divyansh, et al.
Published: (2025)
by: Srivastava, Divyansh, et al.
Published: (2025)
Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation
by: Zhou, Junwei, et al.
Published: (2026)
by: Zhou, Junwei, et al.
Published: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
by: Mai, Ziyang, et al.
Published: (2025)
by: Mai, Ziyang, et al.
Published: (2025)
Safety of Multimodal Large Language Models on Images and Texts
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation
by: Wang, Anbang, et al.
Published: (2026)
by: Wang, Anbang, et al.
Published: (2026)
Distill Gold from Massive Ores: Bi-level Data Pruning towards Efficient Dataset Distillation
by: Xu, Yue, et al.
Published: (2023)
by: Xu, Yue, et al.
Published: (2023)
LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration
by: Zhang, Yuyao, et al.
Published: (2025)
by: Zhang, Yuyao, et al.
Published: (2025)
TRIM: Scalable 3D Gaussian Diffusion Inference with Temporal and Spatial Trimming
by: Yin, Zeyuan, et al.
Published: (2025)
by: Yin, Zeyuan, et al.
Published: (2025)
HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting
by: Liu, Xian, et al.
Published: (2023)
by: Liu, Xian, et al.
Published: (2023)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
by: Xu, Mingjie, et al.
Published: (2025)
by: Xu, Mingjie, et al.
Published: (2025)
Similar Items
-
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
by: Liu, Xinhang, et al.
Published: (2025) -
Multimodal Generation of Animatable 3D Human Models with AvatarForge
by: Liu, Xinhang, et al.
Published: (2025) -
Agentic 3D Scene Generation with Spatially Contextualized VLMs
by: Liu, Xinhang, et al.
Published: (2025) -
ChatCam: Empowering Camera Control through Conversational AI
by: Liu, Xinhang, et al.
Published: (2024) -
ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
by: Wang, Zixuan, et al.
Published: (2025)