Saved in:
| Main Authors: | Pan, Boxiao, Xu, Zhan, Huang, Chun-Hao Paul, Singh, Krishna Kumar, Zhou, Yang, Guibas, Leonidas J., Yang, Jimei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2401.10822 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Animal Pose Labeling Using General-Purpose Point Trackers
by: Pan, Zhuoyang, et al.
Published: (2025)
by: Pan, Zhuoyang, et al.
Published: (2025)
LookOut: Real-World Humanoid Egocentric Navigation
by: Pan, Boxiao, et al.
Published: (2025)
by: Pan, Boxiao, et al.
Published: (2025)
SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes
by: Maillard, Léopold, et al.
Published: (2026)
by: Maillard, Léopold, et al.
Published: (2026)
InfoGaussian: Structure-Aware Dynamic Gaussians through Lightweight Information Shaping
by: Zhang, Yunchao, et al.
Published: (2024)
by: Zhang, Yunchao, et al.
Published: (2024)
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
by: Huang, Ian, et al.
Published: (2024)
by: Huang, Ian, et al.
Published: (2024)
PASTA: Controllable Part-Aware Shape Generation with Autoregressive Transformers
by: Li, Songlin, et al.
Published: (2024)
by: Li, Songlin, et al.
Published: (2024)
Synergistic Global-space Camera and Human Reconstruction from Videos
by: Zhao, Yizhou, et al.
Published: (2024)
by: Zhao, Yizhou, et al.
Published: (2024)
GenAnalysis: Joint Shape Analysis by Learning Man-Made Shape Generators with Deformation Regularizations
by: Yang, Yuezhi, et al.
Published: (2025)
by: Yang, Yuezhi, et al.
Published: (2025)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
by: Kuang, Zhengfei, et al.
Published: (2024)
by: Kuang, Zhengfei, et al.
Published: (2024)
MultiPhys: Multi-Person Physics-aware 3D Motion Estimation
by: Ugrinovic, Nicolas, et al.
Published: (2024)
by: Ugrinovic, Nicolas, et al.
Published: (2024)
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
by: Lei, Jiahui, et al.
Published: (2025)
by: Lei, Jiahui, et al.
Published: (2025)
OCH3R: Object-Centric Holistic 3D Reconstruction
by: Du, Yi, et al.
Published: (2026)
by: Du, Yi, et al.
Published: (2026)
Dynamic Reflections: Probing Video Representations with Text Alignment
by: Zhu, Tyler, et al.
Published: (2025)
by: Zhu, Tyler, et al.
Published: (2025)
NeRF Revisited: Fixing Quadrature Instability in Volume Rendering
by: Uy, Mikaela Angelina, et al.
Published: (2023)
by: Uy, Mikaela Angelina, et al.
Published: (2023)
Template-Free Single-View 3D Human Digitalization with Diffusion-Guided LRM
by: Weng, Zhenzhen, et al.
Published: (2024)
by: Weng, Zhenzhen, et al.
Published: (2024)
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
by: Deng, Boyang, et al.
Published: (2024)
by: Deng, Boyang, et al.
Published: (2024)
AnimateAnywhere: Rouse the Background in Human Image Animation
by: Liu, Xiaoyu, et al.
Published: (2025)
by: Liu, Xiaoyu, et al.
Published: (2025)
Mixture of Contexts for Long Video Generation
by: Cai, Shengqu, et al.
Published: (2025)
by: Cai, Shengqu, et al.
Published: (2025)
Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
by: You, Yang, et al.
Published: (2024)
by: You, Yang, et al.
Published: (2024)
MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
by: Lei, Jiahui, et al.
Published: (2024)
by: Lei, Jiahui, et al.
Published: (2024)
SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
by: Fedele, Elisabetta, et al.
Published: (2025)
by: Fedele, Elisabetta, et al.
Published: (2025)
Hearing Anywhere in Any Environment
by: Liu, Xiulong, et al.
Published: (2025)
by: Liu, Xiulong, et al.
Published: (2025)
Video Perception Models for 3D Scene Synthesis
by: Huang, Rui, et al.
Published: (2025)
by: Huang, Rui, et al.
Published: (2025)
pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
by: Chen, Hansheng, et al.
Published: (2025)
by: Chen, Hansheng, et al.
Published: (2025)
Support-Set Context Matters for Bongard Problems
by: Raghuraman, Nikhil, et al.
Published: (2023)
by: Raghuraman, Nikhil, et al.
Published: (2023)
Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface Networks
by: Dedhia, Bhishma, et al.
Published: (2025)
by: Dedhia, Bhishma, et al.
Published: (2025)
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
by: Zheng, Yang, et al.
Published: (2025)
by: Zheng, Yang, et al.
Published: (2025)
SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
by: Wang, Qianxu, et al.
Published: (2023)
by: Wang, Qianxu, et al.
Published: (2023)
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
by: Deng, Youming, et al.
Published: (2025)
by: Deng, Youming, et al.
Published: (2025)
Refining Pre-Trained Motion Models
by: Sun, Xinglong, et al.
Published: (2024)
by: Sun, Xinglong, et al.
Published: (2024)
ProvNeRF: Modeling per Point Provenance in NeRFs as a Stochastic Field
by: Nakayama, Kiyohiro, et al.
Published: (2024)
by: Nakayama, Kiyohiro, et al.
Published: (2024)
GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance Modeling
by: Zheng, Yang, et al.
Published: (2025)
by: Zheng, Yang, et al.
Published: (2025)
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
by: Ren, Yixuan, et al.
Published: (2024)
by: Ren, Yixuan, et al.
Published: (2024)
The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space
by: Hu, Xia, et al.
Published: (2026)
by: Hu, Xia, et al.
Published: (2026)
Dynamic Gaussian Marbles for Novel View Synthesis of Casual Monocular Videos
by: Stearns, Colton, et al.
Published: (2024)
by: Stearns, Colton, et al.
Published: (2024)
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
by: Lee, Phillip Y., et al.
Published: (2025)
by: Lee, Phillip Y., et al.
Published: (2025)
FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
by: Huang, Ian, et al.
Published: (2025)
by: Huang, Ian, et al.
Published: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Change-Aware Siamese Network for Surface Defects Segmentation under Complex Background
by: Liu, Biyuan, et al.
Published: (2024)
by: Liu, Biyuan, et al.
Published: (2024)
PACE: A Large-Scale Dataset with Pose Annotations in Cluttered Environments
by: You, Yang, et al.
Published: (2023)
by: You, Yang, et al.
Published: (2023)
Similar Items
-
Animal Pose Labeling Using General-Purpose Point Trackers
by: Pan, Zhuoyang, et al.
Published: (2025) -
LookOut: Real-World Humanoid Egocentric Navigation
by: Pan, Boxiao, et al.
Published: (2025) -
SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes
by: Maillard, Léopold, et al.
Published: (2026) -
InfoGaussian: Structure-Aware Dynamic Gaussians through Lightweight Information Shaping
by: Zhang, Yunchao, et al.
Published: (2024) -
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
by: Huang, Ian, et al.
Published: (2024)