Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Shijie, Ren, Hui, Weng, Yijia, Zhang, Shuwang, Wang, Zhen, Xu, Dejia, Fan, Zhiwen, You, Suya, Wang, Zhangyang, Guibas, Leonidas, Kadambi, Achuta |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
by: Zhou, Shijie, et al.
Published: (2023)
by: Zhou, Shijie, et al.
Published: (2023)
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
4K4DGen: Panoramic 4D Generation at 4K Resolution
by: Li, Renjie, et al.
Published: (2024)
by: Li, Renjie, et al.
Published: (2024)
MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
by: Lei, Jiahui, et al.
Published: (2024)
by: Lei, Jiahui, et al.
Published: (2024)
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
by: Jeon, Changwoo, et al.
Published: (2026)
by: Jeon, Changwoo, et al.
Published: (2026)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
by: Weng, Yijia, et al.
Published: (2025)
by: Weng, Yijia, et al.
Published: (2025)
LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS
by: Fan, Zhiwen, et al.
Published: (2023)
by: Fan, Zhiwen, et al.
Published: (2023)
Expressive Gaussian Human Avatars from Monocular RGB Video
by: Hu, Hezhen, et al.
Published: (2024)
by: Hu, Hezhen, et al.
Published: (2024)
Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
by: You, Yang, et al.
Published: (2024)
by: You, Yang, et al.
Published: (2024)
Dynamic Gaussian Marbles for Novel View Synthesis of Casual Monocular Videos
by: Stearns, Colton, et al.
Published: (2024)
by: Stearns, Colton, et al.
Published: (2024)
VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
by: He, Xuehai, et al.
Published: (2025)
by: He, Xuehai, et al.
Published: (2025)
View-Consistent Hierarchical 3D Segmentation Using Ultrametric Feature Fields
by: He, Haodi, et al.
Published: (2024)
by: He, Haodi, et al.
Published: (2024)
Robot Learning from Any Images
by: Zhao, Siheng, et al.
Published: (2025)
by: Zhao, Siheng, et al.
Published: (2025)
4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency
by: Yin, Yuyang, et al.
Published: (2023)
by: Yin, Yuyang, et al.
Published: (2023)
SparseGS: Real-Time 360° Sparse View Synthesis using Gaussian Splatting
by: Xiong, Haolin, et al.
Published: (2023)
by: Xiong, Haolin, et al.
Published: (2023)
SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
by: Wang, Qianxu, et al.
Published: (2023)
by: Wang, Qianxu, et al.
Published: (2023)
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety
by: Jalilian, Laleh, et al.
Published: (2024)
by: Jalilian, Laleh, et al.
Published: (2024)
Zero-Shot Image Feature Consensus with Deep Functional Maps
by: Cheng, Xinle, et al.
Published: (2024)
by: Cheng, Xinle, et al.
Published: (2024)
Solutions to Deepfakes: Can Camera Hardware, Cryptography, and Deep Learning Verify Real Images?
by: Vilesov, Alexander, et al.
Published: (2024)
by: Vilesov, Alexander, et al.
Published: (2024)
Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules
by: Liang, Yiqing, et al.
Published: (2024)
by: Liang, Yiqing, et al.
Published: (2024)
Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
by: Liang, Hanwen, et al.
Published: (2024)
by: Liang, Hanwen, et al.
Published: (2024)
InfoGaussian: Structure-Aware Dynamic Gaussians through Lightweight Information Shaping
by: Zhang, Yunchao, et al.
Published: (2024)
by: Zhang, Yunchao, et al.
Published: (2024)
INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation Processing
by: Abi-Karam, Stefan, et al.
Published: (2023)
by: Abi-Karam, Stefan, et al.
Published: (2023)
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024)
by: Weng, Yijia, et al.
Published: (2024)
Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
by: Wang, Penghao, et al.
Published: (2025)
by: Wang, Penghao, et al.
Published: (2025)
Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D
by: T, Mukund Varma, et al.
Published: (2024)
by: T, Mukund Varma, et al.
Published: (2024)
FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting
by: Zhu, Zehao, et al.
Published: (2023)
by: Zhu, Zehao, et al.
Published: (2023)
InstantRestore: Single-Step Personalized Face Restoration with Shared-Image Attention
by: Zhang, Howard, et al.
Published: (2024)
by: Zhang, Howard, et al.
Published: (2024)
LLM-AutoDiff: Auto-Differentiate Any LLM Workflow
by: Yin, Li, et al.
Published: (2025)
by: Yin, Li, et al.
Published: (2025)
OCH3R: Object-Centric Holistic 3D Reconstruction
by: Du, Yi, et al.
Published: (2026)
by: Du, Yi, et al.
Published: (2026)
PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
by: Deng, Youming, et al.
Published: (2025)
by: Deng, Youming, et al.
Published: (2025)
Comp4D: LLM-Guided Compositional 4D Scene Generation
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
Not Just Streaks: Towards Ground Truth for Single Image Deraining
by: Ba, Yunhao, et al.
Published: (2022)
by: Ba, Yunhao, et al.
Published: (2022)
4KAgent: Agentic Any Image to 4K Super-Resolution
by: Zuo, Yushen, et al.
Published: (2025)
by: Zuo, Yushen, et al.
Published: (2025)
Similar Items
-
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
by: Zhou, Shijie, et al.
Published: (2023) -
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024) -
4K4DGen: Panoramic 4D Generation at 4K Resolution
by: Li, Renjie, et al.
Published: (2024) -
MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
by: Lei, Jiahui, et al.
Published: (2024) -
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
by: Jeon, Changwoo, et al.
Published: (2026)