MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents
Fuente:
arXiv
Saved in:
| Main Authors: | Kwon, Minkyung, Choi, Jinhyeok, Shin, Youngjin, Kim, Jaeyeong, Lee, JongMin, Kim, Seungryong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Projected Representation Conditioning for High-fidelity Novel View Synthesis
by: Kwak, Min-Seop, et al.
Published: (2026)
by: Kwak, Min-Seop, et al.
Published: (2026)
Dense-SfM: Structure from Motion with Dense Consistent Matching
by: Lee, JongMin, et al.
Published: (2025)
by: Lee, JongMin, et al.
Published: (2025)
Repurposing Geometric Foundation Models for Multi-view Diffusion
by: Jang, Wooseok, et al.
Published: (2026)
by: Jang, Wooseok, et al.
Published: (2026)
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
by: Kwon, Minkyung, et al.
Published: (2025)
by: Kwon, Minkyung, et al.
Published: (2025)
Grounding World Simulation Models in a Real-World Metropolis
by: Seo, Junyoung, et al.
Published: (2026)
by: Seo, Junyoung, et al.
Published: (2026)
LatentSwap: An Efficient Latent Code Mapping Framework for Face Swapping
by: Choi, Changho, et al.
Published: (2024)
by: Choi, Changho, et al.
Published: (2024)
APPLE: Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping
by: Kang, Jiwon, et al.
Published: (2026)
by: Kang, Jiwon, et al.
Published: (2026)
CORAL: Correspondence Alignment for Improved Virtual Try-On
by: Kim, Jiyoung, et al.
Published: (2026)
by: Kim, Jiyoung, et al.
Published: (2026)
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
by: Choi, Changho, et al.
Published: (2025)
by: Choi, Changho, et al.
Published: (2025)
SpikeMatch: Semi-Supervised Learning with Temporal Dynamics of Spiking Neural Networks
by: Yang, Jini, et al.
Published: (2025)
by: Yang, Jini, et al.
Published: (2025)
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
by: Seo, Junyoung, et al.
Published: (2023)
by: Seo, Junyoung, et al.
Published: (2023)
MV-TAP: Tracking Any Point in Multi-View Videos
by: Koo, Jahyeok, et al.
Published: (2025)
by: Koo, Jahyeok, et al.
Published: (2025)
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
by: Nam, Jisu, et al.
Published: (2026)
by: Nam, Jisu, et al.
Published: (2026)
Hybrid Video Diffusion Models with 2D Triplane and 3D Wavelet Representation
by: Kim, Kihong, et al.
Published: (2024)
by: Kim, Kihong, et al.
Published: (2024)
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
by: Jin, Siyoon, et al.
Published: (2025)
by: Jin, Siyoon, et al.
Published: (2025)
Exploring Temporally-Aware Features for Point Tracking
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
Domain Generalization Using Large Pretrained Models with Mixture-of-Adapters
by: Lee, Gyuseong, et al.
Published: (2023)
by: Lee, Gyuseong, et al.
Published: (2023)
WaTeRFlow: Watermark Temporal Robustness via Flow Consistency
by: Jeong, Utae, et al.
Published: (2025)
by: Jeong, Utae, et al.
Published: (2025)
S^4M: Boosting Semi-Supervised Instance Segmentation with SAM
by: Yoon, Heeji, et al.
Published: (2025)
by: Yoon, Heeji, et al.
Published: (2025)
PLOT: Pseudo-Labeling via Video Object Tracking for Scalable Monocular 3D Object Detection
by: Lee, Seokyeong, et al.
Published: (2025)
by: Lee, Seokyeong, et al.
Published: (2025)
Geometry-Aware Score Distillation via 3D Consistent Noising and Gradient Consistency Modeling
by: Kwak, Min-Seop, et al.
Published: (2024)
by: Kwak, Min-Seop, et al.
Published: (2024)
InterRVOS: Interaction-aware Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2025)
by: Jin, Woojeong, et al.
Published: (2025)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
by: Shin, Youngwoo, et al.
Published: (2026)
by: Shin, Youngwoo, et al.
Published: (2026)
LightSplat: Fast and Memory-Efficient Open-Vocabulary 3D Scene Understanding in Five Seconds
by: Bang, Jaehun, et al.
Published: (2026)
by: Bang, Jaehun, et al.
Published: (2026)
Proxy-Free Gaussian Splats Deformation with Splat-Based Surface Estimation
by: Kim, Jaeyeong, et al.
Published: (2025)
by: Kim, Jaeyeong, et al.
Published: (2025)
Talk3D: High-Fidelity Talking Portrait Synthesis via Personalized 3D Generative Prior
by: Ko, Jaehoon, et al.
Published: (2024)
by: Ko, Jaehoon, et al.
Published: (2024)
Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
by: Kim, Jin Hyeon, et al.
Published: (2025)
by: Kim, Jin Hyeon, et al.
Published: (2025)
ILV: Iterative Latent Volumes for Fast and Accurate Sparse-View CT Reconstruction
by: Lee, Seungryong, et al.
Published: (2026)
by: Lee, Seungryong, et al.
Published: (2026)
Exploring Conditions for Diffusion models in Robotic Control
by: Shin, Heeseong, et al.
Published: (2025)
by: Shin, Heeseong, et al.
Published: (2025)
Multi-Modal Guided Multi-Source Domain Adaptation for Object Detection
by: Lee, Sangin, et al.
Published: (2026)
by: Lee, Sangin, et al.
Published: (2026)
Emergent Temporal Correspondences from Video Diffusion Transformers
by: Nam, Jisu, et al.
Published: (2025)
by: Nam, Jisu, et al.
Published: (2025)
Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction
by: Kim, Jin Hyeon, et al.
Published: (2026)
by: Kim, Jin Hyeon, et al.
Published: (2026)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2026)
by: Jin, Woojeong, et al.
Published: (2026)
Motion Cues from Image-based Point Tracking for LiDAR Scene Flow Estimation
by: Jang, Youngdong, et al.
Published: (2026)
by: Jang, Youngdong, et al.
Published: (2026)
Spatial-Temporal State Propagation Autoregressive Model for 4D Object Generation
by: Yang, Liying, et al.
Published: (2026)
by: Yang, Liying, et al.
Published: (2026)
Zero4D: Training-Free 4D Video Generation From Single Video Using Off-the-Shelf Video Diffusion
by: Park, Jangho, et al.
Published: (2025)
by: Park, Jangho, et al.
Published: (2025)
Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection
by: Oh, Ahyoung, et al.
Published: (2026)
by: Oh, Ahyoung, et al.
Published: (2026)
Referring Video Object Segmentation via Language-aligned Track Selection
by: Kim, Seongchan, et al.
Published: (2024)
by: Kim, Seongchan, et al.
Published: (2024)
CorGi: Contribution-Guided Block-Wise Interval Caching for Training-Free Acceleration of Diffusion Transformers
by: Son, Yonglak, et al.
Published: (2025)
by: Son, Yonglak, et al.
Published: (2025)
Fine-Tuning Visual Autoregressive Models for Subject-Driven Generation
by: Chung, Jiwoo, et al.
Published: (2025)
by: Chung, Jiwoo, et al.
Published: (2025)
Similar Items
-
Projected Representation Conditioning for High-fidelity Novel View Synthesis
by: Kwak, Min-Seop, et al.
Published: (2026) -
Dense-SfM: Structure from Motion with Dense Consistent Matching
by: Lee, JongMin, et al.
Published: (2025) -
Repurposing Geometric Foundation Models for Multi-view Diffusion
by: Jang, Wooseok, et al.
Published: (2026) -
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
by: Kwon, Minkyung, et al.
Published: (2025) -
Grounding World Simulation Models in a Real-World Metropolis
by: Seo, Junyoung, et al.
Published: (2026)