Orthogonal Spatial-temporal Distributional Transfer for 4D Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Wei, Wu, Shengqiong, Li, Bobo, Zhao, Haoyu, Fei, Hao, Lee, Mong-Li, Hsu, Wynne |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
by: Wu, Lanhu, et al.
Published: (2025)
by: Wu, Lanhu, et al.
Published: (2025)
Multi-Part Object Representations via Graph Structures and Co-Part Discovery
by: Foo, Alex, et al.
Published: (2025)
by: Foo, Alex, et al.
Published: (2025)
SCP: Spatial Causal Prediction in Video
by: Zhao, Yanguang, et al.
Published: (2026)
by: Zhao, Yanguang, et al.
Published: (2026)
TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)
by: Yan, Zehong, et al.
Published: (2025)
Cross-Domain Feature Augmentation for Domain Generalization
by: Liu, Yingnan, et al.
Published: (2024)
by: Liu, Yingnan, et al.
Published: (2024)
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
by: Luo, Meng, et al.
Published: (2025)
by: Luo, Meng, et al.
Published: (2025)
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
by: Luo, Meng, et al.
Published: (2026)
by: Luo, Meng, et al.
Published: (2026)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
by: Wu, Shengqiong, et al.
Published: (2026)
by: Wu, Shengqiong, et al.
Published: (2026)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
Universal Scene Graph Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)
by: Yan, Zehong, et al.
Published: (2025)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024)
by: Qi, Peng, et al.
Published: (2024)
Global Commander and Local Operative: A Dual-Agent Framework for Scene Navigation
by: Jin, Kaiming, et al.
Published: (2026)
by: Jin, Kaiming, et al.
Published: (2026)
Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
by: Liang, Zhengyang, et al.
Published: (2025)
by: Liang, Zhengyang, et al.
Published: (2025)
4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency
by: Yin, Yuyang, et al.
Published: (2023)
by: Yin, Yuyang, et al.
Published: (2023)
MuSLR: Multimodal Symbolic Logical Reasoning
by: Xu, Jundong, et al.
Published: (2025)
by: Xu, Jundong, et al.
Published: (2025)
Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
by: Liang, Hanwen, et al.
Published: (2024)
by: Liang, Hanwen, et al.
Published: (2024)
DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion
by: Liu, Mengmeng, et al.
Published: (2025)
by: Liu, Mengmeng, et al.
Published: (2025)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
by: Fei, Hao, et al.
Published: (2023)
by: Fei, Hao, et al.
Published: (2023)
Towards Robust 3D Pose Transfer with Adversarial Learning
by: Chen, Haoyu, et al.
Published: (2024)
by: Chen, Haoyu, et al.
Published: (2024)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation
by: Liu, Mengmeng, et al.
Published: (2025)
by: Liu, Mengmeng, et al.
Published: (2025)
FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature Banks
by: Li, Jinwei, et al.
Published: (2025)
by: Li, Jinwei, et al.
Published: (2025)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2024)
by: Wu, Shengqiong, et al.
Published: (2024)
Audio-Visual Intelligence in Large Foundation Models
by: Qin, You, et al.
Published: (2026)
by: Qin, You, et al.
Published: (2026)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
by: Wang, Yaoting, et al.
Published: (2025)
by: Wang, Yaoting, et al.
Published: (2025)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency
by: Liu, Tianqi, et al.
Published: (2025)
by: Liu, Tianqi, et al.
Published: (2025)
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
by: Wu, Shengqiong, et al.
Published: (2024)
by: Wu, Shengqiong, et al.
Published: (2024)
OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
by: Zhang, Pengze, et al.
Published: (2026)
by: Zhang, Pengze, et al.
Published: (2026)
HSSDCT: Factorized Spatial-Spectral Correlation for Hyperspectral Image Fusion
by: Lee, Chia-Ming, et al.
Published: (2026)
by: Lee, Chia-Ming, et al.
Published: (2026)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
HFGS: 4D Gaussian Splatting with Emphasis on Spatial and Temporal High-Frequency Components for Endoscopic Scene Reconstruction
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Learning Hierarchical Orthogonal Prototypes for Generalized Few-Shot 3D Point Cloud Segmentation
by: Zhao, Yifei, et al.
Published: (2026)
by: Zhao, Yifei, et al.
Published: (2026)
Modeling Cross-vision Synergy for Unified Large Vision Model
by: Wu, Shengqiong, et al.
Published: (2026)
by: Wu, Shengqiong, et al.
Published: (2026)
Similar Items
-
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024) -
LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
by: Wu, Lanhu, et al.
Published: (2025) -
Multi-Part Object Representations via Graph Structures and Co-Part Discovery
by: Foo, Alex, et al.
Published: (2025) -
SCP: Spatial Causal Prediction in Video
by: Zhao, Yanguang, et al.
Published: (2026) -
TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)