USTEP: Spatio-Temporal Predictive Learning under A Unified View
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Cheng, Wang, Jue, Gao, Zhangyang, Li, Siyuan, Li, Stan Z. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning
by: Tan, Cheng, et al.
Published: (2024)
by: Tan, Cheng, et al.
Published: (2024)
MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning
by: Li, Zhe, et al.
Published: (2024)
by: Li, Zhe, et al.
Published: (2024)
OpenMixup: Open Mixup Toolbox and Benchmark for Visual Representation Learning
by: Li, Siyuan, et al.
Published: (2022)
by: Li, Siyuan, et al.
Published: (2022)
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
by: Tang, Chen, et al.
Published: (2025)
by: Tang, Chen, et al.
Published: (2025)
UniTS: Unified Spatio-Temporal Generative Model for Remote Sensing
by: Zhang, Yuxiang, et al.
Published: (2025)
by: Zhang, Yuxiang, et al.
Published: (2025)
LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving
by: Sun, Qihao, et al.
Published: (2026)
by: Sun, Qihao, et al.
Published: (2026)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
by: Li, Peiyan, et al.
Published: (2026)
by: Li, Peiyan, et al.
Published: (2026)
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
by: Chen, Dubing, et al.
Published: (2025)
by: Chen, Dubing, et al.
Published: (2025)
Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
by: Zheng, Yaozong, et al.
Published: (2025)
by: Zheng, Yaozong, et al.
Published: (2025)
From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper
by: Li, Ling, et al.
Published: (2026)
by: Li, Ling, et al.
Published: (2026)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
SSPA: Split-and-Synthesize Prompting with Gated Alignments for Multi-Label Image Recognition
by: Tan, Hao, et al.
Published: (2024)
by: Tan, Hao, et al.
Published: (2024)
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
by: Li, Siyuan, et al.
Published: (2023)
by: Li, Siyuan, et al.
Published: (2023)
MogaNet: Multi-order Gated Aggregation Network
by: Li, Siyuan, et al.
Published: (2022)
by: Li, Siyuan, et al.
Published: (2022)
FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting
by: Zhu, Zehao, et al.
Published: (2023)
by: Zhu, Zehao, et al.
Published: (2023)
Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
by: Shen, Hao, et al.
Published: (2024)
by: Shen, Hao, et al.
Published: (2024)
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
by: Jin, Yudong, et al.
Published: (2025)
by: Jin, Yudong, et al.
Published: (2025)
SpatioTemporal Difference Network for Video Depth Super-Resolution
by: Wang, Zhengxue, et al.
Published: (2025)
by: Wang, Zhengxue, et al.
Published: (2025)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
CurConMix+: A Unified Spatio-Temporal Framework for Hierarchical Surgical Workflow Understanding
by: Jeon, Yongjun, et al.
Published: (2026)
by: Jeon, Yongjun, et al.
Published: (2026)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
by: Cheng, Zixu, et al.
Published: (2025)
by: Cheng, Zixu, et al.
Published: (2025)
GenURL: A General Framework for Unsupervised Representation Learning
by: Li, Siyuan, et al.
Published: (2021)
by: Li, Siyuan, et al.
Published: (2021)
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
by: Tian, Juanxi, et al.
Published: (2025)
by: Tian, Juanxi, et al.
Published: (2025)
Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
by: Zhang, Zhaoxing, et al.
Published: (2024)
by: Zhang, Zhaoxing, et al.
Published: (2024)
Unify the Views: View-Consistent Prototype Learning for Few-Shot Segmentation
by: Liu, Hongli, et al.
Published: (2026)
by: Liu, Hongli, et al.
Published: (2026)
UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
by: Li, Peiming, et al.
Published: (2025)
by: Li, Peiming, et al.
Published: (2025)
STLight: a Fully Convolutional Approach for Efficient Predictive Learning by Spatio-Temporal joint Processing
by: Alfarano, Andrea, et al.
Published: (2024)
by: Alfarano, Andrea, et al.
Published: (2024)
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
by: Shi, Chen, et al.
Published: (2025)
by: Shi, Chen, et al.
Published: (2025)
Dual-view Spatio-Temporal Feature Fusion with CNN-Transformer Hybrid Network for Chinese Isolated Sign Language Recognition
by: Jing, Siyuan, et al.
Published: (2025)
by: Jing, Siyuan, et al.
Published: (2025)
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
by: Qiu, Tianheng, et al.
Published: (2025)
by: Qiu, Tianheng, et al.
Published: (2025)
Transformer RGBT Tracking with Spatio-Temporal Multimodal Tokens
by: Sun, Dengdi, et al.
Published: (2024)
by: Sun, Dengdi, et al.
Published: (2024)
DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration
by: Chen, Zheng, et al.
Published: (2026)
by: Chen, Zheng, et al.
Published: (2026)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
by: Garg, Aaryan, et al.
Published: (2025)
by: Garg, Aaryan, et al.
Published: (2025)
STDR: Spatio-Temporal Decoupling for Real-Time Dynamic Scene Rendering
by: Li, Zehao, et al.
Published: (2025)
by: Li, Zehao, et al.
Published: (2025)
TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding
by: Li, Jinxuan, et al.
Published: (2025)
by: Li, Jinxuan, et al.
Published: (2025)
Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
UniSTFormer: Unified Spatio-Temporal Lightweight Transformer for Efficient Skeleton-Based Action Recognition
by: Wu, Wenhan, et al.
Published: (2025)
by: Wu, Wenhan, et al.
Published: (2025)
Similar Items
-
Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning
by: Tan, Cheng, et al.
Published: (2024) -
MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning
by: Li, Zhe, et al.
Published: (2024) -
OpenMixup: Open Mixup Toolbox and Benchmark for Visual Representation Learning
by: Li, Siyuan, et al.
Published: (2022) -
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
by: Yuan, Yuqian, et al.
Published: (2025) -
UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
by: Tang, Chen, et al.
Published: (2025)