Reprogramming Vision Foundation Models for Spatio-Temporal Forecasting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Changlu, Liu, Yanbin, Niu, Chaoxi, Chen, Ling, Zhu, Tianqing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
von: Goodge, Adam, et al.
Veröffentlicht: (2025)
von: Goodge, Adam, et al.
Veröffentlicht: (2025)
Cross Space and Time: A Spatio-Temporal Unitized Model for Traffic Flow Forecasting
von: Ruan, Weilin, et al.
Veröffentlicht: (2024)
von: Ruan, Weilin, et al.
Veröffentlicht: (2024)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2024)
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2024)
Spatio-Temporal Side Tuning Pre-trained Foundation Models for Video-based Pedestrian Attribute Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
Context-based Interpretable Spatio-Temporal Graph Convolutional Network for Human Motion Forecasting
von: Medina, Edgar, et al.
Veröffentlicht: (2024)
von: Medina, Edgar, et al.
Veröffentlicht: (2024)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
von: Sha, Lin, et al.
Veröffentlicht: (2026)
von: Sha, Lin, et al.
Veröffentlicht: (2026)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
von: Huang, Wei-Jhe, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Jhe, et al.
Veröffentlicht: (2024)
Playing to Vision Foundation Model's Strengths in Stereo Matching
von: Liu, Chuang-Wei, et al.
Veröffentlicht: (2024)
von: Liu, Chuang-Wei, et al.
Veröffentlicht: (2024)
ViTime: Foundation Model for Time Series Forecasting Powered by Vision Intelligence
von: Yang, Luoxiao, et al.
Veröffentlicht: (2024)
von: Yang, Luoxiao, et al.
Veröffentlicht: (2024)
STeP-Diff: Spatio-Temporal Physics-Informed Diffusion Models for Mobile Fine-Grained Pollution Forecasting
von: Zhou, Nan, et al.
Veröffentlicht: (2025)
von: Zhou, Nan, et al.
Veröffentlicht: (2025)
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
von: Zhu, Yuanbing, et al.
Veröffentlicht: (2024)
von: Zhu, Yuanbing, et al.
Veröffentlicht: (2024)
UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
von: Li, Peiming, et al.
Veröffentlicht: (2025)
von: Li, Peiming, et al.
Veröffentlicht: (2025)
EventAug: Multifaceted Spatio-Temporal Data Augmentation Methods for Event-based Learning
von: Tian, Yukun, et al.
Veröffentlicht: (2024)
von: Tian, Yukun, et al.
Veröffentlicht: (2024)
VFMF: World Modeling by Forecasting Vision Foundation Model Features
von: Boduljak, Gabrijel, et al.
Veröffentlicht: (2025)
von: Boduljak, Gabrijel, et al.
Veröffentlicht: (2025)
STAA: Spatio-Temporal Attention Attribution for Real-Time Interpreting Transformer-based Video Models
von: Wang, Zerui, et al.
Veröffentlicht: (2024)
von: Wang, Zerui, et al.
Veröffentlicht: (2024)
Robust SAM: On the Adversarial Robustness of Vision Foundation Models
von: Long, Jiahuan, et al.
Veröffentlicht: (2025)
von: Long, Jiahuan, et al.
Veröffentlicht: (2025)
X-VORTEX: Spatio-Temporal Contrastive Learning for Wake Vortex Trajectory Forecasting
von: Qu, Zhan, et al.
Veröffentlicht: (2026)
von: Qu, Zhan, et al.
Veröffentlicht: (2026)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
Historical Test-time Prompt Tuning for Vision Foundation Models
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
Data-Efficient Surgical Phase Segmentation in Small-Incision Cataract Surgery: A Controlled Study of Vision Foundation Models
von: Spencer, Lincoln, et al.
Veröffentlicht: (2026)
von: Spencer, Lincoln, et al.
Veröffentlicht: (2026)
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
von: Chetan, Aditya, et al.
Veröffentlicht: (2026)
von: Chetan, Aditya, et al.
Veröffentlicht: (2026)
Physical Prompt Injection Attacks on Large Vision-Language Models
von: Ling, Chen, et al.
Veröffentlicht: (2026)
von: Ling, Chen, et al.
Veröffentlicht: (2026)
Agentic Spatio-Temporal Grounding via Collaborative Reasoning
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment
von: Chen, Weizhi, et al.
Veröffentlicht: (2025)
von: Chen, Weizhi, et al.
Veröffentlicht: (2025)
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
BrainCast: A Spatio-Temporal Forecasting Model for Whole-Brain fMRI Time Series Prediction
von: Gao, Yunlong, et al.
Veröffentlicht: (2026)
von: Gao, Yunlong, et al.
Veröffentlicht: (2026)
IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
von: Chen, Zhihao, et al.
Veröffentlicht: (2023)
von: Chen, Zhihao, et al.
Veröffentlicht: (2023)
Mechanistic Learning with Guided Diffusion Models to Predict Spatio-Temporal Brain Tumor Growth
von: Laslo, Daria, et al.
Veröffentlicht: (2025)
von: Laslo, Daria, et al.
Veröffentlicht: (2025)
TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
von: Carrión, Héctor, et al.
Veröffentlicht: (2025)
von: Carrión, Héctor, et al.
Veröffentlicht: (2025)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2024)
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
von: Guo, Zhongbin, et al.
Veröffentlicht: (2025)
von: Guo, Zhongbin, et al.
Veröffentlicht: (2025)
Temporal Reversal Regularization for Spiking Neural Networks: Hybrid Spatio-Temporal Invariance for Generalization
von: Zuo, Lin, et al.
Veröffentlicht: (2024)
von: Zuo, Lin, et al.
Veröffentlicht: (2024)
The Butterfly Effect in Pathology: Exploring Security in Pathology Foundation Models
von: Liu, Jiashuai, et al.
Veröffentlicht: (2025)
von: Liu, Jiashuai, et al.
Veröffentlicht: (2025)
Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
von: Li, Qiankun, et al.
Veröffentlicht: (2025)
von: Li, Qiankun, et al.
Veröffentlicht: (2025)
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
von: Huang, Wenlong, et al.
Veröffentlicht: (2024)
von: Huang, Wenlong, et al.
Veröffentlicht: (2024)
FireSentry: A Multi-Modal Spatio-temporal Benchmark Dataset for Fine-Grained Wildfire Spread Forecasting
von: Zhou, Nan, et al.
Veröffentlicht: (2025)
von: Zhou, Nan, et al.
Veröffentlicht: (2025)
DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
von: Li, Weihong, et al.
Veröffentlicht: (2025)
von: Li, Weihong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
von: Goodge, Adam, et al.
Veröffentlicht: (2025) -
Cross Space and Time: A Spatio-Temporal Unitized Model for Traffic Flow Forecasting
von: Ruan, Weilin, et al.
Veröffentlicht: (2024) -
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025) -
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2024) -
Spatio-Temporal Side Tuning Pre-trained Foundation Models for Video-based Pedestrian Attribute Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2024)