Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xiangkai, Zhang, Han, Li, Wenzhong, Lu, Sanglu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic-Supervised Spatial-Temporal Fusion for LiDAR-based 3D Object Detection
by: Wang, Chaoqun, et al.
Published: (2025)
by: Wang, Chaoqun, et al.
Published: (2025)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
by: Han, Chaolei, et al.
Published: (2025)
by: Han, Chaolei, et al.
Published: (2025)
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
by: Ding, Zijun, et al.
Published: (2025)
by: Ding, Zijun, et al.
Published: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
Bootstrap Fine-Grained Vision-Language Alignment for Unified Zero-Shot Anomaly Localization
by: Deng, Hanqiu, et al.
Published: (2023)
by: Deng, Hanqiu, et al.
Published: (2023)
Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
OZ-TAL: Online Zero-Shot Temporal Action Localization
by: Han, Chaolei, et al.
Published: (2026)
by: Han, Chaolei, et al.
Published: (2026)
Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
by: Ma, Xiangkai, et al.
Published: (2025)
by: Ma, Xiangkai, et al.
Published: (2025)
EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning
by: Qu, Xiangyan, et al.
Published: (2024)
by: Qu, Xiangyan, et al.
Published: (2024)
Crowded Video Individual Counting Informed by Social Grouping and Spatial-Temporal Displacement Priors
by: Lu, Hao, et al.
Published: (2026)
by: Lu, Hao, et al.
Published: (2026)
Drift-Resilient Temporal Priors for Visual Tracking
by: Huang, Yuqing, et al.
Published: (2026)
by: Huang, Yuqing, et al.
Published: (2026)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
by: Wen, Haokun, et al.
Published: (2026)
by: Wen, Haokun, et al.
Published: (2026)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
by: Yang, Min, et al.
Published: (2024)
by: Yang, Min, et al.
Published: (2024)
Zero-Shot Skeleton-based Action Recognition with Dual Visual-Text Alignment
by: Kuang, Jidong, et al.
Published: (2024)
by: Kuang, Jidong, et al.
Published: (2024)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
by: Zhang, Erhang, et al.
Published: (2025)
by: Zhang, Erhang, et al.
Published: (2025)
A Unified Spatial Alignment Framework for Highly Transferable Transformation-Based Attacks on Spatially Structured Tasks
by: Liang, Jiaming, et al.
Published: (2026)
by: Liang, Jiaming, et al.
Published: (2026)
Two-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video
by: Huang, Jiantang
Published: (2026)
by: Huang, Jiantang
Published: (2026)
A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing
by: Li, Maomao, et al.
Published: (2023)
by: Li, Maomao, et al.
Published: (2023)
Spatial-Temporal-Spectral Unified Modeling for Remote Sensing Dense Prediction
by: Zhao, Sijie, et al.
Published: (2025)
by: Zhao, Sijie, et al.
Published: (2025)
Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives
by: Li, Daiqiang, et al.
Published: (2026)
by: Li, Daiqiang, et al.
Published: (2026)
UniTS: Unified Spatio-Temporal Generative Model for Remote Sensing
by: Zhang, Yuxiang, et al.
Published: (2025)
by: Zhang, Yuxiang, et al.
Published: (2025)
Test-Time Zero-Shot Temporal Action Localization
by: Liberatori, Benedetta, et al.
Published: (2024)
by: Liberatori, Benedetta, et al.
Published: (2024)
Zero-Shot Image Harmonization with Generative Model Prior
by: Chen, Jianqi, et al.
Published: (2023)
by: Chen, Jianqi, et al.
Published: (2023)
A Wave is Worth 100 Words: Investigating Cross-Domain Transferability in Time Series
by: Ma, Xiangkai, et al.
Published: (2024)
by: Ma, Xiangkai, et al.
Published: (2024)
STeInFormer: Spatial-Temporal Interaction Transformer Architecture for Remote Sensing Change Detection
by: Ma, Xiaowen, et al.
Published: (2024)
by: Ma, Xiaowen, et al.
Published: (2024)
ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
by: Yang, Ziteng, et al.
Published: (2025)
by: Yang, Ziteng, et al.
Published: (2025)
Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
Energy-Aware Pattern Disentanglement: A Generalizable Pattern Assisted Architecture for Multi-task Time Series Analysis
by: Ma, Xiangkai, et al.
Published: (2025)
by: Ma, Xiangkai, et al.
Published: (2025)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Graph Network for Sign Language Tasks
by: Gan, Shiwei, et al.
Published: (2025)
by: Gan, Shiwei, et al.
Published: (2025)
Attribute Distribution Modeling and Semantic-Visual Alignment for Generative Zero-shot Learning
by: Pu, Haojie, et al.
Published: (2026)
by: Pu, Haojie, et al.
Published: (2026)
Zero-Shot Temporal Action Localization Through Textual Guidance
by: Liberatori, Benedetta, et al.
Published: (2026)
by: Liberatori, Benedetta, et al.
Published: (2026)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
by: Jiang, Huajie, et al.
Published: (2025)
by: Jiang, Huajie, et al.
Published: (2025)
Cross-Model Transfer of Task Vectors via Few-Shot Orthogonal Alignment
by: Kawamoto, Kazuhiko, et al.
Published: (2025)
by: Kawamoto, Kazuhiko, et al.
Published: (2025)
Temporal-Spatial Object Relations Modeling for Vision-and-Language Navigation
by: Huang, Bowen, et al.
Published: (2024)
by: Huang, Bowen, et al.
Published: (2024)
USTM: Unified Spatial and Temporal Modeling for Continuous Sign Language Recognition
by: Hasanaath, Ahmed Abul, et al.
Published: (2025)
by: Hasanaath, Ahmed Abul, et al.
Published: (2025)
Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
by: Xu, Wenjia, et al.
Published: (2024)
by: Xu, Wenjia, et al.
Published: (2024)
Similar Items
-
Semantic-Supervised Spatial-Temporal Fusion for LiDAR-based 3D Object Detection
by: Wang, Chaoqun, et al.
Published: (2025) -
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
by: Han, Chaolei, et al.
Published: (2025) -
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
by: Ding, Zijun, et al.
Published: (2025) -
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
by: Yang, Shuai, et al.
Published: (2024) -
Bootstrap Fine-Grained Vision-Language Alignment for Unified Zero-Shot Anomaly Localization
by: Deng, Hanqiu, et al.
Published: (2023)