OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Tongjia, Yu, Hongshan, Yang, Zhengeng, Li, Zechuan, Sun, Wei, Chen, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Domain-invariant Prototypes for Semantic Segmentation
by: Yang, Zhengeng, et al.
Published: (2022)
by: Yang, Zhengeng, et al.
Published: (2022)
Soft Masked Transformer for Point Cloud Processing with Skip Attention-Based Upsampling
by: He, Yong, et al.
Published: (2024)
by: He, Yong, et al.
Published: (2024)
PointDiffuse: A Dual-Conditional Diffusion Model for Enhanced Point Cloud Semantic Segmentation
by: He, Yong, et al.
Published: (2025)
by: He, Yong, et al.
Published: (2025)
Deep learning based infrared small object segmentation: Challenges and future directions
by: Yang, Zhengeng, et al.
Published: (2025)
by: Yang, Zhengeng, et al.
Published: (2025)
Learning Sequence Descriptor based on Spatio-Temporal Attention for Visual Place Recognition
by: Zhao, Junqiao, et al.
Published: (2023)
by: Zhao, Junqiao, et al.
Published: (2023)
Deep Learning Based 3D Segmentation: A Survey
by: He, Yong, et al.
Published: (2021)
by: He, Yong, et al.
Published: (2021)
Efficient Point Cloud Processing with High-Dimensional Positional Encoding and Non-Local MLPs
by: Zou, Yanmei, et al.
Published: (2026)
by: Zou, Yanmei, et al.
Published: (2026)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition
by: Kiu, Yu, et al.
Published: (2025)
by: Kiu, Yu, et al.
Published: (2025)
RSEdit: Text-Guided Image Editing for Remote Sensing
by: Zhenyuan, Chen, et al.
Published: (2026)
by: Zhenyuan, Chen, et al.
Published: (2026)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
Are Image-to-Video Models Good Zero-Shot Image Editors?
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
by: Chen, Junyang, et al.
Published: (2025)
by: Chen, Junyang, et al.
Published: (2025)
GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
Knowledge Distillation with Refined Logits
by: Sun, Wujie, et al.
Published: (2024)
by: Sun, Wujie, et al.
Published: (2024)
UniSTFormer: Unified Spatio-Temporal Lightweight Transformer for Efficient Skeleton-Based Action Recognition
by: Wu, Wenhan, et al.
Published: (2025)
by: Wu, Wenhan, et al.
Published: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
Self-Supervised Place Recognition by Refining Temporal and Featural Pseudo Labels from Panoramic Data
by: Chen, Chao, et al.
Published: (2022)
by: Chen, Chao, et al.
Published: (2022)
InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba
by: Wu, Zizhao, et al.
Published: (2025)
by: Wu, Zizhao, et al.
Published: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
by: Guo, Yanan, et al.
Published: (2025)
by: Guo, Yanan, et al.
Published: (2025)
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
by: Zhang, Yan, et al.
Published: (2024)
by: Zhang, Yan, et al.
Published: (2024)
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
by: Aparcedo, Alejandro, et al.
Published: (2026)
by: Aparcedo, Alejandro, et al.
Published: (2026)
TrackNetV5: Residual-Driven Spatio-Temporal Refinement and Motion Direction Decoupling for Fast Object Tracking
by: Tang, Haonan, et al.
Published: (2025)
by: Tang, Haonan, et al.
Published: (2025)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
by: Chen, Boyu, et al.
Published: (2024)
by: Chen, Boyu, et al.
Published: (2024)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
by: Cheng, Zixu, et al.
Published: (2025)
by: Cheng, Zixu, et al.
Published: (2025)
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
by: Chen, Zhennan, et al.
Published: (2024)
by: Chen, Zhennan, et al.
Published: (2024)
DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration
by: Chen, Zheng, et al.
Published: (2026)
by: Chen, Zheng, et al.
Published: (2026)
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
by: Li, Qirui, et al.
Published: (2025)
by: Li, Qirui, et al.
Published: (2025)
STGV: Spatio-Temporal Hash Encoding for Gaussian-based Video Representation
by: Lin, Jierun, et al.
Published: (2026)
by: Lin, Jierun, et al.
Published: (2026)
SpatioTemporal Difference Network for Video Depth Super-Resolution
by: Wang, Zhengxue, et al.
Published: (2025)
by: Wang, Zhengxue, et al.
Published: (2025)
M2OST: Many-to-one Regression for Predicting Spatial Transcriptomics from Digital Pathology Images
by: Wang, Hongyi, et al.
Published: (2024)
by: Wang, Hongyi, et al.
Published: (2024)
Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
by: Peddi, Rohith, et al.
Published: (2026)
by: Peddi, Rohith, et al.
Published: (2026)
Generalized Correspondence Matching via Flexible Hierarchical Refinement and Patch Descriptor Distillation
by: Han, Yu, et al.
Published: (2024)
by: Han, Yu, et al.
Published: (2024)
Video-Language Alignment via Spatio-Temporal Graph Transformer
by: Zhang, Shi-Xue, et al.
Published: (2024)
by: Zhang, Shi-Xue, et al.
Published: (2024)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
STAF: 3D Human Mesh Recovery from Video with Spatio-Temporal Alignment Fusion
by: Yao, Wei, et al.
Published: (2024)
by: Yao, Wei, et al.
Published: (2024)
SVASTIN: Sparse Video Adversarial Attack via Spatio-Temporal Invertible Neural Networks
by: Pan, Yi, et al.
Published: (2024)
by: Pan, Yi, et al.
Published: (2024)
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
Similar Items
-
Domain-invariant Prototypes for Semantic Segmentation
by: Yang, Zhengeng, et al.
Published: (2022) -
Soft Masked Transformer for Point Cloud Processing with Skip Attention-Based Upsampling
by: He, Yong, et al.
Published: (2024) -
PointDiffuse: A Dual-Conditional Diffusion Model for Enhanced Point Cloud Semantic Segmentation
by: He, Yong, et al.
Published: (2025) -
Deep learning based infrared small object segmentation: Challenges and future directions
by: Yang, Zhengeng, et al.
Published: (2025) -
Learning Sequence Descriptor based on Spatio-Temporal Attention for Visual Place Recognition
by: Zhao, Junqiao, et al.
Published: (2023)