CSTA: CNN-based Spatiotemporal Attention for Video Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Son, Jaewon, Park, Jaehun, Kim, Kwangsu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
by: Chen, Tieyuan, et al.
Published: (2025)
by: Chen, Tieyuan, et al.
Published: (2025)
SRVP: Strong Recollection Video Prediction Model Using Attention-Based Spatiotemporal Correlation Fusion
by: Kim, Yuseon, et al.
Published: (2025)
by: Kim, Yuseon, et al.
Published: (2025)
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
COSMOS: Coherent Supergaussian Modeling with Spatial Priors for Sparse-View 3D Splatting
by: Jeong, Chaeyoung, et al.
Published: (2025)
by: Jeong, Chaeyoung, et al.
Published: (2025)
Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
by: Tariq, Omer, et al.
Published: (2026)
by: Tariq, Omer, et al.
Published: (2026)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
by: Wang, Wenjing, et al.
Published: (2023)
by: Wang, Wenjing, et al.
Published: (2023)
Scaling Up Video Summarization Pretraining with Large Language Models
by: Argaw, Dawit Mureja, et al.
Published: (2024)
by: Argaw, Dawit Mureja, et al.
Published: (2024)
FullTransNet: Full Transformer with Local-Global Attention for Video Summarization
by: Lan, Libin, et al.
Published: (2025)
by: Lan, Libin, et al.
Published: (2025)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
Repurposing Video Diffusion Transformers for Robust Point Tracking
by: Son, Soowon, et al.
Published: (2025)
by: Son, Soowon, et al.
Published: (2025)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
by: Kim, Ji-Hoon, et al.
Published: (2025)
by: Kim, Ji-Hoon, et al.
Published: (2025)
Large Model based Sequential Keyframe Extraction for Video Summarization
by: Tan, Kailong, et al.
Published: (2024)
by: Tan, Kailong, et al.
Published: (2024)
CNN-based Multi-In-Multi-Out Model for Efficient Spatiotemporal Prediction
by: Jin, Hyeonseok
Published: (2026)
by: Jin, Hyeonseok
Published: (2026)
ROI-Aware Multiscale Cross-Attention Vision Transformer for Pest Image Identification
by: Kim, Ga-Eun, et al.
Published: (2023)
by: Kim, Ga-Eun, et al.
Published: (2023)
Video Summarization with Large Language Models
by: Lee, Min Jung, et al.
Published: (2025)
by: Lee, Min Jung, et al.
Published: (2025)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
by: Choi, Joonmyung, et al.
Published: (2024)
by: Choi, Joonmyung, et al.
Published: (2024)
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
by: Yune, Sungwoong, et al.
Published: (2026)
by: Yune, Sungwoong, et al.
Published: (2026)
LightSplat: Fast and Memory-Efficient Open-Vocabulary 3D Scene Understanding in Five Seconds
by: Bang, Jaehun, et al.
Published: (2026)
by: Bang, Jaehun, et al.
Published: (2026)
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
by: Truong, Quang-Trung, et al.
Published: (2024)
by: Truong, Quang-Trung, et al.
Published: (2024)
Video Summarization Techniques: A Comprehensive Review
by: Alaa, Toqa, et al.
Published: (2024)
by: Alaa, Toqa, et al.
Published: (2024)
Towards Efficient Vision State Space Models via Token Merging
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
by: Mu, Kuan-Chen, et al.
Published: (2024)
by: Mu, Kuan-Chen, et al.
Published: (2024)
Compensating Spatiotemporally Inconsistent Observations for Online Dynamic 3D Gaussian Splatting
by: Yun, Youngsik, et al.
Published: (2025)
by: Yun, Youngsik, et al.
Published: (2025)
SoGAR: Self-supervised Spatiotemporal Attention-based Social Group Activity Recognition
by: Chappa, Naga VS Raviteja, et al.
Published: (2023)
by: Chappa, Naga VS Raviteja, et al.
Published: (2023)
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Locally Grouped and Scale-Guided Attention for Dense Pest Counting
by: Son, Chang-Hwan
Published: (2024)
by: Son, Chang-Hwan
Published: (2024)
Spatiotemporal Tile-based Attention-guided LSTMs for Traffic Video Prediction
by: Nguyen, Tu
Published: (2019)
by: Nguyen, Tu
Published: (2019)
Video Summarization using Denoising Diffusion Probabilistic Model
by: Shang, Zirui, et al.
Published: (2024)
by: Shang, Zirui, et al.
Published: (2024)
Language-Guided Graph Representation Learning for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
Spiking Variational Graph Representation Inference for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
Enhancing Deepfake Detection using SE Block Attention with CNN
by: Dasgupta, Subhram, et al.
Published: (2025)
by: Dasgupta, Subhram, et al.
Published: (2025)
UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
by: Xia, Tian, et al.
Published: (2024)
by: Xia, Tian, et al.
Published: (2024)
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
by: Park, Jicheol, et al.
Published: (2024)
by: Park, Jicheol, et al.
Published: (2024)
Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
by: Lee, Ji Soo, et al.
Published: (2025)
by: Lee, Ji Soo, et al.
Published: (2025)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
by: Lin, Jingyang, et al.
Published: (2023)
by: Lin, Jingyang, et al.
Published: (2023)
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
Contrast-Phys: Unsupervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast
by: Sun, Zhaodong, et al.
Published: (2022)
by: Sun, Zhaodong, et al.
Published: (2022)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
Guided Slot Attention for Unsupervised Video Object Segmentation
by: Lee, Minhyeok, et al.
Published: (2023)
by: Lee, Minhyeok, et al.
Published: (2023)
Similar Items
-
CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
by: Chen, Tieyuan, et al.
Published: (2025) -
SRVP: Strong Recollection Video Prediction Model Using Attention-Based Spatiotemporal Correlation Fusion
by: Kim, Yuseon, et al.
Published: (2025) -
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025) -
COSMOS: Coherent Supergaussian Modeling with Spatial Priors for Sparse-View 3D Splatting
by: Jeong, Chaeyoung, et al.
Published: (2025) -
Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
by: Tariq, Omer, et al.
Published: (2026)