SparseTem: Boosting the Efficiency of CNN-Based Video Encoders by Exploiting Temporal Continuity
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Kunyun, Yang, Shuo, Zhao, Jieru, Ding, Wenchao, Chen, Quan, Leng, Jingwen, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning High-Frequency Continuous Action Chunks in Latent Space
by: Wang, Kunyun, et al.
Published: (2026)
by: Wang, Kunyun, et al.
Published: (2026)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
by: Ning, Zhenyu, et al.
Published: (2025)
by: Ning, Zhenyu, et al.
Published: (2025)
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
by: Ning, Zhenyu, et al.
Published: (2024)
by: Ning, Zhenyu, et al.
Published: (2024)
Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
by: Wang, Kunyun, et al.
Published: (2025)
by: Wang, Kunyun, et al.
Published: (2025)
Accelerating Sparse DNNs Based on Tiled GEMM
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
by: Miao, Wenxuan, et al.
Published: (2025)
by: Miao, Wenxuan, et al.
Published: (2025)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
by: Gong, Meiqi, et al.
Published: (2025)
by: Gong, Meiqi, et al.
Published: (2025)
STREAMINGGS: Voxel-Based Streaming 3D Gaussian Splatting with Memory Optimization and Architectural Support
by: Zhang, Chenqi, et al.
Published: (2025)
by: Zhang, Chenqi, et al.
Published: (2025)
SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
Design the Quantum Instruction Set with the Cartan Coordinate Analysis Framework
by: Wu, Anbang, et al.
Published: (2024)
by: Wu, Anbang, et al.
Published: (2024)
Exploiting the Semantic Knowledge of Pre-trained Text-Encoders for Continual Learning
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers
by: Liu, Haosong, et al.
Published: (2025)
by: Liu, Haosong, et al.
Published: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
Efficient One-stage Video Object Detection by Exploiting Temporal Consistency
by: Sun, Guanxiong, et al.
Published: (2024)
by: Sun, Guanxiong, et al.
Published: (2024)
Exploiting Spatial-Temporal Context for Interacting Hand Reconstruction on Monocular RGB Video
by: Zhao, Weichao, et al.
Published: (2023)
by: Zhao, Weichao, et al.
Published: (2023)
TSdetector: Temporal-Spatial Self-correction Collaborative Learning for Colonoscopy Video Detection
by: Wang, Kaini, et al.
Published: (2024)
by: Wang, Kaini, et al.
Published: (2024)
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
by: Li, Qirui, et al.
Published: (2025)
by: Li, Qirui, et al.
Published: (2025)
Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
by: Zhang, Junyi, et al.
Published: (2025)
by: Zhang, Junyi, et al.
Published: (2025)
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
by: Rasekh, Ali, et al.
Published: (2025)
by: Rasekh, Ali, et al.
Published: (2025)
Hierarchical Source-to-Post-Route QoR Prediction in High-Level Synthesis with GNNs
by: Gao, Mingzhe, et al.
Published: (2024)
by: Gao, Mingzhe, et al.
Published: (2024)
AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMs
by: Gao, Mingzhe, et al.
Published: (2024)
by: Gao, Mingzhe, et al.
Published: (2024)
ParaTransCNN: Parallelized TransCNN Encoder for Medical Image Segmentation
by: Sun, Hongkun, et al.
Published: (2024)
by: Sun, Hongkun, et al.
Published: (2024)
StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic Termination
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
by: Zhang, Weichuang, et al.
Published: (2024)
by: Zhang, Weichuang, et al.
Published: (2024)
SparseSplat: Towards Applicable Feed-Forward 3D Gaussian Splatting with Pixel-Unaligned Prediction
by: Zhang, Zicheng, et al.
Published: (2026)
by: Zhang, Zicheng, et al.
Published: (2026)
Enhancing Indoor Occupancy Prediction via Sparse Query-Based Multi-Level Consistent Knowledge Distillation
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
HybridGS: High-Efficiency Gaussian Splatting Data Compression using Dual-Channel Sparse Representation and Point Cloud Encoder
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
by: Hu, Xiantao, et al.
Published: (2024)
by: Hu, Xiantao, et al.
Published: (2024)
Progressively Texture-Aware Diffusion for Contrast-Enhanced Sparse-View CT
by: Wang, Tianqi, et al.
Published: (2026)
by: Wang, Tianqi, et al.
Published: (2026)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
by: Park, NaHyeon, et al.
Published: (2024)
by: Park, NaHyeon, et al.
Published: (2024)
R-Sparse R-CNN: SAR Ship Detection Based on Background-Aware Sparse Learnable Proposals
by: Kamirul, Kamirul, et al.
Published: (2025)
by: Kamirul, Kamirul, et al.
Published: (2025)
Sparse R-CNN OBB: Ship Target Detection in SAR Images Based on Oriented Sparse Proposals
by: Kamirul, Kamirul, et al.
Published: (2024)
by: Kamirul, Kamirul, et al.
Published: (2024)
Compressed Image Captioning using CNN-based Encoder-Decoder Framework
by: Ridoy, Md Alif Rahman, et al.
Published: (2024)
by: Ridoy, Md Alif Rahman, et al.
Published: (2024)
Cross-Temporal 3D Gaussian Splatting for Sparse-View Guided Scene Update
by: An, Zeyuan, et al.
Published: (2025)
by: An, Zeyuan, et al.
Published: (2025)
Lumina: Real-Time Mobile Neural Rendering by Exploiting Computational Redundancy
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment
by: wang, Shuo, et al.
Published: (2025)
by: wang, Shuo, et al.
Published: (2025)
T-DEED: Temporal-Discriminability Enhancer Encoder-Decoder for Precise Event Spotting in Sports Videos
by: Xarles, Artur, et al.
Published: (2024)
by: Xarles, Artur, et al.
Published: (2024)
SVASTIN: Sparse Video Adversarial Attack via Spatio-Temporal Invertible Neural Networks
by: Pan, Yi, et al.
Published: (2024)
by: Pan, Yi, et al.
Published: (2024)
Similar Items
-
Learning High-Frequency Continuous Action Chunks in Latent Space
by: Wang, Kunyun, et al.
Published: (2026) -
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
by: Ning, Zhenyu, et al.
Published: (2025) -
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
by: Ning, Zhenyu, et al.
Published: (2024) -
Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
by: Wang, Kunyun, et al.
Published: (2025) -
Accelerating Sparse DNNs Based on Tiled GEMM
by: Guo, Cong, et al.
Published: (2024)