EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Hwang, Sunil, Yoon, Jaehong, Lee, Youngwan, Hwang, Sung Ju |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
Continual Learning: Forget-free Winning Subnetworks for Video Representations
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
Self-Refining Video Sampling
by: Jang, Sangwon, et al.
Published: (2026)
by: Jang, Sangwon, et al.
Published: (2026)
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024)
by: Lee, Youngwan, et al.
Published: (2024)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
by: Yeo, Woongyeong, et al.
Published: (2025)
by: Yeo, Woongyeong, et al.
Published: (2025)
Progressive Fourier Neural Representation for Sequential Video Compilation
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
by: Ki, Taekyung, et al.
Published: (2026)
by: Ki, Taekyung, et al.
Published: (2026)
Concept-skill Transferability-based Data Selection for Large Vision-Language Models
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
by: Kim, Kangsan, et al.
Published: (2024)
by: Kim, Kangsan, et al.
Published: (2024)
KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis
by: Lee, Youngwan, et al.
Published: (2023)
by: Lee, Youngwan, et al.
Published: (2023)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
by: Sung, Yi-Lin, et al.
Published: (2023)
by: Sung, Yi-Lin, et al.
Published: (2023)
Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
by: Kim, Jaihoon, et al.
Published: (2025)
by: Kim, Jaihoon, et al.
Published: (2025)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Semantic Prompting with Image-Token for Continual Learning
by: Han, Jisu, et al.
Published: (2024)
by: Han, Jisu, et al.
Published: (2024)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
Occupancy-Based Dual Contouring
by: Hwang, Jisung, et al.
Published: (2024)
by: Hwang, Jisung, et al.
Published: (2024)
SMCL: Saliency Masked Contrastive Learning for Long-tailed Recognition
by: Park, Sanglee, et al.
Published: (2024)
by: Park, Sanglee, et al.
Published: (2024)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
MCAQ-YOLO: Morphological Complexity-Aware Quantization for Efficient Object Detection with Curriculum Learning
by: Seo, Yoonjae, et al.
Published: (2025)
by: Seo, Yoonjae, et al.
Published: (2025)
SupMAE: Supervised Masked Autoencoders Are Efficient Vision Learners
by: Liang, Feng, et al.
Published: (2022)
by: Liang, Feng, et al.
Published: (2022)
A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
by: Lee, Heejun, et al.
Published: (2024)
by: Lee, Heejun, et al.
Published: (2024)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2026)
by: Lee, Youngwan, et al.
Published: (2026)
MaskVD: Region Masking for Efficient Video Object Detection
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
Robust Representation Learning in Masked Autoencoders
by: Shrivastava, Anika, et al.
Published: (2026)
by: Shrivastava, Anika, et al.
Published: (2026)
Multimodal Representation Learning by Alternating Unimodal Adaptation
by: Zhang, Xiaohui, et al.
Published: (2023)
by: Zhang, Xiaohui, et al.
Published: (2023)
Mask and Restore: Blind Backdoor Defense at Test Time with Masked Autoencoder
by: Sun, Tao, et al.
Published: (2023)
by: Sun, Tao, et al.
Published: (2023)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
by: Kim, Kangsan, et al.
Published: (2026)
by: Kim, Kangsan, et al.
Published: (2026)
Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos
by: Zhou, Jiaheng, et al.
Published: (2025)
by: Zhou, Jiaheng, et al.
Published: (2025)
SARMAE: Masked Autoencoder for SAR Representation Learning
by: Liu, Danxu, et al.
Published: (2025)
by: Liu, Danxu, et al.
Published: (2025)
Structure is Supervision: Multiview Masked Autoencoders for Radiology
by: Laguna, Sonia, et al.
Published: (2025)
by: Laguna, Sonia, et al.
Published: (2025)
SF(DA)$^2$: Source-free Domain Adaptation Through the Lens of Data Augmentation
by: Hwang, Uiwon, et al.
Published: (2024)
by: Hwang, Uiwon, et al.
Published: (2024)
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
by: Min, Dongchan, et al.
Published: (2022)
by: Min, Dongchan, et al.
Published: (2022)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
by: Yoon, Jaehong, et al.
Published: (2024)
by: Yoon, Jaehong, et al.
Published: (2024)
Osteoporosis Prediction from Hand X-ray Images Using Segmentation-for-Classification and Self-Supervised Learning
by: Hwang, Ung, et al.
Published: (2024)
by: Hwang, Ung, et al.
Published: (2024)
Anomaly Score: Evaluating Generative Models and Individual Generated Images based on Complexity and Vulnerability
by: Hwang, Jaehui, et al.
Published: (2023)
by: Hwang, Jaehui, et al.
Published: (2023)
HyperKD: Distilling Cross-Spectral Knowledge in Masked Autoencoders via Inverse Domain Shift with Spatial-Aware Masking and Specialized Loss
by: Matin, Abdul, et al.
Published: (2025)
by: Matin, Abdul, et al.
Published: (2025)
Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors
by: Lee, Jonghyun, et al.
Published: (2024)
by: Lee, Jonghyun, et al.
Published: (2024)
Downstream Task Guided Masking Learning in Masked Autoencoders Using Multi-Level Optimization
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
Attention-Guided Masked Autoencoders For Learning Image Representations
by: Sick, Leon, et al.
Published: (2024)
by: Sick, Leon, et al.
Published: (2024)
Similar Items
-
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
by: Lee, Daeun, et al.
Published: (2024) -
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023) -
Continual Learning: Forget-free Winning Subnetworks for Video Representations
by: Kang, Haeyong, et al.
Published: (2023) -
Self-Refining Video Sampling
by: Jang, Sangwon, et al.
Published: (2026) -
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024)