Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Bing, Lu, Quanhao, Feng, Jiekang, Wang, Qilong, Hu, Qinghua, Zhu, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Occlusion-Aware 3D Hand-Object Pose Estimation with Masked AutoEncoders
by: Yang, Hui, et al.
Published: (2025)
by: Yang, Hui, et al.
Published: (2025)
RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning
by: Su, Kunming, et al.
Published: (2024)
by: Su, Kunming, et al.
Published: (2024)
Distribution Matching Variational AutoEncoder
by: Ye, Sen, et al.
Published: (2025)
by: Ye, Sen, et al.
Published: (2025)
MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
by: Labatie, Antoine, et al.
Published: (2025)
by: Labatie, Antoine, et al.
Published: (2025)
3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud Pretraining
by: Yan, Siming, et al.
Published: (2023)
by: Yan, Siming, et al.
Published: (2023)
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
Visible and Clear: Finding Tiny Objects in Difference Map
by: Cao, Bing, et al.
Published: (2024)
by: Cao, Bing, et al.
Published: (2024)
Appearance Blur-driven AutoEncoder and Motion-guided Memory Module for Video Anomaly Detection
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
Latent Enhancing AutoEncoder for Occluded Image Classification
by: Kotwal, Ketan, et al.
Published: (2024)
by: Kotwal, Ketan, et al.
Published: (2024)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
by: Wu, Jiahe, et al.
Published: (2026)
by: Wu, Jiahe, et al.
Published: (2026)
Improved AutoEncoder with LSTM module and KL divergence
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
SAEN-BGS: Energy-Efficient Spiking AutoEncoder Network for Background Subtraction
by: Zhang, Zhixuan, et al.
Published: (2025)
by: Zhang, Zhixuan, et al.
Published: (2025)
ColorMAE: Exploring data-independent masking strategies in Masked AutoEncoders
by: Hinojosa, Carlos, et al.
Published: (2024)
by: Hinojosa, Carlos, et al.
Published: (2024)
Enhancing SAR Object Detection with Self-Supervised Pre-training on Masked Auto-Encoders
by: Pu, Xinyang, et al.
Published: (2025)
by: Pu, Xinyang, et al.
Published: (2025)
Conditional Controllable Image Fusion
by: Cao, Bing, et al.
Published: (2024)
by: Cao, Bing, et al.
Published: (2024)
$S^3$: Synonymous Semantic Space for Improving Zero-Shot Generalization of Vision-Language Models
by: Yin, Xiaojie, et al.
Published: (2024)
by: Yin, Xiaojie, et al.
Published: (2024)
Robust Multimodal Survival Prediction with the Latent Differentiation Conditional Variational AutoEncoder
by: Zhou, Junjie, et al.
Published: (2025)
by: Zhou, Junjie, et al.
Published: (2025)
Self-Distilled Masked Auto-Encoders are Efficient Video Anomaly Detectors
by: Ristea, Nicolae-Catalin, et al.
Published: (2023)
by: Ristea, Nicolae-Catalin, et al.
Published: (2023)
Reversible Efficient Diffusion for Image Fusion
by: Xu, Xingxin, et al.
Published: (2026)
by: Xu, Xingxin, et al.
Published: (2026)
Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
by: Luo, Liwei, et al.
Published: (2025)
by: Luo, Liwei, et al.
Published: (2025)
Task-Customized Mixture of Adapters for General Image Fusion
by: Zhu, Pengfei, et al.
Published: (2024)
by: Zhu, Pengfei, et al.
Published: (2024)
Detecting AutoEncoder is Enough to Catch LDM Generated Images
by: Vesnin, Dmitry, et al.
Published: (2024)
by: Vesnin, Dmitry, et al.
Published: (2024)
Detecting AI-Generated Images via Contextual Anomaly Estimation in Masked AutoEncoders
by: Jang, Minsuk, et al.
Published: (2025)
by: Jang, Minsuk, et al.
Published: (2025)
NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields
by: Irshad, Muhammad Zubair, et al.
Published: (2024)
by: Irshad, Muhammad Zubair, et al.
Published: (2024)
CE-VAE: Capsule Enhanced Variational AutoEncoder for Underwater Image Enhancement
by: Pucci, Rita, et al.
Published: (2024)
by: Pucci, Rita, et al.
Published: (2024)
H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion
by: Sun, Yiming, et al.
Published: (2026)
by: Sun, Yiming, et al.
Published: (2026)
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
by: Kamenetsky, Ronen, et al.
Published: (2025)
by: Kamenetsky, Ronen, et al.
Published: (2025)
GeneA-SLAM2: Dynamic SLAM with AutoEncoder-Preprocessed Genetic Keypoints Resampling and Depth Variance-Guided Dynamic Region Removal
by: Qing, Shufan, et al.
Published: (2025)
by: Qing, Shufan, et al.
Published: (2025)
Asymmetric Reinforcing against Multi-modal Representation Bias
by: Gao, Xiyuan, et al.
Published: (2025)
by: Gao, Xiyuan, et al.
Published: (2025)
Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion
by: Xu, Xingxin, et al.
Published: (2025)
by: Xu, Xingxin, et al.
Published: (2025)
Bi-directional Self-Registration for Misaligned Infrared-Visible Image Fusion
by: Li, Timing, et al.
Published: (2025)
by: Li, Timing, et al.
Published: (2025)
Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion
by: Sun, Yiming, et al.
Published: (2024)
by: Sun, Yiming, et al.
Published: (2024)
Physics Informed Capsule Enhanced Variational AutoEncoder for Underwater Image Enhancement
by: Martinel, Niki, et al.
Published: (2025)
by: Martinel, Niki, et al.
Published: (2025)
BackMix: Regularizing Open Set Recognition by Removing Underlying Fore-Background Priors
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder
by: Bora, Maheswar, et al.
Published: (2024)
by: Bora, Maheswar, et al.
Published: (2024)
Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
by: Yin, Xiaojie, et al.
Published: (2025)
by: Yin, Xiaojie, et al.
Published: (2025)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
by: Zhu, Wencheng, et al.
Published: (2025)
by: Zhu, Wencheng, et al.
Published: (2025)
Similar Items
-
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025) -
Occlusion-Aware 3D Hand-Object Pose Estimation with Masked AutoEncoders
by: Yang, Hui, et al.
Published: (2025) -
RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning
by: Su, Kunming, et al.
Published: (2024) -
Distribution Matching Variational AutoEncoder
by: Ye, Sen, et al.
Published: (2025) -
MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
by: Labatie, Antoine, et al.
Published: (2025)