DailyMAE: Towards Pretraining Masked Autoencoders in One Day
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Jiantao, Mo, Shentong, Atito, Sara, Feng, Zhenhua, Kittler, Josef, Awais, Muhammad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking Positive Pairs in Contrastive Learning
por: Wu, Jiantao, et al.
Publicado: (2024)
por: Wu, Jiantao, et al.
Publicado: (2024)
Pseudo Labelling for Enhanced Masked Autoencoders
por: Nandam, Srinivasa Rao, et al.
Publicado: (2024)
por: Nandam, Srinivasa Rao, et al.
Publicado: (2024)
Investigating Self-Supervised Methods for Label-Efficient Learning
por: Nandam, Srinivasa Rao, et al.
Publicado: (2024)
por: Nandam, Srinivasa Rao, et al.
Publicado: (2024)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
por: Nazarieh, Fatemeh, et al.
Publicado: (2024)
por: Nazarieh, Fatemeh, et al.
Publicado: (2024)
Probabilistically Aligned View-unaligned Clustering with Adaptive Template Selection
por: Dong, Wenhua, et al.
Publicado: (2024)
por: Dong, Wenhua, et al.
Publicado: (2024)
One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion
por: Cheng, Chunyang, et al.
Publicado: (2025)
por: Cheng, Chunyang, et al.
Publicado: (2025)
C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
por: Li, Rongchang, et al.
Publicado: (2024)
por: Li, Rongchang, et al.
Publicado: (2024)
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
por: Mo, Shentong
Publicado: (2024)
por: Mo, Shentong
Publicado: (2024)
ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
por: Atito, Sara, et al.
Publicado: (2022)
por: Atito, Sara, et al.
Publicado: (2022)
MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
por: Nazarieh, Fatemeh, et al.
Publicado: (2025)
por: Nazarieh, Fatemeh, et al.
Publicado: (2025)
Information theoretic underpinning of self-supervised learning by clustering
por: Kittler, Josef, et al.
Publicado: (2026)
por: Kittler, Josef, et al.
Publicado: (2026)
Channel-Aware Probing for Multi-Channel Imaging
por: Marikkar, Umar, et al.
Publicado: (2026)
por: Marikkar, Umar, et al.
Publicado: (2026)
Domain Adaptation Without the Compute Burden for Efficient Whole Slide Image Analysis
por: Marikkar, Umar, et al.
Publicado: (2026)
por: Marikkar, Umar, et al.
Publicado: (2026)
SupMAE: Supervised Masked Autoencoders Are Efficient Vision Learners
por: Liang, Feng, et al.
Publicado: (2022)
por: Liang, Feng, et al.
Publicado: (2022)
C3R: Channel Conditioned Cell Representations for unified evaluation in microscopy imaging
por: Marikkar, Umar, et al.
Publicado: (2025)
por: Marikkar, Umar, et al.
Publicado: (2025)
CL-MAE: Curriculum-Learned Masked Autoencoders
por: Madan, Neelu, et al.
Publicado: (2023)
por: Madan, Neelu, et al.
Publicado: (2023)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
por: Mo, Shentong
Publicado: (2024)
por: Mo, Shentong
Publicado: (2024)
DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images
por: Marikkar, Umar, et al.
Publicado: (2026)
por: Marikkar, Umar, et al.
Publicado: (2026)
PaCX-MAE: Physiology-Augmented Chest X-Ray Masked Autoencoder
por: Liu, Yancheng, et al.
Publicado: (2026)
por: Liu, Yancheng, et al.
Publicado: (2026)
i-MAE: Are Latent Representations in Masked Autoencoders Linearly Separable?
por: Zhang, Kevin, et al.
Publicado: (2022)
por: Zhang, Kevin, et al.
Publicado: (2022)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
MU-MAE: Multimodal Masked Autoencoders-Based One-Shot Learning
por: Liu, Rex, et al.
Publicado: (2024)
por: Liu, Rex, et al.
Publicado: (2024)
R-MAE: Regions Meet Masked Autoencoders
por: Nguyen, Duy-Kien, et al.
Publicado: (2023)
por: Nguyen, Duy-Kien, et al.
Publicado: (2023)
NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining
por: Zeng, Liang, et al.
Publicado: (2026)
por: Zeng, Liang, et al.
Publicado: (2026)
SenPa-MAE: Sensor Parameter Aware Masked Autoencoder for Multi-Satellite Self-Supervised Pretraining
por: Prexl, Jonathan, et al.
Publicado: (2024)
por: Prexl, Jonathan, et al.
Publicado: (2024)
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
por: Deria, Ankan, et al.
Publicado: (2025)
por: Deria, Ankan, et al.
Publicado: (2025)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition
por: Xiang, Peihao, et al.
Publicado: (2024)
por: Xiang, Peihao, et al.
Publicado: (2024)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
por: Ahamed, Shihab Aaqil, et al.
Publicado: (2025)
por: Ahamed, Shihab Aaqil, et al.
Publicado: (2025)
PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders
por: Zhang, Xiangdong, et al.
Publicado: (2024)
por: Zhang, Xiangdong, et al.
Publicado: (2024)
MV2MAE: Multi-View Video Masked Autoencoders
por: Shah, Ketul, et al.
Publicado: (2024)
por: Shah, Ketul, et al.
Publicado: (2024)
T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
por: Wei, Weijie, et al.
Publicado: (2023)
por: Wei, Weijie, et al.
Publicado: (2023)
Improving Visual Representation Alignment Generation with GRPO
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
TerraMAE: Learning Spatial-Spectral Representations from Hyperspectral Earth Observation Data via Adaptive Masked Autoencoders
por: Faruk, Tanjim Bin, et al.
Publicado: (2025)
por: Faruk, Tanjim Bin, et al.
Publicado: (2025)
Stochastic Siamese MAE Pretraining for Longitudinal Medical Images
por: Emre, Taha, et al.
Publicado: (2025)
por: Emre, Taha, et al.
Publicado: (2025)
Adaptive Hyper-Graph Convolution Network for Skeleton-based Human Action Recognition with Virtual Connections
por: Zhou, Youwei, et al.
Publicado: (2024)
por: Zhou, Youwei, et al.
Publicado: (2024)
Periodic-MAE: Periodic Video Masked Autoencoder for rPPG Estimation
por: Choi, Jiho, et al.
Publicado: (2025)
por: Choi, Jiho, et al.
Publicado: (2025)
Audio-visual Generalized Zero-shot Learning the Easy Way
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Ejemplares similares
-
Rethinking Positive Pairs in Contrastive Learning
por: Wu, Jiantao, et al.
Publicado: (2024) -
Pseudo Labelling for Enhanced Masked Autoencoders
por: Nandam, Srinivasa Rao, et al.
Publicado: (2024) -
Investigating Self-Supervised Methods for Label-Efficient Learning
por: Nandam, Srinivasa Rao, et al.
Publicado: (2024) -
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
por: Nazarieh, Fatemeh, et al.
Publicado: (2024) -
Probabilistically Aligned View-unaligned Clustering with Adaptive Template Selection
por: Dong, Wenhua, et al.
Publicado: (2024)