Self-Guided Masked Autoencoder
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Jeongwoo, Lee, Inseo, Lee, Junho, Lee, Joonseok |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
by: Lee, Inseo, et al.
Published: (2025)
by: Lee, Inseo, et al.
Published: (2025)
Scalable Frame Sampling for Video Classification: A Semi-Optimal Policy Approach with Reduced Search Space
by: Lee, Junho, et al.
Published: (2024)
by: Lee, Junho, et al.
Published: (2024)
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026)
by: Lee, Junho, et al.
Published: (2026)
Equivariant Latent Alignment via Flow Matching under Group Symmetries
by: Kim, Sunghyun, et al.
Published: (2026)
by: Kim, Sunghyun, et al.
Published: (2026)
Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schrödinger Bridge Matching
by: Shin, Jeongwoo, et al.
Published: (2026)
by: Shin, Jeongwoo, et al.
Published: (2026)
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024)
by: Hahm, Jaehoon, et al.
Published: (2024)
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
by: Ha, Seongsu, et al.
Published: (2024)
by: Ha, Seongsu, et al.
Published: (2024)
ArtSplat: Feed-Forward Articulated 3D Gaussian Splatting from Sparse Multi-State Uncalibrated Views
by: Lee, Inseo, et al.
Published: (2026)
by: Lee, Inseo, et al.
Published: (2026)
Modality-Aware Representation Learning for Zero-shot Sketch-based Image Retrieval
by: Lyou, Eunyi, et al.
Published: (2024)
by: Lyou, Eunyi, et al.
Published: (2024)
Towards Scalable Human-aligned Benchmark for Text-guided Image Editing
by: Ryu, Suho, et al.
Published: (2025)
by: Ryu, Suho, et al.
Published: (2025)
Periodic-MAE: Periodic Video Masked Autoencoder for rPPG Estimation
by: Choi, Jiho, et al.
Published: (2025)
by: Choi, Jiho, et al.
Published: (2025)
Text-Guided Video Masked Autoencoder
by: Fan, David, et al.
Published: (2024)
by: Fan, David, et al.
Published: (2024)
Efficient Masked Autoencoders with Self-Consistency
by: Li, Zhaowen, et al.
Published: (2023)
by: Li, Zhaowen, et al.
Published: (2023)
A Unified Masked Autoencoder with Patchified Skeletons for Motion Synthesis
by: Mascaro, Esteve Valls, et al.
Published: (2023)
by: Mascaro, Esteve Valls, et al.
Published: (2023)
Visual Style Prompting with Swapping Self-Attention
by: Jeong, Jaeseok, et al.
Published: (2024)
by: Jeong, Jaeseok, et al.
Published: (2024)
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM
by: Lee, Jeongwoo, et al.
Published: (2024)
by: Lee, Jeongwoo, et al.
Published: (2024)
Towards Motion-aware Referring Image Segmentation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
by: Hwang, Sunil, et al.
Published: (2022)
by: Hwang, Sunil, et al.
Published: (2022)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
by: Lee, Eunsoo, et al.
Published: (2026)
by: Lee, Eunsoo, et al.
Published: (2026)
Continual Self-Supervised Learning with Masked Autoencoders in Remote Sensing
by: Möllenbrok, Lars, et al.
Published: (2025)
by: Möllenbrok, Lars, et al.
Published: (2025)
Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics
by: Röhrich, Nikolai, et al.
Published: (2025)
by: Röhrich, Nikolai, et al.
Published: (2025)
Masked Capsule Autoencoders
by: Everett, Miles, et al.
Published: (2024)
by: Everett, Miles, et al.
Published: (2024)
MAESIL: Masked Autoencoder for Enhanced Self-supervised Medical Image Learning
by: Kim, Kyeonghun, et al.
Published: (2026)
by: Kim, Kyeonghun, et al.
Published: (2026)
Transformer with Leveraged Masked Autoencoder for video-based Pain Assessment
by: Nguyen, Minh-Duc, et al.
Published: (2024)
by: Nguyen, Minh-Duc, et al.
Published: (2024)
Improving Masked Autoencoders by Learning Where to Mask
by: Chen, Haijian, et al.
Published: (2023)
by: Chen, Haijian, et al.
Published: (2023)
OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D Perception
by: Koh, Junho, et al.
Published: (2025)
by: Koh, Junho, et al.
Published: (2025)
Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding
by: Zha, Yixin, et al.
Published: (2025)
by: Zha, Yixin, et al.
Published: (2025)
HyperKD: Distilling Cross-Spectral Knowledge in Masked Autoencoders via Inverse Domain Shift with Spatial-Aware Masking and Specialized Loss
by: Matin, Abdul, et al.
Published: (2025)
by: Matin, Abdul, et al.
Published: (2025)
Attention-Guided Masked Autoencoders For Learning Image Representations
by: Sick, Leon, et al.
Published: (2024)
by: Sick, Leon, et al.
Published: (2024)
Domain-Guided Masked Autoencoders for Unique Player Identification
by: Balaji, Bavesh, et al.
Published: (2024)
by: Balaji, Bavesh, et al.
Published: (2024)
MaskAdapt: Learning Flexible Motion Adaptation via Mask-Invariant Prior for Physics-Based Characters
by: Park, Soomin, et al.
Published: (2026)
by: Park, Soomin, et al.
Published: (2026)
Recurrent Video Masked Autoencoders
by: Zoran, Daniel, et al.
Published: (2025)
by: Zoran, Daniel, et al.
Published: (2025)
VG3T: Visual Geometry Grounded Gaussian Transformer
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Exemplar Masking for Multimodal Incremental Learning
by: Lee, Yi-Lun, et al.
Published: (2024)
by: Lee, Yi-Lun, et al.
Published: (2024)
Attention-Guided Multi-Scale Local Reconstruction for Point Clouds via Masked Autoencoder Self-Supervised Learning
by: Cao, Xin, et al.
Published: (2025)
by: Cao, Xin, et al.
Published: (2025)
Calibrating Panoramic Depth Estimation for Practical Localization and Mapping
by: Kim, Junho, et al.
Published: (2023)
by: Kim, Junho, et al.
Published: (2023)
Similar Items
-
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025) -
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
by: Lee, Inseo, et al.
Published: (2025) -
Scalable Frame Sampling for Video Classification: A Semi-Optimal Policy Approach with Reduced Search Space
by: Lee, Junho, et al.
Published: (2024) -
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025) -
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026)