Masking meets Supervision: A Strong Learning Alliance
Fuente:
arXiv
Saved in:
| Main Authors: | Heo, Byeongho, Kim, Taekyung, Yun, Sangdoo, Han, Dongyoon |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)
by: Kim, Jiwon, et al.
Published: (2023)
Morphing Tokens Draw Strong Masked Image Models
by: Kim, Taekyung, et al.
Published: (2023)
by: Kim, Taekyung, et al.
Published: (2023)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
Token Bottleneck: One Token to Remember Dynamics
by: Kim, Taekyung, et al.
Published: (2025)
by: Kim, Taekyung, et al.
Published: (2025)
RL makes MLLMs see better than SFT
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
Learning with Unmasked Tokens Drives Stronger Vision Learners
by: Kim, Taekyung, et al.
Published: (2023)
by: Kim, Taekyung, et al.
Published: (2023)
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
by: Park, Song, et al.
Published: (2025)
by: Park, Song, et al.
Published: (2025)
Model Stock: All we need is just a few fine-tuned models
by: Jang, Dong-Hwan, et al.
Published: (2024)
by: Jang, Dong-Hwan, et al.
Published: (2024)
Exploring Conditions for Diffusion models in Robotic Control
by: Shin, Heeseong, et al.
Published: (2025)
by: Shin, Heeseong, et al.
Published: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
by: Kim, Wonjae, et al.
Published: (2024)
by: Kim, Wonjae, et al.
Published: (2024)
DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
by: Oh, Changdae, et al.
Published: (2024)
by: Oh, Changdae, et al.
Published: (2024)
Similarity of Neural Architectures using Adversarial Attack Transferability
by: Hwang, Jaehui, et al.
Published: (2022)
by: Hwang, Jaehui, et al.
Published: (2022)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
by: Lee, Minhyun, et al.
Published: (2023)
by: Lee, Minhyun, et al.
Published: (2023)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text Guidance
by: Nam, Giung, et al.
Published: (2024)
by: Nam, Giung, et al.
Published: (2024)
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation
by: Lee, Minhyun, et al.
Published: (2024)
by: Lee, Minhyun, et al.
Published: (2024)
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
by: Chun, Sanghyuk, et al.
Published: (2025)
by: Chun, Sanghyuk, et al.
Published: (2025)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
by: Kwak, Min-Seop, et al.
Published: (2025)
by: Kwak, Min-Seop, et al.
Published: (2025)
When Test-Time Adaptation Meets Self-Supervised Models
by: Han, Jisu, et al.
Published: (2025)
by: Han, Jisu, et al.
Published: (2025)
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024)
by: Chun, Sanghyuk, et al.
Published: (2024)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Angular Gradient Sign Method: Uncovering Vulnerabilities in Hyperbolic Networks
by: Jo, Minsoo, et al.
Published: (2025)
by: Jo, Minsoo, et al.
Published: (2025)
METAVerse: Meta-Learning Traversability Cost Map for Off-Road Navigation
by: Seo, Junwon, et al.
Published: (2023)
by: Seo, Junwon, et al.
Published: (2023)
Selective Masking based Self-Supervised Learning for Image Semantic Segmentation
by: Wang, Yuemin, et al.
Published: (2025)
by: Wang, Yuemin, et al.
Published: (2025)
GeoMask3D: Geometrically Informed Mask Selection for Self-Supervised Point Cloud Learning in 3D
by: Bahri, Ali, et al.
Published: (2024)
by: Bahri, Ali, et al.
Published: (2024)
Structure is Supervision: Multiview Masked Autoencoders for Radiology
by: Laguna, Sonia, et al.
Published: (2025)
by: Laguna, Sonia, et al.
Published: (2025)
Motion-Oriented Compositional Neural Radiance Fields for Monocular Dynamic Human Modeling
by: Kim, Jaehyeok, et al.
Published: (2024)
by: Kim, Jaehyeok, et al.
Published: (2024)
MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
On Improving the Algorithm-, Model-, and Data- Efficiency of Self-Supervised Learning
by: Cao, Yun-Hao, et al.
Published: (2024)
by: Cao, Yun-Hao, et al.
Published: (2024)
The Role of Masking for Efficient Supervised Knowledge Distillation of Vision Transformers
by: Son, Seungwoo, et al.
Published: (2023)
by: Son, Seungwoo, et al.
Published: (2023)
SupMAE: Supervised Masked Autoencoders Are Efficient Vision Learners
by: Liang, Feng, et al.
Published: (2022)
by: Liang, Feng, et al.
Published: (2022)
Downstream Task Guided Masking Learning in Masked Autoencoders Using Multi-Level Optimization
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
by: Kim, Donghu, et al.
Published: (2024)
by: Kim, Donghu, et al.
Published: (2024)
StyleLipSync: Style-based Personalized Lip-sync Video Generation
by: Ki, Taekyung, et al.
Published: (2023)
by: Ki, Taekyung, et al.
Published: (2023)
Soft Equivariance Regularization for Invariant Self-Supervised Learning
by: Lee, Joohyung, et al.
Published: (2026)
by: Lee, Joohyung, et al.
Published: (2026)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
PolMERLIN: Self-Supervised Polarimetric Complex SAR Image Despeckling with Masked Networks
by: Kato, Shunya, et al.
Published: (2024)
by: Kato, Shunya, et al.
Published: (2024)
A Billion-scale Foundation Model for Remote Sensing Images
by: Cha, Keumgang, et al.
Published: (2023)
by: Cha, Keumgang, et al.
Published: (2023)
Similar Items
-
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023) -
Morphing Tokens Draw Strong Masked Image Models
by: Kim, Taekyung, et al.
Published: (2023) -
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024) -
Token Bottleneck: One Token to Remember Dynamics
by: Kim, Taekyung, et al.
Published: (2025) -
RL makes MLLMs see better than SFT
by: Song, Junha, et al.
Published: (2025)