Rethinking Patch Dependence for Masked Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Letian, Lian, Long, Wang, Renhao, Shi, Baifeng, Wang, Xudong, Yala, Adam, Darrell, Trevor, Efros, Alexei A., Goldberg, Ken |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-grounded Video Diffusion Models
by: Lian, Long, et al.
Published: (2023)
by: Lian, Long, et al.
Published: (2023)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
by: Lian, Long, et al.
Published: (2023)
by: Lian, Long, et al.
Published: (2023)
When Do We Not Need Larger Vision Models?
by: Shi, Baifeng, et al.
Published: (2024)
by: Shi, Baifeng, et al.
Published: (2024)
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
by: Wang, Renhao, et al.
Published: (2025)
by: Wang, Renhao, et al.
Published: (2025)
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)
by: Ge, Jiaxin, et al.
Published: (2023)
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
by: Harrington, Anne, et al.
Published: (2025)
by: Harrington, Anne, et al.
Published: (2025)
Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting
by: Bhattad, Anand, et al.
Published: (2025)
by: Bhattad, Anand, et al.
Published: (2025)
Atlas: Multi-Scale Attention Improves Long Context Image Modeling
by: Agrawal, Kumar Krishna, et al.
Published: (2025)
by: Agrawal, Kumar Krishna, et al.
Published: (2025)
Test-Time Training on Video Streams
by: Wang, Renhao, et al.
Published: (2023)
by: Wang, Renhao, et al.
Published: (2023)
Visually Prompted Benchmarks Are Surprisingly Fragile
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)
by: Tang, Zineng, et al.
Published: (2025)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
by: Shi, Baifeng, et al.
Published: (2026)
by: Shi, Baifeng, et al.
Published: (2026)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
Interpreting the Second-Order Effects of Neurons in CLIP
by: Gandelsman, Yossi, et al.
Published: (2024)
by: Gandelsman, Yossi, et al.
Published: (2024)
COLMAP-Free 3D Gaussian Splatting
by: Fu, Yang, et al.
Published: (2023)
by: Fu, Yang, et al.
Published: (2023)
Improving Masked Autoencoders by Learning Where to Mask
by: Chen, Haijian, et al.
Published: (2023)
by: Chen, Haijian, et al.
Published: (2023)
SelfMedHPM: Self Pre-training With Hard Patches Mining Masked Autoencoders For Medical Image Segmentation
by: Lv, Yunhao, et al.
Published: (2025)
by: Lv, Yunhao, et al.
Published: (2025)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024)
by: Niu, Dantong, et al.
Published: (2024)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
Continuous 3D Perception Model with Persistent State
by: Wang, Qianqian, et al.
Published: (2025)
by: Wang, Qianqian, et al.
Published: (2025)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
by: Gan, Yulu, et al.
Published: (2025)
by: Gan, Yulu, et al.
Published: (2025)
Scaling Vision Pre-Training to 4K Resolution
by: Shi, Baifeng, et al.
Published: (2025)
by: Shi, Baifeng, et al.
Published: (2025)
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025)
by: Kutscher, Declan, et al.
Published: (2025)
Text-Guided Video Masked Autoencoder
by: Fan, David, et al.
Published: (2024)
by: Fan, David, et al.
Published: (2024)
Rethinking Remote Sensing Change Detection With A Mask View
by: Ma, Xiaowen, et al.
Published: (2024)
by: Ma, Xiaowen, et al.
Published: (2024)
Masked Autoencoders are Parameter-Efficient Federated Continual Learners
by: He, Yuchen, et al.
Published: (2024)
by: He, Yuchen, et al.
Published: (2024)
Efficient Masked Autoencoders with Self-Consistency
by: Li, Zhaowen, et al.
Published: (2023)
by: Li, Zhaowen, et al.
Published: (2023)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Masked Capsule Autoencoders
by: Everett, Miles, et al.
Published: (2024)
by: Everett, Miles, et al.
Published: (2024)
Contrastive Masked Autoencoders are Stronger Vision Learners
by: Huang, Zhicheng, et al.
Published: (2022)
by: Huang, Zhicheng, et al.
Published: (2022)
Pillar-0: A New Frontier for Radiology Foundation Models
by: Agrawal, Kumar Krishna, et al.
Published: (2025)
by: Agrawal, Kumar Krishna, et al.
Published: (2025)
Learning from the Right Patches: A Two-Stage Wavelet-Driven Masked Autoencoder for Histopathology Representation Learning
by: Younis, Raneen, et al.
Published: (2025)
by: Younis, Raneen, et al.
Published: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
MARMOT: Masked Autoencoder for Modeling Transient Imaging
by: Shen, Siyuan, et al.
Published: (2025)
by: Shen, Siyuan, et al.
Published: (2025)
GPS as a Control Signal for Image Generation
by: Feng, Chao, et al.
Published: (2025)
by: Feng, Chao, et al.
Published: (2025)
Recurrent Video Masked Autoencoders
by: Zoran, Daniel, et al.
Published: (2025)
by: Zoran, Daniel, et al.
Published: (2025)
Self-Guided Masked Autoencoder
by: Shin, Jeongwoo, et al.
Published: (2025)
by: Shin, Jeongwoo, et al.
Published: (2025)
Similar Items
-
LLM-grounded Video Diffusion Models
by: Lian, Long, et al.
Published: (2023) -
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
by: Lian, Long, et al.
Published: (2023) -
When Do We Not Need Larger Vision Models?
by: Shi, Baifeng, et al.
Published: (2024) -
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
by: Wang, Renhao, et al.
Published: (2025) -
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)