Rethink MAE with Linear Time-Invariant Dynamics
Fuente:
arXiv
Saved in:
| Main Author: | Wang, Zice |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SatSwinMAE: Efficient Autoencoding for Multiscale Time-series Satellite Imagery
by: Nakayama, Yohei, et al.
Published: (2024)
by: Nakayama, Yohei, et al.
Published: (2024)
i-MAE: Are Latent Representations in Masked Autoencoders Linearly Separable?
by: Zhang, Kevin, et al.
Published: (2022)
by: Zhang, Kevin, et al.
Published: (2022)
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
by: Feng, Zhe, et al.
Published: (2026)
by: Feng, Zhe, et al.
Published: (2026)
OceanMAE: A Foundation Model for Ocean Remote Sensing
by: Stamer, Viola-Joanna, et al.
Published: (2026)
by: Stamer, Viola-Joanna, et al.
Published: (2026)
NEMESIS: Noise-suppressed Efficient MAE with Enhanced Superpatch Integration Strategy
by: Kim, Kyeonghun, et al.
Published: (2026)
by: Kim, Kyeonghun, et al.
Published: (2026)
Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs
by: Zaazou, Youssef, et al.
Published: (2026)
by: Zaazou, Youssef, et al.
Published: (2026)
SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark Estimation
by: Yin, Kejia, et al.
Published: (2024)
by: Yin, Kejia, et al.
Published: (2024)
OphMAE: Bridging Volumetric and Planar Imaging with a Foundation Model for Adaptive Ophthalmological Diagnosis
by: Chang, Tienyu, et al.
Published: (2026)
by: Chang, Tienyu, et al.
Published: (2026)
ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
by: Wei, Guoyizhe, et al.
Published: (2025)
by: Wei, Guoyizhe, et al.
Published: (2025)
MAE-Based Self-Supervised Pretraining for Data-Efficient Medical Image Segmentation Using nnFormer
by: Sureddi, R. M. Krishna, et al.
Published: (2026)
by: Sureddi, R. M. Krishna, et al.
Published: (2026)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
by: Tao, Ziyuan, et al.
Published: (2025)
by: Tao, Ziyuan, et al.
Published: (2025)
CL-MAE: Curriculum-Learned Masked Autoencoders
by: Madan, Neelu, et al.
Published: (2023)
by: Madan, Neelu, et al.
Published: (2023)
L-MAE: Longitudinal masked auto-encoder with time and severity-aware encoding for diabetic retinopathy progression prediction
by: Zeghlache, Rachid, et al.
Published: (2024)
by: Zeghlache, Rachid, et al.
Published: (2024)
Dynamic Modality-Camera Invariant Clustering for Unsupervised Visible-Infrared Person Re-identification
by: Yang, Yiming, et al.
Published: (2024)
by: Yang, Yiming, et al.
Published: (2024)
Relative-Absolute Fusion: Rethinking Feature Extraction in Image-Based Iterative Method Selection for Solving Sparse Linear Systems
by: Zhang, Kaiqi, et al.
Published: (2025)
by: Zhang, Kaiqi, et al.
Published: (2025)
Learning Non-Linear Invariants for Unsupervised Out-of-Distribution Detection
by: Doorenbos, Lars, et al.
Published: (2024)
by: Doorenbos, Lars, et al.
Published: (2024)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
by: Singh, Mannat, et al.
Published: (2023)
by: Singh, Mannat, et al.
Published: (2023)
Rethinking Unsupervised Domain Adaptation for Semantic Segmentation
by: Wang, Zhijie, et al.
Published: (2022)
by: Wang, Zhijie, et al.
Published: (2022)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
by: Luo, Yuanhao, et al.
Published: (2026)
by: Luo, Yuanhao, et al.
Published: (2026)
Rethinking Intracranial Aneurysm Vessel Segmentation: A Perspective from Computational Fluid Dynamics Applications
by: Xiao, Feiyang, et al.
Published: (2025)
by: Xiao, Feiyang, et al.
Published: (2025)
Rethinking Token Reduction for Large Vision-Language Models
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Rethinking Metrics and Benchmarks of Video Anomaly Detection
by: Liu, Zihao, et al.
Published: (2025)
by: Liu, Zihao, et al.
Published: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
SF-Mamba: Rethinking State Space Model for Vision
by: Yoshimura, Masakazu, et al.
Published: (2026)
by: Yoshimura, Masakazu, et al.
Published: (2026)
Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance
by: Shen, Dazhong, et al.
Published: (2024)
by: Shen, Dazhong, et al.
Published: (2024)
Revisiting MAE pre-training for 3D medical image segmentation
by: Wald, Tassilo, et al.
Published: (2024)
by: Wald, Tassilo, et al.
Published: (2024)
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
by: Picard, David, et al.
Published: (2026)
by: Picard, David, et al.
Published: (2026)
Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation
by: Zhu, Chunzheng, et al.
Published: (2026)
by: Zhu, Chunzheng, et al.
Published: (2026)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
KernelWarehouse: Rethinking the Design of Dynamic Convolution
by: Li, Chao, et al.
Published: (2024)
by: Li, Chao, et al.
Published: (2024)
Generalizable Sensor-Based Activity Recognition via Categorical Concept Invariant Learning
by: Xiong, Di, et al.
Published: (2024)
by: Xiong, Di, et al.
Published: (2024)
Rethinking Normalization Strategies and Convolutional Kernels for Multimodal Image Fusion
by: He, Dan, et al.
Published: (2024)
by: He, Dan, et al.
Published: (2024)
CAR: Contrast-Agnostic Deformable Medical Image Registration with Contrast-Invariant Latent Regularization
by: Wang, Yinsong, et al.
Published: (2024)
by: Wang, Yinsong, et al.
Published: (2024)
OccluTrack: Rethinking Awareness of Occlusion for Enhancing Multiple Pedestrian Tracking
by: Gao, Jianjun, et al.
Published: (2023)
by: Gao, Jianjun, et al.
Published: (2023)
Domain-Invariant Prompt Learning for Vision-Language Models
by: Khoee, Arsham Gholamzadeh, et al.
Published: (2026)
by: Khoee, Arsham Gholamzadeh, et al.
Published: (2026)
V-VIPE: Variational View Invariant Pose Embedding
by: Levy, Mara, et al.
Published: (2024)
by: Levy, Mara, et al.
Published: (2024)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
Camera-Invariant Meta-Learning Network for Single-Camera-Training Person Re-identification
by: Pei, Jiangbo, et al.
Published: (2024)
by: Pei, Jiangbo, et al.
Published: (2024)
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
by: Fan, Jiawei, et al.
Published: (2026)
by: Fan, Jiawei, et al.
Published: (2026)
GATS: Gaussian Aware Temporal Scaling Transformer for Invariant 4D Spatio-Temporal Point Cloud Representation
by: Tian, Jiayi, et al.
Published: (2026)
by: Tian, Jiayi, et al.
Published: (2026)
Similar Items
-
SatSwinMAE: Efficient Autoencoding for Multiscale Time-series Satellite Imagery
by: Nakayama, Yohei, et al.
Published: (2024) -
i-MAE: Are Latent Representations in Masked Autoencoders Linearly Separable?
by: Zhang, Kevin, et al.
Published: (2022) -
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
by: Feng, Zhe, et al.
Published: (2026) -
OceanMAE: A Foundation Model for Ocean Remote Sensing
by: Stamer, Viola-Joanna, et al.
Published: (2026) -
NEMESIS: Noise-suppressed Efficient MAE with Enhanced Superpatch Integration Strategy
by: Kim, Kyeonghun, et al.
Published: (2026)