MetaFormer Baselines for Vision
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Weihao, Si, Chenyang, Zhou, Pan, Luo, Mi, Zhou, Yichen, Feng, Jiashi, Yan, Shuicheng, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InceptionNeXt: When Inception Meets ConvNeXt
by: Yu, Weihao, et al.
Published: (2023)
by: Yu, Weihao, et al.
Published: (2023)
MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic Segmentation
by: Kang, Beoungwoo, et al.
Published: (2024)
by: Kang, Beoungwoo, et al.
Published: (2024)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
MambaOut: Do We Really Need Mamba for Vision?
by: Yu, Weihao, et al.
Published: (2024)
by: Yu, Weihao, et al.
Published: (2024)
Gamba: Marry Gaussian Splatting with Mamba for single view 3D reconstruction
by: Shen, Qiuhong, et al.
Published: (2024)
by: Shen, Qiuhong, et al.
Published: (2024)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
by: Ren, Lingfeng, et al.
Published: (2026)
by: Ren, Lingfeng, et al.
Published: (2026)
Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
by: Keuth, Ron, et al.
Published: (2025)
by: Keuth, Ron, et al.
Published: (2025)
Isomorphic Pruning for Vision Models
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
Seeing World Dynamics in a Nutshell
by: Shen, Qiuhong, et al.
Published: (2025)
by: Shen, Qiuhong, et al.
Published: (2025)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
by: Meng, Weikang, et al.
Published: (2025)
by: Meng, Weikang, et al.
Published: (2025)
Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models
by: Wang, Jeffrey, et al.
Published: (2026)
by: Wang, Jeffrey, et al.
Published: (2026)
MetaFormer-driven Encoding Network for Robust Medical Semantic Segmentation
by: Tran, Le-Anh, et al.
Published: (2026)
by: Tran, Le-Anh, et al.
Published: (2026)
GaussianFormer: Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
mRadNet: A Compact Radar Object Detector with MetaFormer
by: Chen, Huaiyu, et al.
Published: (2025)
by: Chen, Huaiyu, et al.
Published: (2025)
When Training-Free NAS Meets Vision Transformer: A Neural Tangent Kernel Perspective
by: Zhou, Qiqi, et al.
Published: (2024)
by: Zhou, Qiqi, et al.
Published: (2024)
WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave Equation
by: Shu, Zishan, et al.
Published: (2026)
by: Shu, Zishan, et al.
Published: (2026)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2025)
by: Diao, Haiwen, et al.
Published: (2025)
Vision Bridge Transformer at Scale
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models
by: Shi, Zhiyi, et al.
Published: (2025)
by: Shi, Zhiyi, et al.
Published: (2025)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
AdjointDPM: Adjoint Sensitivity Method for Gradient Backpropagation of Diffusion Probabilistic Models
by: Pan, Jiachun, et al.
Published: (2023)
by: Pan, Jiachun, et al.
Published: (2023)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
by: Li, Wenxi, et al.
Published: (2025)
by: Li, Wenxi, et al.
Published: (2025)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
by: Dong, Xinpeng, et al.
Published: (2026)
by: Dong, Xinpeng, et al.
Published: (2026)
HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
by: Wang, Daming, et al.
Published: (2025)
by: Wang, Daming, et al.
Published: (2025)
HGTS-Former: Hierarchical HyperGraph Transformer for Multivariate Time Series Analysis
by: Si, Hao, et al.
Published: (2025)
by: Si, Hao, et al.
Published: (2025)
Active Multimodal Distillation for Few-shot Action Recognition
by: Feng, Weijia, et al.
Published: (2025)
by: Feng, Weijia, et al.
Published: (2025)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
by: Xu, Wanting, et al.
Published: (2024)
by: Xu, Wanting, et al.
Published: (2024)
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
by: Heinrich, Greg, et al.
Published: (2024)
by: Heinrich, Greg, et al.
Published: (2024)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
Make Geometry Matter for Spatial Reasoning
by: Zhang, Shihua, et al.
Published: (2026)
by: Zhang, Shihua, et al.
Published: (2026)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
High-throughput digital twin framework for predicting neurite deterioration using MetaFormer attention
by: Qian, Kuanren, et al.
Published: (2024)
by: Qian, Kuanren, et al.
Published: (2024)
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
A Data-Centric Vision Transformer Baseline for SAR Sea Ice Classification
by: Mike-Ewewie, David, et al.
Published: (2026)
by: Mike-Ewewie, David, et al.
Published: (2026)
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
RealDPO: Real or Not Real, that is the Preference
by: Cheng, Guo, et al.
Published: (2025)
by: Cheng, Guo, et al.
Published: (2025)
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
by: Wang, Weimin, et al.
Published: (2024)
by: Wang, Weimin, et al.
Published: (2024)
Similar Items
-
InceptionNeXt: When Inception Meets ConvNeXt
by: Yu, Weihao, et al.
Published: (2023) -
MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic Segmentation
by: Kang, Beoungwoo, et al.
Published: (2024) -
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024) -
MambaOut: Do We Really Need Mamba for Vision?
by: Yu, Weihao, et al.
Published: (2024) -
Gamba: Marry Gaussian Splatting with Mamba for single view 3D reconstruction
by: Shen, Qiuhong, et al.
Published: (2024)