Efficiency Follows Global-Local Decoupling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhenyu, Pei, Gensheng, Chen, Tao, Zhou, Yichao, Zhou, Tianfei, Yao, Yazhou, Shen, Fumin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914410161766400
author Yang, Zhenyu
Pei, Gensheng
Chen, Tao
Zhou, Yichao
Zhou, Tianfei
Yao, Yazhou
Shen, Fumin
author_facet Yang, Zhenyu
Pei, Gensheng
Chen, Tao
Zhou, Yichao
Zhou, Tianfei
Yao, Yazhou
Shen, Fumin
contents Modern vision models must capture image-level context without sacrificing local detail while remaining computationally affordable. We revisit this tradeoff and advance a simple principle: decouple the roles of global reasoning and local representation. To operationalize this principle, we introduce ConvNeur, a two-branch architecture in which a lightweight neural memory branch aggregates global context on a compact set of tokens, and a locality-preserving branch extracts fine structure. A learned gate lets global cues modulate local features without entangling their objectives. This separation yields subquadratic scaling with image size, retains inductive priors associated with local processing, and reduces overhead relative to fully global attention. On standard classification, detection, and segmentation benchmarks, ConvNeur matches or surpasses comparable alternatives at similar or lower compute and offers favorable accuracy versus latency trade-offs at similar budgets. These results support the view that efficiency follows global-local decoupling.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19567
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Efficiency Follows Global-Local Decoupling
Yang, Zhenyu
Pei, Gensheng
Chen, Tao
Zhou, Yichao
Zhou, Tianfei
Yao, Yazhou
Shen, Fumin
Computer Vision and Pattern Recognition
Modern vision models must capture image-level context without sacrificing local detail while remaining computationally affordable. We revisit this tradeoff and advance a simple principle: decouple the roles of global reasoning and local representation. To operationalize this principle, we introduce ConvNeur, a two-branch architecture in which a lightweight neural memory branch aggregates global context on a compact set of tokens, and a locality-preserving branch extracts fine structure. A learned gate lets global cues modulate local features without entangling their objectives. This separation yields subquadratic scaling with image size, retains inductive priors associated with local processing, and reduces overhead relative to fully global attention. On standard classification, detection, and segmentation benchmarks, ConvNeur matches or surpasses comparable alternatives at similar or lower compute and offers favorable accuracy versus latency trade-offs at similar budgets. These results support the view that efficiency follows global-local decoupling.
title Efficiency Follows Global-Local Decoupling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.19567