Efficiency Follows Global-Local Decoupling
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914410161766400 |
|---|---|
| author | Yang, Zhenyu Pei, Gensheng Chen, Tao Zhou, Yichao Zhou, Tianfei Yao, Yazhou Shen, Fumin |
| author_facet | Yang, Zhenyu Pei, Gensheng Chen, Tao Zhou, Yichao Zhou, Tianfei Yao, Yazhou Shen, Fumin |
| contents | Modern vision models must capture image-level context without sacrificing local detail while remaining computationally affordable. We revisit this tradeoff and advance a simple principle: decouple the roles of global reasoning and local representation. To operationalize this principle, we introduce ConvNeur, a two-branch architecture in which a lightweight neural memory branch aggregates global context on a compact set of tokens, and a locality-preserving branch extracts fine structure. A learned gate lets global cues modulate local features without entangling their objectives. This separation yields subquadratic scaling with image size, retains inductive priors associated with local processing, and reduces overhead relative to fully global attention. On standard classification, detection, and segmentation benchmarks, ConvNeur matches or surpasses comparable alternatives at similar or lower compute and offers favorable accuracy versus latency trade-offs at similar budgets. These results support the view that efficiency follows global-local decoupling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_19567 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Efficiency Follows Global-Local Decoupling Yang, Zhenyu Pei, Gensheng Chen, Tao Zhou, Yichao Zhou, Tianfei Yao, Yazhou Shen, Fumin Computer Vision and Pattern Recognition Modern vision models must capture image-level context without sacrificing local detail while remaining computationally affordable. We revisit this tradeoff and advance a simple principle: decouple the roles of global reasoning and local representation. To operationalize this principle, we introduce ConvNeur, a two-branch architecture in which a lightweight neural memory branch aggregates global context on a compact set of tokens, and a locality-preserving branch extracts fine structure. A learned gate lets global cues modulate local features without entangling their objectives. This separation yields subquadratic scaling with image size, retains inductive priors associated with local processing, and reduces overhead relative to fully global attention. On standard classification, detection, and segmentation benchmarks, ConvNeur matches or surpasses comparable alternatives at similar or lower compute and offers favorable accuracy versus latency trade-offs at similar budgets. These results support the view that efficiency follows global-local decoupling. |
| title | Efficiency Follows Global-Local Decoupling |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.19567 |