SHViT: Single-Head Vision Transformer with Memory Efficient Macro Design
Fuente:
arXiv
Saved in:
| Main Authors: | Yun, Seokju, Ro, Youngmin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Partial Large Kernel CNNs for Efficient Super-Resolution
by: Lee, Dongheon, et al.
Published: (2024)
by: Lee, Dongheon, et al.
Published: (2024)
Emulating Self-attention with Convolution for Efficient Image Super-Resolution
by: Lee, Dongheon, et al.
Published: (2025)
by: Lee, Dongheon, et al.
Published: (2025)
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
by: Yun, Seokju, et al.
Published: (2024)
by: Yun, Seokju, et al.
Published: (2024)
Implicit Grid Convolution for Multi-Scale Image Super-Resolution
by: Lee, Dongheon, et al.
Published: (2024)
by: Lee, Dongheon, et al.
Published: (2024)
SoMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning
by: Yun, Seokju, et al.
Published: (2024)
by: Yun, Seokju, et al.
Published: (2024)
RecycleLoRA: Rank-Revealing QR-Based Dual-LoRA Subspace Adaptation for Domain Generalized Semantic Segmentation
by: Cho, Chanseul, et al.
Published: (2026)
by: Cho, Chanseul, et al.
Published: (2026)
StAR: Segment Anything Reasoner
by: Yun, Seokju, et al.
Published: (2026)
by: Yun, Seokju, et al.
Published: (2026)
Arbitrary-Scale Downscaling of Tidal Current Data Using Implicit Continuous Representation
by: Lee, Dongheon, et al.
Published: (2024)
by: Lee, Dongheon, et al.
Published: (2024)
OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation
by: Moon, Seungjae, et al.
Published: (2026)
by: Moon, Seungjae, et al.
Published: (2026)
Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
by: Hong, Sunghwan, et al.
Published: (2024)
by: Hong, Sunghwan, et al.
Published: (2024)
Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach
by: Zhang, Taolin, et al.
Published: (2024)
by: Zhang, Taolin, et al.
Published: (2024)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Designing Extremely Memory-Efficient CNNs for On-device Vision Tasks
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
ComFe: An Interpretable Head for Vision Transformers
by: Mannix, Evelyn J., et al.
Published: (2024)
by: Mannix, Evelyn J., et al.
Published: (2024)
Efficient Few-Shot Neural Architecture Search by Counting the Number of Nonlinear Functions
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
AAPL: Adding Attributes to Prompt Learning for Vision-Language Models
by: Kim, Gahyeon, et al.
Published: (2024)
by: Kim, Gahyeon, et al.
Published: (2024)
Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
by: Kim, Gahyeon, et al.
Published: (2025)
by: Kim, Gahyeon, et al.
Published: (2025)
Exploring Frequency-Inspired Optimization in Transformer for Efficient Single Image Super-Resolution
by: Li, Ao, et al.
Published: (2023)
by: Li, Ao, et al.
Published: (2023)
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking
by: Chung, Sangyun, et al.
Published: (2024)
by: Chung, Sangyun, et al.
Published: (2024)
UCMNet: Uncertainty-Aware Context Memory Network for Under-Display Camera Image Restoration
by: Kim, Daehyun, et al.
Published: (2026)
by: Kim, Daehyun, et al.
Published: (2026)
Weak Supervision with Arbitrary Single Frame for Micro- and Macro-expression Spotting
by: Yu, Wang-Wang, et al.
Published: (2024)
by: Yu, Wang-Wang, et al.
Published: (2024)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
by: Jo, Sehyeong, et al.
Published: (2025)
by: Jo, Sehyeong, et al.
Published: (2025)
Memory Efficient Transformer Adapter for Dense Predictions
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
by: Yu, Youngjoon, et al.
Published: (2024)
by: Yu, Youngjoon, et al.
Published: (2024)
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
by: Ro, Juneyoung, et al.
Published: (2025)
by: Ro, Juneyoung, et al.
Published: (2025)
Zero Memory Overhead Approach for Protecting Vision Transformer Parameters
by: Baradaran, Fereshteh, et al.
Published: (2025)
by: Baradaran, Fereshteh, et al.
Published: (2025)
Vision Transformers with Hierarchical Attention
by: Liu, Yun, et al.
Published: (2021)
by: Liu, Yun, et al.
Published: (2021)
PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing
by: Oroz, Antonio, et al.
Published: (2025)
by: Oroz, Antonio, et al.
Published: (2025)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Convolutional Initialization for Data-Efficient Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
Multi-Tailed Vision Transformer for Efficient Inference
by: Wang, Yunke, et al.
Published: (2022)
by: Wang, Yunke, et al.
Published: (2022)
DSD-GS: Dynamic-Static Decomposition of Gaussian Splatting for Efficient and High-Fidelity Dynamic Scene Reconstruction
by: Han, Youngtae, et al.
Published: (2026)
by: Han, Youngtae, et al.
Published: (2026)
Residual Attention Single-Head Vision Transformer Network for Rolling Bearing Fault Diagnosis in Noisy Environments
by: Lai, Songjiang, et al.
Published: (2024)
by: Lai, Songjiang, et al.
Published: (2024)
Seurat: From Moving Points to Depth
by: Cho, Seokju, et al.
Published: (2025)
by: Cho, Seokju, et al.
Published: (2025)
Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learning
by: Zhong, Hanwen, et al.
Published: (2025)
by: Zhong, Hanwen, et al.
Published: (2025)
MDeRainNet: An Efficient Macro-pixel Image Rain Removal Network
by: Yan, Tao, et al.
Published: (2024)
by: Yan, Tao, et al.
Published: (2024)
From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
by: Li, Zongzhao, et al.
Published: (2025)
by: Li, Zongzhao, et al.
Published: (2025)
Similar Items
-
Partial Large Kernel CNNs for Efficient Super-Resolution
by: Lee, Dongheon, et al.
Published: (2024) -
Emulating Self-attention with Convolution for Efficient Image Super-Resolution
by: Lee, Dongheon, et al.
Published: (2025) -
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
by: Yun, Seokju, et al.
Published: (2024) -
Implicit Grid Convolution for Multi-Scale Image Super-Resolution
by: Lee, Dongheon, et al.
Published: (2024) -
SoMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning
by: Yun, Seokju, et al.
Published: (2024)