GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hyunju, Oh, Youngmin, Jeon, Jeimin, Baek, Donghyeon, Ham, Bumsub |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Subnet-Aware Dynamic Supernet Training for Neural Architecture Search
by: Jeon, Jeimin, et al.
Published: (2025)
by: Jeon, Jeimin, et al.
Published: (2025)
TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts
by: Jeon, Jeimin, et al.
Published: (2026)
by: Jeon, Jeimin, et al.
Published: (2026)
Efficient Few-Shot Neural Architecture Search by Counting the Number of Nonlinear Functions
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
FYI: Flip Your Images for Dataset Distillation
by: Son, Byunggwan, et al.
Published: (2024)
by: Son, Byunggwan, et al.
Published: (2024)
Scheduling Weight Transitions for Quantization-Aware Training
by: Lee, Junghyup, et al.
Published: (2024)
by: Lee, Junghyup, et al.
Published: (2024)
Relational Feature Caching for Accelerating Diffusion Transformers
by: Son, Byunggwan, et al.
Published: (2026)
by: Son, Byunggwan, et al.
Published: (2026)
Toward INT4 Fixed-Point Training via Exploring Quantization Error for Gradients
by: Kim, Dohyung, et al.
Published: (2024)
by: Kim, Dohyung, et al.
Published: (2024)
AZ-NAS: Assembling Zero-Cost Proxies for Network Architecture Search
by: Lee, Junghyup, et al.
Published: (2024)
by: Lee, Junghyup, et al.
Published: (2024)
AccuQuant: Simulating Multiple Denoising Steps for Quantizing Diffusion Models
by: Lee, Seunghoon, et al.
Published: (2025)
by: Lee, Seunghoon, et al.
Published: (2025)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
by: Cho, Hyeonwoo, et al.
Published: (2026)
by: Cho, Hyeonwoo, et al.
Published: (2026)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
by: Han, Donghoon, et al.
Published: (2023)
by: Han, Donghoon, et al.
Published: (2023)
Maximizing the Position Embedding for Vision Transformers with Global Average Pooling
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection
by: Lee, Sanghoon, et al.
Published: (2026)
by: Lee, Sanghoon, et al.
Published: (2026)
3DPillars: Pillar-based two-stage 3D object detection
by: Noh, Jongyoun, et al.
Published: (2025)
by: Noh, Jongyoun, et al.
Published: (2025)
Disentangled Representations for Short-Term and Long-Term Person Re-Identification
by: Eom, Chanho, et al.
Published: (2024)
by: Eom, Chanho, et al.
Published: (2024)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
by: Li, Zhengang, et al.
Published: (2024)
by: Li, Zhengang, et al.
Published: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
Deeper Inside Deep ViT
by: Hong, Sungrae
Published: (2025)
by: Hong, Sungrae
Published: (2025)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
by: Zhong, Yunshan, et al.
Published: (2023)
by: Zhong, Yunshan, et al.
Published: (2023)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
Instance-Aware Group Quantization for Vision Transformers
by: Moon, Jaehyeon, et al.
Published: (2024)
by: Moon, Jaehyeon, et al.
Published: (2024)
ToaSt: Token Channel Selection and Structured Pruning for Efficient ViT
by: Moon, Hyunchan, et al.
Published: (2026)
by: Moon, Hyunchan, et al.
Published: (2026)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
by: Chattopadhyay, Nandish, et al.
Published: (2026)
by: Chattopadhyay, Nandish, et al.
Published: (2026)
Frequency-Adaptive Discrete Cosine-ViT-ResNet Architecture for Sparse-Data Vision
by: Kang, Ziyue, et al.
Published: (2025)
by: Kang, Ziyue, et al.
Published: (2025)
ViTCAE: ViT-based Class-conditioned Autoencoder
by: Jebraeeli, Vahid, et al.
Published: (2025)
by: Jebraeeli, Vahid, et al.
Published: (2025)
HyTAS: A Hyperspectral Image Transformer Architecture Search Benchmark and Analysis
by: Zhou, Fangqin, et al.
Published: (2024)
by: Zhou, Fangqin, et al.
Published: (2024)
Cerberus: Attribute-based person re-identification using semantic IDs
by: Eom, Chanho, et al.
Published: (2024)
by: Eom, Chanho, et al.
Published: (2024)
Partial Large Kernel CNNs for Efficient Super-Resolution
by: Lee, Dongheon, et al.
Published: (2024)
by: Lee, Dongheon, et al.
Published: (2024)
Rethinking Random Masking in Self-Distillation on ViT
by: Seong, Jihyeon, et al.
Published: (2025)
by: Seong, Jihyeon, et al.
Published: (2025)
YOLO-Former: YOLO Shakes Hand With ViT
by: Khoramdel, Javad, et al.
Published: (2024)
by: Khoramdel, Javad, et al.
Published: (2024)
Your ViT is Secretly an Image Segmentation Model
by: Kerssies, Tommie, et al.
Published: (2025)
by: Kerssies, Tommie, et al.
Published: (2025)
ViT-5: Vision Transformers for The Mid-2020s
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
Learning CNN on ViT: A Hybrid Model to Explicitly Class-specific Boundaries for Domain Adaptation
by: Ngo, Ba Hung, et al.
Published: (2024)
by: Ngo, Ba Hung, et al.
Published: (2024)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
by: Shah, Arya, et al.
Published: (2025)
by: Shah, Arya, et al.
Published: (2025)
Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation
by: Tang, Fenghe, et al.
Published: (2025)
by: Tang, Fenghe, et al.
Published: (2025)
Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
by: Kang, Ben, et al.
Published: (2025)
by: Kang, Ben, et al.
Published: (2025)
HydraViT: Stacking Heads for a Scalable ViT
by: Haberer, Janek, et al.
Published: (2024)
by: Haberer, Janek, et al.
Published: (2024)
Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
by: Yunusa, Haruna, et al.
Published: (2024)
by: Yunusa, Haruna, et al.
Published: (2024)
MSCViT: A Small-size ViT architecture with Multi-Scale Self-Attention Mechanism for Tiny Datasets
by: Zhang, Bowei, et al.
Published: (2025)
by: Zhang, Bowei, et al.
Published: (2025)
Similar Items
-
Subnet-Aware Dynamic Supernet Training for Neural Architecture Search
by: Jeon, Jeimin, et al.
Published: (2025) -
TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts
by: Jeon, Jeimin, et al.
Published: (2026) -
Efficient Few-Shot Neural Architecture Search by Counting the Number of Nonlinear Functions
by: Oh, Youngmin, et al.
Published: (2024) -
FYI: Flip Your Images for Dataset Distillation
by: Son, Byunggwan, et al.
Published: (2024) -
Scheduling Weight Transitions for Quantization-Aware Training
by: Lee, Junghyup, et al.
Published: (2024)