Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Phat, Geng, Xue, Xu, Kaixin, Zhe, Wang, Yang, Xulei, Cheung, Ngai-Man |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI
by: Nguyen, Phat, et al.
Published: (2025)
by: Nguyen, Phat, et al.
Published: (2025)
On the Adversarial Robustness of 3D Large Vision-Language Models
by: Liu, Chao, et al.
Published: (2026)
by: Liu, Chao, et al.
Published: (2026)
LPViT: Low-Power Semi-structured Pruning for Vision Transformers
by: Xu, Kaixin, et al.
Published: (2024)
by: Xu, Kaixin, et al.
Published: (2024)
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024)
by: Wijaya, Robert, et al.
Published: (2024)
CERSA: Cumulative Energy-Retaining Subspace Adaptation for Memory-Efficient Fine-Tuning
by: Ge, Jingze, et al.
Published: (2026)
by: Ge, Jingze, et al.
Published: (2026)
Text to Image Generation and Editing: A Survey
by: Yang, Pengfei, et al.
Published: (2025)
by: Yang, Pengfei, et al.
Published: (2025)
DM3D: Distortion-Minimized Weight Pruning for Lossless 3D Object Detection
by: Xu, Kaixin, et al.
Published: (2024)
by: Xu, Kaixin, et al.
Published: (2024)
Frequency Masking for Universal Deepfake Detection
by: Doloriel, Chandler Timm, et al.
Published: (2024)
by: Doloriel, Chandler Timm, et al.
Published: (2024)
An Efficient 3D Convolutional Neural Network with Channel-wise, Spatial-grouped, and Temporal Convolutions
by: Wang, Zhe, et al.
Published: (2025)
by: Wang, Zhe, et al.
Published: (2025)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
by: Liu, Guimeng, et al.
Published: (2025)
by: Liu, Guimeng, et al.
Published: (2025)
Urban Air Temperature Prediction using Conditional Diffusion Models
by: Dai, Siyang, et al.
Published: (2024)
by: Dai, Siyang, et al.
Published: (2024)
Vision Transformer Neural Architecture Search for Out-of-Distribution Generalization: Benchmark and Insights
by: Ho, Sy-Tuyen, et al.
Published: (2025)
by: Ho, Sy-Tuyen, et al.
Published: (2025)
A Timely Survey on Vision Transformer for Deepfake Detection
by: Wang, Zhikan, et al.
Published: (2024)
by: Wang, Zhikan, et al.
Published: (2024)
Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
by: Nguyen, Ngoc-Bao, et al.
Published: (2025)
by: Nguyen, Ngoc-Bao, et al.
Published: (2025)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
by: Su, Yongyi, et al.
Published: (2025)
by: Su, Yongyi, et al.
Published: (2025)
SegMaFormer: A Hybrid State-Space and Transformer Model for Efficient Segmentation
by: Nguyen, Duy D., et al.
Published: (2026)
by: Nguyen, Duy D., et al.
Published: (2026)
Content-Aware Radiance Fields: Aligning Model Complexity with Scene Intricacy Through Learned Bitwidth Quantization
by: Liu, Weihang, et al.
Published: (2024)
by: Liu, Weihang, et al.
Published: (2024)
Model Inversion Robustness: Can Transfer Learning Help?
by: Ho, Sy-Tuyen, et al.
Published: (2024)
by: Ho, Sy-Tuyen, et al.
Published: (2024)
Residual Attention Single-Head Vision Transformer Network for Rolling Bearing Fault Diagnosis in Noisy Environments
by: Lai, Songjiang, et al.
Published: (2024)
by: Lai, Songjiang, et al.
Published: (2024)
Low-Bitwidth Floating Point Quantization for Efficient High-Quality Diffusion Models
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
by: Li, Zhaoxu, et al.
Published: (2025)
by: Li, Zhaoxu, et al.
Published: (2025)
VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
Vision Transformer with Super Token Sampling
by: Huang, Huaibo, et al.
Published: (2022)
by: Huang, Huaibo, et al.
Published: (2022)
Neighbor-Aware Token Reduction via Hilbert Curve for Vision Transformers
by: Li, Yunge, et al.
Published: (2025)
by: Li, Yunge, et al.
Published: (2025)
Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models
by: Luan, Bozhi, et al.
Published: (2025)
by: Luan, Bozhi, et al.
Published: (2025)
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
by: Shams, Montasir, et al.
Published: (2025)
by: Shams, Montasir, et al.
Published: (2025)
Towards Joint Quantization and Token Pruning of Vision-Language Models
by: Li, Xinqing, et al.
Published: (2026)
by: Li, Xinqing, et al.
Published: (2026)
Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Advancing SEM Based Nano-Scale Defect Analysis in Semiconductor Manufacturing for Advanced IC Nodes
by: Dey, Bappaditya, et al.
Published: (2024)
by: Dey, Bappaditya, et al.
Published: (2024)
Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
by: Tang, Yutao, et al.
Published: (2025)
by: Tang, Yutao, et al.
Published: (2025)
On the Vulnerability of Skip Connections to Model Inversion Attacks
by: Koh, Jun Hao, et al.
Published: (2024)
by: Koh, Jun Hao, et al.
Published: (2024)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
by: Zeng, Fanhu, et al.
Published: (2025)
by: Zeng, Fanhu, et al.
Published: (2025)
Learning Content-Aware Multi-Modal Joint Input Pruning via Bird's-Eye-View Representation
by: Li, Yuxin, et al.
Published: (2024)
by: Li, Yuxin, et al.
Published: (2024)
Wavelet-Based Image Tokenizer for Vision Transformers
by: Zhu, Zhenhai, et al.
Published: (2024)
by: Zhu, Zhenhai, et al.
Published: (2024)
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
by: Mao, Junzhu, et al.
Published: (2025)
by: Mao, Junzhu, et al.
Published: (2025)
Spectral Vision Transformer for Efficient Tokenization with Limited Data
by: Roberts, Alexandra G., et al.
Published: (2026)
by: Roberts, Alexandra G., et al.
Published: (2026)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Similar Items
-
Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI
by: Nguyen, Phat, et al.
Published: (2025) -
On the Adversarial Robustness of 3D Large Vision-Language Models
by: Liu, Chao, et al.
Published: (2026) -
LPViT: Low-Power Semi-structured Pruning for Vision Transformers
by: Xu, Kaixin, et al.
Published: (2024) -
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024) -
CERSA: Cumulative Energy-Retaining Subspace Adaptation for Memory-Efficient Fine-Tuning
by: Ge, Jingze, et al.
Published: (2026)