IPT-V2: Efficient Image Processing Transformer using Hierarchical Attentions
Fuente:
arXiv
Saved in:
| Main Authors: | Tu, Zhijun, Du, Kunpeng, Chen, Hanting, Wang, Hailing, Li, Wei, Hu, Jie, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instruct-IPT: All-in-One Image Processing Transformer via Weight Modulation
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
LIPT: Latency-aware Image Processing Transformer
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
Autoregressive Image Generation with Vision Full-view Prompt
by: Cai, Miaomiao, et al.
Published: (2025)
by: Cai, Miaomiao, et al.
Published: (2025)
ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices
by: Du, Kunpeng, et al.
Published: (2026)
by: Du, Kunpeng, et al.
Published: (2026)
DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution
by: Cai, Miaomiao, et al.
Published: (2025)
by: Cai, Miaomiao, et al.
Published: (2025)
Distilling Semantic Priors from SAM to Efficient Image Restoration Models
by: Zhang, Quan, et al.
Published: (2024)
by: Zhang, Quan, et al.
Published: (2024)
EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution
by: Xie, Haizhen, et al.
Published: (2025)
by: Xie, Haizhen, et al.
Published: (2025)
Data Upcycling Knowledge Distillation for Image Super-Resolution
by: Zhang, Yun, et al.
Published: (2023)
by: Zhang, Yun, et al.
Published: (2023)
From Sequential to Spatial: Reordering Autoregression for Efficient Visual Generation
by: Wang, Siyang, et al.
Published: (2025)
by: Wang, Siyang, et al.
Published: (2025)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
by: Guo, Jialong, et al.
Published: (2024)
by: Guo, Jialong, et al.
Published: (2024)
Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
by: Zou, Zhentao, et al.
Published: (2025)
by: Zou, Zhentao, et al.
Published: (2025)
A Survey on Transformer Compression
by: Tang, Yehui, et al.
Published: (2024)
by: Tang, Yehui, et al.
Published: (2024)
Effective Diffusion Transformer Architecture for Image Super-Resolution
by: Cheng, Kun, et al.
Published: (2024)
by: Cheng, Kun, et al.
Published: (2024)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
by: Guo, Jianyuan, et al.
Published: (2024)
by: Guo, Jianyuan, et al.
Published: (2024)
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation
by: Wu, Xue, et al.
Published: (2025)
by: Wu, Xue, et al.
Published: (2025)
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
by: Hu, Youbing, et al.
Published: (2024)
by: Hu, Youbing, et al.
Published: (2024)
Knowledge Distillation with Multi-granularity Mixture of Priors for Image Super-Resolution
by: Li, Simiao, et al.
Published: (2024)
by: Li, Simiao, et al.
Published: (2024)
One Step Diffusion-based Super-Resolution with Time-Aware Distillation
by: He, Xiao, et al.
Published: (2024)
by: He, Xiao, et al.
Published: (2024)
Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
by: Liu, Jizhihui, et al.
Published: (2025)
by: Liu, Jizhihui, et al.
Published: (2025)
Surgformer: Surgical Transformer with Hierarchical Temporal Attention for Surgical Phase Recognition
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
Vision Transformers with Hierarchical Attention
by: Liu, Yun, et al.
Published: (2021)
by: Liu, Yun, et al.
Published: (2021)
Improved Dense Nested Attention Network Based on Transformer for Infrared Small Target Detection
by: Bao, Chun, et al.
Published: (2023)
by: Bao, Chun, et al.
Published: (2023)
GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization
by: Chen, Yirui, et al.
Published: (2024)
by: Chen, Yirui, et al.
Published: (2024)
Efficient Adaptation of Large Vision Transformer via Adapter Re-Composing
by: Dong, Wei, et al.
Published: (2023)
by: Dong, Wei, et al.
Published: (2023)
TexLiverNet: Leveraging Medical Knowledge and Spatial-Frequency Perception for Enhanced Liver Tumor Segmentation
by: Jiang, Xiaoyan, et al.
Published: (2024)
by: Jiang, Xiaoyan, et al.
Published: (2024)
Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
S^2-Transformer for Mask-Aware Hyperspectral Image Reconstruction
by: Wang, Jiamian, et al.
Published: (2022)
by: Wang, Jiamian, et al.
Published: (2022)
Hi-Mamba: Hierarchical Mamba for Efficient Image Super-Resolution
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
Outlier-Aware Post-Training Quantization for Image Super-Resolution
by: Wang, Hailing, et al.
Published: (2025)
by: Wang, Hailing, et al.
Published: (2025)
Hierarchical Separable Video Transformer for Snapshot Compressive Imaging
by: Wang, Ping, et al.
Published: (2024)
by: Wang, Ping, et al.
Published: (2024)
Alignment-Free RGB-T Salient Object Detection: A Large-scale Dataset and Progressive Correlation Network
by: Wang, Kunpeng, et al.
Published: (2024)
by: Wang, Kunpeng, et al.
Published: (2024)
HAT: Hybrid Attention Transformer for Image Restoration
by: Chen, Xiangyu, et al.
Published: (2023)
by: Chen, Xiangyu, et al.
Published: (2023)
Mixture of Scale Experts for Alignment-free RGBT Video Object Detection and A Unified Benchmark
by: Wang, Qishun, et al.
Published: (2024)
by: Wang, Qishun, et al.
Published: (2024)
HiT-SR: Hierarchical Transformer for Efficient Image Super-Resolution
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
Neural Discrimination-Prompted Transformers for Efficient UHD Image Restoration and Enhancement
by: Wang, Cong, et al.
Published: (2026)
by: Wang, Cong, et al.
Published: (2026)
FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
by: Wang, Huihan, et al.
Published: (2025)
by: Wang, Huihan, et al.
Published: (2025)
Omni-Dimensional Frequency Learner for General Time Series Analysis
by: Chen, Xianing, et al.
Published: (2024)
by: Chen, Xianing, et al.
Published: (2024)
Similar Items
-
Instruct-IPT: All-in-One Image Processing Transformer via Weight Modulation
by: Tian, Yuchuan, et al.
Published: (2024) -
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
by: Tian, Yuchuan, et al.
Published: (2024) -
LIPT: Latency-aware Image Processing Transformer
by: Qiao, Junbo, et al.
Published: (2024) -
Autoregressive Image Generation with Vision Full-view Prompt
by: Cai, Miaomiao, et al.
Published: (2025) -
ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices
by: Du, Kunpeng, et al.
Published: (2026)