Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Yunjie, Xie, Lingxi, Qiu, Jihao, Jiao, Jianbin, Wang, Yaowei, Tian, Qi, Ye, Qixiang |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Keyframe Sampling for Long Video Understanding
by: Tang, Xi, et al.
Published: (2025)
by: Tang, Xi, et al.
Published: (2025)
ChatterBox: Multi-round Multimodal Referring and Grounding
by: Tian, Yunjie, et al.
Published: (2024)
by: Tian, Yunjie, et al.
Published: (2024)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
by: Qiu, Jihao, et al.
Published: (2026)
by: Qiu, Jihao, et al.
Published: (2026)
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
by: Ma, Tianren, et al.
Published: (2024)
by: Ma, Tianren, et al.
Published: (2024)
VMamba: Visual State Space Model
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Artemis: Towards Referential Understanding in Complex Videos
by: Qiu, Jihao, et al.
Published: (2024)
by: Qiu, Jihao, et al.
Published: (2024)
Spatial Transform Decoupling for Oriented Object Detection
by: Yu, Hongtian, et al.
Published: (2023)
by: Yu, Hongtian, et al.
Published: (2023)
T-SiamTPN: Temporal Siamese Transformer Pyramid Networks for Robust and Efficient UAV Tracking
by: Ardi, Hojat, et al.
Published: (2025)
by: Ardi, Hojat, et al.
Published: (2025)
YOLOv12: Attention-Centric Real-Time Object Detectors
by: Tian, Yunjie, et al.
Published: (2025)
by: Tian, Yunjie, et al.
Published: (2025)
Building Vision Models upon Heat Conduction
by: Wang, Zhaozhi, et al.
Published: (2024)
by: Wang, Zhaozhi, et al.
Published: (2024)
Self-supervised Feature-Gate Coupling for Dynamic Network Pruning
by: Shi, Mengnan, et al.
Published: (2021)
by: Shi, Mengnan, et al.
Published: (2021)
Uncertainty-guided Optimal Transport in Depth Supervised Sparse-View 3D Gaussian
by: Sun, Wei, et al.
Published: (2024)
by: Sun, Wei, et al.
Published: (2024)
Depth-guided Texture Diffusion for Image Semantic Segmentation
by: Sun, Wei, et al.
Published: (2024)
by: Sun, Wei, et al.
Published: (2024)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
Color as the Impetus: Transforming Few-Shot Learner
by: Qi, Chaofei, et al.
Published: (2025)
by: Qi, Chaofei, et al.
Published: (2025)
GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions
by: Wang, Junjie, et al.
Published: (2023)
by: Wang, Junjie, et al.
Published: (2023)
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Virtual Classification: Modulating Domain-Specific Knowledge for Multidomain Crowd Counting
by: Guo, Mingyue, et al.
Published: (2024)
by: Guo, Mingyue, et al.
Published: (2024)
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
by: Wang, Zhaozhi, et al.
Published: (2025)
by: Wang, Zhaozhi, et al.
Published: (2025)
Study on Aspect Ratio Variability toward Robustness of Vision Transformer-based Vehicle Re-identification
by: Qiu, Mei, et al.
Published: (2024)
by: Qiu, Mei, et al.
Published: (2024)
Correspondence-Guided SfM-Free 3D Gaussian Splatting for NVS
by: Sun, Wei, et al.
Published: (2024)
by: Sun, Wei, et al.
Published: (2024)
Pyramidal Adaptive Cross-Gating for Multimodal Detection
by: Gu, Zidong, et al.
Published: (2025)
by: Gu, Zidong, et al.
Published: (2025)
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Parameter Efficient Fine-tuning via Cross Block Orchestration for Segment Anything Model
by: Peng, Zelin, et al.
Published: (2023)
by: Peng, Zelin, et al.
Published: (2023)
P2Object: Single Point Supervised Object Detection and Instance Segmentation
by: Chen, Pengfei, et al.
Published: (2025)
by: Chen, Pengfei, et al.
Published: (2025)
AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation
by: Ge, Jiannan, et al.
Published: (2024)
by: Ge, Jiannan, et al.
Published: (2024)
ClickTrack: Towards Real-time Interactive Single Object Tracking
by: Wang, Kuiran, et al.
Published: (2024)
by: Wang, Kuiran, et al.
Published: (2024)
CPR++: Object Localization via Single Coarse Point Supervision
by: Yu, Xuehui, et al.
Published: (2024)
by: Yu, Xuehui, et al.
Published: (2024)
Rethinking Sampling Strategies for Unsupervised Person Re-identification
by: Han, Xumeng, et al.
Published: (2021)
by: Han, Xumeng, et al.
Published: (2021)
GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models
by: Yi, Taoran, et al.
Published: (2023)
by: Yi, Taoran, et al.
Published: (2023)
Regressor-Segmenter Mutual Prompt Learning for Crowd Counting
by: Guo, Mingyue, et al.
Published: (2023)
by: Guo, Mingyue, et al.
Published: (2023)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
by: Baraldi, Lorenzo, et al.
Published: (2023)
by: Baraldi, Lorenzo, et al.
Published: (2023)
A General and Efficient Training for Transformer via Token Expansion
by: Huang, Wenxuan, et al.
Published: (2024)
by: Huang, Wenxuan, et al.
Published: (2024)
SAMConvex: Fast Discrete Optimization for CT Registration using Self-supervised Anatomical Embedding and Correlation Pyramid
by: Li, Zi, et al.
Published: (2023)
by: Li, Zi, et al.
Published: (2023)
Segment Any 3D Gaussians
by: Cen, Jiazhong, et al.
Published: (2023)
by: Cen, Jiazhong, et al.
Published: (2023)
MetaLab: Few-Shot Game Changer for Image Recognition
by: Qi, Chaofei, et al.
Published: (2025)
by: Qi, Chaofei, et al.
Published: (2025)
ZoomNeXt: A Unified Collaborative Pyramid Network for Camouflaged Object Detection
by: Pang, Youwei, et al.
Published: (2023)
by: Pang, Youwei, et al.
Published: (2023)
Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields in Efficient CNNs for Fair Medical Image Classification
by: Wu, Xiao, et al.
Published: (2025)
by: Wu, Xiao, et al.
Published: (2025)
Shallow Deep Learning Can Still Excel in Fine-Grained Few-Shot Learning
by: Qi, Chaofei, et al.
Published: (2025)
by: Qi, Chaofei, et al.
Published: (2025)
Face Pyramid Vision Transformer
by: Islam, Khawar, et al.
Published: (2022)
by: Islam, Khawar, et al.
Published: (2022)
Similar Items
-
Adaptive Keyframe Sampling for Long Video Understanding
by: Tang, Xi, et al.
Published: (2025) -
ChatterBox: Multi-round Multimodal Referring and Grounding
by: Tian, Yunjie, et al.
Published: (2024) -
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
by: Qiu, Jihao, et al.
Published: (2026) -
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
by: Ma, Tianren, et al.
Published: (2024) -
VMamba: Visual State Space Model
by: Liu, Yue, et al.
Published: (2024)