MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Wenzhuo, Zhu, Fei, Ma, Shijie, Liu, Cheng-Lin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-scale Unified Network for Image Classification
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024)
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024)
Any Resolution Any Geometry: From Multi-View To Multi-Patch
di: Cui, Wenqing, et al.
Pubblicazione: (2026)
di: Cui, Wenqing, et al.
Pubblicazione: (2026)
Towards Non-Exemplar Semi-Supervised Class-Incremental Learning
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024)
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024)
Happy: A Debiased Learning Framework for Continual Generalized Category Discovery
di: Ma, Shijie, et al.
Pubblicazione: (2024)
di: Ma, Shijie, et al.
Pubblicazione: (2024)
ViTAR: Vision Transformer with Any Resolution
di: Fan, Qihang, et al.
Pubblicazione: (2024)
di: Fan, Qihang, et al.
Pubblicazione: (2024)
Branch-Tuning: Balancing Stability and Plasticity for Continual Self-Supervised Learning
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024)
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025)
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025)
AnySR: Realizing Image Super-Resolution as Any-Scale, Any-Resource
di: Zhan, Wengyi, et al.
Pubblicazione: (2024)
di: Zhan, Wengyi, et al.
Pubblicazione: (2024)
PILoRA: Prototype Guided Incremental LoRA for Federated Class-Incremental Learning
di: Guo, Haiyang, et al.
Pubblicazione: (2024)
di: Guo, Haiyang, et al.
Pubblicazione: (2024)
Retina Vision Transformer (RetinaViT): Introducing Scaled Patches into Vision Transformers
di: Shu, Yuyang, et al.
Pubblicazione: (2024)
di: Shu, Yuyang, et al.
Pubblicazione: (2024)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
di: Gao, Peng, et al.
Pubblicazione: (2024)
di: Gao, Peng, et al.
Pubblicazione: (2024)
Patch-Fool: Are Vision Transformers Always Robust Against Adversarial Perturbations?
di: Fu, Yonggan, et al.
Pubblicazione: (2022)
di: Fu, Yonggan, et al.
Pubblicazione: (2022)
PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution
di: Liu, Yong, et al.
Pubblicazione: (2024)
di: Liu, Yong, et al.
Pubblicazione: (2024)
Active Generalized Category Discovery
di: Ma, Shijie, et al.
Pubblicazione: (2024)
di: Ma, Shijie, et al.
Pubblicazione: (2024)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
di: Su, Yongyi, et al.
Pubblicazione: (2025)
di: Su, Yongyi, et al.
Pubblicazione: (2025)
AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer
di: Lyu, Jin, et al.
Pubblicazione: (2024)
di: Lyu, Jin, et al.
Pubblicazione: (2024)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
di: Wang, Peng, et al.
Pubblicazione: (2024)
di: Wang, Peng, et al.
Pubblicazione: (2024)
Multi-Scale Implicit Transformer with Re-parameterize for Arbitrary-Scale Super-Resolution
di: Zhu, Jinchen, et al.
Pubblicazione: (2024)
di: Zhu, Jinchen, et al.
Pubblicazione: (2024)
Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
di: Liu, Yujian, et al.
Pubblicazione: (2025)
di: Liu, Yujian, et al.
Pubblicazione: (2025)
AnyTSR: Any-Scale Thermal Super-Resolution for UAV
di: Li, Mengyuan, et al.
Pubblicazione: (2025)
di: Li, Mengyuan, et al.
Pubblicazione: (2025)
TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting
di: Liu, Taorong, et al.
Pubblicazione: (2023)
di: Liu, Taorong, et al.
Pubblicazione: (2023)
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
di: Chowdhury, Md Abtahi Majeed, et al.
Pubblicazione: (2025)
di: Chowdhury, Md Abtahi Majeed, et al.
Pubblicazione: (2025)
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
di: Zeng, Fanhu, et al.
Pubblicazione: (2024)
di: Zeng, Fanhu, et al.
Pubblicazione: (2024)
PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
di: Du, Shian, et al.
Pubblicazione: (2025)
di: Du, Shian, et al.
Pubblicazione: (2025)
Refer to Any Segmentation Mask Group With Vision-Language Prompts
di: Cao, Shengcao, et al.
Pubblicazione: (2025)
di: Cao, Shengcao, et al.
Pubblicazione: (2025)
UniViTAR: Unified Vision Transformer with Native Resolution
di: Qiao, Limeng, et al.
Pubblicazione: (2025)
di: Qiao, Limeng, et al.
Pubblicazione: (2025)
CL-VISTA: Benchmarking Continual Learning in Video Large Language Models
di: Guo, Haiyang, et al.
Pubblicazione: (2026)
di: Guo, Haiyang, et al.
Pubblicazione: (2026)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
di: Liu, Xinyang, et al.
Pubblicazione: (2023)
di: Liu, Xinyang, et al.
Pubblicazione: (2023)
Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark
di: Liu, Jinyuan, et al.
Pubblicazione: (2025)
di: Liu, Jinyuan, et al.
Pubblicazione: (2025)
FlightPatchNet: Multi-Scale Patch Network with Differential Coding for Flight Trajectory Prediction
di: Wu, Lan, et al.
Pubblicazione: (2024)
di: Wu, Lan, et al.
Pubblicazione: (2024)
AniMer+: Unified Pose and Shape Estimation Across Mammalia and Aves via Family-Aware Transformer
di: An, Liang, et al.
Pubblicazione: (2025)
di: An, Liang, et al.
Pubblicazione: (2025)
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
di: Guo, Hailong, et al.
Pubblicazione: (2025)
di: Guo, Hailong, et al.
Pubblicazione: (2025)
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
di: Zhong, Zhizhou, et al.
Pubblicazione: (2025)
di: Zhong, Zhizhou, et al.
Pubblicazione: (2025)
SAMCT: Segment Any CT Allowing Labor-Free Task-Indicator Prompts
di: Lin, Xian, et al.
Pubblicazione: (2024)
di: Lin, Xian, et al.
Pubblicazione: (2024)
MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
di: Mantes, Albert Dominguez, et al.
Pubblicazione: (2026)
di: Mantes, Albert Dominguez, et al.
Pubblicazione: (2026)
GVTNet: Graph Vision Transformer For Face Super-Resolution
di: Yang, Chao, et al.
Pubblicazione: (2025)
di: Yang, Chao, et al.
Pubblicazione: (2025)
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
di: Yang, Zhuoyi, et al.
Pubblicazione: (2024)
di: Yang, Zhuoyi, et al.
Pubblicazione: (2024)
AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
di: Astruc, Guillaume, et al.
Pubblicazione: (2024)
di: Astruc, Guillaume, et al.
Pubblicazione: (2024)
Any to Full: Prompting Depth Anything for Depth Completion in One Stage
di: Zhou, Zhiyuan, et al.
Pubblicazione: (2026)
di: Zhou, Zhiyuan, et al.
Pubblicazione: (2026)
Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency Adapter
di: Zhang, Jianhui, et al.
Pubblicazione: (2025)
di: Zhang, Jianhui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Multi-scale Unified Network for Image Classification
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024) -
Any Resolution Any Geometry: From Multi-View To Multi-Patch
di: Cui, Wenqing, et al.
Pubblicazione: (2026) -
Towards Non-Exemplar Semi-Supervised Class-Incremental Learning
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024) -
Happy: A Debiased Learning Framework for Continual Generalized Category Discovery
di: Ma, Shijie, et al.
Pubblicazione: (2024) -
ViTAR: Vision Transformer with Any Resolution
di: Fan, Qihang, et al.
Pubblicazione: (2024)