UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
Fuente:
arXiv
Saved in:
| Main Authors: | Walmer, Matthew, Suri, Saksham, Aggarwal, Anirud, Shrivastava, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024)
by: Suri, Saksham, et al.
Published: (2024)
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025)
by: Aggarwal, Anirud, et al.
Published: (2025)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026)
by: Agarwal, Vatsal, et al.
Published: (2026)
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
by: Wang, Hanyu, et al.
Published: (2024)
by: Wang, Hanyu, et al.
Published: (2024)
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025)
by: Kumar, Pulkit, et al.
Published: (2025)
Multi-entity Video Transformers for Fine-Grained Video Representation Learning
by: Walmer, Matthew, et al.
Published: (2023)
by: Walmer, Matthew, et al.
Published: (2023)
UVIS: Unsupervised Video Instance Segmentation
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions
by: Padmanabhan, Namitha, et al.
Published: (2024)
by: Padmanabhan, Namitha, et al.
Published: (2024)
Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
by: Seo, Minseok, et al.
Published: (2025)
by: Seo, Minseok, et al.
Published: (2025)
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026)
by: Padmanabhan, Namitha, et al.
Published: (2026)
Weighted Reverse Convolution for Feature Upsampling
by: Li, Wentong, et al.
Published: (2026)
by: Li, Wentong, et al.
Published: (2026)
EAGLES: Efficient Accelerated 3D Gaussians with Lightweight EncodingS
by: Girish, Sharath, et al.
Published: (2023)
by: Girish, Sharath, et al.
Published: (2023)
Edge Detection for Organ Boundaries via Top Down Refinement and SubPixel Upsampling
by: Mehta, Aarav, et al.
Published: (2025)
by: Mehta, Aarav, et al.
Published: (2025)
DyGLNet: Hybrid Global-Local Feature Fusion with Dynamic Upsampling for Medical Image Segmentation
by: Zhao, Yican, et al.
Published: (2025)
by: Zhao, Yican, et al.
Published: (2025)
Utilization of Neighbor Information for Image Classification with Different Levels of Supervision
by: Jayatilaka, Gihan, et al.
Published: (2025)
by: Jayatilaka, Gihan, et al.
Published: (2025)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
AnyUp: Universal Feature Upsampling
by: Wimmer, Thomas, et al.
Published: (2025)
by: Wimmer, Thomas, et al.
Published: (2025)
Efficient and High-Fidelity Omni Modality Retrieval
by: Huynh, Chuong, et al.
Published: (2026)
by: Huynh, Chuong, et al.
Published: (2026)
V-VIPE: Variational View Invariant Pose Embedding
by: Levy, Mara, et al.
Published: (2024)
by: Levy, Mara, et al.
Published: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Mitigating Hallucinations in Diffusion Models through Adaptive Attention Modulation
by: Oorloff, Trevine, et al.
Published: (2025)
by: Oorloff, Trevine, et al.
Published: (2025)
Local Attention Transformers for High-Detail Optical Flow Upsampling
by: Gielisse, Alexander, et al.
Published: (2024)
by: Gielisse, Alexander, et al.
Published: (2024)
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
by: Agarwal, Vatsal, et al.
Published: (2025)
by: Agarwal, Vatsal, et al.
Published: (2025)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
by: Chambon, Loick, et al.
Published: (2025)
by: Chambon, Loick, et al.
Published: (2025)
Towards Understanding Best Practices for Quantization of Vision-Language Models
by: Das, Gautom, et al.
Published: (2026)
by: Das, Gautom, et al.
Published: (2026)
Accelerate High-Quality Diffusion Models with Inner Loop Feedback
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
How to Design and Train Your Implicit Neural Representation for Video Compression
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
DiveUp: Learning Feature Upsampling from Diverse Vision Foundation Models
by: Liu, Xiaoqiong, et al.
Published: (2026)
by: Liu, Xiaoqiong, et al.
Published: (2026)
Improving Feature Stability during Upsampling -- Spectral Artifacts and the Importance of Spatial Context
by: Agnihotri, Shashank, et al.
Published: (2023)
by: Agnihotri, Shashank, et al.
Published: (2023)
Dense Feature Interaction Network for Image Inpainting Localization
by: Yao, Ye, et al.
Published: (2024)
by: Yao, Ye, et al.
Published: (2024)
Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics
by: Maiya, Shishira R, et al.
Published: (2024)
by: Maiya, Shishira R, et al.
Published: (2024)
PLATYPUS: Progressive Local Surface Estimator for Arbitrary-Scale Point Cloud Upsampling
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
From Pixels to Prose: A Large Dataset of Dense Image Captions
by: Singla, Vasu, et al.
Published: (2024)
by: Singla, Vasu, et al.
Published: (2024)
FlowFeat: Pixel-Dense Embedding of Motion Profiles
by: Araslanov, Nikita, et al.
Published: (2025)
by: Araslanov, Nikita, et al.
Published: (2025)
VMatcher: State-Space Semi-Dense Local Feature Matching
by: Youssef, Ali
Published: (2025)
by: Youssef, Ali
Published: (2025)
Efficient Universal Perception Encoder
by: Zhu, Chenchen, et al.
Published: (2026)
by: Zhu, Chenchen, et al.
Published: (2026)
Unified Framework for Open-World Compositional Zero-shot Learning
by: Jayasekara, Hirunima, et al.
Published: (2024)
by: Jayasekara, Hirunima, et al.
Published: (2024)
Similar Items
-
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024) -
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025) -
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026) -
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
by: Wang, Hanyu, et al.
Published: (2024) -
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025)