Weierstrass Positional Encoding for Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Xin, Zhihang, Wang, Rui, Hu, Xitong, Wu, Xiaojun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions
by: Xin, Zhihang, et al.
Published: (2025)
by: Xin, Zhihang, et al.
Published: (2025)
A 2D Semantic-Aware Position Encoding for Vision Transformers
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
by: Vani, Ankit, et al.
Published: (2024)
by: Vani, Ankit, et al.
Published: (2024)
Cameras as Relative Positional Encoding
by: Li, Ruilong, et al.
Published: (2025)
by: Li, Ruilong, et al.
Published: (2025)
PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
by: Durairaju, Raahul Krishna, et al.
Published: (2025)
by: Durairaju, Raahul Krishna, et al.
Published: (2025)
SeqPE: Transformer with Sequential Position Encoding
by: Li, Huayang, et al.
Published: (2025)
by: Li, Huayang, et al.
Published: (2025)
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
by: Chowdhury, Md Abtahi Majeed, et al.
Published: (2025)
by: Chowdhury, Md Abtahi Majeed, et al.
Published: (2025)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
by: Chen, Zhanpeng, et al.
Published: (2025)
by: Chen, Zhanpeng, et al.
Published: (2025)
Do Pre-trained Vision-Language Models Encode Object States?
by: Newman, Kaleb, et al.
Published: (2024)
by: Newman, Kaleb, et al.
Published: (2024)
Amortized-Precision Quantization for Early-Exit Vision Transformers
by: Fang, Rui, et al.
Published: (2026)
by: Fang, Rui, et al.
Published: (2026)
Efficient Point Cloud Processing with High-Dimensional Positional Encoding and Non-Local MLPs
by: Zou, Yanmei, et al.
Published: (2026)
by: Zou, Yanmei, et al.
Published: (2026)
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
by: Chen, Pingyi, et al.
Published: (2025)
by: Chen, Pingyi, et al.
Published: (2025)
LTMSformer: A Local Trend-Aware Attention and Motion State Encoding Transformer for Multi-Agent Trajectory Prediction
by: Yan, Yixin, et al.
Published: (2025)
by: Yan, Yixin, et al.
Published: (2025)
Mitigating Coordinate Prediction Bias from Positional Encoding Failures
by: Tao, Xingjian, et al.
Published: (2025)
by: Tao, Xingjian, et al.
Published: (2025)
Tackling the Abstraction and Reasoning Corpus with Vision Transformers: the Importance of 2D Representation, Positions, and Objects
by: Li, Wenhao, et al.
Published: (2024)
by: Li, Wenhao, et al.
Published: (2024)
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
by: Dong, Linfeng, et al.
Published: (2025)
by: Dong, Linfeng, et al.
Published: (2025)
Vision Bridge Transformer at Scale
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
Equipping Sketch Patches with Context-Aware Positional Encoding for Graphic Sketch Representation
by: Zang, Sicong, et al.
Published: (2024)
by: Zang, Sicong, et al.
Published: (2024)
MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval
by: Zhu, Fengbin, et al.
Published: (2026)
by: Zhu, Fengbin, et al.
Published: (2026)
HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding
by: Shi, Yanzhao, et al.
Published: (2025)
by: Shi, Yanzhao, et al.
Published: (2025)
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
by: Liu, Hanpeng, et al.
Published: (2026)
by: Liu, Hanpeng, et al.
Published: (2026)
Vanishing Depth: A Depth Adapter with Positional Depth Encoding for Generalized Image Encoders
by: Koch, Paul, et al.
Published: (2025)
by: Koch, Paul, et al.
Published: (2025)
Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models
by: Yang, Siyuan, et al.
Published: (2026)
by: Yang, Siyuan, et al.
Published: (2026)
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
by: Thapa, Rahul, et al.
Published: (2024)
by: Thapa, Rahul, et al.
Published: (2024)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
by: Meng, Weikang, et al.
Published: (2025)
by: Meng, Weikang, et al.
Published: (2025)
Language-assisted Vision Model Debugger: A Sample-Free Approach to Finding and Fixing Bugs
by: Jiang, Chaoquan, et al.
Published: (2023)
by: Jiang, Chaoquan, et al.
Published: (2023)
Shared Neural Space: Unified Precomputed Feature Encoding for Multi-Task and Cross Domain Vision
by: Li, Jing, et al.
Published: (2025)
by: Li, Jing, et al.
Published: (2025)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
Uncertainty-Encoded Multi-Modal Fusion for Robust Object Detection in Autonomous Driving
by: Lou, Yang, et al.
Published: (2023)
by: Lou, Yang, et al.
Published: (2023)
DEFormer: DCT-driven Enhancement Transformer for Low-light Image and Dark Vision
by: Yin, Xiangchen, et al.
Published: (2023)
by: Yin, Xiangchen, et al.
Published: (2023)
COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors
by: Yuan, Kaishen, et al.
Published: (2024)
by: Yuan, Kaishen, et al.
Published: (2024)
SAC-ViT: Semantic-Aware Clustering Vision Transformer with Early Exit
by: Hu, Youbing, et al.
Published: (2025)
by: Hu, Youbing, et al.
Published: (2025)
GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
by: Wang, Chun, et al.
Published: (2025)
by: Wang, Chun, et al.
Published: (2025)
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
VORTEX: Challenging CNNs at Texture Recognition by using Vision Transformers with Orderless and Randomized Token Encodings
by: Scabini, Leonardo, et al.
Published: (2025)
by: Scabini, Leonardo, et al.
Published: (2025)
Orthogonal Quadratic Complements for Vision Transformer Feed-Forward Networks
by: Zixian, Wang
Published: (2026)
by: Zixian, Wang
Published: (2026)
Encoding and Controlling Global Semantics for Long-form Video Question Answering
by: Nguyen, Thong Thanh, et al.
Published: (2024)
by: Nguyen, Thong Thanh, et al.
Published: (2024)
Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation
by: Dong, Wei, et al.
Published: (2024)
by: Dong, Wei, et al.
Published: (2024)
DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers
by: Zhang, Hanling, et al.
Published: (2025)
by: Zhang, Hanling, et al.
Published: (2025)
Similar Items
-
Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions
by: Xin, Zhihang, et al.
Published: (2025) -
A 2D Semantic-Aware Position Encoding for Vision Transformers
by: Chen, Xi, et al.
Published: (2025) -
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
by: Vani, Ankit, et al.
Published: (2024) -
Cameras as Relative Positional Encoding
by: Li, Ruilong, et al.
Published: (2025) -
PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
by: Durairaju, Raahul Krishna, et al.
Published: (2025)