Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions
Fuente:
arXiv
Saved in:
| Main Authors: | Xin, Zhihang, Hu, Xitong, Wang, Rui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
PTQ4ViT: Post-training quantization for vision transformers with twin uniform quantization
by: Yuan, Zhihang, et al.
Published: (2021)
by: Yuan, Zhihang, et al.
Published: (2021)
Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID
by: Li, Jiaxuan, et al.
Published: (2026)
by: Li, Jiaxuan, et al.
Published: (2026)
Beyond the final layer: Attentive multilayer fusion for vision transformers
by: Ciernik, Laure, et al.
Published: (2026)
by: Ciernik, Laure, et al.
Published: (2026)
HSFusion: A high-level vision task-driven infrared and visible image fusion network via semantic and geometric domain transformation
by: Jiang, Chengjie, et al.
Published: (2024)
by: Jiang, Chengjie, et al.
Published: (2024)
Multi-encoder nnU-Net outperforms transformer models with self-supervised pretraining
by: Otaghsara, Seyedeh Sahar Taheri, et al.
Published: (2025)
by: Otaghsara, Seyedeh Sahar Taheri, et al.
Published: (2025)
Estimation of geometric transformation matrices using grid-shaped pilot signals
by: Kawano, Rinka, et al.
Published: (2026)
by: Kawano, Rinka, et al.
Published: (2026)
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
by: Chen, Weiming, et al.
Published: (2026)
by: Chen, Weiming, et al.
Published: (2026)
Interpreting the structure of multi-object representations in vision encoders
by: Khajuria, Tarun, et al.
Published: (2024)
by: Khajuria, Tarun, et al.
Published: (2024)
Cross multiscale vision transformer for deep fake detection
by: P, Akhshan, et al.
Published: (2025)
by: P, Akhshan, et al.
Published: (2025)
Animal behavioral analysis and neural encoding with transformer-based self-supervised pretraining
by: Wang, Yanchen, et al.
Published: (2025)
by: Wang, Yanchen, et al.
Published: (2025)
What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers
by: Pawlowsky, Moritz, et al.
Published: (2026)
by: Pawlowsky, Moritz, et al.
Published: (2026)
METER: a mobile vision transformer architecture for monocular depth estimation
by: Papa, L., et al.
Published: (2024)
by: Papa, L., et al.
Published: (2024)
an interpretable vision transformer framework for automated brain tumor classification
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026)
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026)
Hi-End-MAE: Hierarchical encoder-driven masked autoencoders are stronger vision learners for medical image segmentation
by: Tang, Fenghe, et al.
Published: (2025)
by: Tang, Fenghe, et al.
Published: (2025)
Beyond one-hot encoding? Journey into compact encoding for large multi-class segmentation
by: Kujawa, Aaron, et al.
Published: (2025)
by: Kujawa, Aaron, et al.
Published: (2025)
MOSformer: Momentum encoder-based inter-slice fusion transformer for medical image segmentation
by: Huang, De-Xing, et al.
Published: (2024)
by: Huang, De-Xing, et al.
Published: (2024)
A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking
by: Papa, Lorenzo, et al.
Published: (2023)
by: Papa, Lorenzo, et al.
Published: (2023)
Self-supervised pretraining for an iterative image size agnostic vision transformer
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
Beyond sparse denoising in frames: minimax estimation with a scattering transform
by: Cuvelle--Magar, Nathanaël, et al.
Published: (2025)
by: Cuvelle--Magar, Nathanaël, et al.
Published: (2025)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
SAFER: Sharpness Aware layer-selective Finetuning for Enhanced Robustness in vision transformers
by: Gopal, Bhavna, et al.
Published: (2025)
by: Gopal, Bhavna, et al.
Published: (2025)
Low-latency vision transformers via large-scale multi-head attention
by: Gross, Ronit D., et al.
Published: (2025)
by: Gross, Ronit D., et al.
Published: (2025)
LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization
by: Xie, Rui, et al.
Published: (2024)
by: Xie, Rui, et al.
Published: (2024)
FreeZe: Training-free zero-shot 6D pose estimation with geometric and vision foundation models
by: Caraffa, Andrea, et al.
Published: (2023)
by: Caraffa, Andrea, et al.
Published: (2023)
Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving
by: Song, Zhihang, et al.
Published: (2024)
by: Song, Zhihang, et al.
Published: (2024)
High-fidelity Multi-view Normal Integration with Scale-encoded Neural Surface Representation
by: Yang, Tongyu, et al.
Published: (2026)
by: Yang, Tongyu, et al.
Published: (2026)
OPFormer: Object Pose Estimation leveraging foundation model with geometric encoding
by: Moroz, Artem, et al.
Published: (2025)
by: Moroz, Artem, et al.
Published: (2025)
Diffusion-based Synthetic Data Generation for Visible-Infrared Person Re-Identification
by: Dai, Wenbo, et al.
Published: (2025)
by: Dai, Wenbo, et al.
Published: (2025)
Towards Vision-Language Geo-Foundation Model: A Survey
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition
by: Sun, Jian, et al.
Published: (2026)
by: Sun, Jian, et al.
Published: (2026)
Beyond conventional vision: RGB-event fusion for robust object detection in dynamic traffic scenarios
by: Liu, Zhanwen, et al.
Published: (2025)
by: Liu, Zhanwen, et al.
Published: (2025)
VGDFR: Diffusion-based Video Generation with Dynamic Latent Frame Rate
by: Yuan, Zhihang, et al.
Published: (2025)
by: Yuan, Zhihang, et al.
Published: (2025)
Cross-Modal Prototype Allocation: Unsupervised Slide Representation Learning via Patch-Text Contrast in Computational Pathology
by: Chen, Yuxuan, et al.
Published: (2025)
by: Chen, Yuxuan, et al.
Published: (2025)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection
by: Xiang, An, et al.
Published: (2025)
by: Xiang, An, et al.
Published: (2025)
A training-free framework for high-fidelity appearance transfer via diffusion transformers
by: Gu, Shengrong, et al.
Published: (2026)
by: Gu, Shengrong, et al.
Published: (2026)
LaB-GATr: geometric algebra transformers for large biomedical surface and volume meshes
by: Suk, Julian, et al.
Published: (2024)
by: Suk, Julian, et al.
Published: (2024)
Similar Items
-
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026) -
PTQ4ViT: Post-training quantization for vision transformers with twin uniform quantization
by: Yuan, Zhihang, et al.
Published: (2021) -
Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID
by: Li, Jiaxuan, et al.
Published: (2026) -
Beyond the final layer: Attentive multilayer fusion for vision transformers
by: Ciernik, Laure, et al.
Published: (2026) -
HSFusion: A high-level vision task-driven infrared and visible image fusion network via semantic and geometric domain transformation
by: Jiang, Chengjie, et al.
Published: (2024)