A 2D Semantic-Aware Position Encoding for Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xi, Zhou, Shiyang, Huang, Muqi, Feng, Jiaxu, Xiong, Yun, Zhou, Kun, Yang, Biao, Zhang, Yuhui, Bao, Huishuai, Peng, Sijia, Li, Chuan, Shi, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing
by: Yang, Biao, et al.
Published: (2025)
by: Yang, Biao, et al.
Published: (2025)
Rethinking Time Encoding via Learnable Transformation Functions
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Fusion Matters: Length-Aware Analysis of Positional-Encoding Fusion in Transformers
by: Hallam, Mohamed Amine, et al.
Published: (2026)
by: Hallam, Mohamed Amine, et al.
Published: (2026)
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
Dynamic Graph Transformer with Correlated Spatial-Temporal Positional Encoding
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
Age at menopause and cognition in later life – a cross‐nation comparison study between China and India
by: Muqi Guo
Published: (2025)
by: Muqi Guo
Published: (2025)
ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing
by: Ma, Xingwei, et al.
Published: (2025)
by: Ma, Xingwei, et al.
Published: (2025)
Mamba or Transformer for Time Series Forecasting? Mixture of Universals (MoU) Is All You Need
by: Peng, Sijia, et al.
Published: (2024)
by: Peng, Sijia, et al.
Published: (2024)
Diffusion-based Radiotherapy Dose Prediction Guided by Inter-slice Aware Structure Encoding
by: Feng, Zhenghao, et al.
Published: (2023)
by: Feng, Zhenghao, et al.
Published: (2023)
Integrating the Biosynthesis and Genetic Encoding of Noncanonical Amino Acids for Enzyme Design and Catalysis
by: Yuhui Sheng, et al.
Published: (2026)
by: Yuhui Sheng, et al.
Published: (2026)
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
by: Zhao, Liang, et al.
Published: (2023)
by: Zhao, Liang, et al.
Published: (2023)
Semantic Causality-Aware Vision-Based 3D Occupancy Prediction
by: Chen, Dubing, et al.
Published: (2025)
by: Chen, Dubing, et al.
Published: (2025)
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
SAC-ViT: Semantic-Aware Clustering Vision Transformer with Early Exit
by: Hu, Youbing, et al.
Published: (2025)
by: Hu, Youbing, et al.
Published: (2025)
HGFormer: Topology-Aware Vision Transformer with HyperGraph Learning
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
by: Hou, Liang, et al.
Published: (2025)
by: Hou, Liang, et al.
Published: (2025)
On the Geometry of Positional Encodings in Transformers
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Latent Radiance Fields with 3D-aware 2D Representations
by: Zhou, Chaoyi, et al.
Published: (2025)
by: Zhou, Chaoyi, et al.
Published: (2025)
Toward Relative Positional Encoding in Spiking Transformers
by: Lv, Changze, et al.
Published: (2025)
by: Lv, Changze, et al.
Published: (2025)
Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation
by: Zhou, Feng, et al.
Published: (2025)
by: Zhou, Feng, et al.
Published: (2025)
Layered 3D Human Generation via Semantic-Aware Diffusion Model
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
by: Chen, Luoxin, et al.
Published: (2026)
by: Chen, Luoxin, et al.
Published: (2026)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
DTFormer: A Transformer-Based Method for Discrete-Time Dynamic Graph Representation Learning
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Revisiting Multimodal Positional Encoding in Vision-Language Models
by: Huang, Jie, et al.
Published: (2025)
by: Huang, Jie, et al.
Published: (2025)
Local-Global Context Aware Transformer for Language-Guided Video Segmentation
by: Liang, Chen, et al.
Published: (2022)
by: Liang, Chen, et al.
Published: (2022)
Graph Transformers without Positional Encodings
by: Garg, Ayush
Published: (2024)
by: Garg, Ayush
Published: (2024)
Learning in Position-Aware Multinomial Logit Bandits: From Multiplicative to General Position Effects
by: Chen, Xi, et al.
Published: (2026)
by: Chen, Xi, et al.
Published: (2026)
Length Generalization of Causal Transformers without Position Encoding
by: Wang, Jie, et al.
Published: (2024)
by: Wang, Jie, et al.
Published: (2024)
Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
by: Mannes, Mahmoud
Published: (2026)
by: Mannes, Mahmoud
Published: (2026)
HQViT: Hybrid Quantum Vision Transformer for Image Classification
by: Zhang, Hui, et al.
Published: (2025)
by: Zhang, Hui, et al.
Published: (2025)
OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
by: Huang, Ruoxiang, et al.
Published: (2025)
by: Huang, Ruoxiang, et al.
Published: (2025)
DyWPE: Signal-Aware Dynamic Wavelet Positional Encoding for Time Series Transformers
by: Irani, Habib, et al.
Published: (2025)
by: Irani, Habib, et al.
Published: (2025)
Breaking Semantic-Aware Watermarks via LLM-Guided Coherence-Preserving Semantic Injection
by: Gao, Zheng, et al.
Published: (2026)
by: Gao, Zheng, et al.
Published: (2026)
CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping
by: An, Zijian, et al.
Published: (2025)
by: An, Zijian, et al.
Published: (2025)
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
by: Zhou, Shengli, et al.
Published: (2026)
by: Zhou, Shengli, et al.
Published: (2026)
DAM-GT: Dual Positional Encoding-Based Attention Masking Graph Transformer for Node Classification
by: Li, Chenyang, et al.
Published: (2025)
by: Li, Chenyang, et al.
Published: (2025)
DAPE: Data-Adaptive Positional Encoding for Length Extrapolation
by: Zheng, Chuanyang, et al.
Published: (2024)
by: Zheng, Chuanyang, et al.
Published: (2024)
Positional Encoding Field
by: Bai, Yunpeng, et al.
Published: (2025)
by: Bai, Yunpeng, et al.
Published: (2025)
Targeting GPRC5D With CAR‐T Cells in Relapse/Refractory Multiple Myeloma: Case Report and Literature Review
by: Sijia Yan, et al.
Published: (2025)
by: Sijia Yan, et al.
Published: (2025)
Similar Items
-
AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing
by: Yang, Biao, et al.
Published: (2025) -
Rethinking Time Encoding via Learnable Transformation Functions
by: Chen, Xi, et al.
Published: (2025) -
Fusion Matters: Length-Aware Analysis of Positional-Encoding Fusion in Transformers
by: Hallam, Mohamed Amine, et al.
Published: (2026) -
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026) -
Dynamic Graph Transformer with Correlated Spatial-Temporal Positional Encoding
by: Wang, Zhe, et al.
Published: (2024)