Vision Transformer Based User Equipment Positioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Parshwa, Patel, Dhaval K., Soni, Brijesh, López-Benítez, Miguel, Govindasamy, Siddhartan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912702304092160
author Shah, Parshwa
Patel, Dhaval K.
Soni, Brijesh
López-Benítez, Miguel
Govindasamy, Siddhartan
author_facet Shah, Parshwa
Patel, Dhaval K.
Soni, Brijesh
López-Benítez, Miguel
Govindasamy, Siddhartan
contents Recently, Deep Learning (DL) techniques have been used for User Equipment (UE) positioning. However, the key shortcomings of such models is that: i) they weigh the same attention to the entire input; ii) they are not well suited for the non-sequential data e.g., when only instantaneous Channel State Information (CSI) is available. In this context, we propose an attention-based Vision Transformer (ViT) architecture that focuses on the Angle Delay Profile (ADP) from CSI matrix. Our approach, validated on the `DeepMIMO' and `ViWi' ray-tracing datasets, achieves an Root Mean Squared Error (RMSE) of 0.55m indoors, 13.59m outdoors in DeepMIMO, and 3.45m in ViWi's outdoor blockage scenario. The proposed scheme outperforms state-of-the-art schemes by $\sim$ 38\%. It also performs substantially better than other approaches that we have considered in terms of the distribution of error distance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision Transformer Based User Equipment Positioning
Shah, Parshwa
Patel, Dhaval K.
Soni, Brijesh
López-Benítez, Miguel
Govindasamy, Siddhartan
Computer Vision and Pattern Recognition
Networking and Internet Architecture
Recently, Deep Learning (DL) techniques have been used for User Equipment (UE) positioning. However, the key shortcomings of such models is that: i) they weigh the same attention to the entire input; ii) they are not well suited for the non-sequential data e.g., when only instantaneous Channel State Information (CSI) is available. In this context, we propose an attention-based Vision Transformer (ViT) architecture that focuses on the Angle Delay Profile (ADP) from CSI matrix. Our approach, validated on the `DeepMIMO' and `ViWi' ray-tracing datasets, achieves an Root Mean Squared Error (RMSE) of 0.55m indoors, 13.59m outdoors in DeepMIMO, and 3.45m in ViWi's outdoor blockage scenario. The proposed scheme outperforms state-of-the-art schemes by $\sim$ 38\%. It also performs substantially better than other approaches that we have considered in terms of the distribution of error distance.
title Vision Transformer Based User Equipment Positioning
topic Computer Vision and Pattern Recognition
Networking and Internet Architecture
url https://arxiv.org/abs/2511.08549