Building 6G Radio Foundation Models with Transformer Architectures

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Aboulfotouh, Ahmed, Eshaghbeigi, Ashkan, Abou-Zeid, Hatem
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909391148548096
author Aboulfotouh, Ahmed
Eshaghbeigi, Ashkan
Abou-Zeid, Hatem
author_facet Aboulfotouh, Ahmed
Eshaghbeigi, Ashkan
Abou-Zeid, Hatem
contents Foundation deep learning (DL) models are general models, designed to learn general, robust and adaptable representations of their target modality, enabling finetuning across a range of downstream tasks. These models are pretrained on large, unlabeled datasets using self-supervised learning (SSL). Foundation models have demonstrated better generalization than traditional supervised approaches, a critical requirement for wireless communications where the dynamic environment demands model adaptability. In this work, we propose and demonstrate the effectiveness of a Vision Transformer (ViT) as a radio foundation model for spectrogram learning. We introduce a Masked Spectrogram Modeling (MSM) approach to pretrain the ViT in a self-supervised fashion. We evaluate the ViT-based foundation model on two downstream tasks: Channel State Information (CSI)-based Human Activity sensing and Spectrogram Segmentation. Experimental results demonstrate competitive performance to supervised training while generalizing across diverse domains. Notably, the pretrained ViT model outperforms a four-times larger model that is trained from scratch on the spectrogram segmentation task, while requiring significantly less training time, and achieves competitive performance on the CSI-based human activity sensing task. This work demonstrates the effectiveness of ViT with MSM for pretraining as a promising technique for scalable foundation model development in future 6G networks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_09996
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Building 6G Radio Foundation Models with Transformer Architectures
Aboulfotouh, Ahmed
Eshaghbeigi, Ashkan
Abou-Zeid, Hatem
Signal Processing
Artificial Intelligence
Networking and Internet Architecture
Foundation deep learning (DL) models are general models, designed to learn general, robust and adaptable representations of their target modality, enabling finetuning across a range of downstream tasks. These models are pretrained on large, unlabeled datasets using self-supervised learning (SSL). Foundation models have demonstrated better generalization than traditional supervised approaches, a critical requirement for wireless communications where the dynamic environment demands model adaptability. In this work, we propose and demonstrate the effectiveness of a Vision Transformer (ViT) as a radio foundation model for spectrogram learning. We introduce a Masked Spectrogram Modeling (MSM) approach to pretrain the ViT in a self-supervised fashion. We evaluate the ViT-based foundation model on two downstream tasks: Channel State Information (CSI)-based Human Activity sensing and Spectrogram Segmentation. Experimental results demonstrate competitive performance to supervised training while generalizing across diverse domains. Notably, the pretrained ViT model outperforms a four-times larger model that is trained from scratch on the spectrogram segmentation task, while requiring significantly less training time, and achieves competitive performance on the CSI-based human activity sensing task. This work demonstrates the effectiveness of ViT with MSM for pretraining as a promising technique for scalable foundation model development in future 6G networks.
title Building 6G Radio Foundation Models with Transformer Architectures
topic Signal Processing
Artificial Intelligence
Networking and Internet Architecture
url https://arxiv.org/abs/2411.09996