CommGPT: A Graph and Retrieval-Augmented Multimodal Communication Foundation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Feibo, Zhu, Wanyun, Dong, Li, Wang, Kezhi, Yang, Kun, Pan, Cunhua, Dobre, Octavia A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917936543825920
author Jiang, Feibo
Zhu, Wanyun
Dong, Li
Wang, Kezhi
Yang, Kun
Pan, Cunhua
Dobre, Octavia A.
author_facet Jiang, Feibo
Zhu, Wanyun
Dong, Li
Wang, Kezhi
Yang, Kun
Pan, Cunhua
Dobre, Octavia A.
contents Large Language Models (LLMs) possess human-level cognitive and decision-making capabilities, making them a key technology for 6G. However, applying LLMs to the communication domain faces three major challenges: 1) Inadequate communication data; 2) Restricted input modalities; and 3) Difficulty in knowledge retrieval. To overcome these issues, we propose CommGPT, a multimodal foundation model designed specifically for communications. First, we create high-quality pretraining and fine-tuning datasets tailored in communication, enabling the LLM to engage in further pretraining and fine-tuning with communication concepts and knowledge. Then, we design a multimodal encoder to understand and process information from various input modalities. Next, we construct a Graph and Retrieval-Augmented Generation (GRG) framework, efficiently coupling Knowledge Graph (KG) with Retrieval-Augmented Generation (RAG) for multi-scale learning. Finally, we demonstrate the feasibility and effectiveness of the CommGPT through experimental validation.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CommGPT: A Graph and Retrieval-Augmented Multimodal Communication Foundation Model
Jiang, Feibo
Zhu, Wanyun
Dong, Li
Wang, Kezhi
Yang, Kun
Pan, Cunhua
Dobre, Octavia A.
Information Theory
Large Language Models (LLMs) possess human-level cognitive and decision-making capabilities, making them a key technology for 6G. However, applying LLMs to the communication domain faces three major challenges: 1) Inadequate communication data; 2) Restricted input modalities; and 3) Difficulty in knowledge retrieval. To overcome these issues, we propose CommGPT, a multimodal foundation model designed specifically for communications. First, we create high-quality pretraining and fine-tuning datasets tailored in communication, enabling the LLM to engage in further pretraining and fine-tuning with communication concepts and knowledge. Then, we design a multimodal encoder to understand and process information from various input modalities. Next, we construct a Graph and Retrieval-Augmented Generation (GRG) framework, efficiently coupling Knowledge Graph (KG) with Retrieval-Augmented Generation (RAG) for multi-scale learning. Finally, we demonstrate the feasibility and effectiveness of the CommGPT through experimental validation.
title CommGPT: A Graph and Retrieval-Augmented Multimodal Communication Foundation Model
topic Information Theory
url https://arxiv.org/abs/2502.18763