CommGPT: A Graph and Retrieval-Augmented Multimodal Communication Foundation Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917936543825920 |
|---|---|
| author | Jiang, Feibo Zhu, Wanyun Dong, Li Wang, Kezhi Yang, Kun Pan, Cunhua Dobre, Octavia A. |
| author_facet | Jiang, Feibo Zhu, Wanyun Dong, Li Wang, Kezhi Yang, Kun Pan, Cunhua Dobre, Octavia A. |
| contents | Large Language Models (LLMs) possess human-level cognitive and decision-making capabilities, making them a key technology for 6G. However, applying LLMs to the communication domain faces three major challenges: 1) Inadequate communication data; 2) Restricted input modalities; and 3) Difficulty in knowledge retrieval. To overcome these issues, we propose CommGPT, a multimodal foundation model designed specifically for communications. First, we create high-quality pretraining and fine-tuning datasets tailored in communication, enabling the LLM to engage in further pretraining and fine-tuning with communication concepts and knowledge. Then, we design a multimodal encoder to understand and process information from various input modalities. Next, we construct a Graph and Retrieval-Augmented Generation (GRG) framework, efficiently coupling Knowledge Graph (KG) with Retrieval-Augmented Generation (RAG) for multi-scale learning. Finally, we demonstrate the feasibility and effectiveness of the CommGPT through experimental validation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_18763 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CommGPT: A Graph and Retrieval-Augmented Multimodal Communication Foundation Model Jiang, Feibo Zhu, Wanyun Dong, Li Wang, Kezhi Yang, Kun Pan, Cunhua Dobre, Octavia A. Information Theory Large Language Models (LLMs) possess human-level cognitive and decision-making capabilities, making them a key technology for 6G. However, applying LLMs to the communication domain faces three major challenges: 1) Inadequate communication data; 2) Restricted input modalities; and 3) Difficulty in knowledge retrieval. To overcome these issues, we propose CommGPT, a multimodal foundation model designed specifically for communications. First, we create high-quality pretraining and fine-tuning datasets tailored in communication, enabling the LLM to engage in further pretraining and fine-tuning with communication concepts and knowledge. Then, we design a multimodal encoder to understand and process information from various input modalities. Next, we construct a Graph and Retrieval-Augmented Generation (GRG) framework, efficiently coupling Knowledge Graph (KG) with Retrieval-Augmented Generation (RAG) for multi-scale learning. Finally, we demonstrate the feasibility and effectiveness of the CommGPT through experimental validation. |
| title | CommGPT: A Graph and Retrieval-Augmented Multimodal Communication Foundation Model |
| topic | Information Theory |
| url | https://arxiv.org/abs/2502.18763 |