Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Juncheng, Li, Zuchao, Xie, Shuai, Yu, Wei, Li, Shijun, Du, Bo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916197252988928
author Yang, Juncheng
Li, Zuchao
Xie, Shuai
Yu, Wei
Li, Shijun
Du, Bo
author_facet Yang, Juncheng
Li, Zuchao
Xie, Shuai
Yu, Wei
Li, Shijun
Du, Bo
contents The chain-of-thought technique has been received well in multi-modal tasks. It is a step-by-step linear reasoning process that adjusts the length of the chain to improve the performance of generated prompts. However, human thought processes are predominantly non-linear, as they encompass multiple aspects simultaneously and employ dynamic adjustment and updating mechanisms. Therefore, we propose a novel Aggregation-Graph-of-Thought (AGoT) mechanism for soft-prompt tuning in multi-modal representation learning. The proposed AGoT models the human thought process not only as a chain but also models each step as a reasoning aggregation graph to cope with the overlooked multiple aspects of thinking in single-step reasoning. This turns the entire reasoning process into prompt aggregation and prompt flow operations. Experiments show that our multi-modal model enhanced with AGoT soft-prompting achieves good results in several tasks such as text-image retrieval, visual question answering, and image recognition. In addition, we demonstrate that it has good domain generalization performance due to better reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04538
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning
Yang, Juncheng
Li, Zuchao
Xie, Shuai
Yu, Wei
Li, Shijun
Du, Bo
Artificial Intelligence
Computation and Language
The chain-of-thought technique has been received well in multi-modal tasks. It is a step-by-step linear reasoning process that adjusts the length of the chain to improve the performance of generated prompts. However, human thought processes are predominantly non-linear, as they encompass multiple aspects simultaneously and employ dynamic adjustment and updating mechanisms. Therefore, we propose a novel Aggregation-Graph-of-Thought (AGoT) mechanism for soft-prompt tuning in multi-modal representation learning. The proposed AGoT models the human thought process not only as a chain but also models each step as a reasoning aggregation graph to cope with the overlooked multiple aspects of thinking in single-step reasoning. This turns the entire reasoning process into prompt aggregation and prompt flow operations. Experiments show that our multi-modal model enhanced with AGoT soft-prompting achieves good results in several tasks such as text-image retrieval, visual question answering, and image recognition. In addition, we demonstrate that it has good domain generalization performance due to better reasoning.
title Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2404.04538