CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Deng, Xiaoyu, Kang, Zhengjian, Li, Xintao, Zhang, Yongzhe, Guo, Tianmin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913589997076480
author Deng, Xiaoyu
Kang, Zhengjian
Li, Xintao
Zhang, Yongzhe
Guo, Tianmin
author_facet Deng, Xiaoyu
Kang, Zhengjian
Li, Xintao
Zhang, Yongzhe
Guo, Tianmin
contents Graphic visual content helps in promoting information communication and inspiration divergence. However, the interpretation of visual content currently relies mainly on humans' personal knowledge background, thereby affecting the quality and efficiency of information acquisition and understanding. To improve the quality and efficiency of visual information transmission and avoid the limitation of the observer due to the information cocoon, we propose CoVis, a collaborative framework for fine-grained visual understanding. By designing and implementing a cascaded dual-layer segmentation network coupled with a large-language-model (LLM) based content generator, the framework extracts as much knowledge as possible from an image. Then, it generates visual analytics for images, assisting observers in comprehending imagery from a more holistic perspective. Quantitative experiments and qualitative experiments based on 32 human participants indicate that the CoVis has better performance than current methods in feature extraction and can generate more comprehensive and detailed visual descriptions than current general-purpose large models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18764
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding
Deng, Xiaoyu
Kang, Zhengjian
Li, Xintao
Zhang, Yongzhe
Guo, Tianmin
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphic visual content helps in promoting information communication and inspiration divergence. However, the interpretation of visual content currently relies mainly on humans' personal knowledge background, thereby affecting the quality and efficiency of information acquisition and understanding. To improve the quality and efficiency of visual information transmission and avoid the limitation of the observer due to the information cocoon, we propose CoVis, a collaborative framework for fine-grained visual understanding. By designing and implementing a cascaded dual-layer segmentation network coupled with a large-language-model (LLM) based content generator, the framework extracts as much knowledge as possible from an image. Then, it generates visual analytics for images, assisting observers in comprehending imagery from a more holistic perspective. Quantitative experiments and qualitative experiments based on 32 human participants indicate that the CoVis has better performance than current methods in feature extraction and can generate more comprehensive and detailed visual descriptions than current general-purpose large models.
title CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.18764