Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeer, Ahmed, Dogan, Eren, Erdem, Yusuf, Ince, Elif, Shbib, Osama, Uzun, M. Egemen, Uz, Atahan, Yuce, M. Kaan, Kesgin, H. Toprak, Amasyali, M. Fatih
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915047456899072
author Zeer, Ahmed
Dogan, Eren
Erdem, Yusuf
Ince, Elif
Shbib, Osama
Uzun, M. Egemen
Uz, Atahan
Yuce, M. Kaan
Kesgin, H. Toprak
Amasyali, M. Fatih
author_facet Zeer, Ahmed
Dogan, Eren
Erdem, Yusuf
Ince, Elif
Shbib, Osama
Uzun, M. Egemen
Uz, Atahan
Yuce, M. Kaan
Kesgin, H. Toprak
Amasyali, M. Fatih
contents In this study, a Turkish visual instruction model was developed and various model architectures and dataset combinations were analysed to improve the performance of this model. The Cosmos-LLaVA model, which is built by combining different large language models and image coders, is designed to overcome the deficiencies in the Turkish language. In the experiments, the effects of fine-tuning with various datasets on the model performance are analysed in detail. The results show that model architecture and dataset selection have a significant impact on performance. Bu çalışmada bir Türkçe görsel talimat modeli geliştirilerek bu modelin performansını artırmaya yönelik çeşitli model mimarileri ve veri kümesi kombinasyonları derinlemesine incelenmiştir. Farklı büyük dil modelleri ve görüntü kodlayıcılarının bir araya getirilmesiyle oluşturulan Cosmos-LLaVA modeli, Türkçe dilindeki eksiklikleri gidermeye yönelik olarak tasarlanmıştır. Yapılan deneylerde, çeşitli veri kümeleri ile yapılan ince ayarların model performansını nasıl etkilediği detaylı olarak ele alınmıştır. Sonuçlar, model mimarisi ve veri kümesi seçiminin performans üzerinde önemli bir etkiye sahip olduğunu göstermektedir.
format Preprint
id arxiv_https___arxiv_org_abs_2412_02760
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
Zeer, Ahmed
Dogan, Eren
Erdem, Yusuf
Ince, Elif
Shbib, Osama
Uzun, M. Egemen
Uz, Atahan
Yuce, M. Kaan
Kesgin, H. Toprak
Amasyali, M. Fatih
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
In this study, a Turkish visual instruction model was developed and various model architectures and dataset combinations were analysed to improve the performance of this model. The Cosmos-LLaVA model, which is built by combining different large language models and image coders, is designed to overcome the deficiencies in the Turkish language. In the experiments, the effects of fine-tuning with various datasets on the model performance are analysed in detail. The results show that model architecture and dataset selection have a significant impact on performance. Bu çalışmada bir Türkçe görsel talimat modeli geliştirilerek bu modelin performansını artırmaya yönelik çeşitli model mimarileri ve veri kümesi kombinasyonları derinlemesine incelenmiştir. Farklı büyük dil modelleri ve görüntü kodlayıcılarının bir araya getirilmesiyle oluşturulan Cosmos-LLaVA modeli, Türkçe dilindeki eksiklikleri gidermeye yönelik olarak tasarlanmıştır. Yapılan deneylerde, çeşitli veri kümeleri ile yapılan ince ayarların model performansını nasıl etkilediği detaylı olarak ele alınmıştır. Sonuçlar, model mimarisi ve veri kümesi seçiminin performans üzerinde önemli bir etkiye sahip olduğunu göstermektedir.
title Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.02760