Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Soyeon Caren, Cao, Feiqi, Poon, Josiah, Navigli, Roberto
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909340124839936
author Han, Soyeon Caren
Cao, Feiqi
Poon, Josiah
Navigli, Roberto
author_facet Han, Soyeon Caren
Cao, Feiqi
Poon, Josiah
Navigli, Roberto
contents This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understanding of the foundational concepts of multimodality, the evolution of multimodal research, and the key technical challenges addressed by these models. We will cover the latest multimodal datasets and pretrained models, including those beyond vision and language. Additionally, the tutorial will delve into the intricacies of multimodal large models and instruction tuning strategies to optimise performance for specific tasks. Hands-on laboratories will offer practical experience with state-of-the-art multimodal models, demonstrating real-world applications like visual storytelling and visual question answering. This tutorial aims to equip researchers, practitioners, and newcomers with the knowledge and skills to leverage multimodal AI. ACM Multimedia 2024 is the ideal venue for this tutorial, aligning perfectly with our goal of understanding multimodal pretrained and large language models, and their tuning mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05608
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
Han, Soyeon Caren
Cao, Feiqi
Poon, Josiah
Navigli, Roberto
Computation and Language
This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understanding of the foundational concepts of multimodality, the evolution of multimodal research, and the key technical challenges addressed by these models. We will cover the latest multimodal datasets and pretrained models, including those beyond vision and language. Additionally, the tutorial will delve into the intricacies of multimodal large models and instruction tuning strategies to optimise performance for specific tasks. Hands-on laboratories will offer practical experience with state-of-the-art multimodal models, demonstrating real-world applications like visual storytelling and visual question answering. This tutorial aims to equip researchers, practitioners, and newcomers with the knowledge and skills to leverage multimodal AI. ACM Multimedia 2024 is the ideal venue for this tutorial, aligning perfectly with our goal of understanding multimodal pretrained and large language models, and their tuning mechanisms.
title Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
topic Computation and Language
url https://arxiv.org/abs/2410.05608