Grandes Modelos de Linguagem Multimodais (MLLMs): Da Teoria à Prática
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914327558094848 |
|---|---|
| author | da Silva, Neemias Scholz, Júlio C. W. Harrison, John Borges, Marina Ávila, Paulo Santos, Frances A Delgado, Myriam Minetto, Rodrigo Silva, Thiago H |
| author_facet | da Silva, Neemias Scholz, Júlio C. W. Harrison, John Borges, Marina Ávila, Paulo Santos, Frances A Delgado, Myriam Minetto, Rodrigo Silva, Thiago H |
| contents | Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_12302 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Grandes Modelos de Linguagem Multimodais (MLLMs): Da Teoria à Prática da Silva, Neemias Scholz, Júlio C. W. Harrison, John Borges, Marina Ávila, Paulo Santos, Frances A Delgado, Myriam Minetto, Rodrigo Silva, Thiago H Computation and Language Computer Vision and Pattern Recognition Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends. |
| title | Grandes Modelos de Linguagem Multimodais (MLLMs): Da Teoria à Prática |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.12302 |