Grandes Modelos de Linguagem Multimodais (MLLMs): Da Teoria à Prática
Fuente:
arXiv
Saved in:
| Main Authors: | da Silva, Neemias, Scholz, Júlio C. W., Harrison, John, Borges, Marina, Ávila, Paulo, Santos, Frances A, Delgado, Myriam, Minetto, Rodrigo, Silva, Thiago H |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception
by: da Silva, Neemias, et al.
Published: (2026)
by: da Silva, Neemias, et al.
Published: (2026)
Multimodal LLMs See Sentiment
by: da Silva, Neemias B., et al.
Published: (2025)
by: da Silva, Neemias B., et al.
Published: (2025)
Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception
by: da Silva, Neemias B, et al.
Published: (2026)
by: da Silva, Neemias B, et al.
Published: (2026)
Padrão Finder Nível Ômega: Arquitetura de Redes Polifrontes Antifrágeis baseada no Paradoxo de Kjellstadli (2019) e na Filosofia da Mente de Dennett
by: Ferreira, Neemias da Silva
Published: (2026)
by: Ferreira, Neemias da Silva
Published: (2026)
Advancing Multinational License Plate Recognition Through Synthetic and Real Data Fusion: A Comprehensive Evaluation
by: Laroca, Rayson, et al.
Published: (2026)
by: Laroca, Rayson, et al.
Published: (2026)
Towards spatiotemporal integration of bus transit with data-driven approaches
by: Borges, Júlio, et al.
Published: (2024)
by: Borges, Júlio, et al.
Published: (2024)
Leveraging LLMs for On-the-Fly Instruction Guided Image Editing
by: Santos, Rodrigo, et al.
Published: (2024)
by: Santos, Rodrigo, et al.
Published: (2024)
Na Prática, qual IA Entende o Direito? Um Estudo Experimental com IAs Generalistas e uma IA Jurídica
by: Marinho, Marina Soares, et al.
Published: (2025)
by: Marinho, Marina Soares, et al.
Published: (2025)
Narrativas Multimodais: a imagem dos matemáticos em performances matemáticas digitais
by: Ricardo Scucuglia Rodrigues da Silva
Published: (2014)
by: Ricardo Scucuglia Rodrigues da Silva
Published: (2014)
Linking Perception, Confidence and Accuracy in MLLMs
by: Du, Yuetian, et al.
Published: (2026)
by: Du, Yuetian, et al.
Published: (2026)
Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training
by: Santos, Rodrigo, et al.
Published: (2025)
by: Santos, Rodrigo, et al.
Published: (2025)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
by: Jiang, Shixin, et al.
Published: (2024)
by: Jiang, Shixin, et al.
Published: (2024)
Unhackable Temporal Rewarding for Scalable Video MLLMs
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
Práticas educativas e estresse parental de pais de crianças pequenas com desenvolvimento típico e atípico
by: Maria de Fátima Minetto
Published: (2012)
by: Maria de Fátima Minetto
Published: (2012)
Affordance Benchmark for MLLMs
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
CityHood: An Explainable Travel Recommender System for Cities and Neighborhoods
by: Santos, Gustavo H, et al.
Published: (2025)
by: Santos, Gustavo H, et al.
Published: (2025)
Interest Networks (iNETs) for Cities: Cross-Platform Insights and Urban Behavior Explanations
by: Santos, Gustavo H., et al.
Published: (2025)
by: Santos, Gustavo H., et al.
Published: (2025)
Do MLLMs Really Understand the Charts?
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
Can MLLMs Understand the Deep Implication Behind Chinese Images?
by: Zhang, Chenhao, et al.
Published: (2024)
by: Zhang, Chenhao, et al.
Published: (2024)
Clustered Retrieved Augmented Generation (CRAG)
by: Akesson, Simon, et al.
Published: (2024)
by: Akesson, Simon, et al.
Published: (2024)
Sovereign AI-based Public Services are Viable and Affordable
by: Branco, António, et al.
Published: (2026)
by: Branco, António, et al.
Published: (2026)
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
Law of Vision Representation in MLLMs
by: Yang, Shijia, et al.
Published: (2024)
by: Yang, Shijia, et al.
Published: (2024)
Benchmarking Large and Small MLLMs
by: Feng, Xuelu, et al.
Published: (2025)
by: Feng, Xuelu, et al.
Published: (2025)
Redundancy Principles for MLLMs Benchmarks
by: Zhang, Zicheng, et al.
Published: (2025)
by: Zhang, Zicheng, et al.
Published: (2025)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
by: Kim, Jungeun, et al.
Published: (2024)
by: Kim, Jungeun, et al.
Published: (2024)
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
by: Xu, Yichen, et al.
Published: (2026)
by: Xu, Yichen, et al.
Published: (2026)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
by: Huang, Jen-Tse, et al.
Published: (2025)
by: Huang, Jen-Tse, et al.
Published: (2025)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Exploring the Design Space of Visual Context Representation in Video MLLMs
by: Du, Yifan, et al.
Published: (2024)
by: Du, Yifan, et al.
Published: (2024)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024)
by: Yao, Huanjin, et al.
Published: (2024)
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
by: Han, Tianyang, et al.
Published: (2024)
by: Han, Tianyang, et al.
Published: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
by: Yilmaz, Nilay, et al.
Published: (2025)
by: Yilmaz, Nilay, et al.
Published: (2025)
MileBench: Benchmarking MLLMs in Long Context
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
Similar Items
-
Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception
by: da Silva, Neemias, et al.
Published: (2026) -
Multimodal LLMs See Sentiment
by: da Silva, Neemias B., et al.
Published: (2025) -
Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception
by: da Silva, Neemias B, et al.
Published: (2026) -
Padrão Finder Nível Ômega: Arquitetura de Redes Polifrontes Antifrágeis baseada no Paradoxo de Kjellstadli (2019) e na Filosofia da Mente de Dennett
by: Ferreira, Neemias da Silva
Published: (2026) -
Advancing Multinational License Plate Recognition Through Synthetic and Real Data Fusion: A Comprehensive Evaluation
by: Laroca, Rayson, et al.
Published: (2026)