Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities
Fuente:
arXiv
Saved in:
| Main Authors: | Munikoti, Sai, Stewart, Ian, Horawalavithana, Sameera, Kvinge, Henry, Emerson, Tegan, Thompson, Sandra E, Pazdernik, Karl |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
by: Horawalavithana, Sameera, et al.
Published: (2023)
by: Horawalavithana, Sameera, et al.
Published: (2023)
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
by: Horawalavithana, Sameera, et al.
Published: (2026)
by: Horawalavithana, Sameera, et al.
Published: (2026)
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
by: Munikoti, Sai, et al.
Published: (2026)
by: Munikoti, Sai, et al.
Published: (2026)
Benchmarking LLMs for Environmental Review and Permitting
by: Meyur, Rounak, et al.
Published: (2024)
by: Meyur, Rounak, et al.
Published: (2024)
Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing
by: Raab, Reilly, et al.
Published: (2025)
by: Raab, Reilly, et al.
Published: (2025)
Directional Concentration Uncertainty: A representational approach to uncertainty quantification for generative models
by: Chattopadhyay, Souradeep, et al.
Published: (2026)
by: Chattopadhyay, Souradeep, et al.
Published: (2026)
WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain
by: Meyur, Rounak, et al.
Published: (2024)
by: Meyur, Rounak, et al.
Published: (2024)
Reward Design for Physical Reasoning in Vision-Language Models
by: Lilienthal, Derek, et al.
Published: (2026)
by: Lilienthal, Derek, et al.
Published: (2026)
Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains
by: Chaturvedi, Sarthak, et al.
Published: (2025)
by: Chaturvedi, Sarthak, et al.
Published: (2025)
Uncertainty Quantification for Named Entity Recognition via Full-Sequence and Subsequence Conformal Prediction
by: Singer, Matthew, et al.
Published: (2026)
by: Singer, Matthew, et al.
Published: (2026)
A Multiscale Geometric Method for Capturing Relational Topic Alignment
by: Hougen, Conrad D., et al.
Published: (2025)
by: Hougen, Conrad D., et al.
Published: (2025)
A Perspective for Adapting Generalist AI to Specialized Medical AI Applications and Their Challenges
by: Wang, Zifeng, et al.
Published: (2024)
by: Wang, Zifeng, et al.
Published: (2024)
Even with AI, Bijection Discovery is Still Hard: The Opportunities and Challenges of OpenEvolve for Novel Bijection Construction
by: Brown, Davis, et al.
Published: (2025)
by: Brown, Davis, et al.
Published: (2025)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
by: Soni, Aditya Bharat, et al.
Published: (2025)
by: Soni, Aditya Bharat, et al.
Published: (2025)
Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge
by: Yeats, Eric, et al.
Published: (2025)
by: Yeats, Eric, et al.
Published: (2025)
Evaluation of OpenAI o1: Opportunities and Challenges of AGI
by: Zhong, Tianyang, et al.
Published: (2024)
by: Zhong, Tianyang, et al.
Published: (2024)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
by: Vishwanath, Krithik, et al.
Published: (2025)
by: Vishwanath, Krithik, et al.
Published: (2025)
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
by: Vatsa, Mayank, et al.
Published: (2025)
by: Vatsa, Mayank, et al.
Published: (2025)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Towards Building Specialized Generalist AI with System 1 and System 2 Fusion
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation
by: Jamil, Sofia, et al.
Published: (2025)
by: Jamil, Sofia, et al.
Published: (2025)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
by: Mandal, Aishik, et al.
Published: (2025)
by: Mandal, Aishik, et al.
Published: (2025)
GSCo: Towards Generalizable AI in Medicine via Generalist-Specialist Collaboration
by: He, Sunan, et al.
Published: (2024)
by: He, Sunan, et al.
Published: (2024)
On-Device LLMs for SMEs: Challenges and Opportunities
by: Yee, Jeremy Stephen Gabriel, et al.
Published: (2024)
by: Yee, Jeremy Stephen Gabriel, et al.
Published: (2024)
Opportunities and Challenges of LLMs in Education: An NLP Perspective
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
LLMs in Education: Novel Perspectives, Challenges, and Opportunities
by: Alhafni, Bashar, et al.
Published: (2024)
by: Alhafni, Bashar, et al.
Published: (2024)
Leveraging AI to Advance Science and Computing Education across Africa: Challenges, Progress and Opportunities
by: Boateng, George
Published: (2024)
by: Boateng, George
Published: (2024)
LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
by: Hill, Graham, et al.
Published: (2025)
by: Hill, Graham, et al.
Published: (2025)
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts
by: Wu, Jialin, et al.
Published: (2023)
by: Wu, Jialin, et al.
Published: (2023)
A Framework for Situating Innovations, Opportunities, and Challenges in Advancing Vertical Systems with Large AI Models
by: Verma, Gaurav, et al.
Published: (2025)
by: Verma, Gaurav, et al.
Published: (2025)
Challenges and Opportunities in Text Generation Explainability
by: Amara, Kenza, et al.
Published: (2024)
by: Amara, Kenza, et al.
Published: (2024)
UltraMedical: Building Specialized Generalists in Biomedicine
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
Reasoning Over Recall: Evaluating the Efficacy of Generalist Architectures vs. Specialized Fine-Tunes in RAG-Based Mental Health Dialogue Systems
by: Kafi, Md Abdullah Al, et al.
Published: (2026)
by: Kafi, Md Abdullah Al, et al.
Published: (2026)
Multimodal Detection of Fake Reviews using BERT and ResNet-50
by: Veluru, Suhasnadh Reddy, et al.
Published: (2025)
by: Veluru, Suhasnadh Reddy, et al.
Published: (2025)
A Systematic Review of NLP for Dementia -- Tasks, Datasets and Opportunities
by: Peled-Cohen, Lotem, et al.
Published: (2024)
by: Peled-Cohen, Lotem, et al.
Published: (2024)
Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
by: Jiang, Lavender Y., et al.
Published: (2025)
by: Jiang, Lavender Y., et al.
Published: (2025)
NLP for Social Good: A Survey and Outlook of Challenges, Opportunities, and Responsible Deployment
by: Karamolegkou, Antonia, et al.
Published: (2025)
by: Karamolegkou, Antonia, et al.
Published: (2025)
Similar Items
-
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
by: Horawalavithana, Sameera, et al.
Published: (2023) -
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
by: Stewart, Ian, et al.
Published: (2024) -
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
by: Horawalavithana, Sameera, et al.
Published: (2026) -
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
by: Munikoti, Sai, et al.
Published: (2026) -
Benchmarking LLMs for Environmental Review and Permitting
by: Meyur, Rounak, et al.
Published: (2024)