HEMM: Holistic Evaluation of Multimodal Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Paul Pu, Goindani, Akshay, Chafekar, Talha, Mathur, Leena, Yu, Haofei, Salakhutdinov, Ruslan, Morency, Louis-Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
by: Yu, Haofei, et al.
Published: (2023)
by: Yu, Haofei, et al.
Published: (2023)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)
by: Mo, Shentong, et al.
Published: (2023)
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
by: Hu, Jiewen, et al.
Published: (2025)
by: Hu, Jiewen, et al.
Published: (2025)
Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
by: Liang, Paul Pu, et al.
Published: (2023)
by: Liang, Paul Pu, et al.
Published: (2023)
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
by: Mathur, Leena, et al.
Published: (2024)
by: Mathur, Leena, et al.
Published: (2024)
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
by: Mathur, Leena, et al.
Published: (2025)
by: Mathur, Leena, et al.
Published: (2025)
Social Caption: Evaluating Social Understanding in Multimodal Models
by: Thumu, Bhaavanaa, et al.
Published: (2026)
by: Thumu, Bhaavanaa, et al.
Published: (2026)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
Act2See: Emergent Active Visual Perception for Video Reasoning
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
by: Zhang, Ce, et al.
Published: (2025)
by: Zhang, Ce, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
GenSync: A Generalized Talking Head Framework for Audio-driven Multi-Subject Lip-Sync using 3D Gaussian Splatting
by: Agarwal, Anushka, et al.
Published: (2025)
by: Agarwal, Anushka, et al.
Published: (2025)
Dissecting Adversarial Robustness of Multimodal LM Agents
by: Wu, Chen Henry, et al.
Published: (2024)
by: Wu, Chen Henry, et al.
Published: (2024)
Foundations of Multisensory Artificial Intelligence
by: Liang, Paul Pu
Published: (2024)
by: Liang, Paul Pu
Published: (2024)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
by: Kapoor, Raghav, et al.
Published: (2024)
by: Kapoor, Raghav, et al.
Published: (2024)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
by: Wan, Zifu, et al.
Published: (2025)
by: Wan, Zifu, et al.
Published: (2025)
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models
by: Li, Hengzhi, et al.
Published: (2025)
by: Li, Hengzhi, et al.
Published: (2025)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
by: Zhou, Xuhui, et al.
Published: (2023)
by: Zhou, Xuhui, et al.
Published: (2023)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
by: Qiu, Haoyi, et al.
Published: (2024)
by: Qiu, Haoyi, et al.
Published: (2024)
Holistic Evaluation for Interleaved Text-and-Image Generation
by: Liu, Minqian, et al.
Published: (2024)
by: Liu, Minqian, et al.
Published: (2024)
Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation
by: Pennec, Galann, et al.
Published: (2025)
by: Pennec, Galann, et al.
Published: (2025)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
A Vision for Multisensory Intelligence: Sensing, Science, and Synergy
by: Liang, Paul Pu
Published: (2026)
by: Liang, Paul Pu
Published: (2026)
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
by: Zhang, Kaichen, et al.
Published: (2024)
by: Zhang, Kaichen, et al.
Published: (2024)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
by: Li, Jiaang, et al.
Published: (2025)
by: Li, Jiaang, et al.
Published: (2025)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024)
by: Wu, Zhiyong, et al.
Published: (2024)
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
by: Mathur, Suyash Vardhan, et al.
Published: (2024)
by: Mathur, Suyash Vardhan, et al.
Published: (2024)
Multimodal Structured Generation: CVPR's 2nd MMFM Challenge Technical Report
by: Cesista, Franz Louis
Published: (2024)
by: Cesista, Franz Louis
Published: (2024)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
by: Liu, Chengzhi, et al.
Published: (2026)
by: Liu, Chengzhi, et al.
Published: (2026)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
by: Cheng, Zhili, et al.
Published: (2025)
by: Cheng, Zhili, et al.
Published: (2025)
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
by: Diesendruck, Maurice, et al.
Published: (2024)
by: Diesendruck, Maurice, et al.
Published: (2024)
Survey of Video Diffusion Models: Foundations, Implementations, and Applications
by: Wang, Yimu, et al.
Published: (2025)
by: Wang, Yimu, et al.
Published: (2025)
Toward Robust Multimodal Learning using Multimodal Foundational Models
by: Zhao, Xianbing, et al.
Published: (2024)
by: Zhao, Xianbing, et al.
Published: (2024)
Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs
by: Dey, Abhishek, et al.
Published: (2025)
by: Dey, Abhishek, et al.
Published: (2025)
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
by: Zhang, Zhiqian, et al.
Published: (2026)
by: Zhang, Zhiqian, et al.
Published: (2026)
A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
by: Chen, Zhihong, et al.
Published: (2024)
by: Chen, Zhihong, et al.
Published: (2024)
Stylus: Automatic Adapter Selection for Diffusion Models
by: Luo, Michael, et al.
Published: (2024)
by: Luo, Michael, et al.
Published: (2024)
Similar Items
-
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
by: Yu, Haofei, et al.
Published: (2023) -
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023) -
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024) -
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
by: Hu, Jiewen, et al.
Published: (2025) -
Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
by: Liang, Paul Pu, et al.
Published: (2023)