Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
Fuente:
arXiv
Saved in:
| Main Authors: | Surikuchi, Aditya K, Fernández, Raquel, Pezzelle, Sandro |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
by: Surikuchi, Aditya K, et al.
Published: (2024)
by: Surikuchi, Aditya K, et al.
Published: (2024)
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
by: Surikuchi, Aditya K, et al.
Published: (2025)
by: Surikuchi, Aditya K, et al.
Published: (2025)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
by: Chen, Xinyi, et al.
Published: (2023)
by: Chen, Xinyi, et al.
Published: (2023)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024)
by: Takmaz, Ece, et al.
Published: (2024)
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
by: Li, Zhang, et al.
Published: (2023)
by: Li, Zhang, et al.
Published: (2023)
Reasoning Can Hurt the Inductive Abilities of Large Language Models
by: Jin, Haibo, et al.
Published: (2025)
by: Jin, Haibo, et al.
Published: (2025)
Text or Image? What is More Important in Cross-Domain Generalization Capabilities of Hate Meme Detection Models?
by: Aggarwal, Piush, et al.
Published: (2024)
by: Aggarwal, Piush, et al.
Published: (2024)
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
by: Kang, Zhaolu, et al.
Published: (2025)
by: Kang, Zhaolu, et al.
Published: (2025)
ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
by: Zhang, Juntian, et al.
Published: (2025)
by: Zhang, Juntian, et al.
Published: (2025)
Interfacing Foundation Models' Embeddings
by: Zou, Xueyan, et al.
Published: (2023)
by: Zou, Xueyan, et al.
Published: (2023)
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
by: Buckley, Thomas, et al.
Published: (2023)
by: Buckley, Thomas, et al.
Published: (2023)
RADAR: Revealing Asymmetric Development of Abilities in MLLM Pre-training
by: Nie, Yunshuang, et al.
Published: (2026)
by: Nie, Yunshuang, et al.
Published: (2026)
Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities
by: Dutt, Raman, et al.
Published: (2025)
by: Dutt, Raman, et al.
Published: (2025)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Founder effects shape the evolutionary dynamics of multimodality in open LLM families
by: Cebrian, Manuel
Published: (2026)
by: Cebrian, Manuel
Published: (2026)
Moment Sampling in Video LLMs for Long-Form Video QA
by: Chasmai, Mustafa, et al.
Published: (2025)
by: Chasmai, Mustafa, et al.
Published: (2025)
Dual-View Visual Contextualization for Web Navigation
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
GlitchBench: Can large multimodal models detect video game glitches?
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
MAIRA-1: A specialised large multimodal model for radiology report generation
by: Hyland, Stephanie L., et al.
Published: (2023)
by: Hyland, Stephanie L., et al.
Published: (2023)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
EPEE: Towards Efficient and Effective Foundation Models in Biomedicine
by: Zhan, Zaifu, et al.
Published: (2025)
by: Zhan, Zaifu, et al.
Published: (2025)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
by: Sharma, Aditya, et al.
Published: (2024)
by: Sharma, Aditya, et al.
Published: (2024)
Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine
by: Jin, Qiao, et al.
Published: (2024)
by: Jin, Qiao, et al.
Published: (2024)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
by: Padlewski, Piotr, et al.
Published: (2024)
by: Padlewski, Piotr, et al.
Published: (2024)
TVWorld: Foundations for Remote-Control TV Agents
by: Ma, Zhantao, et al.
Published: (2026)
by: Ma, Zhantao, et al.
Published: (2026)
Historical Test-time Prompt Tuning for Vision Foundation Models
by: Zhang, Jingyi, et al.
Published: (2024)
by: Zhang, Jingyi, et al.
Published: (2024)
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
by: Jin, Haibo, et al.
Published: (2025)
by: Jin, Haibo, et al.
Published: (2025)
MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models
by: Li, Hengzhi, et al.
Published: (2025)
by: Li, Hengzhi, et al.
Published: (2025)
Toward Robust Multimodal Learning using Multimodal Foundational Models
by: Zhao, Xianbing, et al.
Published: (2024)
by: Zhao, Xianbing, et al.
Published: (2024)
SCOPE: Sign Language Contextual Processing with Embedding from LLMs
by: Liu, Yuqi, et al.
Published: (2024)
by: Liu, Yuqi, et al.
Published: (2024)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
by: Shangguan, Ziyao, et al.
Published: (2024)
by: Shangguan, Ziyao, et al.
Published: (2024)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
by: Li, Wenbin, et al.
Published: (2026)
by: Li, Wenbin, et al.
Published: (2026)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
by: Bhatia, Mehar, et al.
Published: (2024)
by: Bhatia, Mehar, et al.
Published: (2024)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
by: Zhang, Zhiqian, et al.
Published: (2026)
by: Zhang, Zhiqian, et al.
Published: (2026)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Similar Items
-
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
by: Surikuchi, Aditya K, et al.
Published: (2024) -
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
by: Surikuchi, Aditya K, et al.
Published: (2025) -
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
by: Chen, Xinyi, et al.
Published: (2023) -
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024) -
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
by: Li, Zhang, et al.
Published: (2023)