Evaluating point-light biological motion in multimodal large language models
Fuente:
arXiv
Saved in:
| Main Authors: | Kadambi, Akila, Iacoboni, Marco, Aziz-Zadeh, Lisa, Narayanan, Srini |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embodiment in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025)
by: Zhang, Ruixuan, et al.
Published: (2025)
Teaching large language models to reason like expert diagnosticians
by: Buckley, Thomas A., et al.
Published: (2025)
by: Buckley, Thomas A., et al.
Published: (2025)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Hallucination-aware intermediate representation edit in large vision-language models
by: Suo, Wei, et al.
Published: (2026)
by: Suo, Wei, et al.
Published: (2026)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
by: Padlewski, Piotr, et al.
Published: (2024)
by: Padlewski, Piotr, et al.
Published: (2024)
GlitchBench: Can large multimodal models detect video game glitches?
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
MAIRA-1: A specialised large multimodal model for radiology report generation
by: Hyland, Stephanie L., et al.
Published: (2023)
by: Hyland, Stephanie L., et al.
Published: (2023)
Human-like object concept representations emerge naturally in multimodal large language models
by: Du, Changde, et al.
Published: (2024)
by: Du, Changde, et al.
Published: (2024)
Explaining latent representations of generative models with large multimodal models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
by: Xing, Yang, et al.
Published: (2026)
by: Xing, Yang, et al.
Published: (2026)
From YOLO to VLMs: Advancing Zero-Shot and Few-Shot Detection of Wastewater Treatment Plants Using Satellite Imagery in MENA Region
by: Premarathna, Akila, et al.
Published: (2025)
by: Premarathna, Akila, et al.
Published: (2025)
SmolVLM: Redefining small and efficient multimodal models
by: Marafioti, Andrés, et al.
Published: (2025)
by: Marafioti, Andrés, et al.
Published: (2025)
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
by: Dong, Owen, et al.
Published: (2026)
by: Dong, Owen, et al.
Published: (2026)
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
by: Granite Vision Team, et al.
Published: (2025)
by: Granite Vision Team, et al.
Published: (2025)
Explainable artificial intelligence (XAI): from inherent explainability to large language models
by: Mumuni, Fuseini, et al.
Published: (2025)
by: Mumuni, Fuseini, et al.
Published: (2025)
Semi-supervised classification of dental conditions in panoramic radiographs using large language model and instance segmentation: A real-world dataset evaluation
by: Silva, Bernardo, et al.
Published: (2024)
by: Silva, Bernardo, et al.
Published: (2024)
Parameter-Efficient Active Learning for Foundational models
by: Narayanan, Athmanarayanan Lakshmi, et al.
Published: (2024)
by: Narayanan, Athmanarayanan Lakshmi, et al.
Published: (2024)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
LPOI: Listwise Preference Optimization for Vision Language Models
by: Zadeh, Fatemeh Pesaran, et al.
Published: (2025)
by: Zadeh, Fatemeh Pesaran, et al.
Published: (2025)
Evaluating Large Vision-language Models for Surgical Tool Detection
by: Poudel, Nakul, et al.
Published: (2026)
by: Poudel, Nakul, et al.
Published: (2026)
Vision language models are unreliable at trivial spatial cognition
by: Khemlani, Sangeet, et al.
Published: (2025)
by: Khemlani, Sangeet, et al.
Published: (2025)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
PB-IAD: Utilizing multimodal foundation models for semantic industrial anomaly detection in dynamic manufacturing environments
by: Hofmann, Bernd, et al.
Published: (2025)
by: Hofmann, Bernd, et al.
Published: (2025)
GeoLocator: a location-integrated large multimodal model for inferring geo-privacy
by: Yang, Yifan, et al.
Published: (2023)
by: Yang, Yifan, et al.
Published: (2023)
Multimodal Evaluation of Russian-language Architectures
by: Chervyakov, Artem, et al.
Published: (2025)
by: Chervyakov, Artem, et al.
Published: (2025)
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
Vision language models have difficulty recognizing virtual objects
by: Tran, Tyler, et al.
Published: (2025)
by: Tran, Tyler, et al.
Published: (2025)
Thinker: A vision-language foundation model for embodied intelligence
by: Pan, Baiyu, et al.
Published: (2026)
by: Pan, Baiyu, et al.
Published: (2026)
A multi-scale vision transformer-based multimodal GeoAI model for mapping Arctic permafrost thaw
by: Li, Wenwen, et al.
Published: (2025)
by: Li, Wenwen, et al.
Published: (2025)
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues
by: Ruan, Zanxi, et al.
Published: (2026)
by: Ruan, Zanxi, et al.
Published: (2026)
EHWGesture -- A dataset for multimodal understanding of clinical gestures
by: Amprimo, Gianluca, et al.
Published: (2025)
by: Amprimo, Gianluca, et al.
Published: (2025)
Generalizing vision-language models to novel domains: A comprehensive survey
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
Building and better understanding vision-language models: insights and future directions
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Leveraging AI multimodal geospatial foundation models for improved near-real-time flood mapping at a global scale
by: Tulbure, Mirela G., et al.
Published: (2025)
by: Tulbure, Mirela G., et al.
Published: (2025)
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
Similar Items
-
Embodiment in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025) -
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025) -
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025) -
Teaching large language models to reason like expert diagnosticians
by: Buckley, Thomas A., et al.
Published: (2025) -
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)