Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Saxena, Rohit, Gema, Aryo Pradipta, Minervini, Pasquale |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
di: Saxena, Rohit, et al.
Pubblicazione: (2025)
di: Saxena, Rohit, et al.
Pubblicazione: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
di: Saxena, Rohit, et al.
Pubblicazione: (2026)
di: Saxena, Rohit, et al.
Pubblicazione: (2026)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
di: Attimonelli, Matteo, et al.
Pubblicazione: (2026)
di: Attimonelli, Matteo, et al.
Pubblicazione: (2026)
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2025)
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2025)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
di: Sarto, Sara, et al.
Pubblicazione: (2025)
di: Sarto, Sara, et al.
Pubblicazione: (2025)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
di: Wang, Haochen, et al.
Pubblicazione: (2025)
di: Wang, Haochen, et al.
Pubblicazione: (2025)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
di: Ren, Shuhuai, et al.
Pubblicazione: (2023)
di: Ren, Shuhuai, et al.
Pubblicazione: (2023)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
di: Murphy, Alexander, et al.
Pubblicazione: (2025)
di: Murphy, Alexander, et al.
Pubblicazione: (2025)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
di: Zhang, Jun, et al.
Pubblicazione: (2025)
di: Zhang, Jun, et al.
Pubblicazione: (2025)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
di: Li, Yunxin, et al.
Pubblicazione: (2023)
di: Li, Yunxin, et al.
Pubblicazione: (2023)
MULTI: Multimodal Understanding Leaderboard with Text and Images
di: Zhu, Zichen, et al.
Pubblicazione: (2024)
di: Zhu, Zichen, et al.
Pubblicazione: (2024)
Scientific Reasoning: Assessment of Multimodal Generative LLMs
di: Dreyer, Florian, et al.
Pubblicazione: (2025)
di: Dreyer, Florian, et al.
Pubblicazione: (2025)
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
di: Khayatan, Pegah, et al.
Pubblicazione: (2025)
di: Khayatan, Pegah, et al.
Pubblicazione: (2025)
On Pre-training of Multimodal Language Models Customized for Chart Understanding
di: Fan, Wan-Cyuan, et al.
Pubblicazione: (2024)
di: Fan, Wan-Cyuan, et al.
Pubblicazione: (2024)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
di: Liu, Dongqi, et al.
Pubblicazione: (2025)
di: Liu, Dongqi, et al.
Pubblicazione: (2025)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
di: Zhao, Xiangyu, et al.
Pubblicazione: (2023)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2023)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
di: Deng, Ken, et al.
Pubblicazione: (2026)
di: Deng, Ken, et al.
Pubblicazione: (2026)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
di: Ryan, Yuriel, et al.
Pubblicazione: (2025)
di: Ryan, Yuriel, et al.
Pubblicazione: (2025)
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
di: Chen, Xiaokang, et al.
Pubblicazione: (2025)
di: Chen, Xiaokang, et al.
Pubblicazione: (2025)
FaceLLM: A Multimodal Large Language Model for Face Understanding
di: Shahreza, Hatef Otroshi, et al.
Pubblicazione: (2025)
di: Shahreza, Hatef Otroshi, et al.
Pubblicazione: (2025)
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
di: Zhang, Ge, et al.
Pubblicazione: (2024)
di: Zhang, Ge, et al.
Pubblicazione: (2024)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
di: Song, Tingyu, et al.
Pubblicazione: (2025)
di: Song, Tingyu, et al.
Pubblicazione: (2025)
Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring
di: Onsu, Murat Arda, et al.
Pubblicazione: (2025)
di: Onsu, Murat Arda, et al.
Pubblicazione: (2025)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
di: Zhang, Jiarui, et al.
Pubblicazione: (2024)
di: Zhang, Jiarui, et al.
Pubblicazione: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
di: Xie, Yuxuan, et al.
Pubblicazione: (2024)
di: Xie, Yuxuan, et al.
Pubblicazione: (2024)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
di: LASA Team, et al.
Pubblicazione: (2025)
di: LASA Team, et al.
Pubblicazione: (2025)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
di: Liu, Ziqiang, et al.
Pubblicazione: (2024)
di: Liu, Ziqiang, et al.
Pubblicazione: (2024)
MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding
di: Li, Zekun, et al.
Pubblicazione: (2024)
di: Li, Zekun, et al.
Pubblicazione: (2024)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
di: Broomfield, Julius, et al.
Pubblicazione: (2025)
di: Broomfield, Julius, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
di: Saxena, Rohit, et al.
Pubblicazione: (2025) -
Same Answer, Different Representations: Hidden instability in VLMs
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026) -
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
di: Saxena, Rohit, et al.
Pubblicazione: (2026) -
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
di: Attimonelli, Matteo, et al.
Pubblicazione: (2026) -
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)