The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
Fuente:
arXiv
Salvato in:
| Autori principali: | Ghosh, Samrajnee, Agarwal, Naman, Garg, Hemanshu, Mittal, Chinmay, Mausam, Singla, Parag |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
di: Mittal, Chinmay, et al.
Pubblicazione: (2024)
di: Mittal, Chinmay, et al.
Pubblicazione: (2024)
Give me a hint: Can LLMs take a hint to solve math problems?
di: Agrawal, Vansh, et al.
Pubblicazione: (2024)
di: Agrawal, Vansh, et al.
Pubblicazione: (2024)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
di: Huang, Haoyu, et al.
Pubblicazione: (2026)
di: Huang, Haoyu, et al.
Pubblicazione: (2026)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
di: Ma, David, et al.
Pubblicazione: (2025)
di: Ma, David, et al.
Pubblicazione: (2025)
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
di: Wang, Zhaokai, et al.
Pubblicazione: (2025)
di: Wang, Zhaokai, et al.
Pubblicazione: (2025)
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
di: Yu, En, et al.
Pubblicazione: (2025)
di: Yu, En, et al.
Pubblicazione: (2025)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
di: Zhang, Jiarui, et al.
Pubblicazione: (2025)
di: Zhang, Jiarui, et al.
Pubblicazione: (2025)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
di: Li, Yunxin, et al.
Pubblicazione: (2025)
di: Li, Yunxin, et al.
Pubblicazione: (2025)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
di: Imam, Mohamed Fazli, et al.
Pubblicazione: (2025)
di: Imam, Mohamed Fazli, et al.
Pubblicazione: (2025)
Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions
di: Hu, Zhe, et al.
Pubblicazione: (2024)
di: Hu, Zhe, et al.
Pubblicazione: (2024)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
di: Huang, Yipo, et al.
Pubblicazione: (2024)
di: Huang, Yipo, et al.
Pubblicazione: (2024)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
di: Khan, Shaharukh, et al.
Pubblicazione: (2025)
di: Khan, Shaharukh, et al.
Pubblicazione: (2025)
Reinforced Visual Perception with Tools
di: Zhou, Zetong, et al.
Pubblicazione: (2025)
di: Zhou, Zetong, et al.
Pubblicazione: (2025)
Evaluating Graphical Perception with Multimodal LLMs
di: Nguyen, Rami Huu, et al.
Pubblicazione: (2025)
di: Nguyen, Rami Huu, et al.
Pubblicazione: (2025)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
di: Wang, Yuhao, et al.
Pubblicazione: (2024)
di: Wang, Yuhao, et al.
Pubblicazione: (2024)
On the Perception Bottleneck of VLMs for Chart Understanding
di: Liu, Junteng, et al.
Pubblicazione: (2025)
di: Liu, Junteng, et al.
Pubblicazione: (2025)
Linking Perception, Confidence and Accuracy in MLLMs
di: Du, Yuetian, et al.
Pubblicazione: (2026)
di: Du, Yuetian, et al.
Pubblicazione: (2026)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
di: Chen, Feng, et al.
Pubblicazione: (2024)
di: Chen, Feng, et al.
Pubblicazione: (2024)
Towards Understanding Graphical Perception in Large Multimodal Models
di: Zhang, Kai, et al.
Pubblicazione: (2025)
di: Zhang, Kai, et al.
Pubblicazione: (2025)
Simple Augmentations of Logical Rules for Neuro-Symbolic Knowledge Graph Completion
di: Nandi, Ananjan, et al.
Pubblicazione: (2024)
di: Nandi, Ananjan, et al.
Pubblicazione: (2024)
Teaching Human Behavior Improves Content Understanding Abilities Of LLMs
di: Singh, Somesh, et al.
Pubblicazione: (2024)
di: Singh, Somesh, et al.
Pubblicazione: (2024)
Mitigating Object Hallucination via Robust Local Perception Search
di: Gao, Zixian, et al.
Pubblicazione: (2025)
di: Gao, Zixian, et al.
Pubblicazione: (2025)
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs
di: Liang, Jiafeng, et al.
Pubblicazione: (2026)
di: Liang, Jiafeng, et al.
Pubblicazione: (2026)
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
di: Saxena, Rohit, et al.
Pubblicazione: (2025)
di: Saxena, Rohit, et al.
Pubblicazione: (2025)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
di: Ma, Yunsheng, et al.
Pubblicazione: (2024)
di: Ma, Yunsheng, et al.
Pubblicazione: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
di: Wu, Xueqing, et al.
Pubblicazione: (2026)
di: Wu, Xueqing, et al.
Pubblicazione: (2026)
Parking, Perception, and Retail: Street-Level Determinants of Community Vitality in Harbin
di: Lan, HaoTian
Pubblicazione: (2025)
di: Lan, HaoTian
Pubblicazione: (2025)
$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
di: Fan, Yingqi, et al.
Pubblicazione: (2025)
di: Fan, Yingqi, et al.
Pubblicazione: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
di: Heyward, Joseph, et al.
Pubblicazione: (2024)
di: Heyward, Joseph, et al.
Pubblicazione: (2024)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
di: Sarto, Sara, et al.
Pubblicazione: (2025)
di: Sarto, Sara, et al.
Pubblicazione: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
di: Jain, Jitesh, et al.
Pubblicazione: (2024)
di: Jain, Jitesh, et al.
Pubblicazione: (2024)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
di: Lu, Jinda, et al.
Pubblicazione: (2026)
di: Lu, Jinda, et al.
Pubblicazione: (2026)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
di: Dai, Haocheng, et al.
Pubblicazione: (2024)
di: Dai, Haocheng, et al.
Pubblicazione: (2024)
DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
di: Luo, Run, et al.
Pubblicazione: (2024)
di: Luo, Run, et al.
Pubblicazione: (2024)
Documenti analoghi
-
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
di: Mittal, Chinmay, et al.
Pubblicazione: (2024) -
Give me a hint: Can LLMs take a hint to solve math problems?
di: Agrawal, Vansh, et al.
Pubblicazione: (2024) -
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
di: Huang, Haoyu, et al.
Pubblicazione: (2026) -
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
di: Ma, David, et al.
Pubblicazione: (2025) -
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
di: Li, Yifan, et al.
Pubblicazione: (2025)