BLINK: Multimodal Large Language Models Can See but Not Perceive
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Xingyu, Hu, Yushi, Li, Bangzheng, Feng, Yu, Wang, Haoyu, Lin, Xudong, Roth, Dan, Smith, Noah A., Ma, Wei-Chiu, Krishna, Ranjay |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
por: Hu, Yushi, et al.
Publicado: (2024)
por: Hu, Yushi, et al.
Publicado: (2024)
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
por: Li, Bangzheng, et al.
Publicado: (2023)
por: Li, Bangzheng, et al.
Publicado: (2023)
FamiCom: Further Demystifying Prompts for Language Models with Task-Agnostic Performance Estimation
por: Li, Bangzheng, et al.
Publicado: (2024)
por: Li, Bangzheng, et al.
Publicado: (2024)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
por: Duggal, Shivam, et al.
Publicado: (2025)
por: Duggal, Shivam, et al.
Publicado: (2025)
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
por: Bigverdi, Mahtab, et al.
Publicado: (2025)
por: Bigverdi, Mahtab, et al.
Publicado: (2025)
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
por: Yao, Jihan, et al.
Publicado: (2025)
por: Yao, Jihan, et al.
Publicado: (2025)
DataProphet: Demystifying Supervision Data Generalization in Multimodal LLMs
por: Qi, Xuan, et al.
Publicado: (2026)
por: Qi, Xuan, et al.
Publicado: (2026)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
por: Fu, Xingyu, et al.
Publicado: (2024)
por: Fu, Xingyu, et al.
Publicado: (2024)
Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
por: Liu, Benlin, et al.
Publicado: (2024)
por: Liu, Benlin, et al.
Publicado: (2024)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
por: Hu, Yushi, et al.
Publicado: (2023)
por: Hu, Yushi, et al.
Publicado: (2023)
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
por: Fu, Xingyu, et al.
Publicado: (2025)
por: Fu, Xingyu, et al.
Publicado: (2025)
BIRD: A Trustworthy Bayesian Inference Framework for Large Language Models
por: Feng, Yu, et al.
Publicado: (2024)
por: Feng, Yu, et al.
Publicado: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
por: Zheng, Chenhao, et al.
Publicado: (2024)
por: Zheng, Chenhao, et al.
Publicado: (2024)
Can Multimodal Large Language Model Think Analogically?
por: Guo, Diandian, et al.
Publicado: (2024)
por: Guo, Diandian, et al.
Publicado: (2024)
"I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?
por: Xu, Naen, et al.
Publicado: (2026)
por: Xu, Naen, et al.
Publicado: (2026)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
por: Kamath, Amita, et al.
Publicado: (2026)
por: Kamath, Amita, et al.
Publicado: (2026)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
por: Kamath, Amita, et al.
Publicado: (2025)
por: Kamath, Amita, et al.
Publicado: (2025)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
por: Zhang, Tianyi, et al.
Publicado: (2026)
por: Zhang, Tianyi, et al.
Publicado: (2026)
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
por: Cai, Rui, et al.
Publicado: (2025)
por: Cai, Rui, et al.
Publicado: (2025)
BLINK: Behavioral Latent Modeling of NK Cell Cytotoxicity
por: Nematollahi, Iman, et al.
Publicado: (2026)
por: Nematollahi, Iman, et al.
Publicado: (2026)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
por: Lai, Zhengzhao, et al.
Publicado: (2025)
por: Lai, Zhengzhao, et al.
Publicado: (2025)
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
por: Loakman, Tyler, et al.
Publicado: (2024)
por: Loakman, Tyler, et al.
Publicado: (2024)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
por: Fu, Xingyu, et al.
Publicado: (2023)
por: Fu, Xingyu, et al.
Publicado: (2023)
The Hard Positive Truth about Vision-Language Compositionality
por: Kamath, Amita, et al.
Publicado: (2024)
por: Kamath, Amita, et al.
Publicado: (2024)
Semantic and Expressive Variation in Image Captions Across Languages
por: Ye, Andre, et al.
Publicado: (2023)
por: Ye, Andre, et al.
Publicado: (2023)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
por: Wang, Wenxuan, et al.
Publicado: (2025)
por: Wang, Wenxuan, et al.
Publicado: (2025)
Multilingual Diversity Improves Vision-Language Representations
por: Nguyen, Thao, et al.
Publicado: (2024)
por: Nguyen, Thao, et al.
Publicado: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
por: Ma, Zixian, et al.
Publicado: (2024)
por: Ma, Zixian, et al.
Publicado: (2024)
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
por: Hou, Wenjin, et al.
Publicado: (2026)
por: Hou, Wenjin, et al.
Publicado: (2026)
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
por: Cheng, Long, et al.
Publicado: (2025)
por: Cheng, Long, et al.
Publicado: (2025)
Visual Representations inside the Language Model
por: Liu, Benlin, et al.
Publicado: (2025)
por: Liu, Benlin, et al.
Publicado: (2025)
Multimodal Language Models See Better When They Look Shallower
por: Chen, Haoran, et al.
Publicado: (2025)
por: Chen, Haoran, et al.
Publicado: (2025)
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
por: Cho, Jaemin, et al.
Publicado: (2023)
por: Cho, Jaemin, et al.
Publicado: (2023)
Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language Models
por: Zhang, Baoheng, et al.
Publicado: (2026)
por: Zhang, Baoheng, et al.
Publicado: (2026)
Reinforced Visual Perception with Tools
por: Zhou, Zetong, et al.
Publicado: (2025)
por: Zhou, Zetong, et al.
Publicado: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
por: He, Zhentao, et al.
Publicado: (2025)
por: He, Zhentao, et al.
Publicado: (2025)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
por: Shen, Yixuan, et al.
Publicado: (2026)
por: Shen, Yixuan, et al.
Publicado: (2026)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
por: Wu, Jiaying, et al.
Publicado: (2025)
por: Wu, Jiaying, et al.
Publicado: (2025)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
por: Kang, Wonjune, et al.
Publicado: (2024)
por: Kang, Wonjune, et al.
Publicado: (2024)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
por: Gan, Woody Haosheng, et al.
Publicado: (2025)
por: Gan, Woody Haosheng, et al.
Publicado: (2025)
Ejemplares similares
-
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
por: Hu, Yushi, et al.
Publicado: (2024) -
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
por: Li, Bangzheng, et al.
Publicado: (2023) -
FamiCom: Further Demystifying Prompts for Language Models with Task-Agnostic Performance Estimation
por: Li, Bangzheng, et al.
Publicado: (2024) -
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
por: Duggal, Shivam, et al.
Publicado: (2025) -
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
por: Bigverdi, Mahtab, et al.
Publicado: (2025)