The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Hao, Liu, Jicheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HuMoCon: Concept Discovery for Human Motion Understanding
por: Fang, Qihang, et al.
Publicado: (2025)
por: Fang, Qihang, et al.
Publicado: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
por: Tu, Songjun, et al.
Publicado: (2025)
por: Tu, Songjun, et al.
Publicado: (2025)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
por: Parcalabescu, Letitia, et al.
Publicado: (2021)
por: Parcalabescu, Letitia, et al.
Publicado: (2021)
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
por: Ko, Hanbin, et al.
Publicado: (2025)
por: Ko, Hanbin, et al.
Publicado: (2025)
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
por: Parcalabescu, Letitia, et al.
Publicado: (2022)
por: Parcalabescu, Letitia, et al.
Publicado: (2022)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
por: Cui, Shaoyang, et al.
Publicado: (2026)
por: Cui, Shaoyang, et al.
Publicado: (2026)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
por: Yanambakkam, Hemanth Teja, et al.
Publicado: (2025)
por: Yanambakkam, Hemanth Teja, et al.
Publicado: (2025)
Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
por: Ji, Yikun, et al.
Publicado: (2025)
por: Ji, Yikun, et al.
Publicado: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
por: Shah, Nisarg A., et al.
Publicado: (2025)
por: Shah, Nisarg A., et al.
Publicado: (2025)
VGA: Vision GUI Assistant -- Minimizing Hallucinations through Image-Centric Fine-Tuning
por: Meng, Ziyang, et al.
Publicado: (2024)
por: Meng, Ziyang, et al.
Publicado: (2024)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
GroundCap: A Visually Grounded Image Captioning Dataset
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
por: Freitas, Diogo, et al.
Publicado: (2025)
por: Freitas, Diogo, et al.
Publicado: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
por: He, Wei
Publicado: (2026)
por: He, Wei
Publicado: (2026)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
por: Ji, Yikun, et al.
Publicado: (2025)
por: Ji, Yikun, et al.
Publicado: (2025)
Banana Ripeness Level Classification using a Simple CNN Model Trained with Real and Synthetic Datasets
por: Chuquimarca, Luis, et al.
Publicado: (2025)
por: Chuquimarca, Luis, et al.
Publicado: (2025)
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
por: Dehghani, Mahshid, et al.
Publicado: (2024)
por: Dehghani, Mahshid, et al.
Publicado: (2024)
Does CLIP perceive art the same way we do?
por: Asperti, Andrea, et al.
Publicado: (2025)
por: Asperti, Andrea, et al.
Publicado: (2025)
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
por: Koukounas, Andreas, et al.
Publicado: (2024)
por: Koukounas, Andreas, et al.
Publicado: (2024)
RONA: Pragmatically Diverse Image Captioning with Coherence Relations
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024)
por: Ma, Yueen, et al.
Publicado: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)
por: Dai, Song, et al.
Publicado: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
por: Agarwal, Amit, et al.
Publicado: (2025)
por: Agarwal, Amit, et al.
Publicado: (2025)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
por: Bucher, Martin JJ., et al.
Publicado: (2025)
por: Bucher, Martin JJ., et al.
Publicado: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
por: Anh, Duy Le Dinh, et al.
Publicado: (2024)
por: Anh, Duy Le Dinh, et al.
Publicado: (2024)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
por: Kim, Soyeon, et al.
Publicado: (2026)
por: Kim, Soyeon, et al.
Publicado: (2026)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
por: Sanders, Kate, et al.
Publicado: (2024)
por: Sanders, Kate, et al.
Publicado: (2024)
OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment
por: Zhao, Weiyi, et al.
Publicado: (2025)
por: Zhao, Weiyi, et al.
Publicado: (2025)
Conterfactual Generative Zero-Shot Semantic Segmentation
por: Shen, Feihong, et al.
Publicado: (2021)
por: Shen, Feihong, et al.
Publicado: (2021)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
por: Ouyang, Rongxin, et al.
Publicado: (2024)
por: Ouyang, Rongxin, et al.
Publicado: (2024)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
por: Lyu, Wenbo, et al.
Publicado: (2025)
por: Lyu, Wenbo, et al.
Publicado: (2025)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
por: Fu, Tianyu, et al.
Publicado: (2024)
por: Fu, Tianyu, et al.
Publicado: (2024)
Text-to-Events: Synthetic Event Camera Streams from Conditional Text Input
por: Ott, Joachim, et al.
Publicado: (2024)
por: Ott, Joachim, et al.
Publicado: (2024)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
por: Deichler, Anna, et al.
Publicado: (2026)
por: Deichler, Anna, et al.
Publicado: (2026)
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
por: Malikussaid, et al.
Publicado: (2026)
por: Malikussaid, et al.
Publicado: (2026)
Ejemplares similares
-
HuMoCon: Concept Discovery for Human Motion Understanding
por: Fang, Qihang, et al.
Publicado: (2025) -
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
por: Tu, Songjun, et al.
Publicado: (2025) -
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
por: Parcalabescu, Letitia, et al.
Publicado: (2021) -
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
por: Ko, Hanbin, et al.
Publicado: (2025) -
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
por: Parcalabescu, Letitia, et al.
Publicado: (2022)