COLUMBUS: Evaluating COgnitive Lateral Understanding through Multiple-choice reBUSes
Fuente:
arXiv
Saved in:
| Main Authors: | Kraaijveld, Koen, Jiang, Yifan, Ma, Kaixin, Ilievski, Filip |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Study of Commonsense Reasoning over Visual Object Properties
by: Kolari, Abhishek, et al.
Published: (2025)
by: Kolari, Abhishek, et al.
Published: (2025)
MovieCORE: COgnitive REasoning in Movies
by: Faure, Gueter Josmy, et al.
Published: (2025)
by: Faure, Gueter Josmy, et al.
Published: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025)
by: Liu, Huabin, et al.
Published: (2025)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning
by: Jiang, Yifan, et al.
Published: (2024)
by: Jiang, Yifan, et al.
Published: (2024)
SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense
by: Jiang, Yifan, et al.
Published: (2024)
by: Jiang, Yifan, et al.
Published: (2024)
Exploring Perceptual Limitation of Multimodal Large Language Models
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
AskChart: Universal Chart Understanding through Textual Enhancement
by: Yang, Xudong, et al.
Published: (2024)
by: Yang, Xudong, et al.
Published: (2024)
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
by: Fu, Teng, et al.
Published: (2025)
by: Fu, Teng, et al.
Published: (2025)
Not All Samples Should Be Utilized Equally: Towards Understanding and Improving Dataset Distillation
by: Wang, Shaobo, et al.
Published: (2024)
by: Wang, Shaobo, et al.
Published: (2024)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
by: Jiao, Yang, et al.
Published: (2025)
by: Jiao, Yang, et al.
Published: (2025)
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
by: Xu, Yifan, et al.
Published: (2025)
by: Xu, Yifan, et al.
Published: (2025)
Universal Visuo-Tactile Video Understanding for Embodied Interaction
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
by: Hong, Jack, et al.
Published: (2025)
by: Hong, Jack, et al.
Published: (2025)
FIRE: Food Image to REcipe generation
by: Chhikara, Prateek, et al.
Published: (2023)
by: Chhikara, Prateek, et al.
Published: (2023)
Steal Now and Attack Later: Evaluating Robustness of Object Detection against Black-box Adversarial Attacks
by: Chen, Erh-Chung, et al.
Published: (2024)
by: Chen, Erh-Chung, et al.
Published: (2024)
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
by: Zheng, Yikai, et al.
Published: (2026)
by: Zheng, Yikai, et al.
Published: (2026)
Into the Fog: Evaluating Robustness of Multiple Object Tracking
by: Kirillova, Nadezda, et al.
Published: (2024)
by: Kirillova, Nadezda, et al.
Published: (2024)
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
by: Chen, Qi, et al.
Published: (2024)
by: Chen, Qi, et al.
Published: (2024)
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
CORE-ReID: Comprehensive Optimization and Refinement through Ensemble fusion in Domain Adaptation for person re-identification
by: Nguyen, Trinh Quoc, et al.
Published: (2025)
by: Nguyen, Trinh Quoc, et al.
Published: (2025)
Understanding Visual Feature Reliance through the Lens of Complexity
by: Fel, Thomas, et al.
Published: (2024)
by: Fel, Thomas, et al.
Published: (2024)
VALUED -- Vision and Logical Understanding Evaluation Dataset
by: Saha, Soumadeep, et al.
Published: (2023)
by: Saha, Soumadeep, et al.
Published: (2023)
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
by: Peng, Yingzhe, et al.
Published: (2024)
by: Peng, Yingzhe, et al.
Published: (2024)
Leveraging Gait Patterns as Biomarkers: An attention-guided Deep Multiple Instance Learning Network for Scoliosis Classification
by: Li, Haiqing, et al.
Published: (2025)
by: Li, Haiqing, et al.
Published: (2025)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations
by: Belmecheri, Nassim, et al.
Published: (2024)
by: Belmecheri, Nassim, et al.
Published: (2024)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025)
by: Ma, Chuofan, et al.
Published: (2025)
Evaluating Compositional Scene Understanding in Multimodal Generative Models
by: Fu, Shuhao, et al.
Published: (2025)
by: Fu, Shuhao, et al.
Published: (2025)
Benchmarking Pathology Foundation Models for Spatial Domain Understanding
by: Zhao, Bokai, et al.
Published: (2026)
by: Zhao, Bokai, et al.
Published: (2026)
Evaluating Low-Light Image Enhancement Across Multiple Intensity Levels
by: Pilligua, Maria, et al.
Published: (2025)
by: Pilligua, Maria, et al.
Published: (2025)
Close the Sim2real Gap via Physically-based Structured Light Synthetic Data Simulation
by: Bai, Kaixin, et al.
Published: (2024)
by: Bai, Kaixin, et al.
Published: (2024)
Adapting the re-ID challenge for static sensors
by: Sundaresan, Avirath, et al.
Published: (2024)
by: Sundaresan, Avirath, et al.
Published: (2024)
Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation
by: Cai, Deli, et al.
Published: (2026)
by: Cai, Deli, et al.
Published: (2026)
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
by: Pitta, Elena, et al.
Published: (2025)
by: Pitta, Elena, et al.
Published: (2025)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
by: Pani, Anupam, et al.
Published: (2025)
by: Pani, Anupam, et al.
Published: (2025)
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
by: Li, Yuting, et al.
Published: (2025)
by: Li, Yuting, et al.
Published: (2025)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
by: Peng, Tianhao, et al.
Published: (2025)
by: Peng, Tianhao, et al.
Published: (2025)
Similar Items
-
A Study of Commonsense Reasoning over Visual Object Properties
by: Kolari, Abhishek, et al.
Published: (2025) -
MovieCORE: COgnitive REasoning in Movies
by: Faure, Gueter Josmy, et al.
Published: (2025) -
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025) -
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025) -
MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning
by: Jiang, Yifan, et al.
Published: (2024)