A Study of Commonsense Reasoning over Visual Object Properties
Fuente:
arXiv
Saved in:
| Main Authors: | Kolari, Abhishek, Khojasteh, Mohammadhossein, Jiang, Yifan, Hengst, Floris den, Ilievski, Filip |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025)
by: Liu, Huabin, et al.
Published: (2025)
COLUMBUS: Evaluating COgnitive Lateral Understanding through Multiple-choice reBUSes
by: Kraaijveld, Koen, et al.
Published: (2024)
by: Kraaijveld, Koen, et al.
Published: (2024)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
Enhancing Structural Mapping with LLM-derived Abstractions for Analogical Reasoning in Narratives
by: Khojasteh, Mohammadhossein, et al.
Published: (2026)
by: Khojasteh, Mohammadhossein, et al.
Published: (2026)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
by: Yu, Jiaao, et al.
Published: (2025)
by: Yu, Jiaao, et al.
Published: (2025)
Augmented Commonsense Knowledge for Remote Object Grounding
by: Mohammadi, Bahram, et al.
Published: (2024)
by: Mohammadi, Bahram, et al.
Published: (2024)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
by: Chen, Jiali, et al.
Published: (2024)
by: Chen, Jiali, et al.
Published: (2024)
VCD: A Dataset for Visual Commonsense Discovery in Images
by: Shen, Xiangqing, et al.
Published: (2024)
by: Shen, Xiangqing, et al.
Published: (2024)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
by: Zhou, Kaiwen, et al.
Published: (2023)
by: Zhou, Kaiwen, et al.
Published: (2023)
DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
by: Park, Jun-Hyung, et al.
Published: (2024)
by: Park, Jun-Hyung, et al.
Published: (2024)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
SG-Tailor: Inter-Object Commonsense Relationship Reasoning for Scene Graph Manipulation
by: Shang, Haoliang, et al.
Published: (2025)
by: Shang, Haoliang, et al.
Published: (2025)
Exploring Perceptual Limitation of Multimodal Large Language Models
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
Beyond Visual Appearances: Privacy-sensitive Objects Identification via Hybrid Graph Reasoning
by: Jiang, Zhuohang, et al.
Published: (2024)
by: Jiang, Zhuohang, et al.
Published: (2024)
Towards Commonsense Knowledge based Fuzzy Systems for Supporting Size-Related Fine-Grained Object Detection
by: Zhang, Pu, et al.
Published: (2023)
by: Zhang, Pu, et al.
Published: (2023)
Transparent Visual Reasoning via Object-Centric Agent Collaboration
by: Teoh, Benjamin, et al.
Published: (2025)
by: Teoh, Benjamin, et al.
Published: (2025)
Object Isolated Attention for Consistent Story Visualization
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
by: Jiang, Yankai, et al.
Published: (2026)
by: Jiang, Yankai, et al.
Published: (2026)
Visual Grounding for Object-Level Generalization in Reinforcement Learning
by: Jiang, Haobin, et al.
Published: (2024)
by: Jiang, Haobin, et al.
Published: (2024)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
by: Fu, Xingyu, et al.
Published: (2024)
by: Fu, Xingyu, et al.
Published: (2024)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
by: Zheng, Kaizhi, et al.
Published: (2022)
by: Zheng, Kaizhi, et al.
Published: (2022)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
by: Lv, Guannan, et al.
Published: (2026)
by: Lv, Guannan, et al.
Published: (2026)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
by: Chen, Kesheng, et al.
Published: (2026)
by: Chen, Kesheng, et al.
Published: (2026)
COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
by: Taoudi-Benchekroun, Yassine, et al.
Published: (2025)
by: Taoudi-Benchekroun, Yassine, et al.
Published: (2025)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
Reasoning over the Behaviour of Objects in Video-Clips for Adverb-Type Recognition
by: Seshadri, Amrit Diggavi, et al.
Published: (2023)
by: Seshadri, Amrit Diggavi, et al.
Published: (2023)
MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning
by: Jiang, Yifan, et al.
Published: (2024)
by: Jiang, Yifan, et al.
Published: (2024)
Enhancing Scene Graph Generation with Hierarchical Relationships and Commonsense Knowledge
by: Jiang, Bowen, et al.
Published: (2023)
by: Jiang, Bowen, et al.
Published: (2023)
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
MITracker: Multi-View Integration for Visual Object Tracking
by: Xu, Mengjie, et al.
Published: (2025)
by: Xu, Mengjie, et al.
Published: (2025)
Adversarial Attack for RGB-Event based Visual Object Tracking
by: Chen, Qiang, et al.
Published: (2025)
by: Chen, Qiang, et al.
Published: (2025)
VideoPhy: Evaluating Physical Commonsense for Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation
by: Taioli, Francesco, et al.
Published: (2026)
by: Taioli, Francesco, et al.
Published: (2026)
A Study of Failure Modes in Two-Stage Human-Object Interaction Detection
by: Wang, Lemeng, et al.
Published: (2026)
by: Wang, Lemeng, et al.
Published: (2026)
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
by: Tiwari, Sparsh, et al.
Published: (2026)
by: Tiwari, Sparsh, et al.
Published: (2026)
Predictive Reasoning with Augmented Anomaly Contrastive Learning for Compositional Visual Relations
by: Li, Chengtai, et al.
Published: (2026)
by: Li, Chengtai, et al.
Published: (2026)
Event Stream-based Visual Object Tracking: HDETrack V2 and A High-Definition Benchmark
by: Wang, Shiao, et al.
Published: (2025)
by: Wang, Shiao, et al.
Published: (2025)
Towards Low-Latency Event Stream-based Visual Object Tracking: A Slow-Fast Approach
by: Wang, Shiao, et al.
Published: (2025)
by: Wang, Shiao, et al.
Published: (2025)
Similar Items
-
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025) -
COLUMBUS: Evaluating COgnitive Lateral Understanding through Multiple-choice reBUSes
by: Kraaijveld, Koen, et al.
Published: (2024) -
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025) -
Enhancing Structural Mapping with LLM-derived Abstractions for Analogical Reasoning in Narratives
by: Khojasteh, Mohammadhossein, et al.
Published: (2026) -
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
by: Yu, Jiaao, et al.
Published: (2025)