Seeing Culture: A Benchmark for Visual Reasoning and Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Satar, Burak, Ma, Zhixin, Irawan, Patrick A., Mulyawan, Wilfried A., Jiang, Jing, Lim, Ee-Peng, Ngo, Chong-Wah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
by: Wu, Xiongwei, et al.
Published: (2024)
by: Wu, Xiongwei, et al.
Published: (2024)
Robust Relevance Feedback for Interactive Known-Item Video Search
by: Ma, Zhixin, et al.
Published: (2025)
by: Ma, Zhixin, et al.
Published: (2025)
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
by: Ilaslan, Muhammet Furkan, et al.
Published: (2024)
by: Ilaslan, Muhammet Furkan, et al.
Published: (2024)
Class Agnostic Instance-level Descriptor for Visual Instance Search
by: Sun, Qi-Ying, et al.
Published: (2025)
by: Sun, Qi-Ying, et al.
Published: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
by: Yang, Haibo, et al.
Published: (2024)
by: Yang, Haibo, et al.
Published: (2024)
Seeing Text in the Dark: Algorithm and Benchmark
by: Xu, Chengpei, et al.
Published: (2024)
by: Xu, Chengpei, et al.
Published: (2024)
Efficient Prompt Tuning for Hierarchical Ingredient Recognition
by: Gui, Yinxuan, et al.
Published: (2025)
by: Gui, Yinxuan, et al.
Published: (2025)
Navigating Weight Prediction with Diet Diary
by: Gui, Yinxuan, et al.
Published: (2024)
by: Gui, Yinxuan, et al.
Published: (2024)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
by: Chao, Jianghan, et al.
Published: (2025)
by: Chao, Jianghan, et al.
Published: (2025)
Towards Multimodal Emotional Support Conversation Systems
by: Chu, Yuqi, et al.
Published: (2024)
by: Chu, Yuqi, et al.
Published: (2024)
Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images
by: Jiang, Yuechen, et al.
Published: (2026)
by: Jiang, Yuechen, et al.
Published: (2026)
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
by: Maji, Arijit, et al.
Published: (2025)
by: Maji, Arijit, et al.
Published: (2025)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
by: Zhang, Lei, et al.
Published: (2025)
by: Zhang, Lei, et al.
Published: (2025)
History-Guided Iterative Visual Reasoning with Self-Correction
by: Yang, Xinglong, et al.
Published: (2026)
by: Yang, Xinglong, et al.
Published: (2026)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
Language-Guided Diffusion Model for Visual Grounding
by: Chen, Sijia, et al.
Published: (2023)
by: Chen, Sijia, et al.
Published: (2023)
Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage
by: He, Ziyi, et al.
Published: (2026)
by: He, Ziyi, et al.
Published: (2026)
Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
by: Chen, Yi-Chun
Published: (2025)
by: Chen, Yi-Chun
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
by: Wang, Bingbing, et al.
Published: (2025)
by: Wang, Bingbing, et al.
Published: (2025)
MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
by: Peng, Yuezhang, et al.
Published: (2025)
by: Peng, Yuezhang, et al.
Published: (2025)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
by: Huang, Yuesheng, et al.
Published: (2025)
by: Huang, Yuesheng, et al.
Published: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
The Rhythm of Tai Chi: Revitalizing Cultural Heritage in Virtual Reality through Interactive Visuals
by: Wang, Xianghan
Published: (2025)
by: Wang, Xianghan
Published: (2025)
Causal Debiasing for Visual Commonsense Reasoning
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents
by: Ma, Ziyang, et al.
Published: (2026)
by: Ma, Ziyang, et al.
Published: (2026)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
by: Xu, Pengju, et al.
Published: (2025)
by: Xu, Pengju, et al.
Published: (2025)
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
by: Malakouti, Sina, et al.
Published: (2024)
by: Malakouti, Sina, et al.
Published: (2024)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
by: Lai, Zhengzhao, et al.
Published: (2025)
by: Lai, Zhengzhao, et al.
Published: (2025)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
by: Liao, Liwei, et al.
Published: (2025)
by: Liao, Liwei, et al.
Published: (2025)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
by: Tan, Weiting, et al.
Published: (2025)
by: Tan, Weiting, et al.
Published: (2025)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
by: Gao, Shibo, et al.
Published: (2025)
by: Gao, Shibo, et al.
Published: (2025)
Similar Items
-
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025) -
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025) -
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
by: Wu, Xiongwei, et al.
Published: (2024) -
Robust Relevance Feedback for Interactive Known-Item Video Search
by: Ma, Zhixin, et al.
Published: (2025) -
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024)