CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jianrui, Cai, Mu, Xie, Tengyang, Lee, Yong Jae |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
by: Zhang, Jianrui, et al.
Published: (2024)
by: Zhang, Jianrui, et al.
Published: (2024)
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
by: Bharti, Shubham, et al.
Published: (2024)
by: Bharti, Shubham, et al.
Published: (2024)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
by: Zhang, Le, et al.
Published: (2023)
by: Zhang, Le, et al.
Published: (2023)
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
by: Zou, Bocheng, et al.
Published: (2024)
by: Zou, Bocheng, et al.
Published: (2024)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
by: Shang, Yuzhang, et al.
Published: (2024)
by: Shang, Yuzhang, et al.
Published: (2024)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
by: Yu, Zhuoran, et al.
Published: (2025)
by: Yu, Zhuoran, et al.
Published: (2025)
Matryoshka Multimodal Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
by: Wei, Yana, et al.
Published: (2025)
by: Wei, Yana, et al.
Published: (2025)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
by: Zheng, Ruobing, et al.
Published: (2026)
by: Zheng, Ruobing, et al.
Published: (2026)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
by: Poppi, Tobia, et al.
Published: (2026)
by: Poppi, Tobia, et al.
Published: (2026)
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
by: Zha, Yuheng, et al.
Published: (2025)
by: Zha, Yuheng, et al.
Published: (2025)
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
Scalable Vision Language Model Training via High Quality Data Curation
by: Dong, Hongyuan, et al.
Published: (2025)
by: Dong, Hongyuan, et al.
Published: (2025)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
Redefining <Creative> in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generation
by: Feng, Fu, et al.
Published: (2024)
by: Feng, Fu, et al.
Published: (2024)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Class-Agnostic Visio-Temporal Scene Sketch Semantic Segmentation
by: Kütük, Aleyna, et al.
Published: (2024)
by: Kütük, Aleyna, et al.
Published: (2024)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
by: Rassin, Royi, et al.
Published: (2023)
by: Rassin, Royi, et al.
Published: (2023)
Integrating GAN and Texture Synthesis for Enhanced Road Damage Detection
by: Chen, Tengyang, et al.
Published: (2023)
by: Chen, Tengyang, et al.
Published: (2023)
Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus
by: Guillen-Perez, Antonio
Published: (2025)
by: Guillen-Perez, Antonio
Published: (2025)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
by: Mu, Kuan-Chen, et al.
Published: (2024)
by: Mu, Kuan-Chen, et al.
Published: (2024)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
by: Joshi, Siddharth, et al.
Published: (2025)
by: Joshi, Siddharth, et al.
Published: (2025)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
by: Li, Can, et al.
Published: (2025)
by: Li, Can, et al.
Published: (2025)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
by: Cai, Dexian, et al.
Published: (2025)
by: Cai, Dexian, et al.
Published: (2025)
Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
Rethinking Radiology Report Generation via Causal Inspired Counterfactual Augmentation
by: Song, Xiao, et al.
Published: (2023)
by: Song, Xiao, et al.
Published: (2023)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
by: Loginova, Olga, et al.
Published: (2025)
by: Loginova, Olga, et al.
Published: (2025)
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models
by: Zhang, Letian, et al.
Published: (2023)
by: Zhang, Letian, et al.
Published: (2023)
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
by: Bhattacharya, Amartya
Published: (2026)
by: Bhattacharya, Amartya
Published: (2026)
InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
by: Ganguly, Debargha, et al.
Published: (2025)
by: Ganguly, Debargha, et al.
Published: (2025)
Similar Items
-
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
by: Zhang, Jianrui, et al.
Published: (2024) -
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
by: Oh, Youngtaek, et al.
Published: (2024) -
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
by: Bharti, Shubham, et al.
Published: (2024) -
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
by: Zhang, Le, et al.
Published: (2023) -
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
by: Zou, Bocheng, et al.
Published: (2024)