VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Tianyu, Wang, Suyuchen, Li, Lu, Zhang, Ge, Taslakian, Perouz, Rajeswar, Sai, Fu, Jie, Liu, Bang, Bengio, Yoshua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation
by: Li, Lu, et al.
Published: (2024)
by: Li, Lu, et al.
Published: (2024)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
by: Lo, Chung-Hsiang, et al.
Published: (2026)
by: Lo, Chung-Hsiang, et al.
Published: (2026)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024)
by: Monteiro, Joao, et al.
Published: (2024)
System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
by: Wang, Xiaoqiang, et al.
Published: (2025)
by: Wang, Xiaoqiang, et al.
Published: (2025)
Baking Symmetry into GFlowNets
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
Machine learning and information theory concepts towards an AI Mathematician
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
by: Tsirigotis, Christos, et al.
Published: (2025)
by: Tsirigotis, Christos, et al.
Published: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
by: Wang, Xiaoqiang, et al.
Published: (2026)
by: Wang, Xiaoqiang, et al.
Published: (2026)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
by: Wang, Suyuchen, et al.
Published: (2024)
by: Wang, Suyuchen, et al.
Published: (2024)
Vision Transformer for Robust Occluded Person Reidentification in Complex Surveillance Scenes
by: Li, Bo, et al.
Published: (2025)
by: Li, Bo, et al.
Published: (2025)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR
by: Li, Zhenyang, et al.
Published: (2024)
by: Li, Zhenyang, et al.
Published: (2024)
STRICT: Stress Test of Rendering Images Containing Text
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026)
by: Rodriguez, Juan, et al.
Published: (2026)
Generative Recursive Reasoning
by: Baek, Junyeob, et al.
Published: (2026)
by: Baek, Junyeob, et al.
Published: (2026)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
by: Qi, Yukun, et al.
Published: (2025)
by: Qi, Yukun, et al.
Published: (2025)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
by: Rodriguez, Juan, et al.
Published: (2024)
by: Rodriguez, Juan, et al.
Published: (2024)
Expected flow networks in stochastic environments and two-player zero-sum games
by: Jiralerspong, Marco, et al.
Published: (2023)
by: Jiralerspong, Marco, et al.
Published: (2023)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
by: Pothiraj, Atin, et al.
Published: (2025)
by: Pothiraj, Atin, et al.
Published: (2025)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
by: Rodriguez, Juan A., et al.
Published: (2025)
by: Rodriguez, Juan A., et al.
Published: (2025)
Self-Evolving Curriculum for LLM Reasoning
by: Chen, Xiaoyin, et al.
Published: (2025)
by: Chen, Xiaoyin, et al.
Published: (2025)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
Local Search GFlowNets
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
Efficient Causal Graph Discovery Using Large Language Models
by: Jiralerspong, Thomas, et al.
Published: (2024)
by: Jiralerspong, Thomas, et al.
Published: (2024)
R$^3$Mem: Bridging Memory Retention and Retrieval via Reversible Compression
by: Wang, Xiaoqiang, et al.
Published: (2025)
by: Wang, Xiaoqiang, et al.
Published: (2025)
Hierarchical Retrieval at Scale: Bridging Transparency and Efficiency
by: Gupta, Shubham, et al.
Published: (2025)
by: Gupta, Shubham, et al.
Published: (2025)
Learning to Defer for Causal Discovery with Imperfect Experts
by: Clivio, Oscar, et al.
Published: (2025)
by: Clivio, Oscar, et al.
Published: (2025)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
by: Arefin, Md Rifat, et al.
Published: (2024)
by: Arefin, Md Rifat, et al.
Published: (2024)
Structure Language Models for Protein Conformation Generation
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
Expert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and Characterization
by: Liu, Shengchao, et al.
Published: (2025)
by: Liu, Shengchao, et al.
Published: (2025)
Similar Items
-
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
by: Zhang, Tianyu, et al.
Published: (2025) -
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025) -
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025) -
MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation
by: Li, Lu, et al.
Published: (2024) -
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)