Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xinxin, Xu, Zhaopan, Li, Ming, Wang, Kai, Lee, Yong Jae, Shang, Yuzhang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
di: Zhang, Daoan, et al.
Pubblicazione: (2025)
di: Zhang, Daoan, et al.
Pubblicazione: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
di: Guo, Longteng, et al.
Pubblicazione: (2026)
di: Guo, Longteng, et al.
Pubblicazione: (2026)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
di: Xu, Zhaopan, et al.
Pubblicazione: (2025)
di: Xu, Zhaopan, et al.
Pubblicazione: (2025)
VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning
di: Ji, Yuheng, et al.
Pubblicazione: (2025)
di: Ji, Yuheng, et al.
Pubblicazione: (2025)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
di: Liu, Xinxin, et al.
Pubblicazione: (2026)
di: Liu, Xinxin, et al.
Pubblicazione: (2026)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
di: Lee, Naeun, et al.
Pubblicazione: (2026)
di: Lee, Naeun, et al.
Pubblicazione: (2026)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
di: Shang, Yuzhang, et al.
Pubblicazione: (2024)
di: Shang, Yuzhang, et al.
Pubblicazione: (2024)
Benchmarking and Analyzing Generative Data for Visual Recognition
di: Li, Bo, et al.
Pubblicazione: (2023)
di: Li, Bo, et al.
Pubblicazione: (2023)
KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models
di: Li, Sibo, et al.
Pubblicazione: (2025)
di: Li, Sibo, et al.
Pubblicazione: (2025)
ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
di: Wang, Lihong, et al.
Pubblicazione: (2025)
di: Wang, Lihong, et al.
Pubblicazione: (2025)
ViTCN: Vision Transformer Contrastive Network For Reasoning
di: Song, Bo, et al.
Pubblicazione: (2024)
di: Song, Bo, et al.
Pubblicazione: (2024)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
di: Yu, Zhuoran, et al.
Pubblicazione: (2025)
di: Yu, Zhuoran, et al.
Pubblicazione: (2025)
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
di: Bi, Jing, et al.
Pubblicazione: (2025)
di: Bi, Jing, et al.
Pubblicazione: (2025)
MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
di: Hsin-Ying, Lee, et al.
Pubblicazione: (2026)
di: Hsin-Ying, Lee, et al.
Pubblicazione: (2026)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
di: Qiu, Yansheng, et al.
Pubblicazione: (2025)
di: Qiu, Yansheng, et al.
Pubblicazione: (2025)
Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation
di: Wang, Junjie, et al.
Pubblicazione: (2026)
di: Wang, Junjie, et al.
Pubblicazione: (2026)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
di: Li, Kunyang, et al.
Pubblicazione: (2026)
di: Li, Kunyang, et al.
Pubblicazione: (2026)
ChartBench: A Benchmark for Complex Visual Reasoning in Charts
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2023)
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2023)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
di: Zhu, Chen, et al.
Pubblicazione: (2025)
di: Zhu, Chen, et al.
Pubblicazione: (2025)
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
di: Jahangard, Simindokht, et al.
Pubblicazione: (2025)
di: Jahangard, Simindokht, et al.
Pubblicazione: (2025)
T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
di: Sun, Kaiyue, et al.
Pubblicazione: (2025)
di: Sun, Kaiyue, et al.
Pubblicazione: (2025)
Latent Implicit Visual Reasoning
di: Li, Kelvin, et al.
Pubblicazione: (2025)
di: Li, Kelvin, et al.
Pubblicazione: (2025)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
di: Lee, Yeonkyung, et al.
Pubblicazione: (2026)
di: Lee, Yeonkyung, et al.
Pubblicazione: (2026)
What if? Emulative Simulation with World Models for Situated Reasoning
di: Liu, Ruiping, et al.
Pubblicazione: (2026)
di: Liu, Ruiping, et al.
Pubblicazione: (2026)
RVTBench: A Benchmark for Visual Reasoning Tasks
di: Shen, Yiqing, et al.
Pubblicazione: (2025)
di: Shen, Yiqing, et al.
Pubblicazione: (2025)
ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models
di: Hei, Xusen, et al.
Pubblicazione: (2025)
di: Hei, Xusen, et al.
Pubblicazione: (2025)
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
di: Han, Haonan, et al.
Pubblicazione: (2026)
di: Han, Haonan, et al.
Pubblicazione: (2026)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
di: Wu, Fengyi, et al.
Pubblicazione: (2025)
di: Wu, Fengyi, et al.
Pubblicazione: (2025)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
di: Wu, Linquan, et al.
Pubblicazione: (2026)
di: Wu, Linquan, et al.
Pubblicazione: (2026)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
di: Zhou, Kaiwen, et al.
Pubblicazione: (2023)
di: Zhou, Kaiwen, et al.
Pubblicazione: (2023)
Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement
di: Zheng, Jiakun, et al.
Pubblicazione: (2026)
di: Zheng, Jiakun, et al.
Pubblicazione: (2026)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval
di: Chen, Dongyang, et al.
Pubblicazione: (2026)
di: Chen, Dongyang, et al.
Pubblicazione: (2026)
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
di: Yang, Zheyuan, et al.
Pubblicazione: (2026)
di: Yang, Zheyuan, et al.
Pubblicazione: (2026)
ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
di: Wu, Jay Zhangjie, et al.
Pubblicazione: (2025)
di: Wu, Jay Zhangjie, et al.
Pubblicazione: (2025)
FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
di: Jiang, Yue, et al.
Pubblicazione: (2025)
di: Jiang, Yue, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
di: Zhang, Daoan, et al.
Pubblicazione: (2025) -
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
di: Guo, Longteng, et al.
Pubblicazione: (2026) -
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
di: Xu, Zhaopan, et al.
Pubblicazione: (2025) -
VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning
di: Ji, Yuheng, et al.
Pubblicazione: (2025) -
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)