RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Meng-Hao, Chu, Xuanyu, Yang, Qianrui, Mo, Zhe-Han, Shen, Yiqing, Li, Pei-lin, Lin, Xinjie, Zhang, Jinnian, Chen, Xin-Sheng, Zhang, Yi, Nakayama, Kiyohiro, Geng, Zhengyang, Peng, Houwen, Hu, Han, Hu, Shi-Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
by: Guo, Meng-Hao, et al.
Published: (2025)
by: Guo, Meng-Hao, et al.
Published: (2025)
SafeRBench: Dissecting the Reasoning Safety of Large Language Models
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
by: Sun, Hai-Long, et al.
Published: (2025)
by: Sun, Hai-Long, et al.
Published: (2025)
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation
by: Chen, Yi-Pei, et al.
Published: (2024)
by: Chen, Yi-Pei, et al.
Published: (2024)
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
by: Ma, Qingchuan, et al.
Published: (2026)
by: Ma, Qingchuan, et al.
Published: (2026)
A method for estimating forest carbon storage distribution density via artificial intelligence generated content model
by: Yu, Zhenyu, et al.
Published: (2025)
by: Yu, Zhenyu, et al.
Published: (2025)
Estimating forest carbon stocks from high-resolution remote sensing imagery by reducing domain shift with style transfer
by: Yu, Zhenyu, et al.
Published: (2025)
by: Yu, Zhenyu, et al.
Published: (2025)
ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws
by: Li, Ruihang, et al.
Published: (2024)
by: Li, Ruihang, et al.
Published: (2024)
Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion
by: Hu, Xuanyu
Published: (2025)
by: Hu, Xuanyu
Published: (2025)
Xwin-LM: Strong and Scalable Alignment Practice for LLMs
by: Ni, Bolin, et al.
Published: (2024)
by: Ni, Bolin, et al.
Published: (2024)
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
by: Huang, Benhao, et al.
Published: (2026)
by: Huang, Benhao, et al.
Published: (2026)
Common 7B Language Models Already Possess Strong Math Capabilities
by: Li, Chen, et al.
Published: (2024)
by: Li, Chen, et al.
Published: (2024)
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
by: Ackermann, Jan, et al.
Published: (2025)
by: Ackermann, Jan, et al.
Published: (2025)
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2025)
by: Zhou, Qianrui, et al.
Published: (2025)
Multi-modal Learnable Queries for Image Aesthetics Assessment
by: Xiong, Zhiwei, et al.
Published: (2024)
by: Xiong, Zhiwei, et al.
Published: (2024)
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
by: Zhao, Shijie, et al.
Published: (2025)
by: Zhao, Shijie, et al.
Published: (2025)
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Impact of Returning to Hometown Entrepreneurship Policy on Rural Intergenerational Educational Mobility
by: Pei Zeng, et al.
Published: (2026)
by: Pei Zeng, et al.
Published: (2026)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
by: Chen, Guanjie, et al.
Published: (2024)
by: Chen, Guanjie, et al.
Published: (2024)
Some Discontinuous Galerkin Schemes for Korteweg‐De Vries Equations: Error Estimates and Application
by: Zhilei Wang, et al.
Published: (2025)
by: Zhilei Wang, et al.
Published: (2025)
ProvNeRF: Modeling per Point Provenance in NeRFs as a Stochastic Field
by: Nakayama, Kiyohiro, et al.
Published: (2024)
by: Nakayama, Kiyohiro, et al.
Published: (2024)
GenAnalysis: Joint Shape Analysis by Learning Man-Made Shape Generators with Deformation Regularizations
by: Yang, Yuezhi, et al.
Published: (2025)
by: Yang, Yuezhi, et al.
Published: (2025)
Covariance spectrum in nonlinear recurrent neural networks
by: Shen, Xuanyu, et al.
Published: (2025)
by: Shen, Xuanyu, et al.
Published: (2025)
Bilinear structures of the fourth-order lattice Gel'fand-Dikii equations
by: Zhao, Song-lin, et al.
Published: (2025)
by: Zhao, Song-lin, et al.
Published: (2025)
MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning
by: Tie, Guiyao, et al.
Published: (2025)
by: Tie, Guiyao, et al.
Published: (2025)
TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
by: Hu, Yujie, et al.
Published: (2025)
by: Hu, Yujie, et al.
Published: (2025)
Improved implicit diffusion model with knowledge distillation to estimate the spatial distribution density of carbon stock in remote sensing imagery
by: Yu, Zhenyu, et al.
Published: (2024)
by: Yu, Zhenyu, et al.
Published: (2024)
CoMM: Collaborative Multi-Agent, Multi-Reasoning-Path Prompting for Complex Problem Solving
by: Chen, Pei, et al.
Published: (2024)
by: Chen, Pei, et al.
Published: (2024)
Analysis of Radiation Level and Estimation of Protection Distance of γ Mobile Flaw Detection Source
by: Liu, Zhihui, et al.
Published: (2025)
by: Liu, Zhihui, et al.
Published: (2025)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2026)
by: Zhou, Qianrui, et al.
Published: (2026)
Effects of Non-local Pseudopotentials on the Electrical and Thermal Transport Properties of Aluminum: A Density Functional Theory Study
by: Liu, Qianrui, et al.
Published: (2024)
by: Liu, Qianrui, et al.
Published: (2024)
Garment Particles: A 2D--3D Symmetric Garment Representation for Generation and Editing
by: Nakayama, Kiyohiro, et al.
Published: (2026)
by: Nakayama, Kiyohiro, et al.
Published: (2026)
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
by: Hu, Wenbin, et al.
Published: (2025)
by: Hu, Wenbin, et al.
Published: (2025)
Cohesive Conversations: Enhancing Authenticity in Multi-Agent Simulated Dialogues
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
Exploring and Controlling Diversity in LLM-Agent Conversation
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
by: Zhang, Yiqing, et al.
Published: (2026)
by: Zhang, Yiqing, et al.
Published: (2026)
OSoRA: Output-Dimension and Singular-Value Initialized Low-Rank Adaptation
by: Han, Jialong, et al.
Published: (2025)
by: Han, Jialong, et al.
Published: (2025)
TransGPT: Multi-modal Generative Pre-trained Transformer for Transportation
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
Similar Items
-
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
by: Guo, Meng-Hao, et al.
Published: (2025) -
SafeRBench: Dissecting the Reasoning Safety of Large Language Models
by: Gao, Xin, et al.
Published: (2025) -
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
by: Sun, Hai-Long, et al.
Published: (2025) -
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation
by: Chen, Yi-Pei, et al.
Published: (2024) -
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024)