Benchmarking Large and Small MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Xuelu, Li, Yunsheng, Chen, Dongdong, Gao, Mei, Liu, Mengchen, Yuan, Junsong, Qiao, Chunming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
by: Feng, Xuelu, et al.
Published: (2025)
by: Feng, Xuelu, et al.
Published: (2025)
Pluralistic Salient Object Detection
by: Feng, Xuelu, et al.
Published: (2024)
by: Feng, Xuelu, et al.
Published: (2024)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
by: Zhu, Zixin, et al.
Published: (2024)
by: Zhu, Zixin, et al.
Published: (2024)
GeoRemover: Removing Objects and Their Causal Visual Artifacts
by: Zhu, Zixin, et al.
Published: (2025)
by: Zhu, Zixin, et al.
Published: (2025)
Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network
by: Luan, Tianyu, et al.
Published: (2025)
by: Luan, Tianyu, et al.
Published: (2025)
SRAM: Shape-Realism Alignment Metric for No Reference 3D Shape Evaluation
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
Fully Authentic Visual Question Answering Dataset from Online Communities
by: Chen, Chongyan, et al.
Published: (2023)
by: Chen, Chongyan, et al.
Published: (2023)
ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models
by: Hei, Xusen, et al.
Published: (2025)
by: Hei, Xusen, et al.
Published: (2025)
Show and Segment: Universal Medical Image Segmentation via In-Context Learning
by: Gao, Yunhe, et al.
Published: (2025)
by: Gao, Yunhe, et al.
Published: (2025)
Exploring Invariance in Images through One-way Wave Equations
by: Chen, Yinpeng, et al.
Published: (2023)
by: Chen, Yinpeng, et al.
Published: (2023)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
by: Li, Chentao, et al.
Published: (2026)
by: Li, Chentao, et al.
Published: (2026)
Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition
by: Liu, Mengyuan, et al.
Published: (2026)
by: Liu, Mengyuan, et al.
Published: (2026)
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction
by: Li, Jinnao, et al.
Published: (2026)
by: Li, Jinnao, et al.
Published: (2026)
UAVFF3D: A Geometry-Aware Benchmark for Feed-Forward UAV 3D Reconstruction
by: Yang, Xiang, et al.
Published: (2026)
by: Yang, Xiang, et al.
Published: (2026)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
by: Liu, Shaoyu, et al.
Published: (2025)
by: Liu, Shaoyu, et al.
Published: (2025)
Towards Faithful Reasoning in Comics for Small MLLMs
by: Feng, Chengcheng, et al.
Published: (2026)
by: Feng, Chengcheng, et al.
Published: (2026)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
by: Huang, Jen-Tse, et al.
Published: (2025)
by: Huang, Jen-Tse, et al.
Published: (2025)
Forecasting Future Videos from Novel Views via Disentangled 3D Scene Representation
by: Yarram, Sudhir, et al.
Published: (2024)
by: Yarram, Sudhir, et al.
Published: (2024)
UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding
by: Zhan, Yang, et al.
Published: (2026)
by: Zhan, Yang, et al.
Published: (2026)
FunBench: Benchmarking Fundus Reading Skills of MLLMs
by: Wei, Qijie, et al.
Published: (2025)
by: Wei, Qijie, et al.
Published: (2025)
HyCTAS: Multi-Objective Hybrid Convolution-Transformer Architecture Search for Real-Time Image Segmentation
by: Yu, Hongyuan, et al.
Published: (2024)
by: Yu, Hongyuan, et al.
Published: (2024)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
by: Yuan, Jiakang, et al.
Published: (2025)
by: Yuan, Jiakang, et al.
Published: (2025)
Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
by: Lin, Yuanze, et al.
Published: (2024)
by: Lin, Yuanze, et al.
Published: (2024)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
by: Chen, Shuhang, et al.
Published: (2025)
by: Chen, Shuhang, et al.
Published: (2025)
Multi-Modal Building Change Detection for Large-Scale Small Changes: Benchmark and Baseline
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
by: Gao, Minghe, et al.
Published: (2024)
by: Gao, Minghe, et al.
Published: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024)
by: Ouyang, Kun, et al.
Published: (2024)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
by: Dang, Yunkai, et al.
Published: (2025)
by: Dang, Yunkai, et al.
Published: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
by: Liu, Huan, et al.
Published: (2024)
by: Liu, Huan, et al.
Published: (2024)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
by: Zhu, Xiaorong, et al.
Published: (2025)
by: Zhu, Xiaorong, et al.
Published: (2025)
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
by: Liu, Xianjie, et al.
Published: (2026)
by: Liu, Xianjie, et al.
Published: (2026)
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset
by: Wang, Jinhong, et al.
Published: (2025)
by: Wang, Jinhong, et al.
Published: (2025)
RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
by: Lawrence, Logan, et al.
Published: (2026)
by: Lawrence, Logan, et al.
Published: (2026)
Similar Items
-
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
by: Feng, Xuelu, et al.
Published: (2025) -
Pluralistic Salient Object Detection
by: Feng, Xuelu, et al.
Published: (2024) -
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
by: Zhu, Zixin, et al.
Published: (2024) -
GeoRemover: Removing Objects and Their Causal Visual Artifacts
by: Zhu, Zixin, et al.
Published: (2025) -
Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network
by: Luan, Tianyu, et al.
Published: (2025)