SO-Bench: A Structural Output Evaluation of Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Di, Ma, Kaixin, Nan, Feng, Chen, Haofeng, Zhai, Bohan, Griffiths, David, Gao, Mingfei, Gan, Zhe, Verma, Eshan, Yang, Yinfei, Chen, Zhifeng, Dehghan, Afshin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
by: Amirloo, Elmira, et al.
Published: (2024)
by: Amirloo, Elmira, et al.
Published: (2024)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
by: Daxberger, Erik, et al.
Published: (2025)
by: Daxberger, Erik, et al.
Published: (2025)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
by: Xu, Mingze, et al.
Published: (2025)
by: Xu, Mingze, et al.
Published: (2025)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
Cubify Anything: Scaling Indoor 3D Object Detection
by: Lazarow, Justin, et al.
Published: (2024)
by: Lazarow, Justin, et al.
Published: (2024)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
AToken: A Unified Tokenizer for Vision
by: Lu, Jiasen, et al.
Published: (2025)
by: Lu, Jiasen, et al.
Published: (2025)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
by: Zhang, Haotian, et al.
Published: (2024)
by: Zhang, Haotian, et al.
Published: (2024)
Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping
by: Lazarow, Justin, et al.
Published: (2025)
by: Lazarow, Justin, et al.
Published: (2025)
DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
by: Narayan, Kartik, et al.
Published: (2025)
by: Narayan, Kartik, et al.
Published: (2025)
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
by: Bachmann, Roman, et al.
Published: (2024)
by: Bachmann, Roman, et al.
Published: (2024)
RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
by: Wang, Haofeng, et al.
Published: (2025)
by: Wang, Haofeng, et al.
Published: (2025)
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
by: Fu, Tsu-Jui, et al.
Published: (2025)
by: Fu, Tsu-Jui, et al.
Published: (2025)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
by: Fu, Tsu-Jui, et al.
Published: (2023)
by: Fu, Tsu-Jui, et al.
Published: (2023)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
by: Ye, Hanrong, et al.
Published: (2024)
by: Ye, Hanrong, et al.
Published: (2024)
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
by: Jin, Bohan, et al.
Published: (2025)
by: Jin, Bohan, et al.
Published: (2025)
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
by: Ge, Wentao, et al.
Published: (2023)
by: Ge, Wentao, et al.
Published: (2023)
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
by: Agarwal, Vatsal, et al.
Published: (2025)
by: Agarwal, Vatsal, et al.
Published: (2025)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
by: Li, Zhangheng, et al.
Published: (2024)
by: Li, Zhangheng, et al.
Published: (2024)
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
by: Li, Yanghao, et al.
Published: (2025)
by: Li, Yanghao, et al.
Published: (2025)
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Revealing the Utilized Rank of Subspaces of Learning in Neural Networks
by: Garg, Isha, et al.
Published: (2024)
by: Garg, Isha, et al.
Published: (2024)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
by: Guo, Zichun, et al.
Published: (2026)
by: Guo, Zichun, et al.
Published: (2026)
STEM: Structure-Tracing Evidence Mining for Knowledge Graphs-Driven Retrieval-Augmented Generation
by: Yu, Peng, et al.
Published: (2026)
by: Yu, Peng, et al.
Published: (2026)
From Query to Counsel: Structured Reasoning with a Multi-Agent Framework and Dataset for Legal Consultation
by: Lu, Mingfei, et al.
Published: (2026)
by: Lu, Mingfei, et al.
Published: (2026)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal Models
by: Lin, Hongzhan, et al.
Published: (2024)
by: Lin, Hongzhan, et al.
Published: (2024)
What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and Mitigations
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
by: Wang, Haibo, et al.
Published: (2025)
by: Wang, Haibo, et al.
Published: (2025)
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
by: Chen, Hailin, et al.
Published: (2024)
by: Chen, Hailin, et al.
Published: (2024)
Similar Items
-
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
by: Tian, Rui, et al.
Published: (2025) -
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
by: Amirloo, Elmira, et al.
Published: (2024) -
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
by: Qian, Yusu, et al.
Published: (2024) -
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
by: Tian, Rui, et al.
Published: (2025) -
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
by: Daxberger, Erik, et al.
Published: (2025)