Saved in:
| Main Authors: | Guo, Jindi, Huang, Chaozheng, Fang, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.21277 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
by: Hu, Yijie, et al.
Published: (2025)
by: Hu, Yijie, et al.
Published: (2025)
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
by: Tao, Haoyi, et al.
Published: (2026)
by: Tao, Haoyi, et al.
Published: (2026)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024)
by: Tang, Yolo Y., et al.
Published: (2024)
Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
by: Xi, Suyang, et al.
Published: (2025)
by: Xi, Suyang, et al.
Published: (2025)
Can MLLMs Understand the Deep Implication Behind Chinese Images?
by: Zhang, Chenhao, et al.
Published: (2024)
by: Zhang, Chenhao, et al.
Published: (2024)
Can the Recovery Mechanism Survive AI? Skill Formation, Labor, and What Current Measurement Misses
by: Fan, Aysa Xuemo
Published: (2026)
by: Fan, Aysa Xuemo
Published: (2026)
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
by: Song, Dingjie, et al.
Published: (2026)
by: Song, Dingjie, et al.
Published: (2026)
Can MLLMs generate human-like feedback in grading multimodal short answers?
by: Sil, Pritam, et al.
Published: (2024)
by: Sil, Pritam, et al.
Published: (2024)
UMoE: Unifying Attention and FFN with Shared Experts
by: Yang, Yuanhang, et al.
Published: (2025)
by: Yang, Yuanhang, et al.
Published: (2025)
GeoGS-CE: Learning Delay--Beam Channel Priors with 3D Gaussians for High-Mobility Scenarios
by: Zhang, Yumeng, et al.
Published: (2026)
by: Zhang, Yumeng, et al.
Published: (2026)
Affordance Benchmark for MLLMs
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
Redundancy Principles for MLLMs Benchmarks
by: Zhang, Zicheng, et al.
Published: (2025)
by: Zhang, Zicheng, et al.
Published: (2025)
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
by: Kang, Caixin, et al.
Published: (2026)
by: Kang, Caixin, et al.
Published: (2026)
RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature
by: Li, Hanzheng, et al.
Published: (2025)
by: Li, Hanzheng, et al.
Published: (2025)
VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
by: Huang, Peng, et al.
Published: (2025)
by: Huang, Peng, et al.
Published: (2025)
CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
by: Lam, Man Ho, et al.
Published: (2025)
by: Lam, Man Ho, et al.
Published: (2025)
Multi-Path Collaborative Reasoning via Reinforcement Learning
by: Lv, Jindi, et al.
Published: (2025)
by: Lv, Jindi, et al.
Published: (2025)
MLLMs are Deeply Affected by Modality Bias
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
What Is Missing: Interpretable Ratings for Large Language Model Outputs
by: Stranges, Nicholas, et al.
Published: (2026)
by: Stranges, Nicholas, et al.
Published: (2026)
Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
by: Ulmer, Dennis, et al.
Published: (2025)
by: Ulmer, Dennis, et al.
Published: (2025)
ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
by: Xu, Zhengzhuo, et al.
Published: (2025)
by: Xu, Zhengzhuo, et al.
Published: (2025)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
by: Jiang, Shixin, et al.
Published: (2024)
by: Jiang, Shixin, et al.
Published: (2024)
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
by: Du, Dazhao, et al.
Published: (2026)
by: Du, Dazhao, et al.
Published: (2026)
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
by: Guo, Peizheng, et al.
Published: (2025)
by: Guo, Peizheng, et al.
Published: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
by: Guo, Shengyu, et al.
Published: (2026)
by: Guo, Shengyu, et al.
Published: (2026)
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
by: Li, Zongjie, et al.
Published: (2026)
by: Li, Zongjie, et al.
Published: (2026)
Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
by: Helu, Zhi, et al.
Published: (2025)
by: Helu, Zhi, et al.
Published: (2025)
REPAIR: Robust Editing via Progressive Adaptive Intervention and Reintegration
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
Towards Faithful Reasoning in Comics for Small MLLMs
by: Feng, Chengcheng, et al.
Published: (2026)
by: Feng, Chengcheng, et al.
Published: (2026)
Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning
by: Li, Rongjin, et al.
Published: (2026)
by: Li, Rongjin, et al.
Published: (2026)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
by: Huang, Jincai, et al.
Published: (2026)
by: Huang, Jincai, et al.
Published: (2026)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
by: Gong, Zhantao, et al.
Published: (2025)
by: Gong, Zhantao, et al.
Published: (2025)
HyperNAS: Enhancing Architecture Representation for NAS Predictor via Hypernetwork
by: Lv, Jindi, et al.
Published: (2025)
by: Lv, Jindi, et al.
Published: (2025)
Deploying Models to Non-participating Clients in Federated Learning without Fine-tuning: A Hypernetwork-based Approach
by: Zhou, Yuhao, et al.
Published: (2025)
by: Zhou, Yuhao, et al.
Published: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
by: Li, Shuo, et al.
Published: (2025)
by: Li, Shuo, et al.
Published: (2025)
Mind the Gap: Promoting Missing Modality Brain Tumor Segmentation with Alignment
by: Liu, Tianyi, et al.
Published: (2024)
by: Liu, Tianyi, et al.
Published: (2024)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
by: Zheng, Naishan, et al.
Published: (2025)
by: Zheng, Naishan, et al.
Published: (2025)
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024)
by: Yao, Huanjin, et al.
Published: (2024)
Do MLLMs See What We See? Analyzing Visualization Literacy Barriers in AI Systems
by: Mengli, et al.
Published: (2026)
by: Mengli, et al.
Published: (2026)
Similar Items
-
Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
by: Hu, Yijie, et al.
Published: (2025) -
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
by: Tao, Haoyi, et al.
Published: (2026) -
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024) -
Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
by: Xi, Suyang, et al.
Published: (2025) -
Can MLLMs Understand the Deep Implication Behind Chinese Images?
by: Zhang, Chenhao, et al.
Published: (2024)