RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Zixi, Li, Jiapeng, Diao, Muxi, Jing, Yinuo, Liang, Kongming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
von: Wang, Xinran, et al.
Veröffentlicht: (2024)
von: Wang, Xinran, et al.
Veröffentlicht: (2024)
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025)
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
von: Lin, Junming, et al.
Veröffentlicht: (2024)
von: Lin, Junming, et al.
Veröffentlicht: (2024)
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
von: Dai, Yang, et al.
Veröffentlicht: (2026)
von: Dai, Yang, et al.
Veröffentlicht: (2026)
Robust image representations with counterfactual contrastive learning
von: Roschewitz, Mélanie, et al.
Veröffentlicht: (2024)
von: Roschewitz, Mélanie, et al.
Veröffentlicht: (2024)
Mitigating attribute amplification in counterfactual image generation
von: Xia, Tian, et al.
Veröffentlicht: (2024)
von: Xia, Tian, et al.
Veröffentlicht: (2024)
ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs
von: Wu, Xin, et al.
Veröffentlicht: (2026)
von: Wu, Xin, et al.
Veröffentlicht: (2026)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026)
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026)
HiBug2: Efficient and Interpretable Error Slice Discovery for Comprehensive Model Debugging
von: Chen, Muxi, et al.
Veröffentlicht: (2025)
von: Chen, Muxi, et al.
Veröffentlicht: (2025)
Leveraging counterfactual concepts for debugging and improving CNN model performance
von: Tariq, Syed Ali, et al.
Veröffentlicht: (2025)
von: Tariq, Syed Ali, et al.
Veröffentlicht: (2025)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
Dense Connector for MLLMs
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2025)
PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Image Segmentation
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025)
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025)
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs
von: Jiang, Lifan, et al.
Veröffentlicht: (2026)
von: Jiang, Lifan, et al.
Veröffentlicht: (2026)
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
von: Tang, Yolo Y., et al.
Veröffentlicht: (2024)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2024)
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
von: Diao, Muxi, et al.
Veröffentlicht: (2025)
von: Diao, Muxi, et al.
Veröffentlicht: (2025)
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
SWAT: Sliding Window Adversarial Training for Gradual Domain Adaptation
von: Wang, Zixi, et al.
Veröffentlicht: (2025)
von: Wang, Zixi, et al.
Veröffentlicht: (2025)
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
von: Yu, Zhuang, et al.
Veröffentlicht: (2026)
von: Yu, Zhuang, et al.
Veröffentlicht: (2026)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
von: Gong, Zhantao, et al.
Veröffentlicht: (2025)
von: Gong, Zhantao, et al.
Veröffentlicht: (2025)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
Toward Generalizable Forgery Detection and Reasoning
von: Gao, Yueying, et al.
Veröffentlicht: (2025)
von: Gao, Yueying, et al.
Veröffentlicht: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
von: Liu, Xianjie, et al.
Veröffentlicht: (2025)
von: Liu, Xianjie, et al.
Veröffentlicht: (2025)
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
von: Heng, Yongrui, et al.
Veröffentlicht: (2026)
von: Heng, Yongrui, et al.
Veröffentlicht: (2026)
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
von: Mao, Jiawei, et al.
Veröffentlicht: (2025)
von: Mao, Jiawei, et al.
Veröffentlicht: (2025)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
Towards Faithful Reasoning in Comics for Small MLLMs
von: Feng, Chengcheng, et al.
Veröffentlicht: (2026)
von: Feng, Chengcheng, et al.
Veröffentlicht: (2026)
Explore the Hallucination on Low-level Perception for MLLMs
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
von: Wang, Xinran, et al.
Veröffentlicht: (2024) -
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025) -
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024) -
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
von: Lin, Junming, et al.
Veröffentlicht: (2024) -
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
von: Wang, Xinran, et al.
Veröffentlicht: (2025)