WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Jack, Yan, Shilin, Cai, Jiayin, Jiang, Xiaolong, Hu, Yao, Xie, Weidi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Sanity Check for AI-generated Image Detection
von: Yan, Shilin, et al.
Veröffentlicht: (2024)
von: Yan, Shilin, et al.
Veröffentlicht: (2024)
Progressive Scaling Visual Object Tracking
von: Hong, Jack, et al.
Veröffentlicht: (2025)
von: Hong, Jack, et al.
Veröffentlicht: (2025)
LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant
von: Liu, Yikun, et al.
Veröffentlicht: (2024)
von: Liu, Yikun, et al.
Veröffentlicht: (2024)
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
VISA: Reasoning Video Object Segmentation via Large Language Models
von: Yan, Cilin, et al.
Veröffentlicht: (2024)
von: Yan, Cilin, et al.
Veröffentlicht: (2024)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
von: Shi, Yudi, et al.
Veröffentlicht: (2026)
von: Shi, Yudi, et al.
Veröffentlicht: (2026)
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
von: Tao, Keda, et al.
Veröffentlicht: (2026)
von: Tao, Keda, et al.
Veröffentlicht: (2026)
Real-World Point Tracking with Verifier-Guided Pseudo-Labeling
von: Aydemir, Görkay, et al.
Veröffentlicht: (2026)
von: Aydemir, Görkay, et al.
Veröffentlicht: (2026)
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective
von: Li, Ouxiang, et al.
Veröffentlicht: (2024)
von: Li, Ouxiang, et al.
Veröffentlicht: (2024)
Active Perception Agent for Omnimodal Audio-Video Understanding
von: Tao, Keda, et al.
Veröffentlicht: (2025)
von: Tao, Keda, et al.
Veröffentlicht: (2025)
PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
von: Lu, Xudong, et al.
Veröffentlicht: (2026)
von: Lu, Xudong, et al.
Veröffentlicht: (2026)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
von: Yan, Yibin, et al.
Veröffentlicht: (2024)
von: Yan, Yibin, et al.
Veröffentlicht: (2024)
AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning
von: Qiu, Yilun, et al.
Veröffentlicht: (2026)
von: Qiu, Yilun, et al.
Veröffentlicht: (2026)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
von: Wang, Xiaodong, et al.
Veröffentlicht: (2026)
von: Wang, Xiaodong, et al.
Veröffentlicht: (2026)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
Character-Centric Understanding of Animated Movies
von: Gui, Zhongrui, et al.
Veröffentlicht: (2025)
von: Gui, Zhongrui, et al.
Veröffentlicht: (2025)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
von: Jiang, Tianxiang, et al.
Veröffentlicht: (2025)
von: Jiang, Tianxiang, et al.
Veröffentlicht: (2025)
Towards Universal Soccer Video Understanding
von: Rao, Jiayuan, et al.
Veröffentlicht: (2024)
von: Rao, Jiayuan, et al.
Veröffentlicht: (2024)
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
von: Chen, Jiacheng, et al.
Veröffentlicht: (2024)
von: Chen, Jiacheng, et al.
Veröffentlicht: (2024)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
FaceXBench: Evaluating Multimodal LLMs on Face Understanding
von: Narayan, Kartik, et al.
Veröffentlicht: (2025)
von: Narayan, Kartik, et al.
Veröffentlicht: (2025)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
von: Jiang, Roy, et al.
Veröffentlicht: (2026)
von: Jiang, Roy, et al.
Veröffentlicht: (2026)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
Track-On: Transformer-based Online Point Tracking with Memory
von: Aydemir, Görkay, et al.
Veröffentlicht: (2025)
von: Aydemir, Görkay, et al.
Veröffentlicht: (2025)
MFSR: MeanFlow Distillation for One Step Real-World Image Super Resolution
von: Wang, Ruiqing, et al.
Veröffentlicht: (2026)
von: Wang, Ruiqing, et al.
Veröffentlicht: (2026)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
von: Elobaid, Alaa
Veröffentlicht: (2026)
von: Elobaid, Alaa
Veröffentlicht: (2026)
A Sanity Check on Composed Image Retrieval
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding
von: Titiya, Prasham, et al.
Veröffentlicht: (2025)
von: Titiya, Prasham, et al.
Veröffentlicht: (2025)
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
Appearance-Based Refinement for Object-Centric Motion Segmentation
von: Xie, Junyu, et al.
Veröffentlicht: (2023)
von: Xie, Junyu, et al.
Veröffentlicht: (2023)
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
von: Liang, Yongyuan, et al.
Veröffentlicht: (2025)
von: Liang, Yongyuan, et al.
Veröffentlicht: (2025)
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
von: Li, Hongxiang, et al.
Veröffentlicht: (2025)
von: Li, Hongxiang, et al.
Veröffentlicht: (2025)
SoccerMaster: A Vision Foundation Model for Soccer Understanding
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
Multi-Agent System for Comprehensive Soccer Understanding
von: Rao, Jiayuan, et al.
Veröffentlicht: (2025)
von: Rao, Jiayuan, et al.
Veröffentlicht: (2025)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
von: Zhang, Ruoxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruoxuan, et al.
Veröffentlicht: (2025)
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Sanity Check for AI-generated Image Detection
von: Yan, Shilin, et al.
Veröffentlicht: (2024) -
Progressive Scaling Visual Object Tracking
von: Hong, Jack, et al.
Veröffentlicht: (2025) -
LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant
von: Liu, Yikun, et al.
Veröffentlicht: (2024) -
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025) -
VISA: Reasoning Video Object Segmentation via Large Language Models
von: Yan, Cilin, et al.
Veröffentlicht: (2024)