UAL-Bench: The First Comprehensive Unusual Activity Localization Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdullah, Hasnat Md, Liu, Tian, Wei, Kangda, Kong, Shu, Huang, Ruihong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
CliME: Evaluating Multimodal Climate Discourse on Social Media and the Climate Alignment Quotient (CAQ)
von: Borah, Abhilekh, et al.
Veröffentlicht: (2025)
von: Borah, Abhilekh, et al.
Veröffentlicht: (2025)
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation
von: Shang, Fangxin, et al.
Veröffentlicht: (2025)
von: Shang, Fangxin, et al.
Veröffentlicht: (2025)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
von: Li, Xinze, et al.
Veröffentlicht: (2026)
von: Li, Xinze, et al.
Veröffentlicht: (2026)
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
von: Gao, Siyuan, et al.
Veröffentlicht: (2025)
von: Gao, Siyuan, et al.
Veröffentlicht: (2025)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone
von: Malik, Shaivi, et al.
Veröffentlicht: (2025)
von: Malik, Shaivi, et al.
Veröffentlicht: (2025)
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
von: Ruan, Chenxi, et al.
Veröffentlicht: (2026)
von: Ruan, Chenxi, et al.
Veröffentlicht: (2026)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
von: Li, Siqi, et al.
Veröffentlicht: (2025)
von: Li, Siqi, et al.
Veröffentlicht: (2025)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
von: Ghaboura, Sara, et al.
Veröffentlicht: (2024)
von: Ghaboura, Sara, et al.
Veröffentlicht: (2024)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
von: Zhou, Li, et al.
Veröffentlicht: (2025)
von: Zhou, Li, et al.
Veröffentlicht: (2025)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
von: Su, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Su, Xiaoyan, et al.
Veröffentlicht: (2026)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
von: Wei, Kangda, et al.
Veröffentlicht: (2025) -
CliME: Evaluating Multimodal Climate Discourse on Social Media and the Climate Alignment Quotient (CAQ)
von: Borah, Abhilekh, et al.
Veröffentlicht: (2025) -
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
von: Nie, Yiqi, et al.
Veröffentlicht: (2026) -
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos
von: Wei, Kangda, et al.
Veröffentlicht: (2025) -
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)