JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Onohara, Shota, Miyai, Atsuyuki, Imajuku, Yuki, Egashira, Kazuki, Baek, Jeonghun, Yue, Xiang, Neubig, Graham, Aizawa, Kiyoharu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025)
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution Detection
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2023)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2023)
A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models
von: Noda, Shiho, et al.
Veröffentlicht: (2025)
von: Noda, Shiho, et al.
Veröffentlicht: (2025)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2026)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2026)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
von: Zhang, Ge, et al.
Veröffentlicht: (2024)
von: Zhang, Ge, et al.
Veröffentlicht: (2024)
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
von: Yue, Xiang, et al.
Veröffentlicht: (2023)
von: Yue, Xiang, et al.
Veröffentlicht: (2023)
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
von: Wada, Yuiga, et al.
Veröffentlicht: (2025)
von: Wada, Yuiga, et al.
Veröffentlicht: (2025)
Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark
von: Zou, Kai, et al.
Veröffentlicht: (2025)
von: Zou, Kai, et al.
Veröffentlicht: (2025)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
Investigating the Perception of Facial Anonymization Techniques in 360° Videos
von: Wöhler, Leslie, et al.
Veröffentlicht: (2024)
von: Wöhler, Leslie, et al.
Veröffentlicht: (2024)
Privacy Protection and Video Manipulation in Immersive Media
von: Wöhler, Leslie, et al.
Veröffentlicht: (2024)
von: Wöhler, Leslie, et al.
Veröffentlicht: (2024)
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
von: Li, Yingxuan, et al.
Veröffentlicht: (2023)
von: Li, Yingxuan, et al.
Veröffentlicht: (2023)
Entity-NeRF: Detecting and Removing Moving Entities in Urban Scenes
von: Otonari, Takashi, et al.
Veröffentlicht: (2024)
von: Otonari, Takashi, et al.
Veröffentlicht: (2024)
Guided Image Synthesis via Initial Image Editing in Diffusion Model
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025)
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025)
Responsible Federated LLMs via Safety Filtering and Constitutional AI
von: Noh, Eunchung, et al.
Veröffentlicht: (2025)
von: Noh, Eunchung, et al.
Veröffentlicht: (2025)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
von: Nyandwi, Jean de Dieu, et al.
Veröffentlicht: (2025)
von: Nyandwi, Jean de Dieu, et al.
Veröffentlicht: (2025)
Training-Free Sketch-Guided Diffusion with Latent Optimization
von: Ding, Sandra Zhang, et al.
Veröffentlicht: (2024)
von: Ding, Sandra Zhang, et al.
Veröffentlicht: (2024)
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
von: Jiang, Junfeng, et al.
Veröffentlicht: (2024)
von: Jiang, Junfeng, et al.
Veröffentlicht: (2024)
PerFace: Metric Learning in Perceptual Facial Similarity for Enhanced Face Anonymization
von: Kumagai, Haruka, et al.
Veröffentlicht: (2025)
von: Kumagai, Haruka, et al.
Veröffentlicht: (2025)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025)
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026)
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025)
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025)
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
von: Song, Yueqi, et al.
Veröffentlicht: (2025)
von: Song, Yueqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025) -
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025) -
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026) -
Harnessing PDF Data for Improving Japanese Large Multimodal Models
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025) -
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)