JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sugiura, Issa, Maeda, Koki, Kurita, Shuhei, Oda, Yusuke, Kawahara, Daisuke, Okazaki, Naoaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
von: Sugiura, Issa, et al.
Veröffentlicht: (2026)
von: Sugiura, Issa, et al.
Veröffentlicht: (2026)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
von: Sugiura, Issa, et al.
Veröffentlicht: (2026)
von: Sugiura, Issa, et al.
Veröffentlicht: (2026)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
von: Maeda, Koki, et al.
Veröffentlicht: (2026)
von: Maeda, Koki, et al.
Veröffentlicht: (2026)
Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model
von: Sasagawa, Keito, et al.
Veröffentlicht: (2024)
von: Sasagawa, Keito, et al.
Veröffentlicht: (2024)
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
von: Yashima, Daichi, et al.
Veröffentlicht: (2026)
von: Yashima, Daichi, et al.
Veröffentlicht: (2026)
Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
von: Sasagawa, Keito, et al.
Veröffentlicht: (2025)
von: Sasagawa, Keito, et al.
Veröffentlicht: (2025)
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
von: Yashima, Daichi, et al.
Veröffentlicht: (2026)
von: Yashima, Daichi, et al.
Veröffentlicht: (2026)
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
von: Oi, Masanari, et al.
Veröffentlicht: (2026)
von: Oi, Masanari, et al.
Veröffentlicht: (2026)
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
von: Ukai, Mahiro, et al.
Veröffentlicht: (2025)
von: Ukai, Mahiro, et al.
Veröffentlicht: (2025)
Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos
von: Tsuchiya, Fumihiko, et al.
Veröffentlicht: (2026)
von: Tsuchiya, Fumihiko, et al.
Veröffentlicht: (2026)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
von: Ohi, Masanari, et al.
Veröffentlicht: (2024)
von: Ohi, Masanari, et al.
Veröffentlicht: (2024)
Noisy Label Refinement with Semantically Reliable Synthetic Images
von: Li, Yingxuan, et al.
Veröffentlicht: (2025)
von: Li, Yingxuan, et al.
Veröffentlicht: (2025)
Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness
von: Wan, Hanwen, et al.
Veröffentlicht: (2025)
von: Wan, Hanwen, et al.
Veröffentlicht: (2025)
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
von: Sakamoto, Koya, et al.
Veröffentlicht: (2026)
von: Sakamoto, Koya, et al.
Veröffentlicht: (2026)
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
von: Yoshida, Tomoya, et al.
Veröffentlicht: (2025)
von: Yoshida, Tomoya, et al.
Veröffentlicht: (2025)
Cross3DVG: Cross-Dataset 3D Visual Grounding on Different RGB-D Scans
von: Miyanishi, Taiki, et al.
Veröffentlicht: (2023)
von: Miyanishi, Taiki, et al.
Veröffentlicht: (2023)
DetRefiner: Model-Agnostic Detection Refinement with Feature Fusion Transformer
von: Okazaki, Soichiro, et al.
Veröffentlicht: (2026)
von: Okazaki, Soichiro, et al.
Veröffentlicht: (2026)
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
von: Yokomizo, Hisayuki, et al.
Veröffentlicht: (2026)
von: Yokomizo, Hisayuki, et al.
Veröffentlicht: (2026)
Answerability Fields: Answerable Location Estimation via Diffusion Models
von: Azuma, Daichi, et al.
Veröffentlicht: (2024)
von: Azuma, Daichi, et al.
Veröffentlicht: (2024)
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
Text-driven Affordance Learning from Egocentric Vision
von: Yoshida, Tomoya, et al.
Veröffentlicht: (2024)
von: Yoshida, Tomoya, et al.
Veröffentlicht: (2024)
Toward Reliable VLM: A Fine-Grained Benchmark and Framework for Exposure, Bias, and Inference in Korean Street Views
von: Wang, Xiaonan, et al.
Veröffentlicht: (2025)
von: Wang, Xiaonan, et al.
Veröffentlicht: (2025)
Map-based Modular Approach for Zero-shot Embodied Question Answering
von: Sakamoto, Koya, et al.
Veröffentlicht: (2024)
von: Sakamoto, Koya, et al.
Veröffentlicht: (2024)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
von: Zhao, Yi, et al.
Veröffentlicht: (2026)
von: Zhao, Yi, et al.
Veröffentlicht: (2026)
Benchmarking and Enhancing VLM for Compressed Image Understanding
von: Zhang, Zifu, et al.
Veröffentlicht: (2025)
von: Zhang, Zifu, et al.
Veröffentlicht: (2025)
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
Referring Expression Comprehension for Small Objects
von: Goto, Kanoko, et al.
Veröffentlicht: (2025)
von: Goto, Kanoko, et al.
Veröffentlicht: (2025)
Concepts Worth Having: Refining VLM-Guided Concept Bottleneck Models with Minimal Annotations
von: Debole, Nicola, et al.
Veröffentlicht: (2026)
von: Debole, Nicola, et al.
Veröffentlicht: (2026)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
von: Danish, Muhammad Sohail, et al.
Veröffentlicht: (2024)
von: Danish, Muhammad Sohail, et al.
Veröffentlicht: (2024)
Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
FLARE-SSM: Deep State Space Models with Influence-Balanced Loss for 72-Hour Solar Flare Prediction
von: Takagi, Yusuke, et al.
Veröffentlicht: (2025)
von: Takagi, Yusuke, et al.
Veröffentlicht: (2025)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
CERBERUS: Crack Evaluation & Recognition Benchmark for Engineering Reliability & Urban Stability
von: Reinman, Justin, et al.
Veröffentlicht: (2025)
von: Reinman, Justin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
von: Sugiura, Issa, et al.
Veröffentlicht: (2026) -
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
von: Sugiura, Issa, et al.
Veröffentlicht: (2025) -
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
von: Sugiura, Issa, et al.
Veröffentlicht: (2026) -
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
von: Maeda, Koki, et al.
Veröffentlicht: (2024) -
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
von: Maeda, Koki, et al.
Veröffentlicht: (2026)