WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Sugiura, Issa, Kurita, Shuhei, Oda, Yusuke, Kawahara, Daisuke, Okabe, Yasuo, Okazaki, Naoaki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation
by: Sugiura, Issa, et al.
Published: (2026)
by: Sugiura, Issa, et al.
Published: (2026)
HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
by: Sugiura, Issa, et al.
Published: (2026)
by: Sugiura, Issa, et al.
Published: (2026)
Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model
by: Sasagawa, Keito, et al.
Published: (2024)
by: Sasagawa, Keito, et al.
Published: (2024)
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
by: Sugiura, Issa, et al.
Published: (2026)
by: Sugiura, Issa, et al.
Published: (2026)
Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
by: Sasagawa, Keito, et al.
Published: (2025)
by: Sasagawa, Keito, et al.
Published: (2025)
Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling
by: Sugiura, Issa, et al.
Published: (2025)
by: Sugiura, Issa, et al.
Published: (2025)
llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
by: Sugiura, Issa, et al.
Published: (2025)
by: Sugiura, Issa, et al.
Published: (2025)
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
by: Yashima, Daichi, et al.
Published: (2026)
by: Yashima, Daichi, et al.
Published: (2026)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
by: Maeda, Koki, et al.
Published: (2024)
by: Maeda, Koki, et al.
Published: (2024)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
by: Maeda, Koki, et al.
Published: (2026)
by: Maeda, Koki, et al.
Published: (2026)
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
by: Yashima, Daichi, et al.
Published: (2026)
by: Yashima, Daichi, et al.
Published: (2026)
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer
by: Ma, Youmi, et al.
Published: (2024)
by: Ma, Youmi, et al.
Published: (2024)
JDocQA: Japanese Document Question Answering Dataset for Generative Language Models
by: Onami, Eri, et al.
Published: (2024)
by: Onami, Eri, et al.
Published: (2024)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
Drifting Objectives for Refining Discrete Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Diffusion-State Policy Optimization for Masked Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
by: Miyamoto, Sora, et al.
Published: (2026)
by: Miyamoto, Sora, et al.
Published: (2026)
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching
by: Yashima, Daichi, et al.
Published: (2026)
by: Yashima, Daichi, et al.
Published: (2026)
Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
by: Ichinose, Tatsuya, et al.
Published: (2026)
by: Ichinose, Tatsuya, et al.
Published: (2026)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
by: Oi, Masanari, et al.
Published: (2026)
by: Oi, Masanari, et al.
Published: (2026)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
by: Hida, Rem, et al.
Published: (2024)
by: Hida, Rem, et al.
Published: (2024)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
Kanbun-LM: Reading and Translating Classical Chinese in Japanese Methods by Language Models
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
Text-driven Affordance Learning from Egocentric Vision
by: Yoshida, Tomoya, et al.
Published: (2024)
by: Yoshida, Tomoya, et al.
Published: (2024)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
by: Ukai, Mahiro, et al.
Published: (2025)
by: Ukai, Mahiro, et al.
Published: (2025)
Building a Large Japanese Web Corpus for Large Language Models
by: Okazaki, Naoaki, et al.
Published: (2024)
by: Okazaki, Naoaki, et al.
Published: (2024)
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
by: Yokomizo, Hisayuki, et al.
Published: (2026)
by: Yokomizo, Hisayuki, et al.
Published: (2026)
LegalViz: Legal Text Visualization by Text To Diagram Generation
by: Onami, Eri, et al.
Published: (2025)
by: Onami, Eri, et al.
Published: (2025)
CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
by: Lee, Jungdae, et al.
Published: (2024)
by: Lee, Jungdae, et al.
Published: (2024)
Tokenization as Finite-State Transduction
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Similar Items
-
JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation
by: Sugiura, Issa, et al.
Published: (2026) -
HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
by: Sugiura, Issa, et al.
Published: (2026) -
Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model
by: Sasagawa, Keito, et al.
Published: (2024) -
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
by: Sugiura, Issa, et al.
Published: (2026) -
Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
by: Sasagawa, Keito, et al.
Published: (2025)