WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sugiura, Issa, Kurita, Shuhei, Oda, Yusuke, Kawahara, Daisuke, Okabe, Yasuo, Okazaki, Naoaki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910274529787904
author Sugiura, Issa
Kurita, Shuhei
Oda, Yusuke
Kawahara, Daisuke
Okabe, Yasuo
Okazaki, Naoaki
author_facet Sugiura, Issa
Kurita, Shuhei
Oda, Yusuke
Kawahara, Daisuke
Okabe, Yasuo
Okazaki, Naoaki
contents Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing English-only caption filters and pretraining on global data is effective for improving multicultural performance. We study whether such global pretraining is sufficient for culture-specific understanding, or whether further adaptation with natively sourced data can boost performance beyond what global pretraining alone achieves. To enable this investigation, we present WAON, the largest publicly available native Japanese image-text dataset constructed from native Japanese web content in Common Crawl, containing approximately 155 million examples. We also introduce WAON-Bench, a manually curated Japanese cultural benchmark spanning 374 classes. Through comparative fine-tuning experiments on multiple Japanese image-text datasets, we observe that models fine-tuned on WAON consistently achieve stronger performance on Japanese cultural benchmarks than those fine-tuned on English-to-Japanese translated data. We release our dataset and code.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22276
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
Sugiura, Issa
Kurita, Shuhei
Oda, Yusuke
Kawahara, Daisuke
Okabe, Yasuo
Okazaki, Naoaki
Computer Vision and Pattern Recognition
Computation and Language
Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing English-only caption filters and pretraining on global data is effective for improving multicultural performance. We study whether such global pretraining is sufficient for culture-specific understanding, or whether further adaptation with natively sourced data can boost performance beyond what global pretraining alone achieves. To enable this investigation, we present WAON, the largest publicly available native Japanese image-text dataset constructed from native Japanese web content in Common Crawl, containing approximately 155 million examples. We also introduce WAON-Bench, a manually curated Japanese cultural benchmark spanning 374 classes. Through comparative fine-tuning experiments on multiple Japanese image-text datasets, we observe that models fine-tuned on WAON consistently achieve stronger performance on Japanese cultural benchmarks than those fine-tuned on English-to-Japanese translated data. We release our dataset and code.
title WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.22276