{"schema":"hottub-news/story/v1","story":"st_7etbasuhhy5mpvrodm4q","title":"WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models","title_en":null,"outlets":1,"brief":null,"page":"https://hottub.news/stories/waon-a-large-scale-japanese-image-text-dataset-for-cultural-adaptation-7etbasuhhy","count":1,"items":[{"id":"n_7etbasuhhy5mpvrodm4q","seq":1756503,"kind":"article","title":"WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models","url":"https://arxiv.org/abs/2510.22276","summary":"arXiv:2510.22276v4 Announce Type: replace Abstract: Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing English-only caption filters and pretraining on global data is effective for improving multicultural performance. We study whether such global pretraining is sufficient for culture-specific understanding, or whether further adaptation with natively sourced data can boost performance…","published":"2026-10-05T04:00:00Z","seen":"2026-10-05T04:11:02Z","lang":"en","source":{"id":"kite-arxiv-org-1637f0","name":"arxiv.org","domain":"rss.arxiv.org"},"via":"kite-arxiv-org-1637f0","topics":["photonics","cs.cv","cs.cl"],"authors":["Issa Sugiura, Shuhei Kurita, Yusuke Oda, Daisuke Kawahara, Yasuo Okabe, Naoaki…"],"entities":[{"id":"Q83664782","name":"Rosa 'Waon'","type":"other"}],"mentions":2,"story":"st_7etbasuhhy5mpvrodm4q","category":"arts-culture-entertainment","category_p":0.4099999964237213,"category2":"science-technology","sentiment":"neutral","tone":0.20999999344348907,"political":0.009999999776482582}]}