Seeing Culture: A Benchmark for Visual Reasoning and Grounding

Published in EMNLP 2025 (Main Conference), 2025

Burak Satar, Zhixin Ma, Patrick Amadeus Irawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo

A two-stage benchmark where models must first reason about a cultural artifact, then visually ground it, built across Southeast Asia.

Seeing Culture's two-stage task: a culturally grounded question with four image options, then a segmentation mask marking the artifact the model reasoned about.

What it covers

Seeing Culture benchmark coverage across Southeast AsiaIndonesia 1377 questions across 35 artifacts; Myanmar 717 questions across 45 artifacts; Thailand 344 questions across 26 artifacts; Vietnam 324 questions across 11 artifacts; Philippines 212 questions across 10 artifacts; Malaysia 182 questions across 5 artifacts; Cambodia 22 questions across 2 artifactsIndonesia: 1377 questions, 35 artifactsMyanmar: 717 questions, 45 artifactsThailand: 344 questions, 26 artifactsVietnam: 324 questions, 11 artifactsPhilippines: 212 questions, 10 artifactsMalaysia: 182 questions, 5 artifactsCambodia: 22 questions, 2 artifactsLaos: not in the benchmarkBrunei: not in the benchmarkTimor-Leste: not in the benchmark Singapore: where the benchmark was built

3,178 questions on 134 cultural artifacts across 7 countries, counted from the public question set of the Seeing Culture benchmark. The paper reports 138 artifacts in total; a few do not appear in the released questions.

  • Indonesia1,377 questions · 35 artifacts
  • Myanmar717 questions · 45 artifacts
  • Thailand344 questions · 26 artifacts
  • Vietnam324 questions · 11 artifacts
  • Philippines212 questions · 10 artifacts
  • Malaysia182 questions · 5 artifacts
  • Cambodia22 questions · 2 artifacts

Browse the 1,065 images and 3,178 questions in the dataset viewer on Hugging Face.

The video sequel, Cultural Moment (EMNLP 2026), carries the visual-option design into video and adds free-form temporal localization: name the concept, recognize it among unlabeled video moments, then locate its sub-events in time. Its project page has a playable sample of all three stages.

BibTeX
@inproceedings{satar-etal-2025-seeing,
    title = "Seeing Culture: A Benchmark for Visual Reasoning and Grounding",
    author = "Satar, Burak  and
      Ma, Zhixin  and
      Irawan, Patrick Amadeus  and
      Mulyawan, Wilfried Ariel  and
      Jiang, Jing  and
      Lim, Ee-Peng  and
      Ngo, Chong-Wah",
    editor = "Christodoulopoulos, Christos  and
      Chakraborty, Tanmoy  and
      Rose, Carolyn  and
      Peng, Violet",
    booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.emnlp-main.1131/",
    doi = "10.18653/v1/2025.emnlp-main.1131",
    pages = "22227--22243",
    ISBN = "979-8-89176-332-6"
}