Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia
Published in Under review, 2026
Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo
Extends cultural reasoning and grounding from still images to video, across Southeast Asia.
Status
This work is under review. There is no preprint yet; a link will appear here when one is available.
If you would like to read a draft in the meantime, or discuss the benchmark, email me or book a 30-minute chat.
Related
It builds on Seeing Culture (EMNLP 2025), which asks the same two-stage question — reason about a cultural artifact, then ground it in the image — of still images rather than video.
