Retrieval-Augmented Reasoning Segmentation in Cultural Context

Published in Under review, 2026

Zhixin Ma*, Burak Satar*, Patrick Amadeus Irawan, Wilfried Ariel Mulyawan, Phuong Anh Nguyen, Chong-Wah Ngo (* equal contribution)

Retrieval-augmented reasoning segmentation for culturally grounded scenes, so a model must retrieve the right cultural knowledge before it segments.

Status

This work is under review. There is no preprint yet; a link will appear here when one is available.

For a draft, or to discuss the work, email me or book a 30-minute chat.

Part of the same line of work as Seeing Culture (EMNLP 2025) and the Cultural Moment benchmark, which ask whether a model reasons about a cultural artifact or only sounds confident about it. This one pushes the grounding requirement from a bounding box to a segmentation mask.