Retrieval-Augmented Reasoning Segmentation in Cultural Context
Published in Under review, 2026
Zhixin Ma*, Burak Satar*, Patrick Amadeus Irawan, Wilfried Ariel Mulyawan, Phuong Anh Nguyen, Chong-Wah Ngo (* equal contribution)
Retrieval-augmented reasoning segmentation for culturally grounded scenes, so a model must retrieve the right cultural knowledge before it segments.
Status
This work is under review. There is no preprint yet; a link will appear here when one is available.
For a draft, or to discuss the work, email me or book a 30-minute chat.
Related
Part of the same line of work as Seeing Culture (EMNLP 2025) and the Cultural Moment benchmark, which ask whether a model reasons about a cultural artifact or only sounds confident about it. This one pushes the grounding requirement from a bounding box to a segmentation mask.
