Seeing Culture: A Benchmark for Visual Reasoning and Grounding
A two-stage benchmark where models must first reason about a cultural artifact, then visually ground it, built across Southeast Asia.
Also on Google Scholar. An asterisk marks equal contribution.
Benchmarks and methods that test whether a vision-language model understands a cultural scene or has only learned to sound confident about it. Southeast Asia first, because it is the hardest available test case.
A two-stage benchmark where models must first reason about a cultural artifact, then visually ground it, built across Southeast Asia.
Extends cultural reasoning and grounding from still images to video, across Southeast Asia. EMNLP 2026.
Retrieval-augmented reasoning segmentation for culturally grounded scenes, so a model must retrieve the right cultural knowledge before it segments. Under review for ACM TOMM.
Tracking and segmenting culturally significant objects through video, where the object’s identity depends on cultural knowledge rather than appearance alone.
Video retrieval models take shortcuts: clip length, object co-occurrence, verb priors. My doctoral work identified these shortcuts and removed them with semantic role structure and causal intervention.
Shows that text-video retrieval models exploit clip length as a shortcut, and mitigates the bias with causal intervention.
An extended abstract mapping open challenges in egocentric text-video retrieval.
A role-aware mixture-of-experts transformer that separates verbs, objects and context for text-to-video retrieval.
Joint 3rd place in the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge at CVPR 2022.
Aligns semantic roles between text and video with a correlation transformer for retrieval.
Grounding language-model plans in video, and finding the right moment inside hours of egocentric footage.
Procedural planning with LLMs improves when text and video prompts are visually grounded together.
Finding the right moment inside hours of egocentric video by fusing audio with language-model reasoning over the visual stream.
Vehicle detection and recognition, from my master's.
CNN-based vehicle make and model classification from my MSc research.