Publications

Also on Google Scholar. An asterisk marks equal contribution.

Cultural multimodal reasoning and grounding

Benchmarks and methods that test whether a vision-language model understands a cultural scene or has only learned to sound confident about it. Southeast Asia first, because it is the hardest available test case.

Debiased and semantic text-video retrieval

Video retrieval models take shortcuts: clip length, object co-occurrence, verb priors. My doctoral work identified these shortcuts and removed them with semantic role structure and causal intervention.

Long-video retrieval and multimodal planning

Grounding language-model plans in video, and finding the right moment inside hours of egocentric footage.

Earlier work

Vehicle detection and recognition, from my master's.