Burak Satar

I make Vision-Language Models (VLMs) culturally aware, starting with Southeast Asia, one of the world’s most culturally diverse regions.

I am a Research Scientist at Singapore Management University (SMU), working with Prof Chong-Wah Ngo on multimodal reasoning across image, video, audio and text.

I am actively looking for collaborators and student interns on culturally-aware multimodal AI.
Book a 30-minute chat or email me.

Watch this space: our newest benchmark, Cultural Moment, is under review at EMNLP 2026.

I build tests that vision-language models fail. When a model describes a festival, a dish or a ritual from Southeast Asia, does it understand what it is looking at, or has it only learned what confidence sounds like? Our EMNLP 2025 benchmark, Seeing Culture, makes models show their work: answer a culturally grounded question, then point to the evidence in the image. A model that names the right artifact while highlighting the wrong one did not know the answer; it guessed well. The gaps we measure are systematic, not noise.

Before SMU I spent five years at NTU and A*STAR’s Institute for Infocomm Research on an A*STAR SINGA scholarship, working on semantic, debiased and moment-level text-video retrieval. I grew up in Bursa (Türkiye), studied on exchange in Siena and Naples (Italy), worked in Valencia (Spain) and Istanbul (Türkiye), and have lived in Singapore since 2020. I have often been the person in the room who does not get the reference. The difference is that I knew it. The models I test do not.

My PhD thesis, Towards Semantic, Debiased and Moment Video Retrieval with Multi-modal Features, was supervised by Prof Joo-Hwee Lim, Dr Hongyuan Zhu and Prof Hanwang Zhang. During it I spent three months with Dr Michael Wray in Dima Damen’s group at the University of Bristol. My master’s, on vehicle detection, was supervised by Prof Ahmet Emir Dirik at Uludağ University.

What I work on

  • Cultural multimodal reasoning and grounding 1 published · 3 in progress

    Benchmarks and methods that test whether a vision-language model understands a cultural scene or has only learned to sound confident about it. Southeast Asia first, because it is the hardest available test case.

  • Debiased and semantic text-video retrieval 5 published

    Video retrieval models take shortcuts: clip length, object co-occurrence, verb priors. My doctoral work identified these shortcuts and removed them with semantic role structure and causal intervention.

  • Long-video retrieval and multimodal planning 1 published · 1 in progress

    Grounding language-model plans in video, and finding the right moment inside hours of egocentric footage.

Selected work

All publications →

Awards and honours

  • Joint 3rd Place, EPIC-Kitchens-100 Multi-Instance Retrieval Challenge, CVPR 2022
  • Finalist, Three Minute Thesis (3MT), Nanyang Technological University, 2022
  • SINGA Ph.D. Scholarship, A*STAR, 2020-2024
  • Student Travel Award, European Neural Network Society (ICANN), 2018

Recent news

Full news archive →

Work with me

  • Research collaborators — cultural reasoning and grounding benchmarks, culturally-aware VLMs, Southeast Asia datasets. Book a 30-minute chat.
  • Students — internships and research mentorship at SMU on multimodal AI. Email me with your CV and a short note on what you would like to work on.
  • Industry and talks — invited talks, media, and projects on cultural AI evaluation and model design. Email me or book a slot.