Enhancing Video Corpus Moment Retrieval in Long Ego-centric Videos with LLM and Audio Fusion

Published in In preparation, 2026

Finding the right moment inside hours of egocentric video by fusing audio with language-model reasoning over the visual stream.

Status

This work is under development — no preprint and no results to report yet.

If it overlaps something you are working on, get in touch; I would rather compare notes early than collide at submission.

It continues the video corpus moment retrieval work I began during a research visit to Dima Damen’s group at the University of Bristol, and follows the retrieval line of Frame Length Bias (BMVC 2023) and An Overview of Challenges in Egocentric Text-Video Retrieval (CVPR 2023 Workshop) into longer, untrimmed footage.