Enhancing Video Corpus Moment Retrieval in Long Ego-centric Videos with LLM and Audio Fusion
Published in In preparation, 2026
Finding the right moment inside hours of egocentric video by fusing audio with language-model reasoning over the visual stream.
Status
This work is under development — no preprint and no results to report yet.
If it overlaps something you are working on, get in touch; I would rather compare notes early than collide at submission.
Related
It continues the video corpus moment retrieval work I began during a research visit to Dima Damen’s group at the University of Bristol, and follows the retrieval line of Frame Length Bias (BMVC 2023) and An Overview of Challenges in Egocentric Text-Video Retrieval (CVPR 2023 Workshop) into longer, untrimmed footage.
