Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Apple Machine Learning Research Version 1 original current

Imported from official source

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. …

This version

Version
1 of 1
Recorded
September 20, 2026 19:52
Change
Initial
Content hash
a1e02ddbd66a9ddfd6c0ff7f08c46520
All versions
Revision history

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.