ADVC: Adversarial dense video captioning with unsupervised pretraining

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Dense video captioning involves detecting and describing events that represent a video story in untrimmed videos using sentences. This task holds great promise for various video analytics-related applications. However, the nondeterministic nature of dense video captioning poses challenges in generating realistic events and captions. Recently, with the advent of large-scale video datasets, pretraining approaches have emerged. Nevertheless, these methods still require strict supervision and often lack accurate localization or are tightly coupled with localization and captioning. To address these challenges, this paper introduces ADVC, a novel approach for dense video captioning that combines unsupervised pre-training and adversarial adaptation. ADVC learns from readily available unlabeled videos and text corpora at scale, thereby reducing the need for strict supervision. It achieves realistic outcomes by directly learning the distribution of human-annotated events and captions through adversarial adaptation. Adversarial adaptation allows for the decoupling of localization and captioning subtasks while effectively considering their interdependence. We evaluate the performance of ADVC using multiple benchmark datasets to showcase the efficacy of our unsupervised pre-training and adversarial adaptation approach.

키워드

Dense video captioningGenerative adversarial networksNondeterminismUnsupervised learning
제목
ADVC: Adversarial dense video captioning with unsupervised pretraining
저자
Yoon, JongwonChoi, WangyuChen, Jiasi
DOI
10.1016/j.imavis.2025.105595
발행일
2025-09
유형
Article
저널명
Image and Vision Computing
161
페이지
1 ~ 10