Context-aware video caption generation with VideoBERT and domain-aware adapters
Main Article Content
Abstract
Background. Automated video description is a relevant task for workplace reengineering and business process analysis, yet its real-world efficiency is affected by domain shift between training datasets and deployment environments. Aim. The study aims to develop a context-aware video caption generation architecture combining pretrained video representations with lightweight adaptation mechanisms. Methods. The proposed framework integrates a frozen twelve-layer VideoBERT encoder with a dual-branch Domain-Aware Adapter adaptation module and a three-layer autoregressive Transformer decoder. Visual features are extracted using an ImageNet-pretrained ResNet-50 backbone. The Domain-Invariant Adapter and Domain-Specific Adapter branches process contextualized representations in parallel and are combined through a trainable scalar fusion coefficient. A fifty percent stochastic Masked Language Modeling strategy is applied during training. Results. Evaluation on the Something-Something V2 benchmark shows that the complete VideoBERT, Domain-Aware Adapter and Transformer decoder configuration with fifty percent Masked Language Modeling achieves a Mean BLEU-4 of zero point one four six nine, Mean ROUGE-L of zero point three four seven four, and Mean METEOR of zero point two seven two zero. Component-wise ablation demonstrates improvements from both the dual branch adaptation module and Masked Language Modeling regularization. Conclusions. The proposed architecture provides a modular, parameter-efficient framework for context-aware video caption generation, while explicit cross-domain transfer evaluation remains a direction for future research.

