Context-aware video caption generation with VideoBERT and  domain-aware adapters

Main Article Content

Sergii Volodymyrovych Mashtalir
Mariia Sergiivna Novichonok

Abstract

Background. Automated video description is a relevant task for workplace reengineering and business process analysis, yet its  real-world efficiency is affected by domain shift between training datasets and deployment environments. Aim. The study aims to  develop a context-aware video caption generation architecture combining pretrained video representations with lightweight  adaptation mechanisms. Methods. The proposed framework integrates a frozen twelve-layer VideoBERT encoder with a dual-branch  Domain-Aware Adapter adaptation module and a three-layer autoregressive Transformer decoder. Visual features are extracted using  an ImageNet-pretrained ResNet-50 backbone. The Domain-Invariant Adapter and Domain-Specific Adapter branches process  contextualized representations in parallel and are combined through a trainable scalar fusion coefficient. A fifty percent stochastic  Masked Language Modeling strategy is applied during training. Results. Evaluation on the Something-Something V2 benchmark  shows that the complete VideoBERT, Domain-Aware Adapter and Transformer decoder configuration with fifty percent Masked  Language Modeling achieves a Mean BLEU-4 of zero point one four six nine, Mean ROUGE-L of zero point three four seven four,  and Mean METEOR of zero point two seven two zero. Component-wise ablation demonstrates improvements from both the dual branch adaptation module and Masked Language Modeling regularization. Conclusions. The proposed architecture provides a  modular, parameter-efficient framework for context-aware video caption generation, while explicit cross-domain transfer evaluation  remains a direction for future research.

Downloads

Download data is not yet available.

Article Details

Section

Informatics and intelligent information technologies

Author Biographies

Sergii Volodymyrovych Mashtalir, Kharkiv  National University of Radio Electronics,14, Nauky Ave. Kharkiv, 61166, Ukraine

Doctor of Engineering Science, Professor of Informatics Department

Scopus Author ID: 36183980100

 

Mariia Sergiivna Novichonok, Kharkiv National University of Radio  Electronics,14, Nauky Ave. Kharkiv, 61166, Ukraine

PhD student of Informatics Department

How to Cite

Context-aware video caption generation with VideoBERT and  domain-aware adapters. (2026). Informatics. Culture. Technology, 3(1 (3), 221–235. https://doi.org/10.15276/ict.03.2026.19

References