Analysis of network traffic representation models for training diffusion models in computer network management tasks
Main Article Content
Abstract
The paper examines the use of network traffic representation models as training objects for diffusion generative models in computer network analysis and management tasks. The relevance of the study is determined by the fact that real network data are often characterized by missing values, irregular measurements, distortions, storage limitations, and insufficient representation of particular network operating conditions. Under such circumstances, diffusion models can be used both to generate synthetic samples and, in a conditional setting, to probabilistically reconstruct missing components; however, the way they are applied depends on the structure of the training object. Based on packet-level, flow-level, telemetry, traffic-matrix, and graph-based representations of network data, the correspondence between feature types, temporal and structural organization of the data, and the choice of continuous, discrete, or combined diffusion modeling is analyzed. The process of selecting a representation according to the properties of the network process required to solve a target task is formalized, including the subsequent formation of the training object and dataset, specification of the diffusion modeling setting, model training, verification, and use of the obtained result. Domain-specific constraints of the corresponding representations are analyzed, including protocol and sequential correctness of packet data, consistency of aggregated flow characteristics, semantics of telemetry indicators and counters, spatiotemporal dependencies of traffic matrices and, depending on the target task, their consistency with routing and resource context, as well as the consistency of dynamic characteristics with the topological structure of graph-based representations. A distinction is proposed between homogeneous training objects formed within a single representation and heterogeneous training objects combining components of several network data representations. It is shown that the heterogeneity of a training object is not equivalent to the heterogeneity of feature types and should be considered separately from the use of additional known information as a condition for generation or reconstruction. A probabilistic formulation for reconstructing incomplete network observations from the known part of the data and additional context is formalized. It is substantiated that the reconstructed result should be interpreted as a conditionally plausible network state rather than as a guaranteed recovery of the actual missing values. Three complementary levels of criteria are generalized for verifying the results of diffusion-based generation and reconstruction: statistical, network-domain, and application-level. The obtained results can serve as a methodological basis for selecting a network data representation, forming a training object, and specifying a diffusion modeling setting according to the target computer network analysis or management task.

