Comparison of news flow processing complexity levels in the task of  stock volatility forecasting

Main Article Content

Kateryna O. Nytska

Abstract

The subject of this study is the forecasting of stock volatility using news as a potential source of additional  information to improve the quality of forecasting. Methods for utilizing news flow vary in terms of both implementation  complexity and computational cost; therefore, the objective of this study is to determine the level of complexity in news  processing beyond which the increase in forecasting accuracy ceases. The study utilized data on stocks from companies  in a single industry over a fourteen-year period and approximately forty-nine thousand news reports, with varying  depths of news coverage across assets. The logarithm of the daily variance was forecast, estimated based on four price  values per trading day: the opening, high, low and closing prices of the stocks. The baseline forecasting model was a  heterogeneous autoregression with three main components corresponding to the daily, weekly, and monthly averaging  horizons of the variable. To compare the results, conditional heteroskedasticity models, a moving average model, and a  naive forecast serving as a lower bound for accuracy were also used. The news flow was fed into the model as an  additional regressor at three levels of processing complexity: the first level – the number of news items per day and the  deviation of this number from the expected level; the second – the sentiment of the news, determined using a lexicon based method; the third – the sentiment determined using a financial language model. Sentiment was represented by the  same set of variables in both methods; therefore, the difference between these levels reflects the quality of sentiment  estimation rather than the number of estimated coefficients. Model parameters were estimated using a rolling window,  and forecasts were generated one day in advance for days not included in the estimation. Since the target variable itself  is measured with an error, the quality of the forecast was assessed using a metric that is robust to this error, and the  difference in accuracy between the models was verified using a formal test for comparing predictive accuracy. It was  found that the deviation of news volume from the expected level, rather than the sentiment of the news, was informative  for the forecast. The difference between the lexicon-based method for detecting news sentiment and the financial  language model is statistically insignificant for the stocks under consideration, despite the latter requiring four orders of  magnitude more computational resources. After correction for multiple comparisons, only the simplest specification  significantly improves the forecast; however, its contribution was approximately sixty times smaller than that of the  stock's own volatility history. Prospects for further research include validating the obtained results in market risk  assessment and applying a large language model as the next level of news flow processing for forecasting. 


 

Downloads

Download data is not yet available.

Article Details

Section

Informatics and intelligent information technologies

Author Biography

Kateryna O. Nytska, Національний технічний університет України «Київський політехнічний інститут імені Ігоря Сікорського», пр.  Берестейський, 37. Київ, 03056, Україна

Master's Student, Department of Mathematical Methods of System Analysis. 

How to Cite

Comparison of news flow processing complexity levels in the task of  stock volatility forecasting. (2026). Informatics. Culture. Technology, 3(1 (3), 166–178. https://doi.org/10.15276/ict.03.2026.14

References