RESEARCH AND EDUCATION ISSN: 2181-3191 VOLUME 4 | ISSUE 9 | 2025 Multidisciplinary Scientific Journal October, 2025 64 DOI: https://doi.org/10.5281/zenodo.17503299 STATISTICAL ANALYSIS AND FORECAST OF THE DYNAMICS OF COTTON YIELD Turgunov Abrorjon Makhamatsolievich PhD, Associate Professor, Tashkent University of Information Technologies named after Muhammad al-Khwarizmi,
[email protected] Xudoyqulov Quvondiq Oybek ugli FullStack Software Engineer
[email protected] Abstract: Observations changing over time, conducted in an ordered sequence, are called a time series. In the article, the statistical regularity of the series of dynamics is studied - the average cotton yield in the Namangan region of the Republic of Uzbekistan (according to the data provided by the Central Statistical Office of the Republic of Uzbekistan for 1991-2018); the study was conducted using the method of statistical analysis of time series. Point and interval estimates for the average cotton yield were constructed, with a 95% guarantee, explicit trends were identified, and yield in the region was predicted for subsequent years. It was found, using the statistical Durbin-Watson tests, that the average cotton yield in the region has an autocorrelation relationship. Keywords: discrete, dynamic, series, trend, seasonality, component, linear, hypothesis, autocorrelation, asymmetry, kurtosis. Аннотация: Вақт бўйича ўзгариб турувчи, тартибли равишда ўтказилган кузатувлар вақт қаторлари деб аталади. Мақолада Ўзбекистон Республикаси Наманган вилоятидаги ўртача пахта ҳосилининг динамика қаторининг статистик қонуниятлари ўрганилган (1991-2018 йиллар учун Ўзбекистон Республикаси Марказий статистика идораси томонидан берилган маълумотлар асосида); тадқиқот вақт қаторларини статистик таҳлил қилиш усулида олиб борилган. Ўртача пахта ҳосили учун нуқта ва интервал баҳолари қурилган, 95% ишонч билан аниқ тенденциялар белгиланган ва кейинги йилларда
RESEARCH AND EDUCATION ISSN: 2181-3191 VOLUME 4 | ISSUE 9 | 2025 Multidisciplinary Scientific Journal October, 2025 65 ҳосил прогноз қилинган. Статистик Durbin-Watson тестлари орқали таҳлил қилинганда, вилоятдаги ўртача пахта ҳосилининг автокорреляцион алоқаси борлиги аниқланган. Калит сўзлар: дискрет, динамик, қатор, тенденция, мавсумийлик, компонент, линей, гипотеза, автокорреляция, асимметрия, куртозис. Аннотация: Наблюдения, изменяющиеся во времени и проводимые в упорядоченной последовательности, называются временными рядами. В статье изучается статистическая закономерность ряда динамики средней урожайности хлопка в Наманганской области Республики Узбекистан (по данным, предоставленным Центральным статистическим управлением Республики Узбекистан за 1991-2018 годы); исследование проводилось с использованием метода статистического анализа временных рядов. Были построены точечные и интервальные оценки средней урожайности хлопка с гарантией 95%, выявлены явные тенденции и спрогнозирована урожайность в регионе на последующие годы. С помощью статистических тестов ДарбинаУотсона установлено, что средняя урожайность хлопка в регионе имеет автокорреляционную зависимость. Ключевые слова: дискретный, динамический, ряд, тенденция, сезонность, компонент, линейный, гипотеза, автокорреляция, асимметрия, куртозис. INTRODUCTION In almost every area, there are phenomena that are important to study in their development and change over time. We can, for example, strive to predict the future based on knowledge of the past, control the processes; describe the characteristic features of a series based on a limited amount of information. When processing time series, the methods are mainly based on the methods developed by mathematical statistics for distribution series. To date, statistics has a variety of methods for analyzing time series from the most elementary methods to very complex ones [1-4]. In general cases, time series {𝑦𝑡,𝑡∈𝑇}consists of four components: trend; fluctuations in relation to the trend; seasonality effect; random component [1-5]. The study of the crop productivity of agricultural processes, as a discrete time series and the forecast of their productivity on the basis of experimental data, play an important role in determining the economic efficiency of farms, deckhand farms. In this article, the processing and analysis of cotton yields for the observation period 1991-2018 in the Namangan region of Uzbekistan was conducted, as a discrete time series. Using the methods of statistical analysis of time series, point and interval estimates for the average yield of cotton were constructed, explicit types of trends were
RESEARCH AND EDUCATION ISSN: 2181-3191 VOLUME 4 | ISSUE 9 | 2025 Multidisciplinary Scientific Journal October, 2025 66 determined and the yield was predicted for subsequent years; various statistical hypotheses were tested. The research conducted by Anderson [1], Kendal [2], Tikhomirov [3], Sulaimanov [4] and others was devoted to the study and analysis of time series. Analysis of results and examples The geometric image of the observed data (Table 1, column 3), the coordinate system give a basis for the first-approximation hypothesis that the trend part of the process has a linear dependence (Fig. 1) of form where unknown parameters are determined by the least squares method, i.e. based on experimental data, when solving the following system of normal equations: {𝑎0𝑇+𝑎1∑𝑡=∑𝑦𝑡; 𝑎0∑𝑡+𝑎1∑𝑡2=∑𝑦𝑡𝑡. (1) Using the calculations based on the data of Table 1, we have: ∑𝑦𝑡=740,9c/ha, 𝑎0=1 𝑇∑𝑦𝑡=740.9 28 =26,46c/ha, 𝑎1=1 ∑𝑡2∑𝑦𝑡𝑡=175.2 1834=0,096c/ha. Hence, the equation of the linear trend (tendency) of cotton yield in the Namangan region is 𝑦(𝑡)=0,096𝑡+26,46 (2) Substituting 𝑡=3 into equation (2), we determine the expected cotton yield in the Namangan region in 2021, as equal, on average, to 27 kg/ha. TableI: Calculation of data to determine the trend of the time series 1 2 3 4 5 6 7 NN Years of observation 𝑦𝑡 c/hа 𝑡 𝑡2 𝑦𝑡⋅𝑡 𝑦𝑡⋅𝑡2 1 1991 30,9 -13 169 -401,7 5222,1 2 1992 28,6 -12 144 -343,2 4118,4 3 1993 28,7 -11 121 -315,7 3472,7 4 1994 29,8 -10 100 -298 2980 5 1995 30,6 -9 81 -275,4 2478,6 6 1996 23,6 -8 64 -188,8 1510,4 10 ()y t a t a=+
RESEARCH AND EDUCATION ISSN: 2181-3191 VOLUME 4 | ISSUE 9 | 2025 Multidisciplinary Scientific Journal October, 2025 67 7 1997 29 -7 49 -203 1421 8 1998 22,7 -6 36 -136,2 817,2 9 1999 25,2 -5 25 -126 630 10 2000 29,9 -4 16 -119,6 478,4 11 2001 27,3 -3 9 -81,9 245,7 12 2002 25,9 -2 4 -51,8 103,6 13 2003 18,7 -1 1 -18,7 18,7 14 2004 21,8 0 0 0 0 15 2005 27,3 1 1 27,3 27,3 16 2006 24,9 2 4 49,8 99,6 17 2007 25,2 3 9 75,6 226,8 18 2008 23,6 4 16 94,4 377,6 19 2009 27,6 5 25 138 690 20 2010 28 6 36 168 1008 21 2011 29 7 49 203 1421 22 2012 28,2 8 64 225,6 1804,8 23 2013 28,2 9 81 253,8 2284,2 24 2014 28 10 100 280 2800 25 2015 28,6 11 121 314,6 3460,6 26 2016 23,4 12 144 280,8 3369,6 27 2017 22,5 13 169 292,5 3802,5 28 2018 23,7 14 196 331,8 4645,2 Total 740,9 14 1834 175,2 49514 Finite differences were calculated based on the observed data (Table 2). 1, t t t Y Y Y + = − 21, t t t Y Y Y + = − 3 2 2 1t t t Y Y Y + = −
RESEARCH AND EDUCATION ISSN: 2181-3191 VOLUME 4 | ISSUE 9 | 2025 Multidisciplinary Scientific Journal October, 2025 68 According to Table 2, the coefficients of variation of the differences 𝑉𝑘= ∑(𝛥𝑘𝑦𝑡)2 𝑇 𝑡=𝑘 (𝑇−𝑘)𝐶2𝑘 𝑘 are calculated; and it is determined that 𝑉1≈𝑉2≈𝑉3. Therefore, first-order finite differences eliminate the linear trend. The presence of autocorrelation in the time series of cotton yield is checked using the Durbin - Watson criterion: 𝑑=∑(𝑌𝑡+1−𝑌𝑡)2 𝑇 𝑡=1 /∑𝑌𝑡2. 𝑇 𝑡=1 TableII: Calculating data for determining finite differences 1 2 3 4 5 6 7 8 9 Years of observation 𝑌(𝑡) c/ha 𝑌𝑡2 𝛥𝑌𝑡 𝛥𝑌𝑡2 𝛥2𝑌𝑡 𝛥2𝑌𝑡2 𝛥3𝑌𝑡 𝛥3𝑌𝑡2 1991 30,9 954,81 1992 28,6 817,96 -2,3 5,29 1993 28,7 823,69 0,1 0,01 -2,2 4,84 1994 29,8 888,04 1,1 1,21 1,2 1,44 -1,1 1,21 1995 30,6 936,36 0,8 0,64 1,9 3,61 2 4 1996 23,6 556,96 -7 49 -6,2 38,44 -5,1 26,01 1997 29 841 5,4 29,16 -1,6 2,56 -0,8 0,64 1998 22,7 515,29 -6,3 39,69 -0,9 0,81 -7,9 62,41 1999 25,2 635,04 2,5 6,25 -3,8 14,44 1,6 2,56 2000 29,9 894,01 4,7 22,09 7,2 51,84 0,9 0,81 2001 27,3 745,29 -2,6 6,76 2,1 4,41 4,6 21,16 2002 25,9 670,81 -1,4 1,96 -4 16 0,7 0,49 2003 18,7 349,69 -7,2 51,84 -8,6 73,96 -11,2 125,44 2004 21,8 475,24 3,1 9,61 -4,1 16,81 -5,5 30,25 2005 27,3 745,29 5,5 30,25 8,6 73,96 1,4 1,96 2006 24,9 620,01 -2,4 5,76 3,1 9,61 6,2 38,44 2007 25,2 635,04 0,3 0,09 -2,1 4,41 3,4 11,56 2008 23,6 556,96 -1,6 2,56 -1,3 1,69 -3,7 13,69 2009 27,6 761,76 4 16 2,4 5,76 2,7 7,29 2010 28 784 0,4 0,16 4,4 19,36 2,8 7,84 2011 29 841 1 1 1,4 1,96 5,4 29,16 2012 28,2 795,24 -0,8 0,64 0,2 0,04 0,6 0,36 2013 28,2 795,24 0 0 -0,8 0,64 0,2 0,04 2014 28 784 -0,2 0,04 -0,2 0,04 -1 1 2015 28,6 817,96 0,6 0,36 0,4 0,16 0,4 0,16 2016 23,4 547,56 -5,2 27,04 -4,6 21,16 -4,8 23,04 2017 22,5 506,25 -0,9 0,81 -6,1 37,21 -5,5 30,25 2018 23,7 561,69 1,2 1,44 0,3 0,09 -4,9 24,01 Total 740,9 19856,19 -7,2 51,84 -13,3 405,25 -18,6 463,78
RESEARCH AND EDUCATION ISSN: 2181-3191 VOLUME 4 | ISSUE 9 | 2025 Multidisciplinary Scientific Journal October, 2025 69 𝑑наб=0,0026 is calculated by the formula (3) and compared with the table value 𝑑крит=1,08 [4]. Since 𝑑наб=0,0026<𝑑крит=1,08, the average cotton yield in the region has an autocorrelation relation TableIII: To the calculation of data to determine the indices of autocorrelation relation 1 2 3 4 5 6 7 T t Y 1tt YY + 2tt YY + 3tt YY + 4tt YY + 5tt YY + 1991 30,9 1992 28,6 883,74 1993 28,7 820,82 886,83 1994 29,8 855,26 852,28 920,82 1995 30,6 911,88 878,22 875,16 945,54 1996 23,6 722,16 703,28 677,32 674,96 729,24 1997 29 684,4 887,4 864,2 832,3 829,4 1998 22,7 658,3 535,72 694,62 676,46 651,49 1999 25,2 572,04 730,8 594,72 771,12 750,96 2000 29,9 753,48 678,73 867,1 705,64 914,94 2001 27,3 816,27 687,96 619,71 791,7 644,28 2002 25,9 707,07 774,41 652,68 587,93 751,1 2003 18,7 484,33 510,51 559,13 471,24 424,49 2004 21,8 407,66 564,62 595,14 651,82 549,36 2005 27,3 595,14 510,51 707,07 745,29 816,27 2006 24,9 679,77 542,82 465,63 644,91 679,77 2007 25,2 627,48 687,96 549,36 471,24 652,68 2008 23,6 594,72 587,64 644,28 514,48 441,32 2009 27,6 651,36 695,52 687,24 753,48 601,68 2010 28 772,8 660,8 705,6 697,2 764,4 2011 29 812 800,4 684,4 730,8 722,1 2012 28,2 817,8 789,6 778,32 665,52 710,64 2013 28,2 795,24 817,8 789,6 778,32 665,52 2014 28 789,6 789,6 812 784 772,8 2015 28,6 800,8 806,52 806,52 829,4 800,8 2016 23,4 669,24 655,2 659,88 659,88 678,6 2017 22,5 526,5 643,5 630 634,5 634,5 2018 23,7 533,25 554,58 677,82 663,6 668,34 Total 740,9 18943,11 18233,21 17518,32 16681,33 15854,68
RESEARCH AND EDUCATION ISSN: 2181-3191 VOLUME 4 | ISSUE 9 | 2025 Multidisciplinary Scientific Journal October, 2025 70 Using the data given in Table 3 and formulas from the literature sources [1-4], we determine the values of the autocorrelation coefficients 𝑅𝐿𝑓𝑜𝑟𝐿=1,2,3,4,6 (where 𝐿лаг is the time shift, i.e., the time lag between one phenomenon and another ass 𝑅𝐿=∑𝑌𝑡𝑌𝑡+1 𝑁−𝐿 𝑡=1 −∑𝑌𝑡 𝑁−𝐿 𝑡=1 ∑𝑌𝑡 𝑁 𝑡=𝐿+1 𝑁−𝐿 √[∑𝑌𝑡2− 𝑁−𝐿 𝑡=1 (∑ 𝑌𝑡 𝑁−𝐿 𝑡=1 )2 𝑁−𝐿 ][∑𝑌𝑡2− 𝑁 𝑡=𝐿+1 (∑ 𝑌𝑡 𝑁 𝑡=𝐿+1 )2 𝑁−𝐿 ] The difference of from zero gives reason to believe that there is a significant autocorrelation relationship between the cotton yields. Based on the sample data, using the x7.2019 program package and Excel [5], numerical characteristics were calculated for the average cotton yield of the Namangan region (Table 4): TableIV: Assessment of the main parameters of the time series Sample characteristics Estimates of sample characteristics Average cotton yield 𝑦𝑇 c / ha 26,46 Dispersion 9,31 Root mean square deviation 𝜎𝑇 3,05 The coefficient of variation 𝜈(%) 11,52% Asymmetry 𝐴𝜍 -0,66 Kurtosis 𝐸𝐾𝜍 -0,18 Error of mean value of 𝑦𝑇,𝑚𝑦 𝑚𝑦=𝜎𝑦 √𝑛=0,58 Marginal error 𝑚𝑦 ′ 𝑚𝑦 ′=𝑡𝑚𝑦=2,06⋅0,58= 1,20 Standard deviation error 𝜎𝑇 𝑚𝜎=𝜎 √2𝑛=3,05 7,48=0,41 Interval estimate (95% ) 𝑦𝑇±𝑡𝑚𝑦 for cotton yield 𝑦𝑇±𝑡𝑚𝑦=26,46±1,20 (25,26; 27,66 ) c/hа Statistical Hypothesis Testing 𝐻0:𝑃(𝑋<𝑥)−Φ𝑎,𝜎(𝑥) 95% guarantee of hypothesis 𝐻0 accepted L R t y
RESEARCH AND EDUCATION ISSN: 2181-3191 VOLUME 4 | ISSUE 9 | 2025 Multidisciplinary Scientific Journal October, 2025 71 CONCLUSIONS Based on the above statistical analyses, and defined dynamics 𝑦𝑡 of the cotton yield in the Namangan region as a time series with reliability 𝛾=0,95, the following conclusions could be drawn: 1) point and interval statistical estimates were constructed for the average cotton yield (25.26; 27.66) c/ha; 2) the explicit types of the trend were determined and its linearity was established; 3) it was established by the Durbin-Watson criterion that there exists a significant autocorrelation of the cotton yield. REFERENCES 1. T. Anderson, “Statistical analysis of time series,” Moscow: MIR, 1976, 759 p. 2. M. Kendal, A. Stewart “Multivariate statistical analysis and time series,” Moscow: Nauka, 1976, 736 p. 3. N.P. Tikhomirov, E.Yu. Dorokhina, “Econometrics,” Moscow: Textbook. “Exam”, 2003, 512 p. 4. B.A. Sulaimonov, A.A. Faiziev, Zh.N. Fayziev, “Tajriba malumotlarining statisticians told,” Tashkent: TashDAU, 2015, 124 p. 5. M.U. Achilov, A.A. Fayziev, “The analysis of dynamics of fruits and berry productivity grown in Uzbekistan,” EPRA Int. J. of Research and Development (IJRD), 4(8), 2019, p. 5-9.