scieee AI-readable full text Open interactive document viewer

Methodological Precedence in Health Tech: An Exploratory Case on the Necessity of Descriptive and Inferential Statistics to Validate Inputs for ML/Big Data Epidemiological Models

Roccetti, Marco

Abstract

Abstract: The adoption of sophisticated analytical tools, including Machine Learning (ML) and massive data processing, has accelerated health research. However, a foundational principle asserts that the rigor of these complex methods is dependent on the integrity and validity of the underlying statistical design. We posit that advanced analyses, particularly in epidemiology, must be subsequent to the rigorous verification of basic methodological coherence. This study uses an exploratory case to demonstrate a crucial cautionary principle: complex models amplify, rather than correct, severe methodological flaws. To demonstrate this, we apply standard descriptive and inferential statistical methods (Z-tests, Confidence Intervals, and t-tests) alongside established national epidemiological benchmarks to a recently published cohort study on vaccine outcomes and psychiatric events. Through this approach, we expose multiple, statistically irreconcilable paradoxes within the source data, including implausible incidence rates and profound baseline group imbalances. These findings, proven by inferential statistical evidence, demonstrate that the observed effects (e.g., contradictory Hazard Ratios) are not biological but are mathematical artifacts stemming from uncorrected selection and classification biases in the cohort construction.Our analysis serves as a robust demonstration that the validity of any conclusion drawn from subsequent advanced ML or statistical modeling sourced form health data rests entirely on first passing the test of basic epidemiological consistency.

Full text

1 Research Article (Non-Peer-Reviewed-Preprint) Methodological Precedence in Health Tech: An Exploratory Case on the Necessity of Descriptive and Inferential Statistics to Validate Inputs for ML/Big Data Epidemiological Models Marco Roccetti Department of Computer Science and Engineering University of Bologna, 40126, Italy [email protected] ORCID: 0000-0003-1264-8595 Abstract: The adoption of sophisticated analytical tools, including Machine Learning (ML) and massive data processing, has accelerated health research. However, a foundational principle asserts that the rigor of these complex methods is dependent on the integrity and validity of the underlying statistical design. We posit that advanced analyses, particularly in epidemiology, must be subsequent to the rigorous verification of basic methodological coherence. This study uses an exploratory case to demonstrate a crucial cautionary principle: complex models amplify, rather than correct, severe methodological flaws. To demonstrate this, we apply standard descriptive and inferential statistical methods (Z-tests, Confidence Intervals, and t-tests) alongside established national epidemiological benchmarks to a recently published cohort study on vaccine outcomes and psychiatric events. Through this approach, we expose multiple, statistically irreconcilable paradoxes within the source data, including implausible incidence rates and profound baseline group imbalances. These findings, proven by inferential statistical evidence, demonstrate that the observed effects (e.g., contradictory Hazard Ratios) are not biological but are mathematical artifacts stemming from uncorrected selection and classification biases in the cohort construction.Our analysis serves as a robust demonstration that the validity of any conclusion drawn from subsequent advanced ML or statistical modeling sourced form health data rests entirely on first passing the test of basic epidemiological consistency. 2 Keywords: Biostatistics; ML/Big Data Precedence; Descriptive and Inferential Statistics, Computational Epidemiology, Selection Bias Classification MSC: 62.XX; 62P10; 68T07 Classification ACM: I.2.6; H2; I.6 1. Introduction The modern era of medical research is defined by an escalating reliance on Big Data platforms and Machine/Deep Learning (ML/DL) algorithms [1]. These technologies, ranging from neural networks for medical image analysis to predictive models for population health, promise to uncover subtle associations and forecast patient outcomes with high precision. The prevailing narrative often suggests that the complexity of the computational approach inherently guarantees the robustness and reliability of the conclusions, creating a dangerous methodological inversion where complex modeling precedes basic data validation [2, 3]. In fact, research demonstrates that simpler models, when properly specified, can yield identical or more stable results than overly complex computational methods [4, 5]. However, a core principle of data science remains immutable and should be the sine qua non for any complex analysis: the outcome of any processing, regardless of its computational sophistication, is ultimately constrained by the quality and design integrity of the input data. A flawed methodological foundation, particularly in medical cohort construction or case definition, will not be corrected by the power of complex statistics or ML/DL. Instead, the complexity may mask and amplify the underlying bias. This is why the application of ML/DL and Big Data tools in health research must be rigorously conditioned on the initial validation of the cohort through descriptive and inferential statistical methods and the consistency of its observed epidemiological metrics. For example, in retrospective observational studies utilizing large administrative databases (such as National Health Service cohorts), the crucial challenge lies in achieving covariate balance between comparison groups. Failure to correct for intrinsic, large differences in baseline characteristics (like age, comorbidity status, and health-seeking behavior) introduces severe selection and misclassification bias. The resulting statistical metrics, such as Hazard Ratios (HRs), would then reflect this baseline disparity rather than any true biological effect. 3 In this paper, we leverage the power of basic descriptive and rigorous inferential statistics (e.g., Z-tests and t-tests) alongside established national epidemiological benchmarks to expose a severe, uncorrected selection bias in the cohort construction of a population-based study [6] investigating the association between an intervention (COVID-19 vaccination) and psychiatric outcomes in a South Korean cohort. Our objective is not to perform a complex re-analysis, but to prove that simple methodological scrutiny is the definitive test for validity, a prerequisite that must be satisfied before any subsequent ML/DL or predictive analysis is meaningful. The remainder of the paper is structured as follows: The Materials and Methods section (Section 2) details our consistency check approach, including the use of specific inferential statistical tests. In Section 3, we present the three major epidemiological paradoxes observed in the scrutinized data, which were derived from checks rooted in basic statistical epidemiology. Had these simple consistency checks not been applied, the data could have proceeded to advanced modeling, inevitably producing an invalid model for subsequent clinical inference. Section 4 discusses how advanced models like ML/DL would fail catastrophically on such flawed data, emphasizing the urgent necessity of basic yet robust covariate balancing techniques. Finally, Section 5 provides a summary conclusion of our finding. 2 Materials and Methods 2.1 Data Sources The data analyzed in this report are entirely sourced from the published literature, specifically from the reported incidence rates and baseline characteristics extracted directly from the primary study under examination [6]. This study utilized administrative data to compare outcomes over a three-month follow-up period between a vaccinated group and an unvaccinated control group. Given the nature of the case study, which investigates the association between COVID-19 vaccination and adverse psychiatric events, it is necessary to perform preliminary data validation and calculations to ensure the integrity of subsequent inferences. This analysis will focus specifically on key psychiatric outcomes, specifically: 1) Schizophrenia (ICD-10 F20-F29), 2) Bipolar Disorder (ICD-10 F31), and 3) Anxiety Disorder (ICD-10 F41.x). Instead, the epidemiological benchmarks used for comparison are derived from robust, independent national studies focused on the South Korean population (subject of the 4 investigated case), which provide validated annual prevalence and incidence rates for the psychiatric conditions studied [7, 8, 9]. 2.2 Descriptive Statistics First, we list and explain, here, the descriptive statistics methods, and their corresponding formulas, used for basic epidemiological checks on the study data under scrutiny. 2.2.1 Calculation of Epidemiological Consistency Metrics To assess the validity of the reported incidence rates and Hazard Ratios (HRs), we applied standard statistical methods based on consistency checks and known relationships between epidemiological measures. To begin, all national annual incidence and prevalence rates were normalized to the equivalent three-month period and to the per 10,000 population scale used in the primary study [6] for direct comparison using the conversion formula for annual incidence I(Annual) to estimated quarterly incidence I(Quarterly) expressed as: I(Quarterly), = ,I(Annual) ,/,4,., (1) 2.2.2 Calculation of Expected Upper Bounds For Anxiety Disorders particularly, the reported 12-month prevalence P(Annual) was used to establish an absolute theoretical upper limit for the incidence over three months [9]. Since in computational epidemiology the incidence (new cases) must be lower than its prevalence (total existing cases), the quarterly fraction of the national prevalence serves as the maximum plausible quarterly incidence ( P(Quarterly_max) ) as showhn below: P(Quarterly_max) ,=, P(Annual ),/,4 . (2) 5 2.2.3 Consistency Checks on Hazard Ratios (HRs) The Hazard Ratio (HR) is the ratio of the hazard rates between the vaccinated H(Vaccinated) and unvaccinated H(Unvaccinated) groups. The consistency check examines the simultaneous occurrence of highly disparate HRs (e.g., HR ≫ 1 and HR ≪ 1) for different chronic conditions within the same non-adjusted cohort, to assess if the effect is biological or an artefact due to baseline bias. 2.3 Inferential Statistics After basic descriptive statistics, we list here the inferential statistics methods necessary to develop rigorous hypothesis testing procedures that add inferential confirmation (or simply rejection) to the validity hypotheses of the initial data coming from study [6]. The methods are explained succinctly, but with a listing of the corresponding null and alternative hypotheses [10]. 2.3.1 One-Sample Z-Test for Schizophrenia Incidence The purpose of this inferential test is to statistically assess the validity of the Schizophrenia incidence rate reported in [6] for the vaccinated group, by comparing it against established national epidemiological benchmarks, thereby testing the Null Hypothesis of methodological consistency. The analysis will rely on the following key data and metrics: 1. Observed Rate (i.e., P(Obs)): Extracted directly from [6], showing the 3-month Schizophrenia incidence rate in the vaccinated cohort. 2. Benchmark Rate (i.e., P(Bench)): Derived from the robust national registry study in [7], which establishes the expected annual incidence for Schizophrenia in South Korea. This annual rate is normalized to a 3-month (quarterly) period. 3. Sample Size (N): The exact size of the vaccinated cohort, N(Vac), as reported in [6]. A summary of these figures is reported in Table 1 below. 6 Metric Value (3-month rate/proportion) Source/Reference Calculation of Cases Observed Rate P(Obs) 0.51 / 10,000 = 0.000051 Kim HJ et al. [6] (Vaccinated Cohort) 1,718,999 x 0.000051 ≈, 88 Benchmark Rate P(Bench) ~ 2.1 / 10,000 = 0.00021 Cho et al. [7] (National Incidence Range: 2.02.2/10,000 annually, normalized / 4) 1,718,999 x 0.00021 ≈,361 Cohort Size (N) 1,718,999 Kim et al. [6] (Vaccinated Cohort) - Table 1: Data inputs required to compare the observed 3-month incidence proportion of Schizophrenia in the vaccinated cohort (Kim et al. [6]) against the national epidemiological benchmark (Cho et al. [7]). In this case, the test of interest specifically aims to detect a non-plausible deficit in observed cases, making it a one-tailed Z-test which can be defined as follows: • Null Hypothesis (H0). The observed Schizophrenia incidence proportion in the vaccinated cohort P(Vac) is equal to or greater than the national benchmark proportion P(Bench). This assumes methodological consistency: H0: P(Vac) > = P(Bench). • Alternative Hypothesis (H1). The observed 3-month schizophrenia incidence rate in the vaccinated cohort is significantly lower than the national benchmark rate, suggesting uncorrected bias, H1: P(Vac) < P(Bench). At this point a correct statistical test is needed to decide H0 vs H1. This is the Z-score test statistic which is calculated using the standard formula for comparing a sample proportion to a known population proportion: 𝑍= ,𝑃(𝑂𝑏𝑠)−𝑃(𝐵𝑒𝑛𝑐ℎ) F(𝑃(𝐵𝑒𝑛𝑐ℎ)(1−𝑃(𝐵𝑒𝑛𝑐ℎ)) 𝑁 ⁄⁄ . 2.3.2 Confidence Interval Calculation for Bipolar Disorder Incidence The objective is to establish the statistical precision and plausible range of the observed 3-month Bipolar Disorder (BD) incidence rate reported in the unvaccinated control cohort of [6]. This precision will then be compared against the external 12-month national prevalence rate from [8] to test for methodological consistency. This time, 7 the idea is to use fromal statistical hypothesis testing from inferential statstics to corroborate (or reject) the validation activity begun with simpler techniques of descriptive statistics. We initiate this analysis focusing on the unvaccinated control group, where the anomalous BD incidence rate was observed, as reported in Table 2 below. Metric Value (Proportion or Count) Source/Context Observed Incidence Proportion (P) 0.000139 (from 1.39/10,000) Extracted from [6] (Unvaccinated Control Group, 3month incidence). Prevalence Benchmark (P(Bench)) 0.000106 (from 1.06/10,000) Extracted from [8] (National 12-month Prevalence for similar conditions). Control Cohort Size (N) 308,354 Size of the Unvaccinated Control Group reported in [6]. Observed Cases (O) ≈ 43 Calculated from N x P Table 2: Data inputs comparing the 3-month observed BD incidence proportion (P) in the unvaccinated control cohort [6] against the 12-month BD national prevalence benchmark (P(Bench)) [8]. In this case, we will use the Wald method to calculate the 95% Confidence Interval for the observed BD incidence proportion (P). The formula for the 95% Confidence Interval for the proportion is: 𝐶𝐼(95%)=𝑃,±𝑍!/# ,×,𝑆𝐸, where P is the observed BD incidence proportion (0.000139); 𝑍!/# = 1.96 (i.e., the critical Z-value for a 95% CI) and SE (the Standard Error) which in turn can be calculated as S(%&')' ). At this point, the inferential test involves checking the relative position of the 12-month BD prevalence benchmark (P(Bench)) within the calculated 95% CI of the 3-month BD incidence (P). If the benchmark is statistically consistent with the observed incidence rate, it will fall within the calculated CI. 2.3.3 Two-Sample Independent t-Test for Covariate Balance (Mean Age) The purpose of this inferential test is to check a more general condition that could affect the representativeness of a given cohort. Essentially, we want to determine if the compared cohorts (in [6]) were statistically equivalent on a critical confounding variable, Mean Age, prior to the intervention. Establishing this baseline balance is a necessary prerequisite for valid causal inference in non-randomized studies. In the end, following this way, it will be possible to verify if the scrutinized study [6] yields valid or contradictory HRs. In this latter case, this would reflect the confounding effect of these pre-existing differences, rather than the biological effect of the vaccination 8 intervention. We begin this kind of analysis utilizing the baseline statistics reported in [6] for the Mean Age of the two comparison groups reported in Table 3. Metric Value Source/Context ([6]) Vaccinated Mean Age (𝑋 U(Vac)) 54.67 years Reported mean age for the vaccinated group. Vaccinated SD (SD(Vac)) 16.26 Reported standard deviation for the vaccinated group. Vaccinated Cohort Size (N(Vac)) 1,718,999 Cohort size used for the t-test. Non-Vaccinated Mean Age (𝑋 U(NonVac)) 44.18 years Reported mean age for the non-vaccinated group. Non-Vaccinated (SD(NonVac)) 16.2 Reported standard deviation for the non-vaccinated group. Non-Vaccinated Cohort Size (N(NonVac)) 308,354 Cohort size used for the t-test. Table 3: Empirical baseline characteristics (Mean Age and Standard Deviation) derived from [6], used to test the assumption of covariate balance between the Vaccinated and Non-Vaccinated sub-cohorts. We use Welch's Independent Samples t-Test to compare the means of the two cohorts. Specifically, we structure the following hypothesis test: • Null Hypothesis (H0). There is no statistically significant difference in the mean age between the Vaccinated and Unvaccinated cohorts (i.e., the groups are balanced for age): H0: 𝑋 U(Vac) = 𝑋 U(NonVac) • Alternative Hypothesis (H1). There is a statistically significant difference in the mean age between the two cohorts (i.e., the groups are unbalanced): 𝑋 U(Vac) ≠ 𝑋 U(NonVac). The final t-test statistic is calculated as: 𝑡=(𝑋 U(𝑉𝑎𝑐)−,𝑋 U(𝑁𝑜𝑛𝑉𝑎𝑐))/S*+(,-.)! )(,-. +*+()/0,-.)! )()/0,-.). 9 3. Results The application of basic consistency metrics to the published data of [6] reveals three fundamental statistical paradoxes that challenge the core findings of that study. If these statistical epidemiological checks, which highlighted the three paradoxes discussed below, had not been performed first, any subsequent analyses or inferences would have been invalid, contradictory, or, at a minimum, irrelevant. 3.1 Paradox I: Implausible Protective Effect for Schizophrenia The study [6] reports a Hazard Ratio (HR) of 0.231 for the development of Schizophrenia (ICD-10: F20-F29) in the vaccinated group compared to the unvaccinated control group. Obviously, an HR below 1.0 would suggest a protective effect, and HR of 0.231 is interpreted as an approximately 77% reduction in the risk of developing Schizophrenia (1 - 0.231). This is the first paradox: there is no biological or clinical justification for a COVID-19 vaccine to confer such a profound, immediate protective effect against a chronic, neurodevelopmental disorder like Schizophrenia. To understand the mathematical source of this implausible finding, we must compare the incidence rates used to calculate this HR against known epidemiological benchmarks like in Table 4 below. Condition Group Reported Incidence [6] (per 10,000 over 3 mo) Crude Ratio (Vac / Unvac) National Benchmark [7] (Quarterly Range per 10,000) Schizophrenia Unvaccinated Control 1.98 Not applicable 2.0 - 2.2 Vaccinated (High-Risk) 0.51 0.257 2.0 - 2.2 Table 4: Reported and Benchmark quarterly incidence rates for Schizophrenia (ICD-10 F20-F29), illustrating severe cohort selection bias. The reported HR of 0.231 is extremely close to the crude ratio of the incidence rates (0.51 / 1.98, approx. 0.257). The small difference exists because the HR is derived from a Cox regression model which incorporates time-toevent data and slight adjustments, whereas 0.257 is a simple rate ratio. Unfortunately, both values signify the same magnitude of disparity. Nonetheless, the unequivocal proof of the methodological error is the incidence rate observed in the Vaccinated cohort (0.51 per 10,000). In fact, this rate is: 16 References 1. Habehh, H., Gohel, S. Current genomics, Machine Learning in Healthcare. Current Genomics 2021, 22(4), [291 - 300]. doi: 10.2174/1389202922666210705124359. 2. Roccetti, M., Delnevo, G., Casini, L., Cappiello, G. Is bigger always better? A controversial journey to the center of machine learning design, with uses and misuses of big data for predicting water meter failures. Journal of Big Data 2019, 6(1), 70. Doi: 10.1186/s40537-019-0235-y 3. Alhumaidi, N. H., Dermawan, D., Kamaruzaman, H. F., Alotaiq, N. The Use of Machine Learning for Analyzing Real-World Data in Disease Prediction and Management: Systematic Review. JMIR Med Inform 2025, 13, e68898. doi: 10.2196/68898. 4. Roccetti, M., Cacciapuoti, G. Beyond the Gold Standard: Linear Regression and Poisson GLM Yield Identical Mortality Trends and Deaths Counts for COVID-19 in Italy: 2021–2025. Computation 2025, 13(10), 233. doi: 10.3390/computation13100233 5. Roccetti, M., De Rosa, E. M. A Segmented Linear Regression Study of Seasonal Profiles of COVID-19 Deaths in Italy: September 2021–September 2024. Computation 2025, 13(7), 165. doi: 10.3390/computation13070165 6. Kim HJ, Kim MH, Choi MG, Chun EM. Psychiatric adverse events following COVID-19 vaccination: a population-based cohort study in Seoul, South Korea. Mol Psychiatry. 2024;29(12):3635–3643. doi: 10.1038/s41380-024-02627-0 7. Cho SJ, Kim J, Kang YJ, Lee SY, Seo HY, Park JE, et al. Annual Prevalence and Incidence of Schizophrenia and Similar Psychotic Disorders in the Republic of Korea: A National Health Insurance Data-Based Study. Psychiatry Investig. 2020 Jan 25;17(1):61–70. doi. 10.30773/pi.2019.0041 8. Shin H, Lee HS, Lee BC, Park G, Uranbileg K. The Prevalence and Clinical Characteristics of Borderline Personality Disorder in South Korea Using National Health Insurance Service Customized Database. Yonsei Med J. 2023;64(9):566–72. doi: 10.3349/ymj.2023.0071. 9. Rim SJ, Hahm B-J, Seong SJ, Park JE, Chang SM, Kim B-S, et al. Prevalence of Mental Disorders and Associated Factors in Korean Adults: National Mental Health Survey of Korea 2021. Psychiatry Investig. 2023;20(3):262–272. doi:10.30773/pi.2022.0307 10. Casella G, Berger RL. Statistical inference. 2nd ed. Pacific Grove (CA): Duxbury; 2002. ISBN: 9780534243128 17 11. D’Agostino RB Jr. Propensity score methods for bias reduction in the comparison of a treatment to a non-randomized control group. Stat Med. 1998;17(19):2265–81. doi: 10.1002/(sici)10970258(19981015)17:19<2265::aid-sim918>3.0.co;2-b Contributions MR conducted the data analysis, conceptualized the statistical arguments, and wrote the paper. Corresponding author Correspondence to Marco Roccetti: [email protected]. Competing interests The author declares no competing interests. Funding This research received no external funding Data Availability The data presented here is either included directly or was extracted from the referenced documents. All calculations are easily reproducible based on the definitions provided Ethics approval and consent to participate This study uses publicly available, aggregated data that contains no private information. Therefore, ethical approval is not required Generative AI statement The author declares that no Gen AI was used in the creation of this manuscript