scieee AI-readable full text Open interactive document viewer

Restoring Reliability in Vaccine Safety Surveillance: Correcting Structural Bias in Medical Cohort Construction

Roccetti, Marco

Abstract

AbstractThe modern era of medical research, characterized by the availability of big data and sophisticated statistical methodologies, is paradoxically vulnerable to fundamental structural flaws in experimental cohort construction, despite adherence to rigorous reporting guidelines. Errors often arise from a failure to apply simple base-rate checks, violating core statistical principles derived from the work of pioneers, and leading to contradictory results. This methodological challenge was recently amplified in large-scale COVID-19 vaccine safety surveillance, where a failure to ensure a non-asymmetric distribution of high- risk elderly or vulnerable individuals across compared subgroups (Vaccinated vs. Non-Vaccinated) leads to a systemic and reproducible failure in cohort construction. This generates predictable and spurious Hazard Ratios (HRs). This scenario is powerfully anticipated by historical failures that remind us of the importance of primary checks. For example, the Beta-Agonist Paradox showed to the world that structural bias (Confounding by Indication) could mistake a marker of severity for a causal risk, teaching us once again that external validity is lost when the core cohort structure is compromised. We address this specific issue within two exemplar studies that have alarmed the international scientific community by associating COVID-19 vaccination with increased risks of diseases, including cancer and autoimmune disorders. These studies share a single data source and, critically, this fundamentally flawed methodology. Our analysis aims to reconstruct the cohort data and quantify the exact effect of a demographic asymmetry on reportedly increased risks of cancer and a common autoimmune disease (vitiligo). The applied methodology involves the deconstruction and re-analysis of reported HRs through weighted incidence analysis to produce corrected and plausible risk estimates, thus re-establishing methodological integrity and providing clinicians and the public with reliable information free from unwarranted alarm.

Full text

Restoring Reliability in Vaccine Safety Surveillance: Correcting Structural Bias in Medical Cohort Construction Marco Roccetti1*& ORCID 1 Department of Computer Science and Engineering, University of Bologna, Bologna, Italy *Corresponding author E-mail: [email protected] (MR) &These authors also contributed equally to this work. Abstract The modern era of medical research, characterized by the availability of big data and sophisticated statistical methodologies, is paradoxically vulnerable to fundamental structural flaws in experimental cohort construction, despite adherence to rigorous reporting guidelines. Errors often arise from a failure to apply simple base-rate checks, violating core statistical principles derived from the work of pioneers, and leading to contradictory results. This methodological challenge was recently amplified in large-scale COVID-19 vaccine safety surveillance, where a failure to ensure a non-asymmetric distribution of highrisk elderly or vulnerable individuals across compared subgroups (Vaccinated vs. Non-Vaccinated) leads to a systemic and reproducible failure in cohort construction. This generates predictable and spurious Hazard Ratios (HRs). This scenario is powerfully anticipated by historical failures that remind us of the importance of primary checks. For example, the Beta-Agonist Paradox showed to the world that structural bias (Confounding by Indication) could mistake a marker of severity for a causal risk, teaching us once again that external validity is lost when the core cohort structure is compromised. We address this specific issue within two exemplar studies that have alarmed the international scientific community by associating COVID-19 vaccination with increased risks of diseases, including cancer and 2 autoimmune disorders. These studies share a single data source and, critically, this fundamentally flawed methodology. Our analysis aims to reconstruct the cohort data and quantify the exact effect of a demographic asymmetry on reportedly increased risks of cancer and a common autoimmune disease (vitiligo). The applied methodology involves the deconstruction and re-analysis of reported HRs through weighted incidence analysis to produce corrected and plausible risk estimates, thus re-establishing methodological integrity and providing clinicians and the public with reliable information free from unwarranted alarm. Keywords: Vaccine Safety Surveillance; Data Representativeness; Structural Flaw; Flawed Methodology; Cohort Bias; Hazard Ratio Decomposition Introduction The modern era of medical research is defined by the unprecedented availability of large-scale datasets, encompassing vast quantities of critical patient information. This wealth of big data allows researchers to employ sophisticated analytical techniques, ranging from fundamental descriptive statistics to complex inferential methods and cutting-edge artificial intelligence (AI). Furthermore, the epidemiological community adheres to robust reporting guidelines, such as STROBE [1] and PRISMA [2], for example, specifically designed to ensure the transparency, reproducibility, and ultimate reliability of experimental results. Moreover, the foundation of modern biostatistics rests on the pivotal contributions of figures such as Ronald Fisher (developer of ANOVA, Maximum Likelihood, and randomized design principles), Karl Pearson (pioneer of correlation and the chi-squared test), Jerzy Neyman (formulator of confidence intervals and hypothesis testing), and David Roxbee Cox (the inventor of the proportional hazards model), just to cite a few [3-6]. Their methodologies, including Regression Analysis, Survival Analysis (e.g., Kaplan-Meier), and Multivariable Modeling (e.g., Logistic and Poisson Regression) among others, have enabled breakthroughs ranging from establishing the link between smoking and lung cancer to calculating vaccine effectiveness and identifying complex genetic risk factors. 3 Despite this arsenal of advanced tools, statistical principles, and established protocols, incongruencies still frequently plague the analysis of large, retrospective medical datasets, the most abundant form of evidence in the scientific world. The reliance on these powerful inferential techniques, which are often poorly understood, can sometimes lead to results that violate the very statistical and epidemiological fundamentals they were designed to uphold. Errors originating from fundamental flaws in the organization or structure of this data and the resulting experimental cohorts, particularly when prepared for inferential analysis, can violate core statistical principles and render us unable to extract truthful signals from the data, often leading to contradictory or even self-contradictory findings. At this juncture, the solution is not merely the utilization of more data or superior analytical techniques, like machine/deep learning algorithms [7]; it demands instead a return to the basic principles of statistics, including simple descriptive analysis, and the foundational teachings of the pioneers cited before. This requires recovering the ability, diligence, and most of all objectivity to execute fundamental base-rate checks that ensure a non-contradictory or undistorted use of the data. Such an approach prevents the results from being unduly influenced by the predisposition to always reject the null hypothesis, driven by the conviction that vast data quantities must always conceal a sensational new finding, a notion that is by no means guaranteed. This persistent methodological vulnerability is extremely critical when evaluating research concerning public health outcomes. In particular, the necessity for rigorous methodological review of medical data and relative outcomes has been recently amplified by the sheer scale of the COVID-19 vaccination campaigns. Globally, tens of billions of vaccine doses have been administered since 2021, involving hundreds of millions of people across nearly every demographic and healthcare system. This unprecedented scale means that even rare events, when extrapolated across the entire vaccinated population, necessitate meticulous monitoring and analysis [8]. The resulting observational research, utilizing vast administrative and electronic health record databases, has yielded a highly complex and often contradictory picture regarding potential Serious Adverse Events (SAE) [9]. Many large-scale studies have convincingly found no association between COVID-19 vaccination and numerous serious outcomes, for instance, dispelling early concerns 4 regarding generalized increased risks of stroke, pulmonary embolism, or specific neurological disorders across wide cohorts [10-12]. Conversely, there is a minimal but growing body of recent literature asserting the discovery of associations between vaccination and specific SAEs, such as myocarditis, pericarditis, and also increased risks of certain cancers and autoimmune conditions [13-17]. Importantly, many of these affirmative findings, particularly those published in the last two years, remain at the center of intense and unresolved scientific and methodological debate. Often, the reported signals are inconsistent across different geographic regions or cohort definitions, or they fail to replicate robustly. This ambiguity underscores the current challenge: distinguishing a true, consistent medical signal from a statistical artifact or methodological error inherent in the retrospective analysis of massive, often unbalanced, datasets. The prevalence of such unconvinced or conflicting results mandates that researchers must scrutinize the very foundation of the data, the construction and representativeness of the resulting experimental cohort, before accepting any purported clinical association. Background and Historical Precedents To fully contextualize the gravity of these methodological lapses and demonstrate that the issue of flawed cohort construction is not unique to modern big data medical analysis, it is essential to return to medical history. Before proceeding further, we deem it necessary to provide a historical background illustrating how structural data errors can fundamentally mislead clinical science. Over the modern era of medicine, several landmark observational studies have shaken public confidence with contradictory or ultimately debunked results. In nearly every instance, the underlying motivation for the erroneous conclusions rested upon the incorrect or improper use of supporting data, typically through severe selection or confounding bias. Therefore, we cite three prominent cases where methodological failures, rather than true biological effects, produced misleading signals of risk or benefit. Specifically, the challenge of ensuring external validity (as recognized by the STROBE protocol, principle 21 [1]), that is the degree to which a study cohort accurately represents the target population and its baseline risk, is a recurrent and critical issue in observational epidemiology. When a cohort's baseline risk is systematically skewed due to unmeasured or structural selection differences between exposure groups, it leads to a biased estimate of the effect, often creating a spurious association. This 5 phenomenon, where cohort construction errors lead a study to derive misleading clinical conclusions, is a long-standing concern in medicine, as demonstrated by the following precedents. First, we cite the case of Hormonal Replacement Therapy (HRT) and Cardiovascular Risk (The WHI vs. Observational Studies). For over two decades, in fact, highly influential observational studies (such as the Nurses’ Health Study) concluded that Hormone Replacement Therapy (HRT) drastically reduced the risk of myocardial infarction (up to 40–50%) in menopausal women [18]. These seemingly robust findings pushed HRT to become a global standard treatment for millions. The error seems to lay in a severe healthy user bias: women who chose to take HRT were inherently healthier, wealthier, better educated, maintained healthier lifestyles, and were more likely to seek proactive medical care. This systemic non-random selection artificially lowered their cardiovascular risk before starting the therapy, making the HRT group look protected. The definitive randomized controlled trial (RCT), the Women’s Health Initiative (WHI, 2002), yielded the opposite results: HRT actually increased cardiovascular and thrombotic risk [19]. This case remains the most famous example in medicine of how a structural cohort bias created a protective effect that was entirely non-existent, leading to years of suboptimal clinical practice. We presented this case first because it is one of those that has afflicted evidence-based datadriven medicine for several decades, with the arguments in favor of hormonal therapy recently reasserting themselves, albeit in different forms and aimed at specific patient sectors, in a cycle that seems never to end [20]. The second and perhaps most famous case is the one that became known by the name of: The Wakefield Case, in 1998 (also known as MMR Vaccine and Autism). In 1998, Andrew Wakefield published a study, later retracted and proven fraudulent, that suggested an association between the MMR (Measles, Mumps, and Rubella) vaccine and the onset of autism. While the main issues were data manipulation and conflicts of interest, the central methodological failure was the absence of a valid control group and the use of a tiny, selectively chosen sample (n=12), constituting extreme selection bias. This design flaw ensured that any correlation found was highly vulnerable to confounding. Subsequent large-scale epidemiological studies, involving millions of children across multiple countries, confirmed the total lack of association, demonstrating a relative risk of exactly 1. The MMR–autism case 6 has been determinant in clinical medicine for understanding how the incorrect sample construction and selection can produce devastating public harm by eroding trust in established preventive measures [21]. The third and final case (Asthma Mortality - Beta-agonist Paradox) is an equally famous scenario where fraudulence plays no part, while a theme emerges, almost inadvertently due to insufficient medicoprocedural knowledge, regarding the use of data according to the correct protocols and due precedence. In that case, in the 1980s and 1990s, observational studies had suggested that frequent use of inhaled beta-agonists was associated with an increase in asthma mortality. Those findings raised significant alarms regarding the safety of a first-line treatment. In reality, this was a classic example of confounding by indication: patients who were prescribed and thus received the most intensive and frequent doses of the inhaler were simply those with the most severe, life-threatening asthma. The drug therefore appeared to cause the severity, whereas it was only a marker of pre-existing disease severity and high risk. This structural bias in the cohort produced a spurious Hazard Ratio (HR) that falsely suggested the life-saving medication was harmful, leading to widespread misinterpretation of clinical risk [22, 23]. This case serves as a powerful reminder that procedural errors in establishing priority of causes and effects (indication vs. outcome) can generate highly misleading signals, even without any deliberate manipulation of data. More importantly, this scenario is profoundly instructive as it directly anticipates the central thesis of our study: a problem of generalizability (external validity) is completely obscured by the intense focus on controlling internal validity. In the frantic effort to balance lesser confounding elements, researchers lose sight of the broader meaning of the data, failing to recognize that the core cohort structure itself is fundamentally compromised. Our Contribution These precedents illustrate that when a study’s reference group is systematically unrepresentative of the baseline risk, the resulting HRs or risk associations become fundamentally unreliable. In this context, our aim is, as already anticipated, to address problems stemming from improperly handled medical data in the detection of Serious Adverse Events (SAE) within large cohorts of vaccinated (V) and non-vaccinated (NV) individuals, noting a recent tendency to amplify such signals to the detriment 7 of the Vaccinated (V) group. Specifically, the interpretation of safety/alarm signals derived from observational cohorts and relative data relies fundamentally on the common support assumption that the reference (NV) and exposed (V) groups are comparable in underlying health and risk behavior. Unfortunately what we have frequently observed, and what forms the central focus of the present analysis, is the failure to apply a rigorous quantitative pre-analysis on the data to ensure a nonasymmetric presence of high-risk elderly (or otherwise vulnerable) individuals across the two subgroups (V and NV). This omission directly leads to a systemic and reproducible failure in cohort construction that generates predictable, spurious HRs. In this context, a series of studies have recently emerged that have alarmed the medical world and the international community by associating COVID-19 vaccination with an increased risk of a range of diseases, spanning from psychiatric conditions to autoimmune disorders, and cardiovascular-circulatory issues, but most notably to numerous types of cancers [14-17]. What these studies, which have found acceptance in many respected journals from different editorial groups, have in common is not only their origin from a specific nation or their development by an almost singular research group (factors that hold no relevance for our analysis). Rather, the critical commonality is that all these studies share the same initial data source: a single, specific national database (i.e., the South Korean National Health Insurance Service NHIS database). More importantly, all these studies share a fundamentally flawed construction of the V and NV groups. As we will demonstrate, these groups are systematically unbalanced, favoring the NV group by introducing a significantly lower number of elderly individuals (who have a naturally higher incidence of the reported diseases), thereby leading to a suppressed baseline risk for the NV group, independent of COVID-19 vaccination status. This structural imbalance ensures that any subsequent analytical finding is highly prone to producing spurious HRs. Our present study aims thus to reconstruct the cohort data, quantify the effect of demographic asymmetry on reported HRs, and produce corrected estimates that account for these biases for two of the several cases raised by these studies [14, 15]: specifically, vitiligo and cancer. We focus on vitiligo due to the peculiarity of the finding, as this common autoimmune disorder is not readily implicated in the public or common medical imagination with COVID-19 vaccination, making it a compelling test case for spurious association. Conversely, we focus on cancer for the opposite reason: it is one of the leading causes of death worldwide, and thus, even conceptually, is frequently and widely suspected of being 8 associated with major global events and epidemiological factors, making the reported association particularly resonant and potentially alarming to public health. By doing so, we aim to clarify whether the observed associations reflect true biological risk or are in the end simply statistical mirages. To this aim, we have applied a rigorous quantitative analysis to two cohorts of those high-impact studies by employing a rigorous methodology consisting in deconstruction and re-analysis of reported HRs through weighted incidence analysis, focused on correcting the severe demographic bias (age and beyond) in the non-vaccinated subgroup, to reconstruct more plausible and non-distorted risk estimates. This methodology building on standard survival analysis develops a specific Poisson-like regression decomposition technique for hazard rate models, allowing for a flexible representation of the baseline hazard [24-26]. In essence, we use it to dissect the originally observed HRs into components driven by structural factors and uncorrected methodological bias, proving that the latter is often the dominant force. Our analysis of this data has demonstrated, in the end, that both the observed signals of increased vitiligo and cancer risk are highly likely to be a manifestation of a similar structural selection bias, involving the asymmetric underrepresentation of high-risk elderly individuals in the non-vaccinated group, teaching us once again that similar critical flaws in cohort constructions must be addressed before any conclusion regarding biological causality can be entertained. In closing this Section, it must be underscored that our primary focus is not on the veracity of these findings per se, which remain subject to ongoing demonstration despite the significant attention they have garnered. Rather, our critique centers on the fact that both reported associations, potentially originating from the same research groups, exhibit similar and concerning common patterns in the handling of fundamental medical data. Specifically, they demonstrate a pervasive failure to construct balanced patient cohorts capable of adequately controlling the signals that the data appear to emit, even in retrospective analysis. This methodological weakness persists regardless of the sophistication of the adjustment algorithms employed, such as Propensity Score Matching (PSM) used for cancer, or in cases where such complex balancing techniques were not utilized at all (as observed with the vitiligo endpoint) [27]. This structural inadequacy has definitely suggested that the reported signals are not robust. A closer biostatistical inspection of the underlying data has revealed pronounced deviations from national demographic norms, particularly an underrepresentation of high-risk elderly individuals in 9 the non-vaccinated comparator group. This structural imbalance has mathematically generated spurious HR without implying any causal or even associative relationship. Materials and Methods In this Section, we first provide drawn from international literature the reference data for vitiligo and cancer as the starting point of our analysis, respectively based on the two following sets of South Korean studies, [28-30] and [31-35], which include the essential reference points for the present analysis. We will cross-reference this data, which serves as the national benchmark for those two diseases in terms of incidence disease and age distribution, with the data coming from the scrutinized studies [14, 15], representing it in clear tabular formats. After this, we inform the reader about the main metrics used to calculate and interpret the meaning of risk (i.e., HRs) in contexts of retrospective analysis of medical data, also illustrating in mathematical details the methodology used, starting from the aforementioned reference data, to correct the risk values and obtain those adequate for the starting data. Input Data Summary and Sources Our quantitative comparative analysis starts from the key epidemiological parameters extracted from the two exemplar cohorts under investigation, utilizing the South Korean National Health Insurance Service (NHIS) database [14, 15]. The core parameters extracted are the observed Hazard Ratio HR(Observed), the observed cumulative Incidence Rate in the Vaccinated group, IR(V, Observed), and the corresponding rate in the Non-Vaccinated group, IR(NV, Observed). Importantly, these observed rates are contrasted with independently calculated expected National Incidence rates, IR(Baseline), which serve as the gold standard baseline for validating a demographically representative South Korean cohort [29, 33]. We already anticipate here that the IR(NV, Observed) plays often, in misconstructed cohorts, the role of the failing denominator due to the structural and selection biases inherent in the cohort preparation. We begin by providing the input data summary (HR and IR) for both vitiligo and cancer cases studied in [14, 15] in Table 1. 16 dramatically under-reported relative to the national baseline (a -45% deficit in the high-risk >= 65 age group), indicating a massive selection bias. The V group, although also depressed, was less severely biased (-26.9%). This asymmetry in the bias magnitude is the direct source of the spurious HR(residual), with cohorts almost already clearly non-comparable on underlying health status, necessitating a statistical correction of the baseline risk which we will calculate later under the form of a HR(Corrected). This critical adjustment will neutralize, as shown in the Results Section, the denominator failure caused by the asymmetric selection bias suffered by all the cohort]15], and will provide a non-distorted estimate of the association, directly testing the null hypothesis against the corrected baseline. Ethical Approval and Data Statement This study did not require ethical approval because it involved no humans, animals, plants, relying instead on publicly available, aggregated data containing no private information. The data were included directly in the manuscript's Tables or extracted from cited literature [14, 15, 28-33]. This information has been explicitly reiterated in the relative Statements at the bottom of the paper. Results We now provide our final results in terms of the real risk that the data from the database employed in [14, 15] should have revealed if properly managed. We provide a concise summary of the HRs results in the third Subsection, while in the first two, respectively for vitiligo and cancer, we show how these correctly calculated risk values can be obtained, and discussed as well. Vitiligo Results: Failure and Correction This Section presents the results documenting how we can arrive at the conclusion that either the increased vitiligo risk disappears for the V group, adopting the study's logical framework discussed earlier, or the reasoning developed in [14] can be falsified. We start by making the critical assumptions that the observed cumulative incidence rates of 2.22 (V) and 0.67 (NV) per 10.000 were implicitly annualized and harmonized and are therefore directly comparable to the annual expected incidence rates of [29] which was calculated as equal to 2.473. We 17 also assume that the impact of excluding the under-20 age group is negligible for this primary finding, in any case, remembering that the exclusion from a vitiligo study of the group (< 20 years) that experiences the first peak of the disease is in itself quite unusual. As already anticipated (see Table 2 and Equation 2), the expected incidence rates per 10.000 for both groups (V, NV) of the investigated cohort can be calculated as the two components of the ratio that will lead to the HR(Structural), amounting respectively to 2.6804 (V) and 2.2146 (NV). The ratio of these two components yields the approximated value of HR(structural) = 1.21. This first results shows a first failure caused by the relevant demographic imbalance between the Vaccinated (V) and Non-Vaccinated (NV) groups, because of the use of an insufficient Random Selection matching technique to create the cohort. In some sense, it is said that the risk increases by a 21% only in force of an age disparity between the two subgroups. Simply said, the increased observed risk of 21% is mathematically attributable solely to the difference in the age structure, specifically the higher proportion of older, highrisk individuals in the V cohort. By applying Equation (1), we can derive now the value of the HR(Residual), that is the association that remained after mathematically eliminating the age bias. We achieve a value of 2.24, that is HR(Observed) / HR(Structural) = 2.714 / 1.21 = 2.24. The resulting residual risk of 2.24 confirms a profound failure of the simple random matching technique. Even after adjusting for age, a substantial unexplained difference remained, proving the cohorts were not truly comparable and confirming a qualitative failure in the initial study design. To find the true, non-biased association, we perform the definitive correction. This involved discarding the faulty observed incidence rate of the Non-Vaccinated group and replacing it with the expected National Incidence IR(Baseline), derived from a demographically sound gold standard baseline. This IR(Baseline) however is not drawn from Table 1 (2.473), but adjusted for the age restrictions imposed by the original study cohort of [14], which excluded individuals under 20. The demographically weighted expected incidence for the NV cohort was that already calculated as 2.2146 / 10,000. The corrected Hazard Ratio, HR(Correct), is then calculated by comparing the observed incidence in the Vaccinated group IR(V, Observed) = 2.22 / 10,000, against the verified demographically expected incidence for the 18 control cohort IR(NV, Expected) = 2.2146 per 10,000, yielding 2.22 / 2.2146 = 1.0024 per 10.000 which is very near to the unity. We also emphasize that the value of 2.2146 used for correction represents the demographically weighted incidence, that is the IR(NV, Expected), specifically restricted to the > 20 age cohort analyzed in the original study, thus differing from the overall national IR(Baseline) of 2.473 reported in Table 1. This decomposition conclusively shows that the widely reported HR of 2.714 was almost entirely the product of statistical bias. Once corrected for the structural differences, the true association HR (corrected) has led us to 1, indicating the complete elimination of the alarm signal and confirming that the vaccination was not associated with an increased risk of vitiligo in this population when compared to a proper demographic baseline. In simple words, since the corrected HR is practically 1, it cannot be asserted that NV or V run a greater risk than the other of developing vitiligo following vaccination, given the available data. But also this vitiligo case of [14] has its complexities, because we must ask ourselves what happens if we relax the assumption that the data in Table 1 were actually annualized, and above all, to reintroduce, if one truly wants to discuss the epidemiology of vitiligo, the contribution of those who contract the disease before the age of 20. Let us start with the issue of < 20years. We know from [29] that the annual South Korean incidence rate for this age bracket is 3.4241 (also shown in Table 2), and we also know from [36] that the percentage of individuals < 20 years is approx. 15% of the total population. This leads to a gold standard for the incidence of vitiligo < 20 years equal to 3.4241 x 0.15 = 0.5136. Consequently the national incidence rate ´drawn from [29] for the remaining portion of population will be 2.473 - 0.5136 = 1.9594. If we now annualize the quarterly observed incidence rates in [14] of 2.22 (V) and 0.67 (NV) per 10.000, we achieve respectively 8.88 and 2.68 to be contrasted against 1.9594 per 10.000, like in Table 6 below. 19 Table 6: Inconsistency and Violation of Epidemiological Principles (Vitiligo) Group Observed Rate (Annual) Expected Rate (Gold Standard) Ratio (HR) Inconsistency with Baseline NonVaccinated (NV) 2.68 1.9594 1.37 37% Higher Vaccinated (V) 8.88 1.9594 4.53 453% Higher In essence, the comparison of the annualized observed rates with the national gold standard reveals a profound inconsistency, rendering the outcome of [14] unreliable. It simultaneously obtained two results disastrous for the reliability of the investigated study. On the one hand, it invalidated the control group with an incidence of 2.68 remarkably higher than the national standards (+37% over 1.9594), on the other hand it has contributed to the calculation of a corrected HR of 4.53 which is a result clinically and epidemiologically unsustainable. Ultimately, then, if correct calculations are performed while ignoring a portion of the disease's effect on the South Korean population and considering the observed quarterly incidences, the disappointing result is reached: once the effect of age imbalance in the cohorts is eliminated, the true, correctly recalculated clinical risk is non-existent. Conversely, if incidences are considered on an annualized basis and the substantial portion of the population contracting the disease under the age of 20 is not ignored, paradoxical effects arise, including the invalidation of the control group due to incidence rates exceeding the national average, and the non-vaccinated group is left with an incredibly high portion of risk to be seriously considered clinically, thereby effectively invalidating the entire study. Cancer Results: Failure and Correction The cancer analysis presented a different methodological failure, as the Propensity Score Matching (PSM) successfully eliminated structural demographic bias, thus resulting in an HR(Structural) near unity, without even needing to count (see Table 3). Nonetheless, in this case we are in the presence of 20 a double failure: an internal failure (i.e., residual bias within matched strata, Table 4) and an external failure (non-representativeness of the entire cohort compared to the national population, see the -32.2% deficit in previous Section). Consequently, the entire observed risk is attributed to the residual component as can be derived from the application of Equation 1: HR(Residual) = HR(Observed) / HR(Structural) = 1.27 / 1.00 = 1.27. This result confirms that the observed HR represents the direct quantified effect of an asymmetric selection bias, already documented in Table 5, where the NV group exhibits a significantly deeper deficit from the National Incidence baseline (-45.1% vs. -26.9% for the >= 65 group). This asymmetry proves the cohorts are non-comparable on underlying health status, despite the application of the PSM procedure, due to the well know intereference of the healthy coohort phenomenon [37, 38]. In this case we must also address a problem of external validity where the external nonrepresentativeness of the overall cohort mandates the replacement of the observed incidence rate IR(NV, Observed) = 33.43 with the demographic baseline IR(NV, Baseline) = 55.02 (Table 1). Applying Equation (3) and data from Table 1 provides the corrected Hazard Ratio: HR(Corrected) = IR(V, Observed) / IR(NV, Baseline) = 42.62 / 55.02 = 0.77. In the end, the HR(Corrected) is reduced from the original 1.27 to 0.77. This result not only eliminates the reported positive association but even suggests a protective association after adjusting for the structural deficit of high-risk individuals in the study's overall population compared to the national demographic baseline. This is a resounding confirmation of the fact that an unreasoned use of data leads to results that are not only needlessly alarming but often border on nonsensical. Results Summary Table 7 summarizes the outcomes of the decomposition analysis and the subsequent re-calculation of the corrected risk, contrasting the HR published in [14, 15] with the HR re-calculated using the baseline national incidence as the denominator, for both vitiligo and cancer. 21 Table 7: Summary of Results Case Study HR Observed HR Structural HR Residual HR Corrected Correction Effect Vitiligo (Case I) 2.714 1.21 2.24 Approx. 1 Elimination of signal Vitiligo (Case II) 2.714 x 4 N/A N/A IR(NV, Observed) = 2.68 -> 1.37 / IR(V, Observed) = 8.88 -> 4.53 Clinical Nonsense / Violation of STROBE (21) Cancer 1.27 Approx. 1 1.27 0.77 Reversal to protective signal Discussion The present work has put forth a targeted and quantitative methodological critique on the crucial subject of the reliability of safety signals emerging from post-COVID-19 vaccine surveillance, particularly those derived from retrospective studies on large national databases. Our analysis did not focus on the veracity of the biological findings but on the structural validity of the experimental cohorts from which these results were extracted. The fundamental strength of our argument lies in a return to basic statistical principles and foundational epidemiology, often obscured by the uncritical use of advanced inferential methodologies. We demonstrated how a systemic structural flaw, typically, the asymmetric distribution of high-risk individuals (elderly or vulnerable) between the Vaccinated (V) and Non-Vaccinated (NV) groups, can generate predictable and spurious HRs. This phenomenon is not new, but a direct echo of historical precedents, such as the Beta-Agonist Paradox and the Hormonal Replacement Therapy cases, which warned the scientific community about how confounding by indication or structural selection bias can transform a marker of severity into an apparent risk signal. 22 Our analysis addressed this vulnerability by examining two exemplary studies, focusing on vitiligo and cancers, which share the same data source and, critically, similar methodological defects. Through a rigorous decomposition of the Hazard Ratio (HR), we isolated and quantified the exact contribution of this demographic bias, confirming the established framework for bias quantification [24-26]. In the case of vitiligo [14], we clearly identified that the high observed risk was primarily an artifact of structural flaws. This was twofold: first, much of the risk was mathematically attributable solely to the unbalanced age structure of the cohort; second, a deeper analysis revealed a profound inconsistency in the baseline risk, showing that the non-vaccinated (NV) group's annualized incidence rate was already 37% higher than the national gold standard baseline. This inherent contamination of the control group proved the failure of External Validity (STOBE principle 21). Correcting these dysfunctions in either case, the HR of risk was neutralized. Similarly, in the case of cancer, in spite of the use of a complex balancing technique (Propensity Score Matching or PSM), we revealed an even more insidious failure: the PSM successfully balanced the mean age but failed to balance the baseline risk. The structural underrepresentation of high-risk elderly individuals in the NV group compared to the national benchmark artificially suppressed the baseline risk, creating a defect in the denominator which, once corrected with the expected national incidence, brought the HR back to values close to unity. In both examples, the alarm signal was essentially neutralized when the cohort was corrected to reflect adequate demographic representativeness. The profound strength of our methodology lies in its adherence to the long-established principles of descriptive/inferential statistics and medical protocol design, principles that have guided the proper construction of experimental cohorts for decades. Our approach rigorously confirms that data integrity precedes analytical complexity. We assert that the foundational task of ensuring cohort comparability, a prerequisite for mitigating all forms of internal and external bias, must be satisfied before complex inferential algorithms are deployed. Specifically, by using nationally recognized incidence rates as a gold standard baseline, we were able to diagnose and correct not only internal bias (age asymmetry in the vitiligo study) but also the more subtle external bias, like in the cancer study, where the entire study population was non-representative of the national risk profile, owing to the occurrence of the healthy cohort phenomenon [37, 38]. Our work is a compelling demonstration that statistical rigor, rooted in 23 fundamental checks on data representativeness, provides the necessary corrective mechanism to reestablish methodological integrity when structural flaws compromise sophisticated analyses. Nonetheless, despite the robustness of our reconstruction and quantification, the limitations of our study must be clearly recognized and emphasized. Our analysis is based entirely on the re-edition and reanalysis of the aggregated data and epidemiological parameters published in the original studies [14, 15], cross-referenced with known national incidence data [28-36]. Without direct access to the original South Korean database at the individual level (disaggregated data), our speculations, however rigorously correct and quantified, can never push beyond the level imposed by the initial aggregation. Consequently, we cannot diagnose or exclude the existence of additional residual biases, such as insufficient correction for comorbidities or the absence of other unmeasured confounders, which could only emerge from the analysis of individual records. The impossibility of accessing non-aggregated data remains, by definition, the insurmountable limit of any secondary and retrospective analysis of this type. It is also essential to reiterate that our analysis does not constitute a critique of the biological findings of the studies in question, nor does it intend to challenge the seriousness or integrity of the authors or the journals that published them. Our intent is purely methodological and public health-oriented. We caution the scientific community and the public about the correct and reasoned use of data and the necessity of executing rigorous baseline checks on cohort construction before accepting any signal [39]. Our study unequivocally demonstrates that the analyzed data, once subjected to correction for structural bias, do not provide a reasonable basis for any alarm regarding an increased risk of vitiligo or cancer following Covid-19 vaccination, rendering the original conclusions, although obtained with sophisticated tools, needlessly alarming and bordering on statistical artifact. Conclusion Our study performed a crucial methodological intervention to reassess the reliability of specific adverse event signals reported in COVID-19 vaccine safety surveillance, particularly concerning vitiligo and cancer risks derived from a specific national database. We have demonstrated that the observed associations, which generated significant public alarm, are highly likely to be the result of systemic 24 structural flaws in the construction of the comparison cohorts, rather than true biological signals. The core of the issue lies in the failure to ensure full comparability between the vaccinated and nonvaccinated groups, leading to an asymmetric underrepresentation of high-risk elderly individuals in the non-vaccinated baseline. Through rigorous quantitative deconstruction of the reported Hazard Ratios, our analysis successfully isolated and quantified this structural bias. In both exemplar cases, our corrected risk estimates, derived by utilizing national incidence data as a gold standard baseline, effectively neutralized the alarming signals. Specifically, the observed HRs, originally suggesting increased risk, were reduced to values hovering around unity, demonstrating that the risk disparity disappears when methodological integrity is re-established. This present work serves as a powerful reminder that foundational statistical checks on data representativeness must precede the deployment of complex inferential techniques. The integrity of the study's external validity is compromised when the reference group does not accurately reflect the national risk profile, a failure that even sophisticated algorithms like Propensity Score Matching cannot overcome if the core cohort structure is fundamentally deficient. In summary, our findings strongly suggest that the reported associations between COVID-19 vaccination and increased risks of vitiligo and cancer are statistical artifacts born from methodological bias. While acknowledging the inherent limitation of working exclusively with aggregated data, we caution the scientific community and the public against drawing clinical conclusions or issuing public health alarms based on analyses that fail to adhere to these essential principles of robust cohort construction. Our primary conclusion is that, based on the corrected evidence, no reasonable alarm is warranted from the data under scrutiny. Author Information Marco Roccetti (MR): Department of Computer Science and Engineering, University of Bologna, 40126 Bologna, Italy, marco.roccett[email protected]. ORCID: 0000-0003-1264-8595, sole and corresponding author Author Contributions MR conceived and designed the study, carried out all data collection and analysis, interpreted the quantitative results, and was the sole author responsible for writing and revising the manuscript. The author affirms full responsibility for the integrity of the data and the accuracy of the data analysis presented. 25 Ethics approval This study did not require ethical approval because it involved no humans, animals, plants, relying instead on publicly available, aggregated data containing no private information. Data Availability Statement The data presented here is either included directly or was extracted from the referenced documents and cited literature, specifically [14, 15, 28-33]. All calculations are easily reproducible based on the definitions provided. Funding This research received no specific grant from any funding agency in the public, commercial, or not-forprofit sectors. This study was conducted entirely independently by the author using personal resources. Conflict of Interest The author declares that there is no conflict of interest, financial, personal, or otherwise, that could be construed as influencing the results or the conclusions presented in this paper. Generative AI statement The author declares that no Gen AI was used in the creation of this manuscript. References 1. Elm E, Altman DG, Egger M, et al. (2007) The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) Statement: guidelines for reporting observational studies. PLoS Med., 4(10):e296. DOI: 10.1016/j.jclinepi.2007.11.008