Full text
Facing Methodological Regression: External Validity Failure in Large-Scale Medical Cohorts and the Case of Spurious Cancer Signals in COVID-19 Vaccine Surveillance Marco Roccetti1*& ORCID 1 Department of Computer Science and Engineering, University of Bologna, Bologna, Italy *Corresponding author E-mail: [email protected] (MR) &These authors also contributed equally to this work. Abstract A current crisis in large-scale epidemiology stems from a methodological regression, prioritizing computational complexity over foundational rigor. Large observational studies using national registries have generated alarming risk signals (e.g., cancer) post-COVID-19 vaccination. Although these findings achieve high internal validity via sophisticated balancing algorithms, their interpretation as biological signals constitutes a critical External Validity failure, violating STROBE Principle 21. The error lies in cohort construction: complexity masked a catastrophic structural flaw in the baseline risk assessment. Specifically, the Non-Vaccinated (NV) cohort suffered from asymmetric selection bias, creating an artificially depressed baseline incidence rate, the Failing Denominator, that
2 inflates risk for the Vaccinated (V) subgroup. Applying Blinder-Oaxaca decomposition, we formally prove the entire observed risk signal is attributable solely to this bias. Correcting this structural bias with the national oncological gold standard entirely neutralizes the spurious signal. We conclude that modern epidemiology must restore rigor by mandating quantitative external validity checks, returning to the foundational descriptive science established by the European pioneers of biostatistics. Failure to reassert this rigor risks plunging European medicine and public health discourse into long-term chaos. Keywords: Methodological Regression, Scientific Institutions Crisis, External Validity Failure, Covid-19 Vaccine Safety Surveillance, Asymmetric Selection Bias, Hazard Ratio Decomposition Introduction The contemporary landscape of medical research is defined by a profound and problematic paradox: the unprecedented access to big data and advanced analytical algorithms, ranging from cutting-edge deep learning to sophisticated inferential epidemiology, coexists with a systemic vulnerability to fundamental structural flaws in the very design and construction of experimental cohorts [1]. This crisis of evidence constitutes a substantial regress from the foundational scientific revolution meticulously engineered by European pioneers of biostatistics and experimental design in the 20th century. Before their decisive contributions, much of medical science was constrained by hypothetical-deductive reasoning; associations could be posited and explored, but the robust capacity to subject them to rigorous testing, falsification, and demonstration was largely absent. Key figures such as Ronald Fisher, whose work laid the groundwork for randomization, the Analysis of Variance (ANOVA), and Maximum Likelihood estimation,
3 and later David Roxbee Cox, who developed the proportional hazards model (just to cite a very few), were not merely crafting new mathematical equations. They were formulating principles of objectivity, control, and experimental design designed to neutralize subjective bias and spurious correlation. Their methodologies transformed health and agricultural science from conjecture into an empirically grounded system where findings were experimentable, falsifiable, and demonstrably true [2-4]. Paradoxically, today's sheer volume of data has led to an uncritical and pervasive overreliance on the computational power of complex algorithms, often leading to the neglect or even abandonment of these foundational descriptive and design principles. This methodological regression is particularly perilous because it allows structural flaws to be masked by statistical complexity. The most critical manifestation of this defect is the failure to secure the External Validity of the reference group during a retrospective medical study. External Validity is the degree to which a study cohort accurately reflects the target population, its demographic structure, its risk profile, and its true baseline incidence rate. This essential requirement is not discretionary; it is explicitly enshrined in modern clinical reporting standards, notably Principle 21 of the STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) Statement, which mandates a discussion on generalizability (external validity). This principle exists precisely to ensure that observed findings are not merely internally coherent artifacts but possess genuine relevance for the target population being suitable for clinical application [5]. Unfortunately, this problem of data integrity has recently escalated into a specific public health crisis following the global COVID-19 vaccination efforts, which involved hundreds of millions of people in Europe and globally. This unprecedented mass intervention necessitated
4 large-scale, rapid safety surveillance, often relying on massive national registry databases. This reliance has fostered a prolific body of research, mostly originating from the same research group and database in South Korea, correlating COVID-19 vaccination with a spectrum of increased health risks. These signals span diverse fields, including: psychiatric disorders, autoimmune disturbances, cardiovascular complications, and, most prominently, an increased risk for all-site cancer [6-9]. These studies, often published in high-impact international journals, have garnered significant attention, influencing global public opinion and safety discourse. The critical issue under scrutiny is not the correctness of the statistical calculations within the extracted data, nor the adherence to algorithmic procedures for achieving internal balance. The core question is whether the underlying data, specifically the construction of the retrospective control cohorts, was fundamentally flawed in its inability to represent the broader social and epidemiological reality to which these alarming signals were directed. To be precise: observed associations found within the closed, restricted experimental scope may be statistically true. However, if that experimental design violates the criteria of External Validity, how can one legitimately claim that COVID-19 vaccinated individuals appear to show a higher incidence of developing various diseases when the actual baseline incidence rate used for comparison results to be statistically and clinically significantly lower when compared to the long-term, verifiable national incidence rate? The methodological failure lies in prioritizing the computational ease of comparison over the cultural and statistical integrity of the reference population's true baseline risk. The core issue that a structural bias can generate a spurious signal of an increased risk is not new, but a direct repetition of historical methodological errors, of which we provide two
5 examples historically established in clinical practice. First, the Healthy Worker Effect (HWE), meticulously dissected in [10], is a cornerstone example of selection bias in occupational epidemiology. It systematically shows that the mortality or morbidity rate of an employed cohort is often lower than that of the general population. This effect occurs because individuals who are severely ill, disabled, or too frail are excluded from the working population (a special form of selection bias). When the disease incidence in the worker cohort is compared to the national baseline incidence, the resultant risk ratio is artificially lowered, falsely suggesting that the occupational exposure is protective when, in reality, the comparison group (the working cohort) is simply systematically healthier than the general population. This structural flaw perfectly illustrates how a compromised denominator creates a spurious signal. Secondly, there is the The Beta-Agonist Paradox [11]. Early observational studies suggested that frequent use of short-acting inhaled beta-agonists for asthma was associated with an increased risk of mortality. This alarming signal was largely explained by confounding by indication. In essence, high usage of the medication was not the cause of the mortality, but was merely an indicator that the patient suffered from severe, high-risk, and poorly controlled asthma. In this scenario, the drug exposure was statistically correlated with the severity of the underlying condition, inflating the perceived risk. These precedents demonstrate that when the structure of the reference group or the context of the exposure is compromised by selection or indication bias, the resulting clinical output (the Hazard Ratio, or HR) can transform an indicator of pre-existing severity or, more critically for our case, an indicator of under-representation of risk, into a spurious signal.
6 Our aim with our study is to utilize the recent cancer safety signal (HR=1.27) of [7] as a critical case study to demonstrate this fundamental methodological fragility. We will show that the reported risk, despite being derived from using algorithmic cohort construction procedures on a massive national database, is not a biological signal. Through a rigorous Hazard Ratio (HR) decomposition, based on the principles of the Blinder-Oaxaca methodology [12, 13], and a descriptive comparative analysis, we prove that this association is the direct product of a structural failure of the baseline: an asymmetry in selection bias that rendered the Non-Vaccinated (NV) reference group non-representative of the national oncological risk profile, thus violating STROBE Principle 21 [5]. Ultimately, we reach the central lesson of this history. The profound global impact of findings derived from myriads of national and far databases necessitates an urgent position from scientific institutions, major publishing groups, and the entire peer-review process, regardless of the geographic origin of the data (e.g., a South Korean results influencing policy and public trust in Europe and globally). The slow, minimizing reaction of these bodies, in contrast to the immediate alarm generated in the public sphere, is unsustainable. For those making decisions, there are only two rational paths ahead. We will call the first one as the path to Action and Answer. If the findings of those researches (e.g., the cancer risk [7]) are proven reliable, public health policies must be radically revised concerning the vaccine/drug under consideration. Governments and health authorities must transparently and promptly explain to citizens how this risk touches them and detail plans for support or compensation. The other one is the path to Falsification and Filter. If, conversely, those findings are methodologically unreliable, institutions must redouble efforts to demonstrate this unreliability, impose stricter methodological filters on publication, and confirm the absence of risk, thereby reassuring the public and restoring confidence.
7 The failure to decisively take either of these two roads, the Action or Falsification, leaves open the disastrous third road, that is the scientific chaos. This path is characterized by a process where data, randomly extracted from uncontrolled contexts, determine subjective beliefs rather than supporting rigorously developed theories, based on empirical control and validation. In the sensitive domain of clinical medicine, this descent into chaos is catastrophic, opening the main door to charlatans, undermining professional competence, and jeopardizing the public health and well-being gains won by empirical science over the last 150 years. Methods In a medical experiment, both controlled and retrospectice, as in our case, the foundation of any inferential analysis comparing an exposed group (say V for vaccinated in this specific case) against an unexposed group (say NV for non-vaccinated) rests on the assumption of comparability. When this assumption is violated, the resulting risk measure is distorted, leading to the phenomenon of the so called Failing Denominator. In this context, The Hazard Ratio (HR) is the cornerstone measure in time-to-event (survival) analysis. Unlike the simple Relative Risk (RR) or Odds Ratio (OR), the HR accounts for censoring and the changing rate of events over time [3]. This makes it a powerful, yet sensitive, metric for large-scale longitudinal medical cohort studiesMathematically, the HR is defined, based on Equation 1 below, as the ratio of the hazard function for two groups, the exposed (V) and the unexposed (NV) [12, 13]: HR(t) = hV(t) / hNV(t), (1)
8 where h(t) is the hazard function, representing the instantaneous potential risk of an event (e.g., cancer diagnosis) occurring at time t, given that the event has not occurred up to that time. In order: • If HR=1, the hazard rates are equal; the exposure has no effect. • If HR>1, the exposure (in our case thevaccination) is associated with an increased hazard. • If HR<1, the exposure is associated with a decreased hazard (protective effect). Most importantly, the HR is a relative measure. Its validity and generalizability hinge entirely on the assumption that the baseline hazard function of the unexposed group, hNV(t), accurately represents the counterfactual risk, i.e., the risk the exposed group would have faced had they not been exposed. When observational studies use non-randomly selected cohorts, the true validity of hNV(t) often becomes the weakest link, leading to misleading relative ratios. The goal of observational methodology is precisely to ensure that hNV(t) is a valid representation of the population's natural history of the disease. Take in consideration, now, algorithmic techniques like Propensity Score Matching (PSM) [14] which are designed precisely to stabilize the denominator, hNV(t), by balancing observed covariates between the exposed and unexposed groups [8]. PSM is excellent for achieving Internal Validity which eensures that the results observed within a study are truly due to the intervention being tested, and not to other external or confounding influences. However, PSM fails critically in the domain of External Validity, the extent to which a study's findings can be generalized to other populations, settings, and times, when applied to large registry data, because it cannot account for two factors. The first one amounts to Unmeasured Confounding.
9 In fact, PSM cannot balance unobserved factors (e.g., genetic susceptibility or health-seeking behavior). The decision to vaccinate or not is often an expression of underlying health beliefs and compliance habits, which are potent unmeasured confounders. The second factor that may compromise the efficacy of PSM is the effect of External Bias and the Failing Denominator. In fact, PSM cannot confirm if the matched hNV(t) function is representative of the risk in the general population. If the selection process creates an NV group with an artificially low absolute incidence rate, the HR will be inflated, regardless of how well age or sex were balanced. This failure stems from a violation of the common support principle at the population level. As a method to reason on HRs and their deep meaning we propose to resort to the BlinderOaxaca decomposition methodology [12, 13]. We will use this arsenal to rigorously dissect the source of the observed HR. This decomposition technique, originally developed in economics to explain differences in outcomes (like wage gaps) between groups, in fact was naturally designed to separate the total outcome difference (Δ) into two primary effects based on the following Equation 2: Δ = E + C, (2) where E is the difference explained by endowments (group characteristics) and C is the difference explained by coefficients (the residual effect, or the discriminatory effect of the model). In the context of the Hazard Ratio, we operationalize this decomposition for multiplicative effects following Equation 3: HR(Observed) = HR(Structural) × HR(Residual). (3)
16 The HR decomposition, rooted in the Blinder-Oaxaca framework, we have developed here shows that the sophisticated technique (PSM) failed to remove the residual bias entirely, leaving a result that is a perfect reflection of the underlying structural defect. This reinforces the principle that statistical modeling is only as robust as the data on which it is built. The true innovation lies not in the complexity of the adjustment, but in the integrity of the underlying comparison. Not only that, but the necessity for methodological integrity is codified in contemporary research protocols. The STROBE Statement was introduced in clinical medicine to ensure transparency, reproducibility, and, critically, reliability [5]. However, its application is often focused purely on internal reporting metrics. Our analysis shows that the most severe violation in the study [7] is that of STROBE Principle 21, which governs generalizability (or external validity). We argue that adherence to STROBE 21 must move beyond a simple qualitative statement. It requires quantitative proof of non-distortion. We also should recognize that it is often an issue of qualitative adherence vs. quantitative failure. For example the scrutinized study [7] provided ample detail on internal adjustments (age, sex, comorbidities), thus appearing methodologically robust but our base-rate check exposes its quantitative failure. The -45.1% deficit (Table 2 ) in the high-risk NV group, compared to the national gold standard [19] confirms that the entire study population was systematically unrepresentative of the national risk profile. When the baseline is this fundamentally flawed, the results cannot be generalized; they are confined to a selected, artificially low-risk subsample. The published result of increased cancer risk (of 1.27) is a finding of internal validity (a true result within that flawed sample) but zero external validity.
17 The statistical artifact is a direct consequence of ignoring the quantifiable failure of the control group to meet the standards demanded by the generalizability principle of STROBE 21. In the end, none should be led to believe that it is just a statistical problem The problem of the Failing Denominator is a severe ethical concern in public health. The generation of a positive risk signal that is entirely reducible to statistical bias has a disproportionate impact on vaccine confidence and public trust in scientific institutions. The primary ethical mandate for epidemiologists is to avoid introducing systematic error that could be misinterpreted as a genuine public health threat. Our correction procedure provides the necessary ethical safeguard: it forces the reported finding to be assessed against the true population risk, rather than against an artificially constructed, non-representative control group. Our sentiment is that the ultimate goal for researchers must be to provide an unbiased, generalizable, and demonstrably true risk estimate. The failure to do so allows methodological artifacts to be interpreted as genuine public health threats, undermining the necessary societal reliance on data-driven evidence. We recommend that peer review protocols be updated to explicitly require a quantitative external validity check for any large-scale observational study based on registry data where the risk outcome is the primary focus. We have also to recognize that, while our analysis rigorously demonstrates that the entire observed cancer risk signal is attributable to a quantifiable failure of External Validity our corrective procedure operates exclusively at the aggregate level. We utilized publicly available national oncological gold standard data and published cohort statistics to perform the necessary adjustments and HR decomposition. A key limitation of our investigation is
18 therefore the inability to access the raw, individual-level database used in the original studies (i.e., the South Korean National Health Insurance Service - NHIS data). In fact, our findings, based on macro-level comparisons, highlight a systemic failure in the design of the cohort. However, the definitive proof of causality (or lack thereof) would require a detailed inspection of the individual records. Given the extraordinary size and granularity of the NHIS database, which encompasses a relevant portion of a nation's entire population and healthcare history [20, 21], direct access would definitively resolve the methodological dilemma. The application of our quantitative external validity checks and decomposition methods at the individual patient level would yield a conclusive, non-refutable answer regarding the existence (or non-existence) of this and other hypothesized associations (e.g., cardiovascular or neurological risks). We assert that granting external, independent researchers secure access to these anonymized, large-scale registry databases is not only a matter of transparency but is now the institutional imperative. Such access represents the only rational way to provide a definitive and globally transferable verdict on these health surveillance signals, thus fulfilling the primary ethical and social responsibility of scientific research. In closing, we feel that the persistent generation and publication of methodologically compromised findings, particularly those arising from a limited geographical region and a restricted database, demands decisive action from the established global scientific ecosystem. The minimal or slow response from scientific institutions, major editorial groups, and regulatory bodies, an attitude often perceived as minimization, is inherently destabilizing because it avoids the inevitable choice between the two rational paths. If one follows the first one and believes the risk associations were biologically true, institutions would have the
19 immediate ethical and public health duty to implement policy changes and prepare compensation for affected citizens. If, instead, as we have demonstrated, the finding is a statistical artifact due to a quantified methodological failure (-45.1% deficit in baseline IR), the institutions' duty is to impose stringent methodological oversight, filter such nongeneralizable results, and confirm the unreliability to the public. By failing to commit to either action or refutation, the institutions inadvertently endorse the third road: scientific chaos. This chaos represents a breakdown of the 150-year-old scientific paradigm, replacing controlled inference with arbitrary data extraction that generates unsubstantiated beliefs instead of confirmed theories. In the field of clinical medicine, this chaos has catastrophic implications, risking the public health consensus and opening the door to charlatans, undermining professional competence, and jeopardizing the public health and well-being gains won by empirical science over the last 150 years. Conclusion Our study performed a crucial methodological intervention to re-assess the reliability of a specific cancer adverse event signal. We have conclusively demonstrated that the observed association is highly likely the result of a systemic structural flaw in cohort construction, a failure of External Validity, which fundamentally violates the principles of scientific demonstrability established by the pioneers of biostatistics. The reliance on algorithmic procedures only proved insufficient, as it failed to compensate for the inherent asymmetric selection bias evident in the baseline incidence rates. Our statistical correction, obtained by utilizing national incidence data as a gold standard baseline [22], effectively neutralized the alarming signal, reducing the observed HR of 1.27 to a corrected HR of 0.77. This work serves as a powerful reminder that foundational descriptive checks on data representativeness must precede the deployment of complex inferential techniques. The integrity of a study's external validity is compromised when the reference group does not
20 accurately reflect the national risk profile. The ethical responsibility of the scientific community demands that we reject the notion that sophisticated computational methods can compensate for fundamental flaws in cohort design. Our primary conclusion is that, based on the corrected evidence, no reasonable alarm is warranted from the data under scrutiny. Furthermore, the institutional response to such methodologically compromised findings must move beyond minimization toward definitive action, either confirming the risk or confirming its unreliability. We urge the scientific community to re-establish the methodological rigor that defined the great advances of the 20th-century European statistical tradition, prioritizing the quantitative validation of STROBE Principle 21 to ensure that scientific findings are demonstrably generalizable and reliable, thus protecting the integrity of science itself from the threat of chaos. Author Information Marco Roccetti (MR): Department of Computer Science and Engineering, University of Bologna, 40126 Bologna, Italy, [email protected]. ORCID: 0000-0003-1264-8595, sole and corresponding author Author Contributions MR conceived and designed the study, carried out all data collection and analysis, interpreted the quantitative results, and was the sole author responsible for writing and revising the manuscript. The author affirms full responsibility for the integrity of the data and the accuracy of the data analysis presented.
21 Ethics Approval This study did not require ethical approval because it involved no humans, animals, plants, relying instead on publicly available, aggregated data containing no private information. Data Availability Statement The data presented here is either included directly or was extracted from the referenced documents and cited literature. All calculations are easily reproducible based on the definitions provided. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. This study was conducted entirely independently by the author using personal resources. Conflict of Interest The author declares that there is no conflict of interest, financial, personal, or otherwise, that could be construed as influencing the results or the conclusions presented in this paper. Generative AI statement The author declares that no Gen AI was used in the creation of this manuscript. References 1. Roccetti M, Delnevo G, Casini L, et al. (2019). Is bigger always better? A controversial journey to the center of machine learning design, with uses and misuses
22 of big data for predicting water meter failures. J. Big Data. 6(1):70. DOI: 10.1186/s40537-019-0235-y 2. Fisher RA, (1921) Studies in Crop Variation (I). An Examination of the Yield of Dressed Grain from Broadbalk". J. Agric. Sci. 11(2):107–135. DOI: 10.1017/S0021859600003750 3. Cox DR, (1972). Regression Models and Life-Tables. J. R. Stat. Soc. B, 34(2):187220. DOI: 10.1111/j.2517-6161.1972.tb00970.x 4. Neyman J, Pearson ES, (1933) On the problem of the most efficient tests of statistical hypotheses. Phil. Trans. R. Soc. Lond. A. 231(694–706):289–337. DOI: 10.1098/rsta.1933.0009 5. Elm E, Altman DG, Egger M, et al. (2007) The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) Statement: guidelines for reporting observational studies. PLoS Med., 4(10):e296. DOI: 10.1016/j.jclinepi.2007.11.008 6. Kim HJ, Kim M-H, Park SJ, et al. (2024). Autoimmune adverse events after COVID19 vaccination: a nationwide population-based cohort study in Korea. J Alergy Clin Immunol., 153(6):1711-1720. DOI: 10.1016/j.jaci.2024.01.025 7. Kim HJ, Kim M-H, Choi MG, et al. (2025) 1-year risks of cancers associated with COVID-19 vaccination: a large population-based cohort study in South Korea. Biomark Res. 13(114). DOI: 10.1186/s40364-025-00831-w 8. Kim HJ, Kim MH, Choi MG, et al. (2024) Psychiatric adverse events following COVID-19 vaccination: a population-based cohort study in Seoul, South Korea. Mol Psychiatry. 2024(12):3635–3643. doi: 10.1038/s41380-024-02627-0
23 9. Kim HJ, Kim M-H, Park SJ, et al. (2025). The early impact of COVID-19 vaccines on major events in cardiac, pulmonary, and thromboembolic disease: a population-based study. Korean J Intern Med., 2025(40):801-812. DOI: 10.3904/kjim.2025.056 10. Wen CP, Tsai SP, (1982). Anatomy of the health worker effect - a critique of summary statistics employed in occupational epidemiology. Scand J Work Environ Health, 8(1):48-52. PMID: 7100856 11. Spitzer WO, Suissa S, Ernst P, et al. (1992) The use of beta-agonists and the risk of death and near death from asthma. N Engl J Med., 326(8):501-506. DOI: 10.1056/NEJM199202203260801 12. Oaxaca R, (1973) Male-Female Wage Differentials in Urban Labor Markets. Int. Econ. Rev. 14(3):693-709. doi:10.2307/2525981x 13. Blinder AS, (1973). Wage Discrimination: Reduced Form and Structural Estimates. J. Hum. Resour. 8(4):436-455. DOI:10.2307/144855 14. D’Agostino RB Jr. (1998) Propensity score methods for bias reduction in the comparison of a treatment to a non-randomized control group. Stat Med. 17(19):22652281. DOI: 10.1002/(sici)1097-0258(19981015)17:19<2265::aid-sim918>3.0.co;2-b 15. Chemaitelly H, Ayoub H, Coyle P, et al. (2025) Assessing healthy vaccinee effect in COVID-19 vaccine effectiveness studies: a national cohort study in Qatar. eLife 2025, 14:e103690. DOI: 10.7554/eLife.103690 16. Kang MJ, Jung K-W, Bang SH, et al. (2023). Cancer Statistics in Korea: Incidence, Mortality, Survival, and Prevalence in 2020. Cancer Res Treat., 55(2):385-399. DOI: 10.4143/crt.2023.447 17. Park EH, Jung K-W, Park NJ, et al. (2024). Cancer Statistics in Korea: Incidence, Mortality, Survival, and Prevalence in 2021. Cancer Res Treat., 56(2):357-371. DOI: 10.4143/crt.2024.253
24 18. Park EH, Jung K-W, Park NJ, et al. (2025) Cancer Statistics in Korea: Incidence, Mortality, Survival, and Prevalence in 2022. Cancer Res Treat. 57(2):312-330. DOI: 10.4143/crt.2025.264 19. Statista. South Korea: Cancer crude incidence rate by age, 2022. Statista; 2024. [Accessed 2025 Nov 29]. Available from: https://www.statista.com/statistics/1440818/south-korea-cancer-crude-incidence-rateby-age/ 20. World Bank. Population ages 65 and above (% of total population) - Korea, Rep. World Population Prospects, United Nations (UN). [Accessed 2025 Nov 29]. Available from: https://data.worldbank.org/indicator/SP.POP.65UP.TO.ZS?locations=KR 21. UN World Population Prospects Data, Population Pyramids, [Accessed 2025 Nov 29]. Available from:https://www.populationpyramids.org/southkorea?utm_source=chatgpt.com 22. Roccetti M, Cacciapuoti G,. (2025) Beyond the Gold Standard: Linear Regression and Poisson GLM Yield Identical Mortality Trends and Deaths Counts for COVID-19 in Italy: 2021–2025. Computation, 13(10): 233. DOI: 10.3390/computation13100233.