scieee AI-readable full text Open interactive document viewer

Facing Methodological Regression: External Validity Failure in Large-Scale Medical Cohorts and the Case of Spurious Cancer Signals in COVID-19 Vaccine Surveillance

ROCCETTI, MARCO

Abstract

Abstract: A current crisis in large-scale epidemiology stems from a methodological regression, prioritizing computational complexity over foundational rigor. Large observational studies using national registries have generated alarming risk signals (e.g., cancer) post-COVID-19 vaccination. Although these findings achieve high internal validity via sophisticated balancing algorithms, their interpretation as biological signals constitutes a critical External Validity failure, violating STROBE Principle 21. The error lies in cohort construction: complexity masked a catastrophic structural flaw in the baseline risk assessment. Specifically, the Non-Vaccinated (NV) cohort suffered from asymmetric selection bias, creating an artificially depressed baseline incidence rate, the Failing Denominator, that inflates risk for the Vaccinated (V) subgroup. Applying Blinder-Oaxaca decomposition, we formally prove the entire observed risk signal is attributable solely to this bias. Correcting this structural bias with the national oncological gold standard entirely neutralizes the spurious signal. We conclude that modern epidemiology must restore rigor by mandating quantitative external validity checks, returning to the foundational descriptive science established by the European pioneers of biostatistics. Failure to reassert this rigor risks plunging European medicine and public health discourse into long-term chaos.

Full text

Facing Methodological Regression: External Validity Failure in 1 Large-Scale Medical Cohorts and the Case of Spurious Cancer 2 Signals in COVID-19 Vaccine Surveillance 3 4 5 Marco Roccetti1*& ORCID 6 7 1 Department of Computer Science and Engineering, University of Bologna, Bologna, Italy 8 9 *Corresponding author 10 E-mail: [email protected] (MR) 11 12 &These authors also contributed equally to this work. 13 Abstract 14 A current crisis in large-scale epidemiology stems from a methodological regression, 15 prioritizing computational complexity over foundational rigor. Large observational studies 16 using national registries have generated alarming risk signals (e.g., cancer) post-COVID-19 17 vaccination. Although these findings achieve high internal validity via sophisticated 18 balancing algorithms, their interpretation as biological signals constitutes a critical External 19 Validity failure, violating STROBE Principle 21. The error lies in cohort construction: 20 complexity masked a catastrophic structural flaw in the baseline risk assessment. 21 Specifically, the Non-Vaccinated (NV) cohort suffered from asymmetric selection bias, 22 creating an artificially depressed baseline incidence rate, the Failing Denominator, that 23 2 inflates risk for the Vaccinated (V) subgroup. Applying Blinder-Oaxaca decomposition, we 24 formally prove the entire observed risk signal is attributable solely to this bias. Correcting 25 this structural bias with the national oncological gold standard entirely neutralizes the 26 spurious signal. We conclude that modern epidemiology must restore rigor by mandating 27 quantitative external validity checks, returning to the foundational descriptive science 28 established by the European pioneers of biostatistics. Failure to reassert this rigor risks 29 plunging European medicine and public health discourse into long-term chaos. 30 31 Keywords: Methodological Regression, Scientific Institutions Crisis, External Validity 32 Failure, Covid-19 Vaccine Safety Surveillance, Asymmetric Selection Bias, Hazard Ratio 33 Decomposition 34 Introduction 35 The contemporary landscape of medical research is defined by a profound and problematic 36 paradox: the unprecedented access to big data and advanced analytical algorithms, ranging 37 from cutting-edge deep learning to sophisticated inferential epidemiology, coexists with a 38 systemic vulnerability to fundamental structural flaws in the very design and construction of 39 experimental cohorts [1]. 40 41 This crisis of evidence constitutes a substantial regress from the foundational scientific 42 revolution meticulously engineered by European pioneers of biostatistics and experimental 43 design in the 20th century. Before their decisive contributions, much of medical science was 44 constrained by hypothetical-deductive reasoning; associations could be posited and explored, 45 but the robust capacity to subject them to rigorous testing, falsification, and demonstration 46 was largely absent. Key figures such as Ronald Fisher, whose work laid the groundwork for 47 randomization, the Analysis of Variance (ANOVA), and Maximum Likelihood estimation, 48 3 and later David Roxbee Cox, who developed the proportional hazards model (just to cite a 49 very few), were not merely crafting new mathematical equations. They were formulating 50 principles of objectivity, control, and experimental design designed to neutralize subjective 51 bias and spurious correlation. Their methodologies transformed health and agricultural 52 science from conjecture into an empirically grounded system where findings were 53 experimentable, falsifiable, and demonstrably true [2-4]. 54 55 Paradoxically, today's sheer volume of data has led to an uncritical and pervasive over-56 reliance on the computational power of complex algorithms, often leading to the neglect or 57 even abandonment of these foundational descriptive and design principles. This 58 methodological regression is particularly perilous because it allows structural flaws to be 59 masked by statistical complexity. The most critical manifestation of this defect is the failure 60 to secure the External Validity of the reference group during a retrospective medical study. 61 External Validity is the degree to which a study cohort accurately reflects the target 62 population, its demographic structure, its risk profile, and its true baseline incidence rate. 63 This essential requirement is not discretionary; it is explicitly enshrined in modern clinical 64 reporting standards, notably Principle 21 of the STROBE (Strengthening the Reporting of 65 Observational Studies in Epidemiology) Statement, which mandates a discussion on 66 generalizability (external validity). This principle exists precisely to ensure that observed 67 findings are not merely internally coherent artifacts but possess genuine relevance for the 68 target population being suitable for clinical application [5]. 69 70 Unfortunately, this problem of data integrity has recently escalated into a specific public 71 health crisis following the global COVID-19 vaccination efforts, which involved hundreds of 72 millions of people in Europe and globally. This unprecedented mass intervention necessitated 73 4 large-scale, rapid safety surveillance, often relying on massive national registry databases. 74 This reliance has fostered a prolific body of research, mostly originating from the same 75 research group and database in South Korea, correlating COVID-19 vaccination with a 76 spectrum of increased health risks. These signals span diverse fields, including: psychiatric 77 disorders, autoimmune disturbances, cardiovascular complications, and, most prominently, an 78 increased risk for all-site cancer [6-9]. These studies, often published in high-impact 79 international journals, have garnered significant attention, influencing global public opinion 80 and safety discourse. 81 82 The critical issue under scrutiny is not the correctness of the statistical calculations within the 83 extracted data, nor the adherence to algorithmic procedures for achieving internal balance. 84 The core question is whether the underlying data, specifically the construction of the 85 retrospective control cohorts, was fundamentally flawed in its inability to represent the 86 broader social and epidemiological reality to which these alarming signals were directed. To 87 be precise: observed associations found within the closed, restricted experimental scope may 88 be statistically true. However, if that experimental design violates the criteria of External 89 Validity, how can one legitimately claim that COVID-19 vaccinated individuals appear to 90 show a higher incidence of developing various diseases when the actual baseline incidence 91 rate used for comparison results to be statistically and clinically significantly lower when 92 compared to the long-term, verifiable national incidence rate? The methodological failure lies 93 in prioritizing the computational ease of comparison over the cultural and statistical integrity 94 of the reference population's true baseline risk. 95 96 The core issue that a structural bias can generate a spurious signal of an increased risk is not 97 new, but a direct repetition of historical methodological errors, of which we provide two 98 5 examples historically established in clinical practice. First, the Healthy Worker Effect 99 (HWE), meticulously dissected in [10], is a cornerstone example of selection bias in 100 occupational epidemiology. It systematically shows that the mortality or morbidity rate of an 101 employed cohort is often lower than that of the general population. This effect occurs because 102 individuals who are severely ill, disabled, or too frail are excluded from the working 103 population (a special form of selection bias). When the disease incidence in the worker cohort 104 is compared to the national baseline incidence, the resultant risk ratio is artificially lowered, 105 falsely suggesting that the occupational exposure is protective when, in reality, the 106 comparison group (the working cohort) is simply systematically healthier than the general 107 population. This structural flaw perfectly illustrates how a compromised denominator creates 108 a spurious signal. 109 110 Secondly, there is the The Beta-Agonist Paradox [11]. Early observational studies suggested 111 that frequent use of short-acting inhaled beta-agonists for asthma was associated with an 112 increased risk of mortality. This alarming signal was largely explained by confounding by 113 indication. In essence, high usage of the medication was not the cause of the mortality, but 114 was merely an indicator that the patient suffered from severe, high-risk, and poorly controlled 115 asthma. In this scenario, the drug exposure was statistically correlated with the severity of the 116 underlying condition, inflating the perceived risk. These precedents demonstrate that when 117 the structure of the reference group or the context of the exposure is compromised by 118 selection or indication bias, the resulting clinical output (the Hazard Ratio, or HR) can 119 transform an indicator of pre-existing severity or, more critically for our case, an indicator of 120 under-representation of risk, into a spurious signal. 121 122 6 Our aim with our study is to utilize the recent cancer safety signal (HR=1.27) of [7] as a 123 critical case study to demonstrate this fundamental methodological fragility. We will show 124 that the reported risk, despite being derived from using algorithmic cohort construction 125 procedures on a massive national database, is not a biological signal. Through a rigorous 126 Hazard Ratio (HR) decomposition, based on the principles of the Blinder-Oaxaca 127 methodology [12, 13], and a descriptive comparative analysis, we prove that this association 128 is the direct product of a structural failure of the baseline: an asymmetry in selection bias that 129 rendered the Non-Vaccinated (NV) reference group non-representative of the national 130 oncological risk profile, thus violating STROBE Principle 21 [5]. 131 132 Ultimately, we reach the central lesson of this history. The profound global impact of findings 133 derived from myriads of national and far databases necessitates an urgent position from 134 scientific institutions, major publishing groups, and the entire peer-review process, regardless 135 of the geographic origin of the data (e.g., a South Korean results influencing policy and public 136 trust in Europe and globally). The slow, minimizing reaction of these bodies, in contrast to the 137 immediate alarm generated in the public sphere, is unsustainable. For those making decisions, 138 there are only two rational paths ahead. We will call the first one as the path to Action and 139 Answer. If the findings of those researches (e.g., the cancer risk [7]) are proven reliable, public 140 health policies must be radically revised concerning the vaccine/drug under consideration. 141 Governments and health authorities must transparently and promptly explain to citizens how 142 this risk touches them and detail plans for support or compensation. The other one is the path 143 to Falsification and Filter. If, conversely, those findings are methodologically unreliable, 144 institutions must redouble efforts to demonstrate this unreliability, impose stricter 145 methodological filters on publication, and confirm the absence of risk, thereby reassuring the 146 public and restoring confidence. 147 7 148 The failure to decisively take either of these two roads, the Action or Falsification, leaves 149 open the disastrous third road, that is the scientific chaos. This path is characterized by a 150 process where data, randomly extracted from uncontrolled contexts, determine subjective 151 beliefs rather than supporting rigorously developed theories, based on empirical control and 152 validation. In the sensitive domain of clinical medicine, this descent into chaos is 153 catastrophic, opening the main door to charlatans, undermining professional competence, and 154 jeopardizing the public health and well-being gains won by empirical science over the last 155 150 years. 156 Methods 157 In a medical experiment, both controlled and retrospectice, as in our case, the foundation of 158 any inferential analysis comparing an exposed group (say V for vaccinated in this specific case) 159 against an unexposed group (say NV for non-vaccinated) rests on the assumption of 160 comparability. When this assumption is violated, the resulting risk measure is distorted, leading 161 to the phenomenon of the so called Failing Denominator. 162 163 In this context, The Hazard Ratio (HR) is the cornerstone measure in time-to-event (survival) 164 analysis. Unlike the simple Relative Risk (RR) or Odds Ratio (OR), the HR accounts for 165 censoring and the changing rate of events over time [3]. This makes it a powerful, yet sensitive, 166 metric for large-scale longitudinal medical cohort studiesMathematically, the HR is defined, 167 based on Equation 1 below, as the ratio of the hazard function for two groups, the exposed (V) 168 and the unexposed (NV) [12, 13]: 169 170 HR(t) = hV(t) / hNV(t), (1) 171 172 8 where h(t) is the hazard function, representing the instantaneous potential risk of an event 173 (e.g., cancer diagnosis) occurring at time t, given that the event has not occurred up to that 174 time. In order: 175 176 • If HR=1, the hazard rates are equal; the exposure has no effect. 177 • If HR>1, the exposure (in our case thevaccination) is associated with an increased 178 hazard. 179 • If HR<1, the exposure is associated with a decreased hazard (protective effect). 180 181 Most importantly, the HR is a relative measure. Its validity and generalizability hinge entirely 182 on the assumption that the baseline hazard function of the unexposed group, hNV(t), 183 accurately represents the counterfactual risk, i.e., the risk the exposed group would have 184 faced had they not been exposed. When observational studies use non-randomly selected 185 cohorts, the true validity of hNV(t) often becomes the weakest link, leading to misleading 186 relative ratios. The goal of observational methodology is precisely to ensure that hNV(t) is a 187 valid representation of the population's natural history of the disease. 188 189 Take in consideration, now, algorithmic techniques like Propensity Score Matching (PSM) [14] 190 which are designed precisely to stabilize the denominator, hNV(t), by balancing observed 191 covariates between the exposed and unexposed groups [8]. PSM is excellent for achieving 192 Internal Validity which eensures that the results observed within a study are truly due to the 193 intervention being tested, and not to other external or confounding influences. However, PSM 194 fails critically in the domain of External Validity, the extent to which a study's findings can be 195 generalized to other populations, settings, and times, when applied to large registry data, 196 because it cannot account for two factors. The first one amounts to Unmeasured Confounding. 197 9 In fact, PSM cannot balance unobserved factors (e.g., genetic susceptibility or health-seeking 198 behavior). The decision to vaccinate or not is often an expression of underlying health beliefs 199 and compliance habits, which are potent unmeasured confounders. 200 201 The second factor that may compromise the efficacy of PSM is the effect of External Bias and 202 the Failing Denominator. In fact, PSM cannot confirm if the matched hNV(t) function is 203 representative of the risk in the general population. If the selection process creates an NV group 204 with an artificially low absolute incidence rate, the HR will be inflated, regardless of how well 205 age or sex were balanced. This failure stems from a violation of the common support principle 206 at the population level. 207 208 As a method to reason on HRs and their deep meaning we propose to resort to the Blinder-209 Oaxaca decomposition methodology [12, 13]. We will use this arsenal to rigorously dissect the 210 source of the observed HR. This decomposition technique, originally developed in economics 211 to explain differences in outcomes (like wage gaps) between groups, in fact was naturally 212 designed to separate the total outcome difference (Δ) into two primary effects based on the 213 following Equation 2: 214 Δ = E + C, (2) 215 216 where E is the difference explained by endowments (group characteristics) and C is the 217 difference explained by coefficients (the residual effect, or the discriminatory effect of the 218 model). In the context of the Hazard Ratio, we operationalize this decomposition for 219 multiplicative effects following Equation 3: 220 221 HR(Observed) = HR(Structural) × HR(Residual). (3) 222 16 356 The HR decomposition, rooted in the Blinder-Oaxaca framework, we have developed here 357 shows that the sophisticated technique (PSM) failed to remove the residual bias entirely, 358 leaving a result that is a perfect reflection of the underlying structural defect. This reinforces 359 the principle that statistical modeling is only as robust as the data on which it is built. The 360 true innovation lies not in the complexity of the adjustment, but in the integrity of the 361 underlying comparison. 362 363 Not only that, but the necessity for methodological integrity is codified in contemporary 364 research protocols. The STROBE Statement was introduced in clinical medicine to ensure 365 transparency, reproducibility, and, critically, reliability [5]. However, its application is often 366 focused purely on internal reporting metrics. Our analysis shows that the most severe 367 violation in the study [7] is that of STROBE Principle 21, which governs generalizability (or 368 external validity). We argue that adherence to STROBE 21 must move beyond a simple 369 qualitative statement. It requires quantitative proof of non-distortion. 370 371 We also should recognize that it is often an issue of qualitative adherence vs. quantitative 372 failure. For example the scrutinized study [7] provided ample detail on internal adjustments 373 (age, sex, comorbidities), thus appearing methodologically robust but our base-rate check 374 exposes its quantitative failure. The -45.1% deficit (Table 2 ) in the high-risk NV group, 375 compared to the national gold standard [19] confirms that the entire study population was 376 systematically unrepresentative of the national risk profile. When the baseline is this 377 fundamentally flawed, the results cannot be generalized; they are confined to a selected, 378 artificially low-risk subsample. The published result of increased cancer risk (of 1.27) is a 379 finding of internal validity (a true result within that flawed sample) but zero external validity. 380 17 The statistical artifact is a direct consequence of ignoring the quantifiable failure of the 381 control group to meet the standards demanded by the generalizability principle of STROBE 382 21. 383 384 In the end, none should be led to believe that it is just a statistical problem The problem of the 385 Failing Denominator is a severe ethical concern in public health. The generation of a positive 386 risk signal that is entirely reducible to statistical bias has a disproportionate impact on vaccine 387 confidence and public trust in scientific institutions. The primary ethical mandate for 388 epidemiologists is to avoid introducing systematic error that could be misinterpreted as a 389 genuine public health threat. Our correction procedure provides the necessary ethical 390 safeguard: it forces the reported finding to be assessed against the true population risk, rather 391 than against an artificially constructed, non-representative control group. Our sentiment is 392 that the ultimate goal for researchers must be to provide an unbiased, generalizable, and 393 demonstrably true risk estimate. The failure to do so allows methodological artifacts to be 394 interpreted as genuine public health threats, undermining the necessary societal reliance on 395 data-driven evidence. We recommend that peer review protocols be updated to explicitly 396 require a quantitative external validity check for any large-scale observational study based on 397 registry data where the risk outcome is the primary focus. 398 399 We have also to recognize that, while our analysis rigorously demonstrates that the entire 400 observed cancer risk signal is attributable to a quantifiable failure of External Validity our 401 corrective procedure operates exclusively at the aggregate level. We utilized publicly 402 available national oncological gold standard data and published cohort statistics to perform 403 the necessary adjustments and HR decomposition. A key limitation of our investigation is 404 18 therefore the inability to access the raw, individual-level database used in the original studies 405 (i.e., the South Korean National Health Insurance Service - NHIS data). 406 407 In fact, our findings, based on macro-level comparisons, highlight a systemic failure in the 408 design of the cohort. However, the definitive proof of causality (or lack thereof) would 409 require a detailed inspection of the individual records. Given the extraordinary size and 410 granularity of the NHIS database, which encompasses a relevant portion of a nation's entire 411 population and healthcare history [20, 21], direct access would definitively resolve the 412 methodological dilemma. The application of our quantitative external validity checks and 413 decomposition methods at the individual patient level would yield a conclusive, non-refutable 414 answer regarding the existence (or non-existence) of this and other hypothesized associations 415 (e.g., cardiovascular or neurological risks). We assert that granting external, independent 416 researchers secure access to these anonymized, large-scale registry databases is not only a 417 matter of transparency but is now the institutional imperative. Such access represents the only 418 rational way to provide a definitive and globally transferable verdict on these health 419 surveillance signals, thus fulfilling the primary ethical and social responsibility of scientific 420 research. 421 422 In closing, we feel that the persistent generation and publication of methodologically 423 compromised findings, particularly those arising from a limited geographical region and a 424 restricted database, demands decisive action from the established global scientific ecosystem. 425 The minimal or slow response from scientific institutions, major editorial groups, and 426 regulatory bodies, an attitude often perceived as minimization, is inherently destabilizing 427 because it avoids the inevitable choice between the two rational paths. If one follows the first 428 one and believes the risk associations were biologically true, institutions would have the 429 19 immediate ethical and public health duty to implement policy changes and prepare 430 compensation for affected citizens. If, instead, as we have demonstrated, the finding is a 431 statistical artifact due to a quantified methodological failure (-45.1% deficit in baseline IR), 432 the institutions' duty is to impose stringent methodological oversight, filter such non-433 generalizable results, and confirm the unreliability to the public. By failing to commit to 434 either action or refutation, the institutions inadvertently endorse the third road: scientific 435 chaos. This chaos represents a breakdown of the 150-year-old scientific paradigm, replacing 436 controlled inference with arbitrary data extraction that generates unsubstantiated beliefs 437 instead of confirmed theories. In the field of clinical medicine, this chaos has catastrophic 438 implications, risking the public health consensus and opening the door to charlatans, 439 undermining professional competence, and jeopardizing the public health and well-being 440 gains won by empirical science over the last 150 years. 441 Conclusion 442 Our study performed a crucial methodological intervention to re-assess the reliability of a 443 specific cancer adverse event signal. We have conclusively demonstrated that the observed 444 association is highly likely the result of a systemic structural flaw in cohort construction, a 445 failure of External Validity, which fundamentally violates the principles of scientific 446 demonstrability established by the pioneers of biostatistics. The reliance on algorithmic 447 procedures only proved insufficient, as it failed to compensate for the inherent asymmetric 448 selection bias evident in the baseline incidence rates. Our statistical correction, obtained by 449 utilizing national incidence data as a gold standard baseline [22], effectively neutralized the 450 alarming signal, reducing the observed HR of 1.27 to a corrected HR of 0.77. 451 This work serves as a powerful reminder that foundational descriptive checks on data 452 representativeness must precede the deployment of complex inferential techniques. The 453 integrity of a study's external validity is compromised when the reference group does not 454 20 accurately reflect the national risk profile. The ethical responsibility of the scientific 455 community demands that we reject the notion that sophisticated computational methods can 456 compensate for fundamental flaws in cohort design. Our primary conclusion is that, based on 457 the corrected evidence, no reasonable alarm is warranted from the data under scrutiny. 458 Furthermore, the institutional response to such methodologically compromised findings must 459 move beyond minimization toward definitive action, either confirming the risk or confirming 460 its unreliability. We urge the scientific community to re-establish the methodological rigor 461 that defined the great advances of the 20th-century European statistical tradition, prioritizing 462 the quantitative validation of STROBE Principle 21 to ensure that scientific findings are 463 demonstrably generalizable and reliable, thus protecting the integrity of science itself from 464 the threat of chaos. 465 466 Author Information 467 Marco Roccetti (MR): Department of Computer Science and Engineering, University of 468 Bologna, 40126 Bologna, Italy, [email protected]. ORCID: 0000-0003-1264-8595, 469 sole and corresponding author 470 471 Author Contributions 472 MR conceived and designed the study, carried out all data collection and analysis, interpreted 473 the quantitative results, and was the sole author responsible for writing and revising the 474 manuscript. The author affirms full responsibility for the integrity of the data and the 475 accuracy of the data analysis presented. 476 477 478 479 21 Ethics Approval 480 This study did not require ethical approval because it involved no humans, animals, plants, 481 relying instead on publicly available, aggregated data containing no private information. 482 483 Data Availability Statement 484 The data presented here is either included directly or was extracted from the referenced 485 documents and cited literature. All calculations are easily reproducible based on the 486 definitions provided. 487 488 Funding 489 This research received no specific grant from any funding agency in the public, commercial, 490 or not-for-profit sectors. This study was conducted entirely independently by the author using 491 personal resources. 492 493 Conflict of Interest 494 The author declares that there is no conflict of interest, financial, personal, or otherwise, that 495 could be construed as influencing the results or the conclusions presented in this paper. 496 497 Generative AI statement 498 The author declares that no Gen AI was used in the creation of this manuscript. 499 500 References 501 1. Roccetti M, Delnevo G, Casini L, et al. (2019). Is bigger always better? A 502 controversial journey to the center of machine learning design, with uses and misuses 503 22 of big data for predicting water meter failures. J. Big Data. 6(1):70. DOI: 504 10.1186/s40537-019-0235-y 505 2. Fisher RA, (1921) Studies in Crop Variation (I). An Examination of the Yield of 506 Dressed Grain from Broadbalk". J. Agric. Sci. 11(2):107–135. DOI: 507 10.1017/S0021859600003750 508 3. Cox DR, (1972). Regression Models and Life-Tables. J. R. Stat. Soc. B, 34(2):187-509 220. DOI: 10.1111/j.2517-6161.1972.tb00970.x 510 4. Neyman J, Pearson ES, (1933) On the problem of the most efficient tests of statistical 511 hypotheses. Phil. Trans. R. Soc. Lond. A. 231(694–706):289–337. DOI: 512 10.1098/rsta.1933.0009 513 5. Elm E, Altman DG, Egger M, et al. (2007) The Strengthening the Reporting of 514 Observational Studies in Epidemiology (STROBE) Statement: guidelines for 515 reporting observational studies. PLoS Med., 4(10):e296. DOI: 516 10.1016/j.jclinepi.2007.11.008 517 6. Kim HJ, Kim M-H, Park SJ, et al. (2024). Autoimmune adverse events after COVID-518 19 vaccination: a nationwide population-based cohort study in Korea. J Alergy Clin 519 Immunol., 153(6):1711-1720. DOI: 10.1016/j.jaci.2024.01.025 520 7. Kim HJ, Kim M-H, Choi MG, et al. (2025) 1-year risks of cancers associated with 521 COVID-19 vaccination: a large population-based cohort study in South Korea. 522 Biomark Res. 13(114). DOI: 10.1186/s40364-025-00831-w 523 8. Kim HJ, Kim MH, Choi MG, et al. (2024) Psychiatric adverse events following 524 COVID-19 vaccination: a population-based cohort study in Seoul, South Korea. Mol 525 Psychiatry. 2024(12):3635–3643. doi: 10.1038/s41380-024-02627-0 526 23 9. Kim HJ, Kim M-H, Park SJ, et al. (2025). The early impact of COVID-19 vaccines on 527 major events in cardiac, pulmonary, and thromboembolic disease: a population-based 528 study. Korean J Intern Med., 2025(40):801-812. DOI: 10.3904/kjim.2025.056 529 10. Wen CP, Tsai SP, (1982). Anatomy of the health worker effect - a critique of 530 summary statistics employed in occupational epidemiology. Scand J Work Environ 531 Health, 8(1):48-52. PMID: 7100856 532 11. Spitzer WO, Suissa S, Ernst P, et al. (1992) The use of beta-agonists and the risk of 533 death and near death from asthma. N Engl J Med., 326(8):501-506. DOI: 534 10.1056/NEJM199202203260801 535 12. Oaxaca R, (1973) Male-Female Wage Differentials in Urban Labor Markets. Int. 536 Econ. Rev. 14(3):693-709. doi:10.2307/2525981x 537 13. Blinder AS, (1973). Wage Discrimination: Reduced Form and Structural Estimates. J. 538 Hum. Resour. 8(4):436-455. DOI:10.2307/144855 539 14. D’Agostino RB Jr. (1998) Propensity score methods for bias reduction in the 540 comparison of a treatment to a non-randomized control group. Stat Med. 17(19):2265-541 2281. DOI: 10.1002/(sici)1097-0258(19981015)17:19<2265::aid-sim918>3.0.co;2-b 542 15. Chemaitelly H, Ayoub H, Coyle P, et al. (2025) Assessing healthy vaccinee effect in 543 COVID-19 vaccine effectiveness studies: a national cohort study in Qatar. eLife 2025, 544 14:e103690. DOI: 10.7554/eLife.103690 545 16. Kang MJ, Jung K-W, Bang SH, et al. (2023). Cancer Statistics in Korea: Incidence, 546 Mortality, Survival, and Prevalence in 2020. Cancer Res Treat., 55(2):385-399. DOI: 547 10.4143/crt.2023.447 548 17. Park EH, Jung K-W, Park NJ, et al. (2024). Cancer Statistics in Korea: Incidence, 549 Mortality, Survival, and Prevalence in 2021. Cancer Res Treat., 56(2):357-371. DOI: 550 10.4143/crt.2024.253 551 24 18. Park EH, Jung K-W, Park NJ, et al. (2025) Cancer Statistics in Korea: Incidence, 552 Mortality, Survival, and Prevalence in 2022. Cancer Res Treat. 57(2):312-330. DOI: 553 10.4143/crt.2025.264 554 19. Statista. South Korea: Cancer crude incidence rate by age, 2022. Statista; 2024. 555 [Accessed 2025 Nov 29]. Available from: 556 https://www.statista.com/statistics/1440818/south-korea-cancer-crude-incidence-rate-557 by-age/ 558 20. World Bank. Population ages 65 and above (% of total population) - Korea, Rep. 559 World Population Prospects, United Nations (UN). [Accessed 2025 Nov 29]. 560 Available from: 561 https://data.worldbank.org/indicator/SP.POP.65UP.TO.ZS?locations=KR 562 21. UN World Population Prospects Data, Population Pyramids, [Accessed 2025 Nov 29]. 563 Available from:https://www.populationpyramids.org/south-564 korea?utm_source=chatgpt.com 565 22. Roccetti M, Cacciapuoti G,. (2025) Beyond the Gold Standard: Linear Regression and 566 Poisson GLM Yield Identical Mortality Trends and Deaths Counts for COVID-19 in 567 Italy: 2021–2025. Computation, 13(10): 233. DOI: 10.3390/computation13100233. 568 569