A Biostatistical Reappraisal Unveiling the Mechanism Behind Apparent Cancer Risk Signals in a COVID-19 Vaccinated Cohort Author: Marco Roccetti ABiliation: University of Bologna, Department of Computer Science and Engineering, Bologna, Italy Email:
[email protected] Correspondence to:
[email protected] Abstract Background: A recent large-scale COVID-19 vaccine cohort study has reported statistical associations suggesting increased cancer risk. Our initial application of fundamental demographic and incidence metrics to that specific cohort study suggested instead a pronounced external validity discrepancy: the study’s overall cancer incidence showed a 26% deficit compared to the coherent national rate (South Korea), implying a methodological artifact. Objectives: To determine if the observed external validity discrepancy in the cohort's overall incidence is attributable to an asymmetry in the demographic composition of the exposure groups, and to quantify the resulting impact on the reported statistical association between COVID-19 vaccination and cancer incidence. Methods: We performed a core check of external validity utilizing fundamental demographic and incidence metrics for the country of interest, specifically comparing Age-Stratified Demographic Composition and Crude Incidence Rates (CRs) of the cohort against established national standards. Results: The non-vaccinated subgroup aged >= 65 showed a pronounced, asymmetric undercount of approx. 45% in expected cancer cases relative to the national age-specific rate. This asymmetric underrepresentation of high-risk elderly participants is the specific mechanism that explains the 26% overall incidence suppression, mathematically generating the appearance of excess risk in the vaccinated group. Discussion: The reported association of the scrutinized study is likely to be a statistical artifact due to marked asymmetric selection bias in cohort enrollment, fully demonstrable through fundamental demographic and cancer incidence metrics. Balanced cohorts would predictably show no statistically significant di]erence in cancer incidence between groups.
1. Introduction The global rollout of COVID-19 vaccines has been accompanied by intensive observational research aimed at thoroughly assessing post-vaccination health outcomes, including the crucial endpoint of cancer incidence. While a recent large-scale study from South Korea reported an apparent signal suggesting a higher cancer risk among their vaccinated cohort [1], a closer examination of the baseline data immediately raised a fundamental concern regarding the external validity of their findings. An initial biostatistical analysis, the methodology and findings of which are fully detailed in the subsequent Sections, indicated a striking discrepancy: the reported overall crude cancer incidence rate (CR) was dramatically lower than the corresponding o]icial national average in South Korea. Specifically, the cohort's overall CR (averaged over the period 2020–2022) was historically 22.3% lower than the national average rate for the corresponding period [2-4]. Indeed, to develop a detailed biostatistical reappraisal, we employ the latest o]icial South Korean cancer registry data from 2022 [4] (where the expected national CR is 55.02 per 10,000) as the gold standard. Applying this more recent and rigorous baseline, the suppression of the cohort's overall incidence is even more pronounced, confirming a 26% deficit. This study aims to move beyond merely noting this statistical paradox. Our objective is to provide a comprehensive quantitative explanation by reconstructing the cohort's demographics using publicly available figures. We will demonstrate explicitly how marked, asymmetric selection bias, particularly within the high-risk elderly group, provides a necessary and su]icient condition to mathematically generate the observed, yet spurious, signal of increased cancer risk in the vaccinated population. 2. Materials and Methods 2.1 Data Sources Raw cohort data, including age-stratified participant and case counts, were carefully extracted from the supplementary materials of [1] (Table S4). To establish a robust external validity gold standard, national statistics, including o]icial South Korean cancer incidence rates and demographic reports, were sourced from public registries and o]icial data [2-5]. 2.1.1 South Korea: Population Demography and OBicial Cancer Incidence To create a mathematically coherent national benchmark, we rely on a standard national South Korean population distribution approximation of 18.0% individuals aged >=65 and 82.0% individuals aged < 65, as reported in [6] and shown in Table 1 below.
Age Group Assumed % in National Population >=65 18.0% < 65 82.0% Table 1: National Demographic Information (South Korea, 2022) Table 2, instead, displays the o]icial national crude cancer incidence rates (CR) per 10,000 individuals, which serve as the external validity benchmark for the following analyses. The historical average CR (2020–2022) is shown alongside the rates stratified by age group for the 2022 Korean population [2–5]. Age Group Crude Incidence Rate (CR) per 10,000 Historical Average (2020–2022) 52.46 Total Population (Target, 2022) 55.02 Population >=65 (O]icial, 2022) 155.2 Population < 65 (Calculated Coherent) 33.03 Table 2: O]icial Crude Cancer Incidence Rates (South Korea, 2022) As a first comment on Table 2, it should be noticed that while the Total Population and the Population >=65 figures were directly sourced from [6], the Population < 65 measured was calculated coherently with the other figures. Specifically, the Crude Incidence Rate (CR) for the Population < 65 (33.03 per 10,000) was derived using the aggregate weighted incidence principle, where the total national incidence is the weighted sum of the agespecific incidence rates. We solved the equation using the established demographic weights (18.0% for >= 65 and 82.0% for <65) and the o]icial national CR for the >= 65 group (155.2), thereby ensuring the CR for the under 65 figure is entirely consistent with the o]icially published national total. More importantly, we specify now that the analysis that will follow will center only on the aggregate demographic composition of the total cohort (vaccinated and non-vaccinated combined) contrasted against the national population above. We deliberately avoided further stratifying the population by granular vaccination details (e.g., number of doses), as our core thesis, that the profound demographic bias in the high-risk age group is the defining methodological flaw, is fully demonstrable without introducing this unnecessary complexity. At this point, it is time to show the cohort composition and the relative cancer counts as sourced from [1] and summarized in the following Table 3.
Group Participants <65 Cases <65 Participants >=65 Cases >=65 Total Participants Total Cases Nonvaccinated 522,722 1,373 72,285 616 595,007 1,989 Vaccinated 2,090,888 6,861 289,140 3,283 2,380,028 10,144 Total 2,613,610 8,234 361,425 3,899 2,975,035 12,133 Table 3: Raw Demographic and Cancer Case Counts for the Cohort of [1] 2.1.2 Methods for Calculation of Cancer Incidence For providing the results that will follow in the next Section, we have extensively utilized the following three formulas which are used to compute the Crude cancer incidence (CR), the relative Demographic Deficit/Surplus and the percentage Di]erence between cancer incidence rates. The three key mathematical formulas employed are as follows [7-9]: CR (per 10,000) = (Number of Cases / Population at Risk) x 10,000 (1) Relative Deficit/Surplus % = (Cohort Percentage - Population Percentage) / Population Percentage x 100 (2) Di@erence (% incidence) = (Cohort Incidence Rate - Reference Incidence Rate) / Reference Incidence Rate x 100 (3) We conclude this Section by providing a succint but exhaustive explanation for the paradox mentioned in the Abstract and Introduction. It is su]icient to calculate the CR for the total cohort of [1] (12,133 / 2,975,035 x 10,000, as per Formula 1), yielding 40.78 per 10,000 individuals. Now, when comparing this to the average CR for South Korea during 2020-2022 of 52.46 per 10,000, or the CR for the year 2022 alone of 55.02 per 10,000 (both recoverable from Table 2), calculating the deviation between the pairs of values yields approximately 22.3% in the former case and 26% in the latter. 3. Results We first present results quantifying the demographic discordance, that is the di]erence between the actual national South Korean population structure and the composition represented within the cohort of [1], and the resulting over/under-representation of the two age groups against the established 2022 national standards (Table 4).
Age Group % in Total Cohort % in Population (Assumed 18.0%/82.0%) Absolute DiBerence (Cohort–Population) Relative Deficit/Surplus >=65 12.2% 18.0% -5.8% -32.2% (Deficit) < 65 87.8% 82.0% +5.8% +7% (Surplus) Table 4: Age Representation Comparison for the Cohort of [1] vs. National Population These results clearly indicate a pronounced 5.8% absolute deficit in the total cohort's representation of individuals aged >=65 compared to the coherent national demographic (18.0%). Moreover, the Relative Deficit/Surplus figure powerfully quantifies the severity of the bias. For the high-risk >=65 group, the relative deficit is -32.2%. Since the CR for the >=65 group (155.2 in Table 2) is approximately 4.7 times higher than the CR for the < 65 group (33.03), this severe relative deficit of -32.2% in the age bracket typically most susceptible to cancer is likely to be the overwhelming primary driver of the overall incidence suppression observed in the cohort. Conversely, the less-susceptible < 65 age group shows a relative surplus of almost +7%. We now provide results relative to the Crude Incidence Rate (CR) measured within the cohort of [1], stratified by age groups (Under 65 in Table 5 and Over 65 in Table 6), where each group's cohort incidence is contrasted with the respective South Korean national age-specific incidence rate established earlier in Table 2. Group Cohort CR (/10,000) Coherent National CR (/10,000) DiBerence (% incidence) Nonvaccinated 26.3 33.03 -20.4% Vaccinated 32 33.03 -3.1% Total 31.5 33.03 -4.6% Table 5: Crude Cancer Incidence Comparison for Participants Under 65 (2022 Standard)
Group Cohort CR (/10,000) OBicial National CR (/10,000) DiBerence (% incidence) Nonvaccinated 85.2 155.2 -45.1% Vaccinated 113.5 155.2 -26.9% Total 107.9 155.2 -30.5% Table 6: Crude Cancer Incidence Comparison for Participants Aged 65 and Over (2022 Standard) Importantly, these results show that the critical and dramatically asymmetric di]erence lies particularly in the high-risk group (population aged 65 and over), compared against the o]icial 155.2 rate. In particular, the most pronounced and crucial finding is the asymmetric underrepresentation of cancer cases among non-vaccinated participants aged >= 65: specifically, the -45.1% deficit in the non-vaccinated cohort (Table 5). This marked, unequal undercount in the high-risk, non-vaccinated reference group artificially suppresses the true baseline risk. This fundamental lack of cancer cases in the elderly non-vaccinated population (which would be expected to be substantial given the national incidence rate) most likely explains the spurious result of [1] showing an apparently higher risk for the vaccinated group. This quantitative deficiency is definitely the precise mechanism that has produced the appearance of increased risk in the vaccinated group in [1], providing the comprehensive explanation for the 26% overall incidence suppression. 4. Discussion The primary strength of this reappraisal lies in its strictly quantitative focus on fundamental demographic and incidence metrics. We conclusively demonstrate the existence and the exact magnitude of the demographic bias solely through the careful analysis of publicly available data against established national gold standards. This approach avoids the inherent complexities and controversies of clinical arguments, thereby strengthening the validity of the conclusion. Specifically, our methodology avoids entanglement in endless debates surrounding: • Clinical Causality: The speculative question of whether the vaccine could biologically cause cancer. • Vaccination Status Specificity: Ambiguities regarding the definition of a fully vaccinated or boosted status.
By focusing purely on the pronounced and asymmetric external validity flaw (the 26% incidence suppression and the 32.2% elderly deficit), our findings are robust and fully su]icient to explain the published result as a statistical artifact. The quantitative results, confirming the overall 26% suppression of the cohort's CR and the dramatic 45.1% cancer undercount in the oldest non-vaccinated subgroup, firmly establish that the apparent increased cancer risk is a systematic biostatistical artifact rooted in compromised external validity. The observed di]erence is not driven by a genuine biological signal related to the vaccine but is rather a direct consequence of the highly unequal and biased selection of non-vaccinated, high-risk elderly individuals. The severe deficit in the elderly non-vaccinated subgroup means the "unvaccinated" reference group used in [1] was demographically unrepresentative, consequently skewing the crucial baseline risk estimate importantly downwards. This purely demographic mechanism, when quantified, fully resolves the statistical paradox previously exposed. The root cause of this demographic distortion almost certainly lies in the study's cohort selection process, highly likely involving a flawed statistical matching procedure. This is strongly suggested by the final reported cohort sizes. Given the standard PSM (Propensity Score Matching) procedure adopted by the authors of [1], the smaller group (unvaccinated, approx. 600,000 should have been the control group, and the larger group (vaccinated, approx. 2.4 million) the treatment group, with the PSM defined to balance covariates. However, the exact numerical correspondence reported in [1] (Vaccinated / Unvaccinated approx. 4.00004) indicates that the smaller unvaccinated group was used as the base '1' for the matching, thus yielding a final 4:1 ratio that confirms the inversion of the standard PSM procedure. This fundamental flaw led to the non-vaccinated reference group disproportionately representing the younger, lower-risk segment of the general population. This epidemiological phenomenon, commonly known as the Healthy User Bias has been already widely documented in observational vaccine studies [10]. By artificially suppressing the expected cancer incidence in the oldest non-vaccinated cohort segment, the study created a baseline that was profoundly unrepresentative. Consequently, the less-suppressed incidence rate in the vaccinated group is made to falsely appear as a spurious excess risk. Ultimately, the non-vaccinated baseline exhibits an artificially low cancer incidence, which then causes the less-suppressed incidence rate in the vaccinated group to falsely appear as a spurious excess risk. While this quantitative assessment provides a robust quantitative explanation, it is obviously also subject to an intrinsic limitation that must be transparently acknowledged. Essentially, our entire work constitutes a re-analysis of aggregate published data; we did not perform primary data collection. Lacking access to individuallevel, granular data made it impossible to perform advanced, internal bias correction methods (such as detailed Propensity Score Matching or direct rate standardization for
small sub-groups) to correct the internal selection distortion. Consequently, our strong conclusions rely on the quantitative analysis of external validity (i.e., comparing the cohort to the national gold standard) rather than correcting the cohort's internal structure. Notwithstanding this limitation, the sheer magnitude of the observed demographic and incidence deficit is so substantial and asymmetric that the demonstrated bias is highly likely to be the singular, primary explanation for the flawed results reported in [1]. 5. Conclusion The apparent signal of excess cancer risk reported in vaccinated individuals in [1] is the likely consequence of severe, asymmetric selection bias within the study cohort. Our biostatistical analysis confirms that this demographic distortion, specifically the 45.5% undercount of expected cancer cases in the high-risk non-vaccinated elderly subgroup, artificially suppressed the baseline cancer risk. We conclude that a methodologically rigorous and demographically balanced cohort would predictably show no statistically important di]erence in cancer incidence between vaccinated and non-vaccinated groups, thereby confirming the safety profile of the COVID-19 vaccine with respect to cancer risk. Data Availability Statement All analyzed data are publicly accessible and can be sourced from the study [1] (Table S4),and from South Korean National Cancer Statistics and Demographic Information (2022) [2-6]. Author Contributions MR conceived and designed the study, carried out all data collection and analysis, interpreted the quantitative results, and was the sole author responsible for writing and revising the manuscript. The author a]irms full responsibility for the integrity of the data and the accuracy of the data analysis presented. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. This study was conducted entirely independently by the author using personal resources.
Conflict of Interest The author declares that there is no conflict of interest, financial, personal, or otherwise, that could be construed as influencing the results or the conclusions presented in this paper. Generative AI statement The author declares that no Gen AI was used in the creation of this manuscript. 6. References 1. Kim HJ, Kim M-H, Choi MG, Chun EM. (2025) 1-year risks of cancers associated with COVID-19 vaccination: a large population-based cohort study in South Korea. Biomark Res. 13(114). DOI: 10.1186/s40364-025-00831-w 2. Kang MJ, Jung K-W, Bang SH, et al. (2023). Cancer Statistics in Korea: Incidence, Mortality, Survival, and Prevalence in 2020. Cancer Res Treat., 55(2):385-399. DOI: 10.4143/crt.2023.447 3. Park EH, Jung K-W, Park NJ, et al. (2024). Cancer Statistics in Korea: Incidence, Mortality, Survival, and Prevalence in 2021. Cancer Res Treat., 56(2):357-371. DOI: 10.4143/crt.2024.253 4. Park EH, Jung K-W, Park NJ, et al. (2025) Cancer Statistics in Korea: Incidence, Mortality, Survival, and Prevalence in 2022. Cancer Res Treat. 57(2):312-330. DOI: 10.4143/crt.2025.264 5. Statista. South Korea: Cancer crude incidence rate by age, 2022. Statista; 2024. [Accessed 2025 Nov 2]. Available from: https://www.statista.com/statistics/1440818/south-korea-cancer-crudeincidence-rate-by-age/ 6. World Bank. Population ages 65 and above (% of total population) - Korea, Rep. World Population Prospects, United Nations (UN). [Accessed 2025 Nov 2]. Available from: https://data.worldbank.org/indicator/SP.POP.65UP.TO.ZS?locations=KR 5 7. Rothman KJ, Greenland S, Lash TL. (2008) Measures of Disease Occurrence. In: Modern Epidemiology. 3rd ed. Philadelphia: Lippincott Williams & Wilkins; 2008. ISBN: 9780781755641 8. Casini L, Roccetti M, (2022) Reopening Italy's schools in September 2020: A Bayesian estimation of the change in the growth rate of new SARS-CoV-2 cases BMJ Open. 11(7):e051458. DOI: 10.1136/bmjopen-2021051458