The Wuhan Institute of Virology constructed a novel merbecovirus-MERS Gain of Function chimera that contaminated pre-pandemic rice sequencing datasets from Wuhan Steven E Massey Biology Dept, University of Puerto Rico-Rio Piedras, San Juan, Puerto Rico PR 00901, USA
[email protected] 1
Abstract Previously, an infectious clone of a novel merbecovirus was described contaminating pre-pandemic rice sequencing datasets from Wuhan. Of particular concern was evidence of an additional infectious clone chimera that incorporated the MERS-CoV spike gene into the novel merbecovirus backbone, which would be expected to increase infectivity in humans due to enhanced binding to the human DPP4 receptor and the presence of a furin cleavage site. Sequences flanking the novel merbecovirus genome are shown here to match pBAC-CMV, a plasmid constructed by Dr Zhengli Shi of the Wuhan Institute of Virology. The construct is shown to be circular and so transformable and transfective. The RNA-dependent RNA polymerase (RdRp) sequence of the infectious clone has a 100 % match with a RdRp sequence generated by Shi, ‘isolate 152762’, sampled from lesser bamboo bat Tylonycteris pachypus, collected from Southern China, and published in Latinne et al (2020). Additional evidence is presented of genetic manipulation of the chimera MERS-CoV spike sequence using No See’m technology and the BsaI restriction enzyme. The new information affirms that the infectious clones were made by Shi, prior to contaminating samples from other labs during preparation and sequencing. The construction of the chimera may violate the Biological Weapons Convention, as a protective purpose is elusive. NIAID grant R01AI110964, awarded to Dr Peter Daszak of EcoHealth Alliance, funded the creation of pBAC-CMV and, with USAID PREDICT, sampling of isolate 152762 and generation of its RdRp sequence, thereby contributing directly to dangerous Gain of Function work. These observations confirm undeclared and dangerous Gain of Function experimentation on novel bat coronaviruses from Southern China by Zhengli Shi in Wuhan immediately prior to the COVID-19 pandemic. 2
Introduction There are two competing hypotheses for the origin of the COVID19 pandemic: a natural zoonosis event from horseshoe bats or an intermediate host [1], or escape from a laboratory [2] [3]. The latter lab leak hypothesis implies that a paper trail exists, as experiments are formulated in writing, often in conversation with domain experts, and submitted for funding. Labs keep records of sampling expeditions, experimental work and biosafety approvals. In addition, sequence datasets deposited in publicly accessible databases may contain clues as to experiments that institutions have not been willing to reveal. One type of clue derives from sequence contamination of publicly available datasets, which can reveal otherwise unknown experiments on viruses [4], or suppressed virus sequences from early in outbreaks [5]. Previously, we described an infectious clone containing a novel bat merbecovirus, that had contaminated rice datasets from Wuhan [6]. The datasets were contained in NCBI project PRJNA602160, generated for a project on rice mRNA methylation [7]. The novel merbecovirus was termed ‘HKU4r-HZAU-2020’, ‘HKU4r’ in reference to its phylogenetic relatedness to bat HKU4 coronavirus (HKU4-CoV) [8], ‘HZAU’ in reference to Huazhong Agricultural University, Wuhan, who generated the datasets, and ‘2020’ in reference to the year that the datasets containing the sequence were made available (February 2020). The presence of the HKU4r-HZAU-2020 infectious clone contaminating rice sequences reflects lax biosafety in Wuhan institutions researching coronaviruses. Additionally concerning was evidence for a second infectious clone that had a Middle East Respiratory Syndrome coronavirus (MERS-CoV) spike sequence inserted into the HKU4r-HZAU-2020 backbone, referred to here as HKU4r-HZAU-2020-MERS(S). The insertion of the MERS-CoV spike sequence would be likely to increase human infectivity of the chimera given the enhanced affinity for the human DPP4 receptor exhibited by MERS-CoV receptor binding domain (RBD) compared to the HKU4r-HZAU-2020 RBD [6], and the existence of a furin cleavage site (FCS) at the S1/S2 boundary in MERS-CoV spike [9]. The FCS is absent in both HKU4-CoV [10] and HKU4r-HZAU-2020. Indeed, it has been shown that the lack of a FCS in HKU4-CoV is 3
partly responsible for its lack of infectivity in humans as compared to MERS-CoV [10]. MERS-CoV has a mortality rate of 22-36 % [11] [12], so the addition of known pathogenicity factors from MERS-CoV to novel coronaviruses is a matter of concern. Plasmid sequences that flanked the HKU4r-HZAU-2020 genome sequence were found to match parts of a plasmid from a 2006 patent by Dr Luis Enjuanes and colleagues [13]. Construction of the plasmid was described in [14]. Two constructs were made; one was infectious and contained SARS-CoV (pBAC-SARS-CoV) and another was non-infectious and contained a SARS-CoV replicon (pBAC-SARS-CoV-REP). Interestingly, pBAC-SARS-CoV-REP was used by Dr Zhengli Shi of the Wuhan Institute of Virology (WIV) in a 2008 paper [15]. The Acknowledgements section indicates that the plasmid was provided directly by Enjuanes to Shi, who is thanked for ‘technical advice’. Further analysis presented here determines that the sequences flanking the HKU4r-HZAU-2020 genome exactly match a plasmid constructed by Shi termed ‘pBAC-CMV’, described in a 2016 paper [16]. It is also shown that the HKU4r-HZAU-2020 RNA dependent RNA polymerase (RdRp) gene has a 100 % match to ‘Tylonycteris bat coronavirus HKU4 isolate 152762’, generated by Shi and published in Latinne et al [17]. In addition, further evidence of the use of No See’m cloning in the creation of the HKU4r-HZAU-2020-MERS(S) chimera is presented. These findings indicate that Shi had a direct role in the construction of the HKU4r-HZAU-2020 infectious clone and HKU4r-HZAU-2020-MERS(S) chimera, and had been secretly engaging in dangerous and sloppy Gain of Function work on novel bat coronaviruses in Wuhan prior to the COVID-19 pandemic. The implications for the origin of SARS-CoV-2 and the COVID-19 pandemic should be clear. Methods Straightforward and reproducible bioinformatics methods were used throughout. Sequence alignments were conducted using ClustalW [18]. Restriction sites were identified using NEBcutter [19]. Sequence mapping was conducted using Bowtie2 4
v2.5.4 using the –local option [20]. Sorting of the resulting sam file was conducted using samtools [21], in preparation for visualization with the Integrative Genome Viewer (IGV) [22]. Sequence Data Previously, a trimmed contig, k141_13283, generated from datasets in National Center for Biotechnology Information (NCBI) project PRJNA602160 and containing the 38434 bp HKU4r-HZAU-2020 infectious clone, as described in [6], can be found in the Supplementary Material at https://zenodo.org/records/8351689, The annotated genome of HKU4r-HZAU-2020 is available at the NCBI (Accession BK068617). Other sequences used in this study include the SARS-CoV Urbani genome (Accession AY278741), plasmid pBeloBAC11 (Accession U51113), sequence 11 from patent WO2006136448 (confirmed here as pBAC-SARS-CoV) [13](Accession CS480537), and the MERS-CoV reference genome (Accession NC_019843). A complete sequence for pBAC-CMV, generated in this study, can be found in Supplementary File 2. Results Description of pBAC-CMV from the literature First, there follows a description of the construction and use of pBAC-CMV from the literature. The pBAC-CMV system uses an approach pioneered by Enjuanes, specifically the use of the cytomegalovirus (CMV) promoter upstream of a coronavirus genome, and hepatitis delta virus (HDV) ribozyme and bovine growth hormone (BGH) transcription terminator downstream of the coronavirus genome. These elements were used in the construction of pBAC-SARS-CoV [14], and were added by Shi to the pBelo-BAC11 plasmid [23] to create the pBAC-CMV plasmid [16]. 1) pBAC-CMV was constructed by Shi [16] to clone SARS-related WIV1 [24]. The WIV1 genome sequence was split into eight fragments and ligated internally 5
using BglI. The 5’ and 3’ ends of the genome were ligated into the pBAC-CMV plasmid using SacII and AscI restriction sites, respectively. 2) Shi used pBAC-CMV to construct WIV1-SARSr(S) chimeras, using spike genes from the newly characterized SARSr-CoVs Rs4231 and Rs7327 [25]. Insertion of SARS-related spike genes into the WIV1 backbone was accomplished by ligating all constituent fragments including the plasmid in one go, using a fragment containing the new spike gene to replace that of WIV1 spike. A combination of BsmBI, BsaI and BglI restriction enzymes were used for assembly of the chimera genome fragments. BsmBI and BsaI were used to insert the new spike sequences into the WIV backbone, using a No See’m approach. SacII (5’) and AscI (3’) sites were used to insert the chimeric genome sequences into the pBAC-CMV plasmid. 3) Shi used pBAC-CMV to clone a MERS-CoV replicon, which lacked the spike (S), envelope (E) and membrane (N) genes [26]. The replicon genome was synthesized as 10 fragments, and ligated using a combination of BglI and BsmBI. SacII (5’) and AscI (3’) sites were used to insert the replicon genome sequence into the plasmid. The replicon construct was subsequently provided to Dr Kuanhui Xiang of the Peking University Health Science Center, as reported in a 2025 publication [27]. 4) Shi used a ‘modified pBeloBAC11’ plasmid (presumably pBAC-CMV) in collaboration with Dr Huan Yan of Wuhan University [28] to make an infectious clone of bat HKU5 merbecovirus [8]. It is puzzling why information regarding the identity and provenance of the plasmid was not provided. Seven fragments were assembled using BsaI and BsmBI. AvrII (5’) and AscI (3’) sites were used to insert the virus genome sequence into the plasmid. Analysis of patent sequence 11 The 37971 bp patent sequence 11 was compared to the SARS-CoV Urbani genome. There was a 100 % match from position 1 to 29728. A 25 bp poly(A) tail ran from position 29729 to 29753. This confirms that the patent sequence is the pBAC-SARS-CoV infectious clone [14]. The exact match with the Urbani reference 6
sequence confirms that existing restriction sites in the genome sequence were used for assembly of the genome fragments, rather than a No See’m approach [14]. Such an approach allows amplification of the plasmid containing the genome sequence in bacteria, and excision and replacement of select fragments in order to manipulate the genome sequence, as described in [29]. The remaining portion of the patent sequence from positions 29753 to 37971 was used for comparison with the plasmid sequences identified in contig k141_13283. Matches of k141_13283 to pBeloBAC11 pBeloBAC11 was used as the plasmid backbone for pBAC-CMV (as for pBAC-SARS-CoV) [16]. The leading and trailing sequences of the HKU4r-HZAU-2020 infectious clone (contig k141_13283) were compared to the sequence of pBeloBAC11, with the following findings. 1161 bp of the leading sequence of the HKU4r-HZAU-2020 infectious clone shows a 100 % match to pBeloBAC11 (Supplementary File 1). The 3’ end of the matching region terminates in a BamHI restriction site (GGATCC). The use of pBeloBAC11 and a flanking BamHI site is reported in the construction of pBAC-CMV [16]. However, the 1161 bp region only shows an 85 % match to the pBAC-SARS-CoV plasmid region, indicating that k141_13283 is divergent from the pBAC-SARS-CoV plasmid at the sequence level. 6026 bp of the trailing sequence of k141_13283 (from position 32409 to 38434) shows a 100 % match to pBeloBAC11 (Supplementary File 1). The 5’ start of the matching region begins with a HindIII site (AAGCTT). These features are also reported in the construction of pBAC-CMV [16]. The 6026 bp region also shows a 100 % match to pBAC-SARS-CoV, which is consistent with the observation that pBeloBAC11 plasmid was also used as the basis for the pBAC-SARS-CoV plasmid [30] [14]. Mapping reads from Bioproject PRJNA602160 to the pBeloBAC11 plasmid 7
Previously, reads from four datasets from Bioproject PRJNA602160 were found to map to the HKU4r-HZAU-2020 infectious clone [6]. These were SRR10915173, SRR10915174, SRR10915167 and SRR10915168. Reads from the four datasets were mapped to the pBeloBAC11 plasmid (Figure 1). The reads mapped to the entire plasmid sequence, with a 15 bp gap at position 359 to 384 of the plasmid sequence (Accession U51113). The gap is flanked by BamHI (5’) and HindIII (3’) restriction sites, corresponding to sites that were used to insert the CMV, HDV ribozyme and BGH terminator sequences during the construction of the pBAC-CMV plasmid [16] (described in more detail below). No SNVs compared to the pBeloBAC11 plasmid were observed which means that none of its restriction sites have been altered. These results confirm that the plasmid was circular when sequenced, which means that the infectious clone was able to transform bacteria during sample preparation and sequencing. In addition, circular BAC plasmids are optimal for transfection compared to linear [31], which means the infectious clone was able to transfect mammals (including lab workers) and mammalian cell lines during sample preparation and sequencing. Both these observations have biosafety implications given that the HKU4r-HZAU-2020 infectious clone sequences leaked into other experiments. A complete sequence for pBAC-CMV, generated by this study, can be found in Supplementary File 2. Figure 1 Reads from Bioproject PRJNA602160 mapped to pBeloBAC11 Reads from SRR10915173, SRR10915174, SRR10915167 and SRR10915168, contained in Bioproject PRJNA602160, were mapped to plasmid pBeloBAC11 as described in Methods. Mapped reads are shown below the central horizontal line. A 15 bp gap indicated by the arrow is flanked by BamHI (5’) and HindIII (3’) restriction sites. The figure was generated using the IGV. 8
Match of k141_13283 to the CMV promoter The k141_13283 sequence downstream of the BamHI site from positions 1152 to 1749 had a 100 % match to the CMV promoter of pcDNA3.1(+) (Thermofisher Scientific) (Supplementary File 1), as described in the construction of pBAC-CMV [16]. 3’ of the matching region was a NheI site. The region from position 1152 to 1749 of k141_13283 had a 99.8 % match to the pBAC-SARS-CoV sequence, with a single nucleotide mismatch at position 1152 of contig k141_13283. Upstream of this position, k141_13283 has a BamHI site while pBAC-SARS-CoV has a SfoI site (GGCGCC), as shown in Figure 2. Consistent with this, the CMV promoter in pBAC-CMV is described as being flanked at its 5’ end by BamHI [16], while in pBAC-SARS-CoV it is described as being flanked at its 5’ end by a SfoI site [14]. These observations confirm that the plasmid in k141_13283 is different from that used in pBAC-SARS-CoV. Figure 2 Alignment of 5’ end of the CMV promoter in k141_13283 and pBAC-SARS-CoV 9
T. pachypus HKU4-related viruses are an abundant group of viruses in Southern China, independently sampled by more than one group. The closest matching genome in the NCBI database to HKU4r-HZAU-2020 is Tylonycteris pachypus TpGX16 isolate Q290 (Accession OQ175310), with a 98.6 % match. This genome sequence was reported in a study published in 2023 [35]. However, there is no evidence of collaboration between the authors of the study and Shi (Grok, https://grok.com). There is a 98.0 % match of HKU4r-HZAU-2020 with the Tylonycteris bat coronavirus HKU4 isolate CZ07 genome (Accession MH002338), isolated by Shi in 2012 from Southern China. The associated (unpublished) publication is described as being submitted in February 2018, but the entry date for the genome sequence is March 2020. When the RdRp sequence of HKU4r-HZAU-2020 was examined, it revealed a 100 % match with Tylonycteris bat coronavirus HKU4 isolate 152762, over its length of 399 bp (Accession MN312732). This sequence was generated by Shi and reported in Latinne et al [17]. In its NCBI data entry, the sequence is described as having been generated from the lesser bamboo bat Tylonycteris pachypus. It is highly likely that the HKU4r-HZAU-2020 genome sequence is the full genome sequence of isolate 152762. Latinne et al report that all samples were collected from 2010-2015 [17]. The study was partially funded by NIAID grant R01AI110964 [36], of which Shi was co-PI and Dr Peter Daszak of EcoHealth Alliance was PI. However, the dates of 2010-2015 seem inconsistent with the Annual Reports of grant R01AI110964, which describe sampling from January 2014-October 2018 (Years 1-5) [36] [37]. The Year 1-4 Annual Reports of the grant describe the generation of multiple novel HKU4r-CoV sequences, collected from the Southern Chinese provinces of Yunnan, Guangxi, and Guangdong [36]. The range of T. pachypus is restricted to Southern China [38]. These observations are consistent with HKU4r-HZAU-2020 being generated from a sample from Southern China. 16
Consideration of the MERS-CoV spike insertion into HKU4r-HZAU-2020 Reads spanning the 5’ and 3’ junctions between the HKU4r-HZAU-2020 backbone and the MERS-CoV spike sequence are shown in Supplementary Figure 2. As noted previously, these junctions do not show any 5or 6-cutter restriction sites [6]. This is consistent with the No See’m methodology that Shi used for inserting SARS-related spike proteins into the WIV1 backbone, where the Type IIS restriction enzymes BsmBI and BsaI were used [25]. There is a C25514T SNV in the MERS-CoV spike sequence (numbered in reference to the MERS-CoV genome sequence), present in the four reads that span the 3’ end of the MERS-CoV spike sequence and the HKU4r-HZAU-2020 backbone (Supplementary Figure 2b). The corresponding position in HKU4r-HZAU-2020 is also a T. The SNV is synonymous, and results in the ablation of two restriction sites, DdeI (5 cutter CTNAG) and Hpy166II (6 cutter GTNNAC), and so it may have been deliberate. Alternatively, it may represent a ligation artefact, or a mistake in the sequence design. There is a C21695T SNV present in the MERS-CoV spike reads present in the Bioproject PRJNA602160 dataset, when compared to the MERS-CoV reference sequence (numbered according to the MERS-CoV reference genome) [6]. The C21695T SNV is synonymous. Interestingly, C21695T results in ablation of the sole BsaI restriction site present in the MERS-CoV spike reference sequence (Figure 6) and so was likely deliberate. This observation indicates that BsaI was one of the enzymes used to digest the fragment containing the MERS-CoV spike gene, before ligation into the HKU4r-HZAU-2020 backbone, and confirms that the No See’m technique was used for creation of the chimera, specifically the BsaI enzyme. It may be reasoned that BsaI was used to digest at least the 3’ end of the MER-CoV spike fragment prior to ligation with the HKU4r-HZAU-2020 backbone. This is because there are two BsmBI sites downstream of the HKU4r-HZAU-2020 spike gene (Supplementary Figure 1b). If a BsmBI site was used at the 3’ end of the MERS-CoV spike fragment, this would require a BsmBI site at the 5’ end of the HKU-HZAU-2020 fragment that contains the two downstream BsmBI sites. This is a problem because digestion with BsmBI would result 17
in digestion of the internal fragment BsmBI sites also, resulting in fragmentation. Consequently, the use of BsmBI to digest the 3’ end of the spike fragment would have been avoided. Figure 6 C21695T-mediated disruption of the BsaI site in the MERS-CoV spike sequence The figure shows the MERS-CoV spike reference sequence (‘MERS Spike Ref’) from position 21661 aligned to a matching consensus sequence generated from NCBI Bioproject PRJNA602160 (the read depth at position 21695 is 12 [6]). Position 21695 is in red, the BsaI recognition site GAGACC (the reverse complement of GGTCTC) is in yellow. It may be observed that T21695 disrupts (ablates) the BsaI recognition site. The ligation of the MERS-CoV spike sequence with HKU4r-HZAU-2020 backbone sequences indicates the presence of a second complete infectious clone. This is because in Shi’s pBAC-CMV infectious clone methodology all constituent genome fragments and pBAC-CMV were ligated together in one go, as reported for WIV1 in [16]. Likewise, the methodology used to create WIV1 chimeras was to re-ligate all fragments and plasmid together in one go, replacing the fragment containing the spike gene [25]. This means that the presence of the complete MERS-spike gene ligated to the HKU4r-HZAU-2020 backbone indicates the presence of a complete infectious clone. Lastly, in addition to the enhanced binding of MERS-CoV spike RBD to hDPP4, compared to HKU4r-HZAU-2020, and the presence of a FCS in MERS-CoV spike (absent in the HKU4r-HZAU-2020 spike), there is an additional feature of the MERS-CoV spike that would be expected to cause increased infectivity in humans. MERS-CoV spike has a human endosomal cysteine protease (hECP) cleavage site, 18
AFNH, at amino acid positions 763-766. Cleavage of the hECP site facilitates viral entry into human cells [10]. HKU4-CoV has NYTS at the same position; it has been shown that N is glycosylated and that this prevents cleavage by hECP and hence entry into human cells [10]. HKU4r-HZAU-2020 also possessed NYTS, which also likely prevents cleavage. Insertion of the MERS-CoV spike into the HKU4r-HZAU-2020 backbone would introduce the hECP cleavage site and so is reasonably expected to increase the infectivity of human cells. Shi would have been aware an increase in human infectivity arising from the presence of both the FCS and hECP cleavage site in MERS-CoV spike, compared to HKU4-CoV where they are absent, as she was co-author of the paper that characterized these two infectivity enhancing cleavage sites in MERS-CoV spike, in comparison to HKU4-CoV [10]. Discussion Evidence presented here demonstrates that the plasmid containing the HKU4r-HZAU-2020 genome is pBAC-CMV. pBeloBAC11 was used as the plasmid backbone, as for pBAC-CMV. The restriction sites used to insert the CMV promoter (BamHI), and the HDV ribozyme / BGH terminator (HindIII) into the pBeloBAC11 backbone match those used to construct pBAC-CMV [16]. Likewise, the restriction site 3’ of the HKU4r-HZAU-2020 genome insertion, AscI, matches that of pBAC-CMV. The restriction site 5’ of the HKU4r-HZAU-2020 genome insertion, NheI, does not match the sites reported used for pBAC-CMV (SacII and AvrII), however variable restriction sites were used in that position by Shi. The CMV, HDV ribozyme and BGH terminator sequences match those from the sequence sources reported for the construction of pBAC-CMV [16]. The sequences of the three primer pairs that were used for the construction of pBAC-CMV match those found in the HKU4r-HZAU-2020 plasmid, and indicate the use of OE-PCR to join the HDV ribozyme and BGH terminator sequences, as reported as reported for pBAC-CMV [16]. The 25 bp poly(A) tail of HKU4r-HZAU-2020 is consistent with that of pBAC-CMV infectious clone and replicon constructs reported by Shi [26] [28]. The key features of the HKU4r-HZAU-2020 infectious clone are displayed in Figure 6, for comparison to those of pBAC-CMV. 19
Figure 6 Key features of the HKU4r-HZAU-2020 infectious clone The figure shows the restriction sites, the plasmid sequence components, and the 25 bp poly(A) tail identified in the HKU4r-HZAU-2020 infectious clone (k141_13283). These are all features of pBAC-CMV [16] (with the exception of the NheI site). Exact sequence details can be found in Supplementary File 1. The pBAC-CMV plasmid was published by Shi in 2016. While pBAC-CMV may have been shared with colleagues, the only reference in the literature of its use by an outside group as an infectious clone is in collaboration with Yuan Han [28]. In the author contributions section of the paper it is noted that a member of Shi’s lab, Jing Chen, was responsible for construction of the HKU5 infectious clone, using pBAC-CMV. Jing Chen is also an author on the 2016 paper describing the original construction of pBAC-CMV [16] and on the 2021 paper that used pBAC-CMV to express a MERS-CoV replicon [26]. The 100 % match of HKU4r-HZAU-2020 RdRp to that of HKU4-related isolate 152762, generated by Shi, confirms that the HKU4r-HZAU-2020 construct is a product of Shi’s lab. Confirmation of the use of plasmid pBAC-CMV by Shi for the HKU4r-HZAU-2020 infectious clone, and the publication of its sequence (this work) may inform considerations of restriction site patterns in the SARS-CoV-2 genome [34]. In 2017 Shi published the results of the WIV1-SARSr(S) chimera experiments described in grant R01AI110964 [25]. Chimera experiments are also outlined in the Defense Advanced Research Projects Agency (DARPA) DEFUSE proposal [39], of which Shi was co-PI. The proposal involved the creation of chimeras that involved the insertion of the spike genes of SARS-related coronaviruses into WIV16 [40] and SHC014 [24] 20
backbones (WIV16 and SHC014 are also SARS-related coronaviruses). The existence of the HKU4r-HZAU-2020-MER(S) chimera is consistent with the chimera experiments conducted in NIAID grant R01AI110964 [36]. The grant involved the creation of WIV1-SARSr(S) chimeras and chimeras of MERS-CoV with the receptor binding domain (RBD) sequence of MERS-CoV replaced with those from HKU4-related coronaviruses. These experiments and the results reported in the Year 5 annual report, where reduced infectivity of the chimeras was reported, indicate awareness by Shi that HKU4-related RBDs have lower affinity for hDPP4 than MERS-CoV RBD. These considerations compel us to conclude that she would have been conscious of the danger of inserting the MERS-CoV spike into the HKU4r-HZAU-2020 backbone, due to its higher affinity for hDPP4 (as well as the presence of two cleavage sites absent in HKU4 coronaviruses). A rough timeline for the creation of the HKU4r-HZAU-2020-MERS(S) chimera can be determined. Sampling of HKU4-related coronaviruses was reported as ranging from 2010-2018 in Latinne et al. and grant R01AI110964. Sequencing of partial RdRp sequences was reported in Years 1-5 (2014-2019) of grant R01AI110964. Genome sequencing would have followed the identification of a HKU4r-CoV lineage of interest. The spike protein of HKU4r-HZAU-2020 is 92 % identical to that of the HKU4 spike reference (Accession EF065508) at the nucleotide level, so it appears to have been chosen for its similarity to HKU4-CoV. Given that Bioproject PRJNA602160, which contained the HKU4r-HZAU-2020 infectious clone sequences, was made available on the NCBI in February 2020, HKU4r-HZAU-2020-MERS(S) was likely constructed and sequenced during 2019 (sequencing of the infectious clone would have been conducted in order to ensure sequence correctness in comparison with a pre-existing genome sequence). 2019 corresponds to Year 5 of NIAID grant R01AI110964 (the Year 5 period terminated on 31 May 2019). The Year 5 annual report described a mirror-image experiment to that of the HKU4r-HZAU-2020-MERS(S) chimera: the insertion of novel HKU4r-CoV RBD sequences into a MERS-CoV backbone [37]. After sequence confirmation, the 21
infectious clones would be used experimentally, initially to transfect mammalian cell lines, to generate and recover ‘live’ virus. This work was done at biosafety level 2 (BSL-2) at the WIV for bat coronavirus infectious clones such as WIV1 [16] and HKU5-CoV [28]. There are reports that BSL-2 was also used for studying bat coronavirus chimeras at the WIV [41]. Using BSL-2 to work on the HKU4r-HZAU-2020-MERS(S) chimera, including both its construction and culturing, would have been highly reckless. In the Year 5 annual report of R01AI110964, it is stated that the approach used for making the MERS-CoV chimeras was similar to that used for the construction of SARSr-CoV infectious clones. This refers to the use of pBAC-CMV, initially created for a WIV1 infectious clone [16], and derived chimeras [25], as part of the same grant. The possibility therefore arises that the construction of the HKU4r-HZAU-2020-MERS(S) chimera was funded directly by the grant (in addition to creation of the pBAC-CMV plasmid), however unsurprisingly it wasn’t reported in the Annual Reports. There is a high probability that the HKUJ4r-HZAU-2020-MERS(S) chimera has enhanced infectivity in humans, due to the three MERS-CoV spike pathogenicity-enhancing factors described. There is a possibility it could have a greater mortality rate and/or transmissibility than MERS-CoV itself as the characteristics of chimeras are impossible to predict with certainty. A scientific rationale for the creation of the HKUJ4r-HZAU-2020-MERS(S) chimera is difficult to understand. The MERS-CoV spike gene is the best characterized spike gene in the merbecovirus subgenus. Placing it into a novel HKU4-related backbone is unlikely to reveal any novel property of the MERS-CoV spike itself. While it might be argued that the chimera could be used to test for the existence of MERS-CoV pathogenicity factors outside the spike gene (indicated if the chimera has lower virulence than MERS-CoV itself), the presence of such factors is already known for MERS-CoV (for example, [42]). Such an experiment would not reveal the location of any novel non-spike pathogenicity factors. The standard approach for identifying pathogenicity factors is introducing inactivating changes, such as 22
deletions, into candidate regions of a coronavirus genome, and monitoring the effect on virulence. Any observations of reduced virulence would be complicated by the possibility of potential disruptions of positive epistatic interactions of the MERS-CoV spike with the MERS-CoV backbone that construction of a chimera would entail. In addition, such a chimera would not arise via natural recombination, given several factors, including the large geographical distance between the sources of the two constituent coronaviruses (MERS-CoV in the Arabian Peninsula [43], HKU4r-HZAU-2020 in Southern China, this work). Lastly, use of the HKU4 reference genome would be desirable for scientific work, rather than a new and uncharacterized HKU4-related genome. With this in mind, it is notable that the HKU4r-HZAU-2020-MERS(S) chimera was constructed concurrently with the HKUJ4r-HZAU-2020 infectious clone. This is not a normal scientific approach: typically a novel virus would be characterized first, then further experiments designed taking into account the experimental observations. No published coronavirus chimera infectious clone in which Shi was involved used a previously uncharacterized coronavirus as a backbone [44] [25] [45] [46], in contrast to the HKU4r-HZAU-2020-MERS(S) chimera reported here. The HKU4r-HZAU-2020-MERS(S) chimera is therefore without precedent and represents a significant deviation from Shi’s established routine. Consistent with this observation, the author was unable to find any example from the literature of a coronavirus chimera infectious clone that used a previously uncharacterized coronavirus as a backbone. Consequently, given these considerations the creation of the HKUJ4r-HZAU-2020-MERS(S) chimera may violate Article I of the Biological Weapons Convention [47]. This is because it is hard to see protective or peaceful purposes arising from its development, a stipulate of Article I. In addition, the USA, via NIAID and USAID funding (discussed below), may have violated Article III of the Convention, which states: 23
‘Each State Party to this Convention undertakes not to transfer to any recipient whatsover, directly or indirectly, and not in any way to assist, encourage, or induce any State, group of States or international organisations to manufacture or otherwise acquire any of the agents, toxins, weapons, equipment or means of delivery specified in Article I of the Convention.’ The positive attribution of the HKU4r-HZAU-2020 infectious clone, and particularly the HKU4r-HZAU-2020-MERS(S) chimera, to the WIV may help to explain the reluctance of the WIV to undergo inspection in the wake of the COVID-19 pandemic. It indicates that their Institutional Biosafety Committee (IBC) was apparently willing to approve the creation of the HKU4r-HZAU-2020-MERS(S) chimera. It is an open question if other dangerous Gain of Function experiments on novel MERSor SARS-related coronaviruses were approved. There is however a possibility that the construction of HKU4r-HZAU-2020-MERS(S) went ahead without approval and knowledge of the WIV IBC. It has been noted that the construction of such BAC-based clones presents a hazard to lab personnel (and by extension the wider population) by exposure to the infectious clone DNA (lines 1201-1205 of [48]). That the infectious clones were circular and that the HKU4r-HZAU-2020-MERS(S) chimera leaked into a rice sequence dataset indicates potential exposure of lab personnel to transfectable chimera infectious clone DNA. Finally, it is worth noting that the creation of the pBAC-CMV plasmid was conducted in collaboration with EcoHealth Alliance and was partially funded by NIAID grant R01AI110964 [16]. Given that pBAC-CMV is a component of a dangerous GOF infectious clone ie. the HKU4r-HZAU-2020-MER(S) chimera, it is unequivocal that NIAID funds directly contributed to dangerous GOF at the WIV. In addition, the HKU4r-HZAU-2020 RdRp sequence has a 100 % match to HKU4-related isolate 152762, reported in [17]. The isolate 152762 sample was collected and its RdRp sequence generated using funds from grant R01AI110964, and USAID Emerging Pandemic Threats PREDICT program cooperative agreement GHN-A-OO-09-00010-00 24
[17]. Consequently, USAID-funded virus hunting also contributed directly to the creation of the HKU4r-HZAU-2020-MER(S) chimera and dangerous GOF at the WIV. Conclusions The results reported here confirm that 1. The WIV was making undisclosed infectious clones of novel beta-coranaviruses from Southern China immediately prior to the pandemic 2. The WIV was conducting dangerous Gain of Function on novel and undisclosed beta-coronaviruses, via manipulation of the spike gene using No See’m technology 3. The WIV made a chimera using a spike sequence that had enhanced binding to its human receptor, and which included a furin cleavage site at the S1/S2 boundary, which was lacking in the original spike 4. The WIV leaked transfectable novel coronavirus infectious clones into other experiments 5. WIV and EcoHealth Alliance virus hunting expeditions identified a novel merbecovirus used to create an undisclosed dangerous GOF chimera 6. The WIV used grants from NIAID and USAID PREDICT to fund the creation of a dangerous GOF chimera that incorporated key pathogenicity factors from MERS-CoV 7. The WIV Institutional Biosafety Committee apparently approved dangerous and reckless GOF work on novel bat coronaviruses, which may explain the WIV’s reluctance to undergo an audit 8. Coronavirus infectious clone technological innovations by Western scientists were adopted by WIV researchers in their pursuit of dangerous GOF 9. A scientific rationale for creation of the dangerous HKU4r-HZAU-2020-MER(S) chimera is elusive, and so it may represent a violation of Article I of the Biological Weapons Convention Acknowledgements 25
Supplementary Figure 2 Chimeric reads spanning the MERS-spike and HKU4r-HZAU-2020 backbone sequences Chimeric reads were identified by mapping the combined SRR10915173, SRR10915174, SRR10915167 and SRR10915168 datasets against the MERS-spike gene reference sequence, as described in Methods. a) shows overlapping reads at the 5’ end of the MERS spike sequence while b) shows overlapping reads at the 3’ end of the MERS spike sequence. Nucleotides that match the HKU4r-HZAU-2020 sequence are in green, those that match MERS-CoV spike are in red. Mismatching nucleotides are in blue a) SRR10915168.45695542 0 MERS-spike 1 22 99S51M * 0 0 TTAAATTAAAAGGTACACCAGTACTTCAATTAAAGGAGAGTCAGATTAACGAACTTGTAATTTCTCTCTTGTCACAAGGAAAGTTGCTTATACG TGACAATGATGCACTCAGTGTTTCTACTGATGTTCTTGTTAACGCCTACAGAAAGT AAFFFJJJJJJJJJJJJJJJJJJJJJFFJJJJJJJJJJJJJJJ-<FFJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJ -<AFFJJJAJ7<AJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJ-7AJ-FFFAFFFJ AS:i:91 XN:i:0 XM:i:2 XO:i:0 XG:i:0 NM:i:2 MD:Z:5A32A12 YT:Z:UU SRR10915168.46591921 0 MERS-spike 1 42 29S121M * 0 0 GTCACAAGGAAAGTTGCTTATACGCGACAATGATACACTCAGTGTTTCTACTGATGTTCTTGTTAACACCTACAGAAAGTTACGTTGATGTAGG GCCAGATTCTGTTAAGTCTGCTTGTATTGAGGTTGATATACAACAGACTTTCTTTG AAFFFJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJFJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJFJJJJJJ JJJFJ7FJJFJJJJJJJJJJJFJJJFJJJJJJFFJAFFJFJFJJJJFJJJJJJJJJ AS:i:242 XN:i:0 XM:i:0 XO:i:0 XG:i:0 NM:i:0 MD:Z:121 YT:Z:UU SRR10915173.28724747 0 MERS-spike 1 24 76S74M * 0 0 CTTCAATTAAAGGAGAGTCAAATTAACGAACTTGTAATTTCTCTCTTGTCACAAGGAAAGTTGCTTATACGCGACAATGATACACTCAGTGTTT CTACTGATGTTCTTGTTAACACCTACAGAAAGTTACGTTGATGTAGGGCCAGATTC AAFFFJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJFJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJFJJJJJJJJJJJJ JJJJJJJJJJJJJJJJJJJJJJJJJJJJFJJJJJJJJJJJFJFJFJJJJJ<FAJJJ AS:i:148 XN:i:0 XM:i:0 XO:i:0 XG:i:0 NM:i:0 MD:Z:74 YT:Z:UU SRR10915174.51784135 16 MERS-spike 1 22 110S40M * 0 0 TCGCTTTCCTCTTAAATTAAAAGGTACACCAGTACTTCAATTAAAGGAGAGTCAAATTAACGAACTTGTAATTTCTCTCTTGTCACAAGGAAAG TTGCTTATACGCGACAATGATACACTCAGTGTTTCTACTGATGTTCTTGTTAACAC <JJFJJJJFF<77<AA7-JJJFJJJFJJJJJJFJJFFJJJJJJJJJJJJJJJJJJJFJJJJJFJJJJJJFJFJJFF<-JJJJJFFJFJJJJJJJ JJJJJJJJJJJFJJFJJJJJJJJFJJJJJJJJJJJJJJFJJJJJJJJJJJJFFFAA AS:i:80 XN:i:0 XM:i:0 XO:i:0 XG:i:0 NM:i:0 MD:Z:40 YT:Z:UU b) SRR10915168.17466472 0 MERS-spike 3971 28 92M58S * 0 0 GCACAAACTGTATGGGAAAACTTAAGTGTAATCGTTGTTGTGATAGATACGAGGAATACGACCTCGAGCCGCATAAGGTTCATGTTCATTAAGA TTTTAACGAACTTTTATTAGCAGTATGGTTTCTTTTAACGTTACTGCTATTTTACT AAFFFJJJJJJAFFJJJJJJJJJFJJFJJJJJJJJJJJJJFJJJFJJJJJJJJJJJJAJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJ 32
JJJJJJJJJJJJJJJJJJJJJJJJJJJJFJJJJJJJJJJJJJFJJJJJJJJJJAJJ AS:i:176 XN:i:0 XM:i:1 XO:i:0 XG:i:0 NM:i:1 MD:Z:88C3 YT:Z:UU SRR10915174.29509746 16 MERS-spike 3973 28 90M60S * 0 0 ACAAACTGTATGGGAAAACTTAAGTGTAATCGTTGTTGTGATAGATACGAGGAATACGACCTCGAGCCGCATAAGGTTCATGTTCATTAAGATT TTAACGAACTTTTATTAGCAGTATGGTTTCTTTTAACGTTACTGCTATTTTACTTG JJJJJJJJFAAJJJJJJJJJAJJJJJFAFJJJJJJJJJJJJJJJFJJJJJJJJJJJJJJJJJJJJJJJJJJJJJFJJJJJJJJJJJJJJJJJJF FFJJFJJJJJJJJJJJJJJJJJJJJJJJFJJJJFFJJJJJJJJJJJJJJJJFFF<A AS:i:172 XN:i:0 XM:i:1 XO:i:0 XG:i:0 NM:i:1 MD:Z:86C3 YT:Z:UU SRR10915174.55489561 0 MERS-spike 3999 22 1S64M85S * 0 0 NTAATCGTTGTTGTGATAGATACGAGGAATACGACCTCGAGCCGCATAAGGTTCATGTTCATTAAGATTTTAACGAACTTTTATTAGCAGTATG GTTTCTTTTAACGTTACTGCTATTTTACTTGTATTAATTGCTAATGCTTTTTCTAA #AAFFJJJJJJJJJJJJJJJJFJJJJJAFJJJJFJJJJJJJJJJJJJFJJJJJJJJJJJJJJJFJJFJJJJJJJJJFJJJJJJJJJJF<FJFFF JJJJJJJJJJJJJJJ<FJFJJFFJJJ<JJFFJFJJJFJJJJJA<FFFJJJJJJJ77 AS:i:120 XN:i:0 XM:i:1 XO:i:0 XG:i:0 NM:i:1 MD:Z:60C3 YT:Z:UU SRR10915174.5864550 16 MERS-spike 4017 22 46M104S * 0 0 ATACGAGGAATACGACCTCGAGCCGCATAAGGTTCATGTTCATTAAGATTTTAACGAACTTTTATTAGCAGTATGGTTTCTTTTAACGTTACTG CTATTTTACTTGTATTAATTGCTAATGCTTTTTCTAAACCTTTATATATTCCTGAG JJJJJJJJJJJJJJAJJAJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJ JJJJJJJJJJJJJJJJJJJJJFFJJJJJJJJJJFAJAJJJJFJJJFJJJJJFFFAA AS:i:84 XN:i:0 XM:i:1 XO:i:0 XG:i:0 NM:i:1 MD:Z:42C3 YT:Z:UU 33
Supplementary File 1 Key features of the HKU4r-HZAU-2020 infectious clone sequence in comparison with pBAC-CMV The sequence of the HKU4r-HZAU-2020 infectious clone (k141_13283) was compared to the pBAC-CMV plasmid, described in [16]. Corresponding features are indicated. 34