scieee AI-readable full text Open interactive document viewer

Bayesian phylogenetic and recombination analyses of plum pox virus provide a refined vision of its evolutionary history

Palmisano, Francesco; Kawakubo, Shusuke; Leonetti, Paola; Pantaleo, Vitantonio; Candresse, Thierry; Minafra, Angelantonio

Abstract

Background The discovery of a plum tree isolate of plum pox virus (PPV, Potyvirus plumpoxi), done in Eastern Albania in 2011 in the frame of an EU-funded survey, which represents a divergent strain named PPV-An, proved to be original and informative for the unraveling of PPV evolutionary history. Methods Maximum likelihood and Bayesian phylogenetic methods applied on full-length genomes or selected regions analyzed the affinities of the PPV-An with other PPV strains. Potential recombination events were also evaluated. A refined timeline of PPV evolutionary history integrating recombination events, strains migration and ancestral host state is proposed. Results Altogether, the analyses confirm previous hypotheses that PPV-An corresponds to an ancestral, non-recombinant PPV strain. PPV-An likely served as the origin of the PPV-M and T strains through recombination with isolate(s) of the D strain. Molecular clock analyses dated the most recent common ancestor (TMRCA) of PPV at 4546 years ago and phylogeny separated the main PPV strains from the cherry-adapted strains around 3100 years ago. Meanwhile, the recombination events that gave rise to the M and T strains are estimated to have occurred in the early sixteenth century of the common era (CE). Conclusions The characterization of the PPV-An strain enabled a comprehensive phylogenetic analysis of PPV. PPV-An is confirmed to be the previously unidentified progenitor, which, together with PPV-D led through recombination to the emergence of the currently prevalent and evolutionary successful recombinant strains of European origin (e.g., M, Rec, and T). The low representation of PPV-An in current PPV populations is likely the consequence of a population replacement phenomenon possibly linked to a higher fitness of the recombinant strains deriving from it. These results highlight the PPV-An strain as a key player in PPV evolutionary history and consolidate PPV as one of the promising models to study host-adaptive evolution processes and phylogeography among the most damaging viruses of agricultural systems.

Full text

Palmisanoetal. Virology Journal (2025) 22:319 https://doi.org/10.1186/s12985-025-02892-7 RESEARCH Open Access © The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/. Virology Journal Bayesian phylogenetic andrecombination analyses ofplum pox virus provide arefined vision ofits evolutionary history F. Palmisano1,4†, S. Kawakubo2,5†, M. Chiumenti1, P. Leonetti1, V. Pantaleo1*, T. Candresse3 and A. Minafra1* Abstract Background The discovery of a plum tree isolate of plum pox virus (PPV, Potyvirus plumpoxi), done in Eastern Albania in 2011 in the frame of an EU-funded survey, which represents a divergent strain named PPV-An, proved to be original and informative for the unraveling of PPV evolutionary history. Methods Maximum likelihood and Bayesian phylogenetic methods applied on full-length genomes or selected regions analyzed the affinities of the PPV-An with other PPV strains. Potential recombination events were also evaluated. A refined timeline of PPV evolutionary history integrating recombination events, strains migration and ancestral host state is proposed. Results Altogether, the analyses confirm previous hypotheses that PPV-An corresponds to an ancestral, non-recombinant PPV strain. PPV-An likely served as the origin of the PPV-M and T strains through recombination with isolate(s) of the D strain. Molecular clock analyses dated the most recent common ancestor (TMRCA) of PPV at 4546 years ago and phylogeny separated the main PPV strains from the cherry-adapted strains around 3100 years ago. Meanwhile, the recombination events that gave rise to the M and T strains are estimated to have occurred in the early sixteenth century of the common era (CE). Conclusions The characterization of the PPV-An strain enabled a comprehensive phylogenetic analysis of PPV. PPVAn is confirmed to be the previously unidentified progenitor, which, together with PPV-D led through recombination to the emergence of the currently prevalent and evolutionary successful recombinant strains of European origin (e.g., M, Rec, and T). The low representation of PPV-An in current PPV populations is likely the consequence of a population replacement phenomenon possibly linked to a higher fitness of the recombinant strains deriving from it. These results highlight the PPV-An strain as a key player in PPV evolutionary history and consolidate PPV as one of the promising models to study host-adaptive evolution processes and phylogeography among the most damaging viruses of agricultural systems. Keywords Sharka disease, Plum pox virus, Strain dynamics, Phylogeographic inference, Recombination, Ancestor †F. Palmisano and S. Kawakubo contributed equally to this work. *Correspondence: V. Pantaleo [email protected] A. Minafra [email protected] Full list of author information is available at the end of the article Page 2 of 16 Palmisanoetal. Virology Journal (2025) 22:319 Background Potyvirus plumpoxi (plum pox virus, PPV) is a member of the genus Potyvirus, family Potyviridae. Like other potyviruses, PPV has a single-stranded positive-sense genomic RNA of about 10kb which encodes a single large open reading frame (ORF). Translation of this ORF generates a polyprotein precursor of ∼350 kDa that is in turn processed, giving rise to processing intermediates and to 10 final protein products [1]. Potyvirids also encode a second, small ORF named PIPO, via a ribosomal frameshifting within the P3 cistron. PPV is the agent of the Sharka disease, the most devastating disease of stone fruit trees worldwide [2–7]. Accordingly, PPV is considered as either a quarantine pathogen or a regulated non-quarantine pathogen in a wide range of countries [4, 5, 8]. PPV is transmitted by several species of aphids [9], which acquire the virus when probing infected plants and then transfer it in a non-persistent manner to healthy plants. Once a tree is infected, it can exhibit various symptoms including discolored leaf arabesques and rings, leaf rugosity (wrinkling), and fruit deformations such as ringspots or necrotic spots. Severe fruit drops may occur in the most susceptible varieties. In addition, PPV is transmitted through all vegetative propagation techniques, such as grafting, making the trade of Prunus spp. propagation material responsible for its long-range propagation, including intercontinental spread. The genomes of a broad range of PPV isolates have been completely sequenced [10–16] and PPV has been studied from both phylogenetic and evolutionary perspectives. Researchers have used molecular techniques to analyze the genetic diversity of PPV isolates collected from different geographic areas and host species. These studies provided insights into the evolutionary relationships among different strains of the virus and helped to trace the spread of PPV populations [17–21]. So far, 10 strains of PPV have been recognized and named. The three major strains, which show the broader geographic distribution, are PPV-D (Dideron; [22]), PPV-M (Marcus, [22] and PPV-Rec (Recombinant). The other identified strains show more limited geographic distributions, including the PPV-EA (El Amar; [10, 23]) identified in Egypt, a few other strains with a localized and sporadic presence such as PPV-W (Winona; [12]), PPV-T (Turkey; [13]), and PPV-An (Ancestor; [24]). At least 3 different strains are able to naturally infect sweet and sour cherry PPV-C (Cherry, [11]), PPV-CR (Cherry Russia, [25]), PPV-CV (Cherry Volga, [14]), while several phylogroups of cherryadapted isolates (namely SC, TAT and Y) that may represent further strains have been recently described in the South-Eastern part of Russia [15]. Understanding the phylogenetic relationships between PPV strains is crucial not only for addressing the evolutionary history of PPV but also for developing and implementing effective control strategies, such as deploying resistant cultivars or designing diagnostic tools for early strain-specific detection and spread prevention. Moreover, studying the evolutionary dynamics of PPV can provide insights into how the virus evolves in response to selective pressures, such as host resistance mechanisms or changes in vector populations [26]. The wealth of PPV genetic information available from public repositories and associated metadata makes PPV an interesting model for phylogeographic and evolutionary studies [20, 27, 28]. It is widely accepted that recombination has been a major driving force in PPV evolution. For instance, the PPV-Rec strain originated from recombination between PPV-M and PPV-D, with a breakpoint in the NIb gene 3’region [29]; due to its typical PPV-M coat protein, it was long misidentified as PPV-M [7]. Similarly, a Canadian PPV-W isolate was identified as a complex recombinant involving PPV-W, PPV-M, and PPV-D [30]. Moreover, although the P1 protein is the least conserved among potyviral proteins and is believed to be involved in potyviruses host adaptation [31], the 5’ part of the genome shows high homology among the PPVD, -M, -T, and -Rec strains but not with other strains, suggesting these four strains are linked by (an) ancestral recombination event(s). The precise recombination history linking these strains is complex to unravel, and two tentative recombination breakpoints have been proposed in the HC-Pro and the P3 genes [17]. Two scenarios for the evolutionary history of PPV have been proposed, in which either PPV-D or PPV-M (and PPVT) would be (a) recombinant strain(s), while the other would have contributed as one of the parents involved in the corresponding recombination event [17]. In both scenarios, the existence of an ancestral, nonrecombined form of PPV-D or PPV-M was postulated. Partial genome sequencing of PPV isolates collected during field surveys performed in Albania in 2011 demonstrated the presence of several PPV isolates [32](F. Palmisano and S. Dallot, unpublished) and allowed the identification of an unusual isolate named AL-11pl, which was characterized by a divergent 5’genomic region while the rest of its genome is more typical of already known isolates of the PPV-M and -T strains. These features fit with the properties hypothesized by Glasa & Candresse [17] for a putative ancestral strain which could have contributed, in different recombination events with PPV-D, to the emergence of the -M and -T strains [7]. The discovery and the characterization of the AL-11pl isolate thus provide support for one of the two alternative evolutionary scenarios proposed by Glasa & Candresse [17] and Page 3 of 16 Palmisanoetal. Virology Journal (2025) 22:319 lead to the naming of the corresponding strain as PPV-An (Ancestral of Marcus strain) [24]. Although an initial analysis found faint evidence linking PPV-T to PPV-An [19], a global phylogenetic analysis confirmed a strict association between PPV-An and the PPV-T clade [18]. For the first time, these latter authors [18] proposed a recombination events-based global PPV phylogeny on a dataset of 206 isolates, and estimated the age of the most common recent ancestor at about 800years before common era through sub-tree comparison method. Recent advances in the study of the phylogeny and molecular evolution of potyviruses [33] have integrated Bayesian phylodynamic analysis and dating approaches, as demonstrated in research on PVY [34] and TuMV [35, 36]. Genome characterization and Bayesian phylogenetic inference of georeferenced isolates of these viruses have facilitated the reconstruction of their spread and evolutionary history [37, 38]. In the present study, we used Maximum Likelihood phylogeny and Bayesian approaches for the analysis of molecular evolution and timing of PPV strains differentiation [39–42]. For this, we used the complete AL-11pl sequence and those of an additional 539 PPV isolates to tentatively date the appearance of the PPV-An-derived lineage. From these elaborations, it was possible to propose a hypothetical scenario about the geographical origin and host ancestral state of PPV and for the further spatio-temporal spread of its strains. Methods Recovery frompublic databases ofPPV sequences associated withtheir metadata andreconstruction oftheir phylogeny A total of 609 full length PPV genomes were downloaded on 07 November 2023 from the NCBI Virus database (https:// www. ncbi. nlm. nih. gov/ labs/ virus/ vssi/#/). Sequences showing frameshifts or interruptions of the polyprotein were excluded from further analysis. Metadata (name of isolate, sampling year, host and country) associated with the remaining sequences were downloaded and manually curated. A total of 539 sequences for which complete metadata were available, were kept for subsequent analyses (Additional Table1). A multiple alignment of these full-length genomic sequences was obtained using MAFFT [43]. The aligned complete genomic sequences were further manually checked in Geneious v.2024.0.7 to identify the ends of the large ORF encoding the polyprotein and the correct genome termini. Pairwise identity percentages were calculated from the aligned positions. In addition to the phylogenetic analysis run on the fulllength genome alignment, three distinct genomic regions were selected, avoiding the breakpoints of known recombination events already described and confirmed here (see below and Fig.2D) [7, 17]: (i) a 5’ terminal region (nt 36–1400, the numbering used throughout follows the consensus alignment using all 539 isolates); (ii) an internal 5’ fragment (nt 1500–2600); (iii) a 4 kilobases central region (nt 3500–7500). The best fitting substitution models for the full-length genome and various partial sequences alignments were investigated using MEGA X [44] and ModelFinder implemented in IQ-tree (version 2.3.6 [45];). The substitution model suggested by this latter program was finally applied to the test. Maximum likelihood (ML) phylogenies were then inferred using the IQ-Tree software. Bootstrap values were calculated using 1,000 replicates. Trees were finally visualized in the Interactive Tree of Life (iTOL) [46]. Recombination analysis ofgenomic sequences To reconstruct the evolutionary history of PPV, and to understand the possible origin of the genomic fragments and the role played by recombination in such evolution among the strains, a search for potential recombination breakpoints and the identification of the putative parents was performed using the RDP4 program (version 4.10.1, default program settings) [47] and the above described full-length genomes multiple sequence alignment. Only Table 1 Percentages of identity with PPV-An (isolate AL-11pl) in different genome regions. Values are derived from pairwise identity matrices obtained from a MAFFT multiple alignment of full-length genomes of 539 PPV isolates. Metadata (year, host and country of isolation) for each reported isolate are provided Region (nucleotide position) MAFFT alignment Identity (%) Strain Metadata of strains/sequence 1–1584 max 81.7 T ON745776|P13_ANK|Turkey|Prunus_cerasus|2017 min 73.2 CV MF447179|Tat_2|Russia|Prunus_cerasus|2015 15852758 max 96.3 T MF346246|AnKuAp8|Turkey|Prunus_armeniaca|2014 min 76.8 TAT OK562686|TAT_85|Russia|Prunus_cerasus|2018 2759–7532 max 96.4 M LC494682|Y3|Japan|Prunus_mume|2016 min 77.6 W HQ670746|LV_141pl|Latvia|Prunus_domestica|2010 Page 4 of 16 Palmisanoetal. Virology Journal (2025) 22:319 events detected by at least 4 methods and with corrected p-values < 10–6 were considered as reliable. The identified breakpoints were further manually examined, and BLASTN searches were used to verify the parent/donor strain assignment and their respective sequence homology levels. Ancestral states reconstruction throughBayesian phylogeny The existence of a temporal signal of PPV evolution in the 539 full-length genomic sequences dataset was assessed using the programs TempEst (version 1.5.3; [48]) and TreeTime (version 0.11.4; [49]). For Bayesian phylogenetic analysis, two datasets were used. The first corresponds to the central genomic region (nt 3500–7500) aligned for the 539 isolates representing all strains in our original dataset. This region was shown to be free of recombination events for the isolates used and will be hereafter referred to as the 4K region. The second one corresponds to the 5’ genome fragment (nt 36–1406) from the same PPV full-length genome alignment. The package bModelTest [50] was used to assess the site model and associated substitution model. Marginal likelihood estimation (MLE) through Path sampling/Stepping-stone sampling (PS/SS) analysis [51] was run on BEAST to compare the fit of the relaxed clock model with a strict clock model in the above-mentioned sequence datasets. Based on the available metadata (see paragraph above), for each PPV sequence available discrete location (country of origin) and host of origin states were assigned. A symmetric substitution model was applied for each discrete trait (host and country) and social networks were inferred by Bayesian stochastic search variable selection (BSSVS). Ancestral states were reconstructed for all the considered partitions. The Markov Chain Monte Carlo (MCMC) method in the BEAST package (v10.5.0; [52]), along with BEAGLE, was inferred with the best fit substitution model (GTR + G + I) suggested by bModelTest for the two independent datasets described above. In the case of the 4K region, 13 runs for a total of 1300 million MCMC steps were carried out, while in the case of the 5’ region 3 runs including 800 million steps, were merged using LogCombiner (v10.5.0). The Tracer software (v 1.7.2) was used to confirm that all estimated parameters yielded effective sampling sizes (ESS) greater than 200 and that 10% of the total chain length had been burned-in to reduce the influence of the initial value. The final Bayesian maximum clade credibility (MCC) tree was generated by TreeAnnotator (v10.5.0) and visualized in Figtree (v 1.4.4). A graphical elaboration of the final MCC trees was obtained in RStudio using the packages ggtree and treeio [53, 54]. The resulting Bayesian diffusion states in space and time were calculated with the SPREAD application (version 1.0.7; [55]). The ancestral trait for the host was reconstructed at the root and the oldest nodes for the same MCC trees. To further confirm the molecular clock signals obtained through the previously described BEAST analysis, a synchronous BEAST elaboration for the two datasets (simulating the sampling dates as done all at the same time, i.e. 2017.5) was run with the same parameters described above. Then, date-randomization test was performed with 10 replicates whose sampling date were randomly permutated in the TipDating Beast package. We assumed the presence of temporal structure when the 95% credible interval of the rate estimate from the original dataset did not overlap with the 95% credible interval of any of the rate estimates from the date-randomized replicates. Finally, to understand if the prevalence of PPV-D sequences present in the datasets could have biased the Bayesian calculation of substitution rates and TRMCAs, an additional analysis was performed with the very same parameters described above for the 5’and 4k regions, selecting only PPV-D non recombinant and PPV-Rec strains, for a total of 291 sequences. Results Relevant molecular andserological features ofPPV‑An The AL-11pl isolate (GenBank HF674399), which typifies the PPV-An strain and will be hereafter indicated as An, was identified in a domestic plum tree in Eastern Albania. Serologically, PPV-An reacted to the PPV-M-specific monoclonal antibody (MAb) AL (not shown). Remarkably, it also tested positive with the PPV-D-specific MAb 4DG5, an unusual behaviour previously reported for some PPV-T isolates [56]. The complete genome of the PPV-An isolate is 9,786 nucleotides long, excluding the 3’ terminal polyA tail, and has a GC content of 43.8%. The genomic organization is typical of members of genus Potyvirus, and identical to that of other PPV isolates. A start codon (AUG) is present at positions 147–149, and an amber stop codon at positions 9567–9569, resulting in a single open reading frame (ORF) of 9420 nt/3140 amino acids. In addition, the PIPO ORF [57] putatively encoding a 12kDa protein was identified in the P3 coding region as a + 2 frameshift sequence starting at nucleotide position 2906. Sequence alignments show a conservation of the nine polyprotein cleavage sites as compared to PPV-M isolates, with the exception of mutations observed in the NIa-VPg/NIa-Pro cleavage site (EEVGHE/S in PPV-An and DEVDHE/S in PPV-M isolates) and NIa-Pro/NIb site (EFVHNQ/S vs. EFVYNQ/S). All conserved motifs typical of potyviruses were also identified at their expected locations, including the KITC [58], PTK [59] and DAG motifs associated with aphid transmission. Page 5 of 16 Palmisanoetal. Virology Journal (2025) 22:319 Percentages of nucleotide pairwise identity calculated between PPV-An and all other PPV strains – expressed as their maximal and minimal values in various genomic regions selected from the full-length genome alignment are given in Table 1. The complete genome sequence comparisons indicate that the closest strain to PPV-An is PPV-T, with an overall nucleotide identity of 93.5%. Nevertheless, while PPV-An is closely related to isolates of the T and M strains at the whole genome level, it shows a much lower identity level (74–77%) with these strains for the 5’ non-coding region (5’NCR) and the P1 gene (81.7%, from start codon up to nt 1584). This dissimilarity pattern extends to the HC-Pro gene (nt 1585–2758) in the case of the M strain (82% nt sequence identity) but it is no longer observed in the case of the T strain (96.3% identity; Table1). Nucleotide identity levels with both T and M strains for all the 3’ downstream parts of the genome are higher than 95%. This unusual identity pattern observed for the various regions of the PPV-An, strongly points to its possible implication in recombination events. Analysis ofrecombination events involving PPV‑An asparental ancestor The multiple alignment of 539 full-length genomic sequences of PPV isolates was analysed for recombination events using RDP4. A number of putative recombination events were identified by the program, but most of them involved just one or a few isolates and were detected by only three or less of the methods, with statistically non-significant p-values. However, two recombination events which involve strains sharing a large genomic portion with PPV-An were identified by a strong signature and are discussed in detail here (Additional Table 2). A first statistically significant recombination event (i.e. lowest p-value 1.77 × 10–73) identifies the fragment at nt positions 1561–2735 in the alignment (with 99% confidence intervals [CI 99%] for the breakpoints at 903–1598 and 2678–2765) and affects 130 PPV-M isolates (Fig.1A; event #4, Additional Table2). This event involves a putative major parent belonging to the PPV-T strain (GenBank MF346274), and a minor parent belonging to the PPV-D strain (GenBank KR006730). Given the RDP4 output, strain T isolates are interpreted by the program as non-recombinant major donors of a backbone which received the insertion of the heterologous fragment from a strain D minor parent, thus leading to the current PPV-M isolates. The other relevant recombination event identified, with a breakpoint at nt 2691 of the PPV-An sequence (CI 99%: 2665–2816), involves a PPV-M isolate as the major parent providing the entire 3’ genome part (from nt 2691 up to 3’ end) and an unknown minor parent providing the 5’genome portion (Fig. 1A; event #3, Additional Table2). This recombination event is consistently supported by six methods (i.e. RDP, GENECONV, BootScan, MaxChi, Chimaera, SiScan) with the lowest p-value of 2.23 × 10–115. The breakpoint of this predicted recombination event is in very close proximity to the one (at position 2814) hypothesized by Glasa & Candresse [17] for a recombination event between PPV-D and a putative ancestral isolate leading to the emergence of the PPV-M strain. However, when the 3’ end genomic region of PPVAn (nt 3500–9700) was used as a query for a BLASTN interrogation of GenBank, the closest isolate to PPV-An in that genomic region was PPV-T KrPnPl345 (GenBank MF346272), sharing 96.4% of nt identity. For the 5’ terminal region (1–1561), an unknown minor parent was postulated in the RDP4 output. A BLASTN analysis indicated that the isolate most similar to PPV-An in that region is P93 ANK (GenBank ON745778), another PPV-T isolate, but its identity with PPV-An is only of 85,3%. The #3 and #4 (Additional Table2) events identified here are respectively similar to events X2 and X3 identified by [18], although a reasoning based on phylogenetic analysis led them to conclude that the T strain was derived from a PPV-M strain parent, and not the reverse as predicted by our RDP analysis. However, the potential scenario described by these authors has two weaknesses. First, considering either PPV-M or PPV-T as a non-recombined strain parent to PPV-An fails to provide an explanation for the low divergence in the P1 protein between these two strains and PPV-D strain, despite the fact that P1 is very generally the most divergent potyviral protein. Second, this scenario is not parsimonious since it needs to postulate the existence of a further unknown minor parent providing the 5’terminal portion of PPV-An in the recombination event #3. In contrast, the scenario proposed by [17] and further discussed by [7], which identifies PPV-An and PPV-D as ancestral non recombinant parents and PPV-M and PPV-T as the derived recombinant progeny, is more parsimonious (it does not postulate the existence of another unknown parent) and readily provides an explanation for the high P1 homology observed between PPV-D and M and T strains (Fig.1B, Hypothesis 2). Maximum likelihood phylogenetic analyses onfull length sequences andpartial genomic fragments The maximum likelihood (ML) phylogenetic analysis performed on the full-length genome dataset showed a striking position of PPV-An as a long branch linked to the PPV-T clade (Additional Fig.1). This analysis shows very clearly a separation between the clades of isolates belonging to the European and Mediterranean strains Page 6 of 16 Palmisanoetal. Virology Journal (2025) 22:319 (PPV-M, -T, -An, -D and -EA) from those corresponding to PPV-W and the cherry-adapted strains. As an alternative to the exclusion of the large number of recombinant isolates (PPV-M, T and Rec strains) from the dataset to perform an evaluation of the phylogenetic signals (therefore losing useful information), we decided to use only genome portions known to be free of recombination breakpoints. The best substitution model obtained with ModelFinder for the full-length genomes and the three partial genomic datasets was GTR + F + I + G4, and it was thoroughly applied for the further phylogenetic calculations. A phylogenetic analysis was performed using a 5’genomic fragment (nucleotide positions 36 to 1400), which ends before the border of the first recombination breakpoint suggested by [17] and essentially overlaps with the 5’border of the recombination event #4 discussed above. The resulting ML cladogram showed A B PPV-D PPV-T PPV-M VV VVV PPV-Unknow PPV-M PPV-An VV VVV P1 HC-Pro P3 CI 6K16K2 VPg Pro NIb CP (A) n VPg Hypotesis 1, PPV-An is a recombinant C PPV-An PPV-D PPV-T VV VVV PPV-An PPV-D PPV-M VV VVV Hypotesis 2, PPV-An is NOT a recombinant RDP event 3 RDP event 4 Origin of PPV-T Origin of PPV-M Fig. 1 Schematic representation of the RDP4 output for some recombination events and hypothesis of recombination involving PPV-An strain. A Schematic representation of PPV genomic organization, where the open box represents the translated ORFs and the functional polyprotein fragments are named. B In the hypothesis 1 frame, the representation only pictures two independent recombination events as identified by RDP4 in terms of parental PPV strains and recombined regions and is not related to any phylogenetic assumption. Here, PPV-An is considered as recombinant. Event #3) depicts the PPV-An isolate, AL-11pl, as originated by recombination between an unknown minor parent providing the 5’-end fragment (nt positions 1–2691) and a major parent belonging to M strain and providing the rest of the genome. Event #4) depicts the origin of the M strain isolates through the insertion of a fragment (nt positions 1561–2735) from a D isolate in the backbone of a T isolate (see Additional Table 2). C In the hypothesis 2 frame, the scenario in which An is considered non-recombinant and a parent of M and T strains, when recombining with PPV-D, is illustrated. The different color codes of the strains derive from the different roles as parent or recombinant they cover in the different events Page 7 of 16 Palmisanoetal. Virology Journal (2025) 22:319 a tight clustering of all PPV-D, PPV-M, and PPV-T isolates (61.2% bootstrap), which were separated from a small clade containing the PPV-EA and PPV-An isolates (59.2% bootstrap) (Fig. 2A). In the ML tree reconstructed using the region between the first and second identified recombination breakpoints (nt positions 1500–2600, according to [17] and which spans the C-terminal half of HC-Pro and the beginning of the P3 gene (Fig.2B), PPV-An now forms a loose cluster (41.4% bootstrap) with the PPV-T isolates. Lastly when a large, non-recombined internal genome fragment (nt 3500–7500, referred as the 4K region) was used for ML phylogeny, PPV is split into five phylogenetic groups (Fig.2C) with PPV-An clearly flanking, as a single long branch, the PPV-M and PPV-T clade (65.4% bootstrap). In this tree, a single isolate (SK-111pl, belonging to M strain, GenBank HF585099) has an anomalous position as basal tip at the node which originated all the Anderived strains. This position could most likely be due to the peculiar recombination history of this isolate, which concerns, according the RDP4 analysis, a recombination event involving a small region (from nt 3264 to 4010; Additional Table2, event #5) at the beginning of the 4K region. Overall, the incongruent topologies in the different trees when it comes to the PPV-D, Rec, M, T isolates versus the An one, support the recombination analysis and confirms that recombination events have contributed in a major way to PPV evolutionary history of the Euro-Mediterranean (Euro-Med) PPV strains. Figure2D resumes the current view of recombination events along the PPV genome which characterize the single strains. 0.1 Region 36-1400 0.1 0.1 A Region 1500-2600 B Region 3500-7500 C WCherry EA An D, M, T, Rec W Cherry EA D, M, Rec T D, Rec M, T W Cherr y An An D P1 HC-Pro P3 CI 6K1 6K2 VPg Pro NIb CP (A) n VPg PPV-Cherry PPV-W PPV-EA PPV-M PPV-An PPV-T PPV-Rec PPV-D Fig. 2 Maximum likelihood phylogenetic trees reconstructed using different genome regions of 539 PPV isolates. Strains or phylogroups are indicated as PPV-T (T), PPV-An (An), PPV-M (M), PPV-EA (EA), PPV-W (W), PPV-C (C), PPV-Y (Y), PPV-CV (CV), PPV-SC (SC), PPV (TAT), PPV-CR (CR), PPV-D (D) and PPV-Rec (Rec). Colored spots (when present) correspond to the Prunus species from which any isolate was obtained (see the legend of Additional Fig. 1). All the strains belonging to the same clade and sharing similar sequence identity in the genomic portion under evaluation, are grouped by the same color background. The distance bar is also shown under each tree. A Tree for the 5’ region (nt position 36–1400). B Tree for an internal fragment corresponding to nt positions 1500–2600. C Tree for the 4 K genomic fragment (nt position 3500–7500). D Color-coded schematic representation of the genome of the main PPV strains. Different colors represent the tentative strain origin of the putative recombined fragments. The genome organization of PPV is shown at the top, where the open box represents the translated ORFs and the functional polyprotein fragments are named. The bootstrap values at the main nodes are not reported in the pictures to improve visualization, since most of them are > = 60 Page 8 of 16 Palmisanoetal. Virology Journal (2025) 22:319 PPV‑An‑derived strains diversification andBayesian dating intheglobal PPV evolution andspread The presence of a temporal signal was first investigated on the global PPV phylogeny to determine whether analyzing the evolution of existing lineages could provide a reliable estimate for the time to a common ancestor between An and the other PPV lineages. The TempEst software [48] run on the full-length PPV genomes, recovered a TMRCA intercept at 3604years ago, while the root-to-tip regression using TreeTime put the TMRCA inferred from the same sequence alignment at 4499years ago. The quantitative metrics obtained with TempTest are reported in Additional Table3. This preliminary dating analysis for a retrieval of a temporal signal could incur a bias because of the inclusion of recombinant sequences in the dataset. In fact, according to our evidence of PPV recombination history (Fig.2D) only PPV-D, PPV-An, PPV-EA, most of PPV-W and the cherry-adapted strains are considered non-recombinants, so that a large proportion of the isolates included in the dataset are recombinants. To overcome this problem, Bayesian datation and phylogeographic diffusion analysis were separately performed on two non-recombinant genome regions described above: the 4K central region and the shorter 5’ fragment. First, to evaluate the temporal structure in the dataset, the date-randomization test was performed. The 95% credible intervals of inferred substitution rate for both datasets did not overlap with any of the date-randomized replicates, indicating the presence of temporal signal in our dataset (Additional Fig.3, A and B). Moreover, the presence (strict clock model) and absence (uncorrelated relaxed clock model) of a molecular clock signal were also tested using Bayesian statistics. MLE calculated for the two hypothesized models through PS/SS analysis yielded Bayes factor (BF) values indicating a decisive support in favor of the presence of a molecular clock for both datasets (i.e., 4K BF = 6405.9; 5’ region BF = 6244.3). The bModelTest search for the best substitution model to be applied in the Bayesian trees calculation suggested the GTR + G + I one. The average rate of nucleotide substitution estimated for the central 4K region of all available PPV isolates, calculated after BEAST elaboration, was 6.91 × 10–5 subst./ site/year (with the 95% highest posterior density (HPD) interval of 4.67 × 10–5—9.09 × 10–5). In the Bayesian MCC tree calculated for the 4K region (Fig.3 and Additional Table4 A) the position of the TMRCA was determined at 4546years ago (95% HPD 3141–6099) when the group of Euro-Med strains separated from the branch leading to PPV-W and the cherry-adapted strains. The separation of PPV-EA from the Euro-Med group could have happened 3655years ago (95% HPD 2531–4944) while for its part, the separation of the clade of cherry-adapted strains from the W strain is dated back at 3087years ago (95% HPD 2111–4136). The initial diversification of the Euro-Med group is predicted to have started 1733years ago (95% HPD 1200–2350). The separation of the M and T strains from PPV-An is suggested to have occurred much more recently, between 289 and 279 years ago, respectively (95% HPD 200–394 and 185–400). Using this dataset, the tentative dating of the separation of the D and Rec strains at 251years ago (95% HPD 178–338) could also be hypothesized. Using the 5’ fragment encoding the P1 gene (fragment 36–1406 nt), the BEAST elaboration retrieved a subst./s/y rate of 2.55 × 10–4 for the global alignment of 539 isolates (95% HPD 2.025 × 10–4—3.084 × 10–4), which is about 3.6 times faster than that obtained for the 4K regions and is likely accounted for the lower conservation of the P1. In Fig.4 (and Additional Table4B), the root of the tree dated at 1211years ago (95% HPD 962–1482) and the main diversification of the cherry-adapted strains at 707years ago (95% HPD 545–872). Regardless the difference in substitution rates, the Bayesian analysis on both datasets indicated that the separation of the PPV-An strain is anterior to that of the D strain, and thus to all the subsequent node bifurcation events in the Euro-Med group. Furthermore, the Bayesian elaborations done on both the same regions, but only on a dataset restricted to D isolates (including Rec, for a total of 291 sequences), obtained very similar results, as substitution rates (5’ region: 2.029E-4; 4 K region: 8.688E-5) and TMRCA datation (5’ region: 1409.404years before present; 4K region: 3541.299 years before present), thus indicating that the large predominance of the D isolates in the dataset could substantially have influenced the complete dataset output values. Parallel to these dating inferences, the two datasets were used in synchronous BEAST elaborations assuming (See figure on next page.) Fig. 3 Time-scaled maximum-clade-credibility tree inferred from the 4 K genome region of 539 selected PPV isolates. Isolates in the tree are collapsed by strain and branch lengths are scaled by distance from TMRCA according to time (years from present), as shown by the reference bar at bottom. At the major nodes, a pie chart shows the host set probability with Prunus host species colored according to the legend inset. The pink bar at each node represents the 95% confidence interval (HPD 95%). Specific values for each node date and percentage of host set preference are reported in Additional Table 4A Page 9 of 16 Palmisanoetal. Virology Journal (2025) 22:319 Cherry T EA M Rec D W An P. cerasus P. domestica P. dulcis P. fruticosa P. japonica P. mume P. persica P. persica var. nucifera P. armeniaca P. cerasifera P. tomentosa Prunus Fig. 3 (See legend on previous page.) Page 16 of 16 Palmisanoetal. Virology Journal (2025) 22:319 43. Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 2013;30:772–80. 44. Kumar S, Stecher G, Li M, Knyaz C, Tamura K. MEGA X: molecular evolutionary genetics analysis across computing platforms. Mol Biol Evol. 2018;35:1547–9. 45. Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, Von Haeseler A, et al. IQ-tree 2: new models and efficient methods for phylogenetic inference in the genomic era. Mol Biol Evol. 2020;37:1530–4. 46. Letunic I, Bork P. Interactive tree of life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 2021;49:W293–6. 47. Martin DP, Murrell B, Golden M, Khoosal A, Muhire B. RDP4: Detection and analysis of recombination patterns in virus genomes. Virus Evol. 2015;1:vev003. 48. Rambaut A, Lam TT, Max Carvalho L, Pybus OG. Exploring the temporal structure of heterochronous sequences using TempEst (formerly Path-OGen). Virus Evol. 2016;2: vew007. 49. Sagulenko P, Puller V, Neher RA. TreeTime: Maximum-likelihood phylodynamic analysis. Virus Evol. 2018;4: vex042. 50. Bouckaert RR, Drummond AJ. bModelTest: Bayesian phylogenetic site model averaging and model comparison. BMC Evol Biol. 2017;17:42. 51. Baele G, Lemey P, Bedford T, Rambaut A, Suchard MA, Alekseyenko AV. Improving the accuracy of demographic and molecular clock model comparison while accommodating phylogenetic uncertainty. Mol Biol Evol. 2012;29:2157–67. 52. Suchard MA, Lemey P, Baele G, Ayres DL, Drummond AJ, Rambaut A. Bayesian phylogenetic and phylodynamic data integration using BEAST 1.10. Virus Evol. 2018;4: vey016. 53. Wang L-G, Lam TT-Y, Xu S, Dai Z, Zhou L, Feng T, et al. Treeio: an R package for phylogenetic tree input and output with richly annotated and associated data. Mol Biol Evol. 2020;37:599–603. 54. Yu G, Smith DK, Zhu H, Guan Y, Lam TT. ggtree : an r package for visualization and annotation of phylogenetic trees with their covariates and other associated data. McInerny G, (ed). Methods Ecol Evol. 2017;8:28–36. 55. Bielejec F, Rambaut A, Suchard MA, Lemey P. SPREAD: spatial phylogenetic reconstruction of evolutionary dynamics. Bioinformatics. 2011;27:2910–2. 56. Candresse T, Cambra M, Dallot S, Lanneau M, Asensio M, Gorris MT, et al. Comparison of monoclonal antibodies and polymerase chain reaction assays for the typing of isolates belonging to the D and M serotypes of plum pox potyvirus. Phytopathology®. 1998;88:198–204. 57. Chung BY-W, Miller WA, Atkins JF, Firth AE. An overlapping essential gene in the Potyviridae. Proc Natl Acad Sci USA. 2008;105:5897–902. 58. Urcuqui-Inchima S, Haenni A-L, Bernardi F. Potyvirus proteins: a wealth of functions. Virus Res. 2001;74:157–75. 59. Wang Y, Gal-On A, Huet H, Kadoury D, Raccah B, Peng YH. Mutations in the HC-Pro gene of zucchini yellow mosaic potyvirus: effects on aphid transmission and binding to purified virions. J Gen Virol. 1998;79:897–904. 60. Gürcan K, Teber S, Akbulut M, Çağlayan K. Genetic diversity and a long evolutionary history of plum pox virus strain rec in Turkey. Eur J Plant Pathol. 2021;161:453–61. 61. Cambra M, Asensio M, Gorris MT, Pérez E, Camarasa E, Garciá JA, et al. Detection of plum pox potyvirus using monoclonal antibodies to structural and non-structural proteins1. EPPO Bull. 1994;24:569–77. 62. Boscia D, Zeramdini H, Cambra M, Potere O, Gorris MT, Myrta A, et al. Production and characterization of a monoclonal antibody specific to the M serotype of plum pox potyvirus. Eur J Plant Pathol. 1997;103:477–80. 63. Shan H, Pasin F, Valli A, Castillo C, Rajulu C, Carbonell A, et al. The Potyviridae P1a leader protease contributes to host range specificity. Virology. 2015;476:264–70. 64. Nigam D, LaTourrette K, Souza PF, Garcia-Ruiz H. Genome-wide variation in potyviruses. Front Plant Sci. 2019;10:1439. 65. Sanjuán R, Agudelo-Romero P, Elena SF. Upper-limit mutation rate estimation for a plant RNA virus. Biol Lett. 2009;5:394–6. 66. Jenkins GM, Rambaut A, Pybus OG, Holmes EC. Rates of molecular evolution in RNA viruses: a quantitative phylogenetic analysis. J Mol Evol. 2002;54:156–65. 67. Gibbs AJ, Ohshima K, Phillips MJ, Gibbs MJ. The prehistory of potyviruses: their initial radiation was during the dawn of agriculture. PLoS ONE. 2008;3:e2523. 68. Mohammadi M, Hosseini A, Nasrollanejad S. In silico identification of two novel viruses on Iranian pistachio. Iran J Plant Pathol. 2021;57:81–5. 69. Glasa M, Palkovics L, Komínek P, Labonne G, Pittnerová S, Kúdela O, et al. Geographically and temporally distant natural recombinant isolates of Plum pox virus (PPV) are genetically very similar and form a unique PPV subgroup. J Gen Virol. 2004;85:2671–81. 70. Sheveleva A, Glasa M, Kudryavtseva A, Ivanov P, Chirkov S. Genetic diversity, host range and transmissibility of CR isolates of Plum pox virus. J Gen Plant Pathol. 2019;85:39–43. 71. Dallot S, Karychev R, Dolgikh S, Thébaud G, Jacquot E, Decroocq V. First report of Plum pox virus strain W in Kazakhstan, on Prunus domestica. Plant Dis. 2019;103:2702. 72. Atanasoff D. Mosaic of stone fruits. 1935 . Available from: https:// www. cabid igita llibr ary. org/ doi/ full/ 10. 5555/ 19350 501302. Cited 2025 Feb 27. 73. Moury B, Desbiez C. Host range evolution of potyviruses: a global phylogenetic analysis. Viruses. 2020;12:111. 74. Tsarmpopoulos I, Marais A, Faure C, Theil S, Candresse T. A new potyvirus from hedge mustard (Sisymbrium officinale (L.) Scop.) sheds light on the evolutionary history of turnip mosaic virus. Arch Virol. 2023;168: 14. Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.