scieee AI-readable full text Open interactive document viewer

Global Parkinson's Genetics Program Data Release 11

Hampton Leonard; Mike Nalls; Dan Vitale; Mathew Koretsky; Kristin Levine; Mary Makarious; Zih-Hua Fang; Jones, Lietsel; Solle, J

Abstract

In December 2025, GP2 announced the 11th data release on the Terra and the Verily® Workbench platforms in collaboration with AMP® PD. This release includes 20,842 additional genotyped participants, 17,153 additional WGS participants, and 4,232 additional clinical exomes. The genotype array (NBA) data, including locally-restricted samples, now consists of a total of 103,786 genotyped participants (46,327 PD cases, 28,857 Controls, and 28,602 ‘Other’ phenotypes). The whole genome sequencing (WGS) data now consists of a total of 38,226 sequenced participants (18,219 PD cases, 9,172 Controls, and 10,835 ‘Other’ phenotypes). The clinical exome data now consists of 14,648 samples with PD. Of the 122,317 unique samples with genetic data (NBA, WGS, or clinical exome), 32,897 individuals also have additional extended clinical information. Please see the accompanying blog for further description of this release. To obtain data access, please see https://amp-pd.org/researchers/data-use-agreement. For any publications using data from this release, please reference the DOI number and the following statement: "Data (DOI 10.5281/zenodo.17753486, release 11) and/or code used in the preparation of this article were obtained from Global Parkinson’s Genetics Program (GP2). GP2 is funded by the Aligning Science Across Parkinson’s (ASAP) initiative and implemented by The Michael J. Fox Foundation for Parkinson’s Research (https://gp2.org). For a complete list of GP2 members see https://gp2.org."

Full text

The Components of GP2’s 11th Data Release Tags Research Operations; Research Collaboration; Complex Disease Genetics; Release Authors Hampton Leonard DataTecnica/National Institutes of Health | USA Hampton has a background in data science and machine learning, which she applies to large multi-omic datasets in the neurodegenerative disease space. She is passionate about investigating differences on both clinical and omic levels and how these differences can affect clinical trial outcomes. Mike Nalls DataTecnica/National Institutes of Health | USA Mike founded Data Tecnica in early 2017 after over a decade of experience in large dataset analytics and methods research in healthcare and other scientific fields. Mike has 400+ peer-reviewed publications in the field of applied statistics in large datasets, brain diseases, and genomics. He is a strong advocate of open science, collaboration, and transparency in science. Dan Vitale DataTecnica/National Institutes of Health | USA Dan is a data science consultant for Data Tecnica, consulting primarily for the Laboratory of Neurogenetics and CARD at the National Institute on Aging of the National Institutes of Health. His work is focused on open science, automation, development of genetic analytic pipelines and software, and machine learning. Mathew Koretsky DataTecnica/National Institutes of Health | USA Mat is a data science consultant for Data Tecnica, consulting primarily for CARD at the National Institute on Aging of the National Institutes of Health. He is passionate about pipeline development and meaningful applications of computer science in the biomedical research space. Kristin Levine DataTecnica/National Institutes of Health | USA Kristin works with the Data Tecnica and National Institute on Aging (NIA) teams on data and code sharing plus real-world data analysis of biobanks and healthcare systems. She is also an accomplished writer, now applying her communication skills to scientific domains. Mary B Makarious DataTecnica/National Institutes of Health | USA Mary is a biomedical data scientist committed to open science principles and enhancing diversity in genomic studies. With her background in machine learning, data science, and genetics, she analyzes large-scale multi-omics datasets to develop open, reproducible pipelines and user-friendly notebooks and tools. Her efforts aim to empower others to effectively explore and interpret their own data and to foster a more inclusive and collaborative scientific community. Lietsel Jones DataTecnica/National Institutes of Health | USA Lietsel is an analyst with Data Tecnica with a keen interest in the intersection between epidemiology and genetics. She is also a clinical data manager with GP2 working to collect and harmonize large clinical datasets from worldwide contributors. Zih-Hua Fang German Center for Neurodegenerative Diseases | Germany Zih-Hua leads the whole-genome sequencing data analysis efforts in GP2 and contributes to GP2’s work on monogenic and familial Parkinson’s disease. J Solle Michael J. Fox Foundation for Parkinson’s Research | USA J is the implementation Program Lead for GP2, co-lead for the Operations & Compliance Working Group, and a member of the Operations Committee. On behalf of the GP2 Operations & Compliance, Complex Disease Data Analysis, Monogenic Data Analysis, Clinical Integration, and Data and Code Dissemination Working Groups. Blog Post Overview In December 2025, GP2 announced the 11th data release on the Terra and the Verily® Workbench platforms in collaboration with AMP® PD. This release includes 20,842 additional genotyped participants, 17,153 additional WGS participants, and 4,232 additional clinical exomes. ● The genotype array (NBA) data, including locally-restricted samples, now consists of a total of 103,786 genotyped participants (46,327 PD cases, 28,857 Controls, and 28,602 ‘Other’ phenotypes). ● The whole genome sequencing (WGS) data now consists of a total of 38,226 sequenced participants (18,219 PD cases, 9,172 Controls, and 10,835 ‘Other’ phenotypes). ● The clinical exome data now consists of 14,648 samples with PD. ● Of the 122,317 unique samples with genetic data (NBA, WGS, or clinical exome), 32,897 individuals also have additional extended clinical information. What’s New In This Release? Expanding Genomic Data This release introduces a substantial expansion in the number of participants with available genetic data. We have added: ● 20,842 new participants with genotype array (NBA) data ● 17,153 new participants with whole genome sequencing (WGS) data ● 5915 new participants with extended clinical data ● A family file (and corresponding data dictionary) which reports pairwise kinship estimates between individuals within families. It includes both inferred relationships (with kinship coefficients) and reported relationships. Joint-calling Now Include AMP® PD cohorts ● The jointly-called WGS variant sets now include samples from the following seven AMP® PD cohorts: BioFIND, HBS, PDBP, PPMI, LCC, STEADY-PD3 and SURE-PD3. ○ By processing these samples together with GP2 rather than independently, it minimizes missingness, artifacts, and improves genotype accuracy. ● We have added a column to master key denoting which GP2 samples are also present in the AMP-PD dataset. New Summary Statistics Now Available We’ve made available several additional GWAS summary statistics datasets under Tier 1: ● Blauwendraat et al PD age at onset (AAO) meta-analysis (paper) ● Blauwendraat et al PD autosomal sex-stratified meta-analysis (paper) ● Blauwendraat et al PD GBA1-modifier meta-analysis (paper) Clinical Data This release contains clinical data for a total of 122,317 individuals who have genetic and core clinical data available. Of these, 32,897 have deep clinical phenotyping data available. This information consists of: ● Age at diagnosis and onset ● Primary, current, and latest diagnoses ● Cognitive exams such as the Mini-Mental State Examination (MMSE) and the Montreal Cognitive Assessment (MoCA) ● Movement Disorder Society-Sponsored Revision of the Unified Parkinson's Disease Rating Scale (MDS-UPDRS) ● Detailed “other” phenotypes, such as Lewy body Dementia (LBD) Individual-Level Data We now capture the data from a total of 148 cohorts. Please refer to the GP2 Cohort Dashboard for more information on the cohorts that have been shared. Genetically-determined ancestry of array genotyped GP2 participants are broken into 11 ancestry groups; the tables below provide details of the genetically-determined ancestry of participants in this release that have passed quality control for array data and whole genome sequencing data. These numbers reflect samples from previous releases, reclustered using the updated cluster file and subjected to quality control, as well as newly genotyped samples exclusive to this release. The final table provides information about the genetically-determined ancestry of selected other, non-PD phenotypes. Array Genotyped Data - GP2 Release 11 Ancestry Total PD Control Other African 7,246 2,549 4,367 330 African Admixed 1,382 478 816 88 Ashkenazi Jewish 3,989 1,810 529 1,650 Latino and Indigenous people of the Americas 3,816 2,122 1,487 207 East Asian 7,965 3,083 2,787 2,095 European 71,402 32,952 15,388 23,062 South Asian 1,290 412 405 473 Central Asian 2,801 1,053 1,371 377 Middle Eastern 2,473 973 1,294 206 Finnish 154 110 16 28 Complex Admixture 1,268 785 397 86 Total 103,786 46,327 28,857 28,602 Whole Genome Sequenced Data - GP2 Release 11 Ancestry Total PD Control Other African 3,706 1,349 2,144 213 African Admixed 424 235 160 29 Ashkenazi Jewish 2,148 847 268 1,033 Latino and Indigenous people of the Americas 1,275 824 314 137 East Asian 4,472 1,462 1,058 1,952 European 20,844 11,568 2,784 6,492 South Asian 986 240 344 402 Central Asian 1937 581 986 370 Middle Eastern 1983 852 984 147 Finnish 48 32 6 10 Complex Admixture 402 229 124 50 Total 38,226 18,219 9,172 10,835 “Other” Phenotype - GP2 Release 11 Ancestry Prodromal NBA/ WGS PSP NBA/ WGS AD NBA/ WGS DLB NBA/ WGS MSA NBA/ WGS CBD/CBS NBA/ WGS LBD NBA/ WGS FTD NBA/ WGS African 16/7 12/10 0/0 10/8 14/12 1/0 0/0 1/0 African Admixed 27/7 5/3 1/0 0/8 2/0 1/0 0/0 0/0 Ashkenazi Jewish 311/68 26/14 9/3 13/6 10/6 5/3 2/0 2/2 Latino and Indigenous people of the Americas 38/11 6/0 5/1 2/0 2/0 1/0 0/0 0/0 East Asian 89/4 311/226 5/5 18/0 17/187 13/44 1/1 0/0 European 4672/ 813 1483/ 996 512/ 261 459/ 357 482/400 197/166 191/ 127 75/68 South Asian 3/2 43/34 1/0 5/1 5/8 9/9 0/0 2/2 Central Asian 4/4 16/12 94/89 4/3 5/2 4/2 0/0 0/0 Middle Eastern 17/2 9/5 2/2 1/0 5/5 2/1 0/0 1/1 Finnish 9/0 2/1 2/1 0/0 1/1 0/0 0/0 1/1 Complex Admixture 14/2 7/5 5/4 3/1 1/0 1/0 0/0 1/1 Total 5,200/ 921 1,920/ 1306 636/ 366 515/ 376 544/ 621 234/ 225 194/ 128 83/75 Snapshot of Clinical Data - GP2 Release 11 Clinical Data N, Unique IDs N, IDs with Follow-up Age at Sample Collection 99,247 - Age at Onset 54,489 - Age at Diagnosis 43,283 - Basic Family History 13,135 3,287 Demographics 27,569 5,377 Hoehn & Yahr Stage 11,729 6,217 CISI-PD 594 - UPDRS Part 1 Score 2,359 1,048 UPDRS Part 2 Score 2,338 1,040 UPDRS Part 3 Score 3,492 1,075 UPDRS Part 4 Score 1,739 1,081 MDS UPDRS Part 1 Score 5,292 3,467 MDS UPDRS Part 2 Score 5,367 3,520 MDS UPDRS Part 3 Score 8,045 3,632 MDS UPDRS Part 4 Score 2,649 1,609 MOCA 8,391 2,881 MMSE 2,621 93 SCOPA-Cog 619 272 RBD Score 4,085 3,324 SEADL 5,873 4,250 Head Trauma 5,721 3,781 Vitals 9,357 4,681 Smell 8,047 1,038 Lifestyle/Environmental 15,434 2,161 SAA 128 - Data Access Locality-restricted GDPR samples via the Verily Viewpoint Workbench We are continuing to pilot granting access to locally-restricted samples, otherwise known as samples governed by the General Data Protection Regulation (GDPR) policy, through our collaboration with the Verily Viewpoint Workbench. To gain access to the full release on VWB you must: 1. Have approved GP2 Tier 2 access 2. Fill out the GDPR-governed sample request form Future data releases will continue to grow the diversity of participants available. You can check out our dashboard to see our progress. For users with tier 2 access already, you can explore the data further on our cohort browser, expanded on in a previous blog post. As always, please refer to the README that accompanies each GP2 release for further details regarding recommendations for quality control, pipelines, data, and analyses!