scieee Open visual document viewer

RSAT variation-tools: An accessible and flexible framework to predict the impact of regulatory variants on transcription factor binding

Santana-Garcia, W.; Rocha-Acevedo, M.; Thieffry, D.; van Helden, J.; Medina-Rivera, A.; Mbouamboua, Y.; Ramirez-Navarro, L.; Thomas-Chollier, M.; Contreras-Moreira, B.

Abstract

Gene regulatory regions contain short and degenerated DNA binding sites recognized by transcription factors (TFBS). When TFBS harbor SNPs, the DNA binding site may be affected, thereby altering the transcriptional regulation of the target genes. Such regulatory SNPs have been implicated as causal variants in Genome-Wide Association Study (GWAS) studies. In this study, we describe improved versions of the programs Variation-tools designed to predict regulatory variants, and present four case studies to illustrate their usage and applications. In brief, Variation-tools facilitate i) obtaining variation information, ii) interconversion of variation file formats, iii) retrieval of sequences surrounding variants, and iv) calculating the change on predicted transcription factor affinity scores between alleles, using motif scanning approaches. Notably, the tools support the analysis of haplotypes. The tools are included within the well-maintained suite Regulatory Sequence Analysis Tools (RSAT, http://rsat.eu), and accessible through a web interface that currently enables analysis of five metazoa and ten plant genomes. Variation-tools can also be used in command-line with any locally-installed Ensembl genome. Users can input personal collections of variants and motifs, providing flexibility in the analysis. Santana-Garcia, W.; Rocha-Acevedo, M.; Ramirez-Navarro, L.; Mbouamboua, Y.; Thieffry, D.; Thomas-Chollier, M.; Contreras-Moreira, B.; van Helden, J.; Medina-Rivera, A.

Full text

RSAT a ia ion- ools: An accessible and lexible amewo k o p edic he impac o egula o y a ian s on ansc ip ion ac o binding Wal e San ana-Ga cia a,b , Ma ia Rocha-Ace edo b , Lucia Rami ez-Na a o b , Y on Mbouamboua c, , Denis Thie y a , Mo gane Thomas-Chollie a , B uno Con e as-Mo ei a d,e , Jacques an Helden ,g, ⇑ , Alejand a Medina-Ri e a b,* a Ins i u de Biologie de l’ENS (IBENS), Dépa emen de biologie, École no male supé ieu e, CNRS, INSERM, Uni e si é PSL, 75005 Pa is, F ance b Labo a o io In e nacional de In es igación sob e el Genoma Humano, Uni e sidad Nacional Au ónoma de México, Campus Ju iquilla, Bl d Ju iquilla 3001, San iago de Que é a o 76230, Mexico c Fonda ion Congolaise pou la Reche che Médicale, B azza ille, People’s Republic o Congo d Es ación Expe imen al de Aula Dei-CSIC, Za agoza, Spain e Fundación ARAID, Za agoza, Spain Aix-Ma seille Uni , INSERM UMR S 1090, Theo y and App oaches o Genome Complexi y (TAGC), F-13288 Ma seille, F ance g CNRS, Ins i u F ançais de Bioin o ma ique, IFB-co e, UMS 3601, E y, F ance a icle in o A icle his o y: Recei ed 27 Ap il 2019 Recei ed in e ised o m 22 Sep embe 2019 Accep ed 25 Sep embe 2019 A ailable online 7 No embe 2019 Keywo ds: Regula o y a ian s T ansc ip ion ac o s Posi ion speci ic sco ing ma ix SNPs Binding mo i s abs ac Gene egula o y egions con ain sho and degene a ed DNA binding si es ecognized by ansc ip ion ac o s (TFBS). When TFBS ha bo SNPs, he DNA binding si e may be a ec ed, he eby al e ing he an- sc ip ional egula ion o he a ge genes. Such egula o y SNPs ha e been implica ed as causal a ian s in Genome-Wide Associa ion S udy (GWAS) s udies. In his s udy, we desc ibe imp o ed e sions o he p og ams Va ia ion- ools designed o p edic egula o y a ian s, and p esen ou case s udies o illus- a e hei usage and applica ions. In b ie , Va ia ion- ools acili a e i) ob aining a ia ion in o ma ion, ii) in e con e sion o a ia ion ile o ma s, iii) e ie al o sequences su ounding a ian s, and i ) calcu- la ing he change on p edic ed ansc ip ion ac o a ini y sco es be ween alleles, using mo i scanning app oaches. No ably, he ools suppo he analysis o haplo ypes. The ools a e included wi hin he well-main ained sui e Regula o y Sequence Analysis Tools (RSAT, h p:// sa .eu), and accessible h ough a web in e ace ha cu en ly enables analysis o i e me azoa and en plan genomes. Va ia ion- ools can also be used in command-line wi h any locally-ins alled Ensembl genome. Use s can inpu pe sonal collec ions o a ian s and mo i s, p o iding lexibili y in he analysis. Ó2019 The Au ho s. Published by Else ie B.V. on behal o Resea ch Ne wo k o Compu a ional and S uc u al Bio echnology. This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i e- commons.o g/licenses/by-nc-nd/4.0/). 1. In oduc ion Genomic DNA sequence ha bo s he gene egula o y in o ma- ion necessa y spa ial and empo al gene exp ession pa e ns [38,31]. Gene egula o y egions encompass sho , highly edun- dan DNA mo i s ecognized by ansc ip ion ac o s (TF) [36]. These egula o y egions may con ain gene ic a ian s, Single Nucleo ide Polymo phisms (SNPs) o indels, ha al e he DNA TF binding si e (TFBS), and he eby he binding o TF [20]. Mo e- o e , i has been epo ed ha 93.7% o a ian s ha ha e been associa ed wi h human ai s o diseases ha e been ound o be loca ed in non-coding egions [43,40], and pa icula ly en iched in open ch oma in egions [57], indica ing ha hese a ian s h ps://doi.o g/10.1016/j.csbj.2019.09.009 2001-0370/Ó2019 The Au ho s. Published by Else ie B.V. on behal o Resea ch Ne wo k o Compu a ional and S uc u al Bio echnology. This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i ecommons.o g/licenses/by-nc-nd/4.0/). Abb e ia ions: RSAT, Regula o y Sequence Analysis Tools; SNP, Single Nucleo- ide Polymo phism; TF, T ansc ip ion Fac o ; TFBS, T ansc ip ion Fac o Binding Si e; PSSM, Posi ion Speci ic Sco ing Ma ix; MPRA, Massi ely Pa allel Repo e Assays: MPRA; LD, Linkage Disequilib ium; sID, Re e ence SNP Iden i ie ; SOIs, SNPs o In e es ; GWAS, Genome Wide Associa ion S udies; CRM, Cis-Regula o y Module; eQTL, Exp ession Quan i a i e T ai Loci; ROC, Recei e Ope a ing Cha - ac e is ic; CEU, No he n Eu opeans om U ah. ⇑ Co esponding au ho s a : Labo a o io In e nacional de In es igación sob e el Genoma Humano, Uni e sidad Nacional Au ónoma de México, Campus Ju iquilla, Bl d Ju iquilla 3001, San iago de Que é a o 76230, México (Medina-Ri e a). Aix- Ma seille Uni , INSERM UMR S 1090, Theo y and App oaches o Genome Complexi y (TAGC), F-13288 Ma seille, F ance (J. an Helden ). E-mail add esses: [email p o ec ed] (J. an Helden), amedina@ liigh.unam.mx (A. Medina-Ri e a). Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 jou nal homepage: www.else ie .com/loca e/csbj may a ec ansc ip ional egula o y mechanisms, and he eby explain he obse ed pheno ypes. The Regula o y Sequence Analysis Tools (RSAT, h p:// sa .eu) [47,26] has es ablished i sel in he las 20 yea s as a majo so - wa e sui e dedica ed o he analysis o egula o y egions, wi h i e public se e s suppo ing mo e han 500 euka yo e and 9,000 p oka yo e genomes. Wi h a majo ocus on usabili y and accessi- bili y o use s wi h o wi hou o mal bioin o ma ics aining, RSAT p o ides ools o e ie e sequences, pe o m mo i s analysis, e al- ua e TF mo i quali y, compa e and clus e mo i s, con e ile o - ma s, e c. He e we desc ibe Va ia ion- ools, a subse o ools included in RSAT ha enable use s o analyse egula o y a ian s and assess hei pu a i e impac on TF binding si es. 1.1. Cu en app oaches o de ec ing po en ial egula o y a ian s The ac ha many a ian s a e loca ed in non-coding egions igge ed he de elopmen o bioin o ma ic ools o iden i y he egula o y po en ial o hese gene ic a ian s. S a ing om a lis o SNPs, compu a ional analyses can help o mula ing hypo heses on which TF may be impac ed by a gene ic a ian . Howe e , he e a e nume ous challenges o in silico analysis o un a el he impac o gene ic a ia ions in gene egula o y egions. Se e al ools and esou ces ha e been published, p o iding al e na i e me hods o ackle his p oblem (Table 1). Mos o hem a e ei he based on pa e n-ma ching app oaches o e alua e he impac o alleles on TF binding, o on machine lea ning models buil using unc ional anno a ions o he egula o y egions, e.g. epigenomics and an- sc ip omics da a. S ill, hese esou ces and ools ha e limi a ions hampe ing hei usage in se e al o ganisms [68,28,35], on new anno a ed a ian s [5,63,54], and/o on analyses wi h pe sonal col- lec ions o TF mo i s [68,28,41]. All ools in he Pa e n Ma ching ca ego y, (labeled PM in Table1) use Posi ion-Speci ic Sco ing Ma ices (PSSMs) o e alua e he a ini y o a TF o a gi en sequence wi h an allele. Majo di e - ences be ween hese ools can be ound in (i) hei a ailabili y: web pages [5], command line [12] o bo h [69]; (ii) lexibili y o he use o inpu hei own da a [61]; (iii) usabili y: he possibili y o use se e al a ian o ma s [28]; (i ) esul s ep esen a ion: ig- u es and/o ables [63]; ( ) a ailable o ganisms: only human [62], o o he o ganisms [61]; and ( i) he possibili y o calcula e esul s on- he- ly [41] o access p e-calcula ed ones [5]. Ano he se o ools (labeled ML in Table1) aim o he iden i i- ca ion o po en ial egula o y a ian s by in eg a ing se e al ypes o da a, beyond aking in o accoun po en ial dis up ion o TF bind- ing. Pa icula ly, Lee, e al. [37] in eg a ed DNaseI-seq da a wi h SVM app oaches o iden i y a ian s ha could po en ially dis up TF binding. DeepSea [68] in eg a es unc ional genomic da a om ChIP-seq, DNaseI-seq, RNA-seq and o he unc ional genomic high- h oughpu da a o assess he po en ial damage o a ian s ac oss he human genome. P ecalcula ed esul s o anno a ed a ian s can be accessed on hei websi e. Bo h ools can be ained on o he o ganisms, p o ided ha unc ional genomic da a a e a ailable. The main limi a ion o hese esou ces is he equi ed expe ise in bioin o ma ics and/o com- pu a ional esou ces o use s o analyse hei own da a se s. O he ools iden i y po en ial egula o y e ec s o a a ian by compa ing he measu ed a ini y o a TF o he di e en possible alleles. Ou ool, named a ia ion-scan, alls wi hin his ca ego y. 1.2. Va ia ion- ools In his con ex , we ha e de eloped Va ia ion- ools o add ess he main limi a ions iden i ied in exis ing p og ams (Table 1). Va ia ion- ools a e composed o ou p og ams ha enable (i) e ie al o in o ma ion o Ensembl anno a ed a ian s when a ail- able o a gi en genome in RSAT ( a ia ion-in o), (ii) con e sions be ween a ian ile o ma s (con e - a ia ions), (iii) e ie al o he sequences su ounding a ian s ( e ie e- a ia ion-seq), and (i ) scanning o di e en alleles o a a ian wi h one o se e al mo i s, compa ing he sco es and p- alues in o de o iden i y a ec ed TFBS ( a ia ion-scan)(Fig. 1). Ea lie e sions o hese p o- g ams we e epo ed in 2015 as pa o a RSAT upda e a icle [45], hese i s e sions we e de eloped in pe l and we e e ac o ed and imp o ed o he 2018 upda e [47]. In his a icle we p esen he la es e sions o he ools, wi h op imized memo y usage, and no el suppo o he inclusion o haplo ype in o ma ion. In summa y, RSAT Va ia ion- ools p o ide an accessible esou ce o expe ienced and non-expe use s o analyze egula o y a i- an s in a web in e ace o i een o ganisms ( i e me azoa (h p://me azoa. sa .eu) and en plan s (h p://plan s. sa .eu), wi h lexibili y o upload pe sonal a ian and PSSM collec ions. We desc ibe he e Va ia ion- ools me hodology, along wi h ou case s udies demons a ing he lexibili y o he ools, enabling he anal- ysis o da a se s om di e en o igin (Ensembl a ian s, Genome- Wide Associa ion S udy (GWAS) da a, ChIP-seq egions, e c.), com- plexi y, and o ganisms. 2. Me hods 2.1. Va ia ion- ools: om a ian s o iden i ica ion o egula o y e ec s Va ia ion- ools consis in a subse o ou ools wi hin RSAT de o ed o he iden i ica ion o gene ic a ian s pu a i ely a ec - ing TF binding 1) a ia ion-in o: his ool elies on he Ensembl gene ic a ia- ion in o ma ion [29] anno a ed and ins alled on he co e- sponding se e o each pa icula genome (i.e. human a ian s a e ins alled in he Me azoa se e ). I can ake wo di e en inpu s: 1) a ian sID o 2) genomic loci in bed o ma . This ool will e ie e he in o ma ion o he a ian s ma ching he IDs o he in o ma ion o he a ian s loca ed in he genomic loci. Va ian s ins alled in RSAT se - e s ha e been p ocessed o emo e a ian s wi h incom- ple e anno a ions (no alleles) o ambiguous coo dina es (non ma ching alleles coo dina es). When use s ha e hei own a ian s collec ions, hey can skip his ool and use di ec ly con e - a ia ions. 2) con e - a ia ions: enables he in e con e sion o a ian ile o ma s such as VCF, GVF and a Bed. a Bed is an in e nal o ma o RSAT ha acili a es he e ie al o he sequence su ounding he a ian (Supplemen a y Fig. 1A). 3) e ie e- a ia ion-seq: e ie es he sequence su ounding he a ian , and p oduces one sequence o each allele (Sup- plemen a y Fig. 1B). The ool can ake as inpu a a Bed ile (see con e - a ia ions). Fo o ganisms wi h Ensembl anno- a ed a ian s, i can ake a lis o IDs o a bed ile lis ing genomic loci. The ou pu is p o ided in a o ma named a - Seq, wi h each ow gi ing one allele wi h i s su ounding sequence. Each a ian has a speci ic in e nal ID o accom- moda e se e al a ian s wi h a ious alleles in he same ile. 4) a ia ion-scan: pe o ms he scanning o alleles wi h a PSSM and compa es he sco es and p- alues be ween alleles o assess he pu a i e e ec on TF binding (see de ails below) (Supplemen a y Fig. 2). I equi es as inpu a a Seq ile (see e ie e- a ia ion-seq), a mo i o collec ion o mo i s (o e wen y suppo ed ile o ma s), and a backg ound model ( o me hodological de ails on backg ound model, e e o [59] Box n°3). Di e en backg ound models a e ead- 1416 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 Table 1 Tools simila o a ia ion-scan wi h a ailable implemen a ion. PM s ands o Pa e n Ma ching, ML s ands o Machine Lea ning. Name PMID Sou ce App oach O ganism Inpu Ou pu Ma ix lexibili y Type Las upda e del aSVM 26075791 h p://www. bee lab.o g/ del as m/ Gapped k-me SVM classi ie . Any o ganism DNaseI-seq da a; pu a i e egula o y egions as posi i e aining se and andomized sequences as nega i e aining se . del aSVM, p edic ed impac o a a ian in ch oma in accessibili y which is measu ed by adding up he con ibu ion o all 10-me s in which he SNP is p esen o ch oma in accessibili y. I can only be ained o one TF a a ime. ML, non- s a ic. Las upda e Sep 2015. DeepSea 26301843 h p://deepsea. p ince on.edu/ job/analysis/ c ea e/ Deep con olu ional ne wo k. Human SNPs in VCF o ma . Ch oma in ea u e p obabili ies o e e ence and al e na i e alleles, ch oma in ea u e p obabili y log old changes o each a ian , ch oma in ea u e p obabili y di e ences o each a ian s, e- alues o ch oma in ea u e e ec s, unc ional signi icance sco e o each a ian . The e a e 919 ch oma in ea u es e alua ed. I con ains 690 TF binding p o iles o 160 di e en TFs, bu does no suppo he addi ion o new ma ices. ML, non- s a ic. Las upda e May 2017. a SNP 26092860 h ps:// gi hub.com/ keleslab/a SNP Impo ance sampling algo i hm o p- alue calcula ion, i s - o de Ma ko Model o gene a e andom backg ound sequences. Any o ganism whose genome is included in he Bioconduc o BSGenome package. SNP lis , mo i ile. p- alue o binding a ini y wi h al e na i e and e e ence allele, p- alue o binding a ini y change based on log-likelihood a io and log- ank a io. I also p o ides composi e logo plo s o di ec ly isualizing he SNP e ec s on mo i ma ches. I accep s se e al ma ices, and se e al di e en o ma s. I includes a mo i lib a y o 2,065 PSSMs om ENCODE and JASPAR, bu also allows use -de ined mo i lib a ies. PM, non- s a ic. Las upda e No 2018. BayesPI-BAR 26202972 h p:// olk.uio.no/ junbaiw/BayesPI- BAR/ Biophysical modeling o p o ein-DNA in e ac ion, es ima ion o TF chemical po en ial ( h ough a bayesian nonlinea eg ession model) and di e en ial binding a ini y. Any o ganism ChIP-seq expe imen o TFs o be es ed, DNA sequences o selec ed SNPs,PSSMs o selec ed TFs. Gi en a SNP and a PSSM lis , i p oduces wo lis s so ed by signi icance: one composed o binding mo i s dis up ed by he SNP, and one by si es wi h an inc eased a ini y o he TF caused by he SNP. Can use se e al PSSMs simul aneously. PM, biophysical modeling. Non-s a ic. No upda es lis ed, so wa e c ea ed July 2015. GWAS4D 29771388 h p://mulinlab. mu.edu.cn/ gwas4d/gwas4d/ gwas4d/gwas4d_ se e Va ian p io i iza ion me hod, ollowed by an in eg a i e analysis o genome-wide associa ion. Human Accep s VCF-like, coo dina e only, dbSNP ID and PLINK-like o ma s. Regula o y a ian p io i iza ion able: includes he mos likely a ec ed mo i by al e na i e a ian e ec . The model includes mo i s o 1,480 ansc ip ional egula o s om 13 di e en esou ces. I is no possible o upload use - speci ied ma ices. PM, s a ic Las upda e Sep 2018. sTRAP 20127973 h p:// ap.molgen. mpg.de/cgi-bin/ home.cgi P edic ion o local binding a ini y ollowed by a no maliza ion o binding a ini ies o de e mine di e ence be ween e e ence allele and SNP. O ganisms a ailable in TRANSFAC. Accep s only wo sequences in FASTA o ma . Lis o TFs anked acco ding o changes induced by he SNP. The e is no op ion o use -speci ied ma ices, ma ices om TRANSFAC e sions can be selec ed. PM, non- s a ic No upda es lis ed, so wa e c ea ed in 2011. (con inued on nex page) W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1417 Table 1 (con inued) Name PMID Sou ce App oach O ganism Inpu Ou pu Ma ix lexibili y Type Las upda e SNP2TFBS 27899579 h ps://ccg.ep l. ch//snp2 bs/ Es ima ion based on PSSM model. Human. When wo king wi h he code, he inpu equi ed is he e e ence genome, a SNP ca alogue and a PSSM collec ion. The web in e ace accep s SNP IDs and VCF o ma , as well as a speci ica ion o a genomic egion h ough a bed ile o by speci ying he s a and end posi ions. Lis o a ec ed TFBSs, so ed by he magni ude o he e ec s. On he web in e ace, only ma ices om JASPAR can be used. None heless, i is possible o download he code used o gene a e he da abase and use a di e en inpu . PM, s a ic. Las upda e July 2017. a SNP Sea ch 30534948 h p://a snp. bios a .wisc.edu/ Used a SNP algo i hm wi h dbSNP build 144 o human genome assembly 38 agains JASPAR and Encode mo i s o c ea e a eposi o y wi h all he SNP-mo i combina ions esul ing om he p e ious esou ces. Human. I can ecei e a se o sIDs, a sID and a window size a ound he SOI, genomic coo dina es, a gene symbol and a window size a ound he gene o in e es , o a TF name. Table including p- alues o mo i ma ches o bo h e e ence and al e na e alleles, as well as he change in he mo i ma ching and he di ec ion o said change. Ou pu includes logo plo s, displaying he sequence logos aligned o bes mo i ma ches wi h e e ence and SNP alleles. Only JASPAR o ENCODE ma ices can be selec ed, and i is possible o selec only one ansc ip ion ac o a a ime. PM, s a ic. Las upda e Jan 2018. HaploReg 22064851, 26657631 h ps://pubs. b oadins i u e. o g/mammals/ haplo eg/haplo eg. php I con ains da a om mul iple genome anno a ion esou ces. PSSMs a e sco ed agains e e ence and al e na i e alleles, and change in log-odds is calcula ed. Human Use s can p o ide a lis o sIDs o ch omosome egions. Use s can also selec GWAS s udies om he NHGRI ca alog. P o ides da a on allelic equencies, conse a ion, ch oma in s a es, and nea genes. Fo each o he egula o y mo i s al e ed by he SNP, i p o ides he change in log-odds and a logo. HaploReg con ains a lib a y c ea ed om li e a u e sou ces, TRANSFAC, JASPAR and PBM expe imen s. The e is no op ion o use - speci ied ma ices. PM, s a ic. Las upda e No embe 2015. RegulomeDB 22955989 h p://www. egulomedb.o g/ RegulomeDB uses in o ma ion om se e al da ase s, as well as manual cu a ion and a heu is ic me hod o dis inguish be ween unc ional and non- unc ional a ian s. Human. Use s can p o ide a lis o dbSNP IDs, hg19 coo dina es in BED, VCF o GFF3 o ma , o hg19 ch omosomal egions in he same o ma s. Table so ed by likely unc ionali y, con aining a ian coo dina es, sco e assigned by he algo i hm, and e idence o unc ion including p o ein binding, mo i s, ch oma in s uc u e, eQTLs and his one modi ica ions. RegulomeDB includes all PSSMs om TRANSFAC, JASPAR CORE, and UniP obe. The e is no op ion o use -speci ied ma ices. PM, s a ic. No upda es, lis ed, so wa e c ea ed in Sep 2012. mo i b eakR 26272984 h ps:// gi hub.com/ Simon-Coe zee/ Mo i B eakR I has h ee op ions o algo i hms: he s anda d sum o log p obabili ies, weigh ed sum, and an in o ma ion con en me hod. O ganisms included in BSgenome. SNPs can be impo ed om an R package o p o ided o he algo i hm in BED o VCF o ma . PSSMs can be selec ed om he Mo i Db package o be use - speci ied. Table con aining s a is ics desc ibing he pe cen o maximum sco e o a ma ix and ma ix alues o bo h alleles, as well as he s and. I also epo s whe he he TFBS is dis up ed s ongly o weakly. PSSMs can be impo ed om he Mo i Db package o be use - speci ied. Mo e han one ma ix can be used a a ime. PM, non- s a ic. Las upda e Jul 2018. a ia ion- scan h p:// sa .eu Es ima ion based on PSSM model. web in e ace: ins alled Ensembl o ganisms. command- line: any locally ins alled o ganism. A collec ion o PSSMs and a se o a ian s in a Seq o ma . This o ma can be ob ained using e ie e- a ia ion-seq. A able wi h one line pe pai o alleles pe mo i (i he e a e mo e han wo, he e will be one line pe possible pai ) epo ing he posi ion, weigh and p- alue o each allele, weigh di e ence and p- alue a io. Use s can selec o he collec ions a ailable in RSAT (JASPAR, HOCOMOCO, CisBP), bu hey can also use pe sonal collec ions. PM-non s a ic. Ap il 2019. 1418 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 ily a ailable h ough he web in e ace. Howe e , depending on he biological ques ion and ela ed po en ial biases, we ecommend he c ea ion o a dedica ed backg ound model, which can be done using he RSAT ool c ea e-backg ound, also a ailable ia he RSAT web in e ace. 2.2. Haplo ype p ocessing Gene ic a ian s can be de ec ed using high- h oughpu ech- niques. This has enabled he iden i ica ion o millions o a ian s in he HapMap [30] and 1000 genomes p ojec s [1]. Howe e , he in o ma ion on he a ian s alone is less use ul han knowing which g oups o alleles a e co-loca ed on he same ch omosome (haplo ype). The p ocess o iden i ying he a ian s ha belong o each ch omosome is known as phasing. Including haplo ype phasing in o ma ion acili a es he iden i ica ion o ela ions be ween a ian s [6]. VCF iles can include haplo ype phasing in o ma ion. The ool con e - a ia ions iden i ies and e ie es he phasing in o ma ion o he a ian s, while he ool e ie e- a ia ion-seq buil s he co - esponding haplo ype wi h all he SNPs ha lay wi hin a de ined window (de aul : 30 bp). 2.3. Compu ing binding speci ici y o a ansc ip ion ac o o a DNA sequence a ia ion-scan uses PSSMs o assess he binding speci ici y o a TF o a DNA sequence wi h di e en alleles in a gi en posi ion. The i s s ep o a ia ion-scan (i.e., scanning o he sequences wi h a gi en PSSM) is delega ed o he RSAT ool ma ix-scan. The sco - ing scheme and p- alue calcula ion a e desc ibed in de ail in [59], Box n°1 and Box n°2, espec i ely. In b ie : PSSM a e used o assess he binding speci ici y o a TF. This a ini y is calcula ed as a weigh sco e (Ws). The Ws o a si e in a ia ion-scan is calcula ed using [27]: Ws ¼lnðPðSjMÞ PðSjBÞ  whe e S is a sequence segmen o he same leng h o M, M is he PSSM, and B is he backg ound model. Hence, P(S|M) is he p obabil- i y o he sequence gi en he PSSM and P(S|B) is he p obabili y o he sequence gi en he backg ound model. Ws has been ela ed o he a ini y o he TF o he sequence, as i assesses simila i y o a sequence o a known se o binding si es, p o iding in o ma ion abou he p obabili y o a sequence o be a new ins ance o a bind- ing si e [56]. Mo eo e , i is possible o calcula e he p- alue o a gi en sco e as: P alue ¼PðWwjBÞ whe e he P- alue is calcula ed as he p obabili y o obse ing a sco e o a leas Wgi en a backg ound model w|B. When a sequence is longe han he PSSM, he PSSM is shi ed base by base un il he ull sequence has been sco ed. This scanning s ep is pe o med on he sequences o all epo ed alleles, so ha each allele is compa ed wi h all he posi ions o a gi en mo i . Backg ound models ep esen he nucleo ide composi ion o a se o sequences (whole genome, all p omo e sequences, e c.). These models a e used o es ima e he expec ancy o a nucleo ide being ound. Backg ound models can ep esen dependency be ween nucleo ides in sequences (e.g. aking in o accoun he e- Fig. 1. Schema ic ep esen a ion o Va ia ion- ools: This se o ools, included in he Regula o y Sequence Analysis Tools (RSAT), ocuses on assessing he impac o di e en allelic a ian s on T ansc ip ion ac o binding si es. A) con e - a ia ions allows use s o inpu hei own a ian s and con e hem o o he o ma s (VCF, GVF and a Bed, he la e is he o ma used in he nex s ep), while a ia ion-in o e ie es he anno a ed in o ma ion o Ensembl a ian s ins alled in RSAT se e s. B) The ool e ie e- a ia ion-seq e ie es he su ounding sequence o a ian s (including possible haplo ypes) and gene a es a ex ile wi h one line pe allele and pe a ian o haplo ype ( a Seq o ma ). C) Use s can inpu hei a ian s in a Seq o ma and a collec ion o mo i s (di ec inpu by he use o selec ed om RSAT a ailable collec ions) o a ia ion- scan; he ool hen scans he co esponding sequences wi h all mo i s and pe o m pai wise compa isons be ween he binding sco es o each ansc ip ion ac o on o all alleles o a a ian o haplo ype. W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1419 quencies o dinucleo ides o build a Ma ko model o o de 1 [59] Box n°3). As backg ound models a e used o calcula e weigh sco es o a binding si e (P(S|B)), i is impo an o selec an app o- p ia e model o each analysis. Examples o selec ed backg ound models a e p esen ed in he di e en s udy cases ma ching each pa icula biological ques ion. 2.4. Assessmen o allele e ec on ansc ip ion ac o binding In he second s ep, i.e., e alua ing he impac o SNPs, a ia ion- scan compa es he ob ained Ws (Ws di e ence = Ws_Allele1 – Ws_Alelle2) and he P- alue (P- alue a io = P- alue_Allele1/P- al ue_Allele2) o each o he alleles, posi ion by posi ion h oughou he scanning window. To e alua e indels, a ia ion-scan compa es he highes Ws and i s co esponding P- alue o each sequence o he epo ed alleles. When mo e han wo alleles o a a ian a e epo ed, all alleles a e compa ed o all alleles in a pai wise manne . 2.5. a ia ion-scan pe o mance es 2.5.1. Compu ing e iciency The ools a ia ion-in o and con e - a ia ion a e coded in Pe l, while e ie e- a ia ion-seq and a ia ion-scan a e coded in C, o enable he analysis o la ge numbe s o a ian s om euka yo ic genomes in a easonable ime. To u he imp o e pe o mance, we educed he da a ans e om he ha d d i e o memo y. a ia ion-scan pe o mance was assessed by andomly selec ing a a ian om he 1000 genomes p ojec [1] and a mo i om he RSAT non- edundan mo i s collec ion [8]. The andomly selec ed a ian was used o c ea e se s wi h di e en numbe s o epli- ca es, anging om one housand o nine millions, o es ima e he ela ion be ween unning ime and he amoun o e alua ed a ian s. The p ocesses we e un on a Dell Powe Edge C6145 se e wi h 2 AMD Op e on( m) P ocesso 6386 SE, 16 co es each, P oces- so speed o 2.8–3.5 Ghz, RAM 256 Gb and wi h an ope a ing sys- em Cen OS 7 (7.6.1810). 2.5.2. Da ase : expe imen ally-de e mined egula o y a ian s in ed blood cells The egula o y ac i i y o 2,756 ed blood cell a ian s has been sys ema ically measu ed using Massi ely Pa allel Repo e Assays (MPRA) [60]. MPRA is a high- h oughpu assay in which a lib a y o pu a i e egula o y elemen s, each ollowed by a unique ba - code, is inse ed in o a plasmid, hen ans ec ed in o a cell, and ansc ip s a e hen quan i ied h ough he abundance o ba codes. These a ian s a e known o be in s ong linkage disequilib ium (LD) wi h 75 a ian s associa ed wi h common ai s o his cell ype. Th ee sliding windows pe a ian (le , igh , and cen e ) we e syn hesized, ba coded and used o s udy he e ec o sligh changes in hei genomic con ex . Following me hods desc ibed by Uli sch, e al.[60], o each sequence mRNA/DNA a io was com- pu ed o ob ain a quan i a i e e alua ion o he egula o y e ec o a sequence a ian . 2.5.3. E alua ion o a ia ion-scan The a ian da ase was used as inpu o a ia ion-scan; he a ian s assessed in he ed blood cell assay we e anno a ed wi h he Ensembl GRCh37 human genome elease, and gi en as inpu o con e - a ia ions ollowed by e ie e- a ia ion-seq. Since h ee sliding windows we e used o each a ian in he MPRA, he co - esponding windows we e me ged be o e compu ing a backg ound model using he c ea e-backg ound-model ool. Acco ding o he o iginal s udy [60], binding si es o he ol- lowing TF we e en iched in he sequences o in e es : GATA1, KLF1, DHS, TAL1, ETS, FLI1 and AP-1. The e o e, a o al o 48 PSSMs anno a ed as ela ed o hese TF we e e ie ed om he non- edundan RSAT mo i collec ion [8], and gi en as inpu o a ia ion-scan. A nega i e con ol se o mo i s was c ea ed using he RSAT ool pe mu e-ma ix [47]; i e pemu ed mo i s we e c ea ed o each o he 48 mo i s, gene a ing a collec ion o 240 con ol mo i s. Fo a a ian o be epo ed in a ia ion-scan as posi i e,we eques ed ha a leas one o he allele sequences was e alua ed as a binding si es wi h a p- alue o a mos 10 4 (using he pa am- e e -u h p al 1e-4 in he command line), and ha he p- alue a io was g ea e o equal o en (a change o one o de o magni ude be ween he bes and he wo s allele p- alues) (-l h p al_ a io 10). We compa ed a ia ion-scan o wo o he ools p e iously used o assess he same se o a ian s by Uli sch, e al. [60]: DeepSea [50] and [37] del aSVM. In o de o a oid pe sonal biases when cal- ib a ing ool pa ame e s, we decided o ely on he published ones [60]. Fo his analysis a ia ion-scan was un wi hou h esholds o iden i y he impac o he pa ame e s, pa icula ly he h eshold on p- alue a io. 2.6. Case s udies 2.6.1. Case s udy 1: Iden i ica ion o egula o y a ian s in he ‘‘Pla inum”genomes haplo ypes The se o high-con idence a ian s om he wo CEU (No he n Eu opeans om U ah) human Pla inum Genomes NA12877 and NA12878 [23] we e downloaded h ough he Amazon Web Se ice (AWS) Command Line In e ace om he Illumina Pla inum Gen- omes AWS S3 bucke (h ps://gi hub.com/Illumina/Pla - inumGenomes). The downloaded VCF iles con ained phasing in o ma ion o each CEU indi idual haplo ype con igu a ion. The genome e sion used was GRCh37. We selec ed SNPs in e sec ing wi h he anno a ed DNAseI-seq clus e ed peaks V3 om he ENCODE p ojec [4]. The VCF ile wi h he selec ed SNPs was p ocessed using con e - a ia ions wi h he op ion phased and hen he haplo ype sequences we e econ- s uc ed wi h e ie e- a ia ion-seq. Fo a haplo ype SNP se o single posi ion a ian s o be epo ed in a ia ion-scan, we eques ed ha a leas one o he sequences was e alua ed as a binding si e wi h a p- alue o a mos 10 4 (-u h p al 1e-4) and ha he p- alue a io be ween he wo alleles was g ea e o equal o 100 (a change o wo o de s o mag- ni ude be ween he bes and he wo s alleles p- alues) (-l h p al_ a io 100). In addi ion, we equi e a change o sign be ween he bes and wo s sco e as an addi ional il e . We anno a ed he p edic ed dis up ed TFBS wi h he TF ChIP- seq non- edundan peak collec ion and wi h he Cis-Regula o y Modules (CRM) egions om ReMap [10] using bed ools in e sec e sion 2.27 [49]. We also calcula ed he en ichmen o anno a- ions in he p o enance sequence segmen s o he p edic ed haplo- ypes si es. 2.6.2. Case s udy 2: p edic ion o egula o y a ian s associa ed wi h suscep ibili y o Mycobac e ium ube culosis in ec ion We collec ed SNPs associa ed wi h he pheno ypic ai ‘‘suscep- ibili y o Mycobac e ium ube culosis in ec ion measu emen ” (dis- ease ID EFO_0008407) om he 1.0.2 e sion o he GWAS ca alog [40] (h ps://www.ebi.ac.uk/gwas/). This que y e u ned one s udy [58] wi h 67 dis inc a ian s, o which 48 had a alid e e ence SNP iden i ie ( sID) and could be u he used (deno ed he ea e as disease-associa ed SNPs, o DA-SNPs). To p edic he TF binding si es pu a i ely a ec ed by hese selec ed SNPs, we designed an app oach combining Va ia ion- ools wi h di e en ex e nal esou ces. We u he collec ed om Ensembl REST in e - ace (h p:// es .ensembl.o g/) 564 SNPs in linkage disequilib ium 1420 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 (LD-SNPs) in he Eu opean popula ion [62], wi h a h eshold on he eg ession coe icien ( 2 0.8) and a maximal dis ance o 200 bp. Anno a ions (ch omosomal loca ion, ype o genomic egion) o he esul ing 612 SNPs (48 DA + 564 LD) we e collec ed om Ensembl BioMa [22,21]. We hen es ic ed he selec ion o SNPs in non-coding egions, esul ing in a se o 572 SNPs o in e es (SOIs) o he de ec ion o egula o y a ian s. Using SNPs in LD, we de e mined LD-Block egions. These we e hen anno a ed based on o e laps wi h ChIP-seq peaks collec ed om he ReMap da a- base [10]. We also calcula ed en ichmen o disease anno a ions using he R XGR package [24]. Finally, we used e ie e- a ia ion-seq o e ie e he sequence a ian s a ound each SOI, and p edic ed he impac o he a ia ion on TF binding o each mo i o he JASPAR non- edundan RSAT mo i collec ion [8] using a ia ion-scan, wi h he h esholds o 1e-4 on he p- alue and 100 on he p- alue a io. 2.6.3. Case s udy 3: Assessmen o he egula o y e ec o GWAS epo ed a ian s in p omo e s wi h enhance unc ion The STARR-seq assay [2] is in i s p inciple simila o he MPRA, and helps iden i y sel - ansc ibing ac i e egula o y egions ha ha e enhance po en ial. Using his app oach Dao e al. [17], anal- ysed he enhance po en ial o anno a ed Re Seq p omo e s [48].In he wo cell lines K562 and HELA, hey iden i ied 632 and 493 p o- mo e s wi h enhance po en ial (eP omo e s), espec i ely. Mo e- o e , he au ho s iden i ied en ichmen o eQTL a ian s epo ed by GTEx [25]. To iden i y eP omo e s a ian s ha could be a ec ing TF bind- ing, we e ie ed he GWAS ca alog e sion 1.0 (downloaded on 7/01/19) [40]. Using bed ools o e lap e sion 2.26.0 [49], we com- pu ed he o e lap be ween SNPs and he eP omo e coo dina es epo ed in [17]. PSSMs ep esen ing TF en iched in eP omo e s we e also ob ained om [17], co esponding o SMRC1, JUN, FOS, ATF:MAF:NEF2, YY1, ETS amily, C eb and USF1/2. Using he selec ed GWAS a ian s ha all wi hin eP omo e s and he TF mo i s en iched in hese egions, we applied a ia ion- scan o assess he po en ial egula o y e ec o hese a ian s. a ia ion-scan was un wi h he pa ame e s – l h w_di 1 – l h p al_ a io 10, wi h a backg ound model buil using c ea e- backg ound wi h all Re Seq p omo e sequences. In o de o il a e a ian s wi h he highes pu a i e egula o y dis up ion, we u - he selec ed a ian s ha showed a change o sign in he weigh sco e be ween alleles. 2.6.4. Case s udy 4: iden i ica ion o egula o y a ian s a ec ing VRN1 binding in ba ley The la es e sion o Ho deum ulga e (ba ley) e e ence gen- ome [42] and a panel o mapped gene ic a ian s we e impo ed om Ensembl Genomes elease 42 [34] and ins alled in he RSAT Plan s se e (h p://plan s. sa .eu). We ob ained expe imen ally de e mined binding si es (ChIP-seq) o VRN1 om [19]. Since hese peaks we e o iginally posi ioned wi hin con igs o he 2012 genome assembly [14], hey had o be ma ched o he co espond- ing egions o he cu en assembly wi h BLAST + 2.9.0 (blas n) local alignmen s agains he epea -masked genome sequence (pe ec ma ches) [7]. Using bed ools o e lap e sion 2.26.0 [49], we selec ed a ian s alling wi hin he VRN1 epo ed binding peaks. The selec ed a ian s in VCF o ma we e hen p ocessed using con e - a ia ions and e ie e- a ia ion-seq o ob ain he sequences wi h he al e na i e alleles. The VRN1 DNA mo i used o scan he a ian s was ob ained om he oo p in DB plan collec ion [16] e sion: 2018-06 (h p:// lo es a.eead.csic.es/ oo p in db/index.php?mo i = AY750993:VRN1:EEADanno ). a ia ion-scan was used wi h a p e- compu ed backg ound Ma ko model (o de 1) o ba ley o assess he e ec o a ian s in TF binding, wi h he ollowing pa ame e s: – l h sco e 1 – l h w_di 1 – l h p al_ a io 10 – u h p al 1e-3. 2.7. A ailabili y Va ia ion- ools a e a ailable on he web (Me azoa: h p://me a- zoa. sa .eu/, Plan s: h p://plan s. sa .eu/, Teaching: h p:// each- ing. sa .eu/). The ools can be also ins alled o command-line usage wi h he RSAT sui e (h p://download. sa .eu/). The code and ma e ial o ep oduce he esul s p esen ed in he a icle can be accessed h ough Gi Hub (h ps://gi hub.com/RSAT- doc/supp-ma e ial-publica ions.gi ). 3. Resul s The Va ia ion- ools p o ide complemen a y p og ams enabling he e ie al o a ian s ( a ia ion-in o) and o hei su ounding sequences ( e ie e- a ia ion-seq), as well as in e con e sion be ween ile o ma s (con e - a ia ion). The main p edic i e p o- g am is a ia ion-scan, which can be used wi h any se o a ian s p o ided by he use (in VCF o GVF o ma s) o anno a ed in Ensembl ( om a lis o sIDs o a bed ile o iden i y o e lapping a ian s in genome coo dina es), wi h any se o mo i s selec ed om he collec ions a ailable in RSAT, o p o ided by he use . 3.1. a ia ion-scan accu a ely assesses he e ec o expe imen ally alida ed egula o y a ian s The o iginal e sion o a ia ion-scan [45] equi ed app oxi- ma ely i e hou s o assess he allele e ec o nine millions a i- an s. The no el e sion [47] signi ican ly educes he p ocessing ime o abou one hou (Supplemen a y Fig. 3). To e alua e he pe o mance o a ia ion-scan, we used an expe imen ally alida ed egula o y a ian se ob ained om a MPRA expe imen [60]. Fo all o he assessed allele pai s, we com- pa ed he weigh sco e di e ences compu ed wi h a ia ion-scan wi h he mRNA/DNA a io o he MPRA (see me hods). As shown in Supplemen a y Fig. 4A, we a e able o eco e only 9.37% o he expe imen ally alida ed a ian s wi h a ia ion-scan, as we eques ed a leas one o he alleles o ha e a binding si e o high con idence (p- alue 10 4 ). Focusing on he a ian s epo ed as posi i e in he MPRA da a se , we obse ed a weak co ela ion be ween he weigh di e ence and he MPRA mRNA/DNA a io in posi i e a ian s. Howe e , his co ela ion is no signi ican , as MPRA alues do no scale wi h he a ia ion-scan weigh di e - ences. Ne e heless, all a ian s show a p- alue a io indica i e o allele binding e ec s, showing ha a ia ion-scan gi es accu a e measu emen s o he impac o egula o y a ian s (Fig. 2A). Wi h he p oposed h esholds, we can con iden ly ejec 96.35% o MPRA nega i e sequences, which could be imp o ed using mo e es ic i e pa ame e , wi h a concomi an educ ion in ue posi- i es. No ewo hy, as any high- h oughpu assay, MPRA has i s limi a ions [51] and sequencing biases could inc ease he numbe o alse nega i es. We pe o med a nega i e con ol, consis ing o 240 pe mu ed ma ices ( i e pe mu ed e sions o he 48 mo i s). Wi h his col- lec ion, i was s ill possible o eco e a g oup o a ian s, bu i only ep esen ed 31.2% o he MPRA posi i e a ian s (Supplemen- a y Fig. 4B). We compa ed he pe o mance o a ia ion-scan o wo o he ools ha had been p e iously used by Uli sch, e al [60] o assess he same se o MPRA a ian s: DeepSea [68] and del aSVM [37]. We decided o use he same pa ame e s in o de o a oid pe sonal biases when calib a ing he ools. The e o e, aining weigh s o DNAse I hype sensi i i y si es we e used in he del aSVM analysis. W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1421 As o DeepSea, he web implemen a ion o he ool was used, wi h he me ic Func ional Signi icance Sco e. Tools we e compa ed based on ROC cu es (Fig. 2B), we addi- ionally an a ia ion-scan using a se o pe mu ed ma ices as neg- a i e con ol (Fig. 2B, ed line). The h ee ools show e y simila sensi i i y s speci ici y a he beginning o he cu es, bu only a ia ion-scan and DeepSea u he emain sepa a ed om he neg- a i e con ol. As expec ed DeepSea pe o ms sligh ly be e han a ia ion-scan a he beginning o he cu e, ne e heless his ool equi es aining using epigene ic da a, while a ia ion-scan equi es only a mo i and a se o a ian s. 3.2. Va ia ion- ools case s udies To illus a e he di e se applica ions o Va ia ion- ools o ackle a ious biological ques ions, we designed ou di e en case s udies: 1. Impac o egula o y a ian s in he same haplo ype on TF bind- ing si es. 2. Iden i ica ion o he egula o y po en ial o a ian s epo ed in GWAS. 3. Assessmen o he egula o y po en ial o GWAS a ian s wi hin expe imen ally de e mined egula o y egions. 4. De e mina ion o egula o y a ian s wi hin TF binding egions iden i ied using ChIP-seq [19]. 3.2.1. Genome-wide haplo ype a ian in o ma ion can be used o iden i y se s o egula o y a ian s a ec ing he same TFBS The lowe ing cos s in sequencing ha e made i possible o ob ain whole genome sequences o mo e indi iduals, opening he possibili y o knowing, no only he a ian s o a genome, bu also he haplo ypes, and de e mining which a ian s a e passed linked wi hin he same ch omosome. This enables he assessmen o he egula o y e ec s o se s o a ian s wi hin he same haplo ype in a gi en TFBS. Using he high-con idence SNPs om wo ‘‘Pla inum” Genomes [23], we de e mined haplo ype a ian s ha a e likely o a ec one TFBS. We selec ed a ian s 30bps apa , loca ed in open ch oma in, o be analysed wi h a ia ion-scan using he non- edundan mo i collec ion a RSAT [8]. We de ec ed 7,406 haplo ype si es wi h a leas wo he e ozygous a ian s and a p obable e ec in binding o 361 TFs. O e all he numbe o he e ozygous a ian s wi hin a haplo ype inc eases he measu ed weigh di e ence. This is expec ed as mo e changes in he binding si es a e mo e likely o change TF a ini y (Fig. 3A). To assess he biological ele ance o all he pu a i e dis up ed TFBS p edic ions, we anno a ed 7,485 p edic ed haplo ypes si es con aining wo o mo e a ian s wi h a leas one he e ozygous a ian and 15,396 p edic ed si es con aining a SNP (single ons) wi h he TF ChIP-seq peaks and he Cis-Regula o y Modules (CRM) egions om ReMap [10]. We ound ha almos all he p e- dic ed dis up ed TFBS (~85%) con ain a CRM o peak anno a ion o bo h (Fig. 3B). In e es ingly, we ound en ichmen o CRM and peak anno a ions in he p o enance sequence segmen s o he 7,485 p edic ed haplo ypes si es compa ed o he p o enance sequence segmen s o he single a ian s (Fishe exac es , p- alue < 2.2e- 16). One o hese anno a ed haplo ypes is composed o he mino alleles o wo SNPs ( s2732317 and s2732318), whe e we obse ed a po en ial egula o y e ec likely a ec ing h ee binding mo i s, o EHF/ELF2, ETV4/ELK1/ETS1/FLI1/ELK4/ETS2/FEV/GABP1, and ELK3/ELF1/ERG/GABPA (Fig. 3C). 3.2.2. Gene ic a ian s associa ed wi h Mycobac e ium ube culosis in ec ion show po en ial egula o y e ec s The second case s udy illus a es a knowledge- ee use o Va ia ion- ools o iden i y egula o y a ian s om GWAS s udies o a use -speci ied disease, wi hou p io indica ion abou he po en ially in ol ed ansc ip ion ac o s o binding mo i s. The app oach is based on he p edic ion o egula o y a ian s wi h RSAT Va ia ion- ools, na owed down by selec ing he egula o y SNPs ha o e lap ChIP-seq peaks in ReMap [10], in o de o iden- i y con e gen indica ions o a po en ial impac o he a ian s on he binding o a TF. A) B) 0.00 0.25 0.50 0.75 1.00 0.000.250.500.751.00 speci ici y sensi i i y name pe mu ed del aSVM a ia ion-sca n DeepSea P- alue a io 100 P- alue a io 1000 P- alue a io 10 R=0.12,p=0.5 0 200 400 −9 −6 −3 0 MPRA p− alue o di e en ial ac i i y (log10) Va ia ion−scan p− alue a io Fig. 2. Iden i ica ion o expe imen ally alida ed egula o y a ian s using a ia ion-scan. A) Co ela ion o he Massi ely Pa allel Repo e Assays (MPRA) p- alue o he mRNA/DNA a io o posi i e a ian s and he a ia ion-scan weigh di e ence o he MPRA a ian s wi h signi ican change. B) Recei e Ope a ing Cha ac e is ic (ROC) cu e compa ing he pe o mance when aiming o classi y MPRA expe imen ally analyzed a ian s using a ia ion-scan ( u quoise), DeepSea (pu ple), del aSVM (g een), and a nega i e con ol which consis s o pe mu ed mo i s sco ed wi h a ia ion-scan ( ed). 1422 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 Fig. 3. Haplo ype analysis in high-quali y human genomes. A) The numbe o he e ozygous a ian s (X-axis) wi hin he same pu a i e binding si e end o ha e a g ea e impac on he TF binding p obabili y. This is expec ed as he inc ease o weigh di e ence obse ed on he iolin plo co esponds o he expec ed cumula ed impac o a ia ions a ec ing di e en posi ions o he same binding si e. B) Numbe o p edic ed dis up ed T ansc ip ion Fac o Binding Si es (TFBSs) wi h Cis-Regula o y Modules (CRMs) and TF ChIP-seq peak anno a ion (blue), wi h only peak anno a ion (yellow), and non-anno a ed p edic ions (g ey). C) Uni e si y o Cali o nia San a C uz (UCSC) b owse [48] sc een sho , showing a locus encompassing wo SNPs ha compose an he e ozygous haplo ype in one o he No he n Eu opeans om U ah (CEU) indi iduals. The igu e shows he e e ence genome haplo ype. The a ian s a e loca ed in he FUT10 p omo e ( op). a ia ion-scan p edic s an e ec in h ee mo i s ha ep esen binding si es o GABPA, ETS1 and ELF2, ac o s ha ha e been p o en o ha e binding si es in his egion by he ENCODE p ojec . The a ian s2732317 has been associa ed wi h e ec s in gene exp ession by he GTEx p ojec . W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1423