scieee Science in your language
[en] (orig)

RSAT variation-tools: An accessible and flexible framework to predict the impact of regulatory variants on transcription factor binding

Abstract

Gene regulatory regions contain short and degenerated DNA binding sites recognized by transcription factors (TFBS). When TFBS harbor SNPs, the DNA binding site may be affected, thereby altering the transcriptional regulation of the target genes. Such regulatory SNPs have been implicated as causal variants in Genome-Wide Association Study (GWAS) studies. In this study, we describe improved versions of the programs Variation-tools designed to predict regulatory variants, and present four case studies to illustrate their usage and applications. In brief, Variation-tools facilitate i) obtaining variation information, ii) interconversion of variation file formats, iii) retrieval of sequences surrounding variants, and iv) calculating the change on predicted transcription factor affinity scores between alleles, using motif scanning approaches. Notably, the tools support the analysis of haplotypes. The tools are included within the well-maintained suite Regulatory Sequence Analysis Tools (RSAT, http://rsat.eu), and accessible through a web interface that currently enables analysis of five metazoa and ten plant genomes. Variation-tools can also be used in command-line with any locally-installed Ensembl genome. Users can input personal collections of variants and motifs, providing flexibility in the analysis. Santana-Garcia, W.; Rocha-Acevedo, M.; Ramirez-Navarro, L.; Mbouamboua, Y.; Thieffry, D.; Thomas-Chollier, M.; Contreras-Moreira, B.; van Helden, J.; Medina-Rivera, A.

Read accessible full text

RSAT variation-tools: An accessible and flexible framework to predict the impact of regulatory variants on transcription factor binding

Author: Santana-Garcia, W.; Rocha-Acevedo, M.; Thieffry, D.; van Helden, J.; Medina-Rivera, A.; Mbouamboua, Y.; Ramirez-Navarro, L.; Thomas-Chollier, M.; Contreras-Moreira, B.
Year: 2019
DOI: 10.1016/j.csbj.2019.09.009
Source: https://zaguan.unizar.es/record/99341/files/texto_completo.pdf
RSAT a ia ion- ools: An accessible and lexible amewo k o p edic he
impac o egula o y a ian s on ansc ip ion ac o binding
Wal e San ana-Ga cia
a,b
, Ma ia Rocha-Ace edo
b
, Lucia Rami ez-Na a o
b
, Y on Mbouamboua
c,
,
Denis Thie y
a
, Mo gane Thomas-Chollie
a
, B uno Con e as-Mo ei a
d,e
, Jacques an Helden
,g,
⇑
,
Alejand a Medina-Ri e a
b,*
a
Ins i u de Biologie de l’ENS (IBENS), Dépa emen de biologie, École no male supé ieu e, CNRS, INSERM, Uni e si é PSL, 75005 Pa is, F ance
b
Labo a o io In e nacional de In es igación sob e el Genoma Humano, Uni e sidad Nacional Au ónoma de México, Campus Ju iquilla, Bl d Ju iquilla 3001, San iago de Que é a o
76230, Mexico
c
Fonda ion Congolaise pou la Reche che Médicale, B azza ille, People’s Republic o Congo
d
Es ación Expe imen al de Aula Dei-CSIC, Za agoza, Spain
e
Fundación ARAID, Za agoza, Spain
Aix-Ma seille Uni , INSERM UMR S 1090, Theo y and App oaches o Genome Complexi y (TAGC), F-13288 Ma seille, F ance
g
CNRS, Ins i u F ançais de Bioin o ma ique, IFB-co e, UMS 3601, E y, F ance
a icle in o
A icle his o y:
Recei ed 27 Ap il 2019
Recei ed in e ised o m 22 Sep embe
2019
Accep ed 25 Sep embe 2019
A ailable online 7 No embe 2019
Keywo ds:
Regula o y a ian s
T ansc ip ion ac o s
Posi ion speci ic sco ing ma ix
SNPs
Binding mo i s
abs ac
Gene egula o y egions con ain sho and degene a ed DNA binding si es ecognized by ansc ip ion
ac o s (TFBS). When TFBS ha bo SNPs, he DNA binding si e may be a ec ed, he eby al e ing he an-
sc ip ional egula ion o he a ge genes. Such egula o y SNPs ha e been implica ed as causal a ian s in
Genome-Wide Associa ion S udy (GWAS) s udies. In his s udy, we desc ibe imp o ed e sions o he
p og ams Va ia ion- ools designed o p edic egula o y a ian s, and p esen ou case s udies o illus-
a e hei usage and applica ions. In b ie , Va ia ion- ools acili a e i) ob aining a ia ion in o ma ion,
ii) in e con e sion o a ia ion ile o ma s, iii) e ie al o sequences su ounding a ian s, and i ) calcu-
la ing he change on p edic ed ansc ip ion ac o a ini y sco es be ween alleles, using mo i scanning
app oaches. No ably, he ools suppo he analysis o haplo ypes. The ools a e included wi hin he
well-main ained sui e Regula o y Sequence Analysis Tools (RSAT, h p:// sa .eu), and accessible h ough
a web in e ace ha cu en ly enables analysis o i e me azoa and en plan genomes. Va ia ion- ools can
also be used in command-line wi h any locally-ins alled Ensembl genome. Use s can inpu pe sonal
collec ions o a ian s and mo i s, p o iding lexibili y in he analysis.
Ó2019 The Au ho s. Published by Else ie B.V. on behal o Resea ch Ne wo k o Compu a ional and
S uc u al Bio echnology. This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i e-
commons.o g/licenses/by-nc-nd/4.0/).
1. In oduc ion
Genomic DNA sequence ha bo s he gene egula o y in o ma-
ion necessa y spa ial and empo al gene exp ession pa e ns
[38,31]. Gene egula o y egions encompass sho , highly edun-
dan DNA mo i s ecognized by ansc ip ion ac o s (TF) [36].
These egula o y egions may con ain gene ic a ian s, Single
Nucleo ide Polymo phisms (SNPs) o indels, ha al e he DNA
TF binding si e (TFBS), and he eby he binding o TF [20]. Mo e-
o e , i has been epo ed ha 93.7% o a ian s ha ha e been
associa ed wi h human ai s o diseases ha e been ound o be
loca ed in non-coding egions [43,40], and pa icula ly en iched
in open ch oma in egions [57], indica ing ha hese a ian s
h ps://doi.o g/10.1016/j.csbj.2019.09.009
2001-0370/Ó2019 The Au ho s. Published by Else ie B.V. on behal o Resea ch Ne wo k o Compu a ional and S uc u al Bio echnology.
This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i ecommons.o g/licenses/by-nc-nd/4.0/).
Abb e ia ions: RSAT, Regula o y Sequence Analysis Tools; SNP, Single Nucleo-
ide Polymo phism; TF, T ansc ip ion Fac o ; TFBS, T ansc ip ion Fac o Binding
Si e; PSSM, Posi ion Speci ic Sco ing Ma ix; MPRA, Massi ely Pa allel Repo e
Assays: MPRA; LD, Linkage Disequilib ium; sID, Re e ence SNP Iden i ie ; SOIs,
SNPs o In e es ; GWAS, Genome Wide Associa ion S udies; CRM, Cis-Regula o y
Module; eQTL, Exp ession Quan i a i e T ai Loci; ROC, Recei e Ope a ing Cha -
ac e is ic; CEU, No he n Eu opeans om U ah.
⇑
Co esponding au ho s a : Labo a o io In e nacional de In es igación sob e el
Genoma Humano, Uni e sidad Nacional Au ónoma de México, Campus Ju iquilla,
Bl d Ju iquilla 3001, San iago de Que é a o 76230, México (Medina-Ri e a). Aix-
Ma seille Uni , INSERM UMR S 1090, Theo y and App oaches o Genome
Complexi y (TAGC), F-13288 Ma seille, F ance (J. an Helden ).
E-mail add esses: [email p o ec ed] (J. an Helden), amedina@
liigh.unam.mx (A. Medina-Ri e a).
Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
jou nal homepage: www.else ie .com/loca e/csbj
may a ec ansc ip ional egula o y mechanisms, and he eby
explain he obse ed pheno ypes.
The Regula o y Sequence Analysis Tools (RSAT, h p:// sa .eu)
[47,26] has es ablished i sel in he las 20 yea s as a majo so -
wa e sui e dedica ed o he analysis o egula o y egions, wi h i e
public se e s suppo ing mo e han 500 euka yo e and 9,000
p oka yo e genomes. Wi h a majo ocus on usabili y and accessi-
bili y o use s wi h o wi hou o mal bioin o ma ics aining, RSAT
p o ides ools o e ie e sequences, pe o m mo i s analysis, e al-
ua e TF mo i quali y, compa e and clus e mo i s, con e ile o -
ma s, e c. He e we desc ibe Va ia ion- ools, a subse o ools
included in RSAT ha enable use s o analyse egula o y a ian s
and assess hei pu a i e impac on TF binding si es.
1.1. Cu en app oaches o de ec ing po en ial egula o y a ian s
The ac ha many a ian s a e loca ed in non-coding egions
igge ed he de elopmen o bioin o ma ic ools o iden i y he
egula o y po en ial o hese gene ic a ian s. S a ing om a lis
o SNPs, compu a ional analyses can help o mula ing hypo heses
on which TF may be impac ed by a gene ic a ian . Howe e , he e
a e nume ous challenges o in silico analysis o un a el he impac
o gene ic a ia ions in gene egula o y egions. Se e al ools and
esou ces ha e been published, p o iding al e na i e me hods o
ackle his p oblem (Table 1). Mos o hem a e ei he based on
pa e n-ma ching app oaches o e alua e he impac o alleles on
TF binding, o on machine lea ning models buil using unc ional
anno a ions o he egula o y egions, e.g. epigenomics and an-
sc ip omics da a. S ill, hese esou ces and ools ha e limi a ions
hampe ing hei usage in se e al o ganisms [68,28,35], on new
anno a ed a ian s [5,63,54], and/o on analyses wi h pe sonal col-
lec ions o TF mo i s [68,28,41].
All ools in he Pa e n Ma ching ca ego y, (labeled PM in
Table1) use Posi ion-Speci ic Sco ing Ma ices (PSSMs) o e alua e
he a ini y o a TF o a gi en sequence wi h an allele. Majo di e -
ences be ween hese ools can be ound in (i) hei a ailabili y:
web pages [5], command line [12] o bo h [69]; (ii) lexibili y o
he use o inpu hei own da a [61]; (iii) usabili y: he possibili y
o use se e al a ian o ma s [28]; (i ) esul s ep esen a ion: ig-
u es and/o ables [63]; ( ) a ailable o ganisms: only human [62],
o o he o ganisms [61]; and ( i) he possibili y o calcula e esul s
on- he- ly [41] o access p e-calcula ed ones [5].
Ano he se o ools (labeled ML in Table1) aim o he iden i i-
ca ion o po en ial egula o y a ian s by in eg a ing se e al ypes
o da a, beyond aking in o accoun po en ial dis up ion o TF bind-
ing. Pa icula ly, Lee, e al. [37] in eg a ed DNaseI-seq da a wi h
SVM app oaches o iden i y a ian s ha could po en ially dis up
TF binding. DeepSea [68] in eg a es unc ional genomic da a om
ChIP-seq, DNaseI-seq, RNA-seq and o he unc ional genomic
high- h oughpu da a o assess he po en ial damage o a ian s
ac oss he human genome. P ecalcula ed esul s o anno a ed
a ian s can be accessed on hei websi e.
Bo h ools can be ained on o he o ganisms, p o ided ha
unc ional genomic da a a e a ailable. The main limi a ion o hese
esou ces is he equi ed expe ise in bioin o ma ics and/o com-
pu a ional esou ces o use s o analyse hei own da a se s. O he
ools iden i y po en ial egula o y e ec s o a a ian by compa ing
he measu ed a ini y o a TF o he di e en possible alleles. Ou
ool, named a ia ion-scan, alls wi hin his ca ego y.
1.2. Va ia ion- ools
In his con ex , we ha e de eloped Va ia ion- ools o add ess he
main limi a ions iden i ied in exis ing p og ams (Table 1).
Va ia ion- ools a e composed o ou p og ams ha enable (i)
e ie al o in o ma ion o Ensembl anno a ed a ian s when a ail-
able o a gi en genome in RSAT ( a ia ion-in o), (ii) con e sions
be ween a ian ile o ma s (con e - a ia ions), (iii) e ie al o
he sequences su ounding a ian s ( e ie e- a ia ion-seq), and
(i ) scanning o di e en alleles o a a ian wi h one o se e al
mo i s, compa ing he sco es and p- alues in o de o iden i y
a ec ed TFBS ( a ia ion-scan)(Fig. 1). Ea lie e sions o hese p o-
g ams we e epo ed in 2015 as pa o a RSAT upda e a icle [45],
hese i s e sions we e de eloped in pe l and we e e ac o ed and
imp o ed o he 2018 upda e [47]. In his a icle we p esen he
la es e sions o he ools, wi h op imized memo y usage, and
no el suppo o he inclusion o haplo ype in o ma ion.
In summa y, RSAT Va ia ion- ools p o ide an accessible esou ce
o expe ienced and non-expe use s o analyze egula o y a i-
an s in a web in e ace o i een o ganisms ( i e me azoa
(h p://me azoa. sa .eu) and en plan s (h p://plan s. sa .eu), wi h
lexibili y o upload pe sonal a ian and PSSM collec ions. We
desc ibe he e Va ia ion- ools me hodology, along wi h ou case
s udies demons a ing he lexibili y o he ools, enabling he anal-
ysis o da a se s om di e en o igin (Ensembl a ian s, Genome-
Wide Associa ion S udy (GWAS) da a, ChIP-seq egions, e c.), com-
plexi y, and o ganisms.
2. Me hods
2.1. Va ia ion- ools: om a ian s o iden i ica ion o egula o y
e ec s
Va ia ion- ools consis in a subse o ou ools wi hin RSAT
de o ed o he iden i ica ion o gene ic a ian s pu a i ely a ec -
ing TF binding
1) a ia ion-in o: his ool elies on he Ensembl gene ic a ia-
ion in o ma ion [29] anno a ed and ins alled on he co e-
sponding se e o each pa icula genome (i.e. human
a ian s a e ins alled in he Me azoa se e ). I can ake
wo di e en inpu s: 1) a ian sID o 2) genomic loci in
bed o ma . This ool will e ie e he in o ma ion o he
a ian s ma ching he IDs o he in o ma ion o he a ian s
loca ed in he genomic loci. Va ian s ins alled in RSAT se -
e s ha e been p ocessed o emo e a ian s wi h incom-
ple e anno a ions (no alleles) o ambiguous coo dina es
(non ma ching alleles coo dina es). When use s ha e hei
own a ian s collec ions, hey can skip his ool and use
di ec ly con e - a ia ions.
2) con e - a ia ions: enables he in e con e sion o a ian ile
o ma s such as VCF, GVF and a Bed. a Bed is an in e nal
o ma o RSAT ha acili a es he e ie al o he sequence
su ounding he a ian (Supplemen a y Fig. 1A).
3) e ie e- a ia ion-seq: e ie es he sequence su ounding
he a ian , and p oduces one sequence o each allele (Sup-
plemen a y Fig. 1B). The ool can ake as inpu a a Bed ile
(see con e - a ia ions). Fo o ganisms wi h Ensembl anno-
a ed a ian s, i can ake a lis o IDs o a bed ile lis ing
genomic loci. The ou pu is p o ided in a o ma named a -
Seq, wi h each ow gi ing one allele wi h i s su ounding
sequence. Each a ian has a speci ic in e nal ID o accom-
moda e se e al a ian s wi h a ious alleles in he same ile.
4) a ia ion-scan: pe o ms he scanning o alleles wi h a PSSM
and compa es he sco es and p- alues be ween alleles o
assess he pu a i e e ec on TF binding (see de ails below)
(Supplemen a y Fig. 2). I equi es as inpu a a Seq ile
(see e ie e- a ia ion-seq), a mo i o collec ion o mo i s
(o e wen y suppo ed ile o ma s), and a backg ound
model ( o me hodological de ails on backg ound model,
e e o [59] Box n°3). Di e en backg ound models a e ead-
1416 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
Table 1
Tools simila o a ia ion-scan wi h a ailable implemen a ion. PM s ands o Pa e n Ma ching, ML s ands o Machine Lea ning.
Name PMID Sou ce App oach O ganism Inpu Ou pu Ma ix lexibili y Type Las
upda e
del aSVM 26075791 h p://www.
bee lab.o g/
del as m/
Gapped k-me SVM classi ie . Any o ganism DNaseI-seq da a;
pu a i e egula o y
egions as posi i e
aining se and
andomized sequences as
nega i e aining se .
del aSVM, p edic ed impac o a
a ian in ch oma in accessibili y
which is measu ed by adding up
he con ibu ion o all 10-me s in
which he SNP is p esen o
ch oma in accessibili y.
I can only be ained o
one TF a a ime.
ML, non-
s a ic.
Las
upda e
Sep 2015.
DeepSea 26301843 h p://deepsea.
p ince on.edu/
job/analysis/
c ea e/
Deep con olu ional ne wo k. Human SNPs in VCF o ma . Ch oma in ea u e p obabili ies
o e e ence and al e na i e
alleles, ch oma in ea u e
p obabili y log old changes o
each a ian , ch oma in ea u e
p obabili y di e ences o each
a ian s, e- alues o ch oma in
ea u e e ec s, unc ional
signi icance sco e o each a ian .
The e a e 919 ch oma in ea u es
e alua ed.
I con ains 690 TF binding
p o iles o 160 di e en
TFs, bu does no suppo
he addi ion o new
ma ices.
ML, non-
s a ic.
Las
upda e
May 2017.
a SNP 26092860 h ps://
gi hub.com/
keleslab/a SNP
Impo ance sampling algo i hm
o p- alue calcula ion, i s -
o de Ma ko Model o
gene a e andom backg ound
sequences.
Any o ganism
whose
genome is
included in
he
Bioconduc o
BSGenome
package.
SNP lis , mo i ile. p- alue o binding a ini y wi h
al e na i e and e e ence allele, p-
alue o binding a ini y change
based on log-likelihood a io and
log- ank a io. I also p o ides
composi e logo plo s o di ec ly
isualizing he SNP e ec s on
mo i ma ches.
I accep s se e al
ma ices, and se e al
di e en o ma s. I
includes a mo i lib a y o
2,065 PSSMs om
ENCODE and JASPAR, bu
also allows use -de ined
mo i lib a ies.
PM, non-
s a ic.
Las
upda e
No 2018.
BayesPI-BAR 26202972 h p:// olk.uio.no/
junbaiw/BayesPI-
BAR/
Biophysical modeling o
p o ein-DNA in e ac ion,
es ima ion o TF chemical
po en ial ( h ough a bayesian
nonlinea eg ession model)
and di e en ial binding a ini y.
Any o ganism ChIP-seq expe imen o
TFs o be es ed, DNA
sequences o selec ed
SNPs,PSSMs o selec ed
TFs.
Gi en a SNP and a PSSM lis , i
p oduces wo lis s so ed by
signi icance: one composed o
binding mo i s dis up ed by he
SNP, and one by si es wi h an
inc eased a ini y o he TF caused
by he SNP.
Can use se e al PSSMs
simul aneously.
PM,
biophysical
modeling.
Non-s a ic.
No
upda es
lis ed,
so wa e
c ea ed
July 2015.
GWAS4D 29771388 h p://mulinlab.
mu.edu.cn/
gwas4d/gwas4d/
gwas4d/gwas4d_
se e
Va ian p io i iza ion me hod,
ollowed by an in eg a i e
analysis o genome-wide
associa ion.
Human Accep s VCF-like,
coo dina e only, dbSNP
ID and PLINK-like
o ma s.
Regula o y a ian p io i iza ion
able: includes he mos likely
a ec ed mo i by al e na i e
a ian e ec .
The model includes
mo i s o 1,480
ansc ip ional egula o s
om 13 di e en
esou ces. I is no
possible o upload use -
speci ied ma ices.
PM, s a ic Las
upda e
Sep 2018.
sTRAP 20127973 h p:// ap.molgen.
mpg.de/cgi-bin/
home.cgi
P edic ion o local binding
a ini y ollowed by a
no maliza ion o binding
a ini ies o de e mine
di e ence be ween e e ence
allele and SNP.
O ganisms
a ailable in
TRANSFAC.
Accep s only wo
sequences in FASTA
o ma .
Lis o TFs anked acco ding o
changes induced by he SNP.
The e is no op ion o
use -speci ied ma ices,
ma ices om TRANSFAC
e sions can be selec ed.
PM, non-
s a ic
No
upda es
lis ed,
so wa e
c ea ed in
2011.
(con inued on nex page)
W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1417
Table 1 (con inued)
Name PMID Sou ce App oach O ganism Inpu Ou pu Ma ix lexibili y Type Las
upda e
SNP2TFBS 27899579 h ps://ccg.ep l.
ch//snp2 bs/
Es ima ion based on PSSM
model.
Human. When wo king wi h he
code, he inpu equi ed
is he e e ence genome,
a SNP ca alogue and a
PSSM collec ion.
The web in e ace
accep s SNP IDs and VCF
o ma , as well as a
speci ica ion o a
genomic egion h ough a
bed ile o by speci ying
he s a and end
posi ions.
Lis o a ec ed TFBSs, so ed by
he magni ude o he e ec s.
On he web in e ace,
only ma ices om
JASPAR can be used.
None heless, i is possible
o download he code
used o gene a e he
da abase and use a
di e en inpu .
PM, s a ic. Las
upda e
July 2017.
a SNP
Sea ch
30534948 h p://a snp.
bios a .wisc.edu/
Used a SNP algo i hm wi h
dbSNP build 144 o human
genome assembly 38 agains
JASPAR and Encode mo i s o
c ea e a eposi o y wi h all he
SNP-mo i combina ions
esul ing om he p e ious
esou ces.
Human. I can ecei e a se o
sIDs, a sID and a
window size a ound he
SOI, genomic coo dina es,
a gene symbol and a
window size a ound he
gene o in e es , o a TF
name.
Table including p- alues o mo i
ma ches o bo h e e ence and
al e na e alleles, as well as he
change in he mo i ma ching and
he di ec ion o said change.
Ou pu includes logo plo s,
displaying he sequence logos
aligned o bes mo i ma ches
wi h e e ence and SNP alleles.
Only JASPAR o ENCODE
ma ices can be selec ed,
and i is possible o selec
only one ansc ip ion
ac o a a ime.
PM, s a ic. Las
upda e Jan
2018.
HaploReg 22064851,
26657631
h ps://pubs.
b oadins i u e.
o g/mammals/
haplo eg/haplo eg.
php
I con ains da a om mul iple
genome anno a ion esou ces.
PSSMs a e sco ed agains
e e ence and al e na i e
alleles, and change in log-odds
is calcula ed.
Human Use s can p o ide a lis o
sIDs o ch omosome
egions. Use s can also
selec GWAS s udies om
he NHGRI ca alog.
P o ides da a on allelic
equencies, conse a ion,
ch oma in s a es, and nea genes.
Fo each o he egula o y mo i s
al e ed by he SNP, i p o ides he
change in log-odds and a logo.
HaploReg con ains a
lib a y c ea ed om
li e a u e sou ces,
TRANSFAC, JASPAR and
PBM expe imen s. The e
is no op ion o use -
speci ied ma ices.
PM, s a ic. Las
upda e
No embe
2015.
RegulomeDB 22955989 h p://www.
egulomedb.o g/
RegulomeDB uses in o ma ion
om se e al da ase s, as well as
manual cu a ion and a heu is ic
me hod o dis inguish be ween
unc ional and non- unc ional
a ian s.
Human. Use s can p o ide a lis o
dbSNP IDs, hg19
coo dina es in BED, VCF
o GFF3 o ma , o hg19
ch omosomal egions in
he same o ma s.
Table so ed by likely
unc ionali y, con aining a ian
coo dina es, sco e assigned by he
algo i hm, and e idence o
unc ion including p o ein
binding, mo i s, ch oma in
s uc u e, eQTLs and his one
modi ica ions.
RegulomeDB includes all
PSSMs om TRANSFAC,
JASPAR CORE, and
UniP obe. The e is no
op ion o use -speci ied
ma ices.
PM, s a ic. No
upda es,
lis ed,
so wa e
c ea ed in
Sep 2012.
mo i b eakR 26272984 h ps://
gi hub.com/
Simon-Coe zee/
Mo i B eakR
I has h ee op ions o
algo i hms: he s anda d sum
o log p obabili ies, weigh ed
sum, and an in o ma ion
con en me hod.
O ganisms
included in
BSgenome.
SNPs can be impo ed
om an R package o
p o ided o he algo i hm
in BED o VCF o ma .
PSSMs can be selec ed
om he Mo i Db
package o be use -
speci ied.
Table con aining s a is ics
desc ibing he pe cen o
maximum sco e o a ma ix and
ma ix alues o bo h alleles, as
well as he s and. I also epo s
whe he he TFBS is dis up ed
s ongly o weakly.
PSSMs can be impo ed
om he Mo i Db
package o be use -
speci ied. Mo e han one
ma ix can be used a a
ime.
PM, non-
s a ic.
Las
upda e Jul
2018.
a ia ion-
scan
h p:// sa .eu Es ima ion based on PSSM
model.
web
in e ace:
ins alled
Ensembl
o ganisms.
command-
line: any
locally
ins alled
o ganism.
A collec ion o PSSMs and
a se o a ian s in a Seq
o ma . This o ma can
be ob ained using
e ie e- a ia ion-seq.
A able wi h one line pe pai o
alleles pe mo i (i he e a e mo e
han wo, he e will be one line
pe possible pai ) epo ing he
posi ion, weigh and p- alue o
each allele, weigh di e ence and
p- alue a io.
Use s can selec o he
collec ions a ailable in
RSAT (JASPAR,
HOCOMOCO, CisBP), bu
hey can also use
pe sonal collec ions.
PM-non
s a ic.
Ap il 2019.
1418 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
ily a ailable h ough he web in e ace. Howe e , depending
on he biological ques ion and ela ed po en ial biases, we
ecommend he c ea ion o a dedica ed backg ound model,
which can be done using he RSAT ool c ea e-backg ound,
also a ailable ia he RSAT web in e ace.
2.2. Haplo ype p ocessing
Gene ic a ian s can be de ec ed using high- h oughpu ech-
niques. This has enabled he iden i ica ion o millions o a ian s
in he HapMap [30] and 1000 genomes p ojec s [1]. Howe e , he
in o ma ion on he a ian s alone is less use ul han knowing
which g oups o alleles a e co-loca ed on he same ch omosome
(haplo ype). The p ocess o iden i ying he a ian s ha belong
o each ch omosome is known as phasing. Including haplo ype
phasing in o ma ion acili a es he iden i ica ion o ela ions
be ween a ian s [6].
VCF iles can include haplo ype phasing in o ma ion. The ool
con e - a ia ions iden i ies and e ie es he phasing in o ma ion
o he a ian s, while he ool e ie e- a ia ion-seq buil s he co -
esponding haplo ype wi h all he SNPs ha lay wi hin a de ined
window (de aul : 30 bp).
2.3. Compu ing binding speci ici y o a ansc ip ion ac o o a DNA
sequence
a ia ion-scan uses PSSMs o assess he binding speci ici y o a
TF o a DNA sequence wi h di e en alleles in a gi en posi ion.
The i s s ep o a ia ion-scan (i.e., scanning o he sequences wi h
a gi en PSSM) is delega ed o he RSAT ool ma ix-scan. The sco -
ing scheme and p- alue calcula ion a e desc ibed in de ail in [59],
Box n°1 and Box n°2, espec i ely. In b ie :
PSSM a e used o assess he binding speci ici y o a TF. This
a ini y is calcula ed as a weigh sco e (Ws). The Ws o a si e in
a ia ion-scan is calcula ed using [27]:
Ws ¼lnðPðSjMÞ
PðSjBÞ

whe e S is a sequence segmen o he same leng h o M, M is he
PSSM, and B is he backg ound model. Hence, P(S|M) is he p obabil-
i y o he sequence gi en he PSSM and P(S|B) is he p obabili y o
he sequence gi en he backg ound model. Ws has been ela ed o
he a ini y o he TF o he sequence, as i assesses simila i y o a
sequence o a known se o binding si es, p o iding in o ma ion
abou he p obabili y o a sequence o be a new ins ance o a bind-
ing si e [56].
Mo eo e , i is possible o calcula e he p- alue o a gi en sco e
as:
P
alue ¼PðWwjBÞ
whe e he P- alue is calcula ed as he p obabili y o obse ing a
sco e o a leas Wgi en a backg ound model w|B.
When a sequence is longe han he PSSM, he PSSM is shi ed
base by base un il he ull sequence has been sco ed. This scanning
s ep is pe o med on he sequences o all epo ed alleles, so ha
each allele is compa ed wi h all he posi ions o a gi en mo i .
Backg ound models ep esen he nucleo ide composi ion o a
se o sequences (whole genome, all p omo e sequences, e c.).
These models a e used o es ima e he expec ancy o a nucleo ide
being ound. Backg ound models can ep esen dependency
be ween nucleo ides in sequences (e.g. aking in o accoun he e-
Fig. 1. Schema ic ep esen a ion o Va ia ion- ools: This se o ools, included in he Regula o y Sequence Analysis Tools (RSAT), ocuses on assessing he impac o di e en
allelic a ian s on T ansc ip ion ac o binding si es. A) con e - a ia ions allows use s o inpu hei own a ian s and con e hem o o he o ma s (VCF, GVF and a Bed,
he la e is he o ma used in he nex s ep), while a ia ion-in o e ie es he anno a ed in o ma ion o Ensembl a ian s ins alled in RSAT se e s. B) The ool e ie e-
a ia ion-seq e ie es he su ounding sequence o a ian s (including possible haplo ypes) and gene a es a ex ile wi h one line pe allele and pe a ian o haplo ype
( a Seq o ma ). C) Use s can inpu hei a ian s in a Seq o ma and a collec ion o mo i s (di ec inpu by he use o selec ed om RSAT a ailable collec ions) o a ia ion-
scan; he ool hen scans he co esponding sequences wi h all mo i s and pe o m pai wise compa isons be ween he binding sco es o each ansc ip ion ac o on o all
alleles o a a ian o haplo ype.
W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1419

quencies o dinucleo ides o build a Ma ko model o o de 1 [59]
Box n°3). As backg ound models a e used o calcula e weigh
sco es o a binding si e (P(S|B)), i is impo an o selec an app o-
p ia e model o each analysis. Examples o selec ed backg ound
models a e p esen ed in he di e en s udy cases ma ching each
pa icula biological ques ion.
2.4. Assessmen o allele e ec on ansc ip ion ac o binding
In he second s ep, i.e., e alua ing he impac o SNPs, a ia ion-
scan compa es he ob ained Ws (Ws di e ence = Ws_Allele1 –
Ws_Alelle2) and he P- alue (P- alue a io = P- alue_Allele1/P- al
ue_Allele2) o each o he alleles, posi ion by posi ion h oughou
he scanning window. To e alua e indels, a ia ion-scan compa es
he highes Ws and i s co esponding P- alue o each sequence o
he epo ed alleles. When mo e han wo alleles o a a ian a e
epo ed, all alleles a e compa ed o all alleles in a pai wise
manne .
2.5. a ia ion-scan pe o mance es
2.5.1. Compu ing e iciency
The ools a ia ion-in o and con e - a ia ion a e coded in Pe l,
while e ie e- a ia ion-seq and a ia ion-scan a e coded in C, o
enable he analysis o la ge numbe s o a ian s om euka yo ic
genomes in a easonable ime. To u he imp o e pe o mance,
we educed he da a ans e om he ha d d i e o memo y.
a ia ion-scan pe o mance was assessed by andomly selec ing
a a ian om he 1000 genomes p ojec [1] and a mo i om he
RSAT non- edundan mo i s collec ion [8]. The andomly selec ed
a ian was used o c ea e se s wi h di e en numbe s o epli-
ca es, anging om one housand o nine millions, o es ima e
he ela ion be ween unning ime and he amoun o e alua ed
a ian s. The p ocesses we e un on a Dell Powe Edge C6145 se e
wi h 2 AMD Op e on( m) P ocesso 6386 SE, 16 co es each, P oces-
so speed o 2.8–3.5 Ghz, RAM 256 Gb and wi h an ope a ing sys-
em Cen OS 7 (7.6.1810).
2.5.2. Da ase : expe imen ally-de e mined egula o y a ian s in ed
blood cells
The egula o y ac i i y o 2,756 ed blood cell a ian s has been
sys ema ically measu ed using Massi ely Pa allel Repo e Assays
(MPRA) [60]. MPRA is a high- h oughpu assay in which a lib a y
o pu a i e egula o y elemen s, each ollowed by a unique ba -
code, is inse ed in o a plasmid, hen ans ec ed in o a cell, and
ansc ip s a e hen quan i ied h ough he abundance o ba codes.
These a ian s a e known o be in s ong linkage disequilib ium
(LD) wi h 75 a ian s associa ed wi h common ai s o his cell
ype. Th ee sliding windows pe a ian (le , igh , and cen e )
we e syn hesized, ba coded and used o s udy he e ec o sligh
changes in hei genomic con ex . Following me hods desc ibed
by Uli sch, e al.[60], o each sequence mRNA/DNA a io was com-
pu ed o ob ain a quan i a i e e alua ion o he egula o y e ec o
a sequence a ian .
2.5.3. E alua ion o a ia ion-scan
The a ian da ase was used as inpu o a ia ion-scan; he
a ian s assessed in he ed blood cell assay we e anno a ed wi h
he Ensembl GRCh37 human genome elease, and gi en as inpu
o con e - a ia ions ollowed by e ie e- a ia ion-seq. Since h ee
sliding windows we e used o each a ian in he MPRA, he co -
esponding windows we e me ged be o e compu ing a backg ound
model using he c ea e-backg ound-model ool.
Acco ding o he o iginal s udy [60], binding si es o he ol-
lowing TF we e en iched in he sequences o in e es : GATA1,
KLF1, DHS, TAL1, ETS, FLI1 and AP-1. The e o e, a o al o 48 PSSMs
anno a ed as ela ed o hese TF we e e ie ed om he non-
edundan RSAT mo i collec ion [8], and gi en as inpu o
a ia ion-scan.
A nega i e con ol se o mo i s was c ea ed using he RSAT ool
pe mu e-ma ix [47]; i e pemu ed mo i s we e c ea ed o each o
he 48 mo i s, gene a ing a collec ion o 240 con ol mo i s.
Fo a a ian o be epo ed in a ia ion-scan as posi i e,we
eques ed ha a leas one o he allele sequences was e alua ed
as a binding si es wi h a p- alue o a mos 10
4
(using he pa am-
e e -u h p al 1e-4 in he command line), and ha he p- alue a io
was g ea e o equal o en (a change o one o de o magni ude
be ween he bes and he wo s allele p- alues) (-l h p al_ a io
10).
We compa ed a ia ion-scan o wo o he ools p e iously used
o assess he same se o a ian s by Uli sch, e al. [60]: DeepSea
[50] and [37] del aSVM. In o de o a oid pe sonal biases when cal-
ib a ing ool pa ame e s, we decided o ely on he published ones
[60]. Fo his analysis a ia ion-scan was un wi hou h esholds o
iden i y he impac o he pa ame e s, pa icula ly he h eshold on
p- alue a io.
2.6. Case s udies
2.6.1. Case s udy 1: Iden i ica ion o egula o y a ian s in he
‘‘Pla inum”genomes haplo ypes
The se o high-con idence a ian s om he wo CEU (No he n
Eu opeans om U ah) human Pla inum Genomes NA12877 and
NA12878 [23] we e downloaded h ough he Amazon Web Se ice
(AWS) Command Line In e ace om he Illumina Pla inum Gen-
omes AWS S3 bucke (h ps://gi hub.com/Illumina/Pla -
inumGenomes). The downloaded VCF iles con ained phasing
in o ma ion o each CEU indi idual haplo ype con igu a ion. The
genome e sion used was GRCh37.
We selec ed SNPs in e sec ing wi h he anno a ed DNAseI-seq
clus e ed peaks V3 om he ENCODE p ojec [4]. The VCF ile wi h
he selec ed SNPs was p ocessed using con e - a ia ions wi h he
op ion phased and hen he haplo ype sequences we e econ-
s uc ed wi h e ie e- a ia ion-seq.
Fo a haplo ype SNP se o single posi ion a ian s o be
epo ed in a ia ion-scan, we eques ed ha a leas one o he
sequences was e alua ed as a binding si e wi h a p- alue o a mos
10
4
(-u h p al 1e-4) and ha he p- alue a io be ween he wo
alleles was g ea e o equal o 100 (a change o wo o de s o mag-
ni ude be ween he bes and he wo s alleles p- alues) (-l h
p al_ a io 100). In addi ion, we equi e a change o sign be ween
he bes and wo s sco e as an addi ional il e .
We anno a ed he p edic ed dis up ed TFBS wi h he TF ChIP-
seq non- edundan peak collec ion and wi h he Cis-Regula o y
Modules (CRM) egions om ReMap [10] using bed ools in e sec
e sion 2.27 [49]. We also calcula ed he en ichmen o anno a-
ions in he p o enance sequence segmen s o he p edic ed haplo-
ypes si es.
2.6.2. Case s udy 2: p edic ion o egula o y a ian s associa ed wi h
suscep ibili y o Mycobac e ium ube culosis in ec ion
We collec ed SNPs associa ed wi h he pheno ypic ai ‘‘suscep-
ibili y o Mycobac e ium ube culosis in ec ion measu emen ” (dis-
ease ID EFO_0008407) om he 1.0.2 e sion o he GWAS
ca alog [40] (h ps://www.ebi.ac.uk/gwas/). This que y e u ned
one s udy [58] wi h 67 dis inc a ian s, o which 48 had a alid
e e ence SNP iden i ie ( sID) and could be u he used (deno ed
he ea e as disease-associa ed SNPs, o DA-SNPs). To p edic he
TF binding si es pu a i ely a ec ed by hese selec ed SNPs, we
designed an app oach combining Va ia ion- ools wi h di e en
ex e nal esou ces. We u he collec ed om Ensembl REST in e -
ace (h p:// es .ensembl.o g/) 564 SNPs in linkage disequilib ium
1420 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
(LD-SNPs) in he Eu opean popula ion [62], wi h a h eshold on he
eg ession coe icien (
2
0.8) and a maximal dis ance o 200 bp.
Anno a ions (ch omosomal loca ion, ype o genomic egion) o
he esul ing 612 SNPs (48 DA + 564 LD) we e collec ed om
Ensembl BioMa [22,21]. We hen es ic ed he selec ion o SNPs
in non-coding egions, esul ing in a se o 572 SNPs o in e es
(SOIs) o he de ec ion o egula o y a ian s. Using SNPs in LD,
we de e mined LD-Block egions. These we e hen anno a ed based
on o e laps wi h ChIP-seq peaks collec ed om he ReMap da a-
base [10]. We also calcula ed en ichmen o disease anno a ions
using he R XGR package [24].
Finally, we used e ie e- a ia ion-seq o e ie e he sequence
a ian s a ound each SOI, and p edic ed he impac o he a ia ion
on TF binding o each mo i o he JASPAR non- edundan RSAT
mo i collec ion [8] using a ia ion-scan, wi h he h esholds o
1e-4 on he p- alue and 100 on he p- alue a io.
2.6.3. Case s udy 3: Assessmen o he egula o y e ec o GWAS
epo ed a ian s in p omo e s wi h enhance unc ion
The STARR-seq assay [2] is in i s p inciple simila o he MPRA,
and helps iden i y sel - ansc ibing ac i e egula o y egions ha
ha e enhance po en ial. Using his app oach Dao e al. [17], anal-
ysed he enhance po en ial o anno a ed Re Seq p omo e s [48].In
he wo cell lines K562 and HELA, hey iden i ied 632 and 493 p o-
mo e s wi h enhance po en ial (eP omo e s), espec i ely. Mo e-
o e , he au ho s iden i ied en ichmen o eQTL a ian s epo ed
by GTEx [25].
To iden i y eP omo e s a ian s ha could be a ec ing TF bind-
ing, we e ie ed he GWAS ca alog e sion 1.0 (downloaded on
7/01/19) [40]. Using bed ools o e lap e sion 2.26.0 [49], we com-
pu ed he o e lap be ween SNPs and he eP omo e coo dina es
epo ed in [17]. PSSMs ep esen ing TF en iched in eP omo e s
we e also ob ained om [17], co esponding o SMRC1, JUN, FOS,
ATF:MAF:NEF2, YY1, ETS amily, C eb and USF1/2.
Using he selec ed GWAS a ian s ha all wi hin eP omo e s
and he TF mo i s en iched in hese egions, we applied a ia ion-
scan o assess he po en ial egula o y e ec o hese a ian s.
a ia ion-scan was un wi h he pa ame e s – l h w_di 1 – l h
p al_ a io 10, wi h a backg ound model buil using c ea e-
backg ound wi h all Re Seq p omo e sequences. In o de o il a e
a ian s wi h he highes pu a i e egula o y dis up ion, we u -
he selec ed a ian s ha showed a change o sign in he weigh
sco e be ween alleles.
2.6.4. Case s udy 4: iden i ica ion o egula o y a ian s a ec ing VRN1
binding in ba ley
The la es e sion o Ho deum ulga e (ba ley) e e ence gen-
ome [42] and a panel o mapped gene ic a ian s we e impo ed
om Ensembl Genomes elease 42 [34] and ins alled in he RSAT
Plan s se e (h p://plan s. sa .eu). We ob ained expe imen ally
de e mined binding si es (ChIP-seq) o VRN1 om [19]. Since
hese peaks we e o iginally posi ioned wi hin con igs o he 2012
genome assembly [14], hey had o be ma ched o he co espond-
ing egions o he cu en assembly wi h BLAST + 2.9.0 (blas n)
local alignmen s agains he epea -masked genome sequence
(pe ec ma ches) [7]. Using bed ools o e lap e sion 2.26.0 [49],
we selec ed a ian s alling wi hin he VRN1 epo ed binding
peaks. The selec ed a ian s in VCF o ma we e hen p ocessed
using con e - a ia ions and e ie e- a ia ion-seq o ob ain he
sequences wi h he al e na i e alleles.
The VRN1 DNA mo i used o scan he a ian s was ob ained
om he oo p in DB plan collec ion [16] e sion: 2018-06
(h p:// lo es a.eead.csic.es/ oo p in db/index.php?mo i =
AY750993:VRN1:EEADanno ). a ia ion-scan was used wi h a p e-
compu ed backg ound Ma ko model (o de 1) o ba ley o assess
he e ec o a ian s in TF binding, wi h he ollowing pa ame e s:
– l h sco e 1 – l h w_di 1 – l h p al_ a io 10 – u h p al 1e-3.
2.7. A ailabili y
Va ia ion- ools a e a ailable on he web (Me azoa: h p://me a-
zoa. sa .eu/, Plan s: h p://plan s. sa .eu/, Teaching: h p:// each-
ing. sa .eu/). The ools can be also ins alled o command-line
usage wi h he RSAT sui e (h p://download. sa .eu/).
The code and ma e ial o ep oduce he esul s p esen ed in he
a icle can be accessed h ough Gi Hub (h ps://gi hub.com/RSAT-
doc/supp-ma e ial-publica ions.gi ).
3. Resul s
The Va ia ion- ools p o ide complemen a y p og ams enabling
he e ie al o a ian s ( a ia ion-in o) and o hei su ounding
sequences ( e ie e- a ia ion-seq), as well as in e con e sion
be ween ile o ma s (con e - a ia ion). The main p edic i e p o-
g am is a ia ion-scan, which can be used wi h any se o a ian s
p o ided by he use (in VCF o GVF o ma s) o anno a ed in
Ensembl ( om a lis o sIDs o a bed ile o iden i y o e lapping
a ian s in genome coo dina es), wi h any se o mo i s selec ed
om he collec ions a ailable in RSAT, o p o ided by he use .
3.1. a ia ion-scan accu a ely assesses he e ec o expe imen ally
alida ed egula o y a ian s
The o iginal e sion o a ia ion-scan [45] equi ed app oxi-
ma ely i e hou s o assess he allele e ec o nine millions a i-
an s. The no el e sion [47] signi ican ly educes he p ocessing
ime o abou one hou (Supplemen a y Fig. 3).
To e alua e he pe o mance o a ia ion-scan, we used an
expe imen ally alida ed egula o y a ian se ob ained om a
MPRA expe imen [60]. Fo all o he assessed allele pai s, we com-
pa ed he weigh sco e di e ences compu ed wi h a ia ion-scan
wi h he mRNA/DNA a io o he MPRA (see me hods). As shown
in Supplemen a y Fig. 4A, we a e able o eco e only 9.37% o
he expe imen ally alida ed a ian s wi h a ia ion-scan, as we
eques ed a leas one o he alleles o ha e a binding si e o high
con idence (p- alue 10
4
). Focusing on he a ian s epo ed as
posi i e in he MPRA da a se , we obse ed a weak co ela ion
be ween he weigh di e ence and he MPRA mRNA/DNA a io in
posi i e a ian s. Howe e , his co ela ion is no signi ican , as
MPRA alues do no scale wi h he a ia ion-scan weigh di e -
ences. Ne e heless, all a ian s show a p- alue a io indica i e
o allele binding e ec s, showing ha a ia ion-scan gi es accu a e
measu emen s o he impac o egula o y a ian s (Fig. 2A).
Wi h he p oposed h esholds, we can con iden ly ejec 96.35%
o MPRA nega i e sequences, which could be imp o ed using mo e
es ic i e pa ame e , wi h a concomi an educ ion in ue posi-
i es. No ewo hy, as any high- h oughpu assay, MPRA has i s
limi a ions [51] and sequencing biases could inc ease he numbe
o alse nega i es.
We pe o med a nega i e con ol, consis ing o 240 pe mu ed
ma ices ( i e pe mu ed e sions o he 48 mo i s). Wi h his col-
lec ion, i was s ill possible o eco e a g oup o a ian s, bu i
only ep esen ed 31.2% o he MPRA posi i e a ian s (Supplemen-
a y Fig. 4B).
We compa ed he pe o mance o a ia ion-scan o wo o he
ools ha had been p e iously used by Uli sch, e al [60] o assess
he same se o MPRA a ian s: DeepSea [68] and del aSVM [37].
We decided o use he same pa ame e s in o de o a oid pe sonal
biases when calib a ing he ools. The e o e, aining weigh s o
DNAse I hype sensi i i y si es we e used in he del aSVM analysis.
W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1421
As o DeepSea, he web implemen a ion o he ool was used, wi h
he me ic Func ional Signi icance Sco e.
Tools we e compa ed based on ROC cu es (Fig. 2B), we addi-
ionally an a ia ion-scan using a se o pe mu ed ma ices as neg-
a i e con ol (Fig. 2B, ed line). The h ee ools show e y simila
sensi i i y s speci ici y a he beginning o he cu es, bu only
a ia ion-scan and DeepSea u he emain sepa a ed om he neg-
a i e con ol. As expec ed DeepSea pe o ms sligh ly be e han
a ia ion-scan a he beginning o he cu e, ne e heless his ool
equi es aining using epigene ic da a, while a ia ion-scan
equi es only a mo i and a se o a ian s.
3.2. Va ia ion- ools case s udies
To illus a e he di e se applica ions o Va ia ion- ools o ackle
a ious biological ques ions, we designed ou di e en case
s udies:
1. Impac o egula o y a ian s in he same haplo ype on TF bind-
ing si es.
2. Iden i ica ion o he egula o y po en ial o a ian s epo ed in
GWAS.
3. Assessmen o he egula o y po en ial o GWAS a ian s wi hin
expe imen ally de e mined egula o y egions.
4. De e mina ion o egula o y a ian s wi hin TF binding egions
iden i ied using ChIP-seq [19].
3.2.1. Genome-wide haplo ype a ian in o ma ion can be used o
iden i y se s o egula o y a ian s a ec ing he same TFBS
The lowe ing cos s in sequencing ha e made i possible o
ob ain whole genome sequences o mo e indi iduals, opening he
possibili y o knowing, no only he a ian s o a genome, bu also
he haplo ypes, and de e mining which a ian s a e passed linked
wi hin he same ch omosome. This enables he assessmen o he
egula o y e ec s o se s o a ian s wi hin he same haplo ype
in a gi en TFBS.
Using he high-con idence SNPs om wo ‘‘Pla inum” Genomes
[23], we de e mined haplo ype a ian s ha a e likely o a ec one
TFBS. We selec ed a ian s 30bps apa , loca ed in open ch oma in,
o be analysed wi h a ia ion-scan using he non- edundan mo i
collec ion a RSAT [8]. We de ec ed 7,406 haplo ype si es wi h a
leas wo he e ozygous a ian s and a p obable e ec in binding
o 361 TFs. O e all he numbe o he e ozygous a ian s wi hin a
haplo ype inc eases he measu ed weigh di e ence. This is
expec ed as mo e changes in he binding si es a e mo e likely o
change TF a ini y (Fig. 3A).
To assess he biological ele ance o all he pu a i e dis up ed
TFBS p edic ions, we anno a ed 7,485 p edic ed haplo ypes si es
con aining wo o mo e a ian s wi h a leas one he e ozygous
a ian and 15,396 p edic ed si es con aining a SNP (single ons)
wi h he TF ChIP-seq peaks and he Cis-Regula o y Modules
(CRM) egions om ReMap [10]. We ound ha almos all he p e-
dic ed dis up ed TFBS (~85%) con ain a CRM o peak anno a ion o
bo h (Fig. 3B). In e es ingly, we ound en ichmen o CRM and peak
anno a ions in he p o enance sequence segmen s o he 7,485
p edic ed haplo ypes si es compa ed o he p o enance sequence
segmen s o he single a ian s (Fishe exac es , p- alue < 2.2e-
16).
One o hese anno a ed haplo ypes is composed o he mino
alleles o wo SNPs ( s2732317 and s2732318), whe e we
obse ed a po en ial egula o y e ec likely a ec ing h ee binding
mo i s, o EHF/ELF2, ETV4/ELK1/ETS1/FLI1/ELK4/ETS2/FEV/GABP1,
and ELK3/ELF1/ERG/GABPA (Fig. 3C).
3.2.2. Gene ic a ian s associa ed wi h Mycobac e ium ube culosis
in ec ion show po en ial egula o y e ec s
The second case s udy illus a es a knowledge- ee use o
Va ia ion- ools o iden i y egula o y a ian s om GWAS s udies
o a use -speci ied disease, wi hou p io indica ion abou he
po en ially in ol ed ansc ip ion ac o s o binding mo i s. The
app oach is based on he p edic ion o egula o y a ian s wi h
RSAT Va ia ion- ools, na owed down by selec ing he egula o y
SNPs ha o e lap ChIP-seq peaks in ReMap [10], in o de o iden-
i y con e gen indica ions o a po en ial impac o he a ian s on
he binding o a TF.
A) B)
0.00
0.25
0.50
0.75
1.00
0.000.250.500.751.00 speci ici y
sensi i i y
name
pe mu ed
del aSVM
a ia ion-sca
n
DeepSea
P- alue a io 100
P- alue a io 1000
P- alue a io 10
R=0.12,p=0.5
0
200
400
−9 −6 −3 0
MPRA p− alue o di e en ial ac i i y (log10)
Va ia ion−scan p− alue a io
Fig. 2. Iden i ica ion o expe imen ally alida ed egula o y a ian s using a ia ion-scan. A) Co ela ion o he Massi ely Pa allel Repo e Assays (MPRA) p- alue o he
mRNA/DNA a io o posi i e a ian s and he a ia ion-scan weigh di e ence o he MPRA a ian s wi h signi ican change. B) Recei e Ope a ing Cha ac e is ic (ROC) cu e
compa ing he pe o mance when aiming o classi y MPRA expe imen ally analyzed a ian s using a ia ion-scan ( u quoise), DeepSea (pu ple), del aSVM (g een), and a
nega i e con ol which consis s o pe mu ed mo i s sco ed wi h a ia ion-scan ( ed).
1422 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
Fig. 3. Haplo ype analysis in high-quali y human genomes. A) The numbe o he e ozygous a ian s (X-axis) wi hin he same pu a i e binding si e end o ha e a g ea e
impac on he TF binding p obabili y. This is expec ed as he inc ease o weigh di e ence obse ed on he iolin plo co esponds o he expec ed cumula ed impac o
a ia ions a ec ing di e en posi ions o he same binding si e. B) Numbe o p edic ed dis up ed T ansc ip ion Fac o Binding Si es (TFBSs) wi h Cis-Regula o y Modules
(CRMs) and TF ChIP-seq peak anno a ion (blue), wi h only peak anno a ion (yellow), and non-anno a ed p edic ions (g ey). C) Uni e si y o Cali o nia San a C uz (UCSC)
b owse [48] sc een sho , showing a locus encompassing wo SNPs ha compose an he e ozygous haplo ype in one o he No he n Eu opeans om U ah (CEU) indi iduals.
The igu e shows he e e ence genome haplo ype. The a ian s a e loca ed in he FUT10 p omo e ( op). a ia ion-scan p edic s an e ec in h ee mo i s ha ep esen
binding si es o GABPA, ETS1 and ELF2, ac o s ha ha e been p o en o ha e binding si es in his egion by he ENCODE p ojec . The a ian s2732317 has been associa ed
wi h e ec s in gene exp ession by he GTEx p ojec .
W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1423