A su ey ongene a i e ad e sa ial ne wo ks
o imbalance p oblems incompu e ision
asks
Vignesh Sampa h1,2*, Iñaki Mau ua1, Juan José Aguila Ma ín2 and Ai o Gu ie ez1
In oduc ion
Recen de elopmen s in Con olu ional Neu al Ne wo ks (Con Ne s) ha e led o sub-
s an ial p og ess in he pe o mance o compu e ision asks applied ac oss a i-
ous domains such as sel -d i ing ca s [1], medical imaging [2], ag icul u e [3, 4],
Abs ac
Any compu e ision applica ion de elopmen s a s off by acqui ing images and
da a, hen p ep ocessing and pa e n ecogni ion s eps o pe o m a ask. When he
acqui ed images a e highly imbalanced and no adequa e, he desi ed ask may no be
achie able. Un o una ely, he occu ence o imbalance p oblems in acqui ed image
da ase s in ce ain complex eal-wo ld p oblems such as anomaly de ec ion, emo ion
ecogni ion, medical image analysis, aud de ec ion, me allic su ace de ec de ec ion,
disas e p edic ion, e c., a e ine i able. The pe o mance o compu e ision algo i hms
can significan ly de e io a e when he aining da ase is imbalanced. In ecen yea s,
Gene a i e Ad e sa ial Neu al Ne wo ks (GANs) ha e gained immense a en ion by
esea che s ac oss a a ie y o applica ion domains due o hei capabili y o model
complex eal-wo ld image da a. I is pa icula ly impo an ha GANs can no only be
used o gene a e syn he ic images, bu also i s ascina ing ad e sa ial lea ning idea
showed good po en ial in es o ing balance in imbalanced da ase s.
In his pape , we examine he mos ecen de elopmen s o GANs based echniques
o add essing imbalance p oblems in image da a. The eal-wo ld challenges and
implemen a ions o syn he ic image gene a ion based on GANs a e ex ensi ely co -
e ed in his su ey. Ou su ey fi s in oduces a ious imbalance p oblems in compu e
ision asks and i s exis ing solu ions, and hen examines key concep s such as deep
gene a i e image models and GANs. A e ha , we p opose a axonomy o summa ize
GANs based echniques o add essing imbalance p oblems in compu e ision asks
in o h ee majo ca ego ies: 1. Image le el imbalances in classifica ion, 2. objec le el
imbalances in objec de ec ion and 3. pixel le el imbalances in segmen a ion asks. We
elabo a e he imbalance p oblems o each g oup, and p o ide GANs based solu ions
in each g oup. Reade s will unde s and how GANs based echniques can handle he
p oblem o imbalances and boos pe o mance o he compu e ision algo i hms.
Keywo ds: Gene a i e ad e sa ial neu al ne wo ks, Imbalanced da a, Objec
de ec ion, Segmen a ion, Classifica ion, Deep lea ning, Deep gene a i e model
Open Access
© The Au ho (s) 2021. This a icle is licensed unde a C ea i e Commons A ibu ion 4.0 In e na ional License, which pe mi s use, sha ing,
adap a ion, dis ibu ion and ep oduc ion in any medium o o ma , as long as you gi e app op ia e c edi o he o iginal au ho (s) and
he sou ce, p o ide a link o he C ea i e Commons licence, and indica e i changes we e made. The images o o he hi d pa y ma e ial
in his a icle a e included in he a icle’s C ea i e Commons licence, unless indica ed o he wise in a c edi line o he ma e ial. I ma e ial
is no included in he a icle’s C ea i e Commons licence and you in ended use is no pe mi ed by s a u o y egula ion o exceeds he
pe mi ed use, you will need o ob ain pe mission di ec ly om he copy igh holde . To iew a copy o his licence, isi h p://c ea i eco
mmons .o g/licen ses/by/4.0/.
SURVEY PAPER
Sampa he al. J Big Da a (2021) 8:27
h ps://doi.o g/10.1186/s40537-021-00414-0
*Co espondence:
ignesh.sampa h@ eknike .es
1 Au onomous
and In elligen Sys ems Uni ,
Teknike , Membe o Basque
Resea ch and Technology
Alliance, Eiba , Spain
Full lis o au ho in o ma ion
is a ailable a he end o he
a icle
Page 2 o 59
Sampa he al. J Big Da a (2021) 8:27
manu ac u ing [5], e c. The a ailabili y o big da a [6], oge he wi h inc eased compu -
ing capabili ies is he p edominan eason o he ecen success. Image acquisi ion is he
fi s s ep in he de elopmen o compu e ision algo i hms. When he acqui ed image
is no adequa e, he desi ed ask may no be possible o achie e. Image classifica ion [7],
objec de ec ion [8] and segmen a ion [9] a e he undamen al building blocks o he
compu e ision asks. All hese me hods use deep Con Ne s wi h eno mous laye s and
ha e a e y high numbe o pa ame e s ha need o be uned. The e o e, hey demand
a huge amoun o ep esen a i e da a o imp o e hei pe o mance and gene aliza ion
abili y. While he amoun o isual da a is inc easing exponen ially, many o he eal-
wo ld da ase s suffe om se e al o ms o imbalance. Handling imbalances in he image
da ase is one o he pe asi e challenges in he field o compu e ision.
Image classifica ion is he ask o classi ying an inpu image acco ding o a se o pos-
sible classes. Classifica ion algo i hms lea n o isola e impo an dis inguishing in o -
ma ion abou an objec in an image like shape o colo and igno e i ele an pa s o
an image such as plane backg ound o noise. Se e al popula image classifica ion a chi-
ec u es such as LeNe [7], AlexNe [10], VGG-16 [11], GoogLeNe [12], ResNe [13],
Incep ion-V3 [14], DenseNe [15] ake an inpu image and hen pass i h ough se e al
con olu ional and pooling laye s. Con olu ional laye helps o ex ac ea u es om he
inpu image, while a pooling laye educes he dimension. Se e al successi e con olu-
ional and pooling laye s may ollow, depending on he layou and in en o he a chi-
ec u e. The esul is a se o ea u e maps educed in size om he o iginal image ha
h ough a aining p ocess ha e lea ned o dis ill in o ma ion abou he con en in he
o iginal image. All ex ac ed ea u e maps a e hen ans o med in o a single ec o ha
can be ed in o a se ies o ully connec ed neu al ne wo k o ob ain a p obabili y dis i-
bu ion o class sco es. The p edic ed class o he inpu image can be ex ac ed om his
p obabili y dis ibu ion.
These a chi ec u es a e ypically designed o wo k well wi h balanced da ase s, bu a
common issue wi h eal-wo ld da ase s is he imbalance o obse ed classes. The mos
commonly known imbalance p oblem in a ask o image classifica ion is he class imbal-
ance. Class imbalance in he eal-wo ld image da ase s is ubiqui ous and can ha e an
ad e se effec on he pe o mance o Con Ne s [16]. These da ase s usually all in o ou
ca ego ies in e ms o i s size and imbalance [17]:
1. The ideal da ase s a e he one ha con ain an adequa e and equal o almos equal
numbe o samples wi hin each class. An equal p obabili y is assigned o all classes
du ing aining o upda e pa ame e s o he ne wo k and app oach he minimum
alue o he e o unc ion. A wide ange o s anda d machine lea ning algo i hms
can be applied o he ideal da ase s.
2. The da ase s wi h an adequa e numbe o samples whe e some ins ances o classes
a e a e han o he ins ances o classes a e said o be une en da ase s. E en hough
hese da ase s ha e adequa e numbe o samples, i is cos ly and may no be possible
o expe s o manually inspec huge unlabeled da ase s o anno a e.
3. Tiny da ase s a e no easily a ailable, and hey can be difficul o collec . Such da a-
se s ha e an equal numbe o samples wi hin each class, bu hey a e almos impos-
sible o collec due o p i acy es ic ion and o he easons.
Page 3 o 59
Sampa he al. J Big Da a (2021) 8:27
4. Absolu e a e da ase s ha e a limi ed numbe o samples and subs an ial class imbal-
ance. Reasons o class imbalance in hese da ase s can a y bu commonly he p ob-
lem a ises because o : (a) Ve y limi ed numbe o expe s a ailable o da a collec-
ion; o an example, gene a ion o medical imaging da ase s equi es specialized
equipmen and well ained medical p ac i ione s o da a acquisi ion (b) Eno mous
manual effo equi ed o label da ase s; and (c) Sca ci y o samples o specific class
leading o class imbalance. Consequen ly, he size o he da ase and class imbal-
ance p oblem becomes a bo leneck ha p e en s us om apping he ue po en ial
o Con Ne s. Figu e1 illus a es diffe en ypes o da ase s in e ms o i s size and
imbalance.
Class imbalance in a da ase can s em om ei he be ween classes (in e class imbal-
ance) o wi hin class (in a class imbalance). In e class imbalance occu s when a mino -
i y class con ains a smalle numbe o ins ances when compa ed o ins ances belonging
o he majo i y class. Classifie s buil using in e class imbalanced da ase s a e mos
likely o p edic mino i y class as a e occu ences, e en some imes assumed as ou lie
o noise which esul s in misclassifica ion o mino i y classes [18]. Mino i y classes a e
o en o g ea e in e es and significance, ha needs o be cau iously handled. Fo exam-
ple, in a a e disease medical diagnosis whe e he e is a i al need o dis inguish such a
a e medical condi ion among he no mal popula ions. Any kind o diagnosis e o s will
cause s ess o he pa ien and u he complica ions. I is he e o e e y impo an ha
deep lea ning models [19] buil using such da ase s should be able o achie e a highe
de ec ion a e on mino i y classes.
In a class imbalance in a da ase can also de e io a e he pe o mance o he classi-
fie . An In a-class imbalance can be iewed as he a ibu e bias wi hin a class, in o he
wo ds in e -class imbalance in fine-g ained isual ca ego iza ion. Fo example, a class
o dog samples can be u he ca ego ized by dog colo , pose a ia ions and dog b eeds.
Imbalances in such ca ego ies (in a class imbalance) is an una oidable p oblem in
Fig. 1 Dis ibu ion o diffe en ype o da ase s (a) Da ase wi h adequa e sample (b) Da ase wi h
inadequa e sample
Page 4 o 59
Sampa he al. J Big Da a (2021) 8:27
da ase s o many classifica ion asks such as modali y based medical image classifica ion
[19], fine g ained a ibu e classifica ion [20], pe son e-iden ifica ion [21], age [22] and
pose in a ian ace ecogni ion [23].
Se e al a emp s ha e been made o o e come he p oblem o class imbalance by
using diffe en app oaches and echniques. These echniques can be g ouped in o
da a-le el app oaches, algo i hm le el me hods and hyb id echniques. While da a
le el app oaches modi y he dis ibu ion o aining se o es o e balance by adding o
emo ing ins ances om he aining da ase , algo i hm le el me hods change he objec-
i e unc ion o he classifie o inc ease he impo ance o he mino i y class. Hyb id
echniques combine algo i hm le el me hods wi h da a le el app oaches. Nex ew pa a-
g aphs will in o m eade s abou some o he adi ional echniques a ailable o coun e
he class imbalance p oblem.
• Resampling To coun e ac he class imbalance p oblem, wo ypes o e-sampling
can be applied: One is unde sampling by dele ing samples om he majo i y class
and ano he is o e sampling by duplica ing samples om he mino i y class [24].
Re-sampling me hod balances he da ase bu ails o p o ide any addi ional in o -
ma ion o he aining se . The o he limi a ions o his me hod include: o e sam-
pling esul s in o e fi ing p oblem while unde sampling leads o subs an ial loss
o in o ma ion [25]. The quan i y o unde -sampling and o e sampling is gene ally
de e mined using expe imen al me hods and empi ically es ablished [26]. In o de o
yield addi ional in o ma ion o he aining se , syn he ic o e sampling me hods c e-
a e new samples ins ead o duplica es o add equilib ium o skewed dis ibu ion. The
Syn he ic Mino i y O e sampling Technique (SMOTE) [27] is a popula syn he ic
o e sampling me hod ha aims o gene a e syn he ic samples based on andomly
selec ed K-nea es neighbo s. SMOTE does no ake accoun o he dis ibu ion o
da a be ween he classes. Adap i e syn he ic sampling (ADASYN) app oach [28]
uses a weigh ed dis ibu ion o diffe en mino i y classes acco ding o hei lea ning
difficul ies o adap i ely gene a e syn he ic da a samples. Clus e based o e sampling
[29] echnique di ides he inpu space in o a ious clus e s and hen inco po a es
sampling o al e he sample size. Many adi ional syn he ic o e sampling ech-
niques such as SMOTE o ADASYN a e only sui able o low dimensional abula
da a which es ic s hei applica ion in a high dimensional image da a. In addi ion,
all he a o emen ioned echniques gene a e da a by ei he dele ing o a e aging exis -
ing da a, and hence may ail o imp o e classifica ion pe o mance.
• Augmen a i e o e sampling Da a augmen a ion is ano he commonly used ech-
nique o infla e he size o he aining da ase [30]. Augmen a ion such as ansla-
ion, c opping, padding, o a ion and ho izon al flipping in oduces small modifica-
ions in he image da a, bu no all hese modifica ions will imp o e he pe o mance
o a classifie . The e is no s anda d me hod ha can decide whe he any pa icula
augmen a ion s a egy can imp o e esul s un il he aining p ocess is comple e.
As aining Con Ne s is a ime-consuming p ocess [31], only a es ic ed amoun
o augmen a ion s a egy is likely o be es ed be o e model deploymen . Also, he
di e si y ha can be ob ained om small modifica ions o he images is ela i ely
small. In addi ion o balancing classes by o e sampling, augmen a ion echniques
Page 5 o 59
Sampa he al. J Big Da a (2021) 8:27
also se e as a kind o egula iza ion in deep neu al ne wo k a chi ec u e and hence
educe he chance o o e fi ing. The e is no consensus abou he bes s a egy o
combining diffe en augmen a ion s a egies oge he . The e o e, mo e ad anced
augmen a ion echniques such as mixing images depend on expe knowledge o
alida ion and labelling [32]. A comple e su ey o Image da a augmen a ion o deep
lea ning has been compiled by Sho en e al. [32].
• Semi-supe ised lea ning (SSL) SSL [33] is one o he mos a ac i e ways o imp o e
classifica ion pe o mance whe e we ha e access o small numbe o labeled sam-
ples
x
x
along wi h la ge amoun o unlabeled samples (Une en da ase ). SSL uses
he combina ion o supe ised and unsupe ised lea ning echniques. I makes use
o small labeled samples as he aining se o ain he model in a supe ised man-
ne , and hen use he ained model o p edic on he emaining unlabeled po ion
o he da ase . The p ocess o labeling each sample o unlabeled da a wi h he indi-
idual ou pu s p edic ed o hem using he ained model is known as pseudo labe-
ling. A e labeling he unlabeled da a h ough he pseudo labeling p ocess, classifica-
ion model is ained on bo h he ac ual and pseudo labeled da a. Pseudo labeling is
an in e es ing pa adigm o anno a e la ge-scale unlabeled da a ha po en ially akes
many edious hou s o human labo o manually label hem. Howe e , SSL elies on
assump ions abou he unde lying ma ginal dis ibu ion o inpu da a
p
(x) , bo h he
labeled and unlabeled samples a e assumed o ha e he same ma ginal dis ibu ion.
This ma ginal dis ibu ion
p
(x) should con ain in o ma ion abou he pos e io dis-
ibu ion p(
y
|x) . A comple e lis o semi supe ised lea ning is de ailed in [34].
• Cos sensi i e lea ning Majo i y o he classifica ion algo i hms assume ha misclas-
sifica ion cos s o bo h mino i y and majo i y classes a e he same. Cos -sensi i e
lea ning [35] pays mo e a en ion o misclassifica ion cos s o he mino i y class
h ough a cos ma ix.
The mos s aigh o wa d and commonly used app oach in Con Ne s is he da a
d i en s a egy, because deep Con Ne s wi h eno mous laye s ha e a e y high num-
be o pa ame e s o be uned, i is p one o o e fi ing when ained on a small sized
da ase . Da a le el app oaches infla e he aining da a size ha se es as egula iza ion
and hence educe he chance o o e fi ing in deep neu al ne wo k a chi ec u e. T adi-
ional da a-le el echniques suffe he ollowing d awbacks, pa icula ly when used o
he class imbalance p oblem in high-dimensional image da a.
a. Syn he ic ins ances c ea ed using adi ional da a le el app oaches may no be he
ue ep esen a i e o he aining se .
b. Syn he ic da a gene a ion is achie ed ei he by duplica ion o linea in e pola ion
which does no gene a e new examples ha a e a ypical and puzzle he classifie
decision bounda ies, and hence ail o imp o e o e all pe o mance.
c. In Medical images, augmen a ion echniques a e es ic ed o mino al e a ion on
an image, as hey abide by s ic s anda ds. Addi ionally, he ypes o augmen a ion
one can use a y om p oblem o p oblem. Fo ins ance, hea y augmen a ions such
as geome ic ans o ma ions, andom e asing, and mixing images migh damage
seman ic con en o he medical image.
Page 6 o 59
Sampa he al. J Big Da a (2021) 8:27
d. Applying da a augmen a ion in an absolu e a e da ase may no p o ide he a ia-
ions equi ed o p oduce a dis inc sample o add equilib ium o skewed dis ibu-
ion.
e. Dealing wi h he class imbalance in fine-g ained isual ca ego iza ion is challenging
because i in ol es la ge in a-class a iabili y and small in e -class a iabili y.
. Mos o he echniques a e designed only o bina y classifica ion p oblems. Mul i
class imbalance p oblems a e gene ally conside ed much ha de han hei bina y
equi alen s o many easons. Fo Ins ance, he e can be se e al combina ions o
mino i y-majo i y classes, i.e., hey may include: 1. Few mino i y-Many majo i y
classes, 2. Many mino i y-Few majo i y classes, and 3. Many mino i y-Many majo -
i y classes.
Class imbalance in image classifica ion asks has been widely explo ed and s udied.
In addi ion o class imbalance, he e a e many diffe en o ms o imbalances ha can
impede pe o mance o o he compu e ision asks such as objec de ec ion and image
segmen a ion. Objec de ec ion, which deals wi h localiza ion and classifica ion o mul-
iple objec s in a gi en image, is ano he challenging and significan ask in compu e
ision. The ypical way o localizing an objec in an image is by d awing a bounding box
a ound he objec . This bounding box can be in e p e ed as a collec ion o coo dina es
ha define he box. Nowadays, objec de ec ion algo i hms all in o wo b oad ca ego-
ies: wo-s age de ec o s and single s age de ec o s. On one hand, wo s age de ec o
such as Region-based Con olu ional Neu al Ne wo ks (R-CNN) [8], Fas R-CNN [36],
Fas e R-CNN [37], Mask R-CNN [38], e c. employ a Region P oposal Ne wo k (RPN) o
sea ch objec s in he fi s s age, and hen p ocess hese egion o in e es s o objec clas-
sifica ion and bounding-box eg ession in he second s age. On he o he hand, single
s age de ec o s such as Single Sho De ec ion (SSD) [39], You Only Look Once (YOLO)
[40], e c. pe o m de ec ion on a g id ha a oids spending oo much ime on gene a ing
egion p oposals. Ins ead o loca ing objec s pe ec ly, hey p io i ize speed and ecogni-
ion. The e o e, one s age objec de ec o s a e as and simple, whe eas wo s age de ec-
o s a e mo e accu a e.
Despi e he ecen ad ances, applying objec de ec ion algo i hms o he eal-wo ld
da ase s such as in-ca ideo [41], anspo a ion su eillance images [42] ha con ain
objec s wi h la ge a iance o scales (Objec s scale imbalance) emains challenging.
Physical size o a same objec a diffe en dis ances om he came a would appea as
diffe en size. Singh e al. [43] showed ha objec le el scale a ia ion g ea ly affec s he
o e all pe o mance o objec de ec o s. Many solu ions ha e been p oposed o add ess
he objec scale imbalance. Scale awa e as R-CNN [44] uses an ensemble o wo objec
de ec o s, one o de ec ing he la ge and medium scale objec s and o he o he small
scale objec s, and hen combines hem o p oduce final p edic ions. Mul i-scale Image
Py amids such as SNIP [43] and SNIPER [45] use an image py amid o build mul i scale
ea u e ep esen a ion. Fea u e Py amid Ne wo ks (FPN) [46] combine ea u e hie a -
chies a diffe en scales o p edic objec s a diffe en scales.
Objec s in he eal-wo ld da ase s only occupy a small po ion o he image, while he
es o he image is backg ound. Bo h single and wo s age algo i hms app oxima ely
e alua e abou 104 o 105 loca ions pe image [47], ye jus a ew loca ions ha e objec s.
Page 7 o 59
Sampa he al. J Big Da a (2021) 8:27
The imbalance be ween o eg ound (objec ) and backg ound can also hinde pe o -
mance o he objec de ec ion algo i hm. Fu he mo e, objec de ec ion algo i hms
should be in a ian o de o ma ion and occluded objec s. In Pedes ian de ec ion Da a-
se [48], o ins ance, mo e han 70% o pedes ians a e occluded in a leas one ame o
a ideo clip and abou 19% o pedes ians a e occluded in all ames, whe e he occlu-
sions a e anked as hea y in almos hal o such cases. Dolla e al. [48] highligh ha
he pe o mance o pedes ian de ec ion using s anda d de ec o s declines subs an ially
e en unde pa ial occlusion, and d as ically unde se e e occlusion. Da a augmen a ion
based on andom e asing [49] is a equen ly used echnique ha o ces de ec o s o pay
a en ion o he en i e objec in an image, a he han jus a po ion o i . Ye , his ech-
nique is no gua an eed o be ad an ageous in all he condi ions. Because skewed dis i-
bu ions a ise e en wi hin de o med and occluded objec s as some o he occlusions and
de o ma ions a e uncommon ha hey ha dly occu in p ac ical scena ios [50].
Image segmen a ion ha classifies e e y pixel in an image suffe s om pixel le el
imbalances, as a e o he compu e ision asks.Some o he well-known image segmen-
a ion algo i hms include Fully connec ed ne wo k [9], SegNe [51], U-Ne [52], ResU-
Ne [53] e c. Image segmen a ion is essen ial o a a ie y o asks, including: U ban
scene segmen a ion o au onomous d i ing [54], indus ial inspec ion [55] and cance
cell segmen a ion [56]. Da ase s o all hese asks suffe om pixel le el imbalance. Fo
example, In U ban s ee scene da ase [57], Pixels co esponding o sky, building and
oad a e a nume ous han pixels o pedes ian and bicyclis . This is due o he ac ha
he a ea co e ed by sky, buildings and oads a e mo e han pedes ians and bicyclis s in
he image. Simila ly, In b ain umou image segmen a ion da ase [58], MRI images ha e
mo e heal hy b ain issue pixels han cance ous issue pixels. The mos equen ly used
loss unc ion o image segmen a ion ask is a pixel wise c oss en opy loss [59]. This loss
assigns equal weigh s o all he pixels, e alua es he p edic ion o each pixel indi idually
and hen a e ages o e all pixels. In o de o mi iga e his p oblem, many wo ksha e
been done which modi y he pixel wise c oss en opy loss unc ion. The s anda d c oss
en opy loss is modified in Weigh ed c oss en opy [52], Focal loss [47], Dice Loss [60],
Gene alised Dice Loss [61], T e sky loss [62], Lo ász-So max [63] and Median e-
quency balancing [51], so as o assign highe impo ance o a e pixels. Al hough modi-
fied loss unc ions a e efficien o some imbalances, such unc ions unde go se e e
difficul ies when i comes o highly imbalanced da ase s, as seen wi h medical image
segmen a ions.
In con as o all he adi ional app oaches desc ibed abo e, Gene a i e ad e sa ial
Neu al Ne wo ks (GANs) aim o lea n unde lying ue da a dis ibu ions om he lim-
i ed a ailable images (bo h mino i y and majo i y class), and hen use he lea ned dis-
ibu ions o gene a e syn he ic images. This aises an in e es ing ques ion on whe he
GANs can be used o gene a e syn he ic images o he mino i y class o a ious imbal-
anced da ase s. Indeed, ecen de elopmen s o GANs sugges ha being capable o
ep esen complex and high dimensional da a can be used as a me hod o in elligen
o e sampling. GANs u ilize he abili y o neu al ne wo ks o lea n a unc ion ha
can app oxima e model dis ibu ion as close as possible o ue dis ibu ion. Pa icu-
la ly, hey do no ely on p io assump ions abou he da a dis ibu ion and can gene -
a e syn he ic images wi h high isual fideli y. This significan p ope y allows GANs o
Page 8 o 59
Sampa he al. J Big Da a (2021) 8:27
be applied o any kind o imbalance p oblem in compu e ision asks. GANs can no
only be able o gene a e a ake image, bu also offe a way o change some hing abou
he o iginal image. In o he wo ds, hey can lea n o p oduce any desi ed numbe o
classes (such as, objec s, iden i ies, people, e c.), and ac oss many a ia ions (such as,
iewpoin s, ligh condi ions, scale, backg ounds, and mo e). The e a e a wide a ie y o
GANs epo ed in he li e a u e, each wi h hei own s eng hs o alle ia e imbalance
p oblem in compu e ision asks. Fo ins ance, A GAN [64], IcGAN [65], ResA -
GAN [66], e c. a e a specific a ian o GANs ha a e commonly used o acial a ibu e
edi ing asks. They lea n o syn hesize no only a new ace image wi h desi ed a ibu es
bu also p ese es a ibu e independen de ails. Recen ly, GANs ha e been combined
wi h a wide ange o exis ing objec de ec ion and image segmen a ion algo i hms o
o e come he p oblem o imbalance and imp o e hei pe o mance.
The o iginal GANs a chi ec u e [67] con ains wo diffe en iable unc ions ep esen ed
by wo ne wo ks, a gene a o
G
and a disc imina o
D
. The lea ning p ocedu e o GANs
is o simul aneously ain a disc imina o
D
and a gene a o
G
. I ollows an ad e sa ial
wo-playe , ze o-sum game. An in ui i e way o unde s anding GAN is wi h he police
and he coun e ei e anecdo e. The gene a o ne wo k is like a g oup o coun e ei e s
ying o p oduce ake money and make i look genuine. The police a emp o disco e
coun e ei e s using ake money, ye a he same ime need o le e e y o he pe son
spend hei eal money. O e ime, he police show signs o imp o emen a iden i ying
ake cash, and he o ge s imp o e a aking i . In he end, he coun e ei e s a e com-
pelled o make ideal copies o eal money. High esolu ion and ealis ic mino i y class
images gene a ed using lea ned model dis ibu ion can be used o balance he class dis-
ibu ion and mi iga ing effec o o e fi ing by infla ing he aining da ase size. GANs
sol e he p oblem o gene a ing da a when he e is no enough da a o begin wi h and
hey equi e no human supe ision. GANs can p o ide an efficien way o fill in holes
in he disc e e dis ibu ion o aining da a. In o he wo ds, hey can ans o m he dis-
c e e dis ibu ion o aining da a o con inuous, p o iding an addi ional da a by non-
linea in e pola ion be ween he disc e e poin s. Bowles e al. [68] a gues ha GANs
offe an access o unlock addi ional in o ma ion om a da ase . In ac , Yann LeCun, he
acebook ice p esiden and chie AI scien is , e e ed o GANs as " he mos in e es ing
hing ha has happened o he field o machine lea ning in he las 10yea s".
In his su ey, as opposed o o he ela ed su eys on class imbalance, ha p esen
class imbalance in abula da a, we ocus on wide ange o imbalance in high dimen-
sional image da a by ollowing a sys ema ic app oach wi h a iew o help esea che s
es ablish a de ailed unde s anding o GAN based syn he ic image gene a ion o he
imbalance p oblems in compu e ision asks. Fu he mo e, ou su ey co e s imbal-
ances in a wide ange o compu e ision asks in con as o o he su eys ha a e lim-
i ed o image classifica ion asks.
The key con ibu ions o his su ey a e p esen ed as ollows:
• In his su ey pape , we e iew cu en esea ch wo k on GAN based syn he ic
image gene a ion o he imbalance p oblems in isual ecogni ion asks spanning
om 2014 o 2020. We g oup hese imbalance p oblems in a axonomic ee wi h
h ee main g oups: Classifica ion, Objec de ec ion and Segmen a ion (Fig.2).
Page 9 o 59
Sampa he al. J Big Da a (2021) 8:27
• Also, we p o ide necessa y ma e ial o in o m esea ch communi ies abou he
la es de elopmen and essen ial echnical componen s in he field o GAN based
syn he ic image gene a ion.
• Apa om analyzing diffe en GAN a chi ec u es, ou su ey ocuses hea ily on
eal wo ld applica ions whe e GAN based syn he ic images a e used o alle ia e
imbalances and fills a esea ch gap in he use o syn he ic images o he imbalance
p oblems in isual ecogni ion asks.
The emainde o his pape is o ganized as ollows: “Deep Gene a i e image mod-
els” sec ion gi es eade s necessa y backg ound in o ma ion on gene a i e models.
“Gene a i e ad e sa ial Neu al Ne wo k” sec ion discusses selec ed GAN a ian s
om he a chi ec u e, algo i hm, and aining icks pe spec i e in de ail. In “Taxon-
omy o class imbalance in isual ecogni ion asks” sec ion, we p o ide a b ie expla-
na ion on a ious ypes o imbalances encoun e ed in isual ecogni ion asks and
how he GAN based syn he ic image is used o ebalance, ollowed by GAN a ian s
om he applica ion pe spec i e. “Discussion and Fu u e wo k” sec ion iden ifies and
enume a es ou pe spec i e and possible u u e esea ch di ec ion. Finally, we con-
clude he pape in “Conclusion” sec ion.
Fig. 2 P oposed axonomy o he e iew o imbalanced p oblem in compu e ision asks
Page 16 o 59
Sampa he al. J Big Da a (2021) 8:27
The defini ion o Jensen-Shannon di e gence (
DJ
S
) be ween wo p obabili y dis i-
bu ions p
g
(x
)
and
p
(x
)
is defined as
The e o e, Eq.(10) is equal o
Essen ially, he loss o he gene a o
G
minimizes he Jensen-Shannon di e gence
be ween he gene a ed da a dis ibu ion pg(x
)
and he eal da a dis ibu ion p (x
)
when disc imina o
D
is op imal. Jensen-Shannon di e gence isasmoo h,symme -
ic e siono heKLdi e gence. Husza [110] belie es ha he main eason behind
he g ea success o GANs is eplacing asymme ic KL di e gence loss unc ion in he
classical app oach o symme ic JS di e gence.
Mean squa ed e o used in la en a iable models such as au oencode , a e ages all
he possible ea u es in an image and gene a e blu y images. In con as , ad e sa ial
loss p ese es he ea u es using disc imina o ne wo ks ha de ec an absence o any
ea u es as an un ealis ic image. An example o his is he s udy ca ied ou by Lo -
e e al. [111], in which models ained using mean squa e loss and ad e sa ial loss
o p edic he nex image ame in a ideo sequence a e compa ed. A model ained
using mean squa e loss gene a es blu y images as shown in Fig.6, whe e ea and eyes
a e no sha ply defined as hey could be. Using an addi ional ad e sa ial loss, ea u es
like he eyes and ea emain p ese ed e y well, because an ea is he ecognizable
pa e n, and he disc imina o ne wo k would no accep any sample ha is missing
an ea .
This sec ion has a emp ed o p o ide eade s a b ie in oduc ion o he cu en
s a e o deep gene a i e image models. A quick summa y o his sec ion is depic ed
below in Fig.7.
Despi e ema kable achie emen s in gene a ing sha p and ealis ic images, GANs
suffe om ce ain d awbacks.
• Non con e gence Bo h gene a o and disc imina o ne wo ks in GANs a e ained
simul aneously using g adien descen in a ze o-sum game. As a esul , imp o e-
(11)
DJS(p ||pg)=
1
2DKL(p || p +pg
2)+
1
2DKL(pg|| p +pg
2
)
(12)
G∗=2DJS(p (x)||pg(x))−2log
2
Fig. 6 An illus a ion o he impo ance o an ad e sa ial loss [111]
Page 17 o 59
Sampa he al. J Big Da a (2021) 8:27
men o he gene a o ne wo k comes a he expense o disc imina o and ice
e sa. Hence he e is no gua an ee o GANs con e gence.
• Mode collapse Gene a o ne wo k achie es a s a e whe e i con inues o gene a e
samples wi h li le a ie y, al hough ained on di e se da ase s. This o m o ail-
u e is e e ed o as mode collapse.
• Vanishing g adien p oblems I he disc imina o is pe ec ly ained ea ly in he
aining p ocess, hen he e would be no g adien s le o ain he gene a o due
o anishing g adien s.
The e o e, many GAN- a ian s ha e been p oposed o o e come hese d awbacks.
These GAN- a ian s can be g ouped in o h ee ca ego ies:
1. A chi ec u e a ian s In e ms o a chi ec u e o gene a o and disc imina o ne -
wo ks, he fi s p oposed GANs use he Mul i- laye pe cep on (MLP). Owing o he
ac ha Con Ne s wo k well wi h high esolu ion image da a aking in o accoun o
he spa ial s uc u e o da a, a Deep Con olu ional GAN (DCGAN) [112] eplaced
he MLP wi h he decon olu ional and con olu ional laye s in gene a o and dis-
c imina o ne wo ks espec i ely.
Cu en S a e o Deep
Gene a e image models
Au o eg essi e models La en Va iable models Ad e sa ial models
* FVBN
* NADE
* MADE
* PixelRNN
* PixelCNN
* GATED PixelCNN
* PixelCNN ++
* PixelSNAIL
P os: Simple and s able
aining.
Cons: Image gene a ion
p ocess is na u ally slow.
* VAE
* β-VAE
* VQ-VAE
* VQ-VAE 2.0
* Condi ional VAE
* Va ia ional lossy
au oencode
* Pixel VAE
* DRAW
P os: Pe om bo h gene a ion
and in e ence wi h la en
a iables.
Cons: 1. La en a iable
models need assump ions on a
p io and pos e io
dis ibu ions.
2. Gene a ed images end o
be blu y.
* Gene a i e Ad e sa ial
Ne wo ks and i s a ian s.
P os: 1. Gene a e he sha pes
image sample.
2. Capable o cap u ing he
high- equency pa s o an
image.
Cons: Gene a i e Ad e sa ial
Ne wo ks a e highly uns able
and di icul o con e ge.
Fig. 7 Compa a i e summa y o Deep gene a i e models discussed in “Deep Gene a i e image models”
sec ion
Page 18 o 59
Sampa he al. J Big Da a (2021) 8:27
Au oencode based GANs such as AAE [113], BiGAN [114], VAE-GAN [115],
DEGAN [116], VEEGAN [117] e c., ha e been p oposed o combine hei cons uc-
ion powe o au oencode s wi h he sampling powe o GANs.
Condi ional based GANs like Condi ional GAN (CGAN) [118], Auxilia y Classifie
GAN (ACGAN) [119], VACGAN [120], in oGAN [121], and SCGAN [122] ocused
on con olling mode o da a being gene a ed by condi ioning model on condi ional
a iable.
2. T aining icks GANs a e difficul o ain. Imp o ed ainings icks such as ea u e
ma ching, miniba ch disc imina ion, his o ical a e aging, one-sided label smoo hing,
and Two Time-Scale Upda e Rule ha e been sugges ed o ensu e ha GANs con-
e ge o achie e Nash equilib ium.
3. Objec i e a ian s In o de o imp o e he s abili y and o e come anishing g adien
p oblems, diffe en objec i e unc ions ha e been explo ed in [123–130].
The ollowing sec ion o his e iew mo es on o desc ibe in g ea e de ail he selec ed
GAN a ian s.
Gene a i e ad e sa ial neu al ne wo ks
A chi ec u e a ian s
The pe o mance and aining s abili y o GANs a e highly influenced by he a chi ec-
u e o he gene a o and he disc imina o ne wo ks. Va ious a chi ec u e a ian s o
GANs ha e been p oposed ha adop se e al echniques o imp o e pe o mance and
s abili y.
i. Condi ional based GAN Va ian s
The s anda d GAN [67] a chi ec u e does no ha e any con ol on he modes o da a
being gene a ed. Van den Oo d e al. [89] a gue ha he class condi ioned image
gene a ion can significan ly enhance he quali y o gene a ed images. Se e al
condi ional based GANs ha e been p oposed ha lea n o sample om a condi-
ional dis ibu ion p(x|
y
) ins ead o ma ginal
p
(x)
.
Condi ional based GANs a i-
an s(Fig.8) can be classified in o wo g oups: 1. Supe ised and 2. Unsupe ised
condi ional GANs.
Supe ised condi ional GANs a ian s equi e a pai o images and co esponding p io
in o ma ion such as class label. The p io in o ma ion could be class labels, ex ual
desc ip ions, o da a om o he modali ies.
cGAN Mi za and Osinde o [118] p oposed condi ional Gene a i e Ad e sa ial Ne -
wo k (cGAN), o ha e a con ol on kind o da a being gene a ed by condi ioning
he model on p io in o ma ion
y
. Bo h disc imina o and gene a o in cGAN a e
condi ioned by eeding
y
as addi ional inpu . Using his p io in o ma ion, cGAN
is guided o gene a e ou pu images wi h desi ed p ope ies du ing he gene a ion
p ocess.
Page 19 o 59
Sampa he al. J Big Da a (2021) 8:27
ACGAN Auxilia y classifie Gene a i e Ad e sa ial Ne wo k (ACGAN) [119] is an
ex ension o he cGAN a chi ec u e. The disc imina o in he ACGAN ecei es
only he image, unlike he cGAN ha ge s bo h he image and he class label as
inpu . I is modified o dis inguish eal and ake da a as well as econs uc class
labels. The e o e, in addi ion o eal ake disc imina ion, he disc imina o also p e-
dic s class label o he image using an auxilia y decode ne wo k.
VACGAN The majo p oblem wi h ACGAN is ha i will affec he aining con e -
gence because o mixing he loss o classifie and disc imina o in o a single loss.
Ve sa ile Auxilia y Gene a i e Ad e sa ial Ne wo k (VACGAN) [120] sepa a es
ou classifie loss by in oducing a classifie ne wo k in pa allel o he disc imina-
o .
No p io in o ma ion is used in unsupe ised condi ional GAN a ian s o con ol on
modes o he image being gene a ed. Ins ead, ea u e in o ma ion such as hai
colo , age, gende e c. is lea ned du ing he aining p ocess. The e o e, hey need
an addi ional algo i hm o decompose he la en space in o disen angled la en ec-
o
c
, which con ains he meaning ea u es, and s anda d inpu noise ec o z. The
con en and ep esen a ion o an image is hen con olled by noise ec o z and
disen angled la en ec o
c
espec i ely.
In o-GAN In o ma ion maximizing Gene a i e Ad e sa ial Ne wo k (In o-GAN) [121]
spli s an inpu la en space in o he s anda d noise ec o
z
and addi ional la en
ec o
c
. The la en ec o c is hen made meaning ul disen angled ep esen a-
ion by maximizing he mu ual in o ma ion be ween la en ec o
c
and gene a ed
images G(z,c
)
using addi ional Q ne wo k.
SC-GAN Simila i y cons ain Gene a i e Ad e sa ial Ne wo k (SC-GAN) [122]
a emp s o lea n disen angled la en ep esen a ion by adding he simila i y con-
s ain be ween la en ec o
c
and gene a ed images G(z,c
)
. In o-GAN uses an
ex a ne wo k o lea n disen angle ep esen a ion, while SC-GAN only adds an
addi ional cons ain o a s anda d GAN. The e o e, SCGAN simplifies he a chi-
ec u e o In o-GAN.
ii. Con olu ional based GAN
DCGAN Deep Con olu ional Gene a i e Ad e sa ial Ne wo k (DCGAN) [112] is he
fi s wo k ha deploys con olu ional and anspose-con olu ional laye s in he dis-
c imina o and gene a o , espec i ely. The salien ea u es o he DCGAN a chi-
ec u e a e enume a ed as ollows:
• Fi s , he gene a o in DCGAN consis s o ac ional con olu ional laye s, ba ch no -
maliza ion laye s and ReLU ac i a ion unc ions.
• Second, he disc imina o is composed o s ided con olu ional laye s, ba ch no -
maliza ion laye s and Leaky ReLU ac i a ion unc ions.
• Thi d, uses Adap i e Momen Es ima ion (ADAM) op imize ins ead o s ochas ic
g adien descen wi h momen um.
iii. Mul iple GANs
Page 20 o 59
Sampa he al. J Big Da a (2021) 8:27
In o de o accomplish mo e han one goal, se e al amewo ks ex end he s anda d
GAN o ei he mul iple disc imina o s, gene a o s, o bo h(Fig.9).
P oGAN In an a emp o syn hesize highe esolu ion images P og essi e G owing o
Gene a i e Ad e sa ial Ne wo k (P oGAN) [131] s acks each laye o he gene a o
and disc imina o in a p og essi e manne as aining p og esses.
LAPGAN Laplacian Gene a i e Ad e sa ial Ne wo k (LAPGAN) [132] is p oposed
o he gene a ion o high quali y images. This a chi ec u e uses a cascade o Con-
Ne s wi hin a Laplacian py amid amewo k. LAPGAN u ilizes se e al Gene a-
o -Disc imina o ne wo ks a mul iple le els o a Laplacian Py amid o an image
de ail enhancemen . Mo i a ed by he success o sequen ial gene a ion, Im e al.
[133] in oduced Gene a i e Recu en Ad e sa ial Ne wo ks (GRAN) based on
ecu en ne wo k ha gene a e high quali y images in a sequen ial p ocess, a he
han in one sho .
D2GAN Dual disc imina o Gene a i e Ad e sa ial Ne wo k (D2GAN) [134] employs
wo disc imina o s and one gene a o o add ess he p oblem o mode collapse.
Unlike GANs, D2GAN o mula es a h ee-playe game ha u ilizes wo disc imi-
na o s o minimize he KL and e e se KL di e gences be ween ue da a and he
gene a ed da a dis ibu ion.
MADGAN Mul i-agen di e se Gene a i e Ad e sa ial Ne wo k (MADGAN) [135]
inco po a es mul iple gene a o s ha disco e di e se modes o he da a while
Fig. 8 A schema ic iew o (a) he anilla GAN and (b– ) a ian s o Condi ional GANs
Page 21 o 59
Sampa he al. J Big Da a (2021) 8:27
main aining high quali y o gene a ed images. To ensu e ha diffe en gene a o s
lea n o gene a e images om diffe en modes o he da a, he objec i e o disc im-
ina o is modified o de ec he gene a o which gene a ed he gi en ake image
along wi h disc imina ing he eal and ake images.
CoGAN Coupled GAN(CoGAN) [136] is used o gene a ing pai o like images in wo
diffe en domains. CoGAN is composed o a se o GANs–GAN1 and GAN2, each
accoun able o syn hesizing images in one domain. I leans a join dis ibu ion
om wo-domain images which a e d awn indi idually om he ma ginal dis ibu-
ions.
CycleGAN and DiscoGAN [137] use wo gene a o s and wo disc imina o s o accom-
plish unpai ed image o image ansla ion asks. CycleGAN [138] adop s he con-
cep o cycle consis ency om machine ansla ion, whe e a sen ence ansla ed
om English o Spanish and ansla e i back om Spanish o English should be
iden ical.
i . Au oencode based GAN Va ian s
The s anda d GANs a chi ec u e is unidi ec ional and can only map om la en space
z o da a space
x
, while au oencode s a e bidi ec ional. The la en space lea ned
by encode s is he dis ibu ion ha con ains comp essed ep esen a ion o he eal
images. Se e al a ian s o GANs ha combine GAN and encode a chi ec u e a e
p oposed o make use o he dis ibu ion lea ned by encode s(Fig.10). A ibu es
edi ing o an image di ec ly on da a space
x
is complex as image dis ibu ions a e
highly s uc u ed and high dimensional. In e pola ion on la en space can acili a e
o ende complica ed adjus men s in he da a space
x
.
DEGAN In s anda d GANs a chi ec u e, he inpu o he gene a o ne wo k is he
noise ec o ha is andomly sampled om a Gaussian dis ibu ion N(0, 1
)
, which may
c ea e a de ia ion om he ue dis ibu ion o eal images. Decode Encode Gene a i e
ad e sa ial Ne wo k (DEGAN) [116] adop decode and encode s uc u e om VAE,
p e ained on he eal images. The p e ained decode and encode s uc u e ans o m
Fig. 9 A schema ic iew o Va ian s o GANs wi h mul iple disc imina o s and gene a o s: a LAPGAN, b
MADGAN and c D2GAN
Page 22 o 59
Sampa he al. J Big Da a (2021) 8:27
andom Gaussian noise o dis ibu ion ha con ains in insic in o ma ion o he images
which is used as inpu o he gene a o ne wo k.
VAEGAN Va ia ional au oencode Gene a i e Ad e sa ial Ne wo k (VAEGAN) [115]
join ly ains VAE and GAN by eplacing he decode o VAE wi h GAN amewo k.
VAEGAN employs ea u e wise ad e sa ial loss o GAN in lieu o elemen wise econ-
s uc ion loss o VAE o imp o e quali y o image gene a ed by VAE. In addi ion o
la en loss and ad e sa ial loss, VAEGAN uses con en loss, also known as pe cep ual
loss, which compa es wo images based on high le el ea u e ep esen a ion om p e-
ained VGG Ne wo k [11].
AAE Unlike VAEGAN ha disc imina es in da a space, ad e sa ial au oencode s
(AAE) [113] imposes a disc imina o on he la en space as lea ning he la en code
dis ibu ion is simple han da a dis ibu ion. The disc imina o ne wo k disc imina es
be ween a sample d awn om la en space and om he dis ibu ion
p
(z) ha we a e
ying o model.
ALI and BiGAN In addi ion o gene a o ne wo k, Ad e sa ially Lea ned In e ence
(ALI) [114] model and Bidi ec ional Gene a i e Ad e sa ial Ne wo k (BiGAN) con ain
an encode componen E ha simul aneously lea n in e se mapping o he inpu da a
x
o he la en code
z
. Unlike o he a ian s o GAN whe e he disc imina o ne wo k
ecei es only eal o a ificially gene a ed images, in he BiGAN and ALI model, he dis-
c imina o ne wo k ecei es bo h image and la en code pai .
VEEGAN [117]: add esses he p oblem o mode collapse h ough addi ion o a econ-
s uc ion ne wo k ha e e ses he ac ion o he gene a o ne wo k. Recons uc ion
ne wo k akes in syn he ic images hen ans o ms hem o noise, while gene a o ne -
wo k akes noise as an inpu and econs uc s hem in o syn he ic image. In addi ion
o ad e sa ial loss, diffe ence be ween he econs uc ed noise and ini ial noise is used
o ain he ne wo k. Bo h gene a o and econs uc ion ne wo ks a e join ly ained,
which encou ages gene a o ne wo k o lea n ue dis ibu ion, hence sol ing he mode
collapse p oblem.
Fig. 10 A schema ic iew o Va ian s o GANs based on Encode and decode a chi ec u e: a AAE, b VAEGAN,
c DEGAN and d BIGAN
Page 23 o 59
Sampa he al. J Big Da a (2021) 8:27
Se e al o he GANs ha e been p oposed o image supe esolu ion. The goal o supe
esolu ion is o upsample low esolu ion images o a high esolu ion one. Ledig e al.
p oposed Supe -Resolu ion GAN (SRGAN) [139] o image supe esolu ion,which
akes poo quali y image as inpu , and gene a es high quali y image wi h 4 × esolu ion.
The gene a o o he SRGAN uses e y deep con olu ional laye s wi h esidual blocks. In
addi ion o an ad e sa ial loss, SRGAN includes a con en loss. The con en loss is com-
pu ed as he euclidean dis ance be ween he ea u e maps o he gene a ed high quali y
image and he g ound u h image, whe e ea u e maps a e ob ained om a p e ained
VGG19 [140] ne wo k. Zhang e al. [141] combined a sel a en ion mechanism wi h
GANs (SAGAN) o handle long ange dependencies ha make he gene a ed image look
mo e globally cohe en . Image- o-image ansla ion GANs such as Pix2Pix GAN [142],
Pix2pix HD GAN[143], and CycleGAN [137] lea n o map an inpu image om a sou ce
domain o an ou pu image om a a ge domain.A summa y o a chi ec u al a ian s
o GANs a e summa ized in Table1.
Objec i e a ian s
The main objec i e o GAN is o app oxima e he eal da a dis ibu ion. Hence, mini-
mizing dis ance be ween he eal da a dis ibu ion
(p
) and he GAN gene a ed da a
dis ibu ion
(
p
g)
is a i al pa o aining GAN. As s a ed in “Deep Gene a i e image
models” sec ion, s anda d GAN [67] uses Jensen Shannon di e gence o measu e simi-
la i y be ween eal and gene a ed da a dis ibu ions DJS(p ||p
g
) . Howe e , JS di e gence
ails o measu e dis ance be ween wo dis ibu ions wi h negligible o no o e lap. To
imp o e pe o mance and achie e s able aining o GAN, se e al dis ances o di e -
gence measu es ha e been p oposed ins ead o JS di e gence.
WGAN Wasse s ein Gene a i e Ad e sa ial Ne wo k (WGAN) [123] eplaces JSD
om he s anda d GAN wi h he Ea h mo e Dis ance (EMD). EMD also known as
Wasse s ein Dis ance (WD) can be in e p e ed in o mally as minimum amoun o
wo k o mo e ea h (quan i y o mass) om he shape o one dis ibu ion p(x) o ha o
ano he dis ibu ion q(x) so as o ma ch shape o bo h he dis ibu ions. WD is smoo h
and can p o ide meaning ul dis ance measu e be ween dis ibu ions wi h negligible o
no o e lap. WGAN imposes an addi ional Lipchi z cons ain o use WD as he loss in
he disc imina o , whe e i deploys weigh clipping o en o ce weigh s o he disc imina-
o o sa is y Lipchi z cons ain a e each aining ba ch.
WGAN-GP Weigh clipping in he disc imina o o a WGAN g ea ly diminishes i s
capaci y o lea n and o en ails o con e ge. WGAN-GP [124] is an ex ension o WGAN
ha eplaces weigh clipping wi h g adien penal y o en o ce disc imina o o sa is y
Lipchi z cons ain . Fu he mo e, Pe zka e al. [125] p oposed a new egula iza ion
me hod, also known as WGAN-LP, ha en o ces he Lipschi z cons ain .
LSGAN Leas squa es Gene a i e Ad e sa ial Ne wo k (LSGAN) [126] deploys leas
squa e loss ins ead o he c oss en opy loss in disc imina o o he s anda d GAN o
o e come he p oblem o Vanishing g adien as well as imp o ing quali y o gene a ed
image.
EBGAN Ene gy Based GAN (EBGAN) [127] uses au o-encode a chi ec u e o con-
s uc he disc imina o as an ene gy unc ion ins ead o a classifie . The Ene gy o
EBGAN is he mean squa ed econs uc ion e o o an au oencode , p o iding lowe
Page 24 o 59
Sampa he al. J Big Da a (2021) 8:27
ene gy o he eal images and high ene gy o gene a ed images. EBGAN exhibi s as e
and mo e s able beha io han s anda d GAN du ing aining.
Same as EBGAN, Bounda y Equilib ium GAN (BEGAN) [128], Ma gin adap a ion
GAN [129] and dual agen GAN [130] also deploy an au o-encode a chi ec u e as he
disc imina o . The disc imina o loss o BEGAN uses Wasse s ein dis ance o ma ch he
dis ibu ions o he econs uc ion losses o eal images wi h he gene a ed images.
The e a e also se e al o he objec i e unc ions based on C ame dis ance [144],
Mean/co a iance Minimiza ion [145], Maximum mean disc epancy [146], Chi-squa e
[147] ha e been p oposed o imp o e pe o mance and achie e s able aining o GAN.
Table 1 An o e iew o GANs a ian s discussed in“A chi ec u e a ian s” sec ion
Ca ego ies GAN Type Main A chi ec u al Con ibu ions oGAN
Basic GAN GAN [67] Use Mul ilaye pe cep on in he gene a o and disc imina o
Con olu ional Based GAN DCGAN [112] Employ Con olu ional and anspose-con olu ional laye s in
he disc imina o and gene a o espec i ely
PROGAN [131] P og essi ely g ow laye s o GAN as aining p og esses
Condi ion based GANs cGAN [118] Con ol kind o image being gene a ed using p io in o ma-
ion
ACGAN [119] Add a classifie loss in addi ion o ad e sa ial loss o econ-
s uc class labels
VACGAN [120] Sepa a e ou classifie loss o ACGAN by in oducing sepa a e
classifie ne wo k pa allel o he disc imina o
in oGAN [121] Lea n disen angled la en ep esen a ion by maximizing
mu ual in o ma ion be ween la en ec o and gene a ed
images
SCGAN [122] Lea n disen angled la en ep esen a ion by adding he
simila i y cons ain on he gene a o
La en ep esen a ion based GANs DEGAN [116] U ilize he p e ained decode and encode s uc u e om
VAE o ans o m andom Gaussian noise o dis ibu ion
ha con ains in insic in o ma ion o he eal images
VAEGAN [115] Combine VAE and GAN
AAE [113] Impose disc imina o on he la en space o he au oencode
a chi ec u e
VEEGAN [117] Add econs uc ion ne wo k ha e e se he ac ion o gen-
e a o ne wo k o add ess he p oblem o mode collapse
BiGAN [114] A ach encode componen o lea n in e se mapping o da a
space o la en space
S ack o GANs LAPGAN [132] In oduce Laplacian py amid amewo k o an image de ail
enhancemen
MADGAN [135] Use mul iple gene a o s o disco e di e se modes o he
da a dis ibu ion
D2GAN [134] Employ wo disc imina o s o add ess he p oblem o mode
collapse
CycleGAN [137] Use wo gene a o s and wo disc imina o s o accomplish
unpai ed image o image ansla ion ask
CoGAN [136] Use wo GANs o lea n a join dis ibu ion om wo-domain
images
O he a ian s SAGAN [141] Inco po a e sel -a en ion mechanism o model long ange
dependencies
GRAN [133] Recu en gene a i e model ained using ad e sa ial p ocess
SRGAN [139] Use e y deep con olu ional laye s wi h esidual blocks o
image supe esolu ion
Page 25 o 59
Sampa he al. J Big Da a (2021) 8:27
T aining icks
While esea ch on a ious GANs a chi ec u es and objec i e unc ions con inue o
imp o e he s abili y o aining, he e a e se e al aining icks p oposed in he li e -
a u e in ended o achie e excellen aining pe o mance. Rad o d e al. [112] showed
using leaky ec ified ac i a ion unc ions in bo h gene a o and disc imina o laye s ga e
highe pe o mance o e using o he ac i a ion unc ions. Salimans e al. [148] p oposed
se e al heu is ic app oaches which can imp o e he pe o mance, and aining s abili y
o GANs. Fi s , ea u e ma ching, changes he objec i e o he gene a o o minimize he
s a is ical diffe ence be ween ea u es o he gene a ed and eal images. In his way, he
disc imina o is ained o lea n impo an ea u es o he eal da a. Second, miniba ch
disc imina ion, whe e he disc imina o p ocess ba ch o samples, a he han in isola-
ion ha helps p e en mode collapse, as he disc imina o can iden i y i he gene a o
con inues o gene a e sample wi h li le a ie y. Thi d, his o ical a e aging, ha akes
he unning a e age o pa ame e s in he pas and penalizes i he e is a la ge diffe ence
be ween pa ame e s, which can help he model o con e ge o an equilib ium. Finally,
one-sided label smoo hing p o ides smoo hed labels o he disc imina o ins ead o 0 o
1, which can smoo h he classifica ion bounda y o he disc imina o .
Sønde by e al. [149] p oposed he idea o c ippling he disc imina o by in oducing
noise o he samples a he han labels, which p e en s he disc imina o om o e fi -
ing. Heusel e al. [150] used a sepa a e lea ning a e o gene a o and disc imina o ,
and ained GANs by aTwo Time-ScaleUpda e Rule (TTUR) o ensu e ha model con-
e ge o a s a iona y local Nash equilib ium. To s abilize he aining o he disc imina-
o , Miya o e al. [151] p oposed no maliza ion echnique called spec al no maliza ion.
Taxonomy o class imbalance in isual ecogni ion asks
This sec ion desc ibes diffe en GANs applied o imbalance p oblems in a ious isual
ecogni ion asks. We g oup he imbalance p oblems in a axonomy wi h h ee main
ypes: 1. Image le el imbalances in classifica ion 2. objec le el imbalances in objec
de ec ion and 3. pixel le el imbalances in segmen a ion asks. Unde s anding his axon-
omy o imbalances will p o ide a aluable amewo k o u he esea ch in o syn he ic
image gene a ion using GAN.
Class imbalances inclassi ica ion
Image classifica ion is he ask o classi ying an inpu image acco ding o a se o pos-
sible classes. Classifica ion can be b oken down in o wo sepa a e p oblems: bina y clas-
sifica ion and mul i-class classifica ion. Bina y classifica ion in ol es assigning an inpu
image in o one o wo classes, whe eas in mul i-class classifica ion wo o se e al classes
a e in ol ed. A classic example o a bina y image classifica ion p oblem is he iden ifica-
ion o ca s o dogs in each inpu image. Image da ase wi h high imbalance [152], which
includes in e -class imbalance and in a-classes imbalance, esul s in poo classifica ion
pe o mance.
Page 32 o 59
Sampa he al. J Big Da a (2021) 8:27
and a spa ial a en ion ne wo k (SAN). Gi en a ace image, SAN lea ns o localize he
a ibu e-specific egion and hen AMN edi he ace image wi h he desi ed a ibu es
in he specific egion loca ed by SAN.
The majo downside wi h he cu en app oaches is ha he inpu o GAN should
be on al ace images. I will be in e es ing o explo e a new a chi ec u e ha can be
ained o modi y he a ibu es o side- iew o any a bi a y iews.
Pe son e-iden ifica ion Pe son e-iden ifica ion [172] is ano he challenging ask wo h
men ioning, which a e ad e sely affec ed due o significan in a class imbalance. In a
class a ia ions caused by o a ion ( a ying poses) a e o en la ge han he in e -pe son
dissimila i ies used o diffe en ia e he ace images [173]. Recen ace- ecogni ion su eys
[174, 175] iden ified pose a ia ion as one o he p ominen un esol ed issues in ace- ec-
ogni ion ask. Fo ins ance, in o de o main ain he highes s anda d o secu i y, a sma
ideo sys em needs o be able o de ec a pe son in a ian o pose (Fig.16).
Qian e al. [176] in oduced a pose-no malized GAN model (PN-GAN) o alle ia ing
he effec s o pose a ia ion. Gi en any pedes ian image and a desi able pose as inpu ,
he model u ilized a desi able pose o p oduce a syn he ic image o he same iden i y
wi h he o iginal pose eplaced wi h he desi able pose (Fig.17). A e his, he au ho s
ained he e-iden ifica ion model wi h he o iginal images and gene a ed pose-no -
malized images o ex ac wo se s o ea u es. Finally, hey used he wo ypes o ea-
u es as he final ea u e. As a esul , he ea u es ex ac ed om he syn hesized images
imp o ed he gene aliza ion abili y o he e-iden ifica ion model.
To add ess pe son e-iden ifica ion challenges in complex scena ios, Wei e al. [177]
p oposed a model called Pe son T ans e Gene a i e Ad e sa ial Ne wo k (PTGAN) o
implausible pe son image s yle ans e om sou ce domain o a ge domain, ac oss
da ase s wi h diffe en s yles, such as backg ounds, poses, seasons, ligh ings, e c. The
domain ans e p ocedu e in PTGAN is inspi ed by CycleGAN [138]. Diffe en om
Fig. 16 Example o Pe son eiden ifica ion ask. Pe son eiden ifica ion is a key elemen in ideo su eillance
ha deals wi h ma ching images o same pe son o e many non-o e lapping came a iews
Page 33 o 59
Sampa he al. J Big Da a (2021) 8:27
Cycle-GAN [138], PTGAN inco po a es addi ional cons ain s on he pe son o e-
g ounds o make su e he s abili y o hei iden i ies du ing ans e . Compa ed wi h
Cycle-GAN, PTGAN gene a es high esolu ion pe son images, whe e pe son iden i ies
a e unchanged, and he s yles a e ans o med.
Being a c oss-came a acking and human e ie al ask, pe son e-iden ifica ion
o en suffe s om image s yle a ia ions esul ing om diffe en came as. The e-
o e, Zhong e al. [178] designed a came a s yle adap ion model o adjus ing Con Ne
aining. They ha e used CycleGAN [138] o ans e ing images omone came a o
hes yleo ano he came a. Gi en ha bo h o iginal and s yle ans e ed images, iden-
ifica ion disc imina i e embedding (IDE) is used o ain he Con Ne model. Pa icu-
la ly, au ho s ha e used ResNe -50 p e- ained on ImageNe da ase as backbone and
ollow he fine- uning s a egy.
Pedes ian images suffe om in o ma ion loss when ans e ing om one came a o
hes yleo ano he came a. Deng e al. [179] p esen ed a model, named simila i y p e-
se ing cycle consis en gene a i e ad e sa ial ne wo k (SPGAN), which is composed
o a CycleGAN and a Siamese ne wo k (SiaNe ). CycleGAN lea ns o ansla e pedes-
ian images om one domain o ano he domain, and he con as i e loss induced by
he SiaNe pulls close a ansla ed image and i s coun e pa in he sou ce domain, and
mo es away he ansla ed image and any image in he a ge domain.
Ge e al. [180] p esen ed a Fea u e Dis illing Gene a i e Ad e sa ial Ne wo k (FD-
GAN) ha aims a lea ning iden i y ela ed and pose-un ela ed pe son ep esen a ions.
The p oposed model adop s a Siamese s uc u e wi h mul iple no el disc imina o s
on human poses (pose disc imina o ) and iden i ies (iden i y disc imina o ). The idea
behind FD-GAN is o lea n pose-un ela ed and iden i y- ela ed ea u es o pedes ian
image, hen i can be used o gene a e he same pedes ian image bu wi h diffe en a -
ge poses.
Al hough GAN-based me hods desc ibed abo e ha e achie ed excellen pe o mance
in image-based pe son e-iden ifica ion, i s ill needs conside able effo o ackle he
ideo-based iden ifica ion da ase s. Fu u e wo k seeks o expand o use GAN o gene -
a ing a sequence o images o he ideo-based iden ifica ion da ase s.
Fig. 17 A chi ec u e diag am o pose-no malized GAN p esen ed by Qian e al. [176]
Page 34 o 59
Sampa he al. J Big Da a (2021) 8:27
Vehicle e-iden ifica ion Vehicle Re-iden ifica ion ask is e en mo e challenging as i su -
e s om la ge in a-class diffe ences caused by iewpoin and illumina ions a ia ions,
and in e -class simila i y p ima ily o diffe en iden i ies wi h he simila look (Fig.18).
Zhou e al. [182] p oposed a model called C oss iew GAN o gene a e images in
diffe en iewpoin s o he same ehicle. C oss iew GAN composed o classifica ion,
gene a o , and disc imina o ne wo k. Fi s , classifica ion ne wo k is ained o lea n
ehicle in insic ea u es such as model, colo , and ype in o ma ion. In addi ion o
in insic ea u es, i also lea ns iewpoin ea u es. Then he gene a i e ne wo k is
condi ioned on he a e age ea u e o he expec ed iewpoin and ehicle’s in insic
ea u es o in e images o he same ehicle in o he iewpoin s. The disc imina o
ne wo k lea ns o dis inguish eal images om he gene a ed images, while ensu ing
images a e gene a ed wi h co ec a ibu es.
Wu e al. [183] imp o ed he disc imina i e powe o he ResNe -50 model o he
Vehicle e-ID ask by simul aneously aining wi h ini ial labeled images and DCGAN
gene a ed unlabeled images. They u he explo e he effec i eness o using DCGAN
gene a ed images on a wide ange o ehicle e-ID da ase s and show imp o ed pe -
o mance o ehicle e-iden ifica ion.
Fine-g ained image classifica ion The fine-g ained image classifica ion is also a ib-
u ed o majo a ia ions in he in a-class and mino in e class a ia ions [184]. I is a
difficul ask o wo easons. Fi s , he aining samples o each class a e inadequa e.
Second, he diffe ences be ween diffe en classes o images a e qui e small [185]. As
an example, i is e y difficul o iden i y he images o She land Sheepdog om ha
o Collie dog. Simila ly, he images o Sayo nis and G ay Kingbi d a e qui e difficul o
dis inguish (Fig.19).
Fu e al. [184] de eloped a model called Fine g ained condi ional GAN (F-CGAN)
o sol e fine g ained class dependen image syn hesis p oblems. F-CGAN consis s o
Fig. 18 illus a ion o challenges in ehicle Re-iden ifica ion p o ided by Zheng e al. [181]
Page 35 o 59
Sampa he al. J Big Da a (2021) 8:27
h ee main componen s: 1. a 2-s age GAN, 2. a fine-g ained ea u e p ese e and 3.
a mul i- ask classifica ion model. The 2-s age GAN gene a es high esolu ion images,
he fine-g ained ea u e p ese e a ge s o cap u e fine g ained de ails and he
mul i- ask classifica ion model u ilizes gene a ed image da a o imp o e fine g ained
classifica ion accu acy.
Wang e al. [188] find ha he disc imina o in GANs lea ns a hie a chical iden-
ifica ion ea u es o he fine-g ained classes and disc imina e pa e n o he fine-
g ained aining samples. They use he a chi ec u e pic u ed below o implemen he
fine-g ained Plank on classifica ion ask (Fig. 20). The main idea is o ain a fine-
g ained classifie ha sha es weigh s wi h disc imina o o he DCGAN, which o ces
disc imina o o concen a e on ea u es o small classes. On WHOI-Plank on da ase
[189], F1 sco e o he classifie imp o ed by o e 7%.
Typically, medical image da ase s con ain bo h gene al labels, e.g., “male”, “ emale” and
disease specific de ailed labels [190]. I is men ioned ha he complexi y and na u e o
da a is ha d o lea n by using a single GAN. Hence, T. Koga e al. [190] connec ed wo
GANs in se ies, one o lea ning gene al ea u es and o he o de ailed ea u es. The fi s
GAN gene a es di e se images, which akes a noise ec o and gene al labels as inpu s.
Collie She land sheepdog Sayo nis G ay Kingbi d
Fig. 19 Sample images om he S an o d Dogs da ase [186] and he Cal ech-UCSD Bi ds da ase [187],
which exhibi s mino in e -class a ia ions and majo in a-class a ia ions
Fig. 20 Comple e fine-g ained Plank on classifie a chi ec u e used by Wang e al. [188]
Page 36 o 59
Sampa he al. J Big Da a (2021) 8:27
The second GAN ecei es syn he ic images gene a ed by he fi s GAN, and disease spe-
cific de ailed labels as inpu s, and gene a es he final fine-g ained medical images.
Mul iclass imbalance
In many eal wo ld p oblems such as emo ion classifica ion [191], plan disease classifi-
ca ion [192], medical image classifica ion [193], indus ial de ec classifica ion [194] e c.,
i is mo e likely ha mo e han one class exis s and needs o be ecognized. Mul iclass
classifica ion has been shown o suffe mo e lea ning difficul ies han bina y class clas-
sifica ion, because mul iclass classifica ion inc eases he da a complexi y and in ensifies
he imbalanced dis ibu ion [195]. Th ee ypes o imbalance could occu o he mul i-
class da ase s: ew mino i y-many majo i y classes, many mino i y- ew majo i y classes,
and many mino i y-many majo i y classes. Shuo Wang e al. [196] s udied he impac o
all diffe en ypes o mul iclass imbalances and showed ha hey nega i ely affec mino -
i y class and o e all pe o mance.
An example o ew mino i y-many majo i y class imbalance is an emo ion classifica-
ion, as some classes o emo ions like disgus a e ela i ely uncommon compa ed o
common emo ions like happy o sad. Zhu e al. [197] employed cycle-GAN which can
syn hesize uncommon emo ion classes like disgus ed om he equen classes (Fig.21).
In addi ion o ad e sa ial and cycle consis ency loss, hey use leas squa e loss om
LSGAN o a oid anishing g adien p oblems. Employing cycle-GAN based da a mino -
i y class da a augmen a ion achie ed 5–10% inc ease in he o e all accu acy. They also
ound ha enla ging mino i y classes also inc eases accu acy o o he majo i y classes.
Wea he Image classifica ion is ano he example o ew mino i y-many majo i y class
imbalance, because some ypes o wea he , like snow, is ela i ely a e compa ed o
sunny, hazy and ainy days. Li e al. [198] used DCGAN o gene a e images o mino i y
classes in aining. They ound ha he GAN-based da a augmen a ion echnique led
o ma gin cla i y be ween classes and hence imp o emen in classifica ion pe o mance.
Fig. 21 On emo ion classifica ion ask [197], he images on he le a e o iginal da a and he es a e images
gene a ed by cycle-GAN
Page 37 o 59
Sampa he al. J Big Da a (2021) 8:27
Huang e al. [199] p esen ed an in e es ing idea o combine ensemble lea ning wi h
GANs designed o add ess he class imbalance p oblem in wea he classifica ion. The
p oposed me hod comp ised o h ee ing edien s as depic ed in (Fig.22): 1. DCGAN o
gene a e syn he ic images and balance he aining da ase 2. Nea es neighbo me hod
o emo e any possible ou lie images gene a ed by DCGAN 3. An ensemble lea ning
me hod o combine he classifica ion esul s o he mul iple classifie s so as o achie e
be e esul s.
The use o DCGAN was es ed by Salehinejad e al. [193] in he ask o ches pa hol-
ogy classifica ion. Using ches X- ay images, hey build a deep Con Ne classifie o
classi y 5 diffe en anemic classes. Thei da ase is highly imbalanced, con ains h ee
majo i y and wo mino i y classes (Fig.23a). The syn he ic images gene a ed using
DCGAN we e used o balance and augmen he o iginal imbalanced da ase . They
demons a ed ha a combina ion o he o iginal imbalanced da ase and gene a ed
images imp o es he accu acy o deep Con Ne classifie in compa ison o he same
classifie ained wi h o iginal imbalanced da ase alone. On ches X- ay da ase
[193], a mean classifica ion accu acy imp o ed om 70.87 o 92.10%.
F id-Ada e al. [200] also showed ha gene a ing syn he ic li e lesion images
using DCGAN can imp o e classifica ion esul s. They combined s anda d augmen-
a ion echniques and DCGAN gene a ed syn he ic images o ain a classifie . Thei
li e lesion da ase con ains 182 compu ed omog aphy images (65 hemangiomas, 64
me as ases and 53 cys s). By adding he syn he ic images o s anda d da a augmen-
a ion, hei classifica ion pe o mance inc eased om 78.6% sensi i i y and 88.4%
specifici y using s anda d augmen a ions o 85.7% sensi i i y and 92.4% specifici y
using DCGAN-based syn he ic images.
Fig. 22 Illus a ion om Huang e al. [199] showing how he Ensemble lea ning is in eg a ed wi h GAN
F amewo k
Page 38 o 59
Sampa he al. J Big Da a (2021) 8:27
Rashid e al. [201] es ed he effec i eness o using GANs o gene a e skin lesion
images. Using ISIC 2018 da ase [202], hey buil a CNN classifie o classi y 7 diffe -
en skin lesions as depic ed in Fig.24. These classes a e highly imbalanced, and he
GAN is used as a me hod o in elligen o e sampling.
Nazki e al. [192] used Cycle-GAN o alle ia e mul iclass imbalance p oblem in
oma o plan disease classifica ion. Thei oma o plan disease da ase con ains 2789
images, highly suffe ed om class imbalance in 9 disease ca ego ies (Fig.23b). Using
Cycle-GAN, hey ansla ed images om he heal hy oma o lea es o unde ep-
esen ed diseased oma o lea es. This s udy demons a ed ha he syn he ic image
gene a ed by Cycle-GAN can be used as an augmen ed aining se o imp o e he
pe o mance o classifie .
Bha ia e al. [203] sough ou o compa e syn he ic images gene a ed using WGAN-
GP agains he s anda d da a augmen a ion in he con ex o mul iclass image clas-
sifica ion. They a ificially in oduced class imbalance in wo balanced da ase s o
CIFAR-10 [87] and FMNIST [204], and s udied he effec s o mul iclass imbalance on
classifica ion pe o mance. On he CIFAR-10 [87] da ase , classifica ion pe o mance
imp o ed om 80.84% accu acy and 0.806 F1-sco e using s anda d da a augmen a-
ion o 81.89% accu acy and 0.812 F1-sco e using WGAN-GP. On FMNIST [204]
da ase , pe o mance imp o ed om 91.9% accu acy and 0.921 F1-sco e using aug-
men a ion o 92.8% accu acy and 0.923 F1-sco e using WGAN-GP.
0
1000
2000
3000
4000
5000
6000
AKIEC BCC BKL DF MEL NV VASC
Numbe o images
ab
Fig. 24 a Dis ibu ion o he se en skin lesion class labels o he ISIC 2018 da ase [202]. b Sample images
om each class
0
5000
10000
15000
20000
25000
30000
35000
Imbalanced Balanced
Numbe o images
Ca diomegaly No mal Effusion Edema Pneumo ho ax
DC-GAN
0
100
200
300
400
500
600
700
800
900
Imbalanced Balanced
Numbe o images
Canke G ay Mold Lea Mold
Low empe a u e Mine Nu ional excess
Plague Powde y mildew Whi efly
Cycle-GAN
ab
Fig. 23 The dis ibu ions o (a) Ches X- ay image da ase [193] and (b) oma o plan disease da ase [192],
be o e (le ) and a e class balancing using GANs ( igh )
Page 39 o 59
Sampa he al. J Big Da a (2021) 8:27
An idea o GANs based ans e lea ning echnique o mul iclass imbalance p ob-
lem is p oposed by Fanny e al. [205]. Thei a chi ec u e named class expe gene a-
i e ad e sa ial ne wo k (CE-GAN) makes use o mul iple GANs models, a sepa a e
GANs o each class. Fea u e maps in he main classifie a e a anged in pa allel, wi h
each ea u e maps p e- ained o iden i y he cha ac e is ics o a single class in he
aining da a (Fig. 25). The weigh s o he p e ained ea u e maps a e ans e ed
om disc imina o s o he GANs o main classifie model o u he aining in a
supe ised mode.
The GAN-based syn he ic images se ed as an in elligen o e sampling echnique
and can add ess he p oblem o mul i-class imbalance o a g ea e ex en . Howe e ,
syn he ic images mus be used wi h cau ion because i he quali y o he syn hesized
images is no high, his would lead o addi ional noise o he o iginal da ase s.
Objec le el imbalances inobjec de ec ion
Objec -scale imbalance
One pe asi e challenge in he scale in a ian objec de ec ion is la ge scale a iance
ac oss objec ins ances, and pa icula ly, de ec ing small objec s a e mo e challeng-
ing han medium and la ge-scale objec s. As pe MS COCO defini ion [206], Objec s
Fig. 25 Illus a ion o he class expe gene a i e ad e sa ial ne wo k a chi ec u e [205]
Page 40 o 59
Sampa he al. J Big Da a (2021) 8:27
wi h size less han 32 × 32 pixels a e small, size be ween 32 × 32 o 96 × 96 pixels
a e conside ed as medium and objec s wi h size g ea e han 96 × 96 pixels a e la ge
objec s (Table2). On he one hand, small objec s in MS COCO da ase accoun s o
only 1.23% o o al objec a ea, on he o he hand, medium and la ge-scale objec s a e
o e 98% o objec a ea. Objec de ec ion algo i hms should be able o de ec bo h
small objec s as well as medium and la ge objec s. De ec ing small objec s a e essen-
ial in many eal-wo ld applica ions. Fo ins ance, de ec ing dis an o small objec s in
he high- esolu ion d i ing scene images cap u ed om ca s is essen ial o achie ing
au onomous d i ing. Many dis an objec s, such as affic ligh s o ca s, a e impe -
cep ible as shown in Fig.26. Haoyue e al. [207] measu e he ex en o scale a ia ion
using he coefficien o a ia ion (CV), de e mined as he a io o he s anda d de ia-
ion o he mean o he objec scale. The bigge he CV, he mo e complica ed he
p oblem o scale a ia ion.
The e can be h ee easons why de ec ing small objec s a e mo e complica ed han
la ge one: 1. Small objec s occupy a much smalle a ea, and consequen ly he e exis s
lack o di e si y whe e small objec s a e loca ed in he image, 2. The e a e compa a i ely
less images in he da ase con aining small objec s which may bias any objec de ec ion
algo i hm o concen a e mo e on medium and la ge-scale objec s, and 3. The ac i a-
ions o small objec s become smalle and smalle wi h each pooling laye in a s anda d
con Ne a chi ec u e as i p og essi ely educes he spa ial size o an image.
To o e come he p oblem o scale imbalance, wo diffe en s a egies based on GAN
ha e been p oposed in he li e a u e. Commonly adop ed s a egy is o con e low
esolu ion small objec ea u es in o high esolu ion ea u es [208] using GAN. Di e -
si y o he small objec loca ions in he images a e enhanced by copy-pas ing small
objec ins ances se e al imes in each image h ough ad e sa ial p ocesses [209].
Table 2 The de ini ions ands a is ics o hesmall, medium, andla ge objec s asMS COCO
[206]
Objec ca ego y Spa ial dimension Objec coun % To al
objec
a ea %
Minimum Maximum
Small 0 × 032 × 32 41.43 1.23
Medium 32 × 32 96 × 96 34.32 10.18
La ge 96 × 96 ∞ × ∞24.24 88.59
Fig. 26 Example o scale a ia ion and he scale (objec size) dis ibu ion o he VisD one2019 da ase
objec s in pixels [207]
Page 41 o 59
Sampa he al. J Big Da a (2021) 8:27
Li e al. [208] u ilized a GAN amewo k ha ans o ms poo ep esen a ion o small-
scale objec s o supe - esol ed la ge objec s. The gene a o a emp s o gene a e supe
esolu ion ea u es o he small objec s. The disc imina o in his amewo k is decom-
posed in o wo b anches, namely, a pe cep ual b anch and an ad e sa ial b anch. An
ad e sa ial b anch is ained o disc imina e be ween eal la ge-scale objec s and gen-
e a ed supe esolu ion objec s while a pe cep ual b anch helps o make su e ha he
gene a ed supe - esol ed objec is use ul o he de ec ion (Fig.27b). They es ed he
effec i eness o his amewo k on Tsinghua-Tencen 100k da ase [210], PASCALVOC
da ase [211] and Cal ech pedes ian benchma k [212].On he PASCAL VOC 2007
Fig. 27 A chi ec u e diag am o (a) SOD-MTGAN [213] (b) Pe cep ual GAN [208] and (c) De ec o GAN [209]
Fig. 28 Imbalanced dis ibu ion o occluded, pa ially occluded and hea ily occluded objec s in
VisD one-DET2018 da ase [215]
Page 48 o 59
Sampa he al. J Big Da a (2021) 8:27
is no clea how much syn he ic images mus be blended wi h o iginal images o achie e
he maximum pe o mance o he classifie s. Addi ionally, syn he ic images would lead
o addi ional noise o he o iginal aining da ase i he quali y o he syn hesized images
is poo . The e o e, mos o he su eyed me hods in GANs based in elligen o e sam-
pling me hods [197] ocused mainly on balancing dis ibu ion as well as imp o ing qual-
i y o he gene a ed images.
Image- o-image ansla ion [138] me hods used o in e -class imbalance p oblem
canno be ex ended o sol e in a-class imbalance as i is difficul o acqui e image da a-
se s wi h de ailed labels. The in e es ing way o sol e his p oblem is o employ clus-
e ing echniques in he ea u e space o he GANs o di ide he images in o diffe en
g oups o au oma ic pa e n ecogni ion in he da ase . Imp o ing he pe o mance o
he clus e ing echniques ha clea ly find he diffe ence among clus e s, is an a ea o
u u e wo k.
GANs and encode ne wo k hyb id models ha e a good po en ial o add ess in a class
imbalance p oblem in ace ecogni ion and e-iden ifica ion asks. The key idea o hese
models is o wo k on la en code space a he han he pixel space. This is because o
Table 3 (con inued)
Ca ego y Imbalance ype S udy Applica ion
Objec de ec ion Objec Scale imbalance Pe cep ual GAN [208] T affic sign de ec ion
Objec Scale imbalance SOD-MTGAN [213] Small objec de ec ion
sys em
Objec Scale imbalance De ec o GAN [209] Pedes ian and disease
de ec ion
Imbalance due o occlu-
sions and de o ma ions
Ad e sa ial-Fas -RCNN
[216]
Occluded objec de ec ion
Imbalance due o occlu-
sions and de o ma ions
Ad e sa ial Occlusion-
awa e Face De ec o
[217]
Occluded ace de ec ion
Imbalance due o occlu-
sions and de o ma ions
Cu -Pas e GAN [218] Occluded objec de ec ion
Fo eg ound Backg ound
objec class imbalance
Task-awa e syn he ic da a
gene a ion [219]
Objec de ec ion
Fo eg ound Backg ound
objec class imbalance
Gene-GAN [221] Objec de ec ion
Fo eg ound Backg ound
objec class imbalance
PSIS [220] Objec de ec ion
Segmen a ion Pixel wise Imbalance Sensi i i y condi ional GAN
[118]
Shadow de ec ion
Pixel wise Imbalance Pix2pix HD GAN [143] Imbalanced pedes ian
image segmen a ion
Pixel wise Imbalance Voxel GAN [226] B ain umo segmen a ion
Pixel wise Imbalance GAN + ensemble lea ning
[228]
Medical image seman ic
segmen a ion
Pixel wise Imbalance GAN + Weigh ed ca ego i-
cal loss [227]
Hea image segmen a ion
Imbalance due o occlu-
sions
SeGAN[231] In isible pa gene a ion and
Segmen a ion
Imbalance due o occlu-
sions
Occlusion-Awa e GAN
[232]
Occlusion ee image gen-
e a ion
Page 49 o 59
Sampa he al. J Big Da a (2021) 8:27
manipula ing a fine g ained image ca ego y, e.g., hai colo , he la en code ep esen a-
ion will ope a e only on ha single la en code (hai colo ), whe eas he pixel space will
edi e e y single pixel in an image.
The ascina ing app oaches o use GANs o he p oblem o objec le el imbalances in
objec de ec ion asks all in o wo gene al ca ego ies: 1. Gene a ing mo e a e examples
as in elligen o e sampling used o class imbalance. These gene a ed a e examples a e
in oduced in o he aining da ase o add ess imbalance p oblems. 2. Lea n an ad e -
sa y in combina ion wi h o iginal objec de ec ion algo i hms. This ad e sa y modifies
he ea u es o sol e imbalance p oblems ins ead o gene a ing examples in pixel space.
i.e., o gene a e ha d- o-de ec samples by pe o ming ea u e space manipula ions.
The capabili y o supe - esolu ion GANs a e being used o up-sample small blu ed
objec s in o fine-scale ones and o eco e de ailed spa ial in o ma ion o accu a e small
objec de ec ion. This echnique combines supe - esolu ion GANs wi h objec de ec ion
algo i hms o sol e he imbalances due o objec size. The powe o ad e sa ial p ocess is
being used o inc ease he di e si y o he small objec loca ions in he images by copy-
pas ing small objec ins ances se e al imes a diffe en loca ions.
Making he bes use o GANs and combining hem in o U-Ne a chi ec u es is an
in e es ing way o sol e pixel le el imbalances in segmen a ion asks. These a chi ec u es
o en use a weigh ed loss unc ion o mi iga e he pixel le el imbalances. Combina ion
o image in pain ing GANs wi h U-Ne a chi ec u es has he g ea po en ial use in seg-
men ing hidden objec s. This echnique is no only efficien in segmen a ion asks, bu
also o in e he appea ance o he objec s beyond hei isible pa s. O e all, combining
diffe en deep lea ning models wi h ad e sa ial p ocess can p o ide a way o sol e many
o he open p oblems in he compu e ision field.
Fu u e wo k
E en hough GANs can be used as an effec i e way o unlock addi ional in o ma ion
om a da ase , he syn he ic images gene a ed by GANs canno eplace he eal images
comple ely. Howe e , a blend o diffe en p opo ions o eal and GANs gene a ed
images a e ex emely use ul o imp o e he di e si y o he aining samples and inc ease
pe o mance o he classifie s. Ou u u e wo k in ends o s udy he influences o blend-
ing diffe en p oposi ions o GANs gene a ed images and eal images on he classifica-
ion pe o mance. The e a e a e y limi ed numbe o compa a i e s udies ha compa e
effec i eness o using GAN based syn he ic images wi h o he adi ional me hods o
in a-class imbalances. We also in end o conduc he compa a i e s udy in o de o ali-
da e he effec i eness o using syn he ic images o in a class imbalances.
Infla ing he size o he da ase b ings ano he p oblem: One o he mos significan
limi a ions in compu e ision expe imen s is compu a ional esou ces. Sophis ica ed
compu e ision models ained on infla ed da ase can pe o m complex asks, he p ob-
lem howe e is, how do we deploy such massi e a chi ec u e on edge de ices o ins an
usage. Handling his p oblem using knowledge dis illa ion is non- i ial and an ac i e
field o esea ch. Knowledge dis illa ion is model comp ession echnique in which a
smalle ne wo k is ained wi h he help o he sophis ica ed p e ained model o achie e
he simila accu acy. This aining p ocess is o en e e ed o as " eache -s uden ”, whe e
Page 50 o 59
Sampa he al. J Big Da a (2021) 8:27
he sophis ica ed p e ained model is he eache and he smalle ne wo k is he s uden .
Wang e al. [235] combine GANs and knowledge dis illa ion o imp o e he efficiency o
he s uden ne wo k in objec de ec ion. Simila o his wo k, we will a emp o u he
implemen GANs and knowledge dis illa ion combina ions o o he compu e isions
asks.
As esea ch on GANs a e de eloping and ma u ing, assessmen o pe o mance has
become essen ial. E alua ion me ics helps o quan i a i ely measu e how well GANs
models a e pe o ming, also o assess he ela i e pe o mance o GANs. Ve y o en he
pe o mance o GANs is measu ed by he manual inspec ion o he isual fideli y o gen-
e a ed images. Howe e , he manual inspec ion is cumbe some, subjec i e, ime-con-
suming, and some imes misleading. Lack o uni e sal e alua ion me ics can impede he
de elopmen o GANs. In oducing new pe o mance measu es o e alua e bo h di e -
si y and fideli y o gene a ed images is a e y impo an a ea o u u e wo k.
Manually designing GANs a chi ec u e o a gi en ask is ime-consuming and some-
imes has a endency o e o s. This d awback has led esea che s o mo e on o he nex
s age o au oma ing GANs a chi ec u e in he o m o neu al a chi ec u e sea ch (NAS).
Ano he in e es ing a ea o u he esea ch is o use me a-heu is ic sea ch algo i hms
ha assis a chi ec u al sea ch and find op imal GANs a chi ec u e which ou pe o ms
human c ea ed GANs models.
Achie ing equilib ium be ween he gene a o and disc imina o o he GANs can ake
a long ime ela i e o o he deep neu al ne wo ks. Dis ibu ed aining o GAN h ough
pa alleliza ion and clus e compu ing is ano he impo an a ea o u u e wo k o cu
down he aining ime.
Mos o he applica ions o he GANs so a ha e been o c ea ing syn he ic images.
GANs a e no limi ed o he isual domain and can be also applied o non- isual appli-
ca ions. Fo example, Paganini e al. [236] used GANs o p edic he ou come o high
ene gy pa icle physics expe imen s. Ins ead o using explici Mon e Ca lo simula ion
o he eal physics o e e y s ep, he GANs lea n by example wha ou come is likely o
occu in each si ua ion. The GANs educe he compu a ional cos o high ene gy pa icle
simula ion, enough o sa e millions o dolla s’ wo h o supe compu e ime. We belie e
ha he in en ion o new applica ions using his powe ul ool will be con inued in he
u u e.
Conclusion
This pape su eys a ious GANs a chi ec u es ha ha e been used o add essing he
diffe en imbalance p oblems in compu e ision asks. In his su ey, we fi s p o ided
de ailed backg ound in o ma ion on deep gene a i e models and GAN a ian s om he
a chi ec u e, algo i hm, and aining icks pe spec i e. In o de o p esen a clea oad-
map o a ious imbalance p oblems in compu e ision asks, we in oduced axonomy
o he imbalance p oblems. Following he p oposed axonomy, we discussed each ype o
p oblems sepa a ely in de ail and p esen ed he GANs based solu ions wi h impo an
ea u es o each app oach and hei a chi ec u es. We ocused mainly on he eal-wo ld
applica ions whe e GAN based syn he ic images a e used o alle ia e class imbalance. In
addi ion o he ho ough discussion on he imbalance p oblems and hei solu ions, we
add essed many open issues ha a e c ucial o compu e ision applica ions.
Page 51 o 59
Sampa he al. J Big Da a (2021) 8:27
Syn he ic bu ealis ic images gene a ed using he me hods discussed in his su ey
ha e he po en ial o mi iga e he class imbalance p oblem while p ese ing he ex insic
dis ibu ion. Many o he me hods su eyed in his pape ackled he highly complex
imbalances by combining GANs a chi ec u e wi h diffe en o he deep lea ning ame-
wo ks. Specifically, he use o au oencode s wi h GANs has offe ed an effec i e way o
pe o m ea u e space manipula ions ins ead o complex pixel space ope a ions.
Syn he ic images gene a ed by GANs canno be used as he comple e eplacemen o
eal da ase s. Howe e , he blend o eal and GANs gene a ed images ha e eno mous
po en ial o inc ease he pe o mance o he deep lea ning model. Looking in o he
u u e, GAN- ela ed esea ch in image as well as non-image da a domains o add ess he
p oblem o imbalances and limi ed aining da ase would con inue o expand. We con-
clude ha he u u e o GANs is p omising and he e a e clea ly a lo o oppo uni ies o
u he esea ch and applica ions in many fields.
Abb e ia ions
Con Ne s: Con olu ional neu al ne wo ks; SMOTE: Syn he ic mino i y o e sampling echnique; ADASYN: Adap i e
syn he ic sampling; IHM: Ins ance ha dness measu e; SSL: Semi-supe ised lea ning; R-CNN: Region-based con olu-
ional neu al ne wo ks; RPN: Region p oposal ne wo k; YOLO: You only look once; SSD: Singe sho de ec ion; SNIP: Scale
no maliza ion o image py amids; FPN: Fea u e py amid ne wo ks; RNN: Recu en neu al ne wo ks; LSTM: Long sho -
e m memo y; PCA: P inciple componen analysis; MADE: Masked au oencode densi y es ima o ; ARs: Au o eg essi e
models; FVBNs: Fully isible belie ne wo ks; RGB: Red G een blue; NADE: Neu al au o eg essi e densi y es ima o ; MADE:
Masked au oencode densi y es ima o ; VAEs: Va ia ional au o encode s; CVAE: Condi ional a ia ional au o encode s;
DC-IGN: Deep con olu ional in e se g aphics ne wo k; IWVAE: Impo ance weigh ed Va ia ional Au o Encode s; VQ-VAEs:
Vec o quan ized a ia ional au o encode s; DRAW : Deep ecu en a en i e w i e ; EMD: Ea h mo e Dis ance; TTUR
: Two ime-scale upda e ule; DDSM: Digi al da abase o sc eening mammog aphy; ARU-ne : Ad e sa ially egula ized
U-ne ; AMN: A ibu e manipula ion ne wo k; SiaNe : Siamese ne wo k; CV: Coefficien o a ia ion; AP: A e age p eci-
sion; ASTN: Ad e sa ial spa ial ans o me ne wo k; ASDN: Ad e sa ial spa ial d opou ne wo k; mAP: Mean a e age
p ecision; AOFD: Ad e sa ial occlusion awa e ace de ec ion; PSIS: P og essi e and selec i e ins ance-swi ching; ADAM:
Adap i e momen es ima ion op imize ; ReLU: Rec ified linea uni ; GANs: Gene a i e ad e sa ial neu al ne wo ks;
cGAN: Condi ional gene a i e ad e sa ial ne wo k; ACGAN: Auxilia y classifie gene a i e ad e sa ial ne wo k; VACGAN:
Ve sa ile Auxilia y classifie gene a i e ad e sa ial ne wo k; In oGAN: In o ma ion maximizing gene a i e ad e sa ial
ne wo k; SCGAN: Simila i y cons ain gene a i e ad e sa ial ne wo k; DCGAN: Deep con olu ional gene a i e ad e sa ial
ne wo k; P oGAN: P og essi e g owing o gene a i e ad e sa ial ne wo k; LAPGAN: Laplacian gene a i e ad e sa ial
ne wo k; GRAN: Gene a i e ecu en ad e sa ial ne wo ks; D2GAN: Dual disc imina o gene a i e ad e sa ial ne wo k;
MADGAN: Mul i-agen di e se gene a i e ad e sa ial ne wo k; CoGAN: Coupled gene a i e ad e sa ial ne wo k; DEGAN:
Decode encode gene a i e ad e sa ial ne wo k; VAEGAN: Va ia ional au oencode gene a i e ad e sa ial ne wo k;
AAE: Ad e sa ial au oencode s; ALI: Ad e sa ially lea ned in e ence; BiGAN: Bidi ec ional gene a i e ad e sa ial ne wo k;
SRGAN: Supe - esolu ion gene a i e ad e sa ial ne wo k; SAGAN: Sel -a en ion gene a i e ad e sa ial ne wo k; WGAN:
Wasse s ein gene a i e ad e sa ial ne wo k; WGAN-GP: Wasse s ein gene a i e ad e sa ial ne wo k wi h g adien pen-
al y; LSGAN: Leas squa es gene a i e ad e sa ial ne wo k; EBGAN: Ene gy based gene a i e ad e sa ial ne wo k; BEGAN:
Bounda y equilib ium gene a i e ad e sa ial ne wo k; SD-GAN: Su ace de ec -gene a i e ad e sa ial ne wo k; BAGAN:
Balancing gene a i e ad e sa ial ne wo k; ciGAN: Condi ional infilling gene a i e ad e sa ial ne wo k; IcGAN: In e ible
condi ional gene a i e ad e sa ial ne wo k; PNGAN: Pose-no malized gene a i e ad e sa ial ne wo k; PTGAN: Pe son
ans e gene a i e ad e sa ial ne wo k; SPGAN: Simila i y p ese ing cycle consis en gene a i e ad e sa ial ne wo k;
FD-GAN: Fea u e dis illing gene a i e ad e sa ial ne wo k; F-CGAN: Fine g ained condi ional GAN; CE-GAN: Class expe
gene a i e ad e sa ial ne wo k; ScGAN: Sensi i i y condi ional gene a i e ad e sa ial ne wo k; OAGAN: Occlusion-awa e
gene a i e ad e sa ial ne wo k.
Acknowledgemen s
The au ho s would like o hank he anonymous e iewe s o hei aluable commen s and sugges ions on he pape .
Also, we acknowledge he membe s o he Au onomous and In elligen Sys ems Uni , Teknike , o aluable discussions
and collabo a ions.
Au ho s’ con ibu ions
VS pe o med he p ima y li e a u e e iew and analysis o his su ey, and also d a ed he manusc ip . IM, JJAM and AG
wo ked wi h VS o de elop he a icle’s amewo k and ocus. IM and JJAM double checked he manusc ip and p o ided
se e al ad anced ideas o his manusc ip . All au ho s ead and app o ed he final manusc ip .
Funding
This esea ch wo k was unde aken in he con ex o DIGIMAN4.0 p ojec (“Digi al Manu ac u ing Technologies o Ze o‐
de ec ”, h ps ://www.digim an4-0.mek.d u.dk/). DIGIMAN4.0 is a Eu opean T aining Ne wo k suppo ed by Ho izon 2020,
he EU F amewo k P og amme o Resea ch and Inno a ion (P ojec ID: 814225). This esea ch was also pa ly suppo ed
by he ELKARTEK p ojec KK-2020/00049 3KIA o he Basque Go e nmen .
Page 52 o 59
Sampa he al. J Big Da a (2021) 8:27
A ailabili y o da a and ma e ials
No applicable.
E hics app o al and consen o pa icipa e
No applicable.
Consen o publica ion
No applicable.
Compe ing in e es s
The au ho s decla e ha hey ha e no compe ing in e es s.
Au ho de ails
1 Au onomous and In elligen Sys ems Uni , Teknike , Membe o Basque Resea ch and Technology Alliance, Eiba , Spain.
2 Design and Manu ac u ing Enginee ing Depa men , Uni e sidad de Za agoza, 3 Ma ía de Luna S ee , To es Que edo
Bld, 50018 Za agoza, Spain.
Recei ed: 30 July 2020 Accep ed: 16 Janua y 2021
Re e ences
1. Nug aha BT, Su SF, Fahmizal. Towa ds sel -d i ing ca using con olu ional neu al ne wo k and oad lane de ec o .
P oceedings o he 2nd In e na ional Con e ence on Au oma ion, Cogni i e Science, Op ics, Mic o Elec o-
Mechanical Sys em, and In o ma ion Technology, ICACOMIT 2017. 2017;2018-Janua:65–9.
2. Yada SS, Jadha SM. Deep con olu ional neu al ne wo k based medical image classifica ion o disease diagnosis.
J Big Da a. 2019. h ps ://doi.o g/10.1186/s4053 7-019-0276-2.
3. Gu ie ez A, Ansua egi A, Suspe egi L, Tubío C, Rankić I, Lenža L. A Benchma king o lea ning s a egies o pes
de ec ion and iden ifica ion on oma o plan s o au onomous scou ing obo s using in e nal da abases. J Sen-
so s. 2019. h ps ://doi.o g/10.1155/2019/52194 71.
4. San os L, San os FN, Oli ei a PM, Shinde P. Deep lea ning applica ions in ag icul u e: a sho e iew. Ad ances in
in elligen sys ems and compu ing. Fou h Ibe. 2020. h ps ://doi.o g/10.1007/978-3-030-35990 -4_12.
5. Wang T, Chen Y, Qiao M, Snoussi H. A as and obus con olu ional neu al ne wo k-based de ec de ec ion model
in p oduc quali y con ol. In J Ad Manu ac u Technol. 2018;94:3465–71.
6. Hashemi M. Enla ging smalle images be o e inpu ing in o con olu ional neu al ne wo k: ze o-padding s in e -
pola ion. J Big Da a. 2019. h ps ://doi.o g/10.1186/s4053 7-019-0263-7.
7. Lecun Y, Bo ou L, Bengio Y, Haffne P. G adien -based lea ning applied o documen ecogni ion. P oceedings o
he IEEE . 1998;86:2278–324. h p://ieeex plo e .ieee.o g/docum en /72679 1/
8. Gi shick R, Donahue J, Da ell T, Malik J. Rich ea u e hie a chies o accu a e objec de ec ion and seman ic seg-
men a ion. 2014 IEEE Con e ence on Compu e Vision and Pa e n Recogni ion . IEEE; 2014. p. 580–7. h p://ieeex
plo e .ieee.o g/docum en /69094 75/
9. Long J, Shelhame E, Da ell T. Fully con olu ional ne wo ks o seman ic segmen a ion. 2015 IEEE Con e ence on
Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2015. p. 3431–40. h p://a xi .o g/abs/1605.06211
10. K izhe sky A, Su ske e I, Hin on GE. ImageNe classifica ion wi h deep con olu ional neu al ne wo ks. Ad Neu al
In o ma P ocess Sys . 2012;2:1097–105.
11. Simonyan K, Zisse man A. Ve y deep con olu ional ne wo ks o la ge-scale image ecogni ion. 3 d In e na ional
Con e ence on Lea ning Rep esen a ions, ICLR 2015–Con e ence T ack P oceedings. 2015;1–14.
12. Szegedy C, Liu W, Jia Y, Se mane P, Reed S, Anguelo D, e al. Going Deepe wi h Con olu ions. CoRR . 2014;
abs/1409.4. h ps ://a xi .o g/abs/1409.4842
13. He K, Zhang X, Ren S, Sun J. Deep esidual lea ning o image ecogni ion. P oceedings o he IEEE compu e
socie y con e ence on compu e ision and pa e n ecogni ion. 2016. p. 770–8. h p://a xi .o g/abs/1512.03385
14. Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z. Re hinking he incep ion a chi ec u e o compu e ision. 2016
IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2016. p. 2818–26. h p://a xi .o g/
abs/1512.00567
15. Huang G, Liu Z, Van De Maa en L, Weinbe ge KQ. Densely connec ed con olu ional ne wo ks. 2017 IEEE Con e -
ence on Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2017. p. 2261–9. h p://a xi .o g/abs/1608.06993
16. Buda M, Maki A, Mazu owski MA. A sys ema ic s udy o he class imbalance p oblem in con olu ional neu al
ne wo ks. Neu al Ne w. 2018;106:249–59. h ps ://linki nghub .else ie .com/ e i e e/pii/S0893 60801 83021 07
17. Al-S ouhi S, Reddy CK. T ans e lea ning o class imbalance p oblems wi h inadequa e da a. Knowl In o ma Sys .
2016;48:201–28. h ps ://doi.o g/10.1007/s1011 5-015-0870-3
18. Ali A, Shamsuddin SM, Ralescu AL. Classifica ion wi h class imbalance p oblem: a e iew. In J Ad So Compu
Applica . 2015;7:176–204.
19. Zhang J, Xia Y, Wu Q, Xie Y. Classifica ion o medical images and illus a ions in he biomedical li e a u e using
syne gic deep lea ning. 2017. h p://a xi .o g/abs/1706.09092
20. Dong Q, Gong S, Zhu X. Imbalanced deep lea ning by mino i y class inc emen al ec ifica ion. IEEE T ansac ions
on Pa e n Analysis and Machine In elligence . 2019;41:1367–81. h ps ://ieeex plo e .ieee.o g/docum en /83537 18
21. Zhang Y, Li B, Lu H, I ie A, Ruan X. Sample-Specific SVM lea ning o pe son e-iden ifica ion. 2016 IEEE Con e ence
on Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2016. p. 1278–87. h p://ieeex plo e .ieee.o g/docum
en /77805 12/
22. Sawan MM, Bhu chandi KM. Age in a ian ace ecogni ion: a su ey on acial aging da abases, echniques and
effec o aging. A ific In ell Re . 2019;52:981–1008. h ps ://doi.o g/10.1007/s1046 2-018-9661-z.
Page 53 o 59
Sampa he al. J Big Da a (2021) 8:27
23. Mos a a E, Ali A, Alajlan N, Fa ag A. Pose In a ian App oach o Face Recogni ion a Dis ance. Be lin : Sp inge ;
2012. p. 15–28. h ps ://doi.o g/10.1007/978-3-642-33783 -3_2.
24. Japkowicz N, S ephen S. The class imbalance p oblem: a sys ema ic s udy. In ell Da a Analy. 2002;6:429–49. h ps
://doi.o g/10.5555/12939 51.12939 54.
25. Chawla NV. Da a mining o imbalanced da ase s: an o e iew. da a mining and knowledge disco e y handbook.
New Yo k : Sp inge -Ve lag; 2009. p. 853–67. h ps ://doi.o g/10.1007/0-387-25465 -X_40.
26. Chawla NV, Japkowicz N, Ko cz A. Special Issue on Lea ning om Imbalanced Da a Se s. ACM SIGKDD Explo a ions
Newsle e . 2004; 6: 1–6. h ps ://doi.o g/10.1145/10077 30.10077 33
27. Chawla N V., Bowye KW, Hall LO, Kegelmeye WP. SMOTE: Syn he ic mino i y o e -sampling echnique. J A ific
In ell Res. 2011;16:321–57. h ps ://doi.o g/10.1613/jai .953. h ps ://a xi .o g/abs/1106.1813
28. Haibo He, Yang Bai, Ga cia EA, Shu ao Li. ADASYN: Adap i e syn he ic sampling app oach o imbalanced lea ning.
2008 IEEE In e na ional Join Con e ence on Neu al Ne wo ks (IEEE Wo ld Cong ess on Compu a ional In elli-
gence) . IEEE; 2008. p. 1322–8. h p://ieeex plo e .ieee.o g/docum en /46339 69/
29. Pun umapon K, Rak hamamon T, Waiyamai K. Clus e -based mino i y o e -sampling o imbalanced da ase s.
IEICE T ansac ions on In o ma ion and Sys ems . 2016;E99.D:3101–9. h ps ://www.js ag e.js .go.jp/a ic le/ ans in /
E99.D/12/E99.D_2016E DP713 0/_a ic le
30. Sima d PY, S eink aus D, Pla JC. Bes p ac ices o con olu ional neu al ne wo ks applied o isual documen
analysis. Se en h In e na ional Con e ence on Documen Analysis and Recogni ion, 2003 P oceedings . IEEE
Compu . Soc; p. 958–63. h p://ieeex plo e .ieee.o g/docum en /12278 01/
31. Lemley J, Baz a kan S, Co co an P. Deep Lea ning o Consume De ices and Se ices: Pushing he limi s o
machine lea ning, a ificial in elligence, and compu e ision. IEEE Consume Elec onics Magazine . 2017;6:48–56.
h p://ieeex plo e .ieee.o g/docum en /78794 02/
32. Sho en C, Khoshgo aa TM. A su ey on image da a augmen a ion o deep lea ning. J Big Da a. 2019;6:60. h ps
://doi.o g/10.1186/s4053 7-019-0197-0.
33. Wu H, P asad S. Semi-Supe ised Deep Lea ning Using Pseudo Labels o Hype spec al Image Classifica ion. IEEE
T ansac ions on Image P ocessing . 2018;27:1259–70. h p://ieeex plo e .ieee.o g/docum en /81058 56/
34. an Engelen JE, Hoos HH. A su ey on semi-supe ised lea ning. Mach Lea n. 2020;109:373–440. h ps ://doi.
o g/10.1007/s1099 4-019-05855 -6.
35. Thai-Nghe N, Gan ne Z, Schmid -Thieme L. Cos -sensi i e lea ning me hods o imbalanced da a. The 2010
In e na ional Join Con e ence on Neu al Ne wo ks (IJCNN) . IEEE; 2010. p. 1–8. h p://ieeex plo e .ieee.o g/docum
en /55964 86/
36. Gi shick R. Fas R-CNN. 2015 IEEE In e na ional Con e ence on Compu e Vision (ICCV) . IEEE; 2015. p. 1440–8.
h p://ieeex plo e .ieee.o g/docum en /74105 26/
37. Ren S, He K, Gi shick R, Sun J. Fas e R-CNN: Towa ds Real-Time Objec De ec ion wi h Region P oposal Ne wo ks.
IEEE T ansac ions on Pa e n Analysis and Machine In elligence . 2017;39:1137–49. h p://ieeex plo e .ieee.o g/
docum en /74858 69/
38. He K, Gkioxa i G, Dolla P, Gi shick R. Mask R-CNN. IEEE T ansac ions on pa e n analysis and machine in elligence.
2020;42:386–97. h ps ://ieeex plo e .ieee.o g/docum en /83726 16/
39. Liu W, Anguelo D, E han D, Szegedy C, Reed S, Fu C-Y, e al. SSD: Single Sho Mul iBox De ec o . In: Leibe B,
Ma as J, Sebe N, Welling M, edi o s. Cham: Sp inge In e na ional Publishing; 2016. p. 21–37. Doi: h ps ://doi.
o g/10.1007/978-3-319-46448 -0_2
40. Redmon JSDRGAF. (YOLO) You Only Look Once. C p . 2016;
41. Yan X, Gong H, Jiang Y, Xia S-T, Zheng F, You X, e al. Video scene pa sing: an o e iew o deep lea ning me hods
and da ase s. Compu e Vision and Image Unde s anding . 2020;201:103077. h ps ://linki nghub .else ie .com/ e i
e e/pii/S1077 31422 03011 20
42. Hsu Y-W, Wang T-Y, Pe ng J-W. Passenge flow coun ing in buses based on deep lea ning using su eillance ideo.
Op ik . 2020;202:163675. h ps ://linki nghub .else ie .com/ e i e e/pii/S0030 40261 93157 36
43. Singh B, Da is LS. An analysis o scale in a iance in objec de ec ion–SNIP. 2018 IEEE/CVF Con e ence on compu e
ision and pa e n ecogni ion. IEEE; 2018. p. 3578–87. h ps ://ieeex plo e .ieee.o g/docum en /85784 75/
44. Yang F, Choi W, Lin Y. Exploi All he Laye s: Fas and Accu a e CNN objec de ec o wi h scale dependen pooling
and cascaded ejec ion classifie s. 2016 IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR) .
IEEE; 2016. p. 2129–37. h p://ieeex plo e .ieee.o g/docum en /77806 03/
45. Singh B, Najibi M, Da is LS. SNIPER: Efficien Mul i-Scale T aining. 32nd con e ence on neu al in o ma ion p ocess-
ing sys ems. Mon éal; 2018. h p://a xi .o g/abs/1805.09300
46. Lin T-Y, Dolla P, Gi shick R, He K, Ha iha an B, Belongie S. Fea u e Py amid Ne wo ks o Objec De ec ion. 2017 IEEE
con e ence on compu e ision and pa e n ecogni ion (CVPR). IEEE; 2017. p. 936–44. h p://ieeex plo e .ieee.o g/
docum en /80995 89/
47. Lin T-Y, Goyal P, Gi shick R, He K, Dolla P. Focal Loss o Dense Objec De ec ion. IEEE T ansac ions on Pa e n
Analysis and Machine In elligence. 2020;42:318–27. h ps ://ieeex plo e .ieee.o g/docum en /84179 76/
48. Dolla P, Wojek C, Schiele B, Pe ona P. Pedes ian de ec ion: a benchma k. 2009 IEEE Con e ence on Compu e
Vision and Pa e n Recogni ion . IEEE; 2009. p. 304–11. h ps ://ieeex plo e .ieee.o g/docum en /52066 31/
49. Zhong Z, Zheng L, Kang G, Li S, Yang Y. Random E asing Da a Augmen a ion. 2017. h p://a xi .o g/
abs/1708.04896
50. Wang X, Sh i as a a A, Gup a A. A-Fas -RCNN: Ha d posi i e gene a ion ia ad e sa y o objec de ec ion. 2017
IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR). IEEE; 2017. p. 3039–48. h p://a xi .o g/
abs/1704.03414
51. Bad ina ayanan V, Kendall A, Cipolla R. SegNe : A deep con olu ional encode -decode a chi ec u e o image
segmen a ion. IEEE T ansac ions on Pa e n Analysis and Machine In elligence. 2017;39:2481–95. h p://a xi .o g/
abs/1511.00561
52. Ronnebe ge O, Fische P, B ox T. U-Ne : Con olu ional ne wo ks o biomedical image segmen a ion. 2015. p.
234–41. h p://a xi .o g/abs/1505.04597
Page 54 o 59
Sampa he al. J Big Da a (2021) 8:27
53. Diakogiannis FI, Waldne F, Cacce a P, Wu C. ResUNe -a: A deep lea ning amewo k o seman ic segmen a ion
o emo ely sensed da a. ISPRS Jou nal o Pho og amme y and Remo e Sensing . 2020;162:94–114. h ps ://linki
nghub .else ie .com/ e i e e/pii/S0924 27162 03001 49
54. Yu se e E, Lambe J, Ca ballo A, Takeda K. A su ey o au onomous d i ing: common p ac ices and eme ging
echnologies. 2019. h p://a xi .o g/abs/1906.05113
55. Tabe nik D, Šela S, Sk a č J, Skočaj D. Segmen a ion-based deep-lea ning app oach o su ace-de ec de ec ion.
2019. h p://a xi .o g/abs/1903.08536
56. Rizwan I Haque I, Neube J. Deep lea ning app oaches o biomedical image segmen a ion. In o ma ics in Medi-
cine Unlocked. 2020;18:100297. h ps ://linki nghub .else ie .com/ e i e e/pii/S2352 91481 93021 4X
57. Co d s M, Om an M, Ramos S, Reh eld T, Enzweile M, Benenson R, e al. The ci yscapes da ase o seman ic u ban
scene unde s anding. P oceedings o he IEEE Compu e Socie y Con e ence on Compu e Vision and Pa e n
Recogni ion. 2016;2016-Decem:3213–23.
58. Menze BH, Jakab A, Baue S, Kalpa hy-C ame J, Fa ahani K, Ki by J, e al. The mul imodal b ain umo image
segmen a ion benchma k (BRATS). IEEE T ansac Med Imag. 2015;34:1993–2024. h p://ieeex plo e .ieee.o g/docum
en /69752 10/
59. Mu phy KP. Machine lea ning: a p obabilis ic pe spec i e (Adap i e Compu a ion and Machine Lea ning se ies).
Camb idge: The MIT P ess; 2012.
60. Mille a i F, Na ab N, Ahmadi S-A. V-Ne : Fully con olu ional neu al ne wo ks o olume ic medical image seg-
men a ion. 2016 Fou h In e na ional Con e ence on 3D Vision (3DV) . IEEE; 2016. p. 565–71. h p://ieeex plo e .ieee.
o g/docum en /77851 32/
61. C um WR, Cama a O, Hill DLG. Gene alized O e lap Measu es o E alua ion and Valida ion in Medical Image
Analysis. IEEE T ansac Med Imag. 2006;25:1451–61. h p://ieeex plo e .ieee.o g/docum en /17176 43/
62. Salehi SSM, E dogmus D, Gholipou A. T e sky loss unc ion o image segmen a ion using 3D ully con olu ional
deep ne wo ks. 2017. p. 379–87. h p://a xi .o g/abs/1706.05721
63. Be man M, T iki AR, Blaschko MB. The Lo asz-So max Loss: A ac able su oga e o he op imiza ion o he
in e sec ion-o e -union measu e in neu al ne wo ks. 2018 IEEE/CVF Con e ence on Compu e Vision and Pa e n
Recogni ion . IEEE; 2018. p. 4413–21. h ps ://ieeex plo e .ieee.o g/docum en /85785 62/
64. He Z, Zuo W, Kan M, Shan S, Chen X. A GAN: Facial a ibu e edi ing by only changing wha you wan . IEEE ans-
ac ions on image p ocessing . 2019;28:5464–78. h ps ://ieeex plo e .ieee.o g/docum en /87185 08/
65. Pe a nau G, an de Weije J, Raducanu B, Ál a ez JM. In e ible Condi ional GANs o image edi ing. Con e ence on
Neu al In o ma ion P ocessing Sys ems . 2016. h p://a xi .o g/abs/1611.06355
66. Tao R, Li Z, Tao R, Li B. ResA -GAN: Unpai ed deep esidual a ibu es lea ning o mul i-domain ace image ans-
la ion. IEEE Access . 2019;7:132594–608. h ps ://ieeex plo e .ieee.o g/docum en /88365 02/
67. Good ellow IJ, Pouge -Abadie J, Mi za M, Xu B, Wa de-Fa ley D, Ozai S, e al. Gene a i e ad e sa ial ne s. Ad
Neu al In P ocess Sys . 2014;3:2672–80.
68. Bowles C, Chen L, Gue e o R, Ben ley P, Gunn R, Hamme s A, e al. GAN Augmen a ion: augmen ing aining da a
using gene a i e ad e sa ial ne wo ks. 2018; h p://a xi .o g/abs/1810.10863
69. Oo d A an den, Kalchb enne N, Ka ukcuoglu K. Pixel ecu en neu al ne wo ks. 2016; h p://a xi .o g/
abs/1601.06759
70. Sejnowski MIJTJ. Lea ning and elea ning in bol zmann machines. G aphical models: ounda ions o neu al com-
pu a ion, MITP. 2001;
71. McClelland DERJL. In o ma ion p ocessing in dynamical sys ems: ounda ions o ha mony heo y. pa allel dis ib-
u ed p ocessing: explo a ions in he mic os uc u e o Cogni ion: Founda ions, MITP. 1987;194–281.
72. Hin on GE, Salakhu dino RR. Reducing he dimensionali y o da a wi h neu al ne wo ks. Science. 2006;313:504–7.
73. Salakhu dino R, Hin on G. Deep Bol zmann machines. J Machine Lea n Res. 2009;5:448–55.
74. Lee H, G osse R, Rangana h R, Y. Ng A. Con olu ional deep belie ne wo ks o scalable unsupe ised lea ning o
hie a chical ep esen a ions. Compu e Science Depa men , S an o d Uni e si y . 2009;8. h p:// obo ics.s an o d.
edu/~ang/pape s/icml0 9-Con o lu io nalDe epBel ie Ne wo k s.pd
75. Hin on GE, Osinde o S, Teh Y-W. A as lea ning algo i hm o deep belie ne s. Neu al Compu . 2006;18:1527–54.
h ps ://doi.o g/10.1162/neco.2006.18.7.1527.
76. Ramachand an P, Paine T Le, Kho ami P, Babaeizadeh M, Chang S, Zhang Y, e al. Fas gene a ion o con olu ional
au o eg essi e models. 2017; h p://a xi .o g/abs/1704.06001
77. F ey BJ. G aphical models o machine lea ning and digi al communica ion. Camb idge: MIT P ess; 1998.
78. F ey BJ, Hin on GE, Dayan P. Does he Wake-sleep algo i hm p oduce good densi y es ima o s? Ad ances in neu al
in o ma ion p ocessing sys ems . 1996;13:661–70. h p://www.cs.u o o n o.ca/~hin o n/absps /wspe .pd %5Cnpa
pe s2 ://publi ca io n/uuid/BCC05 47E-7C14-42EC-8693-D800C 5819C 79
79. U ia B, Cô é M-A, G ego K, Mu ay I, La ochelle H. Neu al au o eg essi e dis ibu ion es ima ion. J Mach Lea n Res.
2016;17:1–37. h p://a xi .o g/abs/1605.02226
80. Schulle B, Wöllme M, Moosmay T, Rigoll G. Recogni ion o noisy speech: a compa a i e su ey o obus model
a chi ec u e and ea u e enhancemen . EURASIP J Audio Speech Music P ocess. 2009;2009:942617. h p://asmp.
eu as ipjou nals .com/con e n /2009/1/94261 7
81. Yang S, Lu H, Kang S, Xue L, Xiao J, Su D, e al. On he localness modeling o he sel -a en ion based end- o-end
speech syn hesis. Neu al Ne w. 2020;125:121–30. h ps ://linki nghub .else ie .com/ e i e e/pii/S0893 60802 03004 47
82. Ghosh R, Vamshi C, Kuma P. RNN based online handw i en wo d ecogni ion in De anaga i and Bengali sc ip s
using ho izon al zoning. Pa e n Recogni . 2019;92:203–18. h ps ://linki nghub .else ie .com/ e i e e/pii/S0031
32031 93013 84
83. Chen J, Zhuge H. Ex ac i e summa iza ion o documen s wi h images based on mul i-modal RNN. Fu u e Gen-
e a Compu Sys . 2019;99:186–96. h ps ://linki nghub .else ie .com/ e i e e/pii/S0167 739X1 83268 76
84. Hoch ei e S, Schmidhube J. Long sho - e m memo y. Neu al Compu . 1997;9:1735–80. h ps ://doi.o g/10.1162/
neco.1997.9.8.1735.
Page 55 o 59
Sampa he al. J Big Da a (2021) 8:27
85. Vaswani A, Shazee N, Pa ma N, Uszko ei J, Jones L, Gomez AN, e al. A en ion is all you need. a Xi . 2017; h p://
a xi .o g/abs/1706.03762
86. Theis L, Be hge M. Gene a i e Image Modeling Using Spa ial LSTMs. P oceedings o he 28 h In e na ional Con e -
ence on Neu al In o ma ion P ocessing Sys ems–Volume 2. Camb idge: MIT P ess; 2015. p. 1927–1935.
87. K izhe sky A. Lea ning mul iple laye s o ea u es om iny images . 2009. h p://www.cs. o on o.edu/~k iz/ci a
.h ml
88. Russako sky O, Deng J, Su H, K ause J, Sa heesh S, Ma S, e al. ImageNe la ge scale isual ecogni ion challenge.
In J Compu Vis. 2015;115:211–52. h ps ://doi.o g/10.1007/s1126 3-015-0816-y.
89. Oo d A an den, Kalchb enne N, Vinyals O, Espehol L, G a es A, Ka ukcuoglu K. Condi ional image gene a ion
wi h PixelCNN Decode s. h p://a xi .o g/abs/1606.05328
90. Salimans T, Ka pa hy A, Chen X, Kingma DP. PixelCNN++: Imp o ing he PixelCNN wi h disc e ized logis ic mix u e
likelihood and o he modifica ions. 2017; h p://a xi .o g/abs/1701.05517
91. Chen X, Mish a N, Rohaninejad M, Abbeel P. PixelSNAIL: an imp o ed au o eg essi e gene a i e model. 2017.
h p://a xi .o g/abs/1712.09763
92. Vincen P, La ochelle H, Bengio Y, Manzagol P-A. Ex ac ing and composing obus ea u es wi h denoising au oen-
code s. P oceedings o he 25 h in e na ional con e ence on Machine lea ning - ICML ’08 . New Yo k: ACM P ess;
2008. p. 1096–103. h ps ://linki nghub .else ie .com/ e i e e/pii/S0925 23121 83061 55
93. Baldi P. Au oencode s, unsupe ised lea ning, and deep a chi ec u es . PMLR; 2012. h p://p oce eding s.ml .p ess /
27/baldi 12a.h ml
94. Y. Ng A. Spa se au oencode .h ps ://web.s an o d.edu/class /cs294 a/spa s eAu o encod e .pd
95. Masci J, Meie U, Ci eşan D, Schmidhube J. S acked con olu ional au o-encode s o hie a chical ea u e ex ac-
ion. 2011. p. 52–9. h ps ://doi.o g/10.1007/978-3-642-21735 -7_7
96. Ri ai S, Vincen P, Mulle X, Glo o X, Bengio Y. Con ac i e au o-encode s: explici in a iance du ing ea u e ex ac-
ion. ICML. 2011.
97. Kingma DP, Welling M. Au o-encoding a ia ional bayes. 2013; h p://a xi .o g/abs/1312.6114
98. Tan S, Li B. S acked con olu ional au o-encode s o s eganalysis o digi al images. Signal and In o ma ion P ocess-
ing Associa ion Annual Summi and Con e ence (APSIPA), 2014 Asia-Pacific. IEEE; 2014. p. 1–4.
99. Ge main M, G ego K, Mu ay I, La ochelle H. MADE: Masked au oencode o dis ibu ion es ima ion. 2015. h p://
a xi .o g/abs/1502.03509
100. Schmidhube J. Lea ning ac o ial codes by p edic abili y minimiza ion. Neu al Compu . 1992;4:863–79. h ps ://
doi.o g/10.1162/neco.1992.4.6.863.
101. Sohn K, Yan X, Lee H. Lea ning s uc u ed ou pu ep esen a ion using deep condi ional gene a i e models. Ad
Neu al In o ma P ocess Sys . 2015;2015-Janua:3483–91.
102. Higgins I, Ma hey L, Pal A, Bu gess C, Glo o X, Bo inick M, e al. Β-VAE: Lea ning basic isual concep s wi h a
cons ained a ia ional amewo k. 5 h In e na ional Con e ence on Lea ning Rep esen a ions, ICLR 2017–Con e -
ence T ack P oceedings. 2019;1–13.
103. Kulka ni TD, Whi ney W, Kohli P, Tenenbaum JB. Deep con olu ional in e se g aphics ne wo k. 2015. h p://a xi
.o g/abs/1503.03167
104. Huang C-W, Sanka an K, Dhekane E, Lacos e A, Cou ille A. Hie a chical Impo ance Weigh ed Au oencode s. In:
Chaudhu i K, Salakhu dino R, edi o s. Long Beach, Cali o nia, USA: PMLR; 2019. p. 2869–78. h p://p oce eding
s.ml .p ess / 97/huang 19d.h ml
105. Gul ajani I, Kuma K, Ahmed F, Taiga AA, Visin F, Vazquez D, e al. PixelVAE: A la en a iable model o na u al
images. 2016; Ah p://a xi .o g/abs/1611.05013
106. Chen X, Kingma DP, Salimans T, Duan Y, Dha iwal P, Schulman J, e al. Va ia ional Lossy Au oencode . 2016. h p://
a xi .o g/abs/1611.02731
107. G ego K, Danihelka I, G a es A, Rezende DJ, Wie s a D. DRAW: A ecu en neu al ne wo k o image gene a ion.
2015. h p://a xi .o g/abs/1502.04623
108. Oo d A an den, Vinyals O, Ka ukcuoglu K. Neu al Disc e e Rep esen a ion Lea ning. 31s Con e ence on Neu al
In o ma ion P ocessing Sys ems . Long Beach, Cali o nia, USA; 2017. h p://a xi .o g/abs/1711.00937
109. Raza i A, Oo d A an den, Vinyals O. Gene a ing di e se high-fideli y images wi h VQ-VAE-2. Ad ances in neu al
in o ma ion p ocessing sys ems 32. 2019. h p://a xi .o g/abs/1906.00446
110. Huszá F. How (no ) o T ain you gene a i e model: scheduled sampling, likelihood, ad e sa y? 2015. h p://a xi
.o g/abs/1511.05101
111. Lo e W, K eiman G, Cox D. Deep P edic i e coding ne wo ks o ideo p edic ion and unsupe ised lea ning.
2016. h p://a xi .o g/abs/1605.08104
112. Rad o d A, Me z L, Chin ala S. Unsupe ised ep esen a ion lea ning wi h deep con olu ional gene a i e ad e -
sa ial ne wo ks. 2015. h p://a xi .o g/abs/1511.06434
113. Makhzani A, Shlens J, Jai ly N, Good ellow I, F ey B. Ad e sa ial Au oencode s. 2015; A ailable om: h p://a xi
.o g/abs/1511.05644
114. Dumoulin V, Belghazi I, Poole B, Mas opie o O, Lamb A, A jo sky M, e al. Ad e sa ially Lea ned In e ence. 2016.
h p://a xi .o g/abs/1606.00704
115. La sen ABL, Sønde by SK, La ochelle H, Win he O. Au oencoding beyond pixels using a lea ned simila i y me ic.
2015. h p://a xi .o g/abs/1512.09300
116. Zhong G, Gao W, Liu Y, Yang Y. Gene a i e Ad e sa ial ne wo ks wi h decode -encode ou pu noise. 2018. h p://
a xi .o g/abs/1807.03923
117. S i as a a A, Valko L, Russell C, Gu mann MU, Su on C. VEEGAN: Reducing Mode Collapse in GANs using implici
a ia ional lea ning. 2017. h p://a xi .o g/abs/1705.07761
118. Mi za M, Osinde o S. Condi ional gene a i e ad e sa ial ne s. 2014. h p://a xi .o g/abs/1411.1784
119. Odena A, Olah C, Shlens J. Condi ional image syn hesis wi h auxilia y classifie GANs. 2016. h p://a xi .o g/
abs/1610.09585
Page 56 o 59
Sampa he al. J Big Da a (2021) 8:27
120. Baz a kan S, Co co an P. Ve sa ile auxilia y classifie wi h gene a i e ad e sa ial ne wo k (VAC+GAN), Mul i Class
Scena ios. 2018. h p://a xi .o g/abs/1806.07751
121. Chen X, Duan Y, Hou hoo R, Schulman J, Su ske e I, Abbeel P. In oGAN: In e p e able ep esen a ion lea ning by
in o ma ion maximizing gene a i e ad e sa ial ne s. 2016. h p://a xi .o g/abs/1606.03657
122. Li X, Chen L, Wang L, Wu P, Tong W. SCGAN: disen angled ep esen a ion lea ning by adding simila i y cons ain
on gene a i e ad e sa ial ne s. IEEE Access . 2019;7:147928–38. h ps ://ieeex plo e .ieee.o g/docum en /84762 90/
123. A jo sky M, Chin ala S, Bo ou L. Wasse s ein GAN. 2017. h p://a xi .o g/abs/1701.07875
124. Gul ajani I, Ahmed F, A jo sky M, Dumoulin V, Cou ille A. Imp o ed aining o Wasse s ein GANs. 2017. h p://
a xi .o g/abs/1704.00028
125. Pe zka H, Fische A, Luko nico D. On he egula iza ion o Wasse s ein GANs. 2017. h p://a xi .o g/
abs/1709.08894
126. Mao X, Li Q, Xie H, Lau RYK, Wang Z, Smolley SP. Leas squa es gene a i e ad e sa ial ne wo ks. 2016. h p://a xi
.o g/abs/1611.04076
127. Zhao J, Ma hieu M, LeCun Y. Ene gy-based Gene a i e Ad e sa ial Ne wo k. 2016. h p://a xi .o g/abs/1609.03126
128. Be helo D, Schumm T, Me z L. BEGAN: Bounda y Equilib ium Gene a i e Ad e sa ial Ne wo ks. 2017. h p://a xi
.o g/abs/1703.10717
129. Wang R, Cully A, Chang HJ, Demi is Y. MAGAN: Ma gin adap a ion o gene a i e ad e sa ial ne wo ks. 2017. h p://
a xi .o g/abs/1704.03817
130. Zhao J, Xiong L, Jayash ee K, Li J, Zhao F, Wang Z, e al. Dual-agen GANs o pho o ealis ic and iden i y p ese ing
p ofile ace syn hesis. Ad an Neu al In o ma P ocess Sys . 2017;2017:66–76.
131. Ka as T, Aila T, Laine S, Leh inen J. P og essi e g owing o GANs o imp o ed quali y, s abili y, and a ia ion. 2017;
h p://a xi .o g/abs/1710.10196
132. Den on E, Chin ala S, Szlam A, Fe gus R. Deep gene a i e image models using a laplacian py amid o ad e sa ial
ne wo ks. Ad ances in Neu al In o ma ion P ocessing Sys ems 28 . 2015. h p://a xi .o g/abs/1506.05751
133. Im DJ, Kim CD, Jiang H, Memise ic R. Gene a ing images wi h ecu en ad e sa ial ne wo ks. 2016; h p://a xi
.o g/abs/1602.05110
134. Nguyen TD, Le T, Vu H, Phung D. Dual disc imina o gene a i e ad e sa ial Ne s. 2017; h p://a xi .o g/
abs/1709.03831
135. Ghosh A, Kulha ia V, Namboodi i V, To PHS, Dokania PK. Mul i-agen di e se gene a i e ad e sa ial ne wo ks.
2017. h p://a xi .o g/abs/1704.02906
136. Liu M-Y, Tuzel O. Coupled gene a i e ad e sa ial ne wo ks. con e ence on neu al in o ma ion p ocessing sys ems.
2016. h p://a xi .o g/abs/1606.07536
137. Kim T, Cha M, Kim H, Lee JK, Kim J. Lea ning o disco e c oss-domain ela ions wi h gene a i e ad e sa ial ne -
wo ks. 2017. h p://a xi .o g/abs/1703.05192
138. Zhu J-Y, Pa k T, Isola P, E os AA. Unpai ed Image- o-image ansla ion using cycle-consis en ad e sa ial ne -
wo ks. 2017 IEEE In e na ional Con e ence on Compu e Vision (ICCV) . IEEE; 2017. p. 2242–51. h p://a xi .o g/
abs/1703.10593
139. Ledig C, Theis L, Husza F, Caballe o J, Cunningham A, Acos a A, e al. Pho o- ealis ic single image supe - esolu ion
using a gene a i e ad e sa ial ne wo k. 2016; h p://a xi .o g/abs/1609.04802
140. Simonyan K, Zisse man A. Ve y deep con olu ional ne wo ks o la ge-scale image ecogni ion. 2014; h p://a xi
.o g/abs/1409.1556
141. Zhang H, Good ellow I, Me axas D, Odena A. Sel -A en ion Gene a i e Ad e sa ial Ne wo ks. 2018; h p://a xi
.o g/abs/1805.08318
142. Isola P, Zhu J-Y, Zhou T, E os AA. Image- o-image ansla ion wi h condi ional ad e sa ial ne wo ks. 2017 IEEE
Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR). IEEE; 2017. p. 5967–76. h p://ieeex plo e .ieee.
o g/docum en /81001 15/
143. Wang T-C, Liu M-Y, Zhu J-Y, Tao A, Kau z J, Ca anza o B. High- esolu ion image syn hesis and seman ic manipula-
ion wi h condi ional GANs. 2018 IEEE/CVF Con e ence on Compu e Vision and Pa e n Recogni ion . IEEE; 2018.
p. 8798–807. h ps ://ieeex plo e .ieee.o g/docum en /85790 15/
144. Bellema e MG, Danihelka I, Dabney W, Mohamed S, Lakshmina ayanan B, Hoye S, e al. The c ame dis ance as a
solu ion o biased wasse s ein g adien s. 2017. h p://a xi .o g/abs/1705.10743
145. M oueh Y, Se cu T, Goel V. McGan: mean and co a iance ea u e ma ching GAN. 2017. h p://a xi .o g/
abs/1702.08398
146. Li C-L, Chang W-C, Cheng Y, Yang Y, Póczos B. MMD GAN: owa ds deepe unde s anding o momen ma ching
ne wo k. 2017. h p://a xi .o g/abs/1705.08584
147. M oueh Y, Se cu T. Fishe GAN. 2017. h p://a xi .o g/abs/1705.09675
148. Salimans T, Good ellow I, Za emba W, Cheung V, Rad o d A, Chen X. Imp o ed echniques o aining GANs. 2016.
h p://a xi .o g/abs/1606.03498
149. Sønde by CK, Caballe o J, Theis L, Shi W, Huszá F. Amo ised MAP in e ence o image supe - esolu ion. 2016.
h p://a xi .o g/abs/1610.04490
150. Heusel M, Ramsaue H, Un e hine T, Nessle B, Hoch ei e S. GANs ained by a wo ime-scale upda e ule con-
e ge o a local nash equilib ium. 2017. h p://a xi .o g/abs/1706.08500
151. Miya o T, Ka aoka T, Koyama M, Yoshida Y. Spec al no maliza ion o gene a i e ad e sa ial ne wo ks. 2018. h p://
a xi .o g/abs/1802.05957
152. Hea h M, Bowye K, Kopans D, Moo e R, Kegelmeye WP. Digi al da abase o sc eening mammog aphy . h ps ://
www.mammo image .o g/da ab ases/
153. Shoohi LM, Saud JH. Dcgan o handling imbalanced mala ia da ase based on o e -sampling echnique and
using cnn. Medico-Legal Upda e. 2020;20:1079–85.
154. Niu S, Li B, Wang X, Lin H. De ec image sample gene a ion Wi h GAN o Imp o ing de ec ecogni ion. IEEE T ans-
ac ions on Au oma ion Science and Enginee ing . 2020;1–12. h ps ://ieeex plo e .ieee.o g/docum en /90008 06/
Page 57 o 59
Sampa he al. J Big Da a (2021) 8:27
155. Ma iani G, Scheidegge F, Is a e R, Bekas C, Malossi C. BAGAN: Da a Augmen a ion wi h Balancing GAN. 2018;
h p://a xi .o g/abs/1803.09655
156. Wu E, Wu K, Cox D, Lo e W. Condi ional infilling GANs o da a augmen a ion in mammog am classifica ion. 2018.
p. 98–106. Doi: h ps ://doi.o g/10.1007/978-3-030-00946 -5_11
157. Mu ama su C, Nishio M, Go o T, Oiwa M, Mo i a T, Yakami M, e al. Imp o ing b eas mass classifica ion by sha ed
da a wi h domain ans o ma ion using a gene a i e ad e sa ial ne wo k. Compu Biol Med. 2020;119:103698.
h ps ://linki nghub .else ie .com/ e i e e/pii/S0010 48252 03008 6X
158. Guan S. B eas cance de ec ion using syn he ic mammog ams om gene a i e ad e sa ial ne wo ks in con olu-
ional neu al ne wo ks. J Med Imag. 2019;6:1. h ps ://doi.o g/10.1117/1.JMI.6.3.03141 1. ull.
159. Waheed A, Goyal M, Gup a D, Khanna A, Al-Tu jman F, Pinhei o PR. Co idGAN: Da a augmen a ion using auxilia y
classifie GAN o imp o ed Co id-19 de ec ion. IEEE Access . 2020;8:91916–23. h ps ://ieeex plo e .ieee.o g/docum
en /90938 42/
160. COVID-19 Ches X-Ray da ase ini ia i e. h ps ://gi hu b.com/agchu ng/Figu e1-COVID -ches x ay-da as e
161. Cohen JP, Mo ison P, Dao L, Ro h K, Duong TQ, Ghassemi M. COVID-19 Image da a collec ion: p ospec i e p edic-
ions a e he u u e. 2020. h p://a xi .o g/abs/2006.11988
162. Co id19 adiog aphy da abase. h ps ://www.kaggl e.com/ awsi u a hman/co id 19- adio g aph y-da ab ase
163. Hase N, I o S, Kanaeko N, Sumi K. Da a augmen a ion o in a-class imbalance wi h gene a i e ad e sa ial
ne wo k. In: Cudel C, Bazeille S, Ve ie N, edi o s. Fou een h In e na ional Con e ence on Quali y Con ol by
A ificial Vision . SPIE; 2019. p. 56. A ailable om: h ps://www.spiedigi allib a y.o g/con e ence-p oceedings-o -
spie/11172/2521692/Da a-augmen a ion- o -in a-class-imbalance-wi h-gene a i e-ad e sa ial-ne wo k/h ps ://
doi.o g/10.1117/12.25216 92. ull
164. Donahue C, Lip on ZC, Balsub amani A, McAuley J. Seman ically Decomposing he La en Spaces o Gene a i e
Ad e sa ial Ne wo ks. 2017; h p://a xi .o g/abs/1705.07904
165. Wang Y, Gong D, Zhou Z, Ji X, Wang H, Li Z, e al. O hogonal deep ea u es decomposi ion o age-in a ian ace
ecogni ion. 2018. p. 764–79. h ps ://doi.o g/10.1007/978-3-030-01267 -0_45
166. Gong D, Li Z, Lin D, Liu J, Tang X. Hidden ac o analysis o age in a ian ace ecogni ion. 2013 IEEE In e na ional
Con e ence on Compu e Vision. IEEE; 2013. p. 2872–9. h p://ieeex plo e .ieee.o g/docum en /67514 68/
167. Yin X, Liu X. Mul i- ask con olu ional neu al ne wo k o pose-in a ian ace ecogni ion. IEEE T ansac ions on
Image P ocessing. 2018;27:964–75. h p://ieeex plo e .ieee.o g/docum en /80802 44/
168. Ca cagnì P, Del CM, Cazza o D, Leo M, Dis an e C. A s udy on diffe en expe imen al configu a ions o age, ace,
and gende es ima ion p oblems. EURASIP J Image Video P ocess. 2015;2015:37. h ps ://doi.o g/10.1186/s1364
0-015-0089-y.
169. Ziwei L, Ping L, Xiaogang W, Tang X. La ge-scale CelebFaces a ibu es (CelebA) Da ase . 2018. h p://mmlab .ie.
cuhk.edu.hk/p oje c s/Celeb A.h ml
170. Zhang J, Li A, Liu Y, Wang M. Ad e sa ially Regula ized U-Ne -based GANs o acial a ibu e modifica ion and
gene a ion. IEEE Access . 2019;7:86453–62. h ps ://ieeex plo e .ieee.o g/docum en /87547 28/
171. Zhang G, Kan M, Shan S, Chen X. Gene a i e ad e sa ial ne wo k wi h spa ial a en ion o ace a ibu e edi ing.
2018. p. 422–37. h ps ://doi.o g/10.1007/978-3-030-01231 -1_26
172. Zheng Z, Yang X, Yu Z, Zheng L, Yang Y, Kau z J. join disc imina i e and gene a i e lea ning o pe son e-iden i-
fica ion. 2019 IEEE/CVF Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2019. p. 2133–42.
h ps ://ieeex plo e .ieee.o g/docum en /89542 92/
173. Zhang X, Gao Y. Face ecogni ion ac oss pose: a e iew. pa e n ecogni ion . 2009;42:2876–96. h ps ://linki nghub
.else ie .com/ e i e e/pii/S0031 32030 90015 38
174. Tan X, Chen S, Zhou Z-H, Zhang F. Face ecogni ion om a single image pe pe son: a su ey. pa e n ecogni ion.
2006;39:1725–45. h ps ://linki nghub .else ie .com/ e i e e/pii/S0031 32030 60012 70
175. Zhao W, Chellappa R, Phillips PJ, Rosen eld A. Face ecogni ion. ACM compu ing su eys. 2003;35:399–458. h p://
po a l.acm.o g/ci a ion.c m?doid=95433 9.95434 2
176. Qian X, Fu Y, Xiang T, Wang W, Qiu J, Wu Y, e al. Pose-No malized Image Gene a ion o Pe son Re-iden ifica ion.
2018. p. 661–78. h ps ://doi.o g/10.1007/978-3-030-01240 -3_40
177. Wei L, Zhang S, Gao W, Tian Q. Pe son T ans e GAN o b idge domain gap o pe son e-iden ifica ion. 2018 IEEE/
CVF con e ence on compu e ision and pa e n ecogni ion . IEEE; 2018. p. 79–88. h ps ://ieeex plo e .ieee.o g/
docum en /85781 14/
178. Zhong Z, Zheng L, Zheng Z, Li S, Yang Y. Came a s yle adap a ion o pe son e-iden ifica ion. 2018 IEEE/CVF con-
e ence on compu e ision and pa e n ecogni ion. IEEE; 2018. p. 5157–66. h ps ://ieeex plo e .ieee.o g/docum
en /85786 39/
179. Deng W, Zheng L, Ye Q, Yang Y, Jiao J. Simila i y-p ese ing image-image domain adap a ion o pe son e-
iden ifica ion. 2018; h p://a xi .o g/abs/1811.10551
180. Ge Y, Li Z, Zhao H, Yin G, Yi S, Wang X, e al. FD-GAN: Pose-guided Fea u e Dis illing GAN o obus pe son e-
iden ifica ion. Ad Neu al In o ma P ocess Sys . 2018;2018:1222–33.
181. Zheng A, Lin X, Li C, He R, Tang J. A ibu es guided ea u e lea ning o ehicle e-iden ifica ion. 2019; h p://a xi
.o g/abs/1905.08997
182. Zhou Y, Shao L. C oss-View GAN Based Vehicle Gene a ion o Re-iden ifica ion. P ocedings o he B i ish Machine
Vision Con e ence 2017 . B i ish Machine Vision Associa ion; 2017. h p://www.bm a.o g/bm c/2017/pape s/
pape 186/index .h ml
183. Wu F, Yan S, Smi h JS, Zhang B. Vehicle e-iden ifica ion in s ill images: applica ion o semi-supe ised lea ning and
e- anking. Signal P ocessing: Image Communica ion . 2019;76:261–71. h ps ://linki nghub .else ie .com/ e i e e/
pii/S0923 59651 83058 00
184. Fu Y, Li X, Ye Y. A mul i- ask lea ning model wi h ad e sa ial da a augmen a ion o classifica ion o fine-g ained
images. Neu ocompu ing . 2020;377:122–9. h ps ://linki nghub .else ie .com/ e i e e/pii/S0925 23121 93137 48