scieee Science in your language
[en] (orig)

A survey on generative adversarial networks for imbalance problems in computer vision tasks

Abstract

Any computer vision application development starts off by acquiring images and data, then preprocessing and pattern recognition steps to perform a task. When the acquired images are highly imbalanced and not adequate, the desired task may not be achievable. Unfortunately, the occurrence of imbalance problems in acquired image datasets in certain complex real-world problems such as anomaly detection, emotion recognition, medical image analysis, fraud detection, metallic surface defect detection, disaster prediction, etc., are inevitable. The performance of computer vision algorithms can significantly deteriorate when the training dataset is imbalanced. In recent years, Generative Adversarial Neural Networks (GANs) have gained immense attention by researchers across a variety of application domains due to their capability to model complex real-world image data. It is particularly important that GANs can not only be used to generate synthetic images, but also its fascinating adversarial learning idea showed good potential in restoring balance in imbalanced datasets. In this paper, we examine the most recent developments of GANs based techniques for addressing imbalance problems in image data. The real-world challenges and implementations of synthetic image generation based on GANs are extensively covered in this survey. Our survey first introduces various imbalance problems in computer vision tasks and its existing solutions, and then examines key concepts such as deep generative image models and GANs. After that, we propose a taxonomy to summarize GANs based techniques for addressing imbalance problems in computer vision tasks into three major categories: 1. Image level imbalances in classification, 2. object level imbalances in object detection and 3. pixel level imbalances in segmentation tasks. We elaborate the imbalance problems of each group, and provide GANs based solutions in each group. Readers will understand how GANs based techniques can handle the problem of imbalances and boost performance of the computer vision algorithms. Sampath, V.; Maurtua, I.; Aguilar Martín, J.J.; Gutierrez, A.

Read accessible full text

A survey on generative adversarial networks for imbalance problems in computer vision tasks

Author: Sampath, V.; Aguilar Martín, J.J.; Maurtua, I.; Gutierrez, A.
Year: 2021
DOI: 10.1186/s40537-021-00414-0
Source: https://zaguan.unizar.es/record/99694/files/texto_completo.pdf
A su ey ongene a i e ad e sa ial ne wo ks
o imbalance p oblems incompu e ision
asks
Vignesh Sampa h1,2*, Iñaki Mau ua1, Juan José Aguila Ma ín2 and Ai o Gu ie ez1
In oduc ion
Recen de elopmen s in Con olu ional Neu al Ne wo ks (Con Ne s) ha e led o sub-
s an ial p og ess in he pe o mance o compu e ision asks applied ac oss a i-
ous domains such as sel -d i ing ca s [1], medical imaging [2], ag icul u e [3, 4],
Abs ac
Any compu e ision applica ion de elopmen s a s off by acqui ing images and
da a, hen p ep ocessing and pa e n ecogni ion s eps o pe o m a ask. When he
acqui ed images a e highly imbalanced and no adequa e, he desi ed ask may no be
achie able. Un o una ely, he occu ence o imbalance p oblems in acqui ed image
da ase s in ce ain complex eal-wo ld p oblems such as anomaly de ec ion, emo ion
ecogni ion, medical image analysis, aud de ec ion, me allic su ace de ec de ec ion,
disas e p edic ion, e c., a e ine i able. The pe o mance o compu e ision algo i hms
can significan ly de e io a e when he aining da ase is imbalanced. In ecen yea s,
Gene a i e Ad e sa ial Neu al Ne wo ks (GANs) ha e gained immense a en ion by
esea che s ac oss a a ie y o applica ion domains due o hei capabili y o model
complex eal-wo ld image da a. I is pa icula ly impo an ha GANs can no only be
used o gene a e syn he ic images, bu also i s ascina ing ad e sa ial lea ning idea
showed good po en ial in es o ing balance in imbalanced da ase s.
In his pape , we examine he mos ecen de elopmen s o GANs based echniques
o add essing imbalance p oblems in image da a. The eal-wo ld challenges and
implemen a ions o syn he ic image gene a ion based on GANs a e ex ensi ely co -
e ed in his su ey. Ou su ey fi s in oduces a ious imbalance p oblems in compu e
ision asks and i s exis ing solu ions, and hen examines key concep s such as deep
gene a i e image models and GANs. A e ha , we p opose a axonomy o summa ize
GANs based echniques o add essing imbalance p oblems in compu e ision asks
in o h ee majo ca ego ies: 1. Image le el imbalances in classifica ion, 2. objec le el
imbalances in objec de ec ion and 3. pixel le el imbalances in segmen a ion asks. We
elabo a e he imbalance p oblems o each g oup, and p o ide GANs based solu ions
in each g oup. Reade s will unde s and how GANs based echniques can handle he
p oblem o imbalances and boos pe o mance o he compu e ision algo i hms.
Keywo ds: Gene a i e ad e sa ial neu al ne wo ks, Imbalanced da a, Objec
de ec ion, Segmen a ion, Classifica ion, Deep lea ning, Deep gene a i e model
Open Access
© The Au ho (s) 2021. This a icle is licensed unde a C ea i e Commons A ibu ion 4.0 In e na ional License, which pe mi s use, sha ing,
adap a ion, dis ibu ion and ep oduc ion in any medium o o ma , as long as you gi e app op ia e c edi o he o iginal au ho (s) and
he sou ce, p o ide a link o he C ea i e Commons licence, and indica e i changes we e made. The images o o he hi d pa y ma e ial
in his a icle a e included in he a icle’s C ea i e Commons licence, unless indica ed o he wise in a c edi line o he ma e ial. I ma e ial
is no included in he a icle’s C ea i e Commons licence and you in ended use is no pe mi ed by s a u o y egula ion o exceeds he
pe mi ed use, you will need o ob ain pe mission di ec ly om he copy igh holde . To iew a copy o his licence, isi h p://c ea i eco
mmons .o g/licen ses/by/4.0/.
SURVEY PAPER
Sampa he al. J Big Da a (2021) 8:27
h ps://doi.o g/10.1186/s40537-021-00414-0
*Co espondence:
ignesh.sampa h@ eknike .es
1 Au onomous
and In elligen Sys ems Uni ,
Teknike , Membe o Basque
Resea ch and Technology
Alliance, Eiba , Spain
Full lis o au ho in o ma ion
is a ailable a he end o he
a icle
Page 2 o 59
Sampa he al. J Big Da a (2021) 8:27
manu ac u ing [5], e c. The a ailabili y o big da a [6], oge he wi h inc eased compu -
ing capabili ies is he p edominan eason o he ecen success. Image acquisi ion is he
fi s s ep in he de elopmen o compu e ision algo i hms. When he acqui ed image
is no adequa e, he desi ed ask may no be possible o achie e. Image classifica ion [7],
objec de ec ion [8] and segmen a ion [9] a e he undamen al building blocks o he
compu e ision asks. All hese me hods use deep Con Ne s wi h eno mous laye s and
ha e a e y high numbe o pa ame e s ha need o be uned. The e o e, hey demand
a huge amoun o ep esen a i e da a o imp o e hei pe o mance and gene aliza ion
abili y. While he amoun o isual da a is inc easing exponen ially, many o he eal-
wo ld da ase s suffe om se e al o ms o imbalance. Handling imbalances in he image
da ase is one o he pe asi e challenges in he field o compu e ision.
Image classifica ion is he ask o classi ying an inpu image acco ding o a se o pos-
sible classes. Classifica ion algo i hms lea n o isola e impo an dis inguishing in o -
ma ion abou an objec in an image like shape o colo and igno e i ele an pa s o
an image such as plane backg ound o noise. Se e al popula image classifica ion a chi-
ec u es such as LeNe [7], AlexNe [10], VGG-16 [11], GoogLeNe [12], ResNe [13],
Incep ion-V3 [14], DenseNe [15] ake an inpu image and hen pass i h ough se e al
con olu ional and pooling laye s. Con olu ional laye helps o ex ac ea u es om he
inpu image, while a pooling laye educes he dimension. Se e al successi e con olu-
ional and pooling laye s may ollow, depending on he layou and in en o he a chi-
ec u e. The esul is a se o ea u e maps educed in size om he o iginal image ha
h ough a aining p ocess ha e lea ned o dis ill in o ma ion abou he con en in he
o iginal image. All ex ac ed ea u e maps a e hen ans o med in o a single ec o ha
can be ed in o a se ies o ully connec ed neu al ne wo k o ob ain a p obabili y dis i-
bu ion o class sco es. The p edic ed class o he inpu image can be ex ac ed om his
p obabili y dis ibu ion.
These a chi ec u es a e ypically designed o wo k well wi h balanced da ase s, bu a
common issue wi h eal-wo ld da ase s is he imbalance o obse ed classes. The mos
commonly known imbalance p oblem in a ask o image classifica ion is he class imbal-
ance. Class imbalance in he eal-wo ld image da ase s is ubiqui ous and can ha e an
ad e se effec on he pe o mance o Con Ne s [16]. These da ase s usually all in o ou
ca ego ies in e ms o i s size and imbalance [17]:
1. The ideal da ase s a e he one ha con ain an adequa e and equal o almos equal
numbe o samples wi hin each class. An equal p obabili y is assigned o all classes
du ing aining o upda e pa ame e s o he ne wo k and app oach he minimum
alue o he e o unc ion. A wide ange o s anda d machine lea ning algo i hms
can be applied o he ideal da ase s.
2. The da ase s wi h an adequa e numbe o samples whe e some ins ances o classes
a e a e han o he ins ances o classes a e said o be une en da ase s. E en hough
hese da ase s ha e adequa e numbe o samples, i is cos ly and may no be possible
o expe s o manually inspec huge unlabeled da ase s o anno a e.
3. Tiny da ase s a e no easily a ailable, and hey can be difficul o collec . Such da a-
se s ha e an equal numbe o samples wi hin each class, bu hey a e almos impos-
sible o collec due o p i acy es ic ion and o he easons.
Page 3 o 59
Sampa he al. J Big Da a (2021) 8:27
4. Absolu e a e da ase s ha e a limi ed numbe o samples and subs an ial class imbal-
ance. Reasons o class imbalance in hese da ase s can a y bu commonly he p ob-
lem a ises because o : (a) Ve y limi ed numbe o expe s a ailable o da a collec-
ion; o an example, gene a ion o medical imaging da ase s equi es specialized
equipmen and well ained medical p ac i ione s o da a acquisi ion (b) Eno mous
manual effo equi ed o label da ase s; and (c) Sca ci y o samples o specific class
leading o class imbalance. Consequen ly, he size o he da ase and class imbal-
ance p oblem becomes a bo leneck ha p e en s us om apping he ue po en ial
o Con Ne s. Figu e1 illus a es diffe en ypes o da ase s in e ms o i s size and
imbalance.
Class imbalance in a da ase can s em om ei he be ween classes (in e class imbal-
ance) o wi hin class (in a class imbalance). In e class imbalance occu s when a mino -
i y class con ains a smalle numbe o ins ances when compa ed o ins ances belonging
o he majo i y class. Classifie s buil using in e class imbalanced da ase s a e mos
likely o p edic mino i y class as a e occu ences, e en some imes assumed as ou lie
o noise which esul s in misclassifica ion o mino i y classes [18]. Mino i y classes a e
o en o g ea e in e es and significance, ha needs o be cau iously handled. Fo exam-
ple, in a a e disease medical diagnosis whe e he e is a i al need o dis inguish such a
a e medical condi ion among he no mal popula ions. Any kind o diagnosis e o s will
cause s ess o he pa ien and u he complica ions. I is he e o e e y impo an ha
deep lea ning models [19] buil using such da ase s should be able o achie e a highe
de ec ion a e on mino i y classes.
In a class imbalance in a da ase can also de e io a e he pe o mance o he classi-
fie . An In a-class imbalance can be iewed as he a ibu e bias wi hin a class, in o he
wo ds in e -class imbalance in fine-g ained isual ca ego iza ion. Fo example, a class
o dog samples can be u he ca ego ized by dog colo , pose a ia ions and dog b eeds.
Imbalances in such ca ego ies (in a class imbalance) is an una oidable p oblem in
Fig. 1 Dis ibu ion o diffe en ype o da ase s (a) Da ase wi h adequa e sample (b) Da ase wi h
inadequa e sample
Page 4 o 59
Sampa he al. J Big Da a (2021) 8:27
da ase s o many classifica ion asks such as modali y based medical image classifica ion
[19], fine g ained a ibu e classifica ion [20], pe son e-iden ifica ion [21], age [22] and
pose in a ian ace ecogni ion [23].
Se e al a emp s ha e been made o o e come he p oblem o class imbalance by
using diffe en app oaches and echniques. These echniques can be g ouped in o
da a-le el app oaches, algo i hm le el me hods and hyb id echniques. While da a
le el app oaches modi y he dis ibu ion o aining se o es o e balance by adding o
emo ing ins ances om he aining da ase , algo i hm le el me hods change he objec-
i e unc ion o he classifie o inc ease he impo ance o he mino i y class. Hyb id
echniques combine algo i hm le el me hods wi h da a le el app oaches. Nex ew pa a-
g aphs will in o m eade s abou some o he adi ional echniques a ailable o coun e
he class imbalance p oblem.
• Resampling To coun e ac he class imbalance p oblem, wo ypes o e-sampling
can be applied: One is unde sampling by dele ing samples om he majo i y class
and ano he is o e sampling by duplica ing samples om he mino i y class [24].
Re-sampling me hod balances he da ase bu ails o p o ide any addi ional in o -
ma ion o he aining se . The o he limi a ions o his me hod include: o e sam-
pling esul s in o e fi ing p oblem while unde sampling leads o subs an ial loss
o in o ma ion [25]. The quan i y o unde -sampling and o e sampling is gene ally
de e mined using expe imen al me hods and empi ically es ablished [26]. In o de o
yield addi ional in o ma ion o he aining se , syn he ic o e sampling me hods c e-
a e new samples ins ead o duplica es o add equilib ium o skewed dis ibu ion. The
Syn he ic Mino i y O e sampling Technique (SMOTE) [27] is a popula syn he ic
o e sampling me hod ha aims o gene a e syn he ic samples based on andomly
selec ed K-nea es neighbo s. SMOTE does no ake accoun o he dis ibu ion o
da a be ween he classes. Adap i e syn he ic sampling (ADASYN) app oach [28]
uses a weigh ed dis ibu ion o diffe en mino i y classes acco ding o hei lea ning
difficul ies o adap i ely gene a e syn he ic da a samples. Clus e based o e sampling
[29] echnique di ides he inpu space in o a ious clus e s and hen inco po a es
sampling o al e he sample size. Many adi ional syn he ic o e sampling ech-
niques such as SMOTE o ADASYN a e only sui able o low dimensional abula
da a which es ic s hei applica ion in a high dimensional image da a. In addi ion,
all he a o emen ioned echniques gene a e da a by ei he dele ing o a e aging exis -
ing da a, and hence may ail o imp o e classifica ion pe o mance.
• Augmen a i e o e sampling Da a augmen a ion is ano he commonly used ech-
nique o infla e he size o he aining da ase [30]. Augmen a ion such as ansla-
ion, c opping, padding, o a ion and ho izon al flipping in oduces small modifica-
ions in he image da a, bu no all hese modifica ions will imp o e he pe o mance
o a classifie . The e is no s anda d me hod ha can decide whe he any pa icula
augmen a ion s a egy can imp o e esul s un il he aining p ocess is comple e.
As aining Con Ne s is a ime-consuming p ocess [31], only a es ic ed amoun
o augmen a ion s a egy is likely o be es ed be o e model deploymen . Also, he
di e si y ha can be ob ained om small modifica ions o he images is ela i ely
small. In addi ion o balancing classes by o e sampling, augmen a ion echniques
Page 5 o 59
Sampa he al. J Big Da a (2021) 8:27
also se e as a kind o egula iza ion in deep neu al ne wo k a chi ec u e and hence
educe he chance o o e fi ing. The e is no consensus abou he bes s a egy o
combining diffe en augmen a ion s a egies oge he . The e o e, mo e ad anced
augmen a ion echniques such as mixing images depend on expe knowledge o
alida ion and labelling [32]. A comple e su ey o Image da a augmen a ion o deep
lea ning has been compiled by Sho en e al. [32].
• Semi-supe ised lea ning (SSL) SSL [33] is one o he mos a ac i e ways o imp o e
classifica ion pe o mance whe e we ha e access o small numbe o labeled sam-
ples
x
x
along wi h la ge amoun o unlabeled samples (Une en da ase ). SSL uses
he combina ion o supe ised and unsupe ised lea ning echniques. I makes use
o small labeled samples as he aining se o ain he model in a supe ised man-
ne , and hen use he ained model o p edic on he emaining unlabeled po ion
o he da ase . The p ocess o labeling each sample o unlabeled da a wi h he indi-
idual ou pu s p edic ed o hem using he ained model is known as pseudo labe-
ling. A e labeling he unlabeled da a h ough he pseudo labeling p ocess, classifica-
ion model is ained on bo h he ac ual and pseudo labeled da a. Pseudo labeling is
an in e es ing pa adigm o anno a e la ge-scale unlabeled da a ha po en ially akes
many edious hou s o human labo o manually label hem. Howe e , SSL elies on
assump ions abou he unde lying ma ginal dis ibu ion o inpu da a
p
(x) , bo h he
labeled and unlabeled samples a e assumed o ha e he same ma ginal dis ibu ion.
This ma ginal dis ibu ion
p
(x) should con ain in o ma ion abou he pos e io dis-
ibu ion p(
y
|x) . A comple e lis o semi supe ised lea ning is de ailed in [34].
• Cos sensi i e lea ning Majo i y o he classifica ion algo i hms assume ha misclas-
sifica ion cos s o bo h mino i y and majo i y classes a e he same. Cos -sensi i e
lea ning [35] pays mo e a en ion o misclassifica ion cos s o he mino i y class
h ough a cos ma ix.
The mos s aigh o wa d and commonly used app oach in Con Ne s is he da a
d i en s a egy, because deep Con Ne s wi h eno mous laye s ha e a e y high num-
be o pa ame e s o be uned, i is p one o o e fi ing when ained on a small sized
da ase . Da a le el app oaches infla e he aining da a size ha se es as egula iza ion
and hence educe he chance o o e fi ing in deep neu al ne wo k a chi ec u e. T adi-
ional da a-le el echniques suffe he ollowing d awbacks, pa icula ly when used o
he class imbalance p oblem in high-dimensional image da a.
a. Syn he ic ins ances c ea ed using adi ional da a le el app oaches may no be he
ue ep esen a i e o he aining se .
b. Syn he ic da a gene a ion is achie ed ei he by duplica ion o linea in e pola ion
which does no gene a e new examples ha a e a ypical and puzzle he classifie
decision bounda ies, and hence ail o imp o e o e all pe o mance.
c. In Medical images, augmen a ion echniques a e es ic ed o mino al e a ion on
an image, as hey abide by s ic s anda ds. Addi ionally, he ypes o augmen a ion
one can use a y om p oblem o p oblem. Fo ins ance, hea y augmen a ions such
as geome ic ans o ma ions, andom e asing, and mixing images migh damage
seman ic con en o he medical image.

Page 6 o 59
Sampa he al. J Big Da a (2021) 8:27
d. Applying da a augmen a ion in an absolu e a e da ase may no p o ide he a ia-
ions equi ed o p oduce a dis inc sample o add equilib ium o skewed dis ibu-
ion.
e. Dealing wi h he class imbalance in fine-g ained isual ca ego iza ion is challenging
because i in ol es la ge in a-class a iabili y and small in e -class a iabili y.
. Mos o he echniques a e designed only o bina y classifica ion p oblems. Mul i
class imbalance p oblems a e gene ally conside ed much ha de han hei bina y
equi alen s o many easons. Fo Ins ance, he e can be se e al combina ions o
mino i y-majo i y classes, i.e., hey may include: 1. Few mino i y-Many majo i y
classes, 2. Many mino i y-Few majo i y classes, and 3. Many mino i y-Many majo -
i y classes.
Class imbalance in image classifica ion asks has been widely explo ed and s udied.
In addi ion o class imbalance, he e a e many diffe en o ms o imbalances ha can
impede pe o mance o o he compu e ision asks such as objec de ec ion and image
segmen a ion. Objec de ec ion, which deals wi h localiza ion and classifica ion o mul-
iple objec s in a gi en image, is ano he challenging and significan ask in compu e
ision. The ypical way o localizing an objec in an image is by d awing a bounding box
a ound he objec . This bounding box can be in e p e ed as a collec ion o coo dina es
ha define he box. Nowadays, objec de ec ion algo i hms all in o wo b oad ca ego-
ies: wo-s age de ec o s and single s age de ec o s. On one hand, wo s age de ec o
such as Region-based Con olu ional Neu al Ne wo ks (R-CNN) [8], Fas R-CNN [36],
Fas e R-CNN [37], Mask R-CNN [38], e c. employ a Region P oposal Ne wo k (RPN) o
sea ch objec s in he fi s s age, and hen p ocess hese egion o in e es s o objec clas-
sifica ion and bounding-box eg ession in he second s age. On he o he hand, single
s age de ec o s such as Single Sho De ec ion (SSD) [39], You Only Look Once (YOLO)
[40], e c. pe o m de ec ion on a g id ha a oids spending oo much ime on gene a ing
egion p oposals. Ins ead o loca ing objec s pe ec ly, hey p io i ize speed and ecogni-
ion. The e o e, one s age objec de ec o s a e as and simple, whe eas wo s age de ec-
o s a e mo e accu a e.
Despi e he ecen ad ances, applying objec de ec ion algo i hms o he eal-wo ld
da ase s such as in-ca ideo [41], anspo a ion su eillance images [42] ha con ain
objec s wi h la ge a iance o scales (Objec s scale imbalance) emains challenging.
Physical size o a same objec a diffe en dis ances om he came a would appea as
diffe en size. Singh e al. [43] showed ha objec le el scale a ia ion g ea ly affec s he
o e all pe o mance o objec de ec o s. Many solu ions ha e been p oposed o add ess
he objec scale imbalance. Scale awa e as R-CNN [44] uses an ensemble o wo objec
de ec o s, one o de ec ing he la ge and medium scale objec s and o he o he small
scale objec s, and hen combines hem o p oduce final p edic ions. Mul i-scale Image
Py amids such as SNIP [43] and SNIPER [45] use an image py amid o build mul i scale
ea u e ep esen a ion. Fea u e Py amid Ne wo ks (FPN) [46] combine ea u e hie a -
chies a diffe en scales o p edic objec s a diffe en scales.
Objec s in he eal-wo ld da ase s only occupy a small po ion o he image, while he
es o he image is backg ound. Bo h single and wo s age algo i hms app oxima ely
e alua e abou 104 o 105 loca ions pe image [47], ye jus a ew loca ions ha e objec s.
Page 7 o 59
Sampa he al. J Big Da a (2021) 8:27
The imbalance be ween o eg ound (objec ) and backg ound can also hinde pe o -
mance o he objec de ec ion algo i hm. Fu he mo e, objec de ec ion algo i hms
should be in a ian o de o ma ion and occluded objec s. In Pedes ian de ec ion Da a-
se [48], o ins ance, mo e han 70% o pedes ians a e occluded in a leas one ame o
a ideo clip and abou 19% o pedes ians a e occluded in all ames, whe e he occlu-
sions a e anked as hea y in almos hal o such cases. Dolla e al. [48] highligh ha
he pe o mance o pedes ian de ec ion using s anda d de ec o s declines subs an ially
e en unde pa ial occlusion, and d as ically unde se e e occlusion. Da a augmen a ion
based on andom e asing [49] is a equen ly used echnique ha o ces de ec o s o pay
a en ion o he en i e objec in an image, a he han jus a po ion o i . Ye , his ech-
nique is no gua an eed o be ad an ageous in all he condi ions. Because skewed dis i-
bu ions a ise e en wi hin de o med and occluded objec s as some o he occlusions and
de o ma ions a e uncommon ha hey ha dly occu in p ac ical scena ios [50].
Image segmen a ion ha classifies e e y pixel in an image suffe s om pixel le el
imbalances, as a e o he compu e ision asks.Some o he well-known image segmen-
a ion algo i hms include Fully connec ed ne wo k [9], SegNe [51], U-Ne [52], ResU-
Ne [53] e c. Image segmen a ion is essen ial o a a ie y o asks, including: U ban
scene segmen a ion o au onomous d i ing [54], indus ial inspec ion [55] and cance
cell segmen a ion [56]. Da ase s o all hese asks suffe om pixel le el imbalance. Fo
example, In U ban s ee scene da ase [57], Pixels co esponding o sky, building and
oad a e a nume ous han pixels o pedes ian and bicyclis . This is due o he ac ha
he a ea co e ed by sky, buildings and oads a e mo e han pedes ians and bicyclis s in
he image. Simila ly, In b ain umou image segmen a ion da ase [58], MRI images ha e
mo e heal hy b ain issue pixels han cance ous issue pixels. The mos equen ly used
loss unc ion o image segmen a ion ask is a pixel wise c oss en opy loss [59]. This loss
assigns equal weigh s o all he pixels, e alua es he p edic ion o each pixel indi idually
and hen a e ages o e all pixels. In o de o mi iga e his p oblem, many wo ksha e
been done which modi y he pixel wise c oss en opy loss unc ion. The s anda d c oss
en opy loss is modified in Weigh ed c oss en opy [52], Focal loss [47], Dice Loss [60],
Gene alised Dice Loss [61], T e sky loss [62], Lo ász-So max [63] and Median e-
quency balancing [51], so as o assign highe impo ance o a e pixels. Al hough modi-
fied loss unc ions a e efficien o some imbalances, such unc ions unde go se e e
difficul ies when i comes o highly imbalanced da ase s, as seen wi h medical image
segmen a ions.
In con as o all he adi ional app oaches desc ibed abo e, Gene a i e ad e sa ial
Neu al Ne wo ks (GANs) aim o lea n unde lying ue da a dis ibu ions om he lim-
i ed a ailable images (bo h mino i y and majo i y class), and hen use he lea ned dis-
ibu ions o gene a e syn he ic images. This aises an in e es ing ques ion on whe he
GANs can be used o gene a e syn he ic images o he mino i y class o a ious imbal-
anced da ase s. Indeed, ecen de elopmen s o GANs sugges ha being capable o
ep esen complex and high dimensional da a can be used as a me hod o in elligen
o e sampling. GANs u ilize he abili y o neu al ne wo ks o lea n a unc ion ha
can app oxima e model dis ibu ion as close as possible o ue dis ibu ion. Pa icu-
la ly, hey do no ely on p io assump ions abou he da a dis ibu ion and can gene -
a e syn he ic images wi h high isual fideli y. This significan p ope y allows GANs o
Page 8 o 59
Sampa he al. J Big Da a (2021) 8:27
be applied o any kind o imbalance p oblem in compu e ision asks. GANs can no
only be able o gene a e a ake image, bu also offe a way o change some hing abou
he o iginal image. In o he wo ds, hey can lea n o p oduce any desi ed numbe o
classes (such as, objec s, iden i ies, people, e c.), and ac oss many a ia ions (such as,
iewpoin s, ligh condi ions, scale, backg ounds, and mo e). The e a e a wide a ie y o
GANs epo ed in he li e a u e, each wi h hei own s eng hs o alle ia e imbalance
p oblem in compu e ision asks. Fo ins ance, A GAN [64], IcGAN [65], ResA -
GAN [66], e c. a e a specific a ian o GANs ha a e commonly used o acial a ibu e
edi ing asks. They lea n o syn hesize no only a new ace image wi h desi ed a ibu es
bu also p ese es a ibu e independen de ails. Recen ly, GANs ha e been combined
wi h a wide ange o exis ing objec de ec ion and image segmen a ion algo i hms o
o e come he p oblem o imbalance and imp o e hei pe o mance.
The o iginal GANs a chi ec u e [67] con ains wo diffe en iable unc ions ep esen ed
by wo ne wo ks, a gene a o
G
and a disc imina o
D
. The lea ning p ocedu e o GANs
is o simul aneously ain a disc imina o
D
and a gene a o
G
. I ollows an ad e sa ial
wo-playe , ze o-sum game. An in ui i e way o unde s anding GAN is wi h he police
and he coun e ei e anecdo e. The gene a o ne wo k is like a g oup o coun e ei e s
ying o p oduce ake money and make i look genuine. The police a emp o disco e
coun e ei e s using ake money, ye a he same ime need o le e e y o he pe son
spend hei eal money. O e ime, he police show signs o imp o emen a iden i ying
ake cash, and he o ge s imp o e a aking i . In he end, he coun e ei e s a e com-
pelled o make ideal copies o eal money. High esolu ion and ealis ic mino i y class
images gene a ed using lea ned model dis ibu ion can be used o balance he class dis-
ibu ion and mi iga ing effec o o e fi ing by infla ing he aining da ase size. GANs
sol e he p oblem o gene a ing da a when he e is no enough da a o begin wi h and
hey equi e no human supe ision. GANs can p o ide an efficien way o fill in holes
in he disc e e dis ibu ion o aining da a. In o he wo ds, hey can ans o m he dis-
c e e dis ibu ion o aining da a o con inuous, p o iding an addi ional da a by non-
linea in e pola ion be ween he disc e e poin s. Bowles e al. [68] a gues ha GANs
offe an access o unlock addi ional in o ma ion om a da ase . In ac , Yann LeCun, he
acebook ice p esiden and chie AI scien is , e e ed o GANs as " he mos in e es ing
hing ha has happened o he field o machine lea ning in he las 10yea s".
In his su ey, as opposed o o he ela ed su eys on class imbalance, ha p esen
class imbalance in abula da a, we ocus on wide ange o imbalance in high dimen-
sional image da a by ollowing a sys ema ic app oach wi h a iew o help esea che s
es ablish a de ailed unde s anding o GAN based syn he ic image gene a ion o he
imbalance p oblems in compu e ision asks. Fu he mo e, ou su ey co e s imbal-
ances in a wide ange o compu e ision asks in con as o o he su eys ha a e lim-
i ed o image classifica ion asks.
The key con ibu ions o his su ey a e p esen ed as ollows:
• In his su ey pape , we e iew cu en esea ch wo k on GAN based syn he ic
image gene a ion o he imbalance p oblems in isual ecogni ion asks spanning
om 2014 o 2020. We g oup hese imbalance p oblems in a axonomic ee wi h
h ee main g oups: Classifica ion, Objec de ec ion and Segmen a ion (Fig.2).
Page 9 o 59
Sampa he al. J Big Da a (2021) 8:27
• Also, we p o ide necessa y ma e ial o in o m esea ch communi ies abou he
la es de elopmen and essen ial echnical componen s in he field o GAN based
syn he ic image gene a ion.
• Apa om analyzing diffe en GAN a chi ec u es, ou su ey ocuses hea ily on
eal wo ld applica ions whe e GAN based syn he ic images a e used o alle ia e
imbalances and fills a esea ch gap in he use o syn he ic images o he imbalance
p oblems in isual ecogni ion asks.
The emainde o his pape is o ganized as ollows: “Deep Gene a i e image mod-
els” sec ion gi es eade s necessa y backg ound in o ma ion on gene a i e models.
“Gene a i e ad e sa ial Neu al Ne wo k” sec ion discusses selec ed GAN a ian s
om he a chi ec u e, algo i hm, and aining icks pe spec i e in de ail. In “Taxon-
omy o class imbalance in isual ecogni ion asks” sec ion, we p o ide a b ie expla-
na ion on a ious ypes o imbalances encoun e ed in isual ecogni ion asks and
how he GAN based syn he ic image is used o ebalance, ollowed by GAN a ian s
om he applica ion pe spec i e. “Discussion and Fu u e wo k” sec ion iden ifies and
enume a es ou pe spec i e and possible u u e esea ch di ec ion. Finally, we con-
clude he pape in “Conclusion” sec ion.
Fig. 2 P oposed axonomy o he e iew o imbalanced p oblem in compu e ision asks
Page 16 o 59
Sampa he al. J Big Da a (2021) 8:27
The defini ion o Jensen-Shannon di e gence (
DJ
S
) be ween wo p obabili y dis i-
bu ions p
g
(x
)
and
p
(x
)
is defined as
The e o e, Eq.(10) is equal o
Essen ially, he loss o he gene a o
G
minimizes he Jensen-Shannon di e gence
be ween he gene a ed da a dis ibu ion pg(x
)
and he eal da a dis ibu ion p (x
)
when disc imina o
D
is op imal. Jensen-Shannon di e gence isasmoo h,symme -
ic e siono  heKLdi e gence. Husza [110] belie es ha he main eason behind
he g ea success o GANs is eplacing asymme ic KL di e gence loss unc ion in he
classical app oach o symme ic JS di e gence.
Mean squa ed e o used in la en a iable models such as au oencode , a e ages all
he possible ea u es in an image and gene a e blu y images. In con as , ad e sa ial
loss p ese es he ea u es using disc imina o ne wo ks ha de ec an absence o any
ea u es as an un ealis ic image. An example o his is he s udy ca ied ou by Lo -
e e al. [111], in which models ained using mean squa e loss and ad e sa ial loss
o p edic he nex image ame in a ideo sequence a e compa ed. A model ained
using mean squa e loss gene a es blu y images as shown in Fig.6, whe e ea and eyes
a e no sha ply defined as hey could be. Using an addi ional ad e sa ial loss, ea u es
like he eyes and ea emain p ese ed e y well, because an ea is he ecognizable
pa e n, and he disc imina o ne wo k would no accep any sample ha is missing
an ea .
This sec ion has a emp ed o p o ide eade s a b ie in oduc ion o he cu en
s a e o deep gene a i e image models. A quick summa y o his sec ion is depic ed
below in Fig.7.
Despi e ema kable achie emen s in gene a ing sha p and ealis ic images, GANs
suffe om ce ain d awbacks.
• Non con e gence Bo h gene a o and disc imina o ne wo ks in GANs a e ained
simul aneously using g adien descen in a ze o-sum game. As a esul , imp o e-
(11)
DJS(p ||pg)=
1
2DKL(p || p +pg
2)+
1
2DKL(pg|| p +pg
2
)
(12)
G∗=2DJS(p (x)||pg(x))−2log
2
Fig. 6 An illus a ion o he impo ance o an ad e sa ial loss [111]

Page 17 o 59
Sampa he al. J Big Da a (2021) 8:27
men o he gene a o ne wo k comes a he expense o disc imina o and ice
e sa. Hence he e is no gua an ee o GANs con e gence.
• Mode collapse Gene a o ne wo k achie es a s a e whe e i con inues o gene a e
samples wi h li le a ie y, al hough ained on di e se da ase s. This o m o ail-
u e is e e ed o as mode collapse.
• Vanishing g adien p oblems I he disc imina o is pe ec ly ained ea ly in he
aining p ocess, hen he e would be no g adien s le o ain he gene a o due
o anishing g adien s.
The e o e, many GAN- a ian s ha e been p oposed o o e come hese d awbacks.
These GAN- a ian s can be g ouped in o h ee ca ego ies:
1. A chi ec u e a ian s In e ms o a chi ec u e o gene a o and disc imina o ne -
wo ks, he fi s p oposed GANs use he Mul i- laye pe cep on (MLP). Owing o he
ac ha Con Ne s wo k well wi h high esolu ion image da a aking in o accoun o
he spa ial s uc u e o da a, a Deep Con olu ional GAN (DCGAN) [112] eplaced
he MLP wi h he decon olu ional and con olu ional laye s in gene a o and dis-
c imina o ne wo ks espec i ely.
Cu en S a e o Deep
Gene a e image models
Au o eg essi e models La en Va iable models Ad e sa ial models
* FVBN
* NADE
* MADE
* PixelRNN
* PixelCNN
* GATED PixelCNN
* PixelCNN ++
* PixelSNAIL
P os: Simple and s able
aining.
Cons: Image gene a ion
p ocess is na u ally slow.
* VAE
* β-VAE
* VQ-VAE
* VQ-VAE 2.0
* Condi ional VAE
* Va ia ional lossy
au oencode
* Pixel VAE
* DRAW
P os: Pe om bo h gene a ion
and in e ence wi h la en
a iables.
Cons: 1. La en a iable
models need assump ions on a
p io and pos e io
dis ibu ions.
2. Gene a ed images end o
be blu y.
* Gene a i e Ad e sa ial
Ne wo ks and i s a ian s.
P os: 1. Gene a e he sha pes
image sample.
2. Capable o cap u ing he
high- equency pa s o an
image.
Cons: Gene a i e Ad e sa ial
Ne wo ks a e highly uns able
and di icul o con e ge.
Fig. 7 Compa a i e summa y o Deep gene a i e models discussed in “Deep Gene a i e image models”
sec ion
Page 18 o 59
Sampa he al. J Big Da a (2021) 8:27
Au oencode based GANs such as AAE [113], BiGAN [114], VAE-GAN [115],
DEGAN [116], VEEGAN [117] e c., ha e been p oposed o combine hei cons uc-
ion powe o au oencode s wi h he sampling powe o GANs.
Condi ional based GANs like Condi ional GAN (CGAN) [118], Auxilia y Classifie
GAN (ACGAN) [119], VACGAN [120], in oGAN [121], and SCGAN [122] ocused
on con olling mode o da a being gene a ed by condi ioning model on condi ional
a iable.
2. T aining icks GANs a e difficul o ain. Imp o ed ainings icks such as ea u e
ma ching, miniba ch disc imina ion, his o ical a e aging, one-sided label smoo hing,
and Two Time-Scale Upda e Rule ha e been sugges ed o ensu e ha GANs con-
e ge o achie e Nash equilib ium.
3. Objec i e a ian s In o de o imp o e he s abili y and o e come anishing g adien
p oblems, diffe en objec i e unc ions ha e been explo ed in [123–130].
The ollowing sec ion o his e iew mo es on o desc ibe in g ea e de ail he selec ed
GAN a ian s.
Gene a i e ad e sa ial neu al ne wo ks
A chi ec u e a ian s
The pe o mance and aining s abili y o GANs a e highly influenced by he a chi ec-
u e o he gene a o and he disc imina o ne wo ks. Va ious a chi ec u e a ian s o
GANs ha e been p oposed ha adop se e al echniques o imp o e pe o mance and
s abili y.
i. Condi ional based GAN Va ian s
The s anda d GAN [67] a chi ec u e does no ha e any con ol on he modes o da a
being gene a ed. Van den Oo d e al. [89] a gue ha he class condi ioned image
gene a ion can significan ly enhance he quali y o gene a ed images. Se e al
condi ional based GANs ha e been p oposed ha lea n o sample om a condi-
ional dis ibu ion p(x|
y
) ins ead o ma ginal
p
(x)
.
Condi ional based GANs a i-
an s(Fig.8) can be classified in o wo g oups: 1. Supe ised and 2. Unsupe ised
condi ional GANs.
Supe ised condi ional GANs a ian s equi e a pai o images and co esponding p io
in o ma ion such as class label. The p io in o ma ion could be class labels, ex ual
desc ip ions, o da a om o he modali ies.
cGAN Mi za and Osinde o [118] p oposed condi ional Gene a i e Ad e sa ial Ne -
wo k (cGAN), o ha e a con ol on kind o da a being gene a ed by condi ioning
he model on p io in o ma ion
y
. Bo h disc imina o and gene a o in cGAN a e
condi ioned by eeding
y
as addi ional inpu . Using his p io in o ma ion, cGAN
is guided o gene a e ou pu images wi h desi ed p ope ies du ing he gene a ion
p ocess.
Page 19 o 59
Sampa he al. J Big Da a (2021) 8:27
ACGAN Auxilia y classifie Gene a i e Ad e sa ial Ne wo k (ACGAN) [119] is an
ex ension o he cGAN a chi ec u e. The disc imina o in he ACGAN ecei es
only he image, unlike he cGAN ha ge s bo h he image and he class label as
inpu . I is modified o dis inguish eal and ake da a as well as econs uc class
labels. The e o e, in addi ion o eal ake disc imina ion, he disc imina o also p e-
dic s class label o he image using an auxilia y decode ne wo k.
VACGAN The majo p oblem wi h ACGAN is ha i will affec he aining con e -
gence because o mixing he loss o classifie and disc imina o in o a single loss.
Ve sa ile Auxilia y Gene a i e Ad e sa ial Ne wo k (VACGAN) [120] sepa a es
ou classifie loss by in oducing a classifie ne wo k in pa allel o he disc imina-
o .
No p io in o ma ion is used in unsupe ised condi ional GAN a ian s o con ol on
modes o he image being gene a ed. Ins ead, ea u e in o ma ion such as hai
colo , age, gende e c. is lea ned du ing he aining p ocess. The e o e, hey need
an addi ional algo i hm o decompose he la en space in o disen angled la en ec-
o
c
, which con ains he meaning ea u es, and s anda d inpu noise ec o z. The
con en and ep esen a ion o an image is hen con olled by noise ec o z and
disen angled la en ec o
c
espec i ely.
In o-GAN In o ma ion maximizing Gene a i e Ad e sa ial Ne wo k (In o-GAN) [121]
spli s an inpu la en space in o he s anda d noise ec o
z
and addi ional la en
ec o
c
. The la en ec o c is hen made meaning ul disen angled ep esen a-
ion by maximizing he mu ual in o ma ion be ween la en ec o
c
and gene a ed
images G(z,c
)
using addi ional Q ne wo k.
SC-GAN Simila i y cons ain Gene a i e Ad e sa ial Ne wo k (SC-GAN) [122]
a emp s o lea n disen angled la en ep esen a ion by adding he simila i y con-
s ain be ween la en ec o
c
and gene a ed images G(z,c
)
. In o-GAN uses an
ex a ne wo k o lea n disen angle ep esen a ion, while SC-GAN only adds an
addi ional cons ain o a s anda d GAN. The e o e, SCGAN simplifies he a chi-
ec u e o In o-GAN.
ii. Con olu ional based GAN
DCGAN Deep Con olu ional Gene a i e Ad e sa ial Ne wo k (DCGAN) [112] is he
fi s wo k ha deploys con olu ional and anspose-con olu ional laye s in he dis-
c imina o and gene a o , espec i ely. The salien ea u es o he DCGAN a chi-
ec u e a e enume a ed as ollows:
• Fi s , he gene a o in DCGAN consis s o ac ional con olu ional laye s, ba ch no -
maliza ion laye s and ReLU ac i a ion unc ions.
• Second, he disc imina o is composed o s ided con olu ional laye s, ba ch no -
maliza ion laye s and Leaky ReLU ac i a ion unc ions.
• Thi d, uses Adap i e Momen Es ima ion (ADAM) op imize ins ead o s ochas ic
g adien descen wi h momen um.
iii. Mul iple GANs
Page 20 o 59
Sampa he al. J Big Da a (2021) 8:27
In o de o accomplish mo e han one goal, se e al amewo ks ex end he s anda d
GAN o ei he mul iple disc imina o s, gene a o s, o bo h(Fig.9).
P oGAN In an a emp o syn hesize highe esolu ion images P og essi e G owing o
Gene a i e Ad e sa ial Ne wo k (P oGAN) [131] s acks each laye o he gene a o
and disc imina o in a p og essi e manne as aining p og esses.
LAPGAN Laplacian Gene a i e Ad e sa ial Ne wo k (LAPGAN) [132] is p oposed
o he gene a ion o high quali y images. This a chi ec u e uses a cascade o Con-
Ne s wi hin a Laplacian py amid amewo k. LAPGAN u ilizes se e al Gene a-
o -Disc imina o ne wo ks a mul iple le els o a Laplacian Py amid o an image
de ail enhancemen . Mo i a ed by he success o sequen ial gene a ion, Im e al.
[133] in oduced Gene a i e Recu en Ad e sa ial Ne wo ks (GRAN) based on
ecu en ne wo k ha gene a e high quali y images in a sequen ial p ocess, a he
han in one sho .
D2GAN Dual disc imina o Gene a i e Ad e sa ial Ne wo k (D2GAN) [134] employs
wo disc imina o s and one gene a o o add ess he p oblem o mode collapse.
Unlike GANs, D2GAN o mula es a h ee-playe game ha u ilizes wo disc imi-
na o s o minimize he KL and e e se KL di e gences be ween ue da a and he
gene a ed da a dis ibu ion.
MADGAN Mul i-agen di e se Gene a i e Ad e sa ial Ne wo k (MADGAN) [135]
inco po a es mul iple gene a o s ha disco e di e se modes o he da a while
Fig. 8 A schema ic iew o (a) he anilla GAN and (b– ) a ian s o Condi ional GANs
Page 21 o 59
Sampa he al. J Big Da a (2021) 8:27
main aining high quali y o gene a ed images. To ensu e ha diffe en gene a o s
lea n o gene a e images om diffe en modes o he da a, he objec i e o disc im-
ina o is modified o de ec he gene a o which gene a ed he gi en ake image
along wi h disc imina ing he eal and ake images.
CoGAN Coupled GAN(CoGAN) [136] is used o gene a ing pai o like images in wo
diffe en domains. CoGAN is composed o a se o GANs–GAN1 and GAN2, each
accoun able o syn hesizing images in one domain. I leans a join dis ibu ion
om wo-domain images which a e d awn indi idually om he ma ginal dis ibu-
ions.
CycleGAN and DiscoGAN [137] use wo gene a o s and wo disc imina o s o accom-
plish unpai ed image o image ansla ion asks. CycleGAN [138] adop s he con-
cep o cycle consis ency om machine ansla ion, whe e a sen ence ansla ed
om English o Spanish and ansla e i back om Spanish o English should be
iden ical.
i . Au oencode based GAN Va ian s
The s anda d GANs a chi ec u e is unidi ec ional and can only map om la en space
z o da a space
x
, while au oencode s a e bidi ec ional. The la en space lea ned
by encode s is he dis ibu ion ha con ains comp essed ep esen a ion o he eal
images. Se e al a ian s o GANs ha combine GAN and encode a chi ec u e a e
p oposed o make use o he dis ibu ion lea ned by encode s(Fig.10). A ibu es
edi ing o an image di ec ly on da a space
x
is complex as image dis ibu ions a e
highly s uc u ed and high dimensional. In e pola ion on la en space can acili a e
o ende complica ed adjus men s in he da a space
x
.
DEGAN In s anda d GANs a chi ec u e, he inpu o he gene a o ne wo k is he
noise ec o ha is andomly sampled om a Gaussian dis ibu ion N(0, 1
)
, which may
c ea e a de ia ion om he ue dis ibu ion o eal images. Decode Encode Gene a i e
ad e sa ial Ne wo k (DEGAN) [116] adop decode and encode s uc u e om VAE,
p e ained on he eal images. The p e ained decode and encode s uc u e ans o m
Fig. 9 A schema ic iew o Va ian s o GANs wi h mul iple disc imina o s and gene a o s: a LAPGAN, b
MADGAN and c D2GAN

Page 22 o 59
Sampa he al. J Big Da a (2021) 8:27
andom Gaussian noise o dis ibu ion ha con ains in insic in o ma ion o he images
which is used as inpu o he gene a o ne wo k.
VAEGAN Va ia ional au oencode Gene a i e Ad e sa ial Ne wo k (VAEGAN) [115]
join ly ains VAE and GAN by eplacing he decode o VAE wi h GAN amewo k.
VAEGAN employs ea u e wise ad e sa ial loss o GAN in lieu o elemen wise econ-
s uc ion loss o VAE o imp o e quali y o image gene a ed by VAE. In addi ion o
la en loss and ad e sa ial loss, VAEGAN uses con en loss, also known as pe cep ual
loss, which compa es wo images based on high le el ea u e ep esen a ion om p e-
ained VGG Ne wo k [11].
AAE Unlike VAEGAN ha disc imina es in da a space, ad e sa ial au oencode s
(AAE) [113] imposes a disc imina o on he la en space as lea ning he la en code
dis ibu ion is simple han da a dis ibu ion. The disc imina o ne wo k disc imina es
be ween a sample d awn om la en space and om he dis ibu ion
p
(z) ha we a e
ying o model.
ALI and BiGAN In addi ion o gene a o ne wo k, Ad e sa ially Lea ned In e ence
(ALI) [114] model and Bidi ec ional Gene a i e Ad e sa ial Ne wo k (BiGAN) con ain
an encode componen E ha simul aneously lea n in e se mapping o he inpu da a
x
o he la en code
z
. Unlike o he a ian s o GAN whe e he disc imina o ne wo k
ecei es only eal o a ificially gene a ed images, in he BiGAN and ALI model, he dis-
c imina o ne wo k ecei es bo h image and la en code pai .
VEEGAN [117]: add esses he p oblem o mode collapse h ough addi ion o a econ-
s uc ion ne wo k ha e e ses he ac ion o he gene a o ne wo k. Recons uc ion
ne wo k akes in syn he ic images hen ans o ms hem o noise, while gene a o ne -
wo k akes noise as an inpu and econs uc s hem in o syn he ic image. In addi ion
o ad e sa ial loss, diffe ence be ween he econs uc ed noise and ini ial noise is used
o ain he ne wo k. Bo h gene a o and econs uc ion ne wo ks a e join ly ained,
which encou ages gene a o ne wo k o lea n ue dis ibu ion, hence sol ing he mode
collapse p oblem.
Fig. 10 A schema ic iew o Va ian s o GANs based on Encode and decode a chi ec u e: a AAE, b VAEGAN,
c DEGAN and d BIGAN
Page 23 o 59
Sampa he al. J Big Da a (2021) 8:27
Se e al o he GANs ha e been p oposed o image supe esolu ion. The goal o supe
esolu ion is o upsample low esolu ion images o a high esolu ion one. Ledig e al.
p oposed Supe -Resolu ion GAN (SRGAN) [139] o image supe esolu ion,which
akes poo quali y image as inpu , and gene a es high quali y image wi h 4 × esolu ion.
The gene a o o he SRGAN uses e y deep con olu ional laye s wi h esidual blocks. In
addi ion o an ad e sa ial loss, SRGAN includes a con en loss. The con en loss is com-
pu ed as he euclidean dis ance be ween he ea u e maps o he gene a ed high quali y
image and he g ound u h image, whe e ea u e maps a e ob ained om a p e ained
VGG19 [140] ne wo k. Zhang e  al. [141] combined a sel a en ion mechanism wi h
GANs (SAGAN) o handle long ange dependencies ha make he gene a ed image look
mo e globally cohe en . Image- o-image ansla ion GANs such as Pix2Pix GAN [142],
Pix2pix HD GAN[143], and CycleGAN [137] lea n o map an inpu image om a sou ce
domain o an ou pu image om a a ge domain.A summa y o a chi ec u al a ian s
o GANs a e summa ized in Table1.
Objec i e a ian s
The main objec i e o GAN is o app oxima e he eal da a dis ibu ion. Hence, mini-
mizing dis ance be ween he eal da a dis ibu ion
(p
) and he GAN gene a ed da a
dis ibu ion
(
p
g)
is a i al pa o aining GAN. As s a ed in “Deep Gene a i e image
models” sec ion, s anda d GAN [67] uses Jensen Shannon di e gence o measu e simi-
la i y be ween eal and gene a ed da a dis ibu ions DJS(p ||p
g
) . Howe e , JS di e gence
ails o measu e dis ance be ween wo dis ibu ions wi h negligible o no o e lap. To
imp o e pe o mance and achie e s able aining o GAN, se e al dis ances o di e -
gence measu es ha e been p oposed ins ead o JS di e gence.
WGAN Wasse s ein Gene a i e Ad e sa ial Ne wo k (WGAN) [123] eplaces JSD
om he s anda d GAN wi h he Ea h mo e Dis ance (EMD). EMD also known as
Wasse s ein Dis ance (WD) can be in e p e ed in o mally as minimum amoun o
wo k o mo e ea h (quan i y o mass) om he shape o one dis ibu ion p(x) o ha o
ano he dis ibu ion q(x) so as o ma ch shape o bo h he dis ibu ions. WD is smoo h
and can p o ide meaning ul dis ance measu e be ween dis ibu ions wi h negligible o
no o e lap. WGAN imposes an addi ional Lipchi z cons ain o use WD as he loss in
he disc imina o , whe e i deploys weigh clipping o en o ce weigh s o he disc imina-
o o sa is y Lipchi z cons ain a e each aining ba ch.
WGAN-GP Weigh clipping in he disc imina o o a WGAN g ea ly diminishes i s
capaci y o lea n and o en ails o con e ge. WGAN-GP [124] is an ex ension o WGAN
ha eplaces weigh clipping wi h g adien penal y o en o ce disc imina o o sa is y
Lipchi z cons ain . Fu he mo e, Pe zka e  al. [125] p oposed a new egula iza ion
me hod, also known as WGAN-LP, ha en o ces he Lipschi z cons ain .
LSGAN Leas squa es Gene a i e Ad e sa ial Ne wo k (LSGAN) [126] deploys leas
squa e loss ins ead o he c oss en opy loss in disc imina o o he s anda d GAN o
o e come he p oblem o Vanishing g adien as well as imp o ing quali y o gene a ed
image.
EBGAN Ene gy Based GAN (EBGAN) [127] uses au o-encode a chi ec u e o con-
s uc he disc imina o as an ene gy unc ion ins ead o a classifie . The Ene gy o
EBGAN is he mean squa ed econs uc ion e o o an au oencode , p o iding lowe
Page 24 o 59
Sampa he al. J Big Da a (2021) 8:27
ene gy o he eal images and high ene gy o gene a ed images. EBGAN exhibi s as e
and mo e s able beha io han s anda d GAN du ing aining.
Same as EBGAN, Bounda y Equilib ium GAN (BEGAN) [128], Ma gin adap a ion
GAN [129] and dual agen GAN [130] also deploy an au o-encode a chi ec u e as he
disc imina o . The disc imina o loss o BEGAN uses Wasse s ein dis ance o ma ch he
dis ibu ions o he econs uc ion losses o eal images wi h he gene a ed images.
The e a e also se e al o he objec i e unc ions based on C ame dis ance [144],
Mean/co a iance Minimiza ion [145], Maximum mean disc epancy [146], Chi-squa e
[147] ha e been p oposed o imp o e pe o mance and achie e s able aining o GAN.
Table 1 An o e iew o GANs a ian s discussed in“A chi ec u e a ian s” sec ion
Ca ego ies GAN Type Main A chi ec u al Con ibu ions oGAN
Basic GAN GAN [67] Use Mul ilaye pe cep on in he gene a o and disc imina o
Con olu ional Based GAN DCGAN [112] Employ Con olu ional and anspose-con olu ional laye s in
he disc imina o and gene a o espec i ely
PROGAN [131] P og essi ely g ow laye s o GAN as aining p og esses
Condi ion based GANs cGAN [118] Con ol kind o image being gene a ed using p io in o ma-
ion
ACGAN [119] Add a classifie loss in addi ion o ad e sa ial loss o econ-
s uc class labels
VACGAN [120] Sepa a e ou classifie loss o ACGAN by in oducing sepa a e
classifie ne wo k pa allel o he disc imina o
in oGAN [121] Lea n disen angled la en ep esen a ion by maximizing
mu ual in o ma ion be ween la en ec o and gene a ed
images
SCGAN [122] Lea n disen angled la en ep esen a ion by adding he
simila i y cons ain on he gene a o
La en ep esen a ion based GANs DEGAN [116] U ilize he p e ained decode and encode s uc u e om
VAE o ans o m andom Gaussian noise o dis ibu ion
ha con ains in insic in o ma ion o he eal images
VAEGAN [115] Combine VAE and GAN
AAE [113] Impose disc imina o on he la en space o he au oencode
a chi ec u e
VEEGAN [117] Add econs uc ion ne wo k ha e e se he ac ion o gen-
e a o ne wo k o add ess he p oblem o mode collapse
BiGAN [114] A ach encode componen o lea n in e se mapping o da a
space o la en space
S ack o GANs LAPGAN [132] In oduce Laplacian py amid amewo k o an image de ail
enhancemen
MADGAN [135] Use mul iple gene a o s o disco e di e se modes o he
da a dis ibu ion
D2GAN [134] Employ wo disc imina o s o add ess he p oblem o mode
collapse
CycleGAN [137] Use wo gene a o s and wo disc imina o s o accomplish
unpai ed image o image ansla ion ask
CoGAN [136] Use wo GANs o lea n a join dis ibu ion om wo-domain
images
O he a ian s SAGAN [141] Inco po a e sel -a en ion mechanism o model long ange
dependencies
GRAN [133] Recu en gene a i e model ained using ad e sa ial p ocess
SRGAN [139] Use e y deep con olu ional laye s wi h esidual blocks o
image supe esolu ion
Page 25 o 59
Sampa he al. J Big Da a (2021) 8:27
T aining icks
While esea ch on a ious GANs a chi ec u es and objec i e unc ions con inue o
imp o e he s abili y o aining, he e a e se e al aining icks p oposed in he li e -
a u e in ended o achie e excellen aining pe o mance. Rad o d e al. [112] showed
using leaky ec ified ac i a ion unc ions in bo h gene a o and disc imina o laye s ga e
highe pe o mance o e using o he ac i a ion unc ions. Salimans e al. [148] p oposed
se e al heu is ic app oaches which can imp o e he pe o mance, and aining s abili y
o GANs. Fi s , ea u e ma ching, changes he objec i e o he gene a o o minimize he
s a is ical diffe ence be ween ea u es o he gene a ed and eal images. In his way, he
disc imina o is ained o lea n impo an ea u es o he eal da a. Second, miniba ch
disc imina ion, whe e he disc imina o p ocess ba ch o samples, a he han in isola-
ion ha helps p e en mode collapse, as he disc imina o can iden i y i he gene a o
con inues o gene a e sample wi h li le a ie y. Thi d, his o ical a e aging, ha akes
he unning a e age o pa ame e s in he pas and penalizes i he e is a la ge diffe ence
be ween pa ame e s, which can help he model o con e ge o an equilib ium. Finally,
one-sided label smoo hing p o ides smoo hed labels o he disc imina o ins ead o 0 o
1, which can smoo h he classifica ion bounda y o he disc imina o .
Sønde by e al. [149] p oposed he idea o c ippling he disc imina o by in oducing
noise o he samples a he han labels, which p e en s he disc imina o om o e fi -
ing. Heusel e al. [150] used a sepa a e lea ning a e o gene a o and disc imina o ,
and ained GANs by aTwo Time-ScaleUpda e Rule (TTUR) o ensu e ha model con-
e ge o a s a iona y local Nash equilib ium. To s abilize he aining o he disc imina-
o , Miya o e al. [151] p oposed no maliza ion echnique called spec al no maliza ion.
Taxonomy o class imbalance in isual ecogni ion asks
This sec ion desc ibes diffe en GANs applied o imbalance p oblems in a ious isual
ecogni ion asks. We g oup he imbalance p oblems in a axonomy wi h h ee main
ypes: 1. Image le el imbalances in classifica ion 2. objec le el imbalances in objec
de ec ion and 3. pixel le el imbalances in segmen a ion asks. Unde s anding his axon-
omy o imbalances will p o ide a aluable amewo k o u he esea ch in o syn he ic
image gene a ion using GAN.
Class imbalances inclassi ica ion
Image classifica ion is he ask o classi ying an inpu image acco ding o a se o pos-
sible classes. Classifica ion can be b oken down in o wo sepa a e p oblems: bina y clas-
sifica ion and mul i-class classifica ion. Bina y classifica ion in ol es assigning an inpu
image in o one o wo classes, whe eas in mul i-class classifica ion wo o se e al classes
a e in ol ed. A classic example o a bina y image classifica ion p oblem is he iden ifica-
ion o ca s o dogs in each inpu image. Image da ase wi h high imbalance [152], which
includes in e -class imbalance and in a-classes imbalance, esul s in poo classifica ion
pe o mance.
Page 32 o 59
Sampa he al. J Big Da a (2021) 8:27
and a spa ial a en ion ne wo k (SAN). Gi en a ace image, SAN lea ns o localize he
a ibu e-specific egion and hen AMN edi he ace image wi h he desi ed a ibu es
in he specific egion loca ed by SAN.
The majo downside wi h he cu en app oaches is ha he inpu o GAN should
be on al ace images. I will be in e es ing o explo e a new a chi ec u e ha can be
ained o modi y he a ibu es o side- iew o any a bi a y iews.
Pe son e-iden ifica ion Pe son e-iden ifica ion [172] is ano he challenging ask wo h
men ioning, which a e ad e sely affec ed due o significan in a class imbalance. In a
class a ia ions caused by o a ion ( a ying poses) a e o en la ge han he in e -pe son
dissimila i ies used o diffe en ia e he ace images [173]. Recen ace- ecogni ion su eys
[174, 175] iden ified pose a ia ion as one o he p ominen un esol ed issues in ace- ec-
ogni ion ask. Fo ins ance, in o de o main ain he highes s anda d o secu i y, a sma
ideo sys em needs o be able o de ec a pe son in a ian o pose (Fig.16).
Qian e al. [176] in oduced a pose-no malized GAN model (PN-GAN) o alle ia ing
he effec s o pose a ia ion. Gi en any pedes ian image and a desi able pose as inpu ,
he model u ilized a desi able pose o p oduce a syn he ic image o he same iden i y
wi h he o iginal pose eplaced wi h he desi able pose (Fig.17). A e his, he au ho s
ained he e-iden ifica ion model wi h he o iginal images and gene a ed pose-no -
malized images o ex ac wo se s o ea u es. Finally, hey used he wo ypes o ea-
u es as he final ea u e. As a esul , he ea u es ex ac ed om he syn hesized images
imp o ed he gene aliza ion abili y o he e-iden ifica ion model.
To add ess pe son e-iden ifica ion challenges in complex scena ios, Wei e al. [177]
p oposed a model called Pe son T ans e Gene a i e Ad e sa ial Ne wo k (PTGAN) o
implausible pe son image s yle ans e om sou ce domain o a ge domain, ac oss
da ase s wi h diffe en s yles, such as backg ounds, poses, seasons, ligh ings, e c. The
domain ans e p ocedu e in PTGAN is inspi ed by CycleGAN [138]. Diffe en om
Fig. 16 Example o Pe son eiden ifica ion ask. Pe son eiden ifica ion is a key elemen in ideo su eillance
ha deals wi h ma ching images o same pe son o e many non-o e lapping came a iews

Page 33 o 59
Sampa he al. J Big Da a (2021) 8:27
Cycle-GAN [138], PTGAN inco po a es addi ional cons ain s on he pe son o e-
g ounds o make su e he s abili y o hei iden i ies du ing ans e . Compa ed wi h
Cycle-GAN, PTGAN gene a es high esolu ion pe son images, whe e pe son iden i ies
a e unchanged, and he s yles a e ans o med.
Being a c oss-came a acking and human e ie al ask, pe son e-iden ifica ion
o en suffe s om image s yle a ia ions esul ing om diffe en came as. The e-
o e, Zhong e al. [178] designed a came a s yle adap ion model o adjus ing Con Ne
aining. They ha e used CycleGAN [138] o ans e ing images omone came a o
hes yleo ano he came a. Gi en ha bo h o iginal and s yle ans e ed images, iden-
ifica ion disc imina i e embedding (IDE) is used o ain he Con Ne model. Pa icu-
la ly, au ho s ha e used ResNe -50 p e- ained on ImageNe da ase as backbone and
ollow he fine- uning s a egy.
Pedes ian images suffe om in o ma ion loss when ans e ing om one came a o
hes yleo ano he came a. Deng e al. [179] p esen ed a model, named simila i y p e-
se ing cycle consis en gene a i e ad e sa ial ne wo k (SPGAN), which is composed
o a CycleGAN and a Siamese ne wo k (SiaNe ). CycleGAN lea ns o ansla e pedes-
ian images om one domain o ano he domain, and he con as i e loss induced by
he SiaNe pulls close a ansla ed image and i s coun e pa in he sou ce domain, and
mo es away he ansla ed image and any image in he a ge domain.
Ge e  al. [180] p esen ed a Fea u e Dis illing Gene a i e Ad e sa ial Ne wo k (FD-
GAN) ha aims a lea ning iden i y ela ed and pose-un ela ed pe son ep esen a ions.
The p oposed model adop s a Siamese s uc u e wi h mul iple no el disc imina o s
on human poses (pose disc imina o ) and iden i ies (iden i y disc imina o ). The idea
behind FD-GAN is o lea n pose-un ela ed and iden i y- ela ed ea u es o pedes ian
image, hen i can be used o gene a e he same pedes ian image bu wi h diffe en a -
ge poses.
Al hough GAN-based me hods desc ibed abo e ha e achie ed excellen pe o mance
in image-based pe son e-iden ifica ion, i s ill needs conside able effo o ackle he
ideo-based iden ifica ion da ase s. Fu u e wo k seeks o expand o use GAN o gene -
a ing a sequence o images o he ideo-based iden ifica ion da ase s.
Fig. 17 A chi ec u e diag am o pose-no malized GAN p esen ed by Qian e al. [176]
Page 34 o 59
Sampa he al. J Big Da a (2021) 8:27
Vehicle e-iden ifica ion Vehicle Re-iden ifica ion ask is e en mo e challenging as i su -
e s om la ge in a-class diffe ences caused by iewpoin and illumina ions a ia ions,
and in e -class simila i y p ima ily o diffe en iden i ies wi h he simila look (Fig.18).
Zhou e al. [182] p oposed a model called C oss iew GAN o gene a e images in
diffe en iewpoin s o he same ehicle. C oss iew GAN composed o classifica ion,
gene a o , and disc imina o ne wo k. Fi s , classifica ion ne wo k is ained o lea n
ehicle in insic ea u es such as model, colo , and ype in o ma ion. In addi ion o
in insic ea u es, i also lea ns iewpoin ea u es. Then he gene a i e ne wo k is
condi ioned on he a e age ea u e o he expec ed iewpoin and ehicle’s in insic
ea u es o in e images o he same ehicle in o he iewpoin s. The disc imina o
ne wo k lea ns o dis inguish eal images om he gene a ed images, while ensu ing
images a e gene a ed wi h co ec a ibu es.
Wu e al. [183] imp o ed he disc imina i e powe o he ResNe -50 model o he
Vehicle e-ID ask by simul aneously aining wi h ini ial labeled images and DCGAN
gene a ed unlabeled images. They u he explo e he effec i eness o using DCGAN
gene a ed images on a wide ange o ehicle e-ID da ase s and show imp o ed pe -
o mance o ehicle e-iden ifica ion.
Fine-g ained image classifica ion The fine-g ained image classifica ion is also a ib-
u ed o majo a ia ions in he in a-class and mino in e class a ia ions [184]. I is a
difficul ask o wo easons. Fi s , he aining samples o each class a e inadequa e.
Second, he diffe ences be ween diffe en classes o images a e qui e small [185]. As
an example, i is e y difficul o iden i y he images o She land Sheepdog om ha
o Collie dog. Simila ly, he images o Sayo nis and G ay Kingbi d a e qui e difficul o
dis inguish (Fig.19).
Fu e al. [184] de eloped a model called Fine g ained condi ional GAN (F-CGAN)
o sol e fine g ained class dependen image syn hesis p oblems. F-CGAN consis s o
Fig. 18 illus a ion o challenges in ehicle Re-iden ifica ion p o ided by Zheng e al. [181]
Page 35 o 59
Sampa he al. J Big Da a (2021) 8:27
h ee main componen s: 1. a 2-s age GAN, 2. a fine-g ained ea u e p ese e and 3.
a mul i- ask classifica ion model. The 2-s age GAN gene a es high esolu ion images,
he fine-g ained ea u e p ese e a ge s o cap u e fine g ained de ails and he
mul i- ask classifica ion model u ilizes gene a ed image da a o imp o e fine g ained
classifica ion accu acy.
Wang e al. [188] find ha he disc imina o in GANs lea ns a hie a chical iden-
ifica ion ea u es o he fine-g ained classes and disc imina e pa e n o he fine-
g ained aining samples. They use he a chi ec u e pic u ed below o implemen he
fine-g ained Plank on classifica ion ask (Fig. 20). The main idea is o ain a fine-
g ained classifie ha sha es weigh s wi h disc imina o o he DCGAN, which o ces
disc imina o o concen a e on ea u es o small classes. On WHOI-Plank on da ase
[189], F1 sco e o he classifie imp o ed by o e 7%.
Typically, medical image da ase s con ain bo h gene al labels, e.g., “male”, “ emale” and
disease specific de ailed labels [190]. I is men ioned ha he complexi y and na u e o
da a is ha d o lea n by using a single GAN. Hence, T. Koga e al. [190] connec ed wo
GANs in se ies, one o lea ning gene al ea u es and o he o de ailed ea u es. The fi s
GAN gene a es di e se images, which akes a noise ec o and gene al labels as inpu s.
Collie She land sheepdog Sayo nis G ay Kingbi d
Fig. 19 Sample images om he S an o d Dogs da ase [186] and he Cal ech-UCSD Bi ds da ase [187],
which exhibi s mino in e -class a ia ions and majo in a-class a ia ions
Fig. 20 Comple e fine-g ained Plank on classifie a chi ec u e used by Wang e al. [188]
Page 36 o 59
Sampa he al. J Big Da a (2021) 8:27
The second GAN ecei es syn he ic images gene a ed by he fi s GAN, and disease spe-
cific de ailed labels as inpu s, and gene a es he final fine-g ained medical images.
Mul iclass imbalance
In many eal wo ld p oblems such as emo ion classifica ion [191], plan disease classifi-
ca ion [192], medical image classifica ion [193], indus ial de ec classifica ion [194] e c.,
i is mo e likely ha mo e han one class exis s and needs o be ecognized. Mul iclass
classifica ion has been shown o suffe mo e lea ning difficul ies han bina y class clas-
sifica ion, because mul iclass classifica ion inc eases he da a complexi y and in ensifies
he imbalanced dis ibu ion [195]. Th ee ypes o imbalance could occu o he mul i-
class da ase s: ew mino i y-many majo i y classes, many mino i y- ew majo i y classes,
and many mino i y-many majo i y classes. Shuo Wang e al. [196] s udied he impac o
all diffe en ypes o mul iclass imbalances and showed ha hey nega i ely affec mino -
i y class and o e all pe o mance.
An example o ew mino i y-many majo i y class imbalance is an emo ion classifica-
ion, as some classes o emo ions like disgus a e ela i ely uncommon compa ed o
common emo ions like happy o sad. Zhu e al. [197] employed cycle-GAN which can
syn hesize uncommon emo ion classes like disgus ed om he equen classes (Fig.21).
In addi ion o ad e sa ial and cycle consis ency loss, hey use leas squa e loss om
LSGAN o a oid anishing g adien p oblems. Employing cycle-GAN based da a mino -
i y class da a augmen a ion achie ed 5–10% inc ease in he o e all accu acy. They also
ound ha enla ging mino i y classes also inc eases accu acy o o he majo i y classes.
Wea he Image classifica ion is ano he example o ew mino i y-many majo i y class
imbalance, because some ypes o wea he , like snow, is ela i ely a e compa ed o
sunny, hazy and ainy days. Li e al. [198] used DCGAN o gene a e images o mino i y
classes in aining. They ound ha he GAN-based da a augmen a ion echnique led
o ma gin cla i y be ween classes and hence imp o emen in classifica ion pe o mance.
Fig. 21 On emo ion classifica ion ask [197], he images on he le a e o iginal da a and he es a e images
gene a ed by cycle-GAN
Page 37 o 59
Sampa he al. J Big Da a (2021) 8:27
Huang e al. [199] p esen ed an in e es ing idea o combine ensemble lea ning wi h
GANs designed o add ess he class imbalance p oblem in wea he classifica ion. The
p oposed me hod comp ised o h ee ing edien s as depic ed in (Fig.22): 1. DCGAN o
gene a e syn he ic images and balance he aining da ase 2. Nea es neighbo me hod
o emo e any possible ou lie images gene a ed by DCGAN 3. An ensemble lea ning
me hod o combine he classifica ion esul s o he mul iple classifie s so as o achie e
be e esul s.
The use o DCGAN was es ed by Salehinejad e al. [193] in he ask o ches pa hol-
ogy classifica ion. Using ches X- ay images, hey build a deep Con Ne classifie o
classi y 5 diffe en anemic classes. Thei da ase is highly imbalanced, con ains h ee
majo i y and wo mino i y classes (Fig.23a). The syn he ic images gene a ed using
DCGAN we e used o balance and augmen he o iginal imbalanced da ase . They
demons a ed ha a combina ion o he o iginal imbalanced da ase and gene a ed
images imp o es he accu acy o deep Con Ne classifie in compa ison o he same
classifie ained wi h o iginal imbalanced da ase alone. On ches X- ay da ase
[193], a mean classifica ion accu acy imp o ed om 70.87 o 92.10%.
F id-Ada e  al. [200] also showed ha gene a ing syn he ic li e lesion images
using DCGAN can imp o e classifica ion esul s. They combined s anda d augmen-
a ion echniques and DCGAN gene a ed syn he ic images o ain a classifie . Thei
li e lesion da ase con ains 182 compu ed omog aphy images (65 hemangiomas, 64
me as ases and 53 cys s). By adding he syn he ic images o s anda d da a augmen-
a ion, hei classifica ion pe o mance inc eased om 78.6% sensi i i y and 88.4%
specifici y using s anda d augmen a ions o 85.7% sensi i i y and 92.4% specifici y
using DCGAN-based syn he ic images.
Fig. 22 Illus a ion om Huang e al. [199] showing how he Ensemble lea ning is in eg a ed wi h GAN
F amewo k

Page 38 o 59
Sampa he al. J Big Da a (2021) 8:27
Rashid e al. [201] es ed he effec i eness o using GANs o gene a e skin lesion
images. Using ISIC 2018 da ase [202], hey buil a CNN classifie o classi y 7 diffe -
en skin lesions as depic ed in Fig.24. These classes a e highly imbalanced, and he
GAN is used as a me hod o in elligen o e sampling.
Nazki e  al. [192] used Cycle-GAN o alle ia e mul iclass imbalance p oblem in
oma o plan disease classifica ion. Thei oma o plan disease da ase con ains 2789
images, highly suffe ed om class imbalance in 9 disease ca ego ies (Fig.23b). Using
Cycle-GAN, hey ansla ed images om he heal hy oma o lea es o unde ep-
esen ed diseased oma o lea es. This s udy demons a ed ha he syn he ic image
gene a ed by Cycle-GAN can be used as an augmen ed aining se o imp o e he
pe o mance o classifie .
Bha ia e al. [203] sough ou o compa e syn he ic images gene a ed using WGAN-
GP agains he s anda d da a augmen a ion in he con ex o mul iclass image clas-
sifica ion. They a ificially in oduced class imbalance in wo balanced da ase s o
CIFAR-10 [87] and FMNIST [204], and s udied he effec s o mul iclass imbalance on
classifica ion pe o mance. On he CIFAR-10 [87] da ase , classifica ion pe o mance
imp o ed om 80.84% accu acy and 0.806 F1-sco e using s anda d da a augmen a-
ion o 81.89% accu acy and 0.812 F1-sco e using WGAN-GP. On FMNIST [204]
da ase , pe o mance imp o ed om 91.9% accu acy and 0.921 F1-sco e using aug-
men a ion o 92.8% accu acy and 0.923 F1-sco e using WGAN-GP.
0
1000
2000
3000
4000
5000
6000
AKIEC BCC BKL DF MEL NV VASC
Numbe o images
ab
Fig. 24 a Dis ibu ion o he se en skin lesion class labels o he ISIC 2018 da ase [202]. b Sample images
om each class
0
5000
10000
15000
20000
25000
30000
35000
Imbalanced Balanced
Numbe o images
Ca diomegaly No mal Effusion Edema Pneumo ho ax
DC-GAN
0
100
200
300
400
500
600
700
800
900
Imbalanced Balanced
Numbe o images
Canke G ay Mold Lea Mold
Low empe a u e Mine Nu ional excess
Plague Powde y mildew Whi efly
Cycle-GAN
ab
Fig. 23 The dis ibu ions o (a) Ches X- ay image da ase [193] and (b) oma o plan disease da ase [192],
be o e (le ) and a e class balancing using GANs ( igh )
Page 39 o 59
Sampa he al. J Big Da a (2021) 8:27
An idea o GANs based ans e lea ning echnique o mul iclass imbalance p ob-
lem is p oposed by Fanny e al. [205]. Thei a chi ec u e named class expe gene a-
i e ad e sa ial ne wo k (CE-GAN) makes use o mul iple GANs models, a sepa a e
GANs o each class. Fea u e maps in he main classifie a e a anged in pa allel, wi h
each ea u e maps p e- ained o iden i y he cha ac e is ics o a single class in he
aining da a (Fig. 25). The weigh s o he p e ained ea u e maps a e ans e ed
om disc imina o s o he GANs o main classifie model o u he aining in a
supe ised mode.
The GAN-based syn he ic images se ed as an in elligen o e sampling echnique
and can add ess he p oblem o mul i-class imbalance o a g ea e ex en . Howe e ,
syn he ic images mus be used wi h cau ion because i he quali y o he syn hesized
images is no high, his would lead o addi ional noise o he o iginal da ase s.
Objec le el imbalances inobjec de ec ion
Objec -scale imbalance
One pe asi e challenge in he scale in a ian objec de ec ion is la ge scale a iance
ac oss objec ins ances, and pa icula ly, de ec ing small objec s a e mo e challeng-
ing han medium and la ge-scale objec s. As pe MS COCO defini ion [206], Objec s
Fig. 25 Illus a ion o he class expe gene a i e ad e sa ial ne wo k a chi ec u e [205]
Page 40 o 59
Sampa he al. J Big Da a (2021) 8:27
wi h size less han 32 × 32 pixels a e small, size be ween 32 × 32 o 96 × 96 pixels
a e conside ed as medium and objec s wi h size g ea e han 96 × 96 pixels a e la ge
objec s (Table2). On he one hand, small objec s in MS COCO da ase accoun s o
only 1.23% o o al objec a ea, on he o he hand, medium and la ge-scale objec s a e
o e 98% o objec a ea. Objec de ec ion algo i hms should be able o de ec bo h
small objec s as well as medium and la ge objec s. De ec ing small objec s a e essen-
ial in many eal-wo ld applica ions. Fo ins ance, de ec ing dis an o small objec s in
he high- esolu ion d i ing scene images cap u ed om ca s is essen ial o achie ing
au onomous d i ing. Many dis an objec s, such as affic ligh s o ca s, a e impe -
cep ible as shown in Fig.26. Haoyue e al. [207] measu e he ex en o scale a ia ion
using he coefficien o a ia ion (CV), de e mined as he a io o he s anda d de ia-
ion o he mean o he objec scale. The bigge he CV, he mo e complica ed he
p oblem o scale a ia ion.
The e can be h ee easons why de ec ing small objec s a e mo e complica ed han
la ge one: 1. Small objec s occupy a much smalle a ea, and consequen ly he e exis s
lack o di e si y whe e small objec s a e loca ed in he image, 2. The e a e compa a i ely
less images in he da ase con aining small objec s which may bias any objec de ec ion
algo i hm o concen a e mo e on medium and la ge-scale objec s, and 3. The ac i a-
ions o small objec s become smalle and smalle wi h each pooling laye in a s anda d
con Ne a chi ec u e as i p og essi ely educes he spa ial size o an image.
To o e come he p oblem o scale imbalance, wo diffe en s a egies based on GAN
ha e been p oposed in he li e a u e. Commonly adop ed s a egy is o con e low
esolu ion small objec ea u es in o high esolu ion ea u es [208] using GAN. Di e -
si y o he small objec loca ions in he images a e enhanced by copy-pas ing small
objec ins ances se e al imes in each image h ough ad e sa ial p ocesses [209].
Table 2 The de ini ions ands a is ics o  hesmall, medium, andla ge objec s asMS COCO
[206]
Objec ca ego y Spa ial dimension Objec coun % To al
objec
a ea %
Minimum Maximum
Small 0 × 032 × 32 41.43 1.23
Medium 32 × 32 96 × 96 34.32 10.18
La ge 96 × 96 ∞ × ∞24.24 88.59
Fig. 26 Example o scale a ia ion and he scale (objec size) dis ibu ion o he VisD one2019 da ase
objec s in pixels [207]
Page 41 o 59
Sampa he al. J Big Da a (2021) 8:27
Li e al. [208] u ilized a GAN amewo k ha ans o ms poo ep esen a ion o small-
scale objec s o supe - esol ed la ge objec s. The gene a o a emp s o gene a e supe
esolu ion ea u es o he small objec s. The disc imina o in his amewo k is decom-
posed in o wo b anches, namely, a pe cep ual b anch and an ad e sa ial b anch. An
ad e sa ial b anch is ained o disc imina e be ween eal la ge-scale objec s and gen-
e a ed supe esolu ion objec s while a pe cep ual b anch helps o make su e ha he
gene a ed supe - esol ed objec is use ul o he de ec ion (Fig.27b). They es ed he
effec i eness o his amewo k on Tsinghua-Tencen 100k da ase [210], PASCALVOC
da ase [211] and Cal ech pedes ian benchma k [212].On he PASCAL VOC 2007
Fig. 27 A chi ec u e diag am o (a) SOD-MTGAN [213] (b) Pe cep ual GAN [208] and (c) De ec o GAN [209]
Fig. 28 Imbalanced dis ibu ion o occluded, pa ially occluded and hea ily occluded objec s in
VisD one-DET2018 da ase [215]
Page 48 o 59
Sampa he al. J Big Da a (2021) 8:27
is no clea how much syn he ic images mus be blended wi h o iginal images o achie e
he maximum pe o mance o he classifie s. Addi ionally, syn he ic images would lead
o addi ional noise o he o iginal aining da ase i he quali y o he syn hesized images
is poo . The e o e, mos o he su eyed me hods in GANs based in elligen o e sam-
pling me hods [197] ocused mainly on balancing dis ibu ion as well as imp o ing qual-
i y o he gene a ed images.
Image- o-image ansla ion [138] me hods used o in e -class imbalance p oblem
canno be ex ended o sol e in a-class imbalance as i is difficul o acqui e image da a-
se s wi h de ailed labels. The in e es ing way o sol e his p oblem is o employ clus-
e ing echniques in he ea u e space o he GANs o di ide he images in o diffe en
g oups o au oma ic pa e n ecogni ion in he da ase . Imp o ing he pe o mance o
he clus e ing echniques ha clea ly find he diffe ence among clus e s, is an a ea o
u u e wo k.
GANs and encode ne wo k hyb id models ha e a good po en ial o add ess in a class
imbalance p oblem in ace ecogni ion and e-iden ifica ion asks. The key idea o hese
models is o wo k on la en code space a he han he pixel space. This is because o
Table 3 (con inued)
Ca ego y Imbalance ype S udy Applica ion
Objec de ec ion Objec Scale imbalance Pe cep ual GAN [208] T affic sign de ec ion
Objec Scale imbalance SOD-MTGAN [213] Small objec de ec ion
sys em
Objec Scale imbalance De ec o GAN [209] Pedes ian and disease
de ec ion
Imbalance due o occlu-
sions and de o ma ions
Ad e sa ial-Fas -RCNN
[216]
Occluded objec de ec ion
Imbalance due o occlu-
sions and de o ma ions
Ad e sa ial Occlusion-
awa e Face De ec o
[217]
Occluded ace de ec ion
Imbalance due o occlu-
sions and de o ma ions
Cu -Pas e GAN [218] Occluded objec de ec ion
Fo eg ound Backg ound
objec class imbalance
Task-awa e syn he ic da a
gene a ion [219]
Objec de ec ion
Fo eg ound Backg ound
objec class imbalance
Gene-GAN [221] Objec de ec ion
Fo eg ound Backg ound
objec class imbalance
PSIS [220] Objec de ec ion
Segmen a ion Pixel wise Imbalance Sensi i i y condi ional GAN
[118]
Shadow de ec ion
Pixel wise Imbalance Pix2pix HD GAN [143] Imbalanced pedes ian
image segmen a ion
Pixel wise Imbalance Voxel GAN [226] B ain umo segmen a ion
Pixel wise Imbalance GAN + ensemble lea ning
[228]
Medical image seman ic
segmen a ion
Pixel wise Imbalance GAN + Weigh ed ca ego i-
cal loss [227]
Hea image segmen a ion
Imbalance due o occlu-
sions
SeGAN[231] In isible pa gene a ion and
Segmen a ion
Imbalance due o occlu-
sions
Occlusion-Awa e GAN
[232]
Occlusion ee image gen-
e a ion

Page 49 o 59
Sampa he al. J Big Da a (2021) 8:27
manipula ing a fine g ained image ca ego y, e.g., hai colo , he la en code ep esen a-
ion will ope a e only on ha single la en code (hai colo ), whe eas he pixel space will
edi e e y single pixel in an image.
The ascina ing app oaches o use GANs o he p oblem o objec le el imbalances in
objec de ec ion asks all in o wo gene al ca ego ies: 1. Gene a ing mo e a e examples
as in elligen o e sampling used o class imbalance. These gene a ed a e examples a e
in oduced in o he aining da ase o add ess imbalance p oblems. 2. Lea n an ad e -
sa y in combina ion wi h o iginal objec de ec ion algo i hms. This ad e sa y modifies
he ea u es o sol e imbalance p oblems ins ead o gene a ing examples in pixel space.
i.e., o gene a e ha d- o-de ec samples by pe o ming ea u e space manipula ions.
The capabili y o supe - esolu ion GANs a e being used o up-sample small blu ed
objec s in o fine-scale ones and o eco e de ailed spa ial in o ma ion o accu a e small
objec de ec ion. This echnique combines supe - esolu ion GANs wi h objec de ec ion
algo i hms o sol e he imbalances due o objec size. The powe o ad e sa ial p ocess is
being used o inc ease he di e si y o he small objec loca ions in he images by copy-
pas ing small objec ins ances se e al imes a diffe en loca ions.
Making he bes use o GANs and combining hem in o U-Ne a chi ec u es is an
in e es ing way o sol e pixel le el imbalances in segmen a ion asks. These a chi ec u es
o en use a weigh ed loss unc ion o mi iga e he pixel le el imbalances. Combina ion
o image in pain ing GANs wi h U-Ne a chi ec u es has he g ea po en ial use in seg-
men ing hidden objec s. This echnique is no only efficien in segmen a ion asks, bu
also o in e he appea ance o he objec s beyond hei isible pa s. O e all, combining
diffe en deep lea ning models wi h ad e sa ial p ocess can p o ide a way o sol e many
o he open p oblems in he compu e ision field.
Fu u e wo k
E en hough GANs can be used as an effec i e way o unlock addi ional in o ma ion
om a da ase , he syn he ic images gene a ed by GANs canno eplace he eal images
comple ely. Howe e , a blend o diffe en p opo ions o eal and GANs gene a ed
images a e ex emely use ul o imp o e he di e si y o he aining samples and inc ease
pe o mance o he classifie s. Ou u u e wo k in ends o s udy he influences o blend-
ing diffe en p oposi ions o GANs gene a ed images and eal images on he classifica-
ion pe o mance. The e a e a e y limi ed numbe o compa a i e s udies ha compa e
effec i eness o using GAN based syn he ic images wi h o he adi ional me hods o
in a-class imbalances. We also in end o conduc he compa a i e s udy in o de o ali-
da e he effec i eness o using syn he ic images o in a class imbalances.
Infla ing he size o he da ase b ings ano he p oblem: One o he mos significan
limi a ions in compu e ision expe imen s is compu a ional esou ces. Sophis ica ed
compu e ision models ained on infla ed da ase can pe o m complex asks, he p ob-
lem howe e is, how do we deploy such massi e a chi ec u e on edge de ices o ins an
usage. Handling his p oblem using knowledge dis illa ion is non- i ial and an ac i e
field o esea ch. Knowledge dis illa ion is model comp ession echnique in which a
smalle ne wo k is ained wi h he help o he sophis ica ed p e ained model o achie e
he simila accu acy. This aining p ocess is o en e e ed o as " eache -s uden ”, whe e
Page 50 o 59
Sampa he al. J Big Da a (2021) 8:27
he sophis ica ed p e ained model is he eache and he smalle ne wo k is he s uden .
Wang e al. [235] combine GANs and knowledge dis illa ion o imp o e he efficiency o
he s uden ne wo k in objec de ec ion. Simila o his wo k, we will a emp o u he
implemen GANs and knowledge dis illa ion combina ions o o he compu e isions
asks.
As esea ch on GANs a e de eloping and ma u ing, assessmen o pe o mance has
become essen ial. E alua ion me ics helps o quan i a i ely measu e how well GANs
models a e pe o ming, also o assess he ela i e pe o mance o GANs. Ve y o en he
pe o mance o GANs is measu ed by he manual inspec ion o he isual fideli y o gen-
e a ed images. Howe e , he manual inspec ion is cumbe some, subjec i e, ime-con-
suming, and some imes misleading. Lack o uni e sal e alua ion me ics can impede he
de elopmen o GANs. In oducing new pe o mance measu es o e alua e bo h di e -
si y and fideli y o gene a ed images is a e y impo an a ea o u u e wo k.
Manually designing GANs a chi ec u e o a gi en ask is ime-consuming and some-
imes has a endency o e o s. This d awback has led esea che s o mo e on o he nex
s age o au oma ing GANs a chi ec u e in he o m o neu al a chi ec u e sea ch (NAS).
Ano he in e es ing a ea o u he esea ch is o use me a-heu is ic sea ch algo i hms
ha assis a chi ec u al sea ch and find op imal GANs a chi ec u e which ou pe o ms
human c ea ed GANs models.
Achie ing equilib ium be ween he gene a o and disc imina o o he GANs can ake
a long ime ela i e o o he deep neu al ne wo ks. Dis ibu ed aining o GAN h ough
pa alleliza ion and clus e compu ing is ano he impo an a ea o u u e wo k o cu
down he aining ime.
Mos o he applica ions o he GANs so a ha e been o c ea ing syn he ic images.
GANs a e no limi ed o he isual domain and can be also applied o non- isual appli-
ca ions. Fo example, Paganini e al. [236] used GANs o p edic he ou come o high
ene gy pa icle physics expe imen s. Ins ead o using explici Mon e Ca lo simula ion
o he eal physics o e e y s ep, he GANs lea n by example wha ou come is likely o
occu in each si ua ion. The GANs educe he compu a ional cos o high ene gy pa icle
simula ion, enough o sa e millions o dolla s’ wo h o supe compu e ime. We belie e
ha he in en ion o new applica ions using his powe ul ool will be con inued in he
u u e.
Conclusion
This pape su eys a ious GANs a chi ec u es ha ha e been used o add essing he
diffe en imbalance p oblems in compu e ision asks. In his su ey, we fi s p o ided
de ailed backg ound in o ma ion on deep gene a i e models and GAN a ian s om he
a chi ec u e, algo i hm, and aining icks pe spec i e. In o de o p esen a clea oad-
map o a ious imbalance p oblems in compu e ision asks, we in oduced axonomy
o he imbalance p oblems. Following he p oposed axonomy, we discussed each ype o
p oblems sepa a ely in de ail and p esen ed he GANs based solu ions wi h impo an
ea u es o each app oach and hei a chi ec u es. We ocused mainly on he eal-wo ld
applica ions whe e GAN based syn he ic images a e used o alle ia e class imbalance. In
addi ion o he ho ough discussion on he imbalance p oblems and hei solu ions, we
add essed many open issues ha a e c ucial o compu e ision applica ions.
Page 51 o 59
Sampa he al. J Big Da a (2021) 8:27
Syn he ic bu ealis ic images gene a ed using he me hods discussed in his su ey
ha e he po en ial o mi iga e he class imbalance p oblem while p ese ing he ex insic
dis ibu ion. Many o he me hods su eyed in his pape ackled he highly complex
imbalances by combining GANs a chi ec u e wi h diffe en o he deep lea ning ame-
wo ks. Specifically, he use o au oencode s wi h GANs has offe ed an effec i e way o
pe o m ea u e space manipula ions ins ead o complex pixel space ope a ions.
Syn he ic images gene a ed by GANs canno be used as he comple e eplacemen o
eal da ase s. Howe e , he blend o eal and GANs gene a ed images ha e eno mous
po en ial o inc ease he pe o mance o he deep lea ning model. Looking in o he
u u e, GAN- ela ed esea ch in image as well as non-image da a domains o add ess he
p oblem o imbalances and limi ed aining da ase would con inue o expand. We con-
clude ha he u u e o GANs is p omising and he e a e clea ly a lo o oppo uni ies o
u he esea ch and applica ions in many fields.
Abb e ia ions
Con Ne s: Con olu ional neu al ne wo ks; SMOTE: Syn he ic mino i y o e sampling echnique; ADASYN: Adap i e
syn he ic sampling; IHM: Ins ance ha dness measu e; SSL: Semi-supe ised lea ning; R-CNN: Region-based con olu-
ional neu al ne wo ks; RPN: Region p oposal ne wo k; YOLO: You only look once; SSD: Singe sho de ec ion; SNIP: Scale
no maliza ion o image py amids; FPN: Fea u e py amid ne wo ks; RNN: Recu en neu al ne wo ks; LSTM: Long sho -
e m memo y; PCA: P inciple componen analysis; MADE: Masked au oencode densi y es ima o ; ARs: Au o eg essi e
models; FVBNs: Fully isible belie ne wo ks; RGB: Red G een blue; NADE: Neu al au o eg essi e densi y es ima o ; MADE:
Masked au oencode densi y es ima o ; VAEs: Va ia ional au o encode s; CVAE: Condi ional a ia ional au o encode s;
DC-IGN: Deep con olu ional in e se g aphics ne wo k; IWVAE: Impo ance weigh ed Va ia ional Au o Encode s; VQ-VAEs:
Vec o quan ized a ia ional au o encode s; DRAW : Deep ecu en a en i e w i e ; EMD: Ea h mo e Dis ance; TTUR
: Two ime-scale upda e ule; DDSM: Digi al da abase o sc eening mammog aphy; ARU-ne : Ad e sa ially egula ized
U-ne ; AMN: A ibu e manipula ion ne wo k; SiaNe : Siamese ne wo k; CV: Coefficien o a ia ion; AP: A e age p eci-
sion; ASTN: Ad e sa ial spa ial ans o me ne wo k; ASDN: Ad e sa ial spa ial d opou ne wo k; mAP: Mean a e age
p ecision; AOFD: Ad e sa ial occlusion awa e ace de ec ion; PSIS: P og essi e and selec i e ins ance-swi ching; ADAM:
Adap i e momen es ima ion op imize ; ReLU: Rec ified linea uni ; GANs: Gene a i e ad e sa ial neu al ne wo ks;
cGAN: Condi ional gene a i e ad e sa ial ne wo k; ACGAN: Auxilia y classifie gene a i e ad e sa ial ne wo k; VACGAN:
Ve sa ile Auxilia y classifie gene a i e ad e sa ial ne wo k; In oGAN: In o ma ion maximizing gene a i e ad e sa ial
ne wo k; SCGAN: Simila i y cons ain gene a i e ad e sa ial ne wo k; DCGAN: Deep con olu ional gene a i e ad e sa ial
ne wo k; P oGAN: P og essi e g owing o gene a i e ad e sa ial ne wo k; LAPGAN: Laplacian gene a i e ad e sa ial
ne wo k; GRAN: Gene a i e ecu en ad e sa ial ne wo ks; D2GAN: Dual disc imina o gene a i e ad e sa ial ne wo k;
MADGAN: Mul i-agen di e se gene a i e ad e sa ial ne wo k; CoGAN: Coupled gene a i e ad e sa ial ne wo k; DEGAN:
Decode encode gene a i e ad e sa ial ne wo k; VAEGAN: Va ia ional au oencode gene a i e ad e sa ial ne wo k;
AAE: Ad e sa ial au oencode s; ALI: Ad e sa ially lea ned in e ence; BiGAN: Bidi ec ional gene a i e ad e sa ial ne wo k;
SRGAN: Supe - esolu ion gene a i e ad e sa ial ne wo k; SAGAN: Sel -a en ion gene a i e ad e sa ial ne wo k; WGAN:
Wasse s ein gene a i e ad e sa ial ne wo k; WGAN-GP: Wasse s ein gene a i e ad e sa ial ne wo k wi h g adien pen-
al y; LSGAN: Leas squa es gene a i e ad e sa ial ne wo k; EBGAN: Ene gy based gene a i e ad e sa ial ne wo k; BEGAN:
Bounda y equilib ium gene a i e ad e sa ial ne wo k; SD-GAN: Su ace de ec -gene a i e ad e sa ial ne wo k; BAGAN:
Balancing gene a i e ad e sa ial ne wo k; ciGAN: Condi ional infilling gene a i e ad e sa ial ne wo k; IcGAN: In e ible
condi ional gene a i e ad e sa ial ne wo k; PNGAN: Pose-no malized gene a i e ad e sa ial ne wo k; PTGAN: Pe son
ans e gene a i e ad e sa ial ne wo k; SPGAN: Simila i y p ese ing cycle consis en gene a i e ad e sa ial ne wo k;
FD-GAN: Fea u e dis illing gene a i e ad e sa ial ne wo k; F-CGAN: Fine g ained condi ional GAN; CE-GAN: Class expe
gene a i e ad e sa ial ne wo k; ScGAN: Sensi i i y condi ional gene a i e ad e sa ial ne wo k; OAGAN: Occlusion-awa e
gene a i e ad e sa ial ne wo k.
Acknowledgemen s
The au ho s would like o hank he anonymous e iewe s o hei aluable commen s and sugges ions on he pape .
Also, we acknowledge he membe s o he Au onomous and In elligen Sys ems Uni , Teknike , o aluable discussions
and collabo a ions.
Au ho s’ con ibu ions
VS pe o med he p ima y li e a u e e iew and analysis o his su ey, and also d a ed he manusc ip . IM, JJAM and AG
wo ked wi h VS o de elop he a icle’s amewo k and ocus. IM and JJAM double checked he manusc ip and p o ided
se e al ad anced ideas o his manusc ip . All au ho s ead and app o ed he final manusc ip .
Funding
This esea ch wo k was unde aken in he con ex o DIGIMAN4.0 p ojec (“Digi al Manu ac u ing Technologies o Ze o‐
de ec ”, h ps ://www.digim an4-0.mek.d u.dk/). DIGIMAN4.0 is a Eu opean T aining Ne wo k suppo ed by Ho izon 2020,
he EU F amewo k P og amme o Resea ch and Inno a ion (P ojec ID: 814225). This esea ch was also pa ly suppo ed
by he ELKARTEK p ojec KK-2020/00049 3KIA o he Basque Go e nmen .
Page 52 o 59
Sampa he al. J Big Da a (2021) 8:27
A ailabili y o da a and ma e ials
No applicable.
E hics app o al and consen o pa icipa e
No applicable.
Consen o publica ion
No applicable.
Compe ing in e es s
The au ho s decla e ha hey ha e no compe ing in e es s.
Au ho de ails
1 Au onomous and In elligen Sys ems Uni , Teknike , Membe o Basque Resea ch and Technology Alliance, Eiba , Spain.
2 Design and Manu ac u ing Enginee ing Depa men , Uni e sidad de Za agoza, 3 Ma ía de Luna S ee , To es Que edo
Bld, 50018 Za agoza, Spain.
Recei ed: 30 July 2020 Accep ed: 16 Janua y 2021
Re e ences
1. Nug aha BT, Su SF, Fahmizal. Towa ds sel -d i ing ca using con olu ional neu al ne wo k and oad lane de ec o .
P oceedings o he 2nd In e na ional Con e ence on Au oma ion, Cogni i e Science, Op ics, Mic o Elec o-
Mechanical Sys em, and In o ma ion Technology, ICACOMIT 2017. 2017;2018-Janua:65–9.
2. Yada SS, Jadha SM. Deep con olu ional neu al ne wo k based medical image classifica ion o disease diagnosis.
J Big Da a. 2019. h ps ://doi.o g/10.1186/s4053 7-019-0276-2.
3. Gu ie ez A, Ansua egi A, Suspe egi L, Tubío C, Rankić I, Lenža L. A Benchma king o lea ning s a egies o pes
de ec ion and iden ifica ion on oma o plan s o au onomous scou ing obo s using in e nal da abases. J Sen-
so s. 2019. h ps ://doi.o g/10.1155/2019/52194 71.
4. San os L, San os FN, Oli ei a PM, Shinde P. Deep lea ning applica ions in ag icul u e: a sho e iew. Ad ances in
in elligen sys ems and compu ing. Fou h Ibe. 2020. h ps ://doi.o g/10.1007/978-3-030-35990 -4_12.
5. Wang T, Chen Y, Qiao M, Snoussi H. A as and obus con olu ional neu al ne wo k-based de ec de ec ion model
in p oduc quali y con ol. In J Ad Manu ac u Technol. 2018;94:3465–71.
6. Hashemi M. Enla ging smalle images be o e inpu ing in o con olu ional neu al ne wo k: ze o-padding s in e -
pola ion. J Big Da a. 2019. h ps ://doi.o g/10.1186/s4053 7-019-0263-7.
7. Lecun Y, Bo ou L, Bengio Y, Haffne P. G adien -based lea ning applied o documen ecogni ion. P oceedings o
he IEEE . 1998;86:2278–324. h p://ieeex plo e .ieee.o g/docum en /72679 1/
8. Gi shick R, Donahue J, Da ell T, Malik J. Rich ea u e hie a chies o accu a e objec de ec ion and seman ic seg-
men a ion. 2014 IEEE Con e ence on Compu e Vision and Pa e n Recogni ion . IEEE; 2014. p. 580–7. h p://ieeex
plo e .ieee.o g/docum en /69094 75/
9. Long J, Shelhame E, Da ell T. Fully con olu ional ne wo ks o seman ic segmen a ion. 2015 IEEE Con e ence on
Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2015. p. 3431–40. h p://a xi .o g/abs/1605.06211
10. K izhe sky A, Su ske e I, Hin on GE. ImageNe classifica ion wi h deep con olu ional neu al ne wo ks. Ad Neu al
In o ma P ocess Sys . 2012;2:1097–105.
11. Simonyan K, Zisse man A. Ve y deep con olu ional ne wo ks o la ge-scale image ecogni ion. 3 d In e na ional
Con e ence on Lea ning Rep esen a ions, ICLR 2015–Con e ence T ack P oceedings. 2015;1–14.
12. Szegedy C, Liu W, Jia Y, Se mane P, Reed S, Anguelo D, e al. Going Deepe wi h Con olu ions. CoRR . 2014;
abs/1409.4. h ps ://a xi .o g/abs/1409.4842
13. He K, Zhang X, Ren S, Sun J. Deep esidual lea ning o image ecogni ion. P oceedings o he IEEE compu e
socie y con e ence on compu e ision and pa e n ecogni ion. 2016. p. 770–8. h p://a xi .o g/abs/1512.03385
14. Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z. Re hinking he incep ion a chi ec u e o compu e ision. 2016
IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2016. p. 2818–26. h p://a xi .o g/
abs/1512.00567
15. Huang G, Liu Z, Van De Maa en L, Weinbe ge KQ. Densely connec ed con olu ional ne wo ks. 2017 IEEE Con e -
ence on Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2017. p. 2261–9. h p://a xi .o g/abs/1608.06993
16. Buda M, Maki A, Mazu owski MA. A sys ema ic s udy o he class imbalance p oblem in con olu ional neu al
ne wo ks. Neu al Ne w. 2018;106:249–59. h ps ://linki nghub .else ie .com/ e i e e/pii/S0893 60801 83021 07
17. Al-S ouhi S, Reddy CK. T ans e lea ning o class imbalance p oblems wi h inadequa e da a. Knowl In o ma Sys .
2016;48:201–28. h ps ://doi.o g/10.1007/s1011 5-015-0870-3
18. Ali A, Shamsuddin SM, Ralescu AL. Classifica ion wi h class imbalance p oblem: a e iew. In J Ad So Compu
Applica . 2015;7:176–204.
19. Zhang J, Xia Y, Wu Q, Xie Y. Classifica ion o medical images and illus a ions in he biomedical li e a u e using
syne gic deep lea ning. 2017. h p://a xi .o g/abs/1706.09092
20. Dong Q, Gong S, Zhu X. Imbalanced deep lea ning by mino i y class inc emen al ec ifica ion. IEEE T ansac ions
on Pa e n Analysis and Machine In elligence . 2019;41:1367–81. h ps ://ieeex plo e .ieee.o g/docum en /83537 18
21. Zhang Y, Li B, Lu H, I ie A, Ruan X. Sample-Specific SVM lea ning o pe son e-iden ifica ion. 2016 IEEE Con e ence
on Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2016. p. 1278–87. h p://ieeex plo e .ieee.o g/docum
en /77805 12/
22. Sawan MM, Bhu chandi KM. Age in a ian ace ecogni ion: a su ey on acial aging da abases, echniques and
effec o aging. A ific In ell Re . 2019;52:981–1008. h ps ://doi.o g/10.1007/s1046 2-018-9661-z.
Page 53 o 59
Sampa he al. J Big Da a (2021) 8:27
23. Mos a a E, Ali A, Alajlan N, Fa ag A. Pose In a ian App oach o Face Recogni ion a Dis ance. Be lin : Sp inge ;
2012. p. 15–28. h ps ://doi.o g/10.1007/978-3-642-33783 -3_2.
24. Japkowicz N, S ephen S. The class imbalance p oblem: a sys ema ic s udy. In ell Da a Analy. 2002;6:429–49. h ps
://doi.o g/10.5555/12939 51.12939 54.
25. Chawla NV. Da a mining o imbalanced da ase s: an o e iew. da a mining and knowledge disco e y handbook.
New Yo k : Sp inge -Ve lag; 2009. p. 853–67. h ps ://doi.o g/10.1007/0-387-25465 -X_40.
26. Chawla NV, Japkowicz N, Ko cz A. Special Issue on Lea ning om Imbalanced Da a Se s. ACM SIGKDD Explo a ions
Newsle e . 2004; 6: 1–6. h ps ://doi.o g/10.1145/10077 30.10077 33
27. Chawla N V., Bowye KW, Hall LO, Kegelmeye WP. SMOTE: Syn he ic mino i y o e -sampling echnique. J A ific
In ell Res. 2011;16:321–57. h ps ://doi.o g/10.1613/jai .953. h ps ://a xi .o g/abs/1106.1813
28. Haibo He, Yang Bai, Ga cia EA, Shu ao Li. ADASYN: Adap i e syn he ic sampling app oach o imbalanced lea ning.
2008 IEEE In e na ional Join Con e ence on Neu al Ne wo ks (IEEE Wo ld Cong ess on Compu a ional In elli-
gence) . IEEE; 2008. p. 1322–8. h p://ieeex plo e .ieee.o g/docum en /46339 69/
29. Pun umapon K, Rak hamamon T, Waiyamai K. Clus e -based mino i y o e -sampling o imbalanced da ase s.
IEICE T ansac ions on In o ma ion and Sys ems . 2016;E99.D:3101–9. h ps ://www.js ag e.js .go.jp/a ic le/ ans in /
E99.D/12/E99.D_2016E DP713 0/_a ic le
30. Sima d PY, S eink aus D, Pla JC. Bes p ac ices o con olu ional neu al ne wo ks applied o isual documen
analysis. Se en h In e na ional Con e ence on Documen Analysis and Recogni ion, 2003 P oceedings . IEEE
Compu . Soc; p. 958–63. h p://ieeex plo e .ieee.o g/docum en /12278 01/
31. Lemley J, Baz a kan S, Co co an P. Deep Lea ning o Consume De ices and Se ices: Pushing he limi s o
machine lea ning, a ificial in elligence, and compu e ision. IEEE Consume Elec onics Magazine . 2017;6:48–56.
h p://ieeex plo e .ieee.o g/docum en /78794 02/
32. Sho en C, Khoshgo aa TM. A su ey on image da a augmen a ion o deep lea ning. J Big Da a. 2019;6:60. h ps
://doi.o g/10.1186/s4053 7-019-0197-0.
33. Wu H, P asad S. Semi-Supe ised Deep Lea ning Using Pseudo Labels o Hype spec al Image Classifica ion. IEEE
T ansac ions on Image P ocessing . 2018;27:1259–70. h p://ieeex plo e .ieee.o g/docum en /81058 56/
34. an Engelen JE, Hoos HH. A su ey on semi-supe ised lea ning. Mach Lea n. 2020;109:373–440. h ps ://doi.
o g/10.1007/s1099 4-019-05855 -6.
35. Thai-Nghe N, Gan ne Z, Schmid -Thieme L. Cos -sensi i e lea ning me hods o imbalanced da a. The 2010
In e na ional Join Con e ence on Neu al Ne wo ks (IJCNN) . IEEE; 2010. p. 1–8. h p://ieeex plo e .ieee.o g/docum
en /55964 86/
36. Gi shick R. Fas R-CNN. 2015 IEEE In e na ional Con e ence on Compu e Vision (ICCV) . IEEE; 2015. p. 1440–8.
h p://ieeex plo e .ieee.o g/docum en /74105 26/
37. Ren S, He K, Gi shick R, Sun J. Fas e R-CNN: Towa ds Real-Time Objec De ec ion wi h Region P oposal Ne wo ks.
IEEE T ansac ions on Pa e n Analysis and Machine In elligence . 2017;39:1137–49. h p://ieeex plo e .ieee.o g/
docum en /74858 69/
38. He K, Gkioxa i G, Dolla P, Gi shick R. Mask R-CNN. IEEE T ansac ions on pa e n analysis and machine in elligence.
2020;42:386–97. h ps ://ieeex plo e .ieee.o g/docum en /83726 16/
39. Liu W, Anguelo D, E han D, Szegedy C, Reed S, Fu C-Y, e al. SSD: Single Sho Mul iBox De ec o . In: Leibe B,
Ma as J, Sebe N, Welling M, edi o s. Cham: Sp inge In e na ional Publishing; 2016. p. 21–37. Doi: h ps ://doi.
o g/10.1007/978-3-319-46448 -0_2
40. Redmon JSDRGAF. (YOLO) You Only Look Once. C p . 2016;
41. Yan X, Gong H, Jiang Y, Xia S-T, Zheng F, You X, e al. Video scene pa sing: an o e iew o deep lea ning me hods
and da ase s. Compu e Vision and Image Unde s anding . 2020;201:103077. h ps ://linki nghub .else ie .com/ e i
e e/pii/S1077 31422 03011 20
42. Hsu Y-W, Wang T-Y, Pe ng J-W. Passenge flow coun ing in buses based on deep lea ning using su eillance ideo.
Op ik . 2020;202:163675. h ps ://linki nghub .else ie .com/ e i e e/pii/S0030 40261 93157 36
43. Singh B, Da is LS. An analysis o scale in a iance in objec de ec ion–SNIP. 2018 IEEE/CVF Con e ence on compu e
ision and pa e n ecogni ion. IEEE; 2018. p. 3578–87. h ps ://ieeex plo e .ieee.o g/docum en /85784 75/
44. Yang F, Choi W, Lin Y. Exploi All he Laye s: Fas and Accu a e CNN objec de ec o wi h scale dependen pooling
and cascaded ejec ion classifie s. 2016 IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR) .
IEEE; 2016. p. 2129–37. h p://ieeex plo e .ieee.o g/docum en /77806 03/
45. Singh B, Najibi M, Da is LS. SNIPER: Efficien Mul i-Scale T aining. 32nd con e ence on neu al in o ma ion p ocess-
ing sys ems. Mon éal; 2018. h p://a xi .o g/abs/1805.09300
46. Lin T-Y, Dolla P, Gi shick R, He K, Ha iha an B, Belongie S. Fea u e Py amid Ne wo ks o Objec De ec ion. 2017 IEEE
con e ence on compu e ision and pa e n ecogni ion (CVPR). IEEE; 2017. p. 936–44. h p://ieeex plo e .ieee.o g/
docum en /80995 89/
47. Lin T-Y, Goyal P, Gi shick R, He K, Dolla P. Focal Loss o Dense Objec De ec ion. IEEE T ansac ions on Pa e n
Analysis and Machine In elligence. 2020;42:318–27. h ps ://ieeex plo e .ieee.o g/docum en /84179 76/
48. Dolla P, Wojek C, Schiele B, Pe ona P. Pedes ian de ec ion: a benchma k. 2009 IEEE Con e ence on Compu e
Vision and Pa e n Recogni ion . IEEE; 2009. p. 304–11. h ps ://ieeex plo e .ieee.o g/docum en /52066 31/
49. Zhong Z, Zheng L, Kang G, Li S, Yang Y. Random E asing Da a Augmen a ion. 2017. h p://a xi .o g/
abs/1708.04896
50. Wang X, Sh i as a a A, Gup a A. A-Fas -RCNN: Ha d posi i e gene a ion ia ad e sa y o objec de ec ion. 2017
IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR). IEEE; 2017. p. 3039–48. h p://a xi .o g/
abs/1704.03414
51. Bad ina ayanan V, Kendall A, Cipolla R. SegNe : A deep con olu ional encode -decode a chi ec u e o image
segmen a ion. IEEE T ansac ions on Pa e n Analysis and Machine In elligence. 2017;39:2481–95. h p://a xi .o g/
abs/1511.00561
52. Ronnebe ge O, Fische P, B ox T. U-Ne : Con olu ional ne wo ks o biomedical image segmen a ion. 2015. p.
234–41. h p://a xi .o g/abs/1505.04597

Page 54 o 59
Sampa he al. J Big Da a (2021) 8:27
53. Diakogiannis FI, Waldne F, Cacce a P, Wu C. ResUNe -a: A deep lea ning amewo k o seman ic segmen a ion
o emo ely sensed da a. ISPRS Jou nal o Pho og amme y and Remo e Sensing . 2020;162:94–114. h ps ://linki
nghub .else ie .com/ e i e e/pii/S0924 27162 03001 49
54. Yu se e E, Lambe J, Ca ballo A, Takeda K. A su ey o au onomous d i ing: common p ac ices and eme ging
echnologies. 2019. h p://a xi .o g/abs/1906.05113
55. Tabe nik D, Šela S, Sk a č J, Skočaj D. Segmen a ion-based deep-lea ning app oach o su ace-de ec de ec ion.
2019. h p://a xi .o g/abs/1903.08536
56. Rizwan I Haque I, Neube J. Deep lea ning app oaches o biomedical image segmen a ion. In o ma ics in Medi-
cine Unlocked. 2020;18:100297. h ps ://linki nghub .else ie .com/ e i e e/pii/S2352 91481 93021 4X
57. Co d s M, Om an M, Ramos S, Reh eld T, Enzweile M, Benenson R, e al. The ci yscapes da ase o seman ic u ban
scene unde s anding. P oceedings o he IEEE Compu e Socie y Con e ence on Compu e Vision and Pa e n
Recogni ion. 2016;2016-Decem:3213–23.
58. Menze BH, Jakab A, Baue S, Kalpa hy-C ame J, Fa ahani K, Ki by J, e al. The mul imodal b ain umo image
segmen a ion benchma k (BRATS). IEEE T ansac Med Imag. 2015;34:1993–2024. h p://ieeex plo e .ieee.o g/docum
en /69752 10/
59. Mu phy KP. Machine lea ning: a p obabilis ic pe spec i e (Adap i e Compu a ion and Machine Lea ning se ies).
Camb idge: The MIT P ess; 2012.
60. Mille a i F, Na ab N, Ahmadi S-A. V-Ne : Fully con olu ional neu al ne wo ks o olume ic medical image seg-
men a ion. 2016 Fou h In e na ional Con e ence on 3D Vision (3DV) . IEEE; 2016. p. 565–71. h p://ieeex plo e .ieee.
o g/docum en /77851 32/
61. C um WR, Cama a O, Hill DLG. Gene alized O e lap Measu es o E alua ion and Valida ion in Medical Image
Analysis. IEEE T ansac Med Imag. 2006;25:1451–61. h p://ieeex plo e .ieee.o g/docum en /17176 43/
62. Salehi SSM, E dogmus D, Gholipou A. T e sky loss unc ion o image segmen a ion using 3D ully con olu ional
deep ne wo ks. 2017. p. 379–87. h p://a xi .o g/abs/1706.05721
63. Be man M, T iki AR, Blaschko MB. The Lo asz-So max Loss: A ac able su oga e o he op imiza ion o he
in e sec ion-o e -union measu e in neu al ne wo ks. 2018 IEEE/CVF Con e ence on Compu e Vision and Pa e n
Recogni ion . IEEE; 2018. p. 4413–21. h ps ://ieeex plo e .ieee.o g/docum en /85785 62/
64. He Z, Zuo W, Kan M, Shan S, Chen X. A GAN: Facial a ibu e edi ing by only changing wha you wan . IEEE ans-
ac ions on image p ocessing . 2019;28:5464–78. h ps ://ieeex plo e .ieee.o g/docum en /87185 08/
65. Pe a nau G, an de Weije J, Raducanu B, Ál a ez JM. In e ible Condi ional GANs o image edi ing. Con e ence on
Neu al In o ma ion P ocessing Sys ems . 2016. h p://a xi .o g/abs/1611.06355
66. Tao R, Li Z, Tao R, Li B. ResA -GAN: Unpai ed deep esidual a ibu es lea ning o mul i-domain ace image ans-
la ion. IEEE Access . 2019;7:132594–608. h ps ://ieeex plo e .ieee.o g/docum en /88365 02/
67. Good ellow IJ, Pouge -Abadie J, Mi za M, Xu B, Wa de-Fa ley D, Ozai S, e al. Gene a i e ad e sa ial ne s. Ad
Neu al In P ocess Sys . 2014;3:2672–80.
68. Bowles C, Chen L, Gue e o R, Ben ley P, Gunn R, Hamme s A, e al. GAN Augmen a ion: augmen ing aining da a
using gene a i e ad e sa ial ne wo ks. 2018; h p://a xi .o g/abs/1810.10863
69. Oo d A an den, Kalchb enne N, Ka ukcuoglu K. Pixel ecu en neu al ne wo ks. 2016; h p://a xi .o g/
abs/1601.06759
70. Sejnowski MIJTJ. Lea ning and elea ning in bol zmann machines. G aphical models: ounda ions o neu al com-
pu a ion, MITP. 2001;
71. McClelland DERJL. In o ma ion p ocessing in dynamical sys ems: ounda ions o ha mony heo y. pa allel dis ib-
u ed p ocessing: explo a ions in he mic os uc u e o Cogni ion: Founda ions, MITP. 1987;194–281.
72. Hin on GE, Salakhu dino RR. Reducing he dimensionali y o da a wi h neu al ne wo ks. Science. 2006;313:504–7.
73. Salakhu dino R, Hin on G. Deep Bol zmann machines. J Machine Lea n Res. 2009;5:448–55.
74. Lee H, G osse R, Rangana h R, Y. Ng A. Con olu ional deep belie ne wo ks o scalable unsupe ised lea ning o
hie a chical ep esen a ions. Compu e Science Depa men , S an o d Uni e si y . 2009;8. h p:// obo ics.s an o d.
edu/~ang/pape s/icml0 9-Con o lu io nalDe epBel ie Ne wo k s.pd
75. Hin on GE, Osinde o S, Teh Y-W. A as lea ning algo i hm o deep belie ne s. Neu al Compu . 2006;18:1527–54.
h ps ://doi.o g/10.1162/neco.2006.18.7.1527.
76. Ramachand an P, Paine T Le, Kho ami P, Babaeizadeh M, Chang S, Zhang Y, e al. Fas gene a ion o con olu ional
au o eg essi e models. 2017; h p://a xi .o g/abs/1704.06001
77. F ey BJ. G aphical models o machine lea ning and digi al communica ion. Camb idge: MIT P ess; 1998.
78. F ey BJ, Hin on GE, Dayan P. Does he Wake-sleep algo i hm p oduce good densi y es ima o s? Ad ances in neu al
in o ma ion p ocessing sys ems . 1996;13:661–70. h p://www.cs.u o o n o.ca/~hin o n/absps /wspe .pd %5Cnpa
pe s2 ://publi ca io n/uuid/BCC05 47E-7C14-42EC-8693-D800C 5819C 79
79. U ia B, Cô é M-A, G ego K, Mu ay I, La ochelle H. Neu al au o eg essi e dis ibu ion es ima ion. J Mach Lea n Res.
2016;17:1–37. h p://a xi .o g/abs/1605.02226
80. Schulle B, Wöllme M, Moosmay T, Rigoll G. Recogni ion o noisy speech: a compa a i e su ey o obus model
a chi ec u e and ea u e enhancemen . EURASIP J Audio Speech Music P ocess. 2009;2009:942617. h p://asmp.
eu as ipjou nals .com/con e n /2009/1/94261 7
81. Yang S, Lu H, Kang S, Xue L, Xiao J, Su D, e al. On he localness modeling o he sel -a en ion based end- o-end
speech syn hesis. Neu al Ne w. 2020;125:121–30. h ps ://linki nghub .else ie .com/ e i e e/pii/S0893 60802 03004 47
82. Ghosh R, Vamshi C, Kuma P. RNN based online handw i en wo d ecogni ion in De anaga i and Bengali sc ip s
using ho izon al zoning. Pa e n Recogni . 2019;92:203–18. h ps ://linki nghub .else ie .com/ e i e e/pii/S0031
32031 93013 84
83. Chen J, Zhuge H. Ex ac i e summa iza ion o documen s wi h images based on mul i-modal RNN. Fu u e Gen-
e a Compu Sys . 2019;99:186–96. h ps ://linki nghub .else ie .com/ e i e e/pii/S0167 739X1 83268 76
84. Hoch ei e S, Schmidhube J. Long sho - e m memo y. Neu al Compu . 1997;9:1735–80. h ps ://doi.o g/10.1162/
neco.1997.9.8.1735.
Page 55 o 59
Sampa he al. J Big Da a (2021) 8:27
85. Vaswani A, Shazee N, Pa ma N, Uszko ei J, Jones L, Gomez AN, e al. A en ion is all you need. a Xi . 2017; h p://
a xi .o g/abs/1706.03762
86. Theis L, Be hge M. Gene a i e Image Modeling Using Spa ial LSTMs. P oceedings o he 28 h In e na ional Con e -
ence on Neu al In o ma ion P ocessing Sys ems–Volume 2. Camb idge: MIT P ess; 2015. p. 1927–1935.
87. K izhe sky A. Lea ning mul iple laye s o ea u es om iny images . 2009. h p://www.cs. o on o.edu/~k iz/ci a
.h ml
88. Russako sky O, Deng J, Su H, K ause J, Sa heesh S, Ma S, e al. ImageNe la ge scale isual ecogni ion challenge.
In J Compu Vis. 2015;115:211–52. h ps ://doi.o g/10.1007/s1126 3-015-0816-y.
89. Oo d A an den, Kalchb enne N, Vinyals O, Espehol L, G a es A, Ka ukcuoglu K. Condi ional image gene a ion
wi h PixelCNN Decode s. h p://a xi .o g/abs/1606.05328
90. Salimans T, Ka pa hy A, Chen X, Kingma DP. PixelCNN++: Imp o ing he PixelCNN wi h disc e ized logis ic mix u e
likelihood and o he modifica ions. 2017; h p://a xi .o g/abs/1701.05517
91. Chen X, Mish a N, Rohaninejad M, Abbeel P. PixelSNAIL: an imp o ed au o eg essi e gene a i e model. 2017.
h p://a xi .o g/abs/1712.09763
92. Vincen P, La ochelle H, Bengio Y, Manzagol P-A. Ex ac ing and composing obus ea u es wi h denoising au oen-
code s. P oceedings o he 25 h in e na ional con e ence on Machine lea ning - ICML ’08 . New Yo k: ACM P ess;
2008. p. 1096–103. h ps ://linki nghub .else ie .com/ e i e e/pii/S0925 23121 83061 55
93. Baldi P. Au oencode s, unsupe ised lea ning, and deep a chi ec u es . PMLR; 2012. h p://p oce eding s.ml .p ess /
27/baldi 12a.h ml
94. Y. Ng A. Spa se au oencode .h ps ://web.s an o d.edu/class /cs294 a/spa s eAu o encod e .pd
95. Masci J, Meie U, Ci eşan D, Schmidhube J. S acked con olu ional au o-encode s o hie a chical ea u e ex ac-
ion. 2011. p. 52–9. h ps ://doi.o g/10.1007/978-3-642-21735 -7_7
96. Ri ai S, Vincen P, Mulle X, Glo o X, Bengio Y. Con ac i e au o-encode s: explici in a iance du ing ea u e ex ac-
ion. ICML. 2011.
97. Kingma DP, Welling M. Au o-encoding a ia ional bayes. 2013; h p://a xi .o g/abs/1312.6114
98. Tan S, Li B. S acked con olu ional au o-encode s o s eganalysis o digi al images. Signal and In o ma ion P ocess-
ing Associa ion Annual Summi and Con e ence (APSIPA), 2014 Asia-Pacific. IEEE; 2014. p. 1–4.
99. Ge main M, G ego K, Mu ay I, La ochelle H. MADE: Masked au oencode o dis ibu ion es ima ion. 2015. h p://
a xi .o g/abs/1502.03509
100. Schmidhube J. Lea ning ac o ial codes by p edic abili y minimiza ion. Neu al Compu . 1992;4:863–79. h ps ://
doi.o g/10.1162/neco.1992.4.6.863.
101. Sohn K, Yan X, Lee H. Lea ning s uc u ed ou pu ep esen a ion using deep condi ional gene a i e models. Ad
Neu al In o ma P ocess Sys . 2015;2015-Janua:3483–91.
102. Higgins I, Ma hey L, Pal A, Bu gess C, Glo o X, Bo inick M, e al. Β-VAE: Lea ning basic isual concep s wi h a
cons ained a ia ional amewo k. 5 h In e na ional Con e ence on Lea ning Rep esen a ions, ICLR 2017–Con e -
ence T ack P oceedings. 2019;1–13.
103. Kulka ni TD, Whi ney W, Kohli P, Tenenbaum JB. Deep con olu ional in e se g aphics ne wo k. 2015. h p://a xi
.o g/abs/1503.03167
104. Huang C-W, Sanka an K, Dhekane E, Lacos e A, Cou ille A. Hie a chical Impo ance Weigh ed Au oencode s. In:
Chaudhu i K, Salakhu dino R, edi o s. Long Beach, Cali o nia, USA: PMLR; 2019. p. 2869–78. h p://p oce eding
s.ml .p ess / 97/huang 19d.h ml
105. Gul ajani I, Kuma K, Ahmed F, Taiga AA, Visin F, Vazquez D, e al. PixelVAE: A la en a iable model o na u al
images. 2016; Ah p://a xi .o g/abs/1611.05013
106. Chen X, Kingma DP, Salimans T, Duan Y, Dha iwal P, Schulman J, e al. Va ia ional Lossy Au oencode . 2016. h p://
a xi .o g/abs/1611.02731
107. G ego K, Danihelka I, G a es A, Rezende DJ, Wie s a D. DRAW: A ecu en neu al ne wo k o image gene a ion.
2015. h p://a xi .o g/abs/1502.04623
108. Oo d A an den, Vinyals O, Ka ukcuoglu K. Neu al Disc e e Rep esen a ion Lea ning. 31s Con e ence on Neu al
In o ma ion P ocessing Sys ems . Long Beach, Cali o nia, USA; 2017. h p://a xi .o g/abs/1711.00937
109. Raza i A, Oo d A an den, Vinyals O. Gene a ing di e se high-fideli y images wi h VQ-VAE-2. Ad ances in neu al
in o ma ion p ocessing sys ems 32. 2019. h p://a xi .o g/abs/1906.00446
110. Huszá F. How (no ) o T ain you gene a i e model: scheduled sampling, likelihood, ad e sa y? 2015. h p://a xi
.o g/abs/1511.05101
111. Lo e W, K eiman G, Cox D. Deep P edic i e coding ne wo ks o ideo p edic ion and unsupe ised lea ning.
2016. h p://a xi .o g/abs/1605.08104
112. Rad o d A, Me z L, Chin ala S. Unsupe ised ep esen a ion lea ning wi h deep con olu ional gene a i e ad e -
sa ial ne wo ks. 2015. h p://a xi .o g/abs/1511.06434
113. Makhzani A, Shlens J, Jai ly N, Good ellow I, F ey B. Ad e sa ial Au oencode s. 2015; A ailable om: h p://a xi
.o g/abs/1511.05644
114. Dumoulin V, Belghazi I, Poole B, Mas opie o O, Lamb A, A jo sky M, e al. Ad e sa ially Lea ned In e ence. 2016.
h p://a xi .o g/abs/1606.00704
115. La sen ABL, Sønde by SK, La ochelle H, Win he O. Au oencoding beyond pixels using a lea ned simila i y me ic.
2015. h p://a xi .o g/abs/1512.09300
116. Zhong G, Gao W, Liu Y, Yang Y. Gene a i e Ad e sa ial ne wo ks wi h decode -encode ou pu noise. 2018. h p://
a xi .o g/abs/1807.03923
117. S i as a a A, Valko L, Russell C, Gu mann MU, Su on C. VEEGAN: Reducing Mode Collapse in GANs using implici
a ia ional lea ning. 2017. h p://a xi .o g/abs/1705.07761
118. Mi za M, Osinde o S. Condi ional gene a i e ad e sa ial ne s. 2014. h p://a xi .o g/abs/1411.1784
119. Odena A, Olah C, Shlens J. Condi ional image syn hesis wi h auxilia y classifie GANs. 2016. h p://a xi .o g/
abs/1610.09585
Page 56 o 59
Sampa he al. J Big Da a (2021) 8:27
120. Baz a kan S, Co co an P. Ve sa ile auxilia y classifie wi h gene a i e ad e sa ial ne wo k (VAC+GAN), Mul i Class
Scena ios. 2018. h p://a xi .o g/abs/1806.07751
121. Chen X, Duan Y, Hou hoo R, Schulman J, Su ske e I, Abbeel P. In oGAN: In e p e able ep esen a ion lea ning by
in o ma ion maximizing gene a i e ad e sa ial ne s. 2016. h p://a xi .o g/abs/1606.03657
122. Li X, Chen L, Wang L, Wu P, Tong W. SCGAN: disen angled ep esen a ion lea ning by adding simila i y cons ain
on gene a i e ad e sa ial ne s. IEEE Access . 2019;7:147928–38. h ps ://ieeex plo e .ieee.o g/docum en /84762 90/
123. A jo sky M, Chin ala S, Bo ou L. Wasse s ein GAN. 2017. h p://a xi .o g/abs/1701.07875
124. Gul ajani I, Ahmed F, A jo sky M, Dumoulin V, Cou ille A. Imp o ed aining o Wasse s ein GANs. 2017. h p://
a xi .o g/abs/1704.00028
125. Pe zka H, Fische A, Luko nico D. On he egula iza ion o Wasse s ein GANs. 2017. h p://a xi .o g/
abs/1709.08894
126. Mao X, Li Q, Xie H, Lau RYK, Wang Z, Smolley SP. Leas squa es gene a i e ad e sa ial ne wo ks. 2016. h p://a xi
.o g/abs/1611.04076
127. Zhao J, Ma hieu M, LeCun Y. Ene gy-based Gene a i e Ad e sa ial Ne wo k. 2016. h p://a xi .o g/abs/1609.03126
128. Be helo D, Schumm T, Me z L. BEGAN: Bounda y Equilib ium Gene a i e Ad e sa ial Ne wo ks. 2017. h p://a xi
.o g/abs/1703.10717
129. Wang R, Cully A, Chang HJ, Demi is Y. MAGAN: Ma gin adap a ion o gene a i e ad e sa ial ne wo ks. 2017. h p://
a xi .o g/abs/1704.03817
130. Zhao J, Xiong L, Jayash ee K, Li J, Zhao F, Wang Z, e al. Dual-agen GANs o pho o ealis ic and iden i y p ese ing
p ofile ace syn hesis. Ad an Neu al In o ma P ocess Sys . 2017;2017:66–76.
131. Ka as T, Aila T, Laine S, Leh inen J. P og essi e g owing o GANs o imp o ed quali y, s abili y, and a ia ion. 2017;
h p://a xi .o g/abs/1710.10196
132. Den on E, Chin ala S, Szlam A, Fe gus R. Deep gene a i e image models using a laplacian py amid o ad e sa ial
ne wo ks. Ad ances in Neu al In o ma ion P ocessing Sys ems 28 . 2015. h p://a xi .o g/abs/1506.05751
133. Im DJ, Kim CD, Jiang H, Memise ic R. Gene a ing images wi h ecu en ad e sa ial ne wo ks. 2016; h p://a xi
.o g/abs/1602.05110
134. Nguyen TD, Le T, Vu H, Phung D. Dual disc imina o gene a i e ad e sa ial Ne s. 2017; h p://a xi .o g/
abs/1709.03831
135. Ghosh A, Kulha ia V, Namboodi i V, To PHS, Dokania PK. Mul i-agen di e se gene a i e ad e sa ial ne wo ks.
2017. h p://a xi .o g/abs/1704.02906
136. Liu M-Y, Tuzel O. Coupled gene a i e ad e sa ial ne wo ks. con e ence on neu al in o ma ion p ocessing sys ems.
2016. h p://a xi .o g/abs/1606.07536
137. Kim T, Cha M, Kim H, Lee JK, Kim J. Lea ning o disco e c oss-domain ela ions wi h gene a i e ad e sa ial ne -
wo ks. 2017. h p://a xi .o g/abs/1703.05192
138. Zhu J-Y, Pa k T, Isola P, E os AA. Unpai ed Image- o-image ansla ion using cycle-consis en ad e sa ial ne -
wo ks. 2017 IEEE In e na ional Con e ence on Compu e Vision (ICCV) . IEEE; 2017. p. 2242–51. h p://a xi .o g/
abs/1703.10593
139. Ledig C, Theis L, Husza F, Caballe o J, Cunningham A, Acos a A, e al. Pho o- ealis ic single image supe - esolu ion
using a gene a i e ad e sa ial ne wo k. 2016; h p://a xi .o g/abs/1609.04802
140. Simonyan K, Zisse man A. Ve y deep con olu ional ne wo ks o la ge-scale image ecogni ion. 2014; h p://a xi
.o g/abs/1409.1556
141. Zhang H, Good ellow I, Me axas D, Odena A. Sel -A en ion Gene a i e Ad e sa ial Ne wo ks. 2018; h p://a xi
.o g/abs/1805.08318
142. Isola P, Zhu J-Y, Zhou T, E os AA. Image- o-image ansla ion wi h condi ional ad e sa ial ne wo ks. 2017 IEEE
Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR). IEEE; 2017. p. 5967–76. h p://ieeex plo e .ieee.
o g/docum en /81001 15/
143. Wang T-C, Liu M-Y, Zhu J-Y, Tao A, Kau z J, Ca anza o B. High- esolu ion image syn hesis and seman ic manipula-
ion wi h condi ional GANs. 2018 IEEE/CVF Con e ence on Compu e Vision and Pa e n Recogni ion . IEEE; 2018.
p. 8798–807. h ps ://ieeex plo e .ieee.o g/docum en /85790 15/
144. Bellema e MG, Danihelka I, Dabney W, Mohamed S, Lakshmina ayanan B, Hoye S, e al. The c ame dis ance as a
solu ion o biased wasse s ein g adien s. 2017. h p://a xi .o g/abs/1705.10743
145. M oueh Y, Se cu T, Goel V. McGan: mean and co a iance ea u e ma ching GAN. 2017. h p://a xi .o g/
abs/1702.08398
146. Li C-L, Chang W-C, Cheng Y, Yang Y, Póczos B. MMD GAN: owa ds deepe unde s anding o momen ma ching
ne wo k. 2017. h p://a xi .o g/abs/1705.08584
147. M oueh Y, Se cu T. Fishe GAN. 2017. h p://a xi .o g/abs/1705.09675
148. Salimans T, Good ellow I, Za emba W, Cheung V, Rad o d A, Chen X. Imp o ed echniques o aining GANs. 2016.
h p://a xi .o g/abs/1606.03498
149. Sønde by CK, Caballe o J, Theis L, Shi W, Huszá F. Amo ised MAP in e ence o image supe - esolu ion. 2016.
h p://a xi .o g/abs/1610.04490
150. Heusel M, Ramsaue H, Un e hine T, Nessle B, Hoch ei e S. GANs ained by a wo ime-scale upda e ule con-
e ge o a local nash equilib ium. 2017. h p://a xi .o g/abs/1706.08500
151. Miya o T, Ka aoka T, Koyama M, Yoshida Y. Spec al no maliza ion o gene a i e ad e sa ial ne wo ks. 2018. h p://
a xi .o g/abs/1802.05957
152. Hea h M, Bowye K, Kopans D, Moo e R, Kegelmeye WP. Digi al da abase o sc eening mammog aphy . h ps ://
www.mammo image .o g/da ab ases/
153. Shoohi LM, Saud JH. Dcgan o handling imbalanced mala ia da ase based on o e -sampling echnique and
using cnn. Medico-Legal Upda e. 2020;20:1079–85.
154. Niu S, Li B, Wang X, Lin H. De ec image sample gene a ion Wi h GAN o Imp o ing de ec ecogni ion. IEEE T ans-
ac ions on Au oma ion Science and Enginee ing . 2020;1–12. h ps ://ieeex plo e .ieee.o g/docum en /90008 06/
Page 57 o 59
Sampa he al. J Big Da a (2021) 8:27
155. Ma iani G, Scheidegge F, Is a e R, Bekas C, Malossi C. BAGAN: Da a Augmen a ion wi h Balancing GAN. 2018;
h p://a xi .o g/abs/1803.09655
156. Wu E, Wu K, Cox D, Lo e W. Condi ional infilling GANs o da a augmen a ion in mammog am classifica ion. 2018.
p. 98–106. Doi: h ps ://doi.o g/10.1007/978-3-030-00946 -5_11
157. Mu ama su C, Nishio M, Go o T, Oiwa M, Mo i a T, Yakami M, e al. Imp o ing b eas mass classifica ion by sha ed
da a wi h domain ans o ma ion using a gene a i e ad e sa ial ne wo k. Compu Biol Med. 2020;119:103698.
h ps ://linki nghub .else ie .com/ e i e e/pii/S0010 48252 03008 6X
158. Guan S. B eas cance de ec ion using syn he ic mammog ams om gene a i e ad e sa ial ne wo ks in con olu-
ional neu al ne wo ks. J Med Imag. 2019;6:1. h ps ://doi.o g/10.1117/1.JMI.6.3.03141 1. ull.
159. Waheed A, Goyal M, Gup a D, Khanna A, Al-Tu jman F, Pinhei o PR. Co idGAN: Da a augmen a ion using auxilia y
classifie GAN o imp o ed Co id-19 de ec ion. IEEE Access . 2020;8:91916–23. h ps ://ieeex plo e .ieee.o g/docum
en /90938 42/
160. COVID-19 Ches X-Ray da ase ini ia i e. h ps ://gi hu b.com/agchu ng/Figu e1-COVID -ches x ay-da as e
161. Cohen JP, Mo ison P, Dao L, Ro h K, Duong TQ, Ghassemi M. COVID-19 Image da a collec ion: p ospec i e p edic-
ions a e he u u e. 2020. h p://a xi .o g/abs/2006.11988
162. Co id19 adiog aphy da abase. h ps ://www.kaggl e.com/ awsi u a hman/co id 19- adio g aph y-da ab ase
163. Hase N, I o S, Kanaeko N, Sumi K. Da a augmen a ion o in a-class imbalance wi h gene a i e ad e sa ial
ne wo k. In: Cudel C, Bazeille S, Ve ie N, edi o s. Fou een h In e na ional Con e ence on Quali y Con ol by
A ificial Vision . SPIE; 2019. p. 56. A ailable om: h ps://www.spiedigi allib a y.o g/con e ence-p oceedings-o -
spie/11172/2521692/Da a-augmen a ion- o -in a-class-imbalance-wi h-gene a i e-ad e sa ial-ne wo k/h ps ://
doi.o g/10.1117/12.25216 92. ull
164. Donahue C, Lip on ZC, Balsub amani A, McAuley J. Seman ically Decomposing he La en Spaces o Gene a i e
Ad e sa ial Ne wo ks. 2017; h p://a xi .o g/abs/1705.07904
165. Wang Y, Gong D, Zhou Z, Ji X, Wang H, Li Z, e al. O hogonal deep ea u es decomposi ion o age-in a ian ace
ecogni ion. 2018. p. 764–79. h ps ://doi.o g/10.1007/978-3-030-01267 -0_45
166. Gong D, Li Z, Lin D, Liu J, Tang X. Hidden ac o analysis o age in a ian ace ecogni ion. 2013 IEEE In e na ional
Con e ence on Compu e Vision. IEEE; 2013. p. 2872–9. h p://ieeex plo e .ieee.o g/docum en /67514 68/
167. Yin X, Liu X. Mul i- ask con olu ional neu al ne wo k o pose-in a ian ace ecogni ion. IEEE T ansac ions on
Image P ocessing. 2018;27:964–75. h p://ieeex plo e .ieee.o g/docum en /80802 44/
168. Ca cagnì P, Del CM, Cazza o D, Leo M, Dis an e C. A s udy on diffe en expe imen al configu a ions o age, ace,
and gende es ima ion p oblems. EURASIP J Image Video P ocess. 2015;2015:37. h ps ://doi.o g/10.1186/s1364
0-015-0089-y.
169. Ziwei L, Ping L, Xiaogang W, Tang X. La ge-scale CelebFaces a ibu es (CelebA) Da ase . 2018. h p://mmlab .ie.
cuhk.edu.hk/p oje c s/Celeb A.h ml
170. Zhang J, Li A, Liu Y, Wang M. Ad e sa ially Regula ized U-Ne -based GANs o acial a ibu e modifica ion and
gene a ion. IEEE Access . 2019;7:86453–62. h ps ://ieeex plo e .ieee.o g/docum en /87547 28/
171. Zhang G, Kan M, Shan S, Chen X. Gene a i e ad e sa ial ne wo k wi h spa ial a en ion o ace a ibu e edi ing.
2018. p. 422–37. h ps ://doi.o g/10.1007/978-3-030-01231 -1_26
172. Zheng Z, Yang X, Yu Z, Zheng L, Yang Y, Kau z J. join disc imina i e and gene a i e lea ning o pe son e-iden i-
fica ion. 2019 IEEE/CVF Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR) . IEEE; 2019. p. 2133–42.
h ps ://ieeex plo e .ieee.o g/docum en /89542 92/
173. Zhang X, Gao Y. Face ecogni ion ac oss pose: a e iew. pa e n ecogni ion . 2009;42:2876–96. h ps ://linki nghub
.else ie .com/ e i e e/pii/S0031 32030 90015 38
174. Tan X, Chen S, Zhou Z-H, Zhang F. Face ecogni ion om a single image pe pe son: a su ey. pa e n ecogni ion.
2006;39:1725–45. h ps ://linki nghub .else ie .com/ e i e e/pii/S0031 32030 60012 70
175. Zhao W, Chellappa R, Phillips PJ, Rosen eld A. Face ecogni ion. ACM compu ing su eys. 2003;35:399–458. h p://
po a l.acm.o g/ci a ion.c m?doid=95433 9.95434 2
176. Qian X, Fu Y, Xiang T, Wang W, Qiu J, Wu Y, e al. Pose-No malized Image Gene a ion o Pe son Re-iden ifica ion.
2018. p. 661–78. h ps ://doi.o g/10.1007/978-3-030-01240 -3_40
177. Wei L, Zhang S, Gao W, Tian Q. Pe son T ans e GAN o b idge domain gap o pe son e-iden ifica ion. 2018 IEEE/
CVF con e ence on compu e ision and pa e n ecogni ion . IEEE; 2018. p. 79–88. h ps ://ieeex plo e .ieee.o g/
docum en /85781 14/
178. Zhong Z, Zheng L, Zheng Z, Li S, Yang Y. Came a s yle adap a ion o pe son e-iden ifica ion. 2018 IEEE/CVF con-
e ence on compu e ision and pa e n ecogni ion. IEEE; 2018. p. 5157–66. h ps ://ieeex plo e .ieee.o g/docum
en /85786 39/
179. Deng W, Zheng L, Ye Q, Yang Y, Jiao J. Simila i y-p ese ing image-image domain adap a ion o pe son e-
iden ifica ion. 2018; h p://a xi .o g/abs/1811.10551
180. Ge Y, Li Z, Zhao H, Yin G, Yi S, Wang X, e al. FD-GAN: Pose-guided Fea u e Dis illing GAN o obus pe son e-
iden ifica ion. Ad Neu al In o ma P ocess Sys . 2018;2018:1222–33.
181. Zheng A, Lin X, Li C, He R, Tang J. A ibu es guided ea u e lea ning o ehicle e-iden ifica ion. 2019; h p://a xi
.o g/abs/1905.08997
182. Zhou Y, Shao L. C oss-View GAN Based Vehicle Gene a ion o Re-iden ifica ion. P ocedings o he B i ish Machine
Vision Con e ence 2017 . B i ish Machine Vision Associa ion; 2017. h p://www.bm a.o g/bm c/2017/pape s/
pape 186/index .h ml
183. Wu F, Yan S, Smi h JS, Zhang B. Vehicle e-iden ifica ion in s ill images: applica ion o semi-supe ised lea ning and
e- anking. Signal P ocessing: Image Communica ion . 2019;76:261–71. h ps ://linki nghub .else ie .com/ e i e e/
pii/S0923 59651 83058 00
184. Fu Y, Li X, Ye Y. A mul i- ask lea ning model wi h ad e sa ial da a augmen a ion o classifica ion o fine-g ained
images. Neu ocompu ing . 2020;377:122–9. h ps ://linki nghub .else ie .com/ e i e e/pii/S0925 23121 93137 48