Recei ed 9 Augus 2023, accep ed 6 Oc obe 2023, da e o publica ion 16 Oc obe 2023, da e o cu en e sion 9 No embe 2023.
Digi al Objec Iden i ie 10.1109/ACCESS.2023.3325062
Deep Rein o cemen Lea ning T -Agen -Based
Objec T acking Wi h Vi ual Au onomous D one
in a Game Engine
KHURSHEDJON FARKHODOV 1, SUK-HWAN LEE2,
JAN PLATOS 3, (Membe , IEEE), AND KI-RYONG KWON 1
1Depa men o AI Con e gence, Pukyong Na ional Uni e si y, Busan 48513, Sou h Ko ea
2Depa men o Compu e Enginee ing, Dong-A Uni e si y, Busan 49315, Sou h Ko ea
3Depa men o Elec ical Enginee ing and Compu e Science, VSB—Technical Uni e si y o Os a a, 708 00 Os a a, Czech Republic
Co esponding au ho : Ki-Ryong Kwon ([email p o ec ed])
This wo k was suppo ed in pa by he Minis y o Science and In o ma ion and Communica ion Technology (MSIT), Sou h Ko ea,
h ough he In o ma ion Technology Resea ch Cen e (ITRC) Suppo P og am, Supe ised by he Ins i u e o In o ma ion and
Communica ions Technology Planning and E alua ion (IITP), unde G an IITP-2023-2020-0-01797; and in pa by he MSIT h ough he
ICT Consilience C ea i e P og am, Supe ised by he IITP, unde G an IITP-2023-2016-0-00318.
ABSTRACT The ecen de elopmen o objec - acking amewo ks has a ec ed he pe o mance o
many manu ac u ing and indus ial se ices such as p oduc deli e y, au onomous d i ing sys ems, secu i y
sys ems, mili a y, anspo a ion and e ailing indus ies, sma ci ies, heal hca e sys ems, ag icul u e, e c.
Achie ing accu a e esul s in physical en i onmen s and condi ions emains qui e challenging o he ac ual
objec - acking. Howe e , he p ocess can be expe imen ed wi h using simula ion echniques o pla o ms
o e alua e and check he model’s pe o mance unde di e en simula ion condi ions and wea he changes.
This pape p esen s one o he a ge acking app oaches based on he ein o cemen lea ning echnique
in eg a ed wi h Tenso Flow-Agen ( -agen ) o accomplish he acking p ocess in he Un eal Game Engine
simula ion pla o m Ai Sim Blocks. The p oduc i i y o hese pla o ms can be seen while expe imen ing
in i ual- eali y condi ions wi h i ual d one agen s and pe o ming ine- uning o achie e he bes o
desi ed pe o mance. In his pape , he -agen d one lea ns how o ack an objec in eg a ion wi h a deep
ein o cemen lea ning p ocess o con ol he ac ions, s a es, and acking by ecei ing sequen ial ames
om a simple Blocks en i onmen . The -agen model is ained in he Ai Sim Blocks en i onmen o
adap a ion o he en i onmen and exis ing objec s in a simula ion en i onmen o u he es ing and
e alua ion ega ding he accu acy o acking and speed. We es ed and compa ed wo app oaches, DQN
and PPO acke s, and epo ed esul s in e ms o s abili y, ewa ds, and nume ical pe o mance.
INDEX TERMS Objec acking, objec de ec ion, ein o cemen lea ning, Ai Sim, i ual en i onmen ,
i ual simula ion, -agen , un eal game engine.
I. INTRODUCTION
Unmanned Ae ial Vehicles (UAV) a e widely u ilized in
e e y ield o manu ac u ing and daily se ice p ocedu es [1],
whe e d ones help o make he p ocess easie by au oma -
ing and dec easing he consump ion ime wi h sa e se ice
ac i i ies. Fo example, in ecen yea s, d ones ha e become
a op p io i y o echnical assis ance in he ag icul u al ield
The associa e edi o coo dina ing he e iew o his manusc ip and
app o ing i o publica ion was Gangyi Jiang.
o he medica ion and e iliza ion p ocesses. Addi ionally,
d ones can obse e cul i a ed ield mapping p ope ly [2]
wi h impo an in o ma ion abou any i egula changes in
he g owe ’s ield. In mos cases, i can be bene icial o
inc easing he p oduc i i y o la ge plan a ions wi h congen-
e ous g owe s o de e mine d ainage pa e ns and we -d y
spo s o ield ele a ion ha allow mo e e icien wa e ing
echniques. Mo eo e , he e a e se e al ields in which d ones
a e being applied such as su eillance and con ol o la ge
un eachable a eas [3], medical pu poses such as Au oma ed
VOLUME 11, 2023
2023 The Au ho s. This wo k is licensed unde a C ea i e Commons A ibu ion-NonComme cial-NoDe i a i es 4.0 License.
Fo mo e in o ma ion, see h ps://c ea i ecommons.o g/licenses/by-nc-nd/4.0/ 124129
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
Eme gency D ones (AED) [4], sea ch and escue ope a ions
[5], oad a ic moni o ing applica ions [6], wea he sensing
in an u ban en i onmen applica ion [7], eme gency sys em
applica ions o i e igh e p ocesses [8], and en e ainmen
(pho og aphy and cinema og aphy) [9]. Many ecen inno a-
ions and echnologies ela ed o UAV amewo ks include
su eillance and a ge acking uni s o se e al easons,
including secu i y main enance, con ol, mili a y assis ance,
and a ic moni o ing, besides manu ac u ing op imiza ion
and con ol. The cu en de elopmen o inno a i e ech-
nologies is aking no el ideas om modeling sys ems such
as i ual eali y en i onmen simula o s. C ea ing a i ual
e sion o he physical objec s and p ocess simula ions can
p o ide p ope se ice op imiza ion and allow o ee expe -
imen a ion wi h condi ions o any ac i i y.
A. IMPORTANCE AND DEVELOPMENT OF OBJECT
TRACKING WITH MACHINE LEARNING
In ecen yea s, mos o he UAV-based esea ch communi y
has paid a en ion o imp o ing he pe o mance o isual
acking echniques wi h se e al neu al ne wo ks ela ed o
ne wo k-based a chi ec u es, such as CNN [10], DNN [11],
LSTM [12], RL [13], and o he s. Howe e , hese esea ch
wo ks p esen aining and es ing esul s in physical en i on-
men da ase s, such as image collec ions aken om di e en
came as in public a eas, and image/ ideo se s aken by d one
came as. These objec - acking applica ions wo k wi h ce -
ain objec classes whe e he d one acks dynamic objec s in
an exac pa hway and localizes objec s wi h speci ic me hods
as an addi ional ask o he a ge - acking amewo k. S ill,
he e a e challenges while acking mo ing and s a ic objec s
wi h appa en ly iden ical aspec s o ecognize which one is an
ac ual acking a ge . Recen objec - acking esea ch de el-
opmen has begun wi h in ensi e lea ning wi h i ual eali y
in eg a ion pla o ms [14],[15], so scien is s can in eg a e
hei echnique o algo i hm wi h a i ual eali y pla o m o
es hei p oposals wi h a ious ine- uning pa ame e s and
condi ions. I gi es schola s mo e oppo uni ies o explo e
hei me hods be e and mo e deeply in a ha dwa e- ee en i-
onmen a ze o cos and op imize hem as much as possible.
Addi ionally, he e a e some echniques mo i a ed by obus
ex ension o in eg al schemes o misma ched unce ain non-
linea sys ems p oposed o suppo asymp o ic acking [16].
Asymp o ic acking means ensu ing ha he sys ems’ ou pu
acks a desi ed e e ence ajec o y o e ime wi h negli-
gible acking e o . The main goal is o design a acking
con ol sys em ha gua an ees he ou pu con e ges o he
e e ence ajec o y as ime app oaches in ini y in unce ain
en i onmen s. Ano he model is he ou pu eedback adap-
i e ise con ol echnique [17] used o unce ain nonlinea
sys ems o achie e accu a e acking o desi ed ajec o ies.
The e m ‘‘adap i e’’ indica es ha he con olle pa ame e s
a e upda ed online based on he sys em’s beha io and he
acking e o . A deep Q-lea ning-based [18] app oach has
been sugges ed o i e igh ing si ua ions, which is ob ainable
in some agen obo s o d ones o inding o planning pa hs
and na iga ing h ough i e en i onmen s. In such complex
and haza dous cases, i is equi ed o be mo e ca e ul o con-
ol he si ua ion wi h conc e e plans and ac ions o escuing
inju ed o ic ims o he inciden by coo dina ing si ua ional
awa eness wi h o he escue s, which is an u gen ask. When
he amewo k is ins alled and applied o eal d ones o obo s,
i can ensu e i e igh e s o escue s make he igh decision
in ex eme, panicky, and diso ien ing condi ions.
B. DESCRIPTION OF THE VR-BASED WORK AND MAIN
CONTRIBUTION
Se e al ele an esea ch s udies ha in eg a e i ual simula-
ion pla o ms wi h p oposal algo i hms ha e been published.
Kalidas e al. [19] p esen ed ision-based na iga ion o UAVs
based simply on image da a by employing deep ein o ce-
men lea ning o a oid s a iona y and mo able obs acles
au onomously in disc e e and con inuous ac ion space. W.
Zhao e al. [20] also p oposed a pe cep ion-based hie a chi-
cal ac i e acking con ol o UAVs deploying a high-le el
con olle and ac ion o de s in a V-REP-based en i onmen .
A ained PPO algo i hm [21] wi h ewa d shaping o ai c a
di ec ion o a mo ing des ina ion in a h ee-dimensional con-
inuous space model was sugges ed, wi h he agen -speci ic
a ge guidance in i ual s a e space using a no el ewa d
calcula ion. Using a PPO-based DRL algo i hm [22] was sug-
ges ed o UAV acking wi h he assis ance o ano he UAV,
in oducing he gene alized dis ibu ed deep ein o cemen
lea ning pla o m, which p o ides solu ions o o e come a i-
ous p oblems such as acking, con olling, and mission coo -
dina ion o UAVs. Mo eo e , M. A. B. Abdelkade e al. [23]
p opose RL-based d one ele a ion con ol on a Py hon-Uni y
in eg a ed simula ion amewo k o achie e a s able use
diag am p o ocol (UDP) wi h he sugges ed algo i hm. Çe in
[24] p oposes coun ing d ones in a 3D space wi h se e al
DRL me hods p esen o coun d ones wi h ano he d one in
he en i onmen p o ided by an Ai Sim simula o .
In his s udy, we de eloped an algo i hm based on -agen
d one acking in a Blocks en i onmen whe e he -agen
ac i ely makes decisions o ack he a ge objec in he
un ime en i onmen . The p oposed me hod includes di e -
en ewa d echniques o boos he lea ning, acking, and
decision-making p ocesses ia a TF-agen -based d one in a
simula ion pla o m. The e is some compu a ional conside -
a ion o co ec ly applying pa ame e alues o achie e a
highe accu acy a e, and s a e ep esen a ion was o mula ed
o clea ou unnecessa y losses and cons ain s o he aining
and es ing p ocesses.
The p ima y con ibu ions o his pape can be summa ized
as ollows:
•We in oduce a i ual en i onmen al simula ion-based
objec - acking algo i hm model ha ecei es inpu
images di ec ly om a ealis ic i ual pla o m.
•Di ec access o he ne wo k eedable sou ce images
om he simula ion en i onmen p o ides he amewo k
124130 VOLUME 11, 2023
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
he mo e ad an ages o lea ning and es ing when i
comes o unknown en i onmen al condi ions.
•The expe imen is implemen ed in an Ai Sim-based
basic Blocks en i onmen wi h a andom, pa icula
walking pe son being acked by a i ual d one agen .
•Two di e en me hods we e adop ed and in eg a ed wi h
he i ual simula ion pla o m o demons a e he pe -
o mance o he models.
II. RELATED WORKS
The ecen de elopmen o objec acking ia ein o ce-
men lea ning has imp o ed by in eg a ing i wi h many
a ge acking echniques, which p oduce be e pe o mance
wi h decision-making in acking p ocedu es. Al hough mos
isual acking concep s based on DRL could pe o m be e
in he case o he ep esen a ion model wi h adop ed manne s
o loca ing he a ge objec wi hin a sea ch egion, he inal
es ima ed a ge coo dina es a e ideally cen e ed.
A. OBJECT TRACKING VIA DEEP REINFORCEMENT
LEARNING
The ad ancemen o objec acking ia ein o cemen lea n-
ing is a compa a i ely no el idea, whe e objec localiza ion
and acking in eg a ion wi h a decision-making model [13],
[25],[26],[27],[28],[29],[30] a e applied o he lea ning and
acking p ocess as well. Se e al s udies ha e disco e ed ha
combining deep and RL [31] in a ious se ings con e s many
ad an ages. Visual objec acking [32], localizing empo al
ac i i y [33], iden i ying objec classes [26], objec ecogni-
ion h ough ideo sequence [34], and segmen a ion [35] a e
jus a ew o he compu e ision p oblems ha ha e used
DRL. No ably, isual objec acking ia DRL amewo k
s udies has inc eased in ecen yea s, whe e he DRL was
associa ed wi h se e al echniques o obus he aining and
decision-making abili y while a ge ing objec loca ion. The
agen mus es ima e he a ge posi ion (bounding box) in
e e y sequence ame in he mos ypical use o DRL on
isual objec acking by epea edly selec ing ul ima e i ing
ac ions o ge accu a e acking esul s.
Acco dingly, he s a e ep esen a ion is he ul illmen s a-
us o he gene al ame s a es wi hin a a ge ed bounding box.
In gene al, ac ions a e he ans o ma ion esul o he bound-
ing box while acking ha can shi , scale, and u n ac ions
depending on how he ne wo k lea ned and adap ed o he
en i onmen in aining ime. In DRL-based objec acking,
accu acy (p ecision) is emphasized as a ewa d alue, show-
ing he di e ence be ween he a ge ed ac ion bounding box
and g ound u h alues. In gene al, i is called in e sec ion-
o e -union (IoU). Rewa d alues will change acco ding o he
ac ion alue di e ence wi h g ound u h ou pu , which shows
acking accu acy.
B. VIRTUAL SIMULATION-BASED OBJECT TRACKING VIA
DEEP REINFORCEMENT LEARNING
In he las ew yea s, mos esea ch opics ha e in e ac ed wi h
inno a i e ends in i ual simula ion wo ld en i onmen s
ha allow he simula ion o any ac ion, objec , o p ocess,
enabling expe imen a ion wi h complex condi ions o manage
and op imize esul s. Algo i hm in eg a ion wi h simula ion
pla o ms makes i challenging o conduc es ing and expe -
imen a ion while aking ad an age o simula ion beha io
closely ela ed o eal-wo ld models wi h dynamic and inac-
i e ac ion modes. Se e al simula ion pla o ms ha e lexible
unc ionali y o connec wi h so wa e algo i hms o expe -
imen a ion. The mos widely used open-sou ce pla o ms
cu en ly a e Ai Sim [14] and Uni y [15], which in end o
b idge he gap be ween he i ual and eal wo lds o sup-
po he de elopmen o au onomous con ol and a ealis ic
eplica o he ac ual wo ld. Bo h pla o ms a e ad ancing
hei echnical abili ies wi h high in ensi y o posi i ely in lu-
ence he de elopmen and es ing o da a-d i en machine
in elligence echniques such as ein o cemen lea ning and
deep lea ning. W. Luo e al. sugges ed an ac i e objec
acking echnique [29] ia deep ein o cemen lea ning,
in which a d one agen adop ed he Con Ne -LSTM unc-
ion app oxima o o p edic ing he a ge mo emen using
a ame- o-ac ion s a egy. Besides, hey pe o m addi ional
(ViZDoom and Un eal Engine simula ion) en i onmen aug-
men a ion echniques and a cus omized ewa d unc ion o
boos he aining p ocess o achie e be e a ge acking
pe o mance. Ano he i ual simula ion-based app oach [36]
uses a monocula onboa d came a ia a DRL model o ollow
he de ec ed a ge objec . They s a e ha his echnique is
a mo e accu a e and cos -e icien s a egy o adop ing an
algo i hm in a i ual en i onmen by using mul iple senso
da a poin s om he p e-calcula ed ajec o y. The p oposed
model combines one o he objec de ec ion models called
MobileNe [37] o ge he bounding box in o ma ion om
he image inpu o he lea ning p ocess. The model includes
con e gence-based explo a ion and exploi a ion o adap-
i ely aligning algo i hms wi h he ne wo k.
Mo eo e , J. Schulman e al. sugges ed a ein o cemen
lea ning-based d one ollow-me beha io objec acking
amewo k [38] using he Deep Q-Lea ning (DQN) model
o con ol RL agen s wi h adap i e and lexible beha io .
In his objec - acking model, s acked image ames and he
inclusion o dep h in o ma ion o in eg a ed as inpu ames
o he lea ning and es ing p ocess. The p oposed model
has expe imen ed wi h he di e en le el en i onmen s wi h
se e al s uc u al changes easonably. Expe imen al ou pu
wi h se e al speci ic condi ions showed ha he RL-based
d one ollowing echnique succeeded in i s adap i e and gen-
e alizing beha io .
In ou ecen esea ch, we p oposed i ual simula ion-
based isual objec acking ia a deep ein o cemen lea ning
algo i hm [25], which he Ai Sim d one agen uses o ack
he a ge ed objec class in a un ime i ual simula ion
en i onmen by u ilizing sequen ial ames di ec ly om i .
Addi ionally, he sugges ed model has been es ed wi h a
public da ase o e alua e he pe o mance o ecen esea ch
ou pu s. The main ad an age o a i ual simula ion pla o m
is ha esea che s can conduc expe imen a ion se e al imes
VOLUME 11, 2023 124131
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
wi h di e en ine- uning echniques a no cos un il hey
imp o e hei p oposal wi h high accu acy. Acco dingly, gen-
e a ing new, ake, o augmen ed da a o collec ing o eusing
da a om public se s o ein o ce model lea ning and boos
localiza ion exac ness while dec easing es ima ion ime and
human e o is unnecessa y.
III. PROPOSED METHOD
In his echnique, we c ea ed an algo i hm implemen a-
ion ha includes se e al componen s o he objec acking
amewo k, including aining and acking, and we e alua ed
i in a i ual simula ion en i onmen . The amewo k’s un-
damen al idea is o lea n he ac ion space using di ec inpu
om a i ual simula ion pla o m. Pla o ms allow aining
he ne wo k wi h li le e o spen ac i ely lea ning and
acking ope a ions. Howe e , he me hod mus be co ec ly
linked wi h he Q-lea ning ne wo k model o ge he equi ed
objec ea u e in o ma ion o s udy he en i onmen and aid
in making con inuous ac ion decisions in each ame o he
acking sequence.
A. ALGORITHM BASELINE
Figu e 1shows he baseline o ou p oposed me hod illus a-
ion. Fi s ly, he Ai Sim simula ion pla o m mus be ins alled
and se wi h he equi ed cha ac e is ic pa ame e s o in e-
g a e he designed algo i hm model. We manually inse an
objec in o he simula ion pla o m wi h a de ined walking
ou e a ound he pa icula loca ion speci ied in he i ual
en i onmen pa o he pipeline (Figu e 1). The i ual simu-
la ion pla o m p o ides essen ial inpu ame sequences wi h
ea u e in o ma ion, such as o dina y, segmen ed, and g ay-
scale (nega i e) dep h images, ha could p oceed h ough
-agen DRL ne wo k laye s o lea n and ake ac ion o a ge
acking measu es. We use image dep h o iden i y objec
loca ion and a ge ed class while expe imen ing h ough he
ne wo k o adap a ion o unknown condi ions.
B. TF-AGENT-BASED DRL OBJECT TRACKING MODEL
T acking objec s on a i ual simula ion pla o m di -
e s om ypical s a e-o - he-a a ge - acking amewo k
app oaches. The a ge objec mo es au oma ically ac oss
he simula ion pla o m a ea, occluded by obs acles such as
high walls, se e al di e en -shaped objec s, e c. In his p o-
posal, a andom walking pedes ian was se in o a simula ion
en i onmen o c ea e lea ning and acking condi ions by a
i ual Ai Sim d one agen . As shown in Figu e 2, we in e-
g a ed he simula ion pla o m wi h he sugges ed al e na i e
algo i hm model o join ly op imize ep esen a i es by expe -
imen ing in di e en condi ions. Fi s ly, we eques he
en i onmen simula ion pla o m o ge he ypical dep h
images and he segmen a ion map o ge he pixels wi h
a a ge . In he nex s ep, ames will be g ay-scaled and
no malized o u he ecogni ion o an objec in he i ual
simula ion model. A e ge ing he pixels wi h he a ge
and c ea ing hem, bounding box poin s a e conca ena ed and
ans o med in o ne wo k- eadable alues o he ollowing
p ocess.
1) DQN-BASED TF-AGENT
The DQN agen is sui able o any en i onmen al condi ion
possessed by a disc e e ac ion space o mula ed de e -
minis ically o simplici y and expec a ions o e s ochas ic
en i onmen al ansi ions. The main goal o he DQN agen
in his model is o ain a policy o maximize he discoun ed
cumula i e ewa d (1).
R 0=X∞
= 0
γ − 0 (1)
Tha is also known as he e u n alue R 0. In mos RL-based
ne wo ks, he discoun ac o γshould be a cons an alue
highligh ing he sum o con e ges be ween 0 and 1. I allows
ou agen o gain be e ewa d alue esul s by a oiding
unce ain en i onmen ea u e in o ma ion and iden i ying
which is less ele an han a ai ly con iden one. The Q∗is
o achie e an a o dable ewa d o e u n alue emphasized:
Q∗:S a e ×Ac ion →Rwhen he ac ion is aken in a
gi en s a e, he e u n esul s om a cons uc ed policy o
achie e maximized ewa ds (2).
π∗(s)=a gmin
a
Q∗(s,a)(2)
In ou i ual eali y wo ld simula ion model, we will ha e
access o s a e and ac ion space in o ma ion ela ed o he Q∗
alue unc ion o c ea e and ain he Q-ne wo k. Mos o he
Q unc ions in he case o policy- equi ed condi ions obey
Bellman’s [39] equa ion (3).
Qπ(s,a)=R+γQπ(s′, π(s′)) (3)
The di e ence be ween ini ial and lea ned alues calcula-
ions ollowing he equali y equa ion, also known as Q- alue
upda ing [39] o he empo al di e ence e o (4).
δ=Qnew(s,a)=(1 −α)Q(s,a)
| {z }
old alue
+α
lea ned alue
z }| {
R +1+γmax
a′Qs′,a′(4)
Equa ion (4) abo e calcula es he upda ing Q- alue o he
s a e-ac ion (s,a)pai a ime s ep . I is assumed o be equal
o a weigh ed sum o old and lea ned alues, whe e he ini ial
old alue would be 0 since he agen is expe iencing his pa -
icula s a e-ac ion pai alue. The old alue is mul iplied by
(1−α). αlea ning a e is deno ed and se as α=0.001 o
ou de aul aining ne wo k. Ins ead o o e w i ing he newly
calcula ed Q- alue, he αlea ning a e is se o de e mine
he p e iously compu ed Q- alue amoun o he ini ial s a e-
ac ion pai . To e ain he ecen ly ob ained Q- alue la e ,
we gi e a highe lea ning a e o he equal s a e-ac ion ma ch
o adop he d one agen quickly o he compu ed Q- alue.
Howe e , i should be a a balanced lea ning a e o keep
he ade-o be ween he p e ious and new Q- alues o he
124132 VOLUME 11, 2023
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
FIGURE 1. P oposed DRL-based TF-agen objec acking baseline in eg a ion wi h he Game Engine.
FIGURE 2. The low cha o he DRL-based -agen objec acking model in Game Engine.
FIGURE 3. The p ocess o Q-lea ning upda e.
u he aining p ocess. A lea ned alue is a ewa d R +1 ha
he d one agen ecei es mo ing andomly om he s a ing
s a e poin , plus discoun ed es ima ion γo op imal u u e
Q- alue o a new s a e-ac ion ma ch (s′,a′n 1- ime s eps.
The ou pu o he lea ned alue mul iplica ion by he lea ning
a e αis done o ge he op imal policy alue upda e. The
Q-lea ning p ocess upda e illus a ion is in Figu e 3.
As illus a ed in Figu e 3, he e can be se e al ac ions
h ough he aining o lea ning p ocess whe e an agen
chooses he seemingly op imal ac ions Qπ(s ,a )and
ecei es a ewa d o he agen ’s pe o mance h ough s eps in
a i ual en i onmen . Fo u he lea ning, he agen should
choose an ac ion om he S +1s a e o con inuously lea n and
analyze he en i onmen wi h mo e p o ound ea u e esul s.
He e, he g eedy epsilon op ion is a s aigh o wa d s a -
egy o balancing explo a ion and exploi a ion by andomly
selec ing be ween he wo. The me hod, whe e epsilon is
he likelihood o selec ing o explo e o exploi , de e mines
whe he i p oceeds o explo e he en i onmen wi h a sligh
chance. We can see he ac ion selec ion wi h he epsilon
g eedy me hod ma hema ical o mula ion below he equa ion
(5).
Ac ion a ime ( ) a =(max
aQ (a)1−ϵ
any ac ion (a)ϵ(5)
The ac ion selec ion me hod o u he lea ning can be
de ailed ho oughly when he alue is 1−ϵ, and he agen
uses exploi a ion o ake ad an age o p io knowledge, which
is a bes -es ima ed ewa d; o he wise, ϵi akes explo a ion o
look o new op imal op ions.
The alue o each ac ion mus be speci ied o ou agen
o choose he one ha will esul in he bes ewa d. The
ac ion- alue es ima ion unc ion (6) uses p obabili y heo y
o de ine hese alues. The p edic ed ewa d ecei ed when
VOLUME 11, 2023 124133
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
choosing an ac ion om a lis o all po en ial ac ions e e
o he ‘‘ alue o ha ac ion’’. So, we u ilize he ‘‘sample-
a e age’’ app oach o es ima e he alue o doing an ac ion
since he agen does no know he alue o choosing a pa ic-
ula ac ion.
Q (a)
=Sum o ewa ds when ac ion (a) aken be o e ime ( )
Numbe o imes ac ion (a) aken be o e ime ( )
=P −1
i=1Ri
−1(6)
The agen will hen selec he ac ion wi h he mos ou s and-
ing es ima ed alue, e e ed o as a g eedy ac ion, once he
alue Q(s′,a′)has i s pick a e.
2) PPO-BASED TF-AGENT
P oximal Policy Op imiza ion (PPO) is a s aigh o wa d pol-
icy g adien app oach o RL-based op imiza ion p oblems
ha al e na es be ween op imizing a ‘‘su oga e’’ objec i e
unc ion using s ochas ic g adien ascen and sampling da a
h ough in e ac ion wi h he en i onmen [30]. The main idea
and di e ence be ween p ima y and no el policy g adien
me hods is ha he mul iple-upda e miniba ch objec i e unc-
ion is applied. A he same ime, he s anda d model upda es
he g adien pe da a sample in a single epoch.
A policy g adien echnique known as he P oximal Pol-
icy Op imiza ion (PPO) algo i hm is applied o imp o e he
policy o a ein o cemen lea ning agen . PPO is a se o
algo i hms ha includes PPO1 and PPO2. In his p oposal,
we will ocus on he PPO1 algo i hm. The Clipped Su oga e
Goal is a d op-in subs i u e o he policy g adien objec i e
o inc ease aining s abili y by es ic ing he policy change
a each s ep. To add ess hese and o he di icul ies, we may
limi he amoun , al e he policy, and ensu e i cons an ly
imp o es. Fu he mo e, implemen ing his model helps o
in eg a e wi h a comple e p ocessing algo i hm o achie e
e icien samples om inpu images and minimize hype pa-
ame e uning indica o s. I achie es he same pe o mance
imp o emen s while a oiding complexi y by op imizing he
basics o he Clipped Su oga e Objec i e (7).
LCLIP
(θ)=ˆ
E hmin( (θ)ˆ
A ,clip( (θ),1−ϵ,1+ϵ)ˆ
A )i
(7)
(θ)ˆ
A – iden i ies he same objec i e be o e, inside he
minimiza ion; clip( (θ),1−ϵ,1+ϵ)ˆ
A – his pa o
o mula ion is he same objec i e, bu (θ)is clipped be ween
(1−ϵ,1+ϵ);The comple e min( (θ)ˆ
A ,clip( (θ),
1−ϵ, 1+ϵ)ˆ
A )– episode shows he min o he same objec i e
om be o e and he clipped one;
The main objec i e o clipping su oga es is a egion clip-
ping p ocess ha p e en s he algo i hm om ge ing oo
g eedy and ying o upda e oo much a once while aining
and lea ning o lea e he egion wi h good samples o es i-
ma ion and summa izing. PPO enables us o conduc many
FIGURE 4. P obabili y a io o he su oga e unc ion LCLIP wi h posi i e
A>0 and nega i e A<0 ad an ages. The ed ci cle on each plo shows
he s a ing poin o he op imiza ion, i.e., =1. The sum o he
su oga e unc ion LCLIP can be pe o med o many e ms [38].
g adien ascen epochs on ou da a s eam wi hou igge ing
ha m ully massi e policy modi ica ions. Conduc ing hese
p ocesses helps ge mo e ou o collec ed o s eamed da a
while dec easing sample ine iciency. Mo eo e , he unning
PPO policy uses N pa allel ac o s ha indi idually collec
da a. The da a is collec ed in o mini-ba ches and hen ained
o K epochs using he Clipped Su oga e Objec i e unc ion.
The Clipped Su oga e Objec i e will a ec and op imize
e e y ac ion he agen akes. Upda ing should be s opped i
he ac ion is be e (posi i e) A>0and mo e p obable while
aking he end o g adien s eps. O he wise, when he d one
ac ion is di ec ed in he w ong di ec ion bu he ac ion is
good, i can be edi ec ed o undone om he ini ial s a e.
In he case o lousy ac ion (nega i e) A<0and less p obable
ou comes gained om agen s’ ac ions, an agen needs o ake
sho s eps, o hey do no need o go a s eps in ac ion space.
When i comes o he no malized le el o upda ing, i can be
con olled in he balanced a ea o ge he op imal p obabili y
a io o agen s. The illus a ion o he LCLIP su oga e
unc ion p obabili y can be seen below in a summa izing
Figu e 4.
Figu e 4, illus a es clipped su oga e objec i e unc ions
op imiza ion pa ame e s o he unning lea ning pe iod
wi h p obabili y a io. E ec i ely, his echnique can be
encou aged using signi ican policy changes ac oss lea ning
en i onmen s o inpu da a o be e p obabili y op imiza ion
wi h a model agen . PPO wi h clipped objec echnique shows
he di e ence be ween wo main ained policy ne wo ks, he
cu en πθ(a |s )and he las used policy πθk(a |s )applied
o collec samples. A new policy e alua ion comes om
necessa y sampling, which in ol es collec ing old policy
samples o imp o e e iciency.
The inal loss unc ion o he PPO ac o -c i ic s yle looks
below equa ion (8), a combina ion o he Clipped Su oga e
Objec i e unc ion, Value Loss Func ion, and En opy bonus.
LCLIP+VF+S
(θ)=ˆ
E hLCLIP
(θ)−c1LVF
(θ)i
+c2S[πθ](s )(8)
The gi en equa ion abo e includes se e al pa s ha can
be complex o unde s and, ye i gi es mo e p io i y o
124134 VOLUME 11, 2023
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
achie ing mo e accu a e esul s while applying hem o he
expe imen a ion p ocess. The explained i s pa is a Clipped
Su oga e Objec i e unc ion gi en in he (7) equa ion. In (8),
gi en abo e, c1and c2a e coe icien s o he alue o he
ela ed pa ame e o calcula ion. LVF
(θ)iden i ies as a
squa ed-e o alue loss: (Vθ(s )−V a g
)2. To ensu e he
su icien explo a ion o unknown and complex scena ios in
i ual en i onmen s added an en y as a bonus S[πθ](s ).
IV. EXPERIMENTAL RESULTS
One o he main objec i es and ocuses was o ge he
mos ad an age om he simula ion pla o m o pe o m
expe imen s in di e en condi ions wi h pa ame ic changes.
Many ela ed esea che s used se e al simula ion pla o ms
o es and e alua e hei algo i hms in se e al e alua ion
s udies. The e a e di e en me hods o ge an ad an age
om he ealis ic i ual pla o m. In mos cases, pla o ms
apply o expe imen ing pu poses only. Howe e , i can
also be p io i ized widely in lea ning, aining, and es ing.
The cu en de elopmen o simula ion pla o ms like Uni y,
Un eal Engine, and Cecium gi es g ea oppo uni ies and
ad an ages o p ocess and expe imen wi h s a e-o - he-a
models in mul iple and imp ac ical ci cums ances. One o
he p ime ea u es o he simula o s is he in e connec ion
be ween p og amming languages (Ja asc ip , Py hon, Go,
Ja a, Ko lin, PHP, C#, Swi , e c.) and amewo ks (Angula ,
jQue y, Reac , Ruby, and Rails, Vue, ASP.NET Co e, Django,
Exp ess, e c.). Howe e , building o se ing up his ype o
a chi ec u e and amewo k is qui e icky, and i could only
be success ul in some cases due o hi d-pa y p og ams’
and lib a ies’ con lic s and disp opo ionali y. In his esea ch
wo k, we conduc ed expe imen s wi h di e en pa ame ic
changes and ine uning, as explained in he ollowing chap e
sessions.
A. TRAINING RESULTS
We ha e ained ou p oposed model wi h a simple Bloks
en i onmen by inse ing andomly mo ing objec s o lea n
en i onmen al space and o c ea e a model o u u e es ing
and e alua ion pu poses. We applied wo ypes o -agen
models, DQN and PPO-based -agen s, o achie e mo e
compa able ou pu esul s wi h a 0.001 lea ning a e con ig-
u a ion.
Figu e 5abo e illus a es he minimum ewa d ou pu s o
he ained models in a ypical Blocks en i onmen , whe e
he DQN-based -agen and he PPO-based -agen model
a e ma ked wi h blue and pink, espec i ely. The minimum
ewa d is he smalles alue ha he agen can ecei e as a
ewa d du ing he aining p ocess. The minimum ewa d is
ypically nega i e since mos p oblems in ol e a penal y o
making subop imal decisions – he aining epoch and ewa d
a 2000 and 50, espec i ely. Fu he mo e, he backg ound
was se wi h a plo ab colo in each me hod o show he
o e all pe o mance o he aining agen s. Each agen model
ini ially gained di e en ewa ds, whe eas he DQN-based
FIGURE 5. The minimum ecei ed ewa d ou pu o wo aining models:
DQN-based TF-AGENT and PPO-based TF-AGENT.
FIGURE 6. The ecei ed maximum ewa d ou pu o wo ypes o aining
models: DQN-based TF-AGENT and PPO-based TF-AGENT.
FIGURE 7. The ecei ed a e age ewa ds ou come o aining o DQN
and PPO-based TF-AGENTS.
agen pe o med be e . Ne e heless, a he end o he ain-
ing epochs, he PPO-based agen ecei es be e esul s han
he DQN-based model agen . The whole aining ewa d pe -
o mance illus a ion in Figu e 6abo e, se o 2000 and 50,
aining epoch and ewa d, espec i ely. The maximum poin
is he highes alue ha he agen can ecei e in he aining
p ocess. The maximum ewa d is ypically a posi i e alue
since mos p oblems in ol e a ewa d o making op imal
decisions. The DQN-based TF-AGENT model ini ially gains
a highe ewa d alue in his g aph. Howe e , he PPO-based
model pe o ms be e a e 400 epochs un il he end o he
aining s eps. Unde s anding he ange o possible ewa ds
can help se he hype pa ame e s o he models, such as he
lea ning a e o he discoun ac o . I can also help assess he
pe o mance o he ained agen , as he ewa ds ob ained by
VOLUME 11, 2023 124135
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
FIGURE 8. The es ing ewa d dis ibu ion o he DQN-based TF-AGENT (a) and he PPO-based TF-AGENT (b) in 50 s eps o he episode.
he agen a e compa ed agains he minimum and maximum
possible alues.
The gi en a e age ewa d below e e s o he mean alue o
he ewa ds ecei ed by agen s du ing hei in e ac ions wi h
he en i onmen while aining using DQN and PPO-based
model algo i hms. The a e age ewa d is essen ial o e alua -
ing he agen s’ pe o mance du ing aining. Du ing aining,
agen s y o lea n an op imal policy ha maximizes he cumu-
la i e ewa d ob ained o e ime. Calcula ing he a e age
e alua ion ewa d is done by di iding he sum o he ewa ds
ecei ed du ing all episodes by di iding i by he o al numbe
o episodes. The es ima ed calcula ion is he a e age ewa d
he agen will ecei e when in e ac ing wi h he en i onmen
using he lea ned policy.
By e alua ing he a e age ewa d ecei ed, we can see he
di e ence be ween he DQN and PPO-based model’s pe o -
mance in a ied con igu a ions. Howe e , in some scena ios,
he a e age ewa d may no be he mos sui able me ic o
e alua ing he agen ’s pe o mance.
B. TESTING RESULTS
We ha e es ed ou p oposed DQN and PPO-based model
agen s wi h he same en i onmen al condi ion bu di e en
unseen es episodes o explo e he abili y o he models and
compa e hei pe o mance. As men ioned abo e, he DRL-
based algo i hm’s pe o mance e alua ion di e s om o he
s a e-o - he-a algo i hms in he case o pe o mance me -
ics e alua ions and compa ison echniques. The agen -based
models’ p ecision can be seen o aken as a ecei ed ewa d
alue. As men ioned ea lie , he a e age ewa d ob ained by
he agen du ing aining can help c ea e a model and apply
124136 VOLUME 11, 2023
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
his model o he es ing p ocess as a pe o mance me ic. This
me ic measu es he agen ’s abili y o na iga e he en i on-
men and ob ain expec ed ou pu acking. The diag am below
(Figu e 8) ep esen s he DQN-based TF-AGENT model’s
ou pu wi h he ewa d pe cen age ecei ed om unseen
es ing scena ios. In he es ing session, he ecei ed ewa d
pe cen age was se o 100 in he 50 s eps espec i ely in
e e y episode. The o e all ecei ed in e e y s ep ma ked
wi h a column and ed line illus a es he smoo hed alue o
he DQN-based TF-AGENT es ing esul s ained and es ed
wi h s anda d ewa d in Figu e 8 (a).
Figu e 8 (b) shows he PPO-based TF-AGENT’s ecei ed
pe cen age ewa d es ing esul s in he 50 s eps o he
episode, along wi h smoo hed ed line ou pu . The di e en
esul s be ween he DQN and PPO models ecei ed ewa ds
in e e y aining s ep. As we can see, he es ing esul s show
ha bo h models gi e high accu acy and p ecise lea ning
pe o mance in e e y es ing s ep ou pu wi h an ele a ed
conclusion.
V. CONCLUSION
In his esea ch wo k, we ha e p esen ed a DQN and PPO-
based TF-AGENT model-based objec acking amewo k
in eg a ed wi h a simple Blocks en i onmen o e alua e he
pe o mance o he p oposed algo i hm. I has been in eg a ed
wi h he simula ion pla o m o highligh he algo i hm’s
o e all pe o mance.
The simula ion pla o m p o ides h ee ypes o essen ial
inpu images o expe imen wi h and e alua e he o e all
s a us. While es ing in a i ual- eali y scena io wi h i ual
d one agen s and ine uning o each he bes o desi ed
esul s, he p oduc i i y and eligibili y o hese pla o ms a e
i al. The DQN and PPO-based i ual -agen d ones lea n
how o de ec and ack an objec inse ed in his pla o m
by ob aining consecu i e ames om a p ima y Blocks en i-
onmen and using a DRL ne wo k o manage he ac ions,
s a es, and acking pipeline. Bo h -agen s a e ained in a
Blocks en i onmen o adap o he su oundings and exis ing
objec s in a simula ion condi ion o addi ional es ing, ack-
ing accu acy, and speed assessmen . In he aining p ocess,
bo h models showed p esen able esul s: minimum 49 (PPO)
and 48 (DQN) ewa ds in 2000 epochs; maximum 49 (DQN)
and 49 (PPO) ewa ds in 2000 epochs; a e age 49 ewa ds
we e ecei ed o bo h (PPO and DQN) models. Models pe -
o mance con as ed 50 s eps o one episode es ing se , whe e
he PPO-based -agen ge s i s pick alue ewa d o 97% in
s ep 23, DQN-based agen ecei es i s max alue o 86% in
he 17 h s ep espec i ely. Howe e , he o e all pe o mance
o he ecei ed pe cen age ewa d g aph (Figu e 8,aand b)
indica es ha he DQN-based model sequen pe o ms be e
han he PPO-based one. Rega ding s abili y, ewa d con-
ibu ion, and nume ic g aphical pe o mance, we examined
and compa ed he algo i hm echniques o a ious es ablished
hype pa ame ic changes wi h ein o cemen lea ning-based
ne wo k con ol inco po a ed in o he simula ion p ocess.
In u u e wo k, we a e going o in eg a e ou model wi h
se e al s a e-o - he-a acking echniques o imp o e he
pe o mance o he a ge acking amewo k by es ing i in
mo e complex i ual simula ion en i onmen s.
REFERENCES
[1] The Manu ac u e . (Jun. 29, 2022). The Bene i s o D ones in Manu ac-
u ing. [Online]. A ailable: h ps://www. hemanu ac u e .com/a icles/ he-
bene i s-o -d ones-in-manu ac u ing/
[2] C op acke . (Ap . 26, 2022). D one Technology in Ag icul u e. D ag-
on ly IT. [Online]. A ailable: h ps://www.c op acke .com/blog/d one-
echnology-in-ag icul u e.h ml
[3] G. McNeal. (No . 2014). D ones and Ae ial Su eillance:
Conside a ions o Legisla u es. [Online]. A ailable: h ps://www.
b ookings.edu/ esea ch/d ones-and-ae ial-su eillance-conside a ions-
o -legisla u es/
[4] B. Pu ahong, T. Anuwongpini , A. Juhong, I. Kanjanasu a , and
C. Pin a iooj, ‘‘Medical d one managing sys em o au oma ed ex e nal
de ib illa o deli e y se ice,’’ D ones, ol. 6, no. 4, p. 93, Ap . 2022, doi:
10.3390/d ones6040093.
[5] N. Tuśnio and W. W óblewski, ‘‘The e iciency o d ones usage o sa e y
and escue ope a ions in an open a ea: A case om Poland,’’ Sus ainabili y,
ol. 14, no. 1, p. 327, Dec. 2021, doi: 10.3390/su14010327.
[6] M. Elloumi, R. Dhaou, B. Esc ig, H. Idoudi, and L. A. Saidane,
‘‘Moni o ing oad a ic wi h a UAV-based sys em,’’ in P oc. IEEE
Wi eless Commun. Ne w. Con . (WCNC), Ap . 2018, pp. 1–6, doi:
10.1109/WCNC.2018.8377077.
[7] A. Chodo ek, R. R. Chodo ek, and A. Yas ebo , ‘‘Wea he sensing in
an u ban en i onmen wi h he use o a UAV and WebRTC-based pla -
o m: A pilo s udy,’’ Senso s, ol. 21, no. 21, p. 7113, Oc . 2021, doi:
10.3390/s21217113.
[8] K. Jewani, M. Ka a, D. Mo wani, and G. Je hwani, ‘‘Fi e igh e d one,’’ in
P oc. 1s In . Con . Ad . Sci. Inno . Sci., Eng., Technol. (ICASISET), 2020.
[9] C. Huang, C.-E. Lin, Z. Yang, Y. Kong, P. Chen, X. Yang, and K.-T. Cheng,
‘‘Lea ning o ilm om p o essional human mo ion ideos,’’ in P oc.
IEEE/CVF Con . Compu . Vis. Pa e n Recogni . (CVPR), Jun. 2019,
pp. 4239–4248, doi: 10.1109/CVPR.2019.00437.
[10] A. A. El-Sha ie, M. Zaki, and S. E. D. Habib, ‘‘Fas CNN-based objec
acking using localiza ion laye s and deep ea u es in e pola ion,’’ in P oc.
15 h In . Wi eless Commun. Mobile Compu . Con . (IWCMC), Jun. 2019,
pp. 1476–1481, doi: 10.1109/IWCMC.2019.8766466.
[11] R. Ra ind an, M. J. San o a, and M. M. Jamali, ‘‘Mul i-objec de ec-
ion and acking, based on DNN, o au onomous ehicles: A e iew,’’
IEEE Senso s J., ol. 21, no. 5, pp. 5668–5677, Ma . 2021, doi:
10.1109/JSEN.2020.3041615.
[12] X. Fa hodo , K.-S. Moon, S.-H. Lee, and K.-R. Kwon, ‘‘LSTM ne wo k
wi h acking associa ion o mul i-objec acking,’’ J. Ko ea Mul imedia
Soc., ol. 23, no. 10, pp. 1236–1249, Oc . 2020.
[13] D. Gözen and S. Oze , ‘‘Visual objec acking in d one images wi h deep
ein o cemen lea ning,’’ in P oc. 25 h In . Con . Pa e n Recogni . (ICPR),
Jan. 2021, pp. 10082–10089.
[14] S. Shah, D. Dey, C. Lo e , and A. Kapoo , ‘‘Ai Sim: High- ideli y
isual and physical simula ion o au onomous ehicles,’’ 2017,
a Xi :1705.05065.
[15] A. Juliani, V.-P. Be ges, E. Teng, A. Cohen, J. Ha pe , C. Elion, C. Goy,
Y. Gao, H. Hen y, M. Ma a , and D. Lange, ‘‘Uni y: A gene al pla o m
o in elligen agen s,’’ 2018, a Xi :1809.02627.
[16] G. Yang, ‘‘Asymp o ic acking wi h no el in eg al obus schemes o mis-
ma ched unce ain nonlinea sys ems,’’ In . J. Robus Nonlinea Con ol,
ol. 33, no. 3, pp. 1988–2002, Feb. 2023.
[17] G. Yang, T. Zhu, F. Yang, L. Cui, and H. Wang, ‘‘Ou pu eedback adap i e
RISE con ol o unce ain nonlinea sys ems,’’ Asian J. Con ol, ol. 25,
no. 1, pp. 433–442, Jan. 2023.
[18] M. Bha a ai and M. Ma ínez-Ramón, ‘‘A deep Q-lea ning based pa h
planning and na iga ion sys em o i e igh ing en i onmen s,’’ in P oc.
13 h In . Con . Agen s A i . In ell. (ICAART), 2021, pp. 267–277, doi:
10.5220/0010267102670277.
[19] A. P. Kalidas, C. J. Joshua, A. Q. Md, S. Bashee , S. Mohan, and S. Sak i,
‘‘Deep ein o cemen lea ning o ision-based na iga ion o UAVs in
a oiding s a iona y and mobile obs acles,’’ D ones, ol. 7, no. 4, p. 245,
Ap . 2023, doi: 10.3390/d ones7040245.
VOLUME 11, 2023 124137