scieee Science in your language
[en] (orig)

Deep reinforcement learning Tf-Agent-based object tracking with virtual autonomous drone in a game engine

Abstract

The recent development of object-tracking frameworks has affected the performance of many manufacturing and industrial services such as product delivery, autonomous driving systems, security systems, military, transportation and retailing industries, smart cities, healthcare systems, agriculture, etc. Achieving accurate results in physical environments and conditions remains quite challenging for the actual object-tracking. However, the process can be experimented with using simulation techniques or platforms to evaluate and check the model’s performance under different simulation conditions and weather changes. This paper presents one of the target tracking approaches based on the reinforcement learning technique integrated with TensorFlow-Agent (tf-agent) to accomplish the tracking process in the Unreal Game Engine simulation platform AirSim Blocks. The productivity of these platforms can be seen while experimenting in virtual-reality conditions with virtual drone agents and performing fine-tuning to achieve the best or desired performance. In this paper, the tf-agent drone learns how to track an object integration with a deep reinforcement learning process to control the actions, states, and tracking by receiving sequential frames from a simple Blocks environment. The tf-agent model is trained in the AirSim Blocks environment for adaptation to the environment and existing objects in a simulation environment for further testing and evaluation regarding the accuracy of tracking and speed. We tested and compared two approaches, DQN and PPO trackers, and reported results in terms of stability, rewards, and numerical performance.

Read accessible full text

Deep reinforcement learning Tf-Agent-based object tracking with virtual autonomous drone in a game engine

Author: Farkhodov, Khurshedjon
Publisher: IEEE
Year: 2023
DOI: 10.1109/ACCESS.2023.3325062
Source: https://dspace.vsb.cz/bitstreams/2dfc0053-7bad-4b3b-9c66-d4860948183b/download
Recei ed 9 Augus 2023, accep ed 6 Oc obe 2023, da e o publica ion 16 Oc obe 2023, da e o cu en e sion 9 No embe 2023.
Digi al Objec Iden i ie 10.1109/ACCESS.2023.3325062
Deep Rein o cemen Lea ning T -Agen -Based
Objec T acking Wi h Vi ual Au onomous D one
in a Game Engine
KHURSHEDJON FARKHODOV 1, SUK-HWAN LEE2,
JAN PLATOS 3, (Membe , IEEE), AND KI-RYONG KWON 1
1Depa men o AI Con e gence, Pukyong Na ional Uni e si y, Busan 48513, Sou h Ko ea
2Depa men o Compu e Enginee ing, Dong-A Uni e si y, Busan 49315, Sou h Ko ea
3Depa men o Elec ical Enginee ing and Compu e Science, VSB—Technical Uni e si y o Os a a, 708 00 Os a a, Czech Republic
Co esponding au ho : Ki-Ryong Kwon ([email p o ec ed])
This wo k was suppo ed in pa by he Minis y o Science and In o ma ion and Communica ion Technology (MSIT), Sou h Ko ea,
h ough he In o ma ion Technology Resea ch Cen e (ITRC) Suppo P og am, Supe ised by he Ins i u e o In o ma ion and
Communica ions Technology Planning and E alua ion (IITP), unde G an IITP-2023-2020-0-01797; and in pa by he MSIT h ough he
ICT Consilience C ea i e P og am, Supe ised by he IITP, unde G an IITP-2023-2016-0-00318.
ABSTRACT The ecen de elopmen o objec - acking amewo ks has a ec ed he pe o mance o
many manu ac u ing and indus ial se ices such as p oduc deli e y, au onomous d i ing sys ems, secu i y
sys ems, mili a y, anspo a ion and e ailing indus ies, sma ci ies, heal hca e sys ems, ag icul u e, e c.
Achie ing accu a e esul s in physical en i onmen s and condi ions emains qui e challenging o he ac ual
objec - acking. Howe e , he p ocess can be expe imen ed wi h using simula ion echniques o pla o ms
o e alua e and check he model’s pe o mance unde di e en simula ion condi ions and wea he changes.
This pape p esen s one o he a ge acking app oaches based on he ein o cemen lea ning echnique
in eg a ed wi h Tenso Flow-Agen ( -agen ) o accomplish he acking p ocess in he Un eal Game Engine
simula ion pla o m Ai Sim Blocks. The p oduc i i y o hese pla o ms can be seen while expe imen ing
in i ual- eali y condi ions wi h i ual d one agen s and pe o ming ine- uning o achie e he bes o
desi ed pe o mance. In his pape , he -agen d one lea ns how o ack an objec in eg a ion wi h a deep
ein o cemen lea ning p ocess o con ol he ac ions, s a es, and acking by ecei ing sequen ial ames
om a simple Blocks en i onmen . The -agen model is ained in he Ai Sim Blocks en i onmen o
adap a ion o he en i onmen and exis ing objec s in a simula ion en i onmen o u he es ing and
e alua ion ega ding he accu acy o acking and speed. We es ed and compa ed wo app oaches, DQN
and PPO acke s, and epo ed esul s in e ms o s abili y, ewa ds, and nume ical pe o mance.
INDEX TERMS Objec acking, objec de ec ion, ein o cemen lea ning, Ai Sim, i ual en i onmen ,
i ual simula ion, -agen , un eal game engine.
I. INTRODUCTION
Unmanned Ae ial Vehicles (UAV) a e widely u ilized in
e e y ield o manu ac u ing and daily se ice p ocedu es [1],
whe e d ones help o make he p ocess easie by au oma -
ing and dec easing he consump ion ime wi h sa e se ice
ac i i ies. Fo example, in ecen yea s, d ones ha e become
a op p io i y o echnical assis ance in he ag icul u al ield
The associa e edi o coo dina ing he e iew o his manusc ip and
app o ing i o publica ion was Gangyi Jiang.
o he medica ion and e iliza ion p ocesses. Addi ionally,
d ones can obse e cul i a ed ield mapping p ope ly [2]
wi h impo an in o ma ion abou any i egula changes in
he g owe ’s ield. In mos cases, i can be bene icial o
inc easing he p oduc i i y o la ge plan a ions wi h congen-
e ous g owe s o de e mine d ainage pa e ns and we -d y
spo s o ield ele a ion ha allow mo e e icien wa e ing
echniques. Mo eo e , he e a e se e al ields in which d ones
a e being applied such as su eillance and con ol o la ge
un eachable a eas [3], medical pu poses such as Au oma ed
VOLUME 11, 2023
2023 The Au ho s. This wo k is licensed unde a C ea i e Commons A ibu ion-NonComme cial-NoDe i a i es 4.0 License.
Fo mo e in o ma ion, see h ps://c ea i ecommons.o g/licenses/by-nc-nd/4.0/ 124129
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
Eme gency D ones (AED) [4], sea ch and escue ope a ions
[5], oad a ic moni o ing applica ions [6], wea he sensing
in an u ban en i onmen applica ion [7], eme gency sys em
applica ions o i e igh e p ocesses [8], and en e ainmen
(pho og aphy and cinema og aphy) [9]. Many ecen inno a-
ions and echnologies ela ed o UAV amewo ks include
su eillance and a ge acking uni s o se e al easons,
including secu i y main enance, con ol, mili a y assis ance,
and a ic moni o ing, besides manu ac u ing op imiza ion
and con ol. The cu en de elopmen o inno a i e ech-
nologies is aking no el ideas om modeling sys ems such
as i ual eali y en i onmen simula o s. C ea ing a i ual
e sion o he physical objec s and p ocess simula ions can
p o ide p ope se ice op imiza ion and allow o ee expe -
imen a ion wi h condi ions o any ac i i y.
A. IMPORTANCE AND DEVELOPMENT OF OBJECT
TRACKING WITH MACHINE LEARNING
In ecen yea s, mos o he UAV-based esea ch communi y
has paid a en ion o imp o ing he pe o mance o isual
acking echniques wi h se e al neu al ne wo ks ela ed o
ne wo k-based a chi ec u es, such as CNN [10], DNN [11],
LSTM [12], RL [13], and o he s. Howe e , hese esea ch
wo ks p esen aining and es ing esul s in physical en i on-
men da ase s, such as image collec ions aken om di e en
came as in public a eas, and image/ ideo se s aken by d one
came as. These objec - acking applica ions wo k wi h ce -
ain objec classes whe e he d one acks dynamic objec s in
an exac pa hway and localizes objec s wi h speci ic me hods
as an addi ional ask o he a ge - acking amewo k. S ill,
he e a e challenges while acking mo ing and s a ic objec s
wi h appa en ly iden ical aspec s o ecognize which one is an
ac ual acking a ge . Recen objec - acking esea ch de el-
opmen has begun wi h in ensi e lea ning wi h i ual eali y
in eg a ion pla o ms [14],[15], so scien is s can in eg a e
hei echnique o algo i hm wi h a i ual eali y pla o m o
es hei p oposals wi h a ious ine- uning pa ame e s and
condi ions. I gi es schola s mo e oppo uni ies o explo e
hei me hods be e and mo e deeply in a ha dwa e- ee en i-
onmen a ze o cos and op imize hem as much as possible.
Addi ionally, he e a e some echniques mo i a ed by obus
ex ension o in eg al schemes o misma ched unce ain non-
linea sys ems p oposed o suppo asymp o ic acking [16].
Asymp o ic acking means ensu ing ha he sys ems’ ou pu
acks a desi ed e e ence ajec o y o e ime wi h negli-
gible acking e o . The main goal is o design a acking
con ol sys em ha gua an ees he ou pu con e ges o he
e e ence ajec o y as ime app oaches in ini y in unce ain
en i onmen s. Ano he model is he ou pu eedback adap-
i e ise con ol echnique [17] used o unce ain nonlinea
sys ems o achie e accu a e acking o desi ed ajec o ies.
The e m ‘‘adap i e’’ indica es ha he con olle pa ame e s
a e upda ed online based on he sys em’s beha io and he
acking e o . A deep Q-lea ning-based [18] app oach has
been sugges ed o i e igh ing si ua ions, which is ob ainable
in some agen obo s o d ones o inding o planning pa hs
and na iga ing h ough i e en i onmen s. In such complex
and haza dous cases, i is equi ed o be mo e ca e ul o con-
ol he si ua ion wi h conc e e plans and ac ions o escuing
inju ed o ic ims o he inciden by coo dina ing si ua ional
awa eness wi h o he escue s, which is an u gen ask. When
he amewo k is ins alled and applied o eal d ones o obo s,
i can ensu e i e igh e s o escue s make he igh decision
in ex eme, panicky, and diso ien ing condi ions.
B. DESCRIPTION OF THE VR-BASED WORK AND MAIN
CONTRIBUTION
Se e al ele an esea ch s udies ha in eg a e i ual simula-
ion pla o ms wi h p oposal algo i hms ha e been published.
Kalidas e al. [19] p esen ed ision-based na iga ion o UAVs
based simply on image da a by employing deep ein o ce-
men lea ning o a oid s a iona y and mo able obs acles
au onomously in disc e e and con inuous ac ion space. W.
Zhao e al. [20] also p oposed a pe cep ion-based hie a chi-
cal ac i e acking con ol o UAVs deploying a high-le el
con olle and ac ion o de s in a V-REP-based en i onmen .
A ained PPO algo i hm [21] wi h ewa d shaping o ai c a
di ec ion o a mo ing des ina ion in a h ee-dimensional con-
inuous space model was sugges ed, wi h he agen -speci ic
a ge guidance in i ual s a e space using a no el ewa d
calcula ion. Using a PPO-based DRL algo i hm [22] was sug-
ges ed o UAV acking wi h he assis ance o ano he UAV,
in oducing he gene alized dis ibu ed deep ein o cemen
lea ning pla o m, which p o ides solu ions o o e come a i-
ous p oblems such as acking, con olling, and mission coo -
dina ion o UAVs. Mo eo e , M. A. B. Abdelkade e al. [23]
p opose RL-based d one ele a ion con ol on a Py hon-Uni y
in eg a ed simula ion amewo k o achie e a s able use
diag am p o ocol (UDP) wi h he sugges ed algo i hm. Çe in
[24] p oposes coun ing d ones in a 3D space wi h se e al
DRL me hods p esen o coun d ones wi h ano he d one in
he en i onmen p o ided by an Ai Sim simula o .
In his s udy, we de eloped an algo i hm based on -agen
d one acking in a Blocks en i onmen whe e he -agen
ac i ely makes decisions o ack he a ge objec in he
un ime en i onmen . The p oposed me hod includes di e -
en ewa d echniques o boos he lea ning, acking, and
decision-making p ocesses ia a TF-agen -based d one in a
simula ion pla o m. The e is some compu a ional conside -
a ion o co ec ly applying pa ame e alues o achie e a
highe accu acy a e, and s a e ep esen a ion was o mula ed
o clea ou unnecessa y losses and cons ain s o he aining
and es ing p ocesses.
The p ima y con ibu ions o his pape can be summa ized
as ollows:
•We in oduce a i ual en i onmen al simula ion-based
objec - acking algo i hm model ha ecei es inpu
images di ec ly om a ealis ic i ual pla o m.
•Di ec access o he ne wo k eedable sou ce images
om he simula ion en i onmen p o ides he amewo k
124130 VOLUME 11, 2023
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
he mo e ad an ages o lea ning and es ing when i
comes o unknown en i onmen al condi ions.
•The expe imen is implemen ed in an Ai Sim-based
basic Blocks en i onmen wi h a andom, pa icula
walking pe son being acked by a i ual d one agen .
•Two di e en me hods we e adop ed and in eg a ed wi h
he i ual simula ion pla o m o demons a e he pe -
o mance o he models.
II. RELATED WORKS
The ecen de elopmen o objec acking ia ein o ce-
men lea ning has imp o ed by in eg a ing i wi h many
a ge acking echniques, which p oduce be e pe o mance
wi h decision-making in acking p ocedu es. Al hough mos
isual acking concep s based on DRL could pe o m be e
in he case o he ep esen a ion model wi h adop ed manne s
o loca ing he a ge objec wi hin a sea ch egion, he inal
es ima ed a ge coo dina es a e ideally cen e ed.
A. OBJECT TRACKING VIA DEEP REINFORCEMENT
LEARNING
The ad ancemen o objec acking ia ein o cemen lea n-
ing is a compa a i ely no el idea, whe e objec localiza ion
and acking in eg a ion wi h a decision-making model [13],
[25],[26],[27],[28],[29],[30] a e applied o he lea ning and
acking p ocess as well. Se e al s udies ha e disco e ed ha
combining deep and RL [31] in a ious se ings con e s many
ad an ages. Visual objec acking [32], localizing empo al
ac i i y [33], iden i ying objec classes [26], objec ecogni-
ion h ough ideo sequence [34], and segmen a ion [35] a e
jus a ew o he compu e ision p oblems ha ha e used
DRL. No ably, isual objec acking ia DRL amewo k
s udies has inc eased in ecen yea s, whe e he DRL was
associa ed wi h se e al echniques o obus he aining and
decision-making abili y while a ge ing objec loca ion. The
agen mus es ima e he a ge posi ion (bounding box) in
e e y sequence ame in he mos ypical use o DRL on
isual objec acking by epea edly selec ing ul ima e i ing
ac ions o ge accu a e acking esul s.
Acco dingly, he s a e ep esen a ion is he ul illmen s a-
us o he gene al ame s a es wi hin a a ge ed bounding box.
In gene al, ac ions a e he ans o ma ion esul o he bound-
ing box while acking ha can shi , scale, and u n ac ions
depending on how he ne wo k lea ned and adap ed o he
en i onmen in aining ime. In DRL-based objec acking,
accu acy (p ecision) is emphasized as a ewa d alue, show-
ing he di e ence be ween he a ge ed ac ion bounding box
and g ound u h alues. In gene al, i is called in e sec ion-
o e -union (IoU). Rewa d alues will change acco ding o he
ac ion alue di e ence wi h g ound u h ou pu , which shows
acking accu acy.
B. VIRTUAL SIMULATION-BASED OBJECT TRACKING VIA
DEEP REINFORCEMENT LEARNING
In he las ew yea s, mos esea ch opics ha e in e ac ed wi h
inno a i e ends in i ual simula ion wo ld en i onmen s
ha allow he simula ion o any ac ion, objec , o p ocess,
enabling expe imen a ion wi h complex condi ions o manage
and op imize esul s. Algo i hm in eg a ion wi h simula ion
pla o ms makes i challenging o conduc es ing and expe -
imen a ion while aking ad an age o simula ion beha io
closely ela ed o eal-wo ld models wi h dynamic and inac-
i e ac ion modes. Se e al simula ion pla o ms ha e lexible
unc ionali y o connec wi h so wa e algo i hms o expe -
imen a ion. The mos widely used open-sou ce pla o ms
cu en ly a e Ai Sim [14] and Uni y [15], which in end o
b idge he gap be ween he i ual and eal wo lds o sup-
po he de elopmen o au onomous con ol and a ealis ic
eplica o he ac ual wo ld. Bo h pla o ms a e ad ancing
hei echnical abili ies wi h high in ensi y o posi i ely in lu-
ence he de elopmen and es ing o da a-d i en machine
in elligence echniques such as ein o cemen lea ning and
deep lea ning. W. Luo e al. sugges ed an ac i e objec
acking echnique [29] ia deep ein o cemen lea ning,
in which a d one agen adop ed he Con Ne -LSTM unc-
ion app oxima o o p edic ing he a ge mo emen using
a ame- o-ac ion s a egy. Besides, hey pe o m addi ional
(ViZDoom and Un eal Engine simula ion) en i onmen aug-
men a ion echniques and a cus omized ewa d unc ion o
boos he aining p ocess o achie e be e a ge acking
pe o mance. Ano he i ual simula ion-based app oach [36]
uses a monocula onboa d came a ia a DRL model o ollow
he de ec ed a ge objec . They s a e ha his echnique is
a mo e accu a e and cos -e icien s a egy o adop ing an
algo i hm in a i ual en i onmen by using mul iple senso
da a poin s om he p e-calcula ed ajec o y. The p oposed
model combines one o he objec de ec ion models called
MobileNe [37] o ge he bounding box in o ma ion om
he image inpu o he lea ning p ocess. The model includes
con e gence-based explo a ion and exploi a ion o adap-
i ely aligning algo i hms wi h he ne wo k.
Mo eo e , J. Schulman e al. sugges ed a ein o cemen
lea ning-based d one ollow-me beha io objec acking
amewo k [38] using he Deep Q-Lea ning (DQN) model
o con ol RL agen s wi h adap i e and lexible beha io .
In his objec - acking model, s acked image ames and he
inclusion o dep h in o ma ion o in eg a ed as inpu ames
o he lea ning and es ing p ocess. The p oposed model
has expe imen ed wi h he di e en le el en i onmen s wi h
se e al s uc u al changes easonably. Expe imen al ou pu
wi h se e al speci ic condi ions showed ha he RL-based
d one ollowing echnique succeeded in i s adap i e and gen-
e alizing beha io .
In ou ecen esea ch, we p oposed i ual simula ion-
based isual objec acking ia a deep ein o cemen lea ning
algo i hm [25], which he Ai Sim d one agen uses o ack
he a ge ed objec class in a un ime i ual simula ion
en i onmen by u ilizing sequen ial ames di ec ly om i .
Addi ionally, he sugges ed model has been es ed wi h a
public da ase o e alua e he pe o mance o ecen esea ch
ou pu s. The main ad an age o a i ual simula ion pla o m
is ha esea che s can conduc expe imen a ion se e al imes
VOLUME 11, 2023 124131
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
wi h di e en ine- uning echniques a no cos un il hey
imp o e hei p oposal wi h high accu acy. Acco dingly, gen-
e a ing new, ake, o augmen ed da a o collec ing o eusing
da a om public se s o ein o ce model lea ning and boos
localiza ion exac ness while dec easing es ima ion ime and
human e o is unnecessa y.
III. PROPOSED METHOD
In his echnique, we c ea ed an algo i hm implemen a-
ion ha includes se e al componen s o he objec acking
amewo k, including aining and acking, and we e alua ed
i in a i ual simula ion en i onmen . The amewo k’s un-
damen al idea is o lea n he ac ion space using di ec inpu
om a i ual simula ion pla o m. Pla o ms allow aining
he ne wo k wi h li le e o spen ac i ely lea ning and
acking ope a ions. Howe e , he me hod mus be co ec ly
linked wi h he Q-lea ning ne wo k model o ge he equi ed
objec ea u e in o ma ion o s udy he en i onmen and aid
in making con inuous ac ion decisions in each ame o he
acking sequence.
A. ALGORITHM BASELINE
Figu e 1shows he baseline o ou p oposed me hod illus a-
ion. Fi s ly, he Ai Sim simula ion pla o m mus be ins alled
and se wi h he equi ed cha ac e is ic pa ame e s o in e-
g a e he designed algo i hm model. We manually inse an
objec in o he simula ion pla o m wi h a de ined walking
ou e a ound he pa icula loca ion speci ied in he i ual
en i onmen pa o he pipeline (Figu e 1). The i ual simu-
la ion pla o m p o ides essen ial inpu ame sequences wi h
ea u e in o ma ion, such as o dina y, segmen ed, and g ay-
scale (nega i e) dep h images, ha could p oceed h ough
-agen DRL ne wo k laye s o lea n and ake ac ion o a ge
acking measu es. We use image dep h o iden i y objec
loca ion and a ge ed class while expe imen ing h ough he
ne wo k o adap a ion o unknown condi ions.
B. TF-AGENT-BASED DRL OBJECT TRACKING MODEL
T acking objec s on a i ual simula ion pla o m di -
e s om ypical s a e-o - he-a a ge - acking amewo k
app oaches. The a ge objec mo es au oma ically ac oss
he simula ion pla o m a ea, occluded by obs acles such as
high walls, se e al di e en -shaped objec s, e c. In his p o-
posal, a andom walking pedes ian was se in o a simula ion
en i onmen o c ea e lea ning and acking condi ions by a
i ual Ai Sim d one agen . As shown in Figu e 2, we in e-
g a ed he simula ion pla o m wi h he sugges ed al e na i e
algo i hm model o join ly op imize ep esen a i es by expe -
imen ing in di e en condi ions. Fi s ly, we eques he
en i onmen simula ion pla o m o ge he ypical dep h
images and he segmen a ion map o ge he pixels wi h
a a ge . In he nex s ep, ames will be g ay-scaled and
no malized o u he ecogni ion o an objec in he i ual
simula ion model. A e ge ing he pixels wi h he a ge
and c ea ing hem, bounding box poin s a e conca ena ed and
ans o med in o ne wo k- eadable alues o he ollowing
p ocess.
1) DQN-BASED TF-AGENT
The DQN agen is sui able o any en i onmen al condi ion
possessed by a disc e e ac ion space o mula ed de e -
minis ically o simplici y and expec a ions o e s ochas ic
en i onmen al ansi ions. The main goal o he DQN agen
in his model is o ain a policy o maximize he discoun ed
cumula i e ewa d (1).
R 0=X∞
= 0
γ − 0 (1)
Tha is also known as he e u n alue R 0. In mos RL-based
ne wo ks, he discoun ac o γshould be a cons an alue
highligh ing he sum o con e ges be ween 0 and 1. I allows
ou agen o gain be e ewa d alue esul s by a oiding
unce ain en i onmen ea u e in o ma ion and iden i ying
which is less ele an han a ai ly con iden one. The Q∗is
o achie e an a o dable ewa d o e u n alue emphasized:
Q∗:S a e ×Ac ion →Rwhen he ac ion is aken in a
gi en s a e, he e u n esul s om a cons uc ed policy o
achie e maximized ewa ds (2).
π∗(s)=a gmin
a
Q∗(s,a)(2)
In ou i ual eali y wo ld simula ion model, we will ha e
access o s a e and ac ion space in o ma ion ela ed o he Q∗
alue unc ion o c ea e and ain he Q-ne wo k. Mos o he
Q unc ions in he case o policy- equi ed condi ions obey
Bellman’s [39] equa ion (3).
Qπ(s,a)=R+γQπ(s′, π(s′)) (3)
The di e ence be ween ini ial and lea ned alues calcula-
ions ollowing he equali y equa ion, also known as Q- alue
upda ing [39] o he empo al di e ence e o (4).
δ=Qnew(s,a)=(1 −α)Q(s,a)
| {z }
old alue
+α
lea ned alue
z }| {
R +1+γmax
a′Qs′,a′(4)
Equa ion (4) abo e calcula es he upda ing Q- alue o he
s a e-ac ion (s,a)pai a ime s ep . I is assumed o be equal
o a weigh ed sum o old and lea ned alues, whe e he ini ial
old alue would be 0 since he agen is expe iencing his pa -
icula s a e-ac ion pai alue. The old alue is mul iplied by
(1−α). αlea ning a e is deno ed and se as α=0.001 o
ou de aul aining ne wo k. Ins ead o o e w i ing he newly
calcula ed Q- alue, he αlea ning a e is se o de e mine
he p e iously compu ed Q- alue amoun o he ini ial s a e-
ac ion pai . To e ain he ecen ly ob ained Q- alue la e ,
we gi e a highe lea ning a e o he equal s a e-ac ion ma ch
o adop he d one agen quickly o he compu ed Q- alue.
Howe e , i should be a a balanced lea ning a e o keep
he ade-o be ween he p e ious and new Q- alues o he
124132 VOLUME 11, 2023
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
FIGURE 1. P oposed DRL-based TF-agen objec acking baseline in eg a ion wi h he Game Engine.
FIGURE 2. The low cha o he DRL-based -agen objec acking model in Game Engine.
FIGURE 3. The p ocess o Q-lea ning upda e.
u he aining p ocess. A lea ned alue is a ewa d R +1 ha
he d one agen ecei es mo ing andomly om he s a ing
s a e poin , plus discoun ed es ima ion γo op imal u u e
Q- alue o a new s a e-ac ion ma ch (s′,a′n 1- ime s eps.
The ou pu o he lea ned alue mul iplica ion by he lea ning
a e αis done o ge he op imal policy alue upda e. The
Q-lea ning p ocess upda e illus a ion is in Figu e 3.
As illus a ed in Figu e 3, he e can be se e al ac ions
h ough he aining o lea ning p ocess whe e an agen
chooses he seemingly op imal ac ions Qπ(s ,a )and
ecei es a ewa d o he agen ’s pe o mance h ough s eps in
a i ual en i onmen . Fo u he lea ning, he agen should
choose an ac ion om he S +1s a e o con inuously lea n and
analyze he en i onmen wi h mo e p o ound ea u e esul s.
He e, he g eedy epsilon op ion is a s aigh o wa d s a -
egy o balancing explo a ion and exploi a ion by andomly
selec ing be ween he wo. The me hod, whe e epsilon is
he likelihood o selec ing o explo e o exploi , de e mines
whe he i p oceeds o explo e he en i onmen wi h a sligh
chance. We can see he ac ion selec ion wi h he epsilon
g eedy me hod ma hema ical o mula ion below he equa ion
(5).
Ac ion a ime ( ) a =(max
aQ (a)1−ϵ
any ac ion (a)ϵ(5)
The ac ion selec ion me hod o u he lea ning can be
de ailed ho oughly when he alue is 1−ϵ, and he agen
uses exploi a ion o ake ad an age o p io knowledge, which
is a bes -es ima ed ewa d; o he wise, ϵi akes explo a ion o
look o new op imal op ions.
The alue o each ac ion mus be speci ied o ou agen
o choose he one ha will esul in he bes ewa d. The
ac ion- alue es ima ion unc ion (6) uses p obabili y heo y
o de ine hese alues. The p edic ed ewa d ecei ed when
VOLUME 11, 2023 124133

K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
choosing an ac ion om a lis o all po en ial ac ions e e
o he ‘‘ alue o ha ac ion’’. So, we u ilize he ‘‘sample-
a e age’’ app oach o es ima e he alue o doing an ac ion
since he agen does no know he alue o choosing a pa ic-
ula ac ion.
Q (a)
=Sum o ewa ds when ac ion (a) aken be o e ime ( )
Numbe o imes ac ion (a) aken be o e ime ( )
=P −1
i=1Ri
−1(6)
The agen will hen selec he ac ion wi h he mos ou s and-
ing es ima ed alue, e e ed o as a g eedy ac ion, once he
alue Q(s′,a′)has i s pick a e.
2) PPO-BASED TF-AGENT
P oximal Policy Op imiza ion (PPO) is a s aigh o wa d pol-
icy g adien app oach o RL-based op imiza ion p oblems
ha al e na es be ween op imizing a ‘‘su oga e’’ objec i e
unc ion using s ochas ic g adien ascen and sampling da a
h ough in e ac ion wi h he en i onmen [30]. The main idea
and di e ence be ween p ima y and no el policy g adien
me hods is ha he mul iple-upda e miniba ch objec i e unc-
ion is applied. A he same ime, he s anda d model upda es
he g adien pe da a sample in a single epoch.
A policy g adien echnique known as he P oximal Pol-
icy Op imiza ion (PPO) algo i hm is applied o imp o e he
policy o a ein o cemen lea ning agen . PPO is a se o
algo i hms ha includes PPO1 and PPO2. In his p oposal,
we will ocus on he PPO1 algo i hm. The Clipped Su oga e
Goal is a d op-in subs i u e o he policy g adien objec i e
o inc ease aining s abili y by es ic ing he policy change
a each s ep. To add ess hese and o he di icul ies, we may
limi he amoun , al e he policy, and ensu e i cons an ly
imp o es. Fu he mo e, implemen ing his model helps o
in eg a e wi h a comple e p ocessing algo i hm o achie e
e icien samples om inpu images and minimize hype pa-
ame e uning indica o s. I achie es he same pe o mance
imp o emen s while a oiding complexi y by op imizing he
basics o he Clipped Su oga e Objec i e (7).
LCLIP
(θ)=ˆ
E hmin( (θ)ˆ
A ,clip( (θ),1−ϵ,1+ϵ)ˆ
A )i
(7)
(θ)ˆ
A – iden i ies he same objec i e be o e, inside he
minimiza ion; clip( (θ),1−ϵ,1+ϵ)ˆ
A – his pa o
o mula ion is he same objec i e, bu (θ)is clipped be ween
(1−ϵ,1+ϵ);The comple e min( (θ)ˆ
A ,clip( (θ),
1−ϵ, 1+ϵ)ˆ
A )– episode shows he min o he same objec i e
om be o e and he clipped one;
The main objec i e o clipping su oga es is a egion clip-
ping p ocess ha p e en s he algo i hm om ge ing oo
g eedy and ying o upda e oo much a once while aining
and lea ning o lea e he egion wi h good samples o es i-
ma ion and summa izing. PPO enables us o conduc many
FIGURE 4. P obabili y a io o he su oga e unc ion LCLIP wi h posi i e
A>0 and nega i e A<0 ad an ages. The ed ci cle on each plo shows
he s a ing poin o he op imiza ion, i.e., =1. The sum o he
su oga e unc ion LCLIP can be pe o med o many e ms [38].
g adien ascen epochs on ou da a s eam wi hou igge ing
ha m ully massi e policy modi ica ions. Conduc ing hese
p ocesses helps ge mo e ou o collec ed o s eamed da a
while dec easing sample ine iciency. Mo eo e , he unning
PPO policy uses N pa allel ac o s ha indi idually collec
da a. The da a is collec ed in o mini-ba ches and hen ained
o K epochs using he Clipped Su oga e Objec i e unc ion.
The Clipped Su oga e Objec i e will a ec and op imize
e e y ac ion he agen akes. Upda ing should be s opped i
he ac ion is be e (posi i e) A>0and mo e p obable while
aking he end o g adien s eps. O he wise, when he d one
ac ion is di ec ed in he w ong di ec ion bu he ac ion is
good, i can be edi ec ed o undone om he ini ial s a e.
In he case o lousy ac ion (nega i e) A<0and less p obable
ou comes gained om agen s’ ac ions, an agen needs o ake
sho s eps, o hey do no need o go a s eps in ac ion space.
When i comes o he no malized le el o upda ing, i can be
con olled in he balanced a ea o ge he op imal p obabili y
a io o agen s. The illus a ion o he LCLIP su oga e
unc ion p obabili y can be seen below in a summa izing
Figu e 4.
Figu e 4, illus a es clipped su oga e objec i e unc ions
op imiza ion pa ame e s o he unning lea ning pe iod
wi h p obabili y a io. E ec i ely, his echnique can be
encou aged using signi ican policy changes ac oss lea ning
en i onmen s o inpu da a o be e p obabili y op imiza ion
wi h a model agen . PPO wi h clipped objec echnique shows
he di e ence be ween wo main ained policy ne wo ks, he
cu en πθ(a |s )and he las used policy πθk(a |s )applied
o collec samples. A new policy e alua ion comes om
necessa y sampling, which in ol es collec ing old policy
samples o imp o e e iciency.
The inal loss unc ion o he PPO ac o -c i ic s yle looks
below equa ion (8), a combina ion o he Clipped Su oga e
Objec i e unc ion, Value Loss Func ion, and En opy bonus.
LCLIP+VF+S
(θ)=ˆ
E hLCLIP
(θ)−c1LVF
(θ)i
+c2S[πθ](s )(8)
The gi en equa ion abo e includes se e al pa s ha can
be complex o unde s and, ye i gi es mo e p io i y o
124134 VOLUME 11, 2023
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
achie ing mo e accu a e esul s while applying hem o he
expe imen a ion p ocess. The explained i s pa is a Clipped
Su oga e Objec i e unc ion gi en in he (7) equa ion. In (8),
gi en abo e, c1and c2a e coe icien s o he alue o he
ela ed pa ame e o calcula ion. LVF
(θ)iden i ies as a
squa ed-e o alue loss: (Vθ(s )−V a g
)2. To ensu e he
su icien explo a ion o unknown and complex scena ios in
i ual en i onmen s added an en y as a bonus S[πθ](s ).
IV. EXPERIMENTAL RESULTS
One o he main objec i es and ocuses was o ge he
mos ad an age om he simula ion pla o m o pe o m
expe imen s in di e en condi ions wi h pa ame ic changes.
Many ela ed esea che s used se e al simula ion pla o ms
o es and e alua e hei algo i hms in se e al e alua ion
s udies. The e a e di e en me hods o ge an ad an age
om he ealis ic i ual pla o m. In mos cases, pla o ms
apply o expe imen ing pu poses only. Howe e , i can
also be p io i ized widely in lea ning, aining, and es ing.
The cu en de elopmen o simula ion pla o ms like Uni y,
Un eal Engine, and Cecium gi es g ea oppo uni ies and
ad an ages o p ocess and expe imen wi h s a e-o - he-a
models in mul iple and imp ac ical ci cums ances. One o
he p ime ea u es o he simula o s is he in e connec ion
be ween p og amming languages (Ja asc ip , Py hon, Go,
Ja a, Ko lin, PHP, C#, Swi , e c.) and amewo ks (Angula ,
jQue y, Reac , Ruby, and Rails, Vue, ASP.NET Co e, Django,
Exp ess, e c.). Howe e , building o se ing up his ype o
a chi ec u e and amewo k is qui e icky, and i could only
be success ul in some cases due o hi d-pa y p og ams’
and lib a ies’ con lic s and disp opo ionali y. In his esea ch
wo k, we conduc ed expe imen s wi h di e en pa ame ic
changes and ine uning, as explained in he ollowing chap e
sessions.
A. TRAINING RESULTS
We ha e ained ou p oposed model wi h a simple Bloks
en i onmen by inse ing andomly mo ing objec s o lea n
en i onmen al space and o c ea e a model o u u e es ing
and e alua ion pu poses. We applied wo ypes o -agen
models, DQN and PPO-based -agen s, o achie e mo e
compa able ou pu esul s wi h a 0.001 lea ning a e con ig-
u a ion.
Figu e 5abo e illus a es he minimum ewa d ou pu s o
he ained models in a ypical Blocks en i onmen , whe e
he DQN-based -agen and he PPO-based -agen model
a e ma ked wi h blue and pink, espec i ely. The minimum
ewa d is he smalles alue ha he agen can ecei e as a
ewa d du ing he aining p ocess. The minimum ewa d is
ypically nega i e since mos p oblems in ol e a penal y o
making subop imal decisions – he aining epoch and ewa d
a 2000 and 50, espec i ely. Fu he mo e, he backg ound
was se wi h a plo ab colo in each me hod o show he
o e all pe o mance o he aining agen s. Each agen model
ini ially gained di e en ewa ds, whe eas he DQN-based
FIGURE 5. The minimum ecei ed ewa d ou pu o wo aining models:
DQN-based TF-AGENT and PPO-based TF-AGENT.
FIGURE 6. The ecei ed maximum ewa d ou pu o wo ypes o aining
models: DQN-based TF-AGENT and PPO-based TF-AGENT.
FIGURE 7. The ecei ed a e age ewa ds ou come o aining o DQN
and PPO-based TF-AGENTS.
agen pe o med be e . Ne e heless, a he end o he ain-
ing epochs, he PPO-based agen ecei es be e esul s han
he DQN-based model agen . The whole aining ewa d pe -
o mance illus a ion in Figu e 6abo e, se o 2000 and 50,
aining epoch and ewa d, espec i ely. The maximum poin
is he highes alue ha he agen can ecei e in he aining
p ocess. The maximum ewa d is ypically a posi i e alue
since mos p oblems in ol e a ewa d o making op imal
decisions. The DQN-based TF-AGENT model ini ially gains
a highe ewa d alue in his g aph. Howe e , he PPO-based
model pe o ms be e a e 400 epochs un il he end o he
aining s eps. Unde s anding he ange o possible ewa ds
can help se he hype pa ame e s o he models, such as he
lea ning a e o he discoun ac o . I can also help assess he
pe o mance o he ained agen , as he ewa ds ob ained by
VOLUME 11, 2023 124135
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
FIGURE 8. The es ing ewa d dis ibu ion o he DQN-based TF-AGENT (a) and he PPO-based TF-AGENT (b) in 50 s eps o he episode.
he agen a e compa ed agains he minimum and maximum
possible alues.
The gi en a e age ewa d below e e s o he mean alue o
he ewa ds ecei ed by agen s du ing hei in e ac ions wi h
he en i onmen while aining using DQN and PPO-based
model algo i hms. The a e age ewa d is essen ial o e alua -
ing he agen s’ pe o mance du ing aining. Du ing aining,
agen s y o lea n an op imal policy ha maximizes he cumu-
la i e ewa d ob ained o e ime. Calcula ing he a e age
e alua ion ewa d is done by di iding he sum o he ewa ds
ecei ed du ing all episodes by di iding i by he o al numbe
o episodes. The es ima ed calcula ion is he a e age ewa d
he agen will ecei e when in e ac ing wi h he en i onmen
using he lea ned policy.
By e alua ing he a e age ewa d ecei ed, we can see he
di e ence be ween he DQN and PPO-based model’s pe o -
mance in a ied con igu a ions. Howe e , in some scena ios,
he a e age ewa d may no be he mos sui able me ic o
e alua ing he agen ’s pe o mance.
B. TESTING RESULTS
We ha e es ed ou p oposed DQN and PPO-based model
agen s wi h he same en i onmen al condi ion bu di e en
unseen es episodes o explo e he abili y o he models and
compa e hei pe o mance. As men ioned abo e, he DRL-
based algo i hm’s pe o mance e alua ion di e s om o he
s a e-o - he-a algo i hms in he case o pe o mance me -
ics e alua ions and compa ison echniques. The agen -based
models’ p ecision can be seen o aken as a ecei ed ewa d
alue. As men ioned ea lie , he a e age ewa d ob ained by
he agen du ing aining can help c ea e a model and apply
124136 VOLUME 11, 2023
K. Fa khodo e al.: Deep Rein o cemen Lea ning T -Agen -Based Objec T acking
his model o he es ing p ocess as a pe o mance me ic. This
me ic measu es he agen ’s abili y o na iga e he en i on-
men and ob ain expec ed ou pu acking. The diag am below
(Figu e 8) ep esen s he DQN-based TF-AGENT model’s
ou pu wi h he ewa d pe cen age ecei ed om unseen
es ing scena ios. In he es ing session, he ecei ed ewa d
pe cen age was se o 100 in he 50 s eps espec i ely in
e e y episode. The o e all ecei ed in e e y s ep ma ked
wi h a column and ed line illus a es he smoo hed alue o
he DQN-based TF-AGENT es ing esul s ained and es ed
wi h s anda d ewa d in Figu e 8 (a).
Figu e 8 (b) shows he PPO-based TF-AGENT’s ecei ed
pe cen age ewa d es ing esul s in he 50 s eps o he
episode, along wi h smoo hed ed line ou pu . The di e en
esul s be ween he DQN and PPO models ecei ed ewa ds
in e e y aining s ep. As we can see, he es ing esul s show
ha bo h models gi e high accu acy and p ecise lea ning
pe o mance in e e y es ing s ep ou pu wi h an ele a ed
conclusion.
V. CONCLUSION
In his esea ch wo k, we ha e p esen ed a DQN and PPO-
based TF-AGENT model-based objec acking amewo k
in eg a ed wi h a simple Blocks en i onmen o e alua e he
pe o mance o he p oposed algo i hm. I has been in eg a ed
wi h he simula ion pla o m o highligh he algo i hm’s
o e all pe o mance.
The simula ion pla o m p o ides h ee ypes o essen ial
inpu images o expe imen wi h and e alua e he o e all
s a us. While es ing in a i ual- eali y scena io wi h i ual
d one agen s and ine uning o each he bes o desi ed
esul s, he p oduc i i y and eligibili y o hese pla o ms a e
i al. The DQN and PPO-based i ual -agen d ones lea n
how o de ec and ack an objec inse ed in his pla o m
by ob aining consecu i e ames om a p ima y Blocks en i-
onmen and using a DRL ne wo k o manage he ac ions,
s a es, and acking pipeline. Bo h -agen s a e ained in a
Blocks en i onmen o adap o he su oundings and exis ing
objec s in a simula ion condi ion o addi ional es ing, ack-
ing accu acy, and speed assessmen . In he aining p ocess,
bo h models showed p esen able esul s: minimum 49 (PPO)
and 48 (DQN) ewa ds in 2000 epochs; maximum 49 (DQN)
and 49 (PPO) ewa ds in 2000 epochs; a e age 49 ewa ds
we e ecei ed o bo h (PPO and DQN) models. Models pe -
o mance con as ed 50 s eps o one episode es ing se , whe e
he PPO-based -agen ge s i s pick alue ewa d o 97% in
s ep 23, DQN-based agen ecei es i s max alue o 86% in
he 17 h s ep espec i ely. Howe e , he o e all pe o mance
o he ecei ed pe cen age ewa d g aph (Figu e 8,aand b)
indica es ha he DQN-based model sequen pe o ms be e
han he PPO-based one. Rega ding s abili y, ewa d con-
ibu ion, and nume ic g aphical pe o mance, we examined
and compa ed he algo i hm echniques o a ious es ablished
hype pa ame ic changes wi h ein o cemen lea ning-based
ne wo k con ol inco po a ed in o he simula ion p ocess.
In u u e wo k, we a e going o in eg a e ou model wi h
se e al s a e-o - he-a acking echniques o imp o e he
pe o mance o he a ge acking amewo k by es ing i in
mo e complex i ual simula ion en i onmen s.
REFERENCES
[1] The Manu ac u e . (Jun. 29, 2022). The Bene i s o D ones in Manu ac-
u ing. [Online]. A ailable: h ps://www. hemanu ac u e .com/a icles/ he-
bene i s-o -d ones-in-manu ac u ing/
[2] C op acke . (Ap . 26, 2022). D one Technology in Ag icul u e. D ag-
on ly IT. [Online]. A ailable: h ps://www.c op acke .com/blog/d one-
echnology-in-ag icul u e.h ml
[3] G. McNeal. (No . 2014). D ones and Ae ial Su eillance:
Conside a ions o Legisla u es. [Online]. A ailable: h ps://www.
b ookings.edu/ esea ch/d ones-and-ae ial-su eillance-conside a ions-
o -legisla u es/
[4] B. Pu ahong, T. Anuwongpini , A. Juhong, I. Kanjanasu a , and
C. Pin a iooj, ‘‘Medical d one managing sys em o au oma ed ex e nal
de ib illa o deli e y se ice,’’ D ones, ol. 6, no. 4, p. 93, Ap . 2022, doi:
10.3390/d ones6040093.
[5] N. Tuśnio and W. W óblewski, ‘‘The e iciency o d ones usage o sa e y
and escue ope a ions in an open a ea: A case om Poland,’’ Sus ainabili y,
ol. 14, no. 1, p. 327, Dec. 2021, doi: 10.3390/su14010327.
[6] M. Elloumi, R. Dhaou, B. Esc ig, H. Idoudi, and L. A. Saidane,
‘‘Moni o ing oad a ic wi h a UAV-based sys em,’’ in P oc. IEEE
Wi eless Commun. Ne w. Con . (WCNC), Ap . 2018, pp. 1–6, doi:
10.1109/WCNC.2018.8377077.
[7] A. Chodo ek, R. R. Chodo ek, and A. Yas ebo , ‘‘Wea he sensing in
an u ban en i onmen wi h he use o a UAV and WebRTC-based pla -
o m: A pilo s udy,’’ Senso s, ol. 21, no. 21, p. 7113, Oc . 2021, doi:
10.3390/s21217113.
[8] K. Jewani, M. Ka a, D. Mo wani, and G. Je hwani, ‘‘Fi e igh e d one,’’ in
P oc. 1s In . Con . Ad . Sci. Inno . Sci., Eng., Technol. (ICASISET), 2020.
[9] C. Huang, C.-E. Lin, Z. Yang, Y. Kong, P. Chen, X. Yang, and K.-T. Cheng,
‘‘Lea ning o ilm om p o essional human mo ion ideos,’’ in P oc.
IEEE/CVF Con . Compu . Vis. Pa e n Recogni . (CVPR), Jun. 2019,
pp. 4239–4248, doi: 10.1109/CVPR.2019.00437.
[10] A. A. El-Sha ie, M. Zaki, and S. E. D. Habib, ‘‘Fas CNN-based objec
acking using localiza ion laye s and deep ea u es in e pola ion,’’ in P oc.
15 h In . Wi eless Commun. Mobile Compu . Con . (IWCMC), Jun. 2019,
pp. 1476–1481, doi: 10.1109/IWCMC.2019.8766466.
[11] R. Ra ind an, M. J. San o a, and M. M. Jamali, ‘‘Mul i-objec de ec-
ion and acking, based on DNN, o au onomous ehicles: A e iew,’’
IEEE Senso s J., ol. 21, no. 5, pp. 5668–5677, Ma . 2021, doi:
10.1109/JSEN.2020.3041615.
[12] X. Fa hodo , K.-S. Moon, S.-H. Lee, and K.-R. Kwon, ‘‘LSTM ne wo k
wi h acking associa ion o mul i-objec acking,’’ J. Ko ea Mul imedia
Soc., ol. 23, no. 10, pp. 1236–1249, Oc . 2020.
[13] D. Gözen and S. Oze , ‘‘Visual objec acking in d one images wi h deep
ein o cemen lea ning,’’ in P oc. 25 h In . Con . Pa e n Recogni . (ICPR),
Jan. 2021, pp. 10082–10089.
[14] S. Shah, D. Dey, C. Lo e , and A. Kapoo , ‘‘Ai Sim: High- ideli y
isual and physical simula ion o au onomous ehicles,’’ 2017,
a Xi :1705.05065.
[15] A. Juliani, V.-P. Be ges, E. Teng, A. Cohen, J. Ha pe , C. Elion, C. Goy,
Y. Gao, H. Hen y, M. Ma a , and D. Lange, ‘‘Uni y: A gene al pla o m
o in elligen agen s,’’ 2018, a Xi :1809.02627.
[16] G. Yang, ‘‘Asymp o ic acking wi h no el in eg al obus schemes o mis-
ma ched unce ain nonlinea sys ems,’’ In . J. Robus Nonlinea Con ol,
ol. 33, no. 3, pp. 1988–2002, Feb. 2023.
[17] G. Yang, T. Zhu, F. Yang, L. Cui, and H. Wang, ‘‘Ou pu eedback adap i e
RISE con ol o unce ain nonlinea sys ems,’’ Asian J. Con ol, ol. 25,
no. 1, pp. 433–442, Jan. 2023.
[18] M. Bha a ai and M. Ma ínez-Ramón, ‘‘A deep Q-lea ning based pa h
planning and na iga ion sys em o i e igh ing en i onmen s,’’ in P oc.
13 h In . Con . Agen s A i . In ell. (ICAART), 2021, pp. 267–277, doi:
10.5220/0010267102670277.
[19] A. P. Kalidas, C. J. Joshua, A. Q. Md, S. Bashee , S. Mohan, and S. Sak i,
‘‘Deep ein o cemen lea ning o ision-based na iga ion o UAVs in
a oiding s a iona y and mobile obs acles,’’ D ones, ol. 7, no. 4, p. 245,
Ap . 2023, doi: 10.3390/d ones7040245.
VOLUME 11, 2023 124137