scieee Science in your language
[en] (orig)

Comprehensive review of vision-based fall detection systems

Abstract

Vision-based fall detection systems have experienced fast development over the last years. To determine the course of its evolution and help new researchers, the main audience of this paper, a comprehensive revision of all published articles in the main scientific databases regarding this area during the last five years has been made. After a selection process, detailed in the Materials and Methods Section, eighty-one systems were thoroughly reviewed. Their characterization and classification techniques were analyzed and categorized. Their performance data were also studied, and comparisons were made to determine which classifying methods best work in this field. The evolution of artificial vision technology, very positively influenced by the incorporation of artificial neural networks, has allowed fall characterization to become more resistant to noise resultant from illumination phenomena or occlusion. The classification has also taken advantage of these networks, and the field starts using robots to make these systems mobile. However, datasets used to train them lack real-world data, raising doubts about their performances facing real elderly falls. In addition, there is no evidence of strong connections between the elderly and the communities of researchers. Gutiérrez, J.; Rodríguez, V.; Martin, S.

Read accessible full text

Comprehensive review of vision-based fall detection systems

Author: Gutiérrez, J.; Rodríguez, V.; Martin, S.
Year: 2021
DOI: 10.3390/s21030947
Source: https://zaguan.unizar.es/record/99771/files/texto_completo.pdf
senso s
Re iew
Comp ehensi e Re iew o Vision-Based Fall De ec ion Sys ems
Jesús Gu ié ez 1,*, Víc o Rod íguez 2and Se gio Ma in 1


Ci a ion: Gu ié ez, J.; Rod íguez, V.;
Ma in, S. Comp ehensi e Re iew o
Vision-Based Fall De ec ion Sys ems.
Senso s 2021,21, 947. h ps://
doi.o g/10.3390/s21030947
Recei ed: 18 Decembe 2020
Accep ed: 25 Janua y 2021
Published: 1 Feb ua y 2021
Publishe ’s No e: MDPI s ays neu al
wi h ega d o ju isdic ional claims in
published maps and ins i u ional a il-
ia ions.
Copy igh : © 2021 by he au ho s.
Licensee MDPI, Basel, Swi ze land.
This a icle is an open access a icle
dis ibu ed unde he e ms and
condi ions o he C ea i e Commons
A ibu ion (CC BY) license (h ps://
c ea i ecommons.o g/licenses/by/
4.0/).
1Uni e sidad Nacional de Educación a Dis ancia, Juan Rosal 12, 28040 Mad id, Spain; [email p o ec ed]
2EduQTech, E.U. Poli écnica, Ma ia Lluna 3, 50018 Za agoza, Spain; [email p o ec ed]
*Co espondence: jgu ie [email p o ec ed]
Abs ac :
Vision-based all de ec ion sys ems ha e expe ienced as de elopmen o e he las yea s.
To de e mine he cou se o i s e olu ion and help new esea che s, he main audience o his pape ,
a comp ehensi e e ision o all published a icles in he main scien i ic da abases ega ding his
a ea du ing he las i e yea s has been made. A e a selec ion p ocess, de ailed in he Ma e ials
and Me hods Sec ion, eigh y-one sys ems we e ho oughly e iewed. Thei cha ac e iza ion and
classi ica ion echniques we e analyzed and ca ego ized. Thei pe o mance da a we e also s udied,
and compa isons we e made o de e mine which classi ying me hods bes wo k in his ield. The
e olu ion o a i icial ision echnology, e y posi i ely in luenced by he inco po a ion o a i icial
neu al ne wo ks, has allowed all cha ac e iza ion o become mo e esis an o noise esul an om
illumina ion phenomena o occlusion. The classi ica ion has also aken ad an age o hese ne wo ks,
and he ield s a s using obo s o make hese sys ems mobile. Howe e , da ase s used o ain hem
lack eal-wo ld da a, aising doub s abou hei pe o mances acing eal elde ly alls. In addi ion,
he e is no e idence o s ong connec ions be ween he elde ly and he communi ies o esea che s.
Keywo ds:
a i icial ision; neu al ne wo ks; all de ec ion; all cha ac e iza ion; all classi ica ion;
all da ase
1. In oduc ion
In acco dance wi h he UN epo on he aging popula ion [
1
], he global popula ion
aged o e 60 doubled i s numbe in 2017 compa ed o 1980. I is expec ed o double again
by 2050 when hey exceed he 2 billion ma k. By his ime, hei numbe will be g ea e
han he numbe o eenage s and youngs e s aged 10 o 24.
The phenomenon o popula ion aging is a global one, mo e ad anced in he de eloped
coun ies, bu also p esen in he de eloping ones, whe e wo- hi ds o he wo lds olde
people li e, a numbe which is ising as .
Wi h his pe spec i e, he amoun o esou ces de o ed o elde ly heal h ca e is
inc easingly high and could, in he non-dis an u u e, become one o he mos ele an
wo ld economic sec o s. Because o his, all elde ly heal h- ela ed a eas ha e a ac ed
g ea esea ch a en ion o e he las decades.
One o he a eas imme sed in his body o esea ch has been human all de ec ion, as,
o his communi y, o e 30% o alls cause impo an inju ies, anging om hip ac u e o
b ain concussion, and a good numbe o hem end up causing dea h [2].
The numbe o echnologies used o de ec alls is wide, and a huge numbe o sys ems
able o wo k wi h hem ha e been de eloped by esea che s. These sys ems, in b oad
e ms, can be classi ied as wea able, ambien and came a-based ones [3].
The i s block, he wea able sys ems, inco po a e senso s ca ied by he su eilled
indi idual. The echnologies used by his g oup o sys ems a e nume ous, anging om
accele ome e s o p essu e senso s, including inclinome e s, gy oscopes o mic ophones,
among o he senso s. R. Rucco e al. [
4
] ho oughly e iew hese sys ems and s udy hem
in-dep h. In his a icle, sys ems a e classi ied in acco dance wi h he numbe and ype
o senso s, hei placemen and he cha ac e is ics o he s udy made du ing he sys em
Senso s 2021,21, 947. h ps://doi.o g/10.3390/s21030947 h ps://www.mdpi.com/jou nal/senso s
Senso s 2021,21, 947 2 o 50
e alua ion phase concluding ha mos sys ems inco po a e one o wo accele ome ic
senso s a ached o he unk.
The second block includes sys ems whose senso s a e placed a ound he moni o ed
pe son and include p essu e, acous ic, in a- ed, and adio- equency senso s. The las
block, he objec o his e iew, g oups sys ems able o iden i y alls h ough a i icial ision.
In pa allel, o e he las yea s, a i icial ision has expe ienced as de elopmen ,
mainly due o he use o a i icial neu al ne wo ks and hei abili y o ecognize objec s
and ac ions.
This a i icial ision de elopmen applied o human ac i i y ecogni ion in gene al,
and human all de ec ion in pa icula , has gi en e y ui ul ou comes in he las decade.
Howe e , up o whe e we know, no sys ema ic e iews on he speci ic a ea o ision-
based de ec ion sys ems ha e been made, as all e e ences o his ield ha e been included
in gene ic all de ec ion sys em e iews.
This e iew in ends o shed some ligh on he p ocess o de elopmen ollowed by
ision-based all de ec ion sys ems, so esea che s ge a clea image o wha has been done
in his ield du ing he las i e yea s ha help hem in hei in es iga ion p ocess. In his
s udy, au ho s in end o show he main ad an ages and disad an ages o all p ocesses
and algo i hms used in he e iewed sys ems so new de elope s ge a clea pic u e o he
s a e o he a in he ield o human all de ec ion h ough a i icial ision, an a ea ha
could signi ican ly imp o e li ing s anda ds o he dependen communi y and ha e a
high impac on hei day- o-day li es.
The a icle is o ganized as ollows: In Sec ion 2, Ma e ials and Me hods, cha ac e iza-
ion and classi ica ion echniques a e desc ibed and applied o he p eselec ed sys ems, so
a numbe o hem a e inally decla ed as eligible o be included in his e iew. In Sec ion 3,
Resul s, hose sys ems a e p esen ed and oughly desc ibed, he da abases used o hei
alida ion a e p esen ed, and some pe o mance compa isons a e made. In he nex Sec ion,
Discussion, he algo i hms and p ocesses used by he sys ems a e desc ibed and, in he
las pa o he e iew, Sec ion 5, conclusions a e ex ac ed based on all he p e iously
p esen ed in o ma ion.
2. Ma e ials and Me hods
In his pape , we ocus on a i icial ision sys ems able o de ec human alls. To ul ill
his pu pose, we ha e pe o med a deep e iew o all published pape s p esen in public
da abases o esea ch documen a ion (ScienceDi ec , IEEE Explo e , Senso s da abase).
This documen al sea ch was based on di e en ex s ing sea ches and was execu ed om
May 2020 o Decembe 2020. The ime ame o publica ion was es ablished be ween 2015
and 2020, so he las de elopmen s in he ield can be iden i ied, and he s udy se es o
o ien a e new esea che s. The e ms used in he bibliog aphical Boolean explo a ion we e
“ all de ec ion” and “ ision”. A seconda y sea ch was ca ied ou o comple e he i s one
by using o he sea ch engines o schola ly li e a u e ocused on heal h (PubMed, MedLine).
All sea ches ha e been limi ed o a icles and publica ions in English, language used by
mos a ea esea che s.
A e an ini ial analysis o pape s ul illing hese sea ching c i e ia 81 a icles, desc ib-
ing he same numbe o sys ems we e selec ed. They illus a e how all de ec ion sys ems
based on a i icial ision ha e e ol ed in he las i e yea s.
The selec ion p ocess included an ini ial sc eening made h ough e e ence manage-
men so wa e o gua an ee no duplica ion, and a manual sc eening, whose objec i e was
making su e he a icle co e ed he ield, did no all wi hin he ield o he all p e en ion
o human ac i i y ecogni ion (HAR), did no mix ision echnologies wi h o he ones and
we e no s udies in ending o classi y he human gai as an indica o o all p obabili y.
This way, he e iew is pu ely cen e ed on a i icial ision all de ec ion.
The en i e p ocess is summa ized in he low diag am shown in Figu e 1.
Senso s 2021,21, 947 3 o 50
Senso s 2020, 20, x FOR PEER REVIEW 3 o 50
Figu e 1. Flow diag am o adop ed sea ch and selec ion s a egy o pape selec ion.
All selec ed sys ems we e s udied one-by-one o de e mine hei cha ac e iza ion and
classi ica ion echniques, desc ibing hem in-dep h in he Discussion (Sec ion 4), so a ull
axonomy can be made based on hei cha ac e is ics. In addi ion, pe o mance compa i-
sons a e also included, so conclusions on which ones a e he mos sui able sys ems can be
eached.
3. Resul s
The a icle sea ch and selec ion p ocess s a ed wi h an ini ial iden i ica ion o 929
po en ial a icles. Duplica ed ones and hose whose i le clea ly did no ma ch he equi ed
con en we e disca ded, lea ing 430 a icles ha we e assessed o eligibili y. These a i-
cles we e hen e iewed, and hose ela ed o HAR, all p e en ion, mixed echnologies,
gai s udies and he ones which did no co e he a ea o ision-based all de ec ion we e
disca ded, so; inally, 81 a icles a e conside ed in he e iew.
The selec ed sys ems we e ho oughly e ised and classi ied in acco dance wi h he
used cha ac e iza ion and classi ica ion me hods, as well as he employed ype o signal.
The used da ase o pe o mance de e mina ion and i s indica o s alues ha e also been
s udied. All his in o ma ion is included in Table 1.
Sys em compa ison da a we e used o de elop Table 2, and inally, all main cha ac-
e is ics o publicly accessible da ase s used by any o he sys ems a e included in Table 3.
Figu e 1. Flow diag am o adop ed sea ch and selec ion s a egy o pape selec ion.
All selec ed sys ems we e s udied one-by-one o de e mine hei cha ac e iza ion
and classi ica ion echniques, desc ibing hem in-dep h in he Discussion (Sec ion 4), so
a ull axonomy can be made based on hei cha ac e is ics. In addi ion, pe o mance
compa isons a e also included, so conclusions on which ones a e he mos sui able sys ems
can be eached.
3. Resul s
The a icle sea ch and selec ion p ocess s a ed wi h an ini ial iden i ica ion o 929 po-
en ial a icles. Duplica ed ones and hose whose i le clea ly did no ma ch he equi ed
con en we e disca ded, lea ing 430 a icles ha we e assessed o eligibili y. These a i-
cles we e hen e iewed, and hose ela ed o HAR, all p e en ion, mixed echnologies,
gai s udies and he ones which did no co e he a ea o ision-based all de ec ion we e
disca ded, so; inally, 81 a icles a e conside ed in he e iew.
The selec ed sys ems we e ho oughly e ised and classi ied in acco dance wi h he
used cha ac e iza ion and classi ica ion me hods, as well as he employed ype o signal.
The used da ase o pe o mance de e mina ion and i s indica o s alues ha e also been
s udied. All his in o ma ion is included in Table 1.
Sys em compa ison da a we e used o de elop Table 2, and inally, all main cha ac e -
is ics o publicly accessible da ase s used by any o he sys ems a e included in Table 3.
Senso s 2021,21, 947 4 o 50
Table 1. Vision-based all de ec ion sys ems published 2015–2020.
Re e ence Yea Cha ac e iza ion (Global/Local/Dep h) Classi ica ion Inpu
Signal Used Da ase s Pe o mance
A. Yajai
e al. [5]2015
Skele on join acking model p o ided
by MS Kinec ®is used o ack join s and
build a 2D and 3D bounding box a ound
he body/dep h cha ac e iza ion
Fea u e- h eshold-based.
•
Heigh /wid h a io o he bounding
box
•cen e o g a i y (CG) posi ion in
ela ion o suppo polygon (de ined
by ankle join s)
Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy 98.43%
Speci ici y 98.75%
Recall 98.12%
C. -J. Chong
e al. [6]2015 Pixel clus e ing and backg ound
(Ho p ase )/global cha ac e iza ion
Fea u e- h eshold-based.
Me hod 1:
•Bounding box (BB) aspec a io
•CG posi ion
Me hod 2:
•Ellipse o ien a ion and aspec a io
•Mo ion his o y image (MHI)
Red-
g een-
blue
(RGB)
Speci ic ideo da ase —no public
access a e ision ime
Me hod 1
Sensi i i y 66.7%
Speci ici y 80%
Me hod 2
Sensi i i y 72.2%
Speci ici y 90%
H. Rajabi
e al. [7]2015
Fo eg ound ex ac ion h ough
backg ound sub ac ion (Gaussian mixed
models—GMM) and Sobel il e
applica ion/ global cha ac e iza ion
Fea u e- h eshold-based.
•BB o ien a ion angle
•Change o CG wid h
•Heigh /wid h ela ion o con ou
•Hu momen in a ian s
RGB
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Fall de ec ion success
a e 81%
L. H. Juang
e al. [8]2015
Fo eg ound ex ac ion h ough
backg ound sub ac ion (op ical
low-based) and human join s
iden i ied/global cha ac e iza ion
Suppo ec o machine (SVM) RGB
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy up o 100%
M. A. Mousse
e al. [9]2015
Fo eg ound ex ac ion h ough pixel
colo and b igh ness dis o ion
de e mina ion and in eg a ion o
o eg ound maps h ough
homog aphy/global cha ac e iza ion
Fea u e- h eshold-based.
Ra io obse ed silhoue e a ea/silhoue e
a ea p ojec ed on he g ound plane
RGB—2
OR-
THOGO-
NAL
VIEWS
Mul icam Fall Da ase [10]Sensi i i y 95.8%
Speci ici y 100%
Muza e
Aslan
e al. [11]
2015
Human silhoue e is segmen ed using
dep h in o ma ion, and cu a u e scale
space (CSS) is calcula ed and encoded in
a Fishe ec o /dep h cha ac e iza ion
SVM Dep h SDUFall [12]A e age accu acy
88.01%
Senso s 2021,21, 947 5 o 50
Table 1. Con .
Re e ence Yea Cha ac e iza ion (Global/Local/Dep h) Classi ica ion Inpu
Signal Used Da ase s Pe o mance
Z. Bian
e al. [13]2015
Silhoue e ex ac ion by using dep h
in o ma ion. Human body join s
iden i ied and acked wi h o so
o a ion/dep h cha ac e iza ion
SVM Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Sensi i i y 95.8%
Speci ici y 100%
C. Lin
e al. [14]2016
Fo eg ound ex ac ion h ough
backg ound sub ac ion (GMM)/global
cha ac e iza ion
Fea u e- h eshold-based.
•Ellipse o ien a ion
•Linea and angula accele a ion
•MHI
RGB
This sys em-speci ic ideo
da ase —no public access a
e ision ime
No published
F. Me ouche
e al. [15]2016
Fo eg ound ex ac ion by using he
di e ence be ween dep h ames and
head acking h ough pa icle
il e /dep h cha ac e iza ion
Fea u e- h eshold-based.
•Ra io head e ical posi ion/pe son
heigh
•CG eloci y
Dep h SDUFall [12]
Sensi i i y 90.76%
Speci ici y 93.52%
Accu acy 92.98%
K. G. Gunale
e al. [16]2016
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison)/global cha ac e iza ion
K-nea es neighbo (KNN) RGB Chu e da ase —no public access a
e ision ime
Accu acy
Fall 90%
No all 100%
K. R. Bha ya
e al. [17]2016
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison)/global cha ac e iza ion +
op ical low (OF)/global cha ac e iza ion
KNN on MHI and OF ea u es RGB
This sys em-speci ic ideo
da ase —no public access a
e ision ime
No published
Kun Wang
e al. [18]2016
Segmen a ion h ough ibe [19] and
his og am o o ien ed g adien s (HOG)
and local bina y pa e n (LBP)/global
cha ac e iza ion + ea u e maps ob ained
h ough con olu ional neu al ne wo k
(CNN)/ local cha ac e iza ion
SVM-linea ke nel RGB
Mul icam Fall Da ase [10] and
SIMPLE Fall De ec ion Da ase [20]
and This sys em-speci ic ideo
da ase —no public access a
e ision ime
Sensi i i y 93.7%
Speci ici y 92%
U. P a ap
e al. [21]2016
Fo eg ound ex ac ion h ough
backg ound sub ac ion (GMM)/global
cha ac e iza ion
Fea u e- h eshold-based.
•Silhoue e CG s a iona y o e a
h eshold ime limi
RGB Speci ic ideo da ase s—no public
access a e ision ime
Fall de ec ion a e
92%
False ala m a e
6.25%
X. Wang
e al. [22]2016
Segmen a ion h ough ibe [19] and
uppe body da abase popula ed and
spa se OF de e mined/global
cha ac e iza ion
Fea u e- h eshold-based.
•Body a io wid h/heigh
•Ve ical eloci y de i ed om OF
•Uppe body posi ion his o y
RGB LE2I [23]A e age p ecision
81.55%

Senso s 2021,21, 947 6 o 50
Table 1. Con .
Re e ence Yea Cha ac e iza ion (Global/Local/Dep h) Classi ica ion Inpu
Signal Used Da ase s Pe o mance
A. Y. Alaoui
e al. [24]2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison)/global cha ac e iza ion +
OF/global cha ac e iza ion
No classi ica ion algo i hm epo ed RGB CHARFI2012 Da ase [25]P ecision 91%
Sensi i i y 86.66%
Apiche Yajai
e al. [26]2017 Skele on join acking model p o ided
by MS Kinec ®/dep h cha ac e iza ion
Fea u e- h eshold-based.
Aspec a ios:
•Bounding box
•CoG
•Bounding box diagonal s. max.
heigh
•Bounding box heigh s. max.
heigh
Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy 98.15%
Sensi i i y 97.75%
Speci ici y 98.25%
B.
Lewandowski
e al. [27]
2017
oxels a ound he poin cloud a e
calcula ed. The ones classi ied as human
a e clus e ed, and IRON ea u es a e
calcula ed/local cha ac e iza ion
Fea u e- h eshold-based.
•Mahalanobis dis ance be ween
clus e IRON ea u es and he
dis ibu ion o IRON ea u es om
allen bodies
Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Sensi i i y in
ope a ional
en i onmen s 99%
F. Ha ou
e al. [28]2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison)/dep h cha ac e iza ion
Mul i a ia e exponen ially weigh ed
mo ing a e age (MEWMA)-SVM
KNN
A i icial neu al ne wo k (ANN)
Naï e Bayes (NB)
RGB UR Fall De ec ion [29] &
Fall De ec ion Da ase [30]
Accu acy
KNN 91.94%
ANN 95.15%
NB 93.55%
NEWMA-SVM
96.66%
G. M.
Basa a aj
e al. [31]
2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (median)/global
cha ac e iza ion
Fea u e- h eshold-based.
•Ellipse eccen ici y and o ien a ion
•MHI
RGB
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy
Fall 86.66%
Non- all 90%
K. Adhika i
e al. [30]2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison) using bo h RGB echniques
and dep h ones and Fea u e maps
ob ained h ough CNN/local and dep h
cha ac e iza ion
So max based on ea u es ec o om
CNN Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
O e all, accu acy
74%
Sys em sensi i i y o
lying pose 99%
Senso s 2021,21, 947 7 o 50
Table 1. Con .
Re e ence Yea Cha ac e iza ion (Global/Local/Dep h) Classi ica ion Inpu
Signal Used Da ase s Pe o mance
Koldo De
Miguel
e al. [32]
2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (GMM) + Spa se
OF de e mined/global cha ac e iza ion
KNN on silhoue e and OF ea u es RGB
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy 96.9%
Sensi i i y 96%
Speci ici y 97.6%
Leiyue Yao
e al. [33]2017 Skele on join acking model p o ided
by MS Kinec ®/dep h cha ac e iza ion
Fea u e- h eshold-based
•To so angle
•Cen oid heigh
Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy 97.5%
T ue posi i e a e
98%
T ue nega i e a e
97%
M. An onello
e al. [34]2017
oxels a ound he poin cloud a e
calcula ed. Then hey a e segmen ed in
homogeneous pa ches and he ones
classi ied as human a e ga he ed and
classi ied o no as a human lying
body/dep h cha ac e iza ion
SVM— adial-based ke nel Dep h IASLAB-RGBD allen pe son
Da ase [35]
Se A
Accu acy: single
iew (SV)
0.87/SV+map
e i ica ion (MV)
0.92
P ecision: SV
0.73/SV+MV 0.85
Recall: SV
0.85/SV+MV 0.85
Se B
Accu acy: SV
0.88/SV+MV 0.9
P ecision: SV
0.8/SV+MV 0.87
Recall: SV
0.86/SV+MV 0.81
M. N. H.
Mohd
e al. [36]
2017
Skele on join acking model p o ided
by MS Kinec ®is used o de e mine join
posi ions and speeds/dep h
cha ac e iza ion
SVM based on join s speeds and
ule-based decision-based on join s
posi ion in ela ion o knees
Dep h
TST Fall De ec ion [37] and UR Fall
De ec ion [
29
] and Falling De ec ion
[38]
Accu acy 97.39%
Speci ici y 96.61%
Sensi i i y 100%
N. B. Joshi
e al. [39]2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (GMM)/global
cha ac e iza ion
Fea u e- h eshold-based.
•BB wid h/heigh a io
•CG posi ion
•O ien a ion
•Hu momen s
RGB LE2I [23]Speci ici y 92.98%
Accu acy 91.89%
Senso s 2021,21, 947 8 o 50
Table 1. Con .
Re e ence Yea Cha ac e iza ion (Global/Local/Dep h) Classi ica ion Inpu
Signal Used Da ase s Pe o mance
N. O anasap
e al. [40]2017 Skele on join acking model p o ided
by MS Kinec ®/dep h cha ac e iza ion
Fea u e- h eshold-based.
•Head eloci y
•CG posi ion in ela ion o ankle
join s
Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Sensi i i y 97%
Accu acy 100%
Q. Feng
e al. [41]2017
CNN is used o de ec and ack people,
and Sub-MHI a e co ela ed o each
pe son BB/local cha ac e iza ion
SVM RGB UR Fall De ec ion [29]
P ecision 96.8%
Recall 98.1%
F197.4%
S. He nandez-
Mendez
e al. [42]
2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison) and silhoue e acking.
Then cen oid and ea u es a e
de e mined/dep h cha ac e iza ion
Fea u e- h eshold-based.
•Angles and a io heigh /wid h o
he BB
Dep h
Dep h And Accele ome ic Da ase
[43] and his sys em-speci ic ideo
da ase —no public access a
e ision ime
The allen pose is
de ec ed co ec ly on
100% o occasions.
S. Kas u i
e al. [44]2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison)/dep h cha ac e iza ion
SVM Dep h UR Fall De ec ion [29]Sensi i i y 100%
Speci ici y 88.33%
S. Kas u i
e al. [45]2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison)/dep h cha ac e iza ion
SVM Dep h UR Fall De ec ion [29]
Accu acy
To al es ing accu acy
96.34%
S. Pa amase
e al. [46]2017
Body ec o cons uc ion and CG
iden i ica ion aking as s a ing poin 16
pa s o he human body/dep h
cha ac e iza ion
Fea u e- h eshold-based.
•CG accele a ion
•Body ec o / e ical angle
Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy 100%
Sajjad
Tagh aei
e al. [47]
2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion/dep h
cha ac e iza ion
Hidden Ma ko model (HMM) Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy 84.72%
Y. M. Gal ão
e al. [48]2017 Median squa e e o (MSE) e e y 3
ames/global cha ac e iza ion
Mul ilaye pe cep on (MLP)
KNN
SVM—polynomial ke nel
RGB UR Fall De ec ion [29]
F1 sco e:
MLP 0.991
KNN 0.988
SVM—polynomial
ke nel 0.988
Senso s 2021,21, 947 9 o 50
Table 1. Con .
Re e ence Yea Cha ac e iza ion (Global/Local/Dep h) Classi ica ion Inpu
Signal Used Da ase s Pe o mance
Thanh-Hai
T an e al. [
49
]
2017
Skele on join acking model p o ided
by MS Kinec
®
/dep h cha ac e iza ion o
Mo ion map ex ac ion om RGB images
and g adien ke nel desc ip o
calcula ed/global cha ac e iza ion
Fea u e- h eshold-based.
•Heigh o hip join
•Ve ical body eloci y
O
•SVM classi ica ion
Dep h o
RGB
UR Fall De ec ion [
29
] and LE2I [
23
]
and Mul imodal Mul i iew Da ase
o Human Ac i i ies [50]
UR Da ase
Sensi i i y 100%
Speci ici y 99.23%
LE2I Da ase
Sensi i i y 97.95%
Speci ici y 97.87%
MULTIMODAL
Da ase (A e age)
Sensi i i y 92.62%
Speci ici y 100%
X. Li
e al. [51]2017
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison) and ea u e maps ob ained
h ough CNN/ local cha ac e iza ion
So max based on ea u es ec o om
CNN RGB UR Fall De ec ion [29]
Sensi i i y 100%
Speci ici y 99.98%
Accu acy 99.98%
Yaxiang Fan
e al. [52]2017
Fea u e maps ob ained h ough CNN
om dynamic images/local
cha ac e iza ion
Classi ica ion made by ully connec ed
las laye s o CNNs RGB
Mul icam Fall Da ase [10] & LE2I
[23] and High-Quali y Da ase [53]
and This sys em-speci ic ideo
da ase —no public access a
e ision ime
Sensi i i y
LE2I 98.43%
Mul icam 97.1%
HIGH-QUALITY
FALL SIM 74.2%
SYSTEM Da ase
63.7%
A. Abobak
e al. [54]2018
Silhoue e ex ac ion by using dep h
in o ma ion. A ea u e ec o o di e en
body pixels based on dep h di e ence
be ween pai s o poin s is c ea ed/dep h
cha ac e iza ion
Random decision o es o pose
ecogni ion and SVM o mo emen
iden i ica ion
Dep h
UR Fall De ec ion [29] and CMU
G aphics Lab—mo ion cap u e
lib a y [55]
Accu acy 96%
P ecision 91%
Sensi i i y 100%
Speci ici y 93%
B. Dai
e al. [56]2018
Fo eg ound ex ac ion h ough
backg ound sub ac ion (di ec
compa ison)/global cha ac e iza ion
Fea u e- h eshold-based.
•BB segmen ed a eas occupancy.
•CG/heigh a io
•CG e ical speed
RGB
UR Fall De ec ion [29] and This
sys em-speci ic ideo da ase —no
public access a e ision ime
Sensi i i y 95%
Speci ici y 96.7%
Senso s 2021,21, 947 16 o 50
Table 1. Con .
Re e ence Yea Cha ac e iza ion (Global/Local/Dep h) Classi ica ion Inpu
Signal Used Da ase s Pe o mance
Qingzhen Xu
e al. [98]2020
Human keypoin s iden i ied by
OpenPose (con olu ional pose machines
and human body ec o cons uc ion)
and CNN used o ea u e maps
c ea ion/local cha ac e iza ion
So max based on ea u es ec o om
CNN implemen ed in i s las laye RGB
UR Fall De ec ion [29] and
Mul icam Fall Da ase [10] and
NTU RGB+D Da ase [99]
Accu acy a e 91.7%
Swe N. H un
e al. [100]2020
Fo eg ound ex ac ion h ough
backg ound sub ac ion (GMM)/global
cha ac e iza ion
Hidden Ma ko model (HMM) based
onObse able da a:
•Silhoue e su ace
•Cen oid heigh
•Bounding box aspec a io
RGB LE2I [23]
P ecision 99.05%
Recall 98.37%
Accu acy 99.8%
T. Kalinga
e al. [101]2020
Skele on join acking model p o ided
by MS Kinec ®is used o de e mine join
speeds and angles o di e en body
pa s/dep h cha ac e iza ion
Fea u e- h eshold-based.
•Join speeds and angles o body
pa s Dep h
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy 92.5%
Sensi i i y 95.45%
Speci ici y 88%
Weiming
Chen
e al. [102]
2020
Human keypoin s iden i ied by
OpenPose (con olu ional pose machines
and human body ec o
cons uc ion)/local cha ac e iza ion
Fea u e- h eshold-based
•Hip e ical eloci y
•Spine/g ound plane angle
•BB aspec a io
RGB
This sys em-speci ic ideo
da ase —no public access a
e ision ime
Accu acy 97%
Sensi i i y 98.3%
Speci ici y 95%
X. Cai
e al. [103]2020
Fea u e maps ob ained h ough hou glass
con olu ional au o-encode (HCAE)
ANN/local cha ac e iza ion
So max based on ea u es ec o om
HCAE RGB UR Fall De ec ion [29]
Sensi i i y 100%
Speci ici y 93%
Accu acy 96.2%
Y. Chen
e al. [104]2020
Fo eg ound ex ac ion h ough CNN and
Bi-LSTM ANN/local cha ac e iza ion
So max based on ea u es ec o om
RNN-Bi-LSTM RGB
UR Fall De ec ion [29] and This
sys em-speci ic ideo da ase —no
public access a e ision ime
URFD
P ecision 0.897
Recall 0.813
F10.852
Speci ic da ase
P ecision 0.981
Recall 0.923
F10.948

Senso s 2021,21, 947 17 o 50
Table 1. Con .
Re e ence Yea Cha ac e iza ion (Global/Local/Dep h) Classi ica ion Inpu
Signal Used Da ase s Pe o mance
Yuxi Chen
e al. [105]2020
Fea u e maps ob ained h ough 3
di e en CNNs (LeNe , AlexNe y
GoogLeNe )/dep h cha ac e iza ion
Classi ica ion made by ully connec ed
las laye s o CNNs Dep h Video da ase de eloped o he
sys em in [84]
A e age alues
Lene
Sensi i i y 82.78%
Speci ici y 98.07%
AlexNe
Sensi i i y 86.84%
Speci ici y 98.41%
GoogLeNe
Sensi i i y 92.87%
Speci ici y 99%
X. Wang
e al. [106]2020
Fea u e maps ob ained h ough
con olu ional laye s o an ANN/local
cha ac e iza ion
Logis ic unc ion o iden i y
ame-by- ame wo classes in he
p edic ion laye (pe son and allen)
RGB UR Fall De ec ion [29] &
Fall De ec ion Da ase [30]
A e age p ecision
(AP) o allen 0.97
mean a e age
p ecision (mAP) o
bo h classes 0.83
Table 2. Sys em pe o mance compa ison.
Re e ence Yea Inpu Signal ANN/Classi ie s and Pe o mance
C. -J. Chong
e al. [6]2015 RGB
Me hod 1 BB aspec a io and CG posi ion
Sensi i i y 66.7%
Speci ici y 80%
Me hod 2 Ellipse o ien a ion and aspec a io + MHI
Sensi i i y 72.2%
Speci ici y 90%
F. Ha ou
e al. [28]2017 RGB
Accu acy Sensi i i y Speci ici y
KNN 91.94% 100% 86.00%
ANN 95.15% 100% 91.00%
NB 93.55% 100% 88.60%
MEWMA-SVM 96.66% 100% 94.93%
Senso s 2021,21, 947 18 o 50
Table 2. Con .
Re e ence Yea Inpu Signal ANN/Classi ie s and Pe o mance
Y. M. Gal ão
e al. [48]2017 RGB
F1 sco e
Mul ilaye pe cep on (MLP) 0.991
K-nea es neighbo s (KNN) 0.988
SVM—polynomial ke nel 0.988
Leila Panahi
e al. [60]2018 Dep h
A e age esul s
SVM
Sensi i i y 98.52%
Speci ici y 97.35%
Th eshold-based decision
Sensi i i y 98.52%
Speci ici y 97.35%
K. Sehai i
e al. [58]2018 RGB
Accu acy
SVM-RBF 99.27%
KNN 98.91%
ANN 99.61%
Chao Ma e al. [
70
]
2019 RGB + IR
Au oencode
Sensi i i y 93.3%
Speci ici y 92.8%
SVM
Sensi i i y 90.8%
Speci ici y 89.6%
F. Ha ou
e al. [74]2019 RGB
Accu acy:
K-NN 91.94%
ANN 95.16%
Naï e Bayes 93.55%
Decision ee 90.48%
SVM 96.66%
Senso s 2021,21, 947 19 o 50
Table 2. Con .
Re e ence Yea Inpu Signal ANN/Classi ie s and Pe o mance
Rica do Espinosa
e al. [79]2019 RGB
Sensi i i y Speci ici y
So max 97.95% 83.08%
SVM 14.10% 90.03%
RF 14.30% 91.26%
MLP 11.03% 93.65%
KNN 14.35% 90.96%
Xiangbo Kong
e al. [84]2019 Dep h
HOG+SVM LeNe AlexNe GoogLeNe ETDA-Ne
A e age accu acy 89.48% 88.28% 93.53% 96.59% 95.66%
A e age speci ici y
95.43% 97.18% 97.56% 98.76% 99.35%
A e age
sensi i i y 83.75% 74.54% 87.10% 88.74% 91.87%
B. Wang e al. [87]2020 RGB
F1 sco e
Falling s a e
GDBT 95.69%
DT 84.85%
RF 95.92%
SVM 96.1%
KNN 93.78%
MLP 97.41%
Fallen s a e
GDBT 95.27%
DT 95.45%
RF 96.8%
SVM 95.22%
KNN 94.22%
MLP 94.46%
Senso s 2021,21, 947 20 o 50
Table 2. Con .
Re e ence Yea Inpu Signal ANN/Classi ie s and Pe o mance
C. Zhong e al. [
89
]
2020 IR
F1 sco e
RBFNN 89.57 (+/−0.62)
SVM 88.74% (+/−1.75)
So max 87.37% (+/−1.4)
DT 88.9% (+/−0.68)
C. Menacho
e al. [88]2020 RGB
Accu acy
VGG-16 87.81%
VGG-19 88.66%
Incep ion V3 92.57%
ResNe 50 92.57%
Xcep ion 92.57%
ANN p oposed in his sys em 88.55%
G. Sun e al. [90]2020 RGB
Sensi i i y Speci ici y
SVM 92.50% 93.70%
KNN 93.80% 92.30%
SVDD 94.60% 93.80%
Yuxi Chen
e al. [105]2020 Dep h
A e age alues
Lene
Sensi i i y 82.78%
Speci ici y 98.07%
AlexNe
Sensi i i y 86.84%
Speci ici y 98.41%
GoogLeNe
Sensi i i y 92.87%
Speci ici y 99%
Senso s 2021,21, 947 21 o 50
Table 3. Sys em pe o mance e alua ion da ase s.
Signal Type Da ase Name Cha ac e is ics
Accele ome ic and
elec oencephalog am (EEG) and RGB
and passi e in a ed (IR)
Up all [80]17 olun ee s execu e alls and ac i i ies o daily li e (ADL) o di e en ypes eco ded by an
accele ome e , EEG, RGB and passi e IR sys ems
Dep h and Accele ome ic
Dep h and accele ome ic da ase [43] Volun ee s execu e se e al ac i i ies, and alls a e eco ded by a dep h sys em and accele ome e s.
TST all de ec ion [37]11 olun ee s execu e 4 all ypes and 4 ADLs eco ded by RGB-dep h (RGB-D) and accele ome e
sys ems
UR all de ec ion [29] 30 alls and 40 ADLs eco ded by RGB-D and accele ome e sys ems
RGB
Cen e o digi al home da a se —MMU [68] 20 ideos, including 31 alls and se e al ADLs
LE2I [23] 191 di e en ac i i ies, including ADLs and 143 alls
Cha i2012 da ase [25]
250 ideo sequences in ou di e en loca ions, 192 con aining alls, and 57 con aining ADLs.
Ac o s, unde di e en ligh condi ions, mo e in en i onmen s whe e occlusion exi s and clu e ed
and ex u ed backg ound is common
High-quali y da ase [53]
I is a all de ec ion da ase ha a emp s o app oach he quali y o a eal-li e all da ase . I has
ealis ic se ings and all scena ios. In de ail, 55 all scena ios and 17 no mal ac i i y scena ios we e
ilmed by i e web-came as in a oom simila o one in a nu sing home
Mul icam all da ase [10]
The ideo da a se is composed o se e al simula ed no mal daily ac i i ies and alls iewed om 8
di e en came as and pe o med by one subjec in 24 scena ios
Simple all de ec ion da ase [20]The da ase con ains 30 daily ac i i ies such as walking, si ing down, squa ing down, and 21 all
ac i i ies such as o wa d alls, backwa d alls and sideway alls
MO da ase [72]
MOT da ase in ends o be a amewo k o he ai e alua ion o mul iple people acking
algo i hms. In his amewo k, he designe s p o ide:
•De ec ions o all he sequences;
•A common e alua ion ool p o iding se e al measu es, om ecall o p ecision o unning
ime;
•An easy way o compa e he pe o mance o s a e-o - he-a acking me hods;
•
Se e al challenges wi h subse s o da a o speci ic asks such as 3D acking and su eillance.
COCO da ase [73]COCO is a la ge-scale objec de ec ion, segmen a ion, and cap ioning da ase designed o show
common objec s in con ex
Pi opo [96]
Mul iple ac i i ies eco ded in wo di e en scena ios wi h bo h con en ional and ish eye came as

Senso s 2021,21, 947 22 o 50
Table 3. Con .
Signal Type Da ase Name Cha ac e is ics
Dep h
IASLAB-RGB allen pe son da ase [35]I consis s o se e al s a ic and dynamic sequences wi h 15 di e en people and 2 di e en
en i onmen s
Mul imodal mul i iew da ase o human
ac i i ies [50]
I consis s o 2 da ase s eco ded simul aneously by 2 Kinec sys ems including ADLs and alls in a
li ing oom equipped wi h a bed, a cupboa d, a chai and su ounding o ice objec s illumina ed by
neon lamps on he ceiling o by sunligh
Sdu all [12] 10 olun ee s de elop 6 ac i i ies eco ded by RGB-D sys ems
Falling de ec ion [38] 6 olun ee s pe o m 26 alls and simila ac i i ies eco ded by RGB-D sys ems.
Fall de ec ion da ase [30] 5 olun ee s execu e 5 di e en ypes o all
NTU RGB+ da ase [99]
I is a la ge-scale da ase o human ac ion ecogni ion.
I con ains 56,880 ac ion samples and includes 4 di e en modali ies o da a o each sample: RGB
ideos, dep h map sequences, 3D skele al da a and IR ideos
Syn he ic Mo emen Da abases CMU G aphics Lab—mo ion cap u e
lib a y [55]Lib a y ha cap u es syn he ic mo emen s h ough mo emen cap u e (MoCap) echnology
Senso s 2021,21, 947 23 o 50
4. Discussion
The s udied sys ems illus a e isual-based all de ec ion e olu ion in he las i e
yea s. These sys ems ollow a pa allel pa h o o he human ac i i y ecogni ion sys ems,
wi h inc easingly in ense use o a i icial neu al ne wo ks (ANN) and a clea endency
owa ds cloud compu ing sys ems, excep o he ones moun ed on obo s.
All s udied sys ems ollow, wi h nuances, a h ee-s ep app oach o all de ec ion
h ough a i icial ision.
The i s s ep, in oduced in Sec ion 4.1 and no always needed, includes ideo signal
p ep ocessing in o de o op imize i as much as possible.
Cha ac e iza ion is he second s ep, s udied in Sec ion 4.2, whe e image ea u es a e
abs ac ed, so wha happens in he images can be exp essed in he o m o desc ip o s ha
will be classi ied in he las s ep o he p ocess.
The hi d p ocess s ep, explained in Sec ion 4.3, in ends o ag he obse ed ac ions,
which main ea u es a e cha ac e ized by abs ac desc ip o s, as a all e en o one which
is no , so measu es can be aken o help he allen pe son as as as possible.
Some o he s udied sys ems ollow a ame-by- ame app oach whe e he sole sys em
goal is classi ying human pose as allen o no , lea ing aside he all mo ion i sel . Fo hose
sys ems ying o de e mine i a speci ic mo emen may be a all, silhoue e acking is a
basic suppo ope a ion de eloped h ough di e en p ocesses. T acking echniques used
by he s udied sys ems a e explained in Sec ion 4.4.
Finally, a compa ison in classi ying algo i hm pe o mance and alida ion da ase s is
p esen ed in Sec ions 4.5 and 4.6.
4.1. P ep ocessing
The inal objec i e o his phase is ei he dis o ion and noise educ ion o o ma adap-
a ion, so downs eam sys em blocks can ex ac cha ac e is ic ea u es wi h classi ica ion
pu poses. Image complexi y educ ion could also be an objec i e du ing he p ep ocessing
phase in some sys ems, so he compu a ional cos can be educed, o ideo s eaming
bandwid h use can be diminished.
The echniques g ouped in his Sec ion o dec easing noise a e nume ous and
ange om Gaussian smoo hing used in [
31
] o he mo phological ope a ions execu ed
in
[17,31,74]
o [
24
]. They a e in oduced in subsequen Sec ion as a pa o he o eg ound
segmen a ion p ocess.
Fo ma adap a ion p ocesses a e p esen in se e al o he s udied sys ems, as is he
case in [
48
], whe e images a e con e ed o g ayscale and ha e hei his og ams equalized
be o e being ans e ed o he cha ac e iza ion p ocess.
Image bina iza ion, as in [
89
], is also in oduced as a pa o he sys ema ic e o o
educe noise du ing he segmen a ion p ocess, while some o he sys ems, like he one p e-
sen ed in [
56
], pu sue image complexi y dec easing by ans o ming ideo signals om ed,
g een and blue (RGB) o black and whi e and hen applying a median il e , an algo i hm
which assigns new alues o image pixels based on he median o he su ounding ones.
Image complexi y educ ion is a goal pu sued by some sys ems, as he one p oposed
in [
91
], which in oduces comp essed sensing (CS), an algo i hm i s p oposed by Donoho
e al. [
107
] used in signal p ocessing o acqui e and econs uc a signal. Th ough his
echnique, signals, spa se in some domain, a e sampled a a es much lowe han equi ed
by he Nyquis –Shannon sampling heo em. The sys em uses a h ee-laye ed app oach
o CS by applying i o ideo signals, which allows p i acy p ese a ion and bandwid h
use educ ion. This echnique, howe e , in oduces noise and o e -smoo hs edges, espe-
cially hose in low con as egions, leading o in o ma ion loss and image low- esolu ion.
The e o e, image complexi y educ ion ea u e cha ac e iza ion o en becomes a challenge.
4.2. Cha ac e iza ion
The second p ocess s ep in ends o exp ess human pose and/o human mo ion as
abs ac ea u es in a quali a i e app oach, o quan i y hei in ensi y in an ul e io quan i y
Senso s 2021,21, 947 24 o 50
app oach. These quan i ied ea u es a e hen used wi h classi ying pu poses in he las s ep
o he all de ec ion sys em.
These abs ac pose/ac ion desc ip o s can globally be classi ied in o h ee main
g oups: global, local and dep h.
Global desc ip o s analyze images as a block, segmen ing o eg ound om back-
g ound, ex ac ing desc ip o s ha de ine i and encoding hem as a whole.
Local desc ip o s app oach he abs ac ion p oblem om a di e en pe spec i e and, in-
s ead o segmen ing he block o in e es , p ocess he images as a collec ion o
local desc ip o s
.
Dep h cha ac e iza ion is an al e na i e way o de ine desc ip o s om images con-
aining dep h in o ma ion by ei he using dep h maps o skele on da a ex ac ed om a
join acking p ocess.
4.2.1. Global
Global desc ip o s y o ex ac abs ac in o ma ion om he o eg ound once i has
been segmen ed om he backg ound and encode i as a whole.
This kind o ac i i y desc ip o s was e y commonly used in a i icial ision ap-
p oaches o human ac i i y ecogni ion in gene al and o all de ec ion in pa icula . How-
e e , o e ime, hey ha e been displaced by local desc ip o s o used in combina ion wi h
hem, as hese ones a e less sensi i e o noise, occlusions and iewpoin changes.
Fo eg ound segmen a ion is execu ed in a numbe o di e en ways. Some app oaches
o his concep es ablish a speci ic backg ound and sub ac i om he o iginal image;
some o he s loca e egions o in e es by iden i ying he silhoue e edges o use he op ical
low, gene a ed as a consequence o body mo emen s, as a desc ip o . Some global cha ac-
e iza ion me hods segmen he human silhoue e o e ime o o m a space– ime olume
which cha ac e izes he mo emen . Some o he me hods ex ac ea u es om images in
a di ec way, as in he case o he sys em desc ibed in [
48
], whe e e e y h ee ames, he
mean squa e e o (MSE) is de e mined and used as an indica o o image simila i y.
Silhoue e Segmen a ion
Human shape segmen a ion can be execu ed h ough a numbe o echniques, bu
all o hem equi e backg ound iden i ica ion and sub ac ion. This p ocess, known as
backg ound ex ac ion, is p obably he mos isually in ui i e one, as i s p oduc is a
human silhoue e.
Backg ound es ima ion is he mos impo an s ep o he p ocess, and i is add essed
in di e en ways.
In [
17
,
24
,
56
,
74
], as he backg ound is supposed cons an , an image o i is aken
du ing sys em ini ializa ion, and a di ec compa ison allows segmen a ion o any new
objec p esen in he ideo. This echnique is easy and powe ul; howe e , i is ex emely
sensi i e o ligh changes. To mi iga e his law, he sys em desc ibed in [
31
], whe e he
backg ound is also supposed s able, a median h oughou ime is calcula ed o e e y pixel
posi ion in e e y colo channel. Then, i is di ec ly sub ac ed om he obse ed image
ame-by- ame.
Despi e e e y hing, he ob ained p oduc s ill con ains a subs an ial amoun o noise
associa ed wi h shadows and illumina ion. To educe i , mo phological ope a o s can be
used as in [
17
,
24
,
31
,
74
]. Dila ion and/o e osion ope a ions a e pe o med by p obing
he image a all possible places wi h a s uc u ing elemen . In he dila ion ope a ion, his
elemen wo ks as a local maximum il e and, he e o e, adds a laye o pixels o bo h inne
and ou e bounda y a eas. In e osion ope a ions, he elemen wo ks as a local minimum
il e and, as a consequence, s ips away a laye o pixels om bo h egions. Noise educ ion
a e segmen a ion can also be pe o med h ough Kalman il e ing, as in [
92
], whe e his
il e ing me hod is success ully used wi h his pu pose.
An al e na i e op ion o backg ound es ima ion and sub ac ion is he applica ion o
Gaussian mix u e models (GMM), a echnique used in [
7
,
11
,
14
,
78
,
92
], among o he s, ha
models he alues associa ed wi h speci ic pixels as a mix o Gaussian dis ibu ions.
Senso s 2021,21, 947 25 o 50
A di e en app oach is used in [
6
], whe e he Ho p ase me hod [
108
] is applied o
backg ound sub ac ion. I uses a compu a ional colo model ha sepa a es he b igh ness
om he ch oma ici y componen . By doing i , i is possible o segmen he o eg ound
much mo e e icien ly when ligh dis u bances a e p esen han wi h p e ious me hods,
diminishing his way ligh change sensi i eness. In his pa icula sys em, pixels a e also
clus e ed by simila i y, so compu a ional complexi y can be educed.
Some sys ems, like he one p esen ed in [
7
], apply a il e o de e mine silhoue e
con ou s. In his pa icula case, a Sobel il e is used, which de e mines a wo-dimensional
g adien o e e y image pixel.
O he segmen a ion me hods, like ibe [
19
], used in [
22
,
94
], s o e, associa ed wi h
speci ic pixels, p e ious alues o he pixel i sel and i s icini y o de e mine whe he i s
cu en alue should be ca ego ized as o eg ound o backg ound. Then, he backg ound
model is adap ed by andomly choosing which alues should be subs i u ed and which no ,
a clea ly di e en pe spec i e om o he echniques, which gi e p e e ence o new alues.
On op o ha , pixel alues decla ed as backg ound a e p opaga ed in o neighbo ing pixels
pa o he backg ound model.
The sys em in [
8
] segmen s he o eg ound using he echnique p oposed in [
109
],
whe e he op ical low (OF), which a e p esen ed in la e Sec ions, is calcula ed o de e mine
wha objec s a e in mo ion in he image, ea u e used o o eg ound segmen a ion. In
a subsequen s ep, o educe noise, images a e bina ized and mo phological ope a o s
a e applied. Finally, he poin s ma king he cen e o he head and he ee a e linked
by lines composing a iangle whose a ea/heigh a io will be used as he cha ac e is ic
classi ica ion ea u e.
Some algo i hms, like he illumina ion change- esis an independen componen
analysis (ICA), p oposed in [
95
], combine ea u es o di e en segmen a ion echniques,
like GMM and sel -o ganizing maps, a well-known g oup o ANN able o classi y in o
low dimensional classes e y high dimensional ec o s, o o e come he p oblems o
silhoue e segmen a ion associa ed wi h illumina ion phenomena. This algo i hm is able o
success ully ackle segmen a ion e o s associa ed wi h sudden illumina ion changes due
o any kind o ligh sou ce, bo h in images aken wi h omnidi ec ional diop ic came as
and in plain ones.
ICA and ibe a e compa ed in [
94
] by using a da ase speci ically de eloped o ha
sys em wi h be e esul s o he ICA algo i hm.
In [
9
], o eg ound ex ac ion is execu ed in acco dance wi h he p ocedu e desc ibed
in [
110
]. This me hod in eg a es he egion-based in o ma ion on colo and b igh ness in a
codewo d, and he collec ion o all codewo ds a e g ouped in an en i y called codebook.
Pixels a e hen checked in e e y single new ame and, when i s colo o b igh ness does
no ma ch he egion codewo d, which encodes a ea b igh ness and colo bands, i is
decla ed as o eg ound. O he wise, he codewo d is upda ed, and he pixel is decla ed
as a ea backg ound. Once pixels a e agged as o eg ound, hey a e clus e ed oge he ,
and codebooks a e upda ed o each one o hem. Finally, hese egions a e app oxima ed
by polygons.
Some sys ems, like he one in [
9
], use o hogonal came as and use o eg ound maps by
using homog aphy. This way, noise associa ed wi h illumina ion a ia ions and occlusion
is g ea ly educed. The sys em also calcula es he obse ed polygon a ea/g ound p ojec ed
polygon a ea a e as he main ea u e o de e mine whe he a all e en has aken place.
Sel -o ganizing maps is a echnique, well desc ibed in [
111
], used wi h segmen a ion
pu poses in [
58
]. When applied, ini ial backg ound es ima ion is made based on he i s
ame a sys em s a up. E e y pixel o his ini ial image is associa ed wi h a neu on in an
ANN h ough a weigh . Those weigh s a e cons an ly upda ed as new ames low in o he
sys em and, he e o e, he backg ound model changes. Sel -o ganizing maps ha e been
success ully used o sub ac o eg ound om backg ound, and hey ha e p o ed a good
esilience o he ligh a ia ion noise.
Senso s 2021,21, 947 32 o 50
pendulum o a ion ene gy and i s gene alized o ce sequences. These ea u es a e hen
codi ied in a ec o and used o classi ica ion pu poses.
The sys em in [
105
] uses se e al ANNs and selec s he mos sui able one as a unc ion
o he en i onmen and he cha ac e is ics o he acked people. In addi ion, i uploads
w ongly ca ego ized images which a e used o e ain he used models.
4.2.3. Dep h
Desc ip o s based on dep h in o ma ion ha e gained g ound hanks o he de elop-
men o low-cos dep h senso s, such as Mic oso Kinec
®
. This a o dable sys em coun s
wi h a so wa e de elopmen ki (SDK) and applica ions able o de ec and ack join s and
cons uc human body ec o models. These elemen s, oge he wi h he dep h in o ma-
ion om s e eoscopic scene obse a ion, ha e aised g ea in e es among he a i icial
ision esea ch communi y in gene al and he human all de ec ion sys em de elope s
in pa icula .
A good numbe o he s udied sys ems use dep h in o ma ion, solely o oge he wi h
RGB one, as he da a sou ce in he abs ac ion p ocess leading o image desc ip o cons uc-
ion. These sys ems ha e p o ed o be able o segmen o eg ound, g ea ly diminishing
in e e ence due o illumina ion in e e ences up o he dis ance whe e s e eoscopic ision
p ocedu es a e able o in e dep h da a. Fall de ec ion sys ems use his in o ma ion ei he
as dep h maps o skele on ec o models.
Dep h Map Rep esen a ion
Dep h maps, unlike RGB ideo signals, con ain di ec h ee-dimensional in o ma ion
on objec s in he image. The e o e, dep h map ideo signals in eg a e aw 3D in o ma ion,
so h ee-dimensional cha ac e iza ion ea u es can be di ec ly ex ac ed om hem.
This way, he sys em desc ibed in [
46
] iden i ies 16 egions o he human body ma ked
wi h ed ape and posi ion hem in space h ough s e eoscopic echniques. Taking ha
in o ma ion as a base, he sys em builds he body ec o (aligned wi h spine o ien a ion)
and iden i ies i s cen e o g a i y (CG). Accele a ion o CG and body ec o angle on a
e ical axis will be used as ea u es o classi ica ion.
Fo eg ound segmen a ion o human silhoue e is made by hese sys ems h ough
dep h in o ma ion, by compa ing dep h da a om images and a e e ence es ablished a
sys em s a up. This way, pixels appea ing in an image a a dis ance di e en om he
one s o ed o ha pa icula pixel in he e e ence a e decla ed as o eg ound. This is he
p ocess ollowed by [44] o segmen he human silhoue e. In an ul e io s ep, desc ip o s
based on bounding box, cen oid, a ea and o ien a ion o he silhoue e a e ex ac ed.
O he sys ems, like he one in [
101
], ex ac backg ound by using he same p ocess
and he silhoue e is de e mined as he majo connec ed body in he esul ing image. Then,
an ellipse is es ablished a ound i , and classi ica ion will be made as a unc ion o i s aspec
a io and cen oid posi ion. A simila p ocess is ollowed in [
60
], whe e, a e backg ound
sub ac ion, an ellipse is es ablished a ound he silhoue e, and i s cen oid ele a ion and
eloci y, as well as i s aspec a io, a e used as classi ica ion ea u es.
The sys em in [
57
] uses dep h maps o segmen silhoue es as well and c ea es a
bounding box a ound hem. Box op coo dina es a e used o de e mine he head eloci y
p o ile du ing a all e en , and i s Hausdo dis ance o head ajec o ies eco ded du ing
eal all e en s is used o de e mine whe he a all has aken place. The Hausdo dis ance
quan i ies how a wo subse s o a me ic space a e om each o he . The no el y o his
sys em, lea ing aside he in oduc ion o he Hausdo dis ance as desc ibed in [
130
], is
he use o a mo ing cap u e (MoCap) echnique o d i e a human model using so wa e
o simula e i s mo ion (OpenSim), so p o iles o head e ical eloci ies can be cap u ed
in ADLs, and a da abase can be buil . This da abase is used, by he in oduc ion o he
Hausdo dis ance, o assess alls.
The sys em in [
85
], a e o eg ound ex ac ion by using dep h in o ma ion as in he
p e ious sys ems, ans o ms he image o a black and whi e o ma and, a e de-noising i

Senso s 2021,21, 947 33 o 50
h ough il e ing, calcula es he HOG. To do i , he sys em de e mines he g adien ec o
and i s di ec ion o each image pixel. Then, a his og am is cons uc ed, which in eg a es
all pixels’ in o ma ion. This is he ea u e used o classi ica ion pu poses.
In [
42
], silhoue es a e acked by using a p opo ional-in eg al-di e en ial (PID) con-
olle . A bounding box is c ea ed a ound he silhoue e, and ea u es a e ex ac ed in
acco dance wi h [
131
]. A all will be called i h esholds es ablished o ea u es a e ex-
ceeded. Faces a e sea ched, and when iden i ied, he acking will be biased
owa ds hem.
Some o he sys ems, like he one in [
15
], sub ac s backg ound by di ec use o dep h
in o ma ion con ained in sequen ial images, so he di e ence be ween consecu i e dep h
ames is used o segmen a ion. Then, he head is acked, so he head e ical posi-
ion/pe son heigh a io can be de e mined, which, oge he wi h CG eloci y, is used as a
classi ica ion ea u e.
In [
54
], all backg ound is se o a ixed dep h dis ance. Then, a g oup o 2000 body
pixels is andomly chosen, and o each o hem, a ec o o 2000 alues, calcula ed as
a unc ion o he dep h di e ence be ween pai s o poin s, is c ea ed. These pai s a e
de e mined by es ablishing 2000 pixel o se se s. The ob ained 2000- alue ec o is used as
a cha ac e is ic ea u e o pose classi ica ion.
The sys em in oduced in [
11
], a e he human silhoue e is segmen ed by using
dep h in o ma ion h ough a GMM p ocess, calcula es i s cu a u e scale space (CSS)
ea u es by using he p ocedu es desc ibed in [
12
]. CSS calcula ion me hod con olu es a
pa ame ic ep esen a ion o a plana cu e, silhoue e edge in his case, wi h a Gaussian
unc ion. This way, a ep esen a ion o he a c leng h s. cu a u e is ob ained. Then,
silhoue es ea u es a e encoded, oge he wi h he Gaussian mix u e model used in he
a o emen ioned CSS p ocess, in a single Fishe ec o , which will be used, a e being
no malized, o classi ica ion pu poses.
Finally, a block o sys ems c ea es olumes based on no mal dis ibu ions cons uc ed
a ound poin clouds. These dis ibu ions, called oxels, a e g ouped oge he , and desc ip-
o s a e ex ac ed ou o oxel clus e s o de e mine, i s , whe he hey ep esen a human
body and hen o assess i i is in a allen s a e.
This way, he sys em p esen ed in [
27
] i s es ima es he g ound plane by assuming
ha mos o he pixels belonging o e e y ho izon al line a e pa o he g ound plane.
The g ound can hen be es ima ed, line pe line, a ending o he pixel dep h alues as
explained in he p ocedu e desc ibed in [
132
]. To clean up he pic u es, all pixels below he
g ound plane a e disca ded. Then, no mal dis ibu ions ans o m (NDT) maps a e c ea ed
as a cloud o poin s su ounded by no mal dis ibu ions wi h he physical appea ance o
an ellipsoid. These dis ibu ions, c ea ed a ound a minimum numbe o poin s, a e called
oxels and, in his sys em, a e gi en ixed dimensions. Then, ea u es ha desc ibe he local
cu a u e and shape o he local neighbo hood a e ex ac ed om he dis ibu ions. These
ea u es, known as IRON [
133
], allow oxel classi ica ion as being pa o a human body
o no and, his way, oxels agged as human a e clus e ed oge he . IRON ea u es a e
hen calcula ed o he clus e ep esen ing a human body, and he Mahalanobis dis ance
be ween ha ec o and he dis ibu ion associa ed wi h allen bodies is calcula ed. I he
dis ance is below a h eshold, he all s a e is decla ed.
A simila p ocess is used in [
34
], whe e, a e he poin cloud is unca ed by emo ing
all poin s no con ained in he a ea in be ween he g ound plane and a pa allel one 0.7 m
o e i by applying he RANSAC p ocedu e [
134
], NDTs a e c ea ed and hen segmen ed
in pa ches o equal dimensions. A suppo ec o machine (SVM) classi ie de e mines
which ones o hose pa ches belong o a human body as a unc ion o hei geome ic
cha ac e is ics. Close pa ches agged as humans a e clus e ed, and a bounding box is
c ea ed a ound. A second SVM de e mines whe he clus e s should be decla ed as a allen
pe son. This classi ica ion is e ined, aking da a om a da abase o obs acles o he a ea,
so i he clus e is decla ed as a allen pe son, bu i is con ained in he obs acle da abase,
he decla a ion is skipped.
Senso s 2021,21, 947 34 o 50
Skele on Rep esen a ion
Sys ems implemen ing his ep esen a ion a e able o de ec and ack join s and, based
on ha in o ma ion, hey can build a human body ec o model. This block o echniques,
as he p e ious one, s ongly diminishes he noise associa ed wi h illumina ion bu ha e
p oblems o build a co ec model when occlusion appea s, bo h he one gene a ed by
obs acles and he one p oduc o pe spec i e au o-occlusions.
A good numbe o hese sys ems a e buil o e he Mic oso Kinec
®
sys em and
ake ad an age o bo h de SDK and he applica ions de eloped o i . This is he case
o he sys em in oduced in [
40
], whe e h ee Kinec
®
sys ems co e he same a ea om
di e en pe spec i es, and join s a e, he e o e, ollowed om di e en angles, educing
his way he acking p oblems associa ed wi h occlusion. In his sys em, human mo emen
is cha ac e ized h ough wo main ea u es, head speed and CG si ua ion e e enced o
ankles posi ion.
The Kinec
®
sys em is also used in [
65
] o ollow join s and es ima e he e ical
dis ance o he g ound plane. Then, he angle be ween he e ical and he o so ec o ,
which links he neck and spine base, is de e mined and used o iden i y a s a key ame
(SKF), whe e a all s a s, and an end key ame (EKF), whe e i ends. Du ing his pe iod,
e ical dis ance o he g ound plane and e ical eloci y o ollowed uppe join s will be
he inpu o classi ica ion. A e y simila app oach is ollowed in [
33
], whe e o so/ e ical
angle and cen oid heigh a e he key ea u es used o classi ica ion.
This sys em is used as well in [
5
] o build, a ound iden i ied join s, bo h 2D and
3D bounding boxes aligned wi h he spine di ec ion. Then, he a io wid h/heigh is
de e mined, and he ela ion H
CG
/P
CG
, being he o me de ele a ion o he CG o e he
g ound plane and he la e de dis ance be ween he CG p ojec ion on he g ound and he
suppo polygon de ined by ankles posi ion, is calcula ed. Those ea u es will be he base
o e en classi ica ion.
In [
135
], human body key poin s a e iden i ied by a CNN whose inpu is a 2D RGB
ideo signal complemen ed by dep h in o ma ion. Based on hose key poin s, he sys em
builds a human body ec o model. A il e was de eloped o gene a e digi al e ain
models om da a cap u ed by ai bo ne sys ems [
136
], and he dep h da a we e hen used
o es ima e he g ound plane. The sys em uses all ha in o ma ion o calcula e he dis ance
om he body CG and he body egion o e he shoulde s o he g ound. These dis ances
will se e o cha ac e ize he human pose.
A CNN is also used in [
61
] o gene a e ea u e maps ou o he dep h images. This
ne wo k s acks con olu ion laye s o ex ac ea u es and pooling laye s o educe map
complexi y, wi h a philosophy iden ical o he one used in he RGB local cha ac e iza ion.
The ou pu map goes h ough wo laye s o ully connec ed laye s o classi y he eco ded
ac i i y, and a So max unc ion is implemen ed in he las laye o he ANN, which
de e mines whe he a all has aken place.
In [84], p io o inpu images in a CNN o gene a e ea u e maps, which will be used
o classi ica ion, he backg ound is sub ac ed h ough an algo i hm ha combines dep h
maps and 2D images o enhance segmen a ion pe o mance. This way, i he pixels o he
segmen ed 2D silhoue e expe imen sha p changes, bu pixels in he dep h map do no ,
pixels subjec o hose changes a e ega ded as noise. The sys em mixes in o ma ion om
bo h sou ces, allowing a be e ack on segmen ed silhoue es and a quick ack egain in
case i is los .
The sys em in [
13
]—a e iden i ying human body join s as he key ea u es whose
ajec o y will be used o de e mine whe he a alling e en has aken place—p oposes
o a ing he o so so i is always e ical. This way, join ex ac ion becomes pose in a ian , a
echnique used in he sys em wi h posi i e esul s in o de o deal wi h he noise associa ed
wi h join iden i ica ion as a esul o apid mo emen and occlusion, cha ac e is ic o alls.
Senso s 2021,21, 947 35 o 50
4.3. Classi ica ion
Once pose/mo emen abs ac desc ip o s ha e been ex ac ed om ideo images,
he nex s ep o he all de ec ion p ocess is classi ica ion. In b oad e ms, du ing his phase,
he sys em classi ies mo emen and o pose as a all o a allen s a e h ough an algo i hm
ha is pa o one o hese wo ca ego ies; gene a i e o disc imina i e models.
Disc imina i e models a e able o de e mine bounda ies be ween classes, ei he by
explici ly being gi en hose bounda ies o by se ing hem hemsel es using se s o p e-
classi ied desc ip o s.
Gene a i e models app oach he classi ica ion p oblem in a o ally di e en way, as
hey explici ly model he dis ibu ion o each class and hen use he Bayes heo em o link
desc ip o s o he mos likely class, which, in his case, can only be a all o a no all s a e.
4.3.1. Disc imina i e Models
The inal goal o any classi ie is assigning a class o a gi en se o desc ip o s. The
disc imina i e models a e able o es ablish he bounda ies sepa a ing classes, so he p oba-
bili y o a desc ip o belonging o a speci ic class can be gi en. In o he e ms, gi en
α
as a
class, and [A] as he ma ix o desc ip o alues associa ed wi h a pose o mo emen , his
amily o classi ie s is able o de e mine he p obabili y P(α|[A]).
Fea u e-Th eshold-Based
Fea u e- h eshold-based classi ica ion models a e b oadly used in he s udied sys ems.
This app oach is easy and in ui i e, as he esea che es ablishes h eshold alues o he
desc ip o s, so hei associa ed e en s can be assigned o a speci ic class in case hose
h esholds a e exceeded.
This is he case o he sys em p oposed in [
31
]. I classi ies he ac ion as a all o a
non- all in acco dance wi h a double a ionale. On one hand, i es ablishes h esholds
o ellipse ea u es o es ima e whe he he pose i s a allen s a e; on he o he , an MHI
ea u e exceeding a ce ain alue indica es a as mo emen and, he e o e, a po en ial all.
The sys em p oposed in [
14
] adds accele a ion o he o me ea u es and, in [
40
], head
speed o e a ce ain h eshold and CG posi ion ou o he segmen de ined by ankles a e
indica i es o a all.
Simila app oaches, whe e h eshold alues a e de e mined by sys em de elope s
based on p e ious expe imen a ion, a e implemen ed in a good numbe o he s udied
sys ems, as hey a e simple, in ui i e and compu a ionally inexpensi e.
Mul i a ia e Exponen ially Weigh ed Mo ing A e age
Mul i a ia e exponen ially weigh ed mo ing a e age (MEWMA) is a s a is ical p ocess
con ol o moni o a iables ha use he en i e his o y o alues o a se o a iables. This
echnique allows designe s o gi e a weigh ing alue o all eco ded a iable ou pu s, so
he mos ecen ones a e gi en highe weigh alues, and he olde ones a e weigh ed
ligh e . This way, he las alue is weigh ed
λ
(being
λ
a numbe be ween 0 and 1) and
p e ious
β
alues a e weigh ed
λβ
. Limi s o he alue o ha weigh ed ou pu a e
es ablished, aking as a basis he expec ed mean and s anda d de ia ion o he p ocess.
Ce ain sys ems, like [
28
], use his echnique o classi ica ion pu poses. Howe e , as i is
unable o dis inguish be ween alling e en s and o he simila ones, e en s agged as all
by he MEWMA classi ie need o go h ough an ul e io suppo ec o machine classi ie .
Suppo Vec o Machines
Suppo ec o machines (SVM) a e a se o supe ised lea ning algo i hms i s
in oduced by Vapnik e al. [137].
SVMs a e used o eg ession and classi ica ion p oblems. They c ea e hype planes in
high dimension spaces ha sepa a e classes nonlinea ly. To ul ill his ask, SVMs, simila
o a i icial neu al ne wo ks, use ke nel unc ions o di e en ypes.
A s anda d SVM bounda y de ini ion is shown in Figu e 4.
Senso s 2021,21, 947 36 o 50
Senso s 2021, 21, x FOR PEER REVIEW 36 o 50
Figu e 4. Suppo ec o machine bounda y de ini ion.
In [74], linea , polynomial, and adial ke nels a e used o ob ain he hype planes; in
[67], adial ones a e implemen ed, and in [48], polynomial ke nels a e used o achie e
nonlinea classi ica ions.
The suppo ec o da a desc ip ion (SVDD), in oduced by Tax e al. [138], is a clas-
si ying algo i hm inspi ed by he suppo ec o machine classi ie , able o ob ain a sphe -
ically shaped bounda y a ound a da ase and, analogously o SVMs, i can use di e en
ke nel unc ions. The me hod is made obus agains ou lie s in he aining se and is
capable o igh ening classi ica ion by using nega i e examples. SVDDs classi ying algo-
i hms a e used in [90].
SVMs ha e been e y used in he s udied sys ems as hey ha e p oo ed o be e y
e ec i e; howe e , hey equi e high compu a ional loads, some hing inapp op ia e o
edge compu ing sys ems.
K-Nea es Neighbo
K-nea es neighbo (KNN) is an algo i hm able o model he condi ional p obabili y
o a sample belonging o a speci ic class. I is used o classi ica ion pu poses in
[16,17,48,74] among o he s.
KNNs assume ha classi ica ion can be success ully made based on he class o he
nea es neighbo s. This way, i o a speci ic ea u e, all µ closes sample neighbo s a e
pa o a de e mined class, he p obabili y o he sample being pa o ha class will be
assessed as e y high. This s udy is epea ed o e e y ea u e con ained in he desc ip o ,
so a inal assessmen based on all ea u es can be made. The algo i hm usually gi es di -
e en weigh s o he neighbo s, and hea ie weigh s a e assigned o he closes ones. On
op o ha , i also assigns di e en weigh s o e e y ea u e. This way, he ones assessed
as mos ele an ge hea ie weigh s.
Decision T ee
Decision ees (DT) a e algo i hms used bo h in eg ession and classi ica ion. I is an
in ui i e ool o make decisions and explici ly ep esen s decision-making. Classi ica ion
DTs use ca ego ical a iables associa ed wi h classes. T ees a e buil by using lea es,
which ep esen class labels, and b anches, which ep esen cha ac e is ic ea u es o hose
Figu e 4. Suppo ec o machine bounda y de ini ion.
In [
74
], linea , polynomial, and adial ke nels a e used o ob ain he hype planes;
in [
67
], adial ones a e implemen ed, and in [
48
], polynomial ke nels a e used o achie e
nonlinea classi ica ions.
The suppo ec o da a desc ip ion (SVDD), in oduced by Tax e al. [
138
], is a
classi ying algo i hm inspi ed by he suppo ec o machine classi ie , able o ob ain
a sphe ically shaped bounda y a ound a da ase and, analogously o SVMs, i can use
di e en ke nel unc ions. The me hod is made obus agains ou lie s in he aining se
and is capable o igh ening classi ica ion by using nega i e examples. SVDDs classi ying
algo i hms a e used in [90].
SVMs ha e been e y used in he s udied sys ems as hey ha e p oo ed o be e y
e ec i e; howe e , hey equi e high compu a ional loads, some hing inapp op ia e o
edge compu ing sys ems.
K-Nea es Neighbo
K-nea es neighbo (KNN) is an algo i hm able o model he condi ional p obabili y o
a sample belonging o a speci ic class. I is used o classi ica ion pu poses in [
16
,
17
,
48
,
74
]
among o he s.
KNNs assume ha classi ica ion can be success ully made based on he class o he
nea es neighbo s. This way, i o a speci ic ea u e, all
µ
closes sample neighbo s a e pa
o a de e mined class, he p obabili y o he sample being pa o ha class will be assessed
as e y high. This s udy is epea ed o e e y ea u e con ained in he desc ip o , so a
inal assessmen based on all ea u es can be made. The algo i hm usually gi es di e en
weigh s o he neighbo s, and hea ie weigh s a e assigned o he closes ones. On op o
ha , i also assigns di e en weigh s o e e y ea u e. This way, he ones assessed as mos
ele an ge hea ie weigh s.
Decision T ee
Decision ees (DT) a e algo i hms used bo h in eg ession and classi ica ion. I is an
in ui i e ool o make decisions and explici ly ep esen s decision-making. Classi ica ion
DTs use ca ego ical a iables associa ed wi h classes. T ees a e buil by using lea es,
which ep esen class labels, and b anches, which ep esen cha ac e is ic ea u es o hose
classes. DTs buil p ocess is i e a i e, wi h a selec ion o ea u es co ec ly o de ed o
de e mine he spli poin s ha minimize a cos unc ion ha measu es he compu a ional
Senso s 2021,21, 947 37 o 50
equi emen s o he algo i hm. These algo i hms a e p one o o e i ing, as se ing he
co ec numbe o b anches pe lea is usually e y challenging. To educe he complexi y
o he ees, and he e o e, hei compu a ional cos , b anches a e p uned when he ela ion
cos -sa ing/accu acy loss is sa is ac o y. This ype o classi ie is used in [87,89].
Random o es (RF), like he one used in [
54
,
87
], is an agg ega ion echnique o DT,
in oduced by B aiman [
139
], which main objec i e is a oiding o e i ing. To accomplish
his ask, he aining da ase is di ided in o subg oups, and he e o e, a inal numbe o
DTs, equal o he numbe o da ase subg oups, is ob ained. All o hem a e used in he
p ocess, so he inal classi ica ion decision is ac ually a combina ion o he classi ica ion o
all DTs.
G adien boos ing decision ees (GBDT) is ano he DT agg ega ion echnique whose
algo i hm was i s in oduced by F iedman [
140
] whe e simple DTs a e buil and, o each
one o hem, a classi ica ion e o in aining ime is de e mined. An e o unc ion based
on calcula ed indi idual e o s is de e mined, and i s g adien is minimized by combining
indi idual DT classi ica ions in a p ope way. This agg ega ion echnique, speci ically
de eloped o DTs, is ac ually pa o a b oade amily ha will be mo e ex ensi ely
p esen ed in he nex sec ion.
Bo h echniques, RF and GBDT, a e used in [87].
Boos Classi ie
Boos classi ie algo i hms a e a amily o classi ie building echniques ha c ea e
s ong classi ie s by g ouping weak ones. I is done by adding up models c ea ed om he
aining da a un il he sys em is pe ec ly p edic ed o a maximum numbe o models is
eached.
This is done by building a model om he aining da a. Then, a second model is
c ea ed o co ec he e o s om he i s one. Models a e added un il he aining se is
well p edic ed o a maximum numbe o hem is added. Du ing he boos ing p ocess, he
i s model is ained on he en i e da abase while he es a e i ed o he esiduals o he
p e ious ones.
Adaboos , used in [
23
], can be u ilized o inc ease pe o mances wi h any classi ica ion
echnique, bu i is mos commonly used wi h one-le el decision ees.
In [
64
], boos ing echniques a e used on a J48 algo i hm, a ee-based echnique, simila
o andom o es , which is used o c ea e uni a ia e decision ees.
Spa se Rep esen a ion Classi ie
Spa se ep esen a ions classi ica ion (SRC) is a echnique used o image classi ica ion
wi h a e y good deg ee o pe o mance.
Na u al images a e usually ich in ex u e and o he s uc u es ha end o be ecu en .
Fo his eason, spa se ep esen a ion can be success ully applied o image p ocessing. This
phenomenon is known as pa ch ecu ence and, because o i , eal-wo ld digi al images
can be ecognized by p ope ly ained dic iona ies.
SRCs a e able o ecognize hose pa ches, as hey can be exp essed as a linea combi-
na ion o a limi ed numbe o elemen s ha a e con ained in he classi ie dic iona ies.
This is he case o he SRC p esen ed in [24].
Logis ic Reg ession
Logis ic eg ession is a s a is ical model used o classi ica ion. I is able o implemen
a bina y classi ie , like he one needed o decide whe he a all e en has aken place. Fo
such a pu pose, a logis ic unc ion is used. I can be adjus ed by using classi ying ea u es
associa ed wi h e en s agged as all o no all.
This me hod is used in sys ems like [
93
], whe e a logis ic classi ying algo i hm is
employed o classi y e en s as all o no a all, based on a ec o ha encodes he empo al
se ies o o a ion ene gy and gene alized o ce.

Senso s 2021,21, 947 38 o 50
Some a i icial neu al ne wo ks implemen a logis ic eg ession unc ion o classi i-
ca ion, like he one desc ibed in [
106
], whe e a CNN uses his unc ion o de e mine he
de ec ion p obabili y o each de ined class.
Deep Lea ning Models
In [
83
], he las laye s o he ANN implemen a So max unc ion, a gene aliza ion
o he logis ic unc ion used o mul inomial logis ic eg ession. This unc ion is used as
he ac i a ion unc ion o he nodes o he las laye o a neu al ne wo k, so i s ou pu is
no malized o a p obabili y dis ibu ion o e he di e en ou pu classes. So max is also
implemen ed in he las laye s o he a i icial neu al ne wo ks used in [
75
,
103
], among
o he s udied sys ems.
Mul ilaye pe cep on (MLP) is a ype o mul ilaye ed ANN wi h hidden laye s
be ween he en ance and he exi ones able o so ou classes non linea ly sepa able. Each
node o his ne wo k is a neu on ha uses a nonlinea ac i a ion unc ion, and i is used o
classi ica ion pu poses in [48,87].
Radial basis unc ion neu al ne wo ks (RBFNN) a e used in he las laye o [
89
] o
classi y he ea u e ec o s coming om p e ious CNN laye s. This ANN is cha ac e ized
by using adial basis unc ions as ac i a ion unc ions and yields be e gene aliza ion
capabili ies han o he a chi ec u es, such as So max, as i is ained ia minimizing he
gene alized e o es ima ed by a localized-gene aliza ion e o model (L-GEM).
O en, he las laye s o ANN a chi ec u es a e ully connec ed ones, as in [
58
,
76
,
86
],
whe e all nodes o a laye a e connec ed o all nodes in he nex one. In hese s uc u es,
he inpu laye is used o la en ou pu s om p e ious laye s and ans o m hem in o a
single ec o , while subsequen laye s apply weigh s o de e mine a p ope agging and,
he e o e, success ully classi y e en s.
Finally, ano he ANN s uc u e use ul o classi ica ion is he au oencode one, used
in [
70
]. Au oencode s a e ANNs ained o gene a e ou pu s equal o inpu s. I s in e nal
s uc u e includes a hidden laye whe e all neu ons a e connec ed o e e y inpu and ou pu
node. This way, au oencode s ge high dimensional ec o s and encode hei ea u es.
Then, hese ea u es a e decoded back. As he numbe o dimensions o he ou pu ec o
may be educed, his kind o ANNs can be used o classi ica ion pu poses by educing he
numbe o ou pu dimensions o he numbe o inal expec ed classes.
4.3.2. Gene a i e Models
The app oach o gene a i e models o he classi ica ion p oblem is comple ely di e en
om he one ollowed by he disc imina i e ones.
Gene a i e models explici ly model he dis ibu ion o each class. This way, gi en
α
as a class, and [A] as he ma ix o desc ip o alues associa ed wi h a pose o mo emen ,
i bo h P ([A]|
α
) and P (
α
) can be de e mined, i will be possible, by di ec applica ion o
he Bayes heo em, o ob ain P (α|[A]), which will sol e he classi ica ion p oblem.
Hidden Ma ko Model
Classi ica ion using he hidden Ma ko model (HMM) algo i hm is one o he h ee
ypical p oblems ha can be sol ed h ough his p ocedu e. I was i s p oposed wi h
his pu pose by Rabine e al. [
141
] o sol e he speech ecogni ion p oblem, and i is used
in [100] o classi y he ea u e ec o s associa ed wi h a silhoue e.
HMMs a e s ochas ic models used o ep esen sys ems whose s a e a iables change
andomly o e ime. Unlike o he s a is ical p ocedu es, like Ma ko chains, which deal
wi h ully obse able sys ems, HMMs ackle pa ially obse able sys ems. This way, he
inal objec i e o he HMM classi ying p oblem esolu ion will be decided, on he basis o
he obse able da a ( ea u e ec o ), whe he a all has occu ed (hidden sys em s a e).
The sys em p oposed in [
100
] de e mines, using an HMM as a classi ie , on he basis
o silhoue e su ace, cen oid posi ion and bounding box aspec a io, whe he a all akes
place o no . To do i , and o ake as a e e ence eco ded alls, a p obabili y is assigned o
Senso s 2021,21, 947 39 o 50
he wo possible sys em s a es ( all/no all) based on alue and a ia ion along he e en
ime ame pe iod o he ea u e ec o . This classi ying echnique is used wi h success in
his sys em, hough in [
142
], a b ie summa y o he nume ous limi a ions o his basic
HMM app oach is p esen ed, and se e al mo e e icien ex ensions o he algo i hm, such
as a iable ansi ion HMM o he hidden semi-Ma ko model, a e in oduced. These
algo i hm a ia ions a e de eloped as he basic HMM p ocess is conside ed ill-sui ed o
modeling sys ems whe e in e ac ing elemen s a e ep esen ed h ough a ec o o single
s a e a iables.
A simila classi ica ion app oach using an HMM classi ie is used in [
47
], whe e u u e
s a es p edic ed by an au o eg essi e-mo ing-a e age (ARMA) algo i hm a e classi ied as
all o no - all e en s. ARMA models a e able o p edic u u e s a es o a sys em based on a
p e ious ime-se ies. The model in eg a es wo modules, an au o eg essi e one, which uses
a linea combina ion o weigh ed p e ious sys em s a e alues, and a mo ing a e age one,
which linea ly combines weigh ed p e ious e o s be ween sys em s a e eal alues and
p edic ed ones. In he model, e o s a e assumed o be andom alues ha i a Gaussian
dis ibu ion o mean 0 and a iance σ2.
4.4. T acking
A good numbe o he e iewed sys ems iden i y objec s h ough ANN o ex ac
silhoue es om he backg ound. Then, ele an ea u es a e associa ed wi h he al eady
segmen ed objec s. This assignmen equi es a cons an upda e, and, he e o e, objec
co ela ion needs o be es ablished om ame- o- ame. This co ela ion is made h ough
objec acking, and a good numbe o di e en echniques a e used o such a pu pose.
4.4.1. Mo ing A e age Fil e
The double mo ing a e age il e used in [
65
] smoo hs e ical dis ance om join s o
he g ound plane. This il e de e mines wice he mean alue o he las n samples, ac ing
his way as a low pass il e , elimina ing high- equency signal componen s associa ed
wi h noise.
4.4.2. PID Fil e
The sys em p oposed in [
42
] uses a p opo ional-in eg al-di e en ial (PID) il e o
main ain acking on silhoue es segmen ed om he backg ound. Cons an s o he il e
o gua an ee smoo h acking, educing o e shoo s and s eady-s a e e o s, a e calcula ed
h ough a gene ic algo i hm. This algo i hm, inspi ed by he heo y o na u al e olu ion,
is a heu is ic sea ch whe e se s o alues a e selec ed o disca ded based on i s abili y o
educe o a minimum he absolu e e o unc ion and, he e o e, minimize o e shoo s and
s eady e o s.
4.4.3. Kalman Fil e
Kalman il e , i s in oduced by R. E. Kalman in [
143
], is a ecu si e algo i hm
ha allows imp o emen s in he de e mina ion o sys em a iable alues by combining
se e al se s o indi ec sys em a iable obse a ions con aining inaccu acies. The esul ing
es ima ion is mo e p ecise han any o he ones which could be in e ed om a single
indi ec obse a ion se .
This way, in [
40
], he acking o join s, ollowed by h ee independen Kinec
®
sys ems,
is used by a Kalman il e . The esul ing join posi ion is es ima ed by in eg a ing in o ma-
ion om he h ee sys ems and is mo e accu a e han one o any o he
indi idual sys ems.
A pa icula a ia ion in he use o Kalman il e ing is he one in [
97
], whe e a p oce-
du e call deep-so , p esen ed in [
129
], is used. In his p ocess, a Kalman algo i hm is used
o es ima e he nex loca ion o he acked pe son, and hen he Mahalanobis dis ance is
calcula ed be ween he de ec ed pe son in he ollowing ame and i s es ima ed posi ion.
By measu ing his dis ance, unce ain y in he ack co ela ion can be quan i ied. This
il e pe o mance is deeply a ec ed by occlusion. To mi iga e his p oblem, he unce ain y
Senso s 2021,21, 947 40 o 50
alue is associa ed wi h he ack desc ip o and, o keep acks a e long occlusion pe iods,
he p ocess sa es hose desc ip o s o 100 ames.
Al hough his il e ing algo i hm wo ks e y well o main ain acks in linea sys ems,
human bodies in ol ed in a all end o beha e nonlinea ly, subs an ially deg ading i s
abili y o main ain acking.
4.4.4. Pa icle Fil e
This me hod, used in [
15
], is a Mon e Ca lo algo i hm used o objec acking in ideo
signals. In oduced in 1993 by Go don [
144
] as a Bayesian ecu si e il e , i is able o
de e mine u u e sys em s a es, in his case, u u e posi ions o he acked objec .
The il e algo i hm ollows an i e a i e app oach. This way, a e a cloud o pa icles,
image pixels, in his case, ha e been selec ed, weigh s a e assigned o hem. Those weigh
alues a e a unc ion o he p obabili y o being pa o he acked objec . Then, he
ini ial pa icle cloud is upda ed by using he weigh alues. Based on objec cinema ic,
i s mo emen is p opaga ed o he pa icle cloud, p edic ing, his way, he u u e objec
si ua ion. The p ocess con inues wi h a new upda e phase o gua an ee he p edic ed cloud
ma ches he acked objec .
This algo i hm, al hough a ec ed by occlusion, has p o en o be highly capable o
main aining acks on objec s mo ing nonlinea ly and, he e o e, he esul is adequa e o
ack human bodies du ing all e en s.
Rao–Blackwellized pa icle il e (RBPF), like he one used in [
63
], is a ype o pa icle
il e acking algo i hm used in linea /nonlinea scena ios whe e a pu ely Gaussian
app oach is inadequa e.
This algo i hm di ides pa icles in o wo se s. Those which can be analy ically e alu-
a ed and hose which canno . This way, he il e ing equa ions a e sepa a ed in o wo se s,
so wo di e en app oaches can be used o calcula e hem. The i s se , which includes
linea mo ing pa icles, is sol ed by using a Kalman il e app oach, while he second one,
whose pa icles mo e nonlinea ly, is sol ed by employing a Mon e Ca lo
sampling me hod.
4.4.5. Fused Images
In [
9
], a using cen e uses images aken om o hogonal iews, and he ob ained
objec is agged wi h a numbe . Objec s iden i ied in he nex ame a e co ela ed o
p e ious ones i hey mee he minimum dis ance es ablished h eshold. This way, he
acking is main ained.
4.4.6. Camshi
This algo i hm, in eg a ed in o OpenCV and used in [
59
], i s con e s images RGB o
hue-sa u a ion- alue (HSV) and, s a ing wi h ames whe e a CNN has c ea ed a bounding
box (BB) a ound a de ec ed pe son, i de e mines he hue his og am in each BB. Then,
mo phological ope a ions a e applied o educe noise associa ed wi h illumina ion. In he
consecu i e ame, he a ea which be e i s he eco ded Hue his og am is es ablished and
compa ed wi h de ec ed BBs. Tha way, a co ela ion can be es ablished and, he e o e, a
ack on a pe son.
4.4.7. Deep Lea ning A chi ec u es
DeepSORT is a CNN used o ack mul iple objec s a he same ime, as shown in [
87
].
The sys em p esen ed in [
71
] acks images using an algo i hm as ollows: Fi s , in
e e y new ame, a YoLO con olu ional a chi ec u e is used o iden i y people. Once all
people in he ame ha e been iden i ied, a Siamese CNN is used o i s de e mine he
cha ac e is ic ea u es o e e y pe son iden i ied in he ame and hen compa e hem
wi h he ones associa ed wi h people iden i ied in p e ious ames, looking o simila i ies.
A he same ime, an LSTM ANN is used o p edic people’s mo ion, so associa ions o
main ain ack o people om ame- o- ame can be made. Based on ea u e simila i y and
mo emen associa ion, a ack can be es ablished on people p esen in consecu i e ideo
Senso s 2021,21, 947 41 o 50
ames o can be s a ed when a new pe son appea s o he i s ime in a ideo sequence.
An almos equal p ocess is used in [
97
] o keep ack o people wi h wo CNNs wo king in
pa allel, a i s one o iden i y people and a second one o ex ac cha ac e is ic ea u es ou
o hem. Tha way, acks can be es ablished.
In [
41
], a CNN is used o de ec people in e e y ame. A BB is es ablished a ound,
and dis ances om cen al poin BBs o consecu i e ames a e de e mined. Boxes mee ing
minimum dis ance c i e ia in consecu i e ames a e co ela ed and, his way, acking
is es ablished.
4.5. Classi ying Algo i hms Pe o mances
A numbe o he e iewed sys ems es ablish compa isons wi h o he ones. Many
o hem base ha compa ison on pe o mance igu es ob ained on di e en da ase s,
while some o he s es ablish a sys em- o-sys em compa ison based on he same da abase.
Howe e , sys ems a e, in b oad e ms, an agg ega ion o wo main blocks, he i s one
whose mission is in e ing desc ip o s om images and a second one ha classi ies hose
ea u es. This way, sys em compa ison, e en on he same da ase , compa es wo agg ega ed
blocks so, compa isons on pe o mances o a speci ic block is di icul o assess, as i is
in luenced by he o he one.
To a oid hese p oblems, hese compa isons ha e been igno ed. The only ones aken
in o conside a ion ha e been hose ha compa e one o he blocks and a e based on he
same da ase . The esul s a e shown in Table 2. In global e ms, SVMs and deep lea ning
classi ie s a e he ones wi h be e pe o mances. The bes wo king classi ying deep
lea ning a chi ec u es a e MLP, au oencode s and hose implemen ing So max algo i hms
like GoogLeNe . I is also ele an ha in acco dance wi h C.J. Chong e al. [
3
], sys ems
whose desc ip o s a e dynamic and, he e o e, include e e ences o he ime a iable,
ha e be e pe o mances han hose o he ones whose desc ip o s do no inco po a e
ha a iable.
4.6. Valida ion Da ase s
The sys ems included in his esea ch ha e been es ed by using da ase s. On many
occasions, hose da ase s ha e been speci ically de eloped by he esea che s o es and
alida e hei sys ems, so hei pe o mances can be de e mined. These da ase s, al hough
b ie ly discussed in he a icles p esen ing he sys ems, a e no usually publicly accessible.
Howe e , he e a e also a g oup o da ase s used in he sys em alida ion and pe o -
mance de e mina ion phases ha a e public. Mos o hem a e also accessible h ough he
In e ne , so de elope s can download and use hem o esea ch pu poses. All he da ase s
belonging o his ca ego y used in he de elopmen o he sys ems con ained in his e iew
a e collec ed in Table 3.
Da ase s associa ed wi h he e iewed sys ems, bo h he publicly accessible ones and
he ones ha a e no , a e eco ded ei he by olun ee s o ac o s young and i enough o
gua an ee ha a simula ed all will no ha m hem. In some o hem, ac o s a e ad ised by
he apis s, so hey can imi a e how an elde ly pe son mo es o alls. Finally, none o he
da abases include elde ly eal alls o daily li e ac i i ies pe o med by elde ly people.
The da ase s a e g ouped by collec ed signal ype, so i e big g oups a e iden i ied.
1.
The i s g oup is in eg a ed by a single da ase . I collec s alls and ac i i ies o daily
li e (ADL) execu ed by olun ee s whose esul s a e eco ded using di e en senso s,
included RGB and IR came as. I is used by a single sys em o alida ion pu poses;
2.
The second g oup, which includes h ee da ase s, inco po a es dep h and accele o-
me ic da a. By i s ele ance and numbe o e iewed sys ems using i in hei
pe o mance e alua ion, one da ase is especially impo an , UR all de ec ion [
29
].
This da ase , employed by o e a hi d o all s udied sys ems, includes 30 alls and
40 ADLs eco ded by wo dep h sys ems, one p o iding on al images and a second
came a eco ding e ical ones. This in o ma ion is accompanied by accele ome ic
da a and was eleased in 2015;
Senso s 2021,21, 947 48 o 50
88.
Menacho, C.; O donez, J. Fall de ec ion based on CNN models implemen ed on a mobile obo . In P oceedings o he 2020 17 h
In e na ional Con e ence on Ubiqui ous Robo s (UR), Kyo o, Japan, 22–26 June 2020; pp. 284–289.
89.
Zhong, C.; Ng, W.W.Y.; Zhang, S.; Nugen , C.; Shewell, C.; Medina-Que o, J. Mul i-occupancy Fall De ec ion using Non-In asi e
The mal Vision Senso . IEEE Sens. J. 2020,21, 1. [C ossRe ]
90.
Sun, G.; Wang, Z. Fall de ec ion algo i hm o he elde ly based on human pos u e es ima ion. In P oceedings o he 2020
Asia-Paci ic Con e ence on Image P ocessing, Elec onics and Compu e s (IPEC), Busan, Ko ea, 13–16 Oc obe 2020; pp. 172–176.
91.
Liu, J.-X.; Tan, R.; Sun, N.; Han, G.; Li, X.-F. Fall De ec ion unde P i acy P o ec ion Using Mul i-laye Comp essed Sensing. In
P oceedings o he 2020 3 d In e na ional Con e ence on A i icial In elligence and Big Da a (ICAIBD), Chengdu, China, 28–31
May 2020; pp. 247–251.
92.
Thummala, J.; Pum in, S. Fall De ec ion using Mo ion His o y Image and Shape De o ma ion. In P oceedings o he 2020 8 h
In e na ional Elec ical Enginee ing Cong ess (iEECON), Chiang Mai, Thailand, 4–6 Ma ch 2020; pp. 1–4. [C ossRe ]
93.
Zhang, J.; Wu, C.; Wang, Y. Human Fall De ec ion Based on Body Pos u e Spa io-Tempo al E olu ion. Senso s
2020
,20, 946.
[C ossRe ]
94.
Ko a i, K.N.; Delibasis, K.; Maglogiannis, I. Real-Time Fall De ec ion Using Uncalib a ed Fisheye Came as. IEEE T ans. Cogn.
De . Sys . 2019,12, 588–600. [C ossRe ]
95.
Delibasis, K.; Goudas, T.; Maglogiannis, I. A no el obus app oach o handling illumina ion changes in ideo segmen a ion.
Eng. Appl. A i . In ell. 2016,49, 43–60. [C ossRe ]
96.
PIROPO (People in Indoo ROoms wi h Pe spec i e and Omnidi ec ional Came as). A ailable online: h ps://www.g i.ss .upm.
es/ esea ch/g i-da a/da abases (accessed on 27 Janua y 2020).
97.
Feng, Q.; Gao, C.; Wang, L.; Zhao, Y.; Song, T.; Li, Q. Spa io- empo al all e en de ec ion in complex scenes using a en ion
guided LSTM. Pa e n Recogni . Le . 2020,130, 242–249. [C ossRe ]
98.
Xu, Q.; Huang, G.; Yu, M.; Guo, Y.; Huang, G. Fall p edic ion based on key poin s o human bones. Phys. A S a . Mech. i s Appl.
2020,540, 123205. [C ossRe ]
99.
Shah oudy, A.; Liu, J.; Ng, T.-T.; Wang, G. NTU RGB+D: A La ge Scale Da ase o 3D Human Ac i i y Analysis. In P oceedings
o he 2016 IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp.
1010–1019.
100.
H un, S.N.; Zin, T.T.; Tin, P. Image P ocessing Technique and Hidden Ma ko Model o an Elde ly Ca e Moni o ing Sys em. J.
Imaging 2020,6, 49. [C ossRe ]
101.
Kalinga, T.; Si i hunge, C.; Buddhika, A.; Jayaseka a, P.; Pe e a, I. A Fall De ec ion and Eme gency No i ica ion Sys em o Elde ly.
In P oceedings o he 2020 6 h In e na ional Con e ence on Con ol, Au oma ion and Robo ics (ICCAR), Singapo e, 20–23 Ap il
2020; pp. 706–712.
102.
Chen, W.; Jiang, Z.; Guo, H.; Ni, X. Fall De ec ion Based on Key Poin s o Human-Skele on Using OpenPose. Symme y
2020
,12,
744. [C ossRe ]
103.
Cai, X.; Li, S.; Liu, X.; Han, G. Vision-Based Fall De ec ion Wi h Mul i-Task Hou glass Con olu ional Au o-Encode . IEEE Access
2020,8, 44493–44502. [C ossRe ]
104.
Chen, Y.; Li, W.; Wang, L.; Hu, J.; Ye, M. Vision-Based Fall E en De ec ion in Complex Backg ound Using A en ion Guided
Bi-Di ec ional LSTM. IEEE Access 2020,8, 161337–161348. [C ossRe ]
105.
Chen, Y.; Kong, X.; Meng, L.; Tomiyama, H. An Edge Compu ing Based Fall De ec ion Sys em o Elde ly Pe sons. P ocedia
Compu . Sci. 2020,174, 9–14. [C ossRe ]
106.
Wang, X.; Jia, K. Human Fall De ec ion Algo i hm Based on YOLO 3. In P oceedings o he 2020 IEEE 5 h In e na ional
Con e ence on Image, Vision and Compu ing (ICIVC), Qingdao, China, 23–25 July 2020; pp. 50–54.
107. Donoho, D.L. Comp essed sensing. IEEE T ans. In . Theo y 2006,52, 1289–1306. [C ossRe ]
108.
Ho p ase , T.; Ha wood, D.; Da is, L.S. A S a is ical App oach o Real- ime Robus Backg ound Sub ac ion and Shadow
De ec ion. In P oceedings o he IEEE ICCV’99 FRAME-RATE Wo kshop, Ke ky a, G eece, 20 Sep embe 1999.
109.
Mi al, A.; Pa agios, N. Mo ion-based backg ound sub ac ion using adap i e ke nel densi y es ima ion. In P oceedings o he
2004 IEEE Compu e Socie y Con e ence on Compu e Vision and Pa e n Recogni ion, CVPR 2004, Washing on, DC, USA, 27
June–2 July 2004.
110.
Mousse, M.A.; Mo amed, C.; Ezin, E.C. Fas Mo ing Objec De ec ion om O e lapping Came as. In P oceedings o he 12 h
In e na ional Con e ence on In o ma ics in Con ol, Au oma ion and Robo ics, Colma , F ance, 21–23 July 2015; pp. 296–303.
111.
Ma io, I.; Chacon, M.; Se gio, G.D.; Ja ie , V.P. Simpli ied SOM-neu al model o ideo segmen a ion o mo ing objec s. In
P oceedings o he 2009 In e na ional Join Con e ence on Neu al Ne wo ks, A lan a, GA, USA, 14–19 June 2009; pp. 474–480.
112.
Nguyen, V.-T.; Le, T.-L.; T an, T.-H.; Mullo , R.; Cou boulay, V.; Van-Toi, N. A new hand ep esen a ion based on ke nels o
hand pos u e ecogni ion. In P oceedings o he 2015 11 h IEEE In e na ional Con e ence and Wo kshops on Au oma ic Face and
Ges u e Recogni ion (FG), Ljubljana, Slo enia, 4–8 May 2015; Volume 1, pp. 1–6.
113.
Kennedy, J.; Ebe ha , R. Pa icle swa m op imiza ion. In P oceedings o he ICNN’95-In e na ional Con e ence on Neu al
Ne wo ks, Pe h, Aus alia, 27 No embe –1 Decembe 1995; Volume 4, pp. 1942–1948. [C ossRe ]
114.
Bobick, A.F.; Da is, J.W. The ecogni ion o human mo emen using empo al empla es. IEEE T ans. Pa e n Anal. Mach. In ell.
2001,23, 257–267. [C ossRe ]

Senso s 2021,21, 947 49 o 50
115.
Lucas, B.D.; Kanade, T. An i e a i e image egis a ion echnique wi h an applica ion o s e eo ision. In P oceedings o he
Imaging Unde s anding Wo kshop, Vancoube , Canada, 24–28 Augus 1981; pp. 121–130.
116. Shi, J.; Tomasi, C. Good Fea u es o T ack. In P oceedings o he IEEE Con e ence on Compu e Vision and Pa e n Recogni ion,
Sea le, DC, USA, 21–23 June 1994; pp. 593–600.
117. Candès, E.; Li, X.; Ma, Y.; W igh , J. Robus p incipal componen analysis? J. ACM 2011,58, 1–37. [C ossRe ]
118.
Dalal, N.; T iggs, B. His og ams o o ien ed g adien s o human de ec ion. In P oceedings o he 2005 IEEE Compu e Socie y
Con e ence on Compu e Vision and Pa e n Recogni ion, CVPR, San Diego, CA, USA, 20–25 June 2005.
119.
Nizam, Y.; Haji Mohd, M.N.; Abdul Jamil, M.M. A S udy on Human Fall De ec ion Sys ems: Daily Ac i i y Classi ica ion and
Sensing Techniques. In . J. In eg . Eng. 2016,8, 35–43.
120.
Kali a, S.; Ka maka , A.; Haza ika, S.M. E icien ex ac ion o spa ial ela ions o ex ended objec s is-à- is human ac i i y
ecogni ion in ideo. Appl. In ell. 2018,48, 204–219. [C ossRe ]
121.
Rosenbla , F. The pe cep on: A p obabilis ic model o in o ma ion s o age and o ganiza ion in he b ain. Psychol. Re .
1958
,65,
386–408. [C ossRe ] [PubMed]
122.
Hop ield, J.J. Neu al ne wo ks and physical sys ems wi h eme gen collec i e compu a ional abili ies. P oc. Na l. Acad. Sci. USA
1982,79, 2554–2558. [C ossRe ] [PubMed]
123. Hoch ei e , S.; Schmidhube , J. Long Sho -Te m Memo y. Neu al Compu . 1997,9, 1735–1780. [C ossRe ] [PubMed]
124. Su ske e , I.; Vinyals, O.; Le, Q.V. Sequence o sequence lea ning wi h neu al ne wo ks. a Xi 2014, a Xi :1409.3215.
125.
Hubel, D.H.; Wiesel, T.N. Recep i e ields, binocula in e ac ion and unc ional a chi ec u e in he ca ’s isual co ex. J. Physiol.
1962,160, 106–154. [C ossRe ]
126.
Fukushima, K. Neocogni on: A sel -o ganizing neu al ne wo k model o a mechanism o pa e n ecogni ion una ec ed by shi
in posi ion. Biol. Cybe n. 1980,36, 193–202. [C ossRe ]
127.
LeCun, Y.; Bo ou, L.; Bengio, Y.; Ha ne , P. G adien -based lea ning applied o documen ecogni ion. P oc. IEEE
1998
,86,
2278–2324. [C ossRe ]
128.
Tyge , M.; B una, J.; Chin ala, S.; LeCun, Y.; Pian ino, S.; Szlam, A. A Ma hema ical Mo i a ion o Complex-Valued Con olu ional
Ne wo ks. Neu al Compu . 2016,28, 815–825. [C ossRe ]
129.
Wojke, N.; Bewley, A.; Paulus, D. Simple online and eal ime acking wi h a deep associa ion me ic. In P oceedings o he 2017
IEEE In e na ional Con e ence on Image P ocessing (ICIP), Beijing, China, 17–20 Sep embe 2017; pp. 3645–3649.
130. Junejo, I.N.; Fo oosh, H. Euclidean pa h modeling o ideo su eillance. Image Vis. Compu . 2008,26, 512–528. [C ossRe ]
131.
Maldonado, C.; Rios-Figue oa, H.V.; Mezu a-Mon es, E.; Ma in, A.; Ma in-He nandez, A. Fea u e selec ion o de ec allen
pose using dep h images. In P oceedings o he 2016 In e na ional Con e ence on Elec onics, Communica ions and Compu e s
(CONIELECOMP), Cholula, Mexico, 24–26 Feb ua y 2016; pp. 94–100.
132.
Labay ade, R.; Aube , D.; Ta el, J.-P. Real ime obs acle de ec ion in s e eo ision on non la oad geome y h ough " -dispa i y"
ep esen a ion. In P oceedings o he In elligen Vehicle Symposium, Ve sailles, F ance, 17–21 June 2002.
133.
Schmiedel, T.; Einho n, E.; G oss, H.-M. IRON: A as in e es poin desc ip o o obus NDT-map ma ching and i s applica ion
o obo localiza ion. In P oceedings o he 2015 IEEE/RSJ In e na ional Con e ence on In elligen Robo s and Sys ems (IROS),
Hambu g, Ge many, 28 Sep embe –2 Oc obe 2015; pp. 3144–3151.
134.
Fischle , M.A.; Bolles, R.C. Random Sample Consensus: A Pa adigm o Model Fi ing wi h Applica ions o Image Analysis and
Au oma ed Ca og aphy. Read. Compu . Vis. 1987,24, 726–740. [C ossRe ]
135.
Solbach, M.D.; Tso sos, J.K. Vision-Based Fallen Pe son De ec ion o he Elde ly. In P oceedings o he 2017 IEEE In e na ional
Con e ence on Compu e Vision Wo kshops (ICCVW), Venice, I aly, 22–29 Oc obe 2017; pp. 1433–1442.
136.
Zhang, K.; Chen, S.-C.; Whi man, D.; Shyu, M.-L.; Yan, J.; Zhang, C. A p og essi e mo phological il e o emo ing nong ound
measu emen s om ai bo ne LIDAR da a. IEEE T ans. Geosci. Remo e Sens. 2003,41, 872–882. [C ossRe ]
137. Co es, C.; Vapnik, V. Suppo - ec o ne wo ks. Mach. Lea n. 1995,20, 273–297. [C ossRe ]
138. Tax, D.M.J.; Duin, R.P.W. Suppo Vec o Da a Desc ip ion. Mach. Lea n. 2004,54, 45–66. [C ossRe ]
139. B eiman, L. Random Fo es s. Mach. Lea n. 2001,45, 5–32. [C ossRe ]
140. F iedman, J.H. G eedy unc ion app oxima ion: A g adien boos ing machine. Ann. S a . 2011,29, 1189–1232. [C ossRe ]
141. Rabine , L.; Juang, B. An in oduc ion o hidden Ma ko models. IEEE ASSP Mag. 1986,3, 4–16. [C ossRe ]
142.
Na a ajan, P.; Ne a ia, R. Online, Real- ime T acking and Recogni ion o Human Ac ions. In P oceedings o he 2008 IEEE
Wo kshop on Mo ion and ideo Compu ing, Coppe Moun ain, CO, USA, 8–9 Janua y 2008; pp. 1–8. [C ossRe ]
143. Kalman, R.E. A New App oach o Linea Fil e ing and P edic ion P oblems. J. Basic Eng. 1960,82, 35–45. [C ossRe ]
144.
Go don, N. Bayesian Me hods o T acking. Ph.D. Thesis, Ma hema ics Depa men Impe ial College, London, UK, 1993.
A ailable online: h ps://spi al.impe ial.ac.uk/bi s eam/10044/1/7783/1/NeilGo don-1994-PhD-Thesis.pd (accessed on 27
Janua y 2020).
145.
Kangas, M.; Vikman, I.; Nybe g, L.; Ko pelainen, R.; Lindblom, J.; Jämsä, T. Compa ison o eal-li e acciden al alls in olde people
wi h expe imen al alls in middle-aged es subjec s. Gai Pos u e 2012,35, 500–505. [C ossRe ] [PubMed]
146.
Klenk, J.; Becke , C.; Lieken, F.; Nicolai, S.; Mae zle , W.; Al , W.; Zijls a, W.; Hausdo , J.M.; Van Lummel, R.C.; Chia i, L.; e al.
Compa ison o accele a ion signals o simula ed and eal-wo ld backwa d alls. Med. Eng. Phys.
2011
,33, 368–373. [C ossRe ]
[PubMed]
Senso s 2021,21, 947 50 o 50
147.
Thilo, F.J.S.; Hahn, S.; Hal ens, R.; Schols, J.M. Usabili y o a wea able all de ec ion p o o ype om he pe spec i e o olde
people–A eal ield es ing app oach. J. Clin. Nu s. 2018,28, 310–320. [C ossRe ]
148.
Demi is, G.; Chaudhu i, S.; Thompson, H.J. Olde Adul s’ Expe ience wi h a No el Fall De ec ion De ice. Telemed. J. E. Heal h
2016,22, 726–732. [C ossRe ]
149.
Ren, L.; Peng, Y. Resea ch o Fall De ec ion and Fall P e en ion Technologies: A Sys ema ic Re iew. IEEE Access
2019
,7,
77702–77722. [C ossRe ]