scieee Science in your language
[en] (orig)

Repositorio Institucional de Documentos

Abstract

La percepción visual 3D se refiere al conjunto de problemas que engloban la reunión de información a través de un sensor visual y la estimación la posición tridimensional y estructura de los objetos y formaciones al rededor del sensor. Algunas funcionalidades como la estimación de la ego moción o construcción de mapas are esenciales para otras tareas de más alto nivel como conducción autónoma o realidad aumentada. En esta tesis se han atacado varios desafíos en la percepción 3D, todos ellos útiles desde la perspectiva de SLAM (Localización y Mapeo Simultáneos) que en si es un problema de percepción 3D.<br />Localización y Mapeo Simultáneos –SLAM– busca realizar el seguimiento de la posición de un dispositivo (por ejemplo de un robot, un teléfono o unas gafas de realidad virtual) con respecto al mapa que está construyendo simultáneamente mientras la plataforma explora el entorno. SLAM es una tecnología muy relevante en distintas aplicaciones como realidad virtual, realidad aumentada o conducción autónoma. SLAM Visual es el termino utilizado para referirse al problema de SLAM resuelto utilizando unicamente sensores visuales. Muchas de las piezas del sistema ideal de SLAM son, hoy en día, bien conocidas, maduras y en muchos casos presentes en aplicaciones. Sin embargo, hay otras piezas que todavía presentan desafíos de investigación significantes. En particular, en los que hemos trabajado en esta tesis son la estimación de la estructura 3D al rededor de una cámara a partir de una sola imagen, reconocimiento de lugares ya visitados bajo cambios de apariencia drásticos, reconstrucción de alto nivel o SLAM en entornos dinámicos; todos ellos utilizando redes neuronales profundas.<br />Estimación de profundidad monocular is la tarea de percibir la distancia a la cámara de cada uno de los pixeles en la imagen, utilizando solo la información que obtenemos de una única imagen. Este es un problema mal condicionado, y por lo tanto es muy difícil de inferir la profundidad exacta de los puntos en una sola imagen. Requiere conocimiento de lo que se ve y del sensor que utilizamos. Por ejemplo, si podemos saber que un modelo de coche tiene cierta altura y también sabemos el tipo de cámara que hemos utilizado (distancia focal, tamaño de pixel...); podemos decir que si ese coche tiene cierta altura en la imagen, por ejemplo 50 pixeles, esta a cierta distancia de la cámara. Para ello nosotros presentamos el primer trabajo capaz de estimar profundidad a partir de una sola vista que es capaz de obtener un funcionamiento razonable con múltiples tipos de cámara; como un teléfono o una cámara de video.<br />También presentamos como estimar, utilizando una sola imagen, la estructura de una habitación o el plan de la habitación. Para este segundo trabajo, aprovechamos imágenes esféricas tomadas por una cámara panorámica utilizando una representación equirectangular. Utilizando estas imágenes recuperamos el plan de la habitación, nuestro objetivo es reconocer las pistas en la imagen que definen la estructura de una habitación. Nos centramos en recuperar la versión más simple, que son las lineas que separan suelo, paredes y techo.<br />Localización y mapeo a largo plazo requiere dar solución a los cambios de apariencia en el entorno; el efecto que puede tener en una imagen tomarla en invierno o verano puede ser muy grande. Introducimos un modelo multivista invariante a cambios de apariencia que resuelve el problema de reconocimiento de lugares de forma robusta. El reconocimiento de lugares visual trata de identificar un lugar que ya hemos visitado asociando pistas visuales que se ven en las imágenes; la tomada en el pasado y la tomada en el presente. Lo preferible es ser invariante a cambios en punto de vista, iluminación, objetos dinámicos y cambios de apariencia a largo plazo como el día y la noche, las estaciones o el clima.<br />Para tener funcionalidad a largo plazo también presentamos DynaSLAM, un sistema de SLAM que distingue las partes estáticas y dinámicas de la escena. Se asegura de estimar su posición unicamente basándose en las partes estáticas y solo reconstruye el mapa de las partes estáticas. De forma que si visitamos una escena de nuevo, nuestro mapa no se ve afectado por la presencia de nuevos objetos dinámicos o la desaparición de los anteriores.<br />En resumen, en esta tesis contribuimos a diferentes problemas de percepción 3D; todos ellos resuelven problemas del SLAM Visual.<br /> <br /> Fácil Ledesma, José María; Civera Sancho, Javier; Montesano del Campo, Luis

Read accessible full text

Repositorio Institucional de Documentos

Publisher: Universidad de Zaragoza, Prensas de la Universidad
Year: 2021
Source: https://zaguan.unizar.es/record/100732/files/TESIS-2021-089.pdf
2021
89
José Ma ía Fácil Ledesma
Deep Lea ning o 3D
Visual Pe cep ion
Di ec o /es
Ci e a Sancho, Ja ie
Mon esano del Campo, Luis
© Uni e sidad de Za agoza
Se icio de Publicaciones
ISSN 2254-7606
José Ma ía Fácil Ledesma
DEEP LEARNING FOR 3D VISUAL PERCEPTION
Di ec o /es
Ci e a Sancho, Ja ie
Mon esano del Campo, Luis
Tesis Doc o al
Au o
2021
UNIVERSIDAD DE ZARAGOZA
Escuela de Doc o ado
P og ama de Doc o ado en Ingenie ía de Sis emas e In o má ica
Reposi o io de la Uni e sidad de Za agoza – Zaguan h p://zaguan.uniza .es
Deep Lea ning o 3D Visual Pe cep ion
José M. Fácil Ledesma
Ad iso s: Ja ie Ci e a & Luis Mon esano
Depa men o de In o má ica e Ingenie ía de Sis emas
Uni e sidad de Za agoza
This disse a ion is submi ed o he deg ee o
Doc o o Philosophy
Oc obe 2020

A mi he mana,
mis pad es,
mis abuelas
y
Mª Luisa.
Acknowledgemen s
I s a ed my hesis ou yea s ago bu I made he decisions ha b ough me he e a ound a
decade ago. Na u ally, du ing all hese yea s he e has been many people in luencing my
li e, suppo ing and helping me e en i hey o I did no no ice. I is going o be impossible
o men ion hem all; hey simply a e oo many. I hank you all in ad ance. The e a e some
majo playe s I would like o men ion, o whom I am deeply hank ul and in deb .
Fi s , I wan o sa e a special men ion o hose people o who I owe such amazing
oppo uni ies and o whom I ha e g ea espec . To my ad iso s Ja ie and Luis o
mo i a ing and ad ising me h ough all hese yea s. To Thomas B ox o le ing me join his
lab o a ew mon hs. To Alejo o always eaching and helping me since my i s days; e en
i he did no ha e o. To Lina, o le ing me be pa o he eam. To my high school eache
Ramon o opening his doo o me.
Thanks o he ex e nal e iewe s Albe o and Manuel and o he PhD commi ee Ruben,
Jesus, Gab iel, Ana C is and Taihú.
I would no like o o ge all he PhD s uden s and P o esso s a Uniza and specially
hose ha I ha e di ec ly wo k wi h. I would like o specially men ion all my co-au ho s
Alejo, Cla a, Be a, Alejand o, Josechu, Jose Nei a and my wo supe -ad iso s. The e a e
some ha I would like o ha e included in he p e ious lis – I s ill wan o – and ha I wan
o hank o he e discussions and opinions: Jose, Iñigo, Jesus, Seong and Ca los. My o me
lab ma es Jason, Edu and he wo Ca los and also my h ee ou s anding Bachello and Mas e
s uden s I ha e co-supe ised: Dani, Mikel and Juan o who I hope he bes in he nex
p o essional s eps.
The day which comes o mind, while w i ing hese lines, as one o my bes days o wo k
was wi h Jose, Cla a and Alejand o. This was he bes wine-ending con e ence deadline, o
CVPR 2019. Exhaus ing expe ience, ha inished a ound 5AM, bu one o he bes momen s
o eam wo k I ha e e e expe ienced. I I ha e hal as good eamma es and iends as you in
xii
po ejemplo 50 pixeles, es a a cie a dis ancia de la cáma a. Pa a ello noso os p esen amos
el p ime abajo capaz de es ima p o undidad a pa i de una sola is a que es capaz de
ob ene un uncionamien o azonable con múl iples ipos de cáma a; como un elé ono o una
cáma a de ideo.
También p esen amos como es ima , u ilizando una sola imagen, la es uc u a de una
habi ación o el plan de la habi ación. Pa a es e segundo abajo, ap o echamos imágenes
es é icas omadas po una cáma a pano ámica u ilizando una ep esen ación equi ec angula .
U ilizando es as imágenes ecupe amos el plan de la habi ación, nues o obje i o es econo-
ce las pis as en la imagen que de inen la es uc u a de una habi ación. Nos cen amos en
ecupe a la e sión más simple, que son las lineas que sepa an suelo, pa edes y echo.
Localización y mapeo a la go plazo equie e da solución a los cambios de apa iencia
en el en o no; el e ec o que puede ene en una imagen oma la en in ie no o e ano puede
se muy g ande. In oducimos un modelo mul i is a in a ian e a cambios de apa iencia que
esuel e el p oblema de econocimien o de luga es de o ma obus a. El econocimien o de
luga es isual a a de iden i ica un luga que ya hemos isi ado asociando pis as isuales
que se en en las imágenes; la omada en el pasado y la omada en el p esen e. Lo p e e ible
es se in a ian e a cambios en pun o de is a, iluminación, obje os dinámicos y cambios de
apa iencia a la go plazo como el día y la noche, las es aciones o el clima.
Pa a ene uncionalidad a la go plazo ambién p esen amos DynaSLAM, un sis ema de
SLAM que dis ingue las pa es es á icas y dinámicas de la escena. Se asegu a de es ima
su posición unicamen e basándose en las pa es es á icas y solo econs uye el mapa de las
pa es es á icas. De o ma que si isi amos una escena de nue o, nues o mapa no se e
a ec ado po la p esencia de nue os obje os dinámicos o la desapa ición de los an e io es.
En esumen, en es a esis con ibuimos a di e en es p oblemas de pe cepción 3D; odos
ellos esuel en p oblemas del SLAM Visual.

Table o con en s
1 In oduc ion 1
1.1 3D Visual Pe cep ion and Visual SLAM . . . . . . . . . . . . . . . . . . . 1
1.2 How Deep Lea ning is Imp o ing Visual Pe cep ion . . . . . . . . . . . . . 3
1.3 Ou con ibu ions in 3D Visual Pe cep ion . . . . . . . . . . . . . . . . . . 5
1.3.1 Visual Mapping wi hou Mo ion . . . . . . . . . . . . . . . . . . . 5
1.3.2 Visual Mapping wi h Li le Mo ion . . . . . . . . . . . . . . . . . 6
1.3.3 Place Recogni ion unde Appea ance Changes . . . . . . . . . . . 9
1.3.4 Visual Recons uc ion o High-Le el S uc u es . . . . . . . . . . . 10
1.3.5 Visual SLAM on Dynamic En i onmen s . . . . . . . . . . . . . . 11
1.4 Lis o Publica ions.............................. 12
1.5 CodeReleased ................................ 14
1.6 Manusc ip O ganiza ion . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
2 Came a-Awa e Mul i-Scale Con olu ions o Single-View Dep h 17
2.1 In oduc ion.................................. 17
2.2 Rela edWo k ................................. 19
2.3 Came a-Awa e Mul i-scale Con olu ions . . . . . . . . . . . . . . . . . . . 20
2.3.1 Focal Leng h No maliza ion . . . . . . . . . . . . . . . . . . . . . 23
2.4 ModelandT aining.............................. 23
2.4.1 Ne wo k A chi ec u e . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.4.2 Losses................................. 27
2.4.3 T ainingSchedule .......................... 28
2.5 Mul i-Came a Expe imen s and Resul s . . . . . . . . . . . . . . . . . . . 29
2.5.1 Expe imen al Se up . . . . . . . . . . . . . . . . . . . . . . . . . . 29
2.5.2 In luence o con ex . . . . . . . . . . . . . . . . . . . . . . . . . . 30
2.5.3 O e i ing o s anda d ne wo ks . . . . . . . . . . . . . . . . . . . 30
2.5.4 Robus Gene aliza ion wi h CAM-Con s . . . . . . . . . . . . . . 34
2.5.5 Expe imen s on Mul iple Da ase s . . . . . . . . . . . . . . . . . . 36
xi Table o con en s
2.6 Conclusions.................................. 38
3 Combining Single-View Deep Lea ning Dep h wi h Mul i-View Dep h 41
3.1 In oduc ion.................................. 42
3.2 Rela edWo k ................................. 43
3.2.1 Mul i-ViewDep h .......................... 43
3.2.2 Single-ViewDep h.......................... 44
3.3 Single and Mul i-View Dep h Fusion . . . . . . . . . . . . . . . . . . . . . 45
3.3.1 Mul i- iewDep h........................... 46
3.3.2 Single- iewDep h .......................... 47
3.3.3 Dep hFusion............................. 48
3.3.4 Mul i- iew Low-E o Poin Selec ion . . . . . . . . . . . . . . . . 50
3.4 Expe imen alResul s............................. 52
3.5 Conclusions.................................. 55
4 Condi ion-In a ian Place Recogni ion 59
4.1 In oduc ion.................................. 59
4.2 The Pa i ioned No dland Da ase . . . . . . . . . . . . . . . . . . . . . . 61
4.2.1 Da a P e-p ocessing . . . . . . . . . . . . . . . . . . . . . . . . . 63
4.2.2 Da ase Pa i ions........................... 63
4.2.3 Placelabels.............................. 63
4.3 Rela edWo k ................................. 63
4.3.1 Desc ip o s.............................. 64
4.3.2 Visual Place Re ie al . . . . . . . . . . . . . . . . . . . . . . . . 66
4.3.3 Mul i-View Place Recogni ion using Mul i-View Desc ip o . . . . 67
4.4 Ne wo kA chi ec u es ............................ 67
4.4.1 Single-View ResNe -50 . . . . . . . . . . . . . . . . . . . . . . . . 67
4.4.2 Desc ip o G ouping . . . . . . . . . . . . . . . . . . . . . . . . . 68
4.4.3 Desc ip o Fusion........................... 68
4.4.4 Recu en Desc ip o s . . . . . . . . . . . . . . . . . . . . . . . . 68
4.5 T aining.................................... 69
4.5.1 Con en ion o Same Place . . . . . . . . . . . . . . . . . . . . . . 69
4.5.2 Model aining ............................ 70
4.6 Expe imen alResul s............................. 71
4.6.1 Pa i ioned No dland Da ase . . . . . . . . . . . . . . . . . . . . . 72
4.6.2 Alde ley................................ 75
4.6.3 Mul i-View E alua ion . . . . . . . . . . . . . . . . . . . . . . . . 76
Table o con en s x
4.6.4 Execu ion ime ............................ 79
4.7 Conclusions.................................. 81
5 Co ne P edic ion o Layou Recons uc ion 83
5.1 In oduc ion.................................. 84
5.2 Rela edWo k ................................. 85
5.3 Co ne s o Layou .............................. 87
5.3.1 Ne wo k a chi ec u e . . . . . . . . . . . . . . . . . . . . . . . . . 87
5.3.2 T aining................................ 89
5.3.3 F om Co ne Maps o 3D Layou . . . . . . . . . . . . . . . . . . . 92
5.4 Equi ec angula Con olu ions . . . . . . . . . . . . . . . . . . . . . . . . 94
5.4.1 EquiCon sDe ails .......................... 95
5.5 Expe imen s.................................. 98
5.5.1 Da ase s................................ 98
5.5.2 Implemen a ion de ails . . . . . . . . . . . . . . . . . . . . . . . . 98
5.5.3 Ne wo k’s ou pu e alua ion . . . . . . . . . . . . . . . . . . . . . 99
5.5.4 Robus ness analysis . . . . . . . . . . . . . . . . . . . . . . . . . 101
5.5.5 3D Layou compa ison . . . . . . . . . . . . . . . . . . . . . . . . 103
5.5.6 Ex a Quali a i e Resul s . . . . . . . . . . . . . . . . . . . . . . . 103
5.6 Conclusions.................................. 106
6 Monocula and RGB-D SLAM on Dynamic En i onmen s 111
6.1 In oduc ion.................................. 111
6.2 Rela edWo k ................................. 114
6.3 DynaSLAM Sys em Desc ip ion . . . . . . . . . . . . . . . . . . . . . . . 115
6.3.1 Segmen a ion o Po en ially Dynamic Con en using a CNN . . . . 116
6.3.2 Low-Cos T acking . . . . . . . . . . . . . . . . . . . . . . . . . . 116
6.3.3
Segmen a ion o Dynamic Con en using
Mask R-CNN
and Mul i-
iewGeome y............................ 117
6.3.4 T acking and Mapping . . . . . . . . . . . . . . . . . . . . . . . . 119
6.3.5 Backg ound Inpain ing . . . . . . . . . . . . . . . . . . . . . . . . 119
6.4 Expe imen alResul s............................. 121
6.4.1 TUMDa ase ............................. 121
6.4.2 KITTIDa ase ............................ 125
6.4.3 TimingAnalysis ........................... 126
6.5 Conclusions.................................. 127
x i Table o con en s
7 Conclusions 129
7.1 Limi a ions and Fu u e Wo k . . . . . . . . . . . . . . . . . . . . . . . . . 131
Re e ences 133
Appendix A De ailed Expe imen s o Came a-Awa e Con olu ions 149
A.1 Expe imen s on S an o d Da ase . . . . . . . . . . . . . . . . . . . . . . . 149
A.1.1 2D-3D Seman ics S an o d Da ase . . . . . . . . . . . . . . . . . 149
A.1.2 No a ion ............................... 149
A.1.3 In luence o Con ex in Single View Dep h . . . . . . . . . . . . . 151
A.1.4 Focal Leng h O e i ing . . . . . . . . . . . . . . . . . . . . . . . 153
A.1.5 Senso Size O e i ing . . . . . . . . . . . . . . . . . . . . . . . . 153
A.1.6 Gene aliza ion wi h CAM-Con s . . . . . . . . . . . . . . . . . . 158
A.2 NYUExpe imen ............................... 161
A.2.1 T aining................................ 162
A.2.2 Tes ing ................................ 163
Chap e 1
In oduc ion
Compu e ision has expe ienced a huge g ow h du ing he las decades. I can be b oadly de-
ined as ex ac ing in o ma ion om images wi h simila goals o humans, making compu e s
see as we do. In o de o mee his ambi ious goal, compu e ision add esses a as numbe
o di e en pe cep ion p oblems like 3D econs uc ion, objec de ec ion o objec acking.
And i does ha by using many di e en models, s a egies, algo i hms and echniques, such
as epipola geome y, p obabilis ic models o machine lea ning o name a ew examples.
Compu e ision is e e y ime mo e p esen in ou li es and we ha e g own o na u alize i ,
bene i ing om i s ad an ages and ge ing used o i s ubiqui ous p esence in many mode n
echnologies like ace ecogni ion in phones, lane de ec ion in ca s o came a acking in
i ual and augmen ed eali y (VR/AR) headse s.
1.1 3D Visual Pe cep ion and Visual SLAM
3D isual pe cep ion consis s on eco e ing in o ma ion o he 3D s uc u e behind he images
aken by came as. The e is a wide ange o esea ch p oblems g ouped unde he gene al
e m pe cep ion, ha goes om ecognizing he objec s a ound he came a o es ima ing he
ego-mo ion o he came a i sel . Some imes we may dispose o mul iple iews o a scene
aken a he same ime (e.g. s e eo came as) o aken sequen ially a di e en imes; and
some imes we will pe cei e ou en i onmen om jus one single image.
Many di e en applica ions may bene i om 3D pe cep ion. Some include, ca s wi h
pedes ian de ec ion as ex a sa e y ea u e, medical image o p ecise su ge y o people
acking in ideo igilance among many o he s. In he las yea s, many o hese pe cep ion
solu ions ha e been ans e ed o comme cial p oduc s and hey a e e e y ime mo e p esen .
Howe e , he e a e s ill many poin s o imp o e bo h in esolu ion, like comple e scene

2In oduc ion
unde s anding, and p ecision, like objec classi ica ion o dense dep h es ima ion. Some o
hese p oblems ha e no ye been add essed success ully o s ill ha e a poo pe o mance.
The e a e wo pa icula pe cep ion p oblems ha humans esol e e y well, which
a e he es ima ion o ou ajec o y – acking ou mo ion – and he 3D pe cep ion o ou
en i onmen and elemen s a ound us – mapping he scene. In he obo ics communi y
his p oblem is usually e e ed o as SLAM (which s ands o Simul aneous Localiza ion
And Mapping). Ano he ac onym is used by he compu e ision communi y o a simila
p oblem, S M (which s ands o S uc u e om Mo ion). The di e ence lies in ha SLAM,
di e en ly o S M, assumes ha he da a will appea sequen ially and aims o comple e he
compu a ion wi hin he sampling ime o he senso (s) ( eal ime). SLAM is a e y ele an
p oblem ha is essen ial o a wide ange o applica ions such as i ual, augmen ed o mixed
eali y (Newcombe e al., 2011, Klein and Mu ay, 2007), Mic o Ae ial Vehicles (MAVs)
na iga ion (Shen e al., 2011), au onomous d i ing (McManus e al., 2013) and assis ed
su ge y (Lama ca e al., 2020). Visual SLAM add esses he localiza ion and mapping asks
using only isual senso s, e.g., RGB came as.
Visual SLAM is cu en ly a ma u e p oblem, bu s ill a ele an and challenging esea ch
a ea (Cadena e al., 2016). Li e a u e has p oposed many di e en app oaches o SLAM, he
wo main amilies being he ea u e-based models – using salien poin s and ma ching hem
ac oss iews – (Klein and Mu ay, 2007, Mu -A al e al., 2015) and di ec app oaches – ha
ely di ec ly on he pixel colo s – (Newcombe e al., 2011, Engel e al., 2014). Al hough
he e a e excellen app oaches o he gene al SLAM p oblem om isual clues, he p oblem
is no ully sol ed ye . The e a e many open challenges ha can ac ually appea e y o en:
su aces wi h no ex u e and he e o e di icul o ma ch be ween di e en iews (Concha and
Ci e a, 2015a), small came a mo ions ha do no c ea e enough pa allax o iangula e poin s,
e isi ing places a e hei appea ances ha e changed (e.g. day and nigh ) (Gomez-Ojeda
e al., 2015), dynamic elemen s o he scene (e.g., people walking on he s ee ) ha iola e
he usual igidi y cons ain s o he algo i hm (Alcan a illa e al., 2012) o adding seman ic
in o ma ion o ce ain pa s o he scene (e.g. oad lanes o oom layou ) (Fe nandez-Lab ado
e al., 2018b).
As he echnology o SLAM has become mo e known and p esen in applica ions like
VR/AR, i s challenges ha e become mo e e iden . Despi e ackling he gene al p oblem
wi h g ea pe o mance, la ge e o s due o cases like day/nigh o small mo ions limi he
po en ial o he echnology in c i ical applica ions, o example au onomous d i ing. In o de
o ha e a eliable pe cep ion sys em, hese cases mus be add essed as well. Concu en ly
wi h he g ow h in use o SLAM sys ems, he echnology o eco ding and s o ing da a has
1.2 How Deep Lea ning is Imp o ing Visual Pe cep ion 3
become be e and mo e accessible. As esul o his, mo e da a wi h be e g ound u h is
a ailable. This has allowed he communi y o c ea e be e benchma ks and e alua ion ools.
The la ge amoun o da a a ailable, in his as in many o he p oblems, has made machine
lea ning algo i hms gain ele ance. A machine lea ning algo i hm is an algo i hm ha
lea ns om da a (Good ellow e al., 2016). Speci ically, i lea ns o imp o e i s pe o mance
(measu ed by a ce ain me ic) on a ask, based on expe ience. An example o a ask can be
objec classi ica ion whe e p e ious expe ience a e images ha ha e been labeled (o no )
and he measu emen o pe o mance migh be he classi ica ion accu acy.
The machine lea ning echnique ha has bene i ed mo e om ha ing a as amoun o
da a, as well as mo e compu a ional powe , is deep lea ning. Deep lea ning is a amily o
algo i hms ha o ms pa o he machine lea ning and a i icial in elligence ield. Deep
lea ning is based on a i icial neu al ne wo ks, being deep eed- o wa d ne wo ks o mul i-
laye pe cep on he quin essen ial model o deep lea ning (Good ellow e al., 2016). Deep
neu al ne wo ks (DNNs) a e a se o in e connec ed p ocessing uni s able o app oxima e
complex unc ions based on a pe o mance measu emen no mally e e ed o as loss.
I is in his con ex o deep lea ning g owing in popula i y and being used o ackle
pe cep ion p oblems ha his hesis akes place. In pa icula , we ha e con ibu ed o se e al
challenges in di e en 3D pe cep ion p oblems: dense monocula dep h es ima ion, place
ecogni ion unde d as ic appea ance changes, seman ic econs uc ion and ego-mo ion
es ima ion in dynamic en i onmen s. All o hem a e connec ed o isual SLAM o a g ea e
o lesse ex en .
1.2 How Deep Lea ning is Imp o ing Visual Pe cep ion
The inc ease o da a and compu a ional powe has made DNNs e y popula in many di e en
applica ions. Some o he mos popula examples a e ound in he na u al language p ocessing
(NLP) domain whe e we can di e en ia e se e al asks: speech ecogni ion (Amodei e al.,
2016, Wang e al., 2019), sen imen analysis (Zhang e al., 2018) o ha e speech de ec ion
(Gomez e al., 2020) among many o he s.
Deep neu al ne wo ks ha e also been widely used in compu e ision o many di e en
asks, using mainly he con olu ional laye s in oduced in LeCun e al. (1989). Con olu ional
neu al ne wo ks (CNNs) a e a specialized ype o neu al ne wo k o p ocessing da a ha has
a ixed-size g id-like opology. They a e only locally connec ed, sha ing he weigh on each
local connec ion ha is no mally e e ed as ke nel and allow he same pa e n o be lea ned
in any pa o he g id. O he ne wo k a chi ec u es we e la e used wi h a CNN backbone, o
example Gene a i e Ad e sa ial Ne wo ks (GANs) in oduced by Good ellow e al. (2014).
4In oduc ion
GANs a e a p oposal o aining gene a i e models ha consis in wo ne wo ks compe ing
in a game whe e he success o one implies he o he one’s loss. These models ha e been
used in many applica ions wi h di e en pu poses, such as aining con ex encode s as
unsupe ised p e aining o objec classi ica ion and seman ic segmen a ion (Pa hak e al.,
2016). In pa allel, Ronnebe ge e al. (2015) in oduced he U-Ne a chi ec u e, adding ex a
connec ions a di e en le els be ween encode and decode , consequen ly be e p opaga ing
high-le el high- esolu ion ea u es. Ano he ele an a chi ec u al design a e he Siamese
ne wo ks, used o asks such as image e ie al (Go do e al., 2016). A Siamese neu al
ne wo k consis s o wo DNNs ha sha e hei weigh s du ing aining, while being ed wi h
wo di e en inpu images o compu e compa able ou pu ec o s.
When wo king wi h da a ha has a empo al dimension, like language o ideo (Donahue
e al., 2014), ecu en neu al ne wo ks (RNN) ha e demons a ed a good pe o mance.
Di e en ly om a no mal DNN, a RNN de ines a di ec ed g aph wi h empo al connec ions
along a sequence o inpu s (Hoch ei e and Schmidhube , 1997). These connec ions allow
he ne wo k o cap u e empo al pa e ns om sequen ial da a, and hey ha e been used in
asks like speech ecogni ion, ideo cap ioning and handw i ing ecogni ion among o he s.
Visual pe cep ion has g ea ly bene i ed om he use o deep lea ning models. Some
popula p oblems ha ha e made use o deep ne wo ks a e objec classi ica ion (Simonyan
and Zisse man, 2014, He e al., 2016), objec de ec ion (Liu e al., 2020), image and ideo
seman ic segmen a ion (Ga cia-Ga cia e al., 2018), and ace ecogni ion (Pa khi e al., 2015).
Mo e aligned wi h he pu pose o his hesis, in 3D isual pe cep ion many asks ha e also
been add essed using deep lea ning, s a ing om pu e monocula dep h es ima ion (Eigen
e al., 2014), came a pose es ima ion (Kendall e al., 2015), objec pose acking (Tompson
e al., 2014) and low es ima ion (Doso i skiy e al., 2015). Geome ic models began o be
added o he ne wo k a chi ec u es and losses, o sel -supe ision (Goda d e al., 2017) o o
be lea ned (Ummenho e e al., 2017).
In his hesis we ha e iden i ied se e al challenges ha needed o be add essed inside
3D pe cep ion. We ha e p oposed solu ions and con ibu ed o di e en a eas, one o hem
being monocula dep h es ima ion. Along his manusc ip we will p esen wo ks ha discuss
some o hese a chi ec u al models men ioned be o e. Fo ins ance, RNN ha e been used o
gene a ing dis inc i e desc ip o s o sequence o images o pe o m isual place ecogni ion.
Siamese ne wo ks ha e been used o compa e hese desc ip o s o se e al images and also
o ain single- iew dep h es ima ion wi h images o di e en sizes. In he nex sec ion we
b ie ly explain each one o hese p oblems and ou con ibu ions.
1.2 How Deep Lea ning is Imp o ing Visual Pe cep ion 5
1.3 Ou con ibu ions in 3D Visual Pe cep ion
“I suppose i is emp ing,
i all you ha e is a hamme , e e y hing looks like a nail”
— Ab aham Maslow, The Psychology o Science (1966)
We ha e ackled he ollowing signi ican challenges on 3D isual pe cep ion: dense
monocula dep h es ima ion, place ecogni ion unde d as ic appea ance changes, seman ic
econs uc ion and ego-mo ion es ima ion in dynamic en i onmen s. We con ibu e in all o
hese challenges using deep lea ning. We make use o con olu ional ne wo ks o wo k wi h
images and add ess hese challenges by ex ac ing and p ocessing isual clues p esen in
he inpu images. We ha e also con ibu ed o make deep lea ning and speci ically CNNs
mo e adap able o di e en came as o image ep esen a ions by in oducing CAM-Con s
and EquiCon s, wo ypes o con olu ions ha will be u he de ailed in Chap e s 2 and 5.
In he ollowing sec ions we in oduce he esea ch p oblems, he ela ed li e a u e and ou
con ibu ions o all he di e en 3D isual pe cep ion challenges we ha e add essed.
1.3.1 Visual Mapping wi hou Mo ion
Visual SLAM is mainly based on ma ching poin s ac oss di e en iews and join ly es ima e
he ela i e mo emen o he came a and he iangula ion o hese poin s (T iggs e al., 1999).
Howe e , ha p ocedu e’s accu acy depends on he pa allax angle gene a ed by he came a
mo ion. When he came a does no mo e, he ajec o y does no need o be calcula ed bu
he es ima ion o he 3D o he scene becomes challenging as i is an ill-posed p oblem. This
p oblem is e e ed o in he compu e ision communi y as single- iew dep h es ima ion
o monocula dep h es ima ion. Dep h pe cep ion om one single image has usually been
add essed using machine lea ning algo i hms (Saxena e al., 2009).
One o he mos impac ul ad ances in single- iew dep h es ima ion was he use o deep
lea ning by Eigen e al. (2014), ha ained in a supe ised manne a deep neu al ne wo k
o p edic dep h om single RGB images. Thei p oposal is a con olu ional neu al ne wo k
(CNN) ained wi h RGB-D examples o p edic a dep h alue o each RGB pixel.
Following wo ks a he imp o ed p edic ions by using deepe models and ne wo ks
p e ained in o he asks, mainly classi ica ion (Eigen and Fe gus, 2015, Laina e al., 2016).
Di e en aining p ocedu es allowed o ain wi h di e en da ase s despi e no being
supe ised – meaning ha hey do no ha e dep h alues o all he pixels. Goda d e al.
(2017) p oposed a sel -supe ised scheme by o cing le - igh s e eo came a consis ency
12 In oduc ion
i we use hem o es ima e ou pose. As humans, we expe imen a simila e ec when we si
inside a ain and h ough he window we see ano he ain in he s a ion. Some imes one o
he wo ains s a s o mo e and we do no eally unde s and which one o he wo is mo ing.
Some imes we w ongly pe cei e as i we we e mo ing because we use a mo ing image o
loca e ou sel es. Fo isual acking we loca e he came a wi h espec o he pa s o he
scene we a e iewing, bu i we a e no ca e ul o using only pa s o he scene ha a e s a ic
as an ancho o es ima e he pose o he de ice, he eco e ed pose can be e oneous. Riazuelo
e al. (2017) p oposed a RGB-D SLAM sys em wi h an embedded eal- ime human acke
module. Li and Lee (2017) also p oposed a simila RGB-D SLAM sys em, bu di e en ly o
p e ious app oaches, hey use dep h edge poin s as an indica o o he p obabili y ha pa s
o he image belong o a dynamic objec .
We p opose, in Chap e 6, a combina ion o adi ional geome y and deep lea ning ech-
niques o de ec and conside dynamic objec s du ing acking and mapping o monocula ,
s e eo and RGB-D SLAM. This sys em combines he use o a seman ic segmen a ion CNN
wi h a classical ea u e-based SLAM sys em. Mo e ecen wo ks ha e ad anced mo e in he
opic o SLAM wi h dynamic objec s. Yang and Sche e (2019) build an RGB-D SLAM
sys em whe e by e ie ing 3D bounding boxes o objec s hey join ly op imize he poses o
he came a, objec s and poin s among di e en iews. Xu e al. (2019) a e able o main ain a
map acking all he ins ances o objec s by building an objec -le el oc ee-based olume ic
ep esen a ion, p o iding a obus RGB-D came a acking as well.
1.4 Lis o Publica ions
The wo k de eloped du ing his PhD hesis has p oduced he ollowing publica ions.
•
Facil, J. M., Concha, A., Mon esano, L., & Ci e a, J. (2017).
Single- iew and Mul i-
iew Dep h usion.
IEEE Robo ics and Au oma ion Le e s, 2(4), (pp. 1994-2001).
Wi h O al P esen a ion a In e na ional Con e ence on In elligen Robo s and Sys ems
(IROS) 2017.
•
Facil, J. M., Ummenho e , B., Zhou, H., Mon esano, L., B ox, T., & Ci e a, J. (2019).
CAM-Con s: CAM-Con s: Came a-Awa e Mul i-scale Con olu ions o Single-
View Dep h.
IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR)
2019 (pp. 11826-11835).
•
Olid, D., Fácil, J. M., & Ci e a, J. (2018).
Single- iew place ecogni ion unde sea-
sonal changes.
Planning, Pe cep ion and Na iga ion o In elligen Vehicles Wo kshop
a In e na ional Con e ence on In elligen Robo s and Sys ems (IROS) 2018.

1.4 Lis o Publica ions 13
Came a
T ajec o y
Dynamic
Elemen s
Fig. 1.6 This image shows an example o how dynamic objec s a e de ec ed o e ime and
aken in o accoun when acking he came a posi ion. A he same ime, he map could be
buil , no conside ing he dynamic objec s because hey should no be pa o i .
•
Facil, J. M., Olid, D., Mon esano, L., & Ci e a, J. (2019).
Condi ion-In a ian Mul i-
View Place Recogni ion. Technical Repo 2019 – a Xi p ep in a Xi :1902.09516.
•
Fe nandez-Lab ado , C., Facil, J. M., Pe ez-Yus, A., Demonceaux, C., & Gue e o, J.
J. (2018).
PanoRoom: F om he Sphe e o he 3D layou
. 3D Mee s Seman ics a
Eu opean Con e ence in Compu e Vision (ECCV) 2018.
•
Fe nandez-Lab ado , C.*, Facil, J. M.*, Pe ez-Yus, A., Demonceaux, C., Ci e a, J.,
& Gue e o, J. J. (2020).
Th eeSix y End- o-End Layou Reco e y
. Woman in
Compu e Vision a IEEE Con e ence on Compu e Vision and Pa e n Recogni ion
(CVPR) 2019. * - Equal Con ibu ion
•
Fe nandez-Lab ado , C.*, Facil, J. M.*, Pe ez-Yus, A., Demonceaux, C., Ci e a, J., &
Gue e o, J. J. (2020).
Co ne s o layou : End- o-end layou eco e y om 360
images
. IEEE Robo ics and Au oma ion Le e s, 5(2), (pp. 1255-1262). * - Equal
Con ibu ion
14 In oduc ion
•
Bescos, B., Facil J.M., Ci e a, J. & Nei a, J.,
De ec ing, T acking and Elimina -
ing Dynamic Objec s in 3D Mapping using Deep Lea ning and Inpain ing
, O al
and Pos e P esen a ion wi hin he Wo kshop a ICRA 2018: Rep esen ing a Com-
plex Wo ld: Pe cep ion, In e ence, and Lea ning o Join Seman ic, Geome ic, and
Physical Unde s anding
•
Bescos, B., Facil J.M., Ci e a, J. & Nei a, J.,
Robus and Accu a e 3D Mapping
by combining Geome y and Machine Lea ning o deal wi h Dynamic Objec s
,
Pos e P esen a ion wi hin he Wo kshop a IROS 2017: Lea ning o Localiza ion and
Mapping
•
Bescos, B., Facil J.M., Ci e a, J. & Nei a, J.,
DynaSLAM: T acking, Mapping and
Inpain ing in Dynamic Scenes
, IEEE Robo ics and Au oma ion Le e s 3 (4), (pp.
4076 - 4083). Wi h O al Spo ligh and Pos e a In e na ional Con e ence on In elligen
Robo s and Sys ems (IROS) 2018.
1.5 Code Released
Du ing he ealiza ion o his hesis we ha e eleased he ollowing code eposi o ies.
•
Sou ce code o CAM-Con s: Came a-Awa e Mul i-Scale Con olu ions o Single-
View Dep h in Tenso Flow 1.4 and 2.0.
h ps://gi hub.com/jm acil/
camcon s
•
Sou ce code o Single-View Place Recogni ion in Ca e.
h ps://gi hub.com/
jm acil/single- iew-place- ecogni ion
•
Sou ce code o Co ne s- o -Layou and EquiCon s (EquiRec angula Con olu ions)
o Tenso Flow 1.4. h ps://gi hub.com/c e nandezlab/CFL
•
Sou ce code o DynaSLAM, build upon ORB-SLAM2, T acking, Mapping and In-
pain ing in Dynamic Scenes o Monocula , S e eo and RGB-D Came as.
h ps:
//gi hub.com/Be aBescos/DynaSLAM
1.6 Manusc ip O ganiza ion
The s uc u e o his Ph.D. hesis is as ollows. In he Chap e 2 we p esen CAM-Con s o
monocula single-image dep h es ima ion wi h di e en came as. In Chap e 3 we p opose a
1.6 Manusc ip O ganiza ion 15
usion be ween single and mul i- iew dep h o imp o e mapping wi h e y li le mo ion and
ex u e o Visual SLAM. Chap e 4 p esen s se e al deep neu al ne wo k a chi ec u es o
obus isual place ecogni ion conside ing appea ance changes. In Chap e 5 we in oduce
Co ne s o Layou (CFL), an end- o-end ne wo k o ex ac co ne s o indoo ooms, and he
Equi ec angula Con olu ions (EquiCon s), a ype o con olu ion ha allows CNNs o adap
o equi ec angula dis o ions. Finally, Chap e 6 p esen s DynaSLAM, a monocula , s e eo
and RGB-D SLAM sys em ha is awa e o dynamic objec s on he scene and igno es hem
when acking he came a pose and mapping he scene.
Chap e 2
Came a-Awa e Mul i-Scale Con olu ions
o Single-View Dep h
As we ha e men ioned in he in oduc ion chap e , s uc u e wi hou mo ion is an ill-posed
p oblem ha is no sol able in gene al wi h adi ional geome ic echniques. Fo ha
eason many lea ning app oaches ha e been p esen ed o e he las yea s ying o ackle
his p oblem. Deep lea ning, as in many o he p oblems, ha e p o en o show he bes
esul s. Howe e , single- iew dep h es ima ion su e s om he p oblem ha a ne wo k
ained on images om one came a does no gene alize o images aken wi h a di e en
came a model. Thus, changing he came a model equi es collec ing an en i ely new aining
da ase . In his chap e , we p opose a new ype o con olu ion ha can ake he came a
pa ame e s in o accoun , hus allowing neu al ne wo ks o lea n calib a ion-awa e pa e ns.
Ou expe imen s con i m ha his imp o es he gene aliza ion capabili ies o dep h p edic ion
ne wo ks conside ably, and clea ly ou pe o ms he s a e o he a when he ain and es
images a e acqui ed wi h di e en came as.
2.1 In oduc ion
Reco e ing 3D in o ma ion om 2D images is one o he undamen al p oblems in compu e
ision ha , due o ecen ad ances and applica ions, is ecei ing nowadays a enewed
a en ion. Among o he s, he e has been ecen ele an esul s on p oblems such as 6D
objec pose de ec ion (Kehl e al., 2017, Sh i as a a e al., 2017, Sunde meye e al., 2018),
3D model econs uc ion Fan e al. (2017), Ta a chenko e al. (2017), dep h es ima ion om
single (Laina e al., 2016, Fu e al., 2018, Li e al., 2018) and mul iple iews (Ummenho e
e al., 2017, Huang e al., 2018), 6D came a pose eco e y (Kendall e al., 2015, 2017) o
came a acking and mapping (Zhou e al., 2018, Bloesch e al., 2018, Tang and Tan, 2018,

18 Came a-Awa e Mul i-Scale Con olu ions o Single-View Dep h
*
CAM-Con
Encode Decode
Came a Model
Fig. 2.1 CAM-Con s allows e icien specializa ion o a came a-gene ic ne wo k o a ious
came a models by eeding came a-speci ic pa ame e s in o he ne wo k.
Ta eno e al., 2017). While adi ional mul i- iew me hods (Schonbe ge and F ahm, 2016)
a e mos ly based on geome y and op imiza ion and, hus, a e la gely independen o he
da a, hese ecen deep lea ning app oaches depend on aining da a ha demons a es he
mapping om images o dep h.
The common s a egy o collec such da a is by using an RGBD senso , like he Kinec
came a, which con enien ly p o ides bo h he RGB image and wha can be conside ed
g ound u h dep h. I is implici ly assumed ha aining on his ype o da a will gene alize
o o he RGB senso s ha do no p o ide dep h. Howe e , he e alua ion o ecen lea ning-
based me hods elies la gely on public benchma ks whe e images ha e been eco ded wi h
he same RGBD came a as he aining da a. Thus, e alua ion on hese benchma ks does no
e eal whe he a dep h es ima ion me hod gene alizes o RGB images om ano he came a.
O e i ing o a benchma k is a common p oblem in compu e ision esea ch. To alba
and E os (2011) ha e shown ha da ase s may ha e s ong biases ha make esea che s
o e -con iden ega ding he pe o mance o hei me hod. In pa icula , ain- es di isions
o he same kind o da a a e no enough o p o e gene aliza ion. In his wo k we show ha ,
indeed, s a e-o - he-a single- iew dep h p edic ion ne wo ks do no gene alize when he
came a pa ame e s o he es images a e di e en om he aining ones.
Mo eo e , we show ha o single- iew dep h p edic ion he p oblem o missing gene al-
iza ion o images om di e en came as is e en mo e se e e: i canno be sol ed by aining
on images om a di e se se o came as wi h di e en pa ame e s. Fo p esen me hods o
adap o a di e en came a model, hey equi e changes in he a chi ec u e.
We p esen a deep neu al ne wo k o single- iew dep h p edic ion ha , o he i s ime,
add esses he a iabili y on he came a’s in e nal pa ame e s. We show ha his allows o
2.2 Rela ed Wo k 19
use images om di e en came as a ain and es ime wi hou a pe o mance deg ada ion.
This is o pa icula in e es , as i enables he exploi a ion o images om any came a o
aining he da a-hung y deep ne wo ks. Speci ically, wi hin ou p oposed ne wo k, ou
main con ibu ion is a no el ype o con olu ion, ha we name as CAM-Con s (Came a-
Awa e Mul i-scale Con olu ions), ha conca ena es he came a in e nal pa ame e s o he
ea u e maps, and hence allows he ne wo k o lea n he dependence o he dep h om hese
pa ame e s. Figu e 2.1 shows an illus a ion o how CAM-Con s ac in he ypical encode -
decode dep h es ima ion pipeline. The ne wo k can be ained wi h a mix u e o images
om di e en came as wi hou o e i ing o speci ic in insics. We show ha he ne wo k
gene alizes also o images om came as i has no been ained on. A compa ison wi h he
s a e o he a in single-image dep h es ima ion demons a es ha he be e gene aliza ion
p ope ies do no educe he accu acy o he dep h es ima es.
2.2 Rela ed Wo k
Es ima ing 3D s uc u e and 6 deg ees-o - eedom mo ion using deep lea ning has been
add essed ecen ly om se e al angles: Supe ised (Laina e al., 2016) and unsupe ised
(Zhou e al., 2017), om single (Eigen and Fe gus, 2015) and mul iple iews (Tang and Tan,
2018), using end- o-end ne wo ks (Laina e al., 2016) o using wi h mul i- iew geome y
(Fácil e al., 2017), comple ing dep h maps (Zhang e al., 2018, Wee aseke a e al., 2018),
and es ima ing geoloca ion (Weyand e al., 2016, Kendall e al., 2015), ela i e mo ion
(Ummenho e e al., 2017), isual odome y (Wang e al., 2017, 2018), and simul aneous
localiza ion and mapping (SLAM) (Ta eno e al., 2017, Bloesch e al., 2018, Zhou e al.,
2018).
In his wo k we deal wi h single- iew supe ised dep h lea ning, so we will ocus ou
li e a u e e iew in his case. Among he pionee ing wo k we can e e ence Hoiem e al.
(2005), ha simila ly o pop-up illus a ions, cu and old a 2D image based on a segmen a ion
in o geome ic classes and some geome ic assump ions. Saxena e al. (2009) is ano he
seminal wo k ha , wi h minimal assump ions on he scene, lea ned a model based on a
MRF. Eigen e al. (2014) was he i s pape ha used deep lea ning o single- iew dep h
p edic ion, p oposing a mul i-scale dep h ne wo k. I s esul s we e imp o ed la e by Eigen
and Fe gus (2015), Liu e al. (2015b), Laina e al. (2016), Chak aba i e al. (2016) and He
e al. (2018).
Many me hods ocus on speci ic da ase s which enable o ain lea ning-based me hods
o speci ic asks. Fo ins ance, Eigen and Fe gus (2015) ex end he mul i-scale a chi ec u e
in Eigen e al. (2014) o he p edic ion o su ace no mals and seman ic labels on he NYU
20 Came a-Awa e Mul i-Scale Con olu ions o Single-View Dep h
da ase (Silbe man e al., 2012). Simila ly, Wang e al. (2015) ain a ne wo k ha join ly
p edic s dep h and segmen a ion on he same da ase . Fo dep h, Laina e al. (2016), Liu
e al. (2015b) and Eigen and Fe gus (2015) show ha hei me hods can be adap ed o o he
da ase s like Make3D (Saxena e al., 2009) o KITTI (Geige e al., 2012). Howe e , hey ea
da ase s like di e en asks and equi e e aining o each da ase o achie e s a e-o - he-a
pe o mance.
Chen e al. (2016), inspi ed by Zo an e al. (2015), in oduce he Dep h in he Wild
da ase and ain a CNN using o dinal ela ions be ween poin pai s. While he images
s em om in e ne pho o collec ions aken wi h many di e en came as, hey do no make
use o he came a pa ame e s du ing aining. Li and Sna ely (2018) use a s uc u e om
mo ion pipeline o ex ac dep h om in e ne pho o collec ions and use his o ain a CNN
p edic ing dep h up o a scale ac o . Again, in o ma ion abou came a pa ame e s is no
exploi ed and gene aliza ion is solely d i en by la ge di e se da ase s. Ex insic pa ame e s
ha e been conside ed o o he asks such as s e eo es ima es (Ummenho e e al., 2017) o
syn hesis o iew poin changes (Zhou e al., 2016). In insic pa ame e s a e usually le ou
in deep lea ning pipelines, wi h he excep ion o He e al. (2018). They embed ocal leng h
in o ma ion in a ully-connec ed app oach, making i impossible o ain and es in di e en
image sizes, while ou p oposal is lexible and can deal wi h di e en image sizes. Pos e io
o he publica ion o his wo k, López-An eque a e al. (2020) ha e p oposed a canonical
came a model, hey show ha con e ing e e y image o his model i is a simple me hod o
ain han he one we p opose, despi e o ha ing some d awbacks like o cing he images o
ha e same size and op ical cen e can o ce o d op pa s o hem.
In he nex sec ion we desc ibe how o explici ly implemen he in e nal came a pa ame e s
in o he ne wo k and he eby imp o e gene aliza ion by CAM-Con s.
2.3 Came a-Awa e Mul i-scale Con olu ions
CAM-Con s (s anding o Came a-Awa e Mul i-scale Con olu ions), is he a ian o he
con olu ion ope a ion ha we p esen in his hesis. CAM-Con s include he came a in insics
in he con olu ions, allowing he ne wo k o lea n and p edic dep h pa e ns ha depend on
he came a calib a ion. Speci ically, we add CAM-Con s in he mapping om RGB ea u es
o 3D in o ma ion–e.g. dep h, no mals–, ha is, be ween he encode and he decode . As
shown in Figu e 2.2, we add hem a e e y le el, such ha we include CAM-Con s on e e y
skip-connec ion oo. No ice ha all he CAM-Con s a e added a e he encode , allowing
he use o p e ained models.
2.3 Came a-Awa e Mul i-scale Con olu ions 21
*
*
*
*
*
*CAM-Con
Fig. 2.2 Adding CAM-Con s o an Encode -Decode U-Ne a chi ec u e.
The basics o CAM-Con s a e as ollows: We p e-compu e pixel-wise coo dina es and
ield-o - iew maps and eed hem along wi h he inpu ea u es o he con olu ion ope a ion.
CAM-Con s use he idea behind Coo d-Con s p esen ed in Liu e al. (2018), on adding
no malized coo dina es pe pixel, bu inco po a ing in o ma ion on he came a calib a ion.
An illus a i e scheme o how CAM-Con s ex a channels wo k is shown in Figu e 2.3. The
di e en maps included a e compu ed using he came a in insic pa ame e s ( ocal leng h
and p incipal poin coo dina es (cx,cy)) and he senso size (wid h wand heigh h):
Cen e ed Coo dina es (cc):
To add he in o ma ion o he p incipal poin loca ion o he
con olu ions, we include
ccx
and
ccy
coo dina e channels cen e ed a he p incipal poin –i.e.
he p incipal poin has coo dina es (0,0). Speci ically, he channels a e
ccx=




0−cx
1−cx
.
.
.
w−cx




w×1
·




1
1
.
.
.
1





⊺
h×1
=


−cx··· w−cx
.
.
.....
.
.
−cx··· w−cx


(2.1)
ccy=




1
1
.
.
.
1




w×1
·




0−cy
1−cy
.
.
.
h−cy





⊺
h×1
=


−cy··· −cy
.
.
.....
.
.
h−cy··· h−cy


.(2.2)
We esize hese maps o he inpu ea u e size using bilinea in e pola ion and conca ena e
hem as new inpu channels. These channels a e sensi i e o he senso size and esolu ion
(pixel size) o he came a, as hei alues depend on i . We assume he senso size is
measu ed in pixels. In Figu e 2.3 we ep esen
cc
wi h a colo g adien om ed ( o nega i e
coo dina es) o blue ( o posi i e coo dina es), whi e o 0. No ice in he igu e how
cc
alues
change when came a senso size, p incipal poin o pixel size change.
28 Came a-Awa e Mul i-Scale Con olu ions o Single-View Dep h
No mal Loss:
Fo he no mal loss, we use he L2 no m. The g ound u h o he no mals
(ˆ
n) is de i ed om he g ound u h dep h image. The loss o he no mals is as ollows:
Ln=∑
i,jn(i,j)−ˆ
n(i,j)|2.(2.10)
To al Loss:
The indi idual losses a e weigh ed by ac o s ob ained empi ically, so he
o al loss Lis
L=λ1Ld+λ2Lg+λ3Lc+λ4Ln,(2.11)
whe e λ1,λ2,λ3and λ4a e 150, 100, 50 and 25 espec i ely.
2.4.3 T aining Schedule
We ain all ou ne wo ks using he Tenso Flow amewo k (Abadi e al., 2016). We s a
om ResNe -50 (p e- ained) and a andomly ini ialized decode . Fo op imiza ion we use
Adam Op imize (Kingma and Ba, 2014) wi h a momen um o
0.9
. The comple e aining o
he ne wo k is composed by h ee di e en s ages, each adding mo e laye s and p edic ions
o he decode (SR,MR and HR espec i ely in Table 2.1).
1s s age.
We ain un il he i s wo esolu ions o he decode (LR-1 in Table 2.1). This
s age is he sho es one, ained only o
10k
i e a ions wi h ba ch size
16
. We only ain
he encode laye s and he decode un il he smalles p edic ion. We do no apply he
scale-in a ian loss o his esolu ion.
2nd s age.
We ain un il he nex wo p edic ions (LR-1, MR-1 and MR-2 in Table 2.1). As
in he p e ious s age we ain only he laye s ha a ec he ou pu s. This s age is ained o
50k
i e a ions wi h ba ch size
16
. We apply a scale-in a ian loss o p edic ion MR-2 a e
25ki e a ions.
3 d s age.
We ain he whole ne wo k, o
200k
i e a ions wi h ba ch size
16
. We apply a
scale-in a ian loss o p edic ions MR-2, HR-1 and HR-2 a e 25ki e a ions.
Lea ning a e.
The lea ning a e policy o he h ee s ages is shown in Figu e 2.5. The base
lea ning a es o he h ee s ages a e
1e−3
,
5e−4
and
1e−4
. The lea ning a e d ops along
he i e a ions, ne e being less han a minimum o
1e−6
. As we ain mul iple models we
use a ixed au oma ic lea ning a e decay.
Weigh ing losses.
Fo all s ages we minimize he losses o se e al esolu ions. We scale he
losses acco ding o he esolu ion le el. Speci ically, we mul iply he losses by a ac o
1
k
,
whe e
k
deno es he esolu ion le el. S a ing wi h he ines esolu ion o he ac i e s age. I.e
in he 1
s
s age LR-1 p edic ion loss would be mul iplied by 1. While in he 3
d
s age LR-1

2.5 Mul i-Came a Expe imen s and Resul s 29
10000 60000 260000
I e a ion s
10 6
10 5
10 4
10 3
Le a ning Ra e
1s s age
2nd s age
3 d s age
Fig. 2.5 Lea ning a e policy o he h ee-s ages aining.
would be mul iplied by
1
5
, MR-1 would be mul iplied by
1
4
and so on un il HR-2 ha would
be mul iplied by 1.
2.5 Mul i-Came a Expe imen s and Resul s
Mos o he single- iew dep h p edic ion ne wo ks ha e been ained and es ed using he
same o e y simila came a models. Gene alizing o di e en came a models has se e al
implica ions ha a e no s aigh o wa d. Fo his eason, we i s p esen a ho ough
analysis on he gene aliza ion capabili ies o cu en app oaches. To his end we apply
naï e gene aliza ion echniques ( ocal no maliza ion and image esizing) du ing aining on
a ne wo k wi hou ou special con olu ions (as Figu e 2.4 bu wi hou CAM-Con s) and
examine he limi a ions. Finally, we ain and e alua e ou ne wo k wi h CAM-Con s (as
Figu e 2.4) and show he imp o ed gene aliza ion pe o mance wi h espec o di e en
came a pa ame e s.
2.5.1 Expe imen al Se up
The majo pa o ou expe imen s a e done on he 2D-3D Seman ics Da ase (A meni e al.,
2017), ha con ains RGB-D equi- ec angula images. This da ase allows us o gene a e
images wi h di e en came a in insics bu he same con en . We ha e obse ed ha dep h
es ima ion ne wo ks o e i o he came a pa ame e s and he image con en dis ibu ion
( he la e being di e en in indoo s and ou doo s da ase s, o example). In his manne we
elimina e he con en dis ibu ion ac o and isola e he e ec o he came a pa ame e s.
All he expe imen s we e done using he 3- old c oss- alida ion sugges ed by A meni
e al. (2017). In his sec ion we p esen median alues o he mos ele an expe imen s.
To see he comple e esul s, mo e de ails on he da ase and image gene a ion p ocess and
addi ional expe imen s we e e he eade o he Appendix A.
30 Came a-Awa e Mul i-Scale Con olu ions o Single-View Dep h
Name s1s2s3
Senso 256×192 192×256 224×224
Name s4s5sSsK
Senso 128×96 320×320 256×192 384×128
Name 72 128 64 n
Focal 72 128 64 100
Table 2.3 No a ion o di e en senso sizes and ocal leng hs.
The no a ion o senso sizes and ocal leng hs used du ing he e alua ion is in Table 2.3.
As an example, i a ne wo k has been ained wi h senso sizes
192×256
and
224×224
,
and ocal leng h 72, we will deno e his model as
s2s3 72
. In some expe imen s we use a
andom dis ibu ion o he ocal leng h. As an example, i he syn hesized ocal leng hs a e
uni o mly dis ibu ed be ween 72 and 128, he model will be deno ed as U 72 128.
We e alua e he pe o mance on bo h dep h and in e se dep h. All he e o me ics we
used in ou expe imen s a e s anda d om he li e a u e. In addi ion we use ela i e me ics
and he scale-in a ian me ic p esen ed by Eigen e al. (2014), which a e widely used in
dep h es ima ion.
2.5.2 In luence o con ex
Modi ying he came a pa ame e s a ec s he ield o iew, and hence he amoun o con ex
he image is cap u ing. We e alua e he in luence o he con ex in he dep h p edic ion o a
s anda d U-Ne encode -decode a chi ec u e (ne wo k in Figu e 2.4 wi hou CAM-Con s)
wi h wo di e en expe imen s. Fi s , we compa e wo ne wo ks ained wi h images wi h
senso size
s1
and wo di e en ocal leng hs
128
and
64
(Table 2.4). Second, we compa e
wo ne wo ks wi h images wi h he same ocal leng h bu di e en senso sizes:
s1
and
s4
(Table 2.5).
As expec ed, con ex helps. The pe o mance is be e o he smalles ocal
64
, which
esul s in a wide FOV and hence mo e con ex . Also he pe o mance is be e o he bigge
senso size
s1
, which also p o ides mo e con ex . To emo e he con ex dependency in
ou analysis, o some o he expe imen s in nex subsec ions we will gene a e images wi h
uni o mly dis ibu ed ocal leng hs.
2.5.3 O e i ing o s anda d ne wo ks
In his expe imen we e alua e he pe o mance o a s anda d U-Ne a chi ec u e o a ia ions
o he came a pa ame e s on he aining and es se s. We will ocus he s udy on wo
2.5 Mul i-Came a Expe imen s and Resul s 31
Tes T ain abs. el mse sc.in sq. el
: 1 m lg(m): 1
s1 64 s1 64 0.17 0.378 0.0347 0.048
s1 128 s1 128 0.195 0.51 0.0387 0.0606
smalle is be e
Table 2.4 In luence o con ex , di e en ocal leng hs.
Tes T ain abs. el mse sc.in sq. el
: 1 m lg(m): 1
s1 64 s1 64 0.17 0.378 0.0347 0.048
s4 64 s4 64 0.204 0.54 0.0384 0.0637
smalle is be e
Table 2.5 In luence o con ex , di e en senso sizes.
pa ame e s: (a) ocal leng h and (b) senso size. Fi s we will ix he senso size o
s1
and
we will es on images wi h ocal leng hs
64
,
72
and
128
( i s h ee es se s in Table 2.6).
Second we will sample andom ocal leng hs om a uni o m dis ibu ion be ween
72
and
128
and we will e alua e on images wi h senso sizes
s1
and
s2
(las wo es se s in Table
2.6). Fo e e y es se he e a e
4
o
5
di e en ain se s ( e e ed in he 2
nd
column o
he able). Fo e e y es se we will e e o he case whe e he came as om he aining
and es se a e he same as he same-came a baseline. T aining se s whe e we did no use
ocal leng h no maliza ion a e deno ed wi h a ’*’. Ne wo ks ained on ain se s wi h wo
senso sizes ha e been ained ei he as Siamese ne wo ks wi h weigh sha ing o wi h image
esizing o size s1(deno ed wi h a ’†’).
I is impo an o ema k ha , o all he expe imen s, he es and aining da a was
gene a ed om he exac same images and he ne wo ks ha e he same a chi ec u e and
we e ained o he same numbe o i e a ions. Any pe o mance a ia ion, hen, should be
a ibu ed o he a ia ions in he came a in insics and he naï e solu ions we analyze. No ice
in Table 2.6 ha , in gene al, he same-came a baseline ou pe o ms he es , demons a ing
he o e i o he came a pa ame e s.
The conclusions o hese expe imen s a e as ollows.
(a) Single- ocal aining o e i s.
The pe o mance o a dep h ne wo k deg ades when
ained on images om a pa icula came a and es ed on images om di e en came as. See,
o example, he d op in pe o mance be ween he 1
s
ow ( es :
s1 64
, ain:
s1 64∗
) and he
2nd ( es : s1 64, ain: s1 72) and 3 d ( es : s1 64, ain: s1 128) ows in all me ics.
Mul i- ocal aining wi h no maliza ion helps.
The esul s imp o e when he aining
se con ains images wi h di e en ocal leng hs and is done wi h ocal no maliza ion. See,
o example, ha he esul s on es se
s1 64
wi h aining se
s1 72 128
is close o he
same-came a baseline. No ice, howe e , ha he mul i- ocal ain se does no each he
32 Came a-Awa e Mul i-Scale Con olu ions o Single-View Dep h
Tes se T ain se l1.in mse sc.in
pixels pixels 1/m m lg(m)
s1 64
s1 64*0.184 0.378 0.0347
s1 72 0.193 0.395 0.0354
s1 128 0.318 0.572 0.0483
s1 72 128*0.659 0.864 0.0614
s1 72 128 0.189 0.387 0.0361
s1 72
s1 72*0.17 0.4 0.0354
s1 128 0.272 0.564 0.0459
s1 72 128*0.552 0.888 0.0609
s1 72 128 0.175 0.404 0.0364
s1 128
s1 128*0.141 0.51 0.0387
s1 72 0.133 0.524 0.0411
s1 72 128*0.208 0.813 0.063
s1 72 128 0.132 0.504 0.038
s1U 72 128
s1U 72 128 0.15 0.46 0.037
s2U 72 128 0.175 0.51 0.0422
s1s2U 72 128 0.153 0.484 0.0401
s1s2s3U 72 128†0.179 0.742 0.064
s2U 72 128
s1U 72 128 0.151 0.44 0.038
s2U 72 128 0.133 0.412 0.0323
s1s2U 72 128 0.139 0.436 0.0352
s1s2s3U 72 128†0.16 0.622 0.0514
smalle is be e
* ained wi hou ocal leng h no maliza ion.
†images esized o s1du ing aining.
Table 2.6 O e i o came a pa ame e s o s anda d encode -decode a chi ec u es. Ne wo ks
ained om images wi h a ia ions in hei in insics pe o m wo se han he same-came a
baseline.
pe o mance o he same-came a baseline. In sec ion 2.5.4 we will show how CAM-Con s
a e able o ou pe o m he same-came a baseline e en when he aining da a does no con ain
he es ocal leng h.
The pe o mance deg ades wi hou ocal no maliza ion. Compa e, o example, he e o
me ics o he ain se s
s1 72 128∗
and
s1 72 128
. Ne wo ks ained on
72 128∗
, in ac , did
no con e ge easily.
Limi a ions o ocal no maliza ion.
Two hings should be no iced ega ding ocal no mal-
iza ion: Fi s , i does no model he changes on he senso size and he esolu ion, and we
will see now how changes on hem deg ade he pe o mance. And second, Equa ion 2.4 only
holds i he pixel size is he same o e e y came a in he aining and es se s, which in
gene al is no he case.
2.5 Mul i-Came a Expe imen s and Resul s 33
Tes T ain abs. el mse.in sc.in sq. el
% 1/km lg(m)100 %
sK
sK9.16 10.54 13.3 2.33
sSsK24.58 36.82 26.51 9.28
sSsK†9.08 10.55 13.98 2.56
abs. el l1.in mse.in sq. el
: 1 1/m1/m: 1
sS
sS0.12 0.09 0.12 0.03
sSsK0.26 0.13 0.16 0.18
sSsK†0.12 0.09 0.12 0.03
smalle is be e
†
senso size has been esized o he i s one in he lis .
Table 2.7 Naï e ain and es on KITTI Uh ig e al. (2017) and ScanNe Dai e al. (2017a).
See ha aining FCN in mul iple image sizes (
sKsS
) does no gene alize. Resizing wo ks,
bu only in his pa icula case, because o he small o e lap o isual ea u es.
(b) Single-senso size aining o e i s.
Ne wo ks ained on a senso size and es ed on
o he senso sizes do no pe o m as well as he same-came a baseline. This can be seen in
Table 2.6 in he las wo es se s
s1U 72 128
and
s2U 72 128
. Single- iew dep h es ima ion
is a con ex -dependen ask, and he ne wo k o e i s o he amoun o con ex in he aining
senso size.
Mul i-senso size aining wi h weigh sha ing does no gene alize.
T aining wi h mul i-
ple senso sizes wo ks be e han aining wi h he w ong senso size bu canno each he
same pe o mance as same-came a baselines. Fu he , aining a s ack o weigh sha ing
ne wo ks also does no scale o la ge numbe s o di e en senso sizes.
Resizing does no wo k.
As a naï e app oach, which scales o mul iple senso sizes, we
use esizing (deno ed wi h ’
†
’ in Table 2.6), which con e s all he images o size (
s1
) du ing
aining. No ice ha esizing changes he aspec a io. I also implies he ecalcula ion o
a new a e age ocal leng h
= x+ y
2
o no maliza ion. The pe o mance deg ada ion
in oduced by esizing is no iceable. Resizing c ea es inconsis en da a in ain and es ing,
which leads o lea ning and con e gence di icul ies.
Resizing helps only in a pa icula case
(non-o e lapping dis ibu ions o isual ea u es).
Table 2.7 shows an expe imen , simila o he p e ious one, on wo public da ase s: KITTI
(Uh ig e al., 2017), wi h senso size
sK
, and ScanNe (Dai e al., 2017a), wi h senso size
sS
.
In his case, aining wi h bo h senso sizes (by weigh sha ing) dec eased he pe o mance.
Howe e , esizing educed he e o o he le el o he same-came a baselines. The eason
o his is he comple ely di e en dis ibu ion o he wo da ase s, wi h null in e sec ion o

34 Came a-Awa e Mul i-Scale Con olu ions o Single-View Dep h
Tes T ain abs. el l1.in mse sc.in
: 1 1/m m lg(m)
s1U 72 128 s1U 72 128 0.189 0.15 0.46 0.037
CAM-C‡0.175 0.144 0.433 0.0312
s2U 72 128 s2U 72 128 0.166 0.133 0.412 0.0323
CAM-C‡0.158 0.131 0.39 0.0265
s3U 72 128
s3U 72 128 0.174 0.14 0.425 0.0336
s1U 72 128 0.184 0.143 0.44 0.0357
s2U 72 128 0.177 0.145 0.435 0.0356
s1s2U 72 128 0.178 0.143 0.451 0.0365
CAM-C‡0.164 0.134 0.402 0.0283
s5 64
s5 64 0.163 0.227 0.309 0.0356
s1 64 0.245 0.292 0.337 0.0598
s1s2U 72 128 0.369 0.369 0.44 0.0427
CAM-C‡0.177 0.236 0.289 0.0362
smalle is be e
‡
T ained wi h weigh sha ing in senso sizes
s1
,
s2
and
U 72 128.
Table 2.8 Came a pa ame e gene aliza ion wi h EquiCon s. Resul s on aining and es ing
on di e en came as.
1s
column: came a pa ame e s o es se .
2nd
column: came a
pa ame e s seen du ing aining. This is a con inua ion o Table 2.6. No ice how he
ne wo k wi h EquiCon s is he only model ha gene alize ge ing be e pe o mance han
he same-came a baseline on mos es se s.
isual ea u es (e.g. he e a e no chai s on KITTI and no ca s on ScanNe ). This is, howe e ,
a e y pa icula case, esizing deg ades signi ican ly he accu acy in gene al.
2.5.4 Robus Gene aliza ion wi h CAM-Con s
In his expe imen we show ha CAM-Con s gene alize o di e en came a models. In o de
o e alua e he in luence o CAM-Con s we ained ou model wi h wo di e en senso sizes
(
s1
and
s2
) and weigh sha ing. Focal leng h du ing aining is sampled andomly om a
uni o m dis ibu ion
U 72 128
. We e alua ed he ained model in ou di e en es se s,
see Table 2.8. The i s wo include he came a model he ne wo k was ained wi h, he
hi d has a senso size unseen du ing aining, and he las (
s5 64
) was gene a ed om a
came a comple ely di e en om he aining ones wi h bigge senso size and smalle ocal
leng h. This case augmen s conside ably he con ex –e.g ield o iew–which p o ed o be
he ha des case in p e ious expe imen s (see ne wo k ained wi h s1 128 in Table 2.6).
CAM-Con s gene alize o e came a in insics, ou pe o ming he same-came a base-
line.
Resul s on he es se s
s1U 72 128
and
s2U 72 128
in Table 2.8 show ha he ne wo k
2.5 Mul i-Came a Expe imen s and Resul s 35
CAM-CONVS NO CAM-CONVS
G ound-T u h
Inpu
Fig. 2.6 Quali a i e esul s o he es se
s5 64
.
1s column
: RGB inpu .
2nd column:
G ound u h dep h.
3 d column:
P edic ion wi h ou ne wo k using CAM-Con s ained on
s1s2U 72 128
.
4 h column:
P edic ion o a ne wo k ained wi hou CAM-Con s. No ice
ha he es came a pa ame e s a e signi ican ly di e en om he aining se and images
ha e a much wide ield o iew. Despi e he la ge di e ence in he came a pa ame e s he
ne wo k wi h CAM-Con s p oduces sha p dep h maps on which oom co ne s a e clea ly
isible.
wi h CAM-Con s ained on images o wo sizes clea ly ou pe o ms he baselines, which
was ained on he exac es size. The addi ion o CAM-Con s allowed he ne wo k o lea n
he dependence o he image ea u es om he calib a ion pa ame e s.
CAM-Con s gene alize o senso sizes unseen du ing aining.
Rema kably, he ne wo k
wi h CAM-Con s also ou pe o ms he same-came a baseline on he es se wi h senso size
s3
( hi d es se in Table 2.8), which is no included in he aining da a. Fu he , i gene alizes
be e han a ne wo k ained on he exac same condi ions bu wi hou CAM-Con s (see
s1s2U 72 128 in he able).
CAM-Con s gene alize o came as unseen du ing aining.
Wi h he las es se (
s5 64
)
in Table 2.8 we e alua e ou ne wo k on an ex eme case o came a pa ame e s wi h a e y
wide ield o iew and e y di e en senso size om he aining ones. Table 2.8 shows
ha CAM-Con s imp o e conside ably he gene aliza ion o new unseen came as o e he
naï e app oaches. Figu e 2.6 shows a quali a i e compa ison be ween ou ne wo k wi h
CAM-Con s and he ne wo k wi hou CAM-Con s (s1s2U 72 128) in he es se s5 64.
36 Came a-Awa e Mul i-Scale Con olu ions o Single-View Dep h
0.3
0.25
0.2
abs ela i e
0.2
0.15
0.1
l1_in e se
0.9
0.8
0.7
0.6
mse
Laina
ou s
0.4 0.3 1.3
Fig. 2.7 E o dis ibu ion on he es se o NYU 2 wi h 6 di e en came a pa ame e s. In
o ange, ou ne wo k wi h CAM-Con s, ained on se e al da ase s no including NYU 2. In
blue, Laina e al. (2016), ained on NYU 2.
2.5.5 Expe imen s on Mul iple Da ase s
In ou las expe imen we demons a e how CAM-Con s can gene alize ac oss da ases s by
aining on ou da ase s wi h di e en came as (KITTI (Uh ig e al., 2017), ScanNe (Dai
e al., 2017a), MegaDep h (Li and Sna ely, 2018) and Sun3D(Xiao e al., 2013) and es ing
on a di e en one (NYU 2 (Silbe man e al., 2012)).
T aining:
We ained ou ne wo k o h ee di e en senso sizes (
320×320
,
256×256
and
224×224
) using weigh sha ing. We augmen ed he aining da a by scaling he images
and shi ing he p incipal poin o inc ease he a ia ion o he came a pa ame e s and hen
c op o image o one o he a ge senso sizes. We did no use ocal leng h no maliza ion in
his expe imen , as we canno ensu e cons an pixel size ac oss da ase s. As MegaDep h has
only up- o-scale g ound u h, we applied only scale-in a ian losses and added he scale-
in a ian cos unc ion o Eigen e al. (2014). The same ne wo k wi hou CAM-Con s, and
hence wi h no came a in o ma ion, did no con e ge du ing aining. The lack o calib a ion
in o ma ion c ea es inconsis encies (e.g. same-size objec s may ha e di e en dep hs due o
di e en ocal leng hs).
Tes ing:
We e alua ed ou ne wo k on he o icial es se o NYU 2 and compa ed
agains he s a e o he a (Laina e al., 2016) (simila ne wo k wi hou CAM-Con s) .
No e ha he ne wo k o Laina e al. (2016) was ained exclusi ely on NYU 2, while ou
ne wo k was ained on a se o da ase s excluding NYU 2 wi h di e en came as and da a
dis ibu ions (some o he da ase s a e ou doo s, see Figu e 2.9 and Figu e 2.10). This is
impo an since ou model canno bene i om he da ase bias (To alba and E os, 2011).
We p edic ed dep hs o images om 6 di e en came as: he o iginal came a o he NYU 2
da ase and 5 simula ed ones by c opping ( o shi p incipal poin and educe senso size) and
esizing ( o change ocal leng h).
2.5 Mul i-Came a Expe imen s and Resul s 37
128x320
640x480
128x320
256x256
128x320
320x224
640x480
Inpu GT Ou s Laina
Fig. 2.8 Quali a i e esul s, NYU 2 es se wi h in insics a ia ions.
1s column
: Inpu
RGB images. Each ow shows he o iginal one and scaled and c opped e sions
2nd column
:
Dep h g ound u h.
3 d column
: P edic ion om ou ne wo k wi h CAM-Con s, ained on
se e al da ase s NOT including NYU 2. Ou ne wo k p oduces consis en dep h close o he
g ound u h o all images.
4 h column
: Laina e al. (2016), ained exclusi ely on NYU 2.
I s e o s a e low on he aining esolu ion bu does no gene alize o new in insics.
Figu e 2.7 shows he dis ibu ion o he mean e o o he usual me ics ob ained o
he 6 di e en came as. Since Laina e al. (2016) was ained on he NYU 2 da ase , i
wo ks sligh ly be e when i p edic s he images om he came a i was ained on ( he
poin wi h he smalles e o ). Howe e , pe o mance deg ades when he came a changes and
CAM-Con s ha e always smalle e o and a iance. Figu e 2.8 illus a e how CAM-Con s
dep h p edic ions a e s able o di e en came as, while p edic ions o Laina e al. (2016)
a y signi ican ly. Recall ha CAM-Con s we e no ained on NYU 2, which indica es ha
hey a e able o gene alize o e di e en came a models and ou pe o m Laina e al. (2016)
al hough hey ained on he same da ase .
Figu es 2.8, 2.9 and 2.10 show dep h p edic ions o images (and c opped/ esized e -
sions) om he NYU 2, KITTI and MegaDep h es se s. Again, no e he excellen pe o -
mance ac oss da ase s wi h di e en da a dis ibu ions and came a in insics. All p edic ions
44 Combining Single-View Deep Lea ning Dep h wi h Mul i-View Dep h
and he egula iza ion o he mul i- iew es ima ion by adding he o al a ia ion (TV) no m
o he cos unc ion.
TV egula iza ion has low accu acy o la ge ex u eless a eas, as shown by Concha e al.
(2014), Pinies e al. (2015), Piniés e al. (2015) among o he s. In o de o o e come his
Concha e al. (2014) p opose a piecewise-plana egula iza ion; he plane pa ame e s coming
om mul i- iew supe pixel iangula ion (Concha and Ci e a, 2014) o layou es ima ion
(Hedau e al., 2009b). Pinies e al. (2015) p opose highe -o de egula iza ion e ms ha
en o ce piecewise a ine cons ain s e en in sepa a ed pixels. Piniés e al. (2015) selec s
he bes egula iza ion unc ion among a se using spa se lase da a. Building upon Concha
e al. (2014), Concha e al. (2015) adds he spa se da a-d i en 3D p imi i es o Fouhey e al.
(2013) as a egula iza ion p io . Compa ed o hese wo ks, ou usion is he i s one whe e
he in o ma ion added o he mul i- iew dep h is ully dense, da a-d i en and single- iew;
and hence i does no ely on addi ional senso s, pa allax o Manha an and piecewise-plana
assump ions. I only elies on he ne wo k capabili ies o he cu en domain, assuming ha
he es da a ollows he same dis ibu ion ha he da a used o aining.
Due o he di icul y o es ima ing an accu a e and ully dense map om monocula
iews he e a e se e al app oaches ha es ima e only he dep h o he highes -g adien pixels
(Engel e al., 2014). While his app oach p oduces maps o highe densi y han he mo e
adi ional ea u e-based ones (Mu -A al e al., 2015), hey a e s ill incomple e models o
he scene and hence hei applicabili y migh be mo e limi ed.
3.2.2 Single-View Dep h
Fo a mo e de ailed and upda ed e ision o he s a e o he a on single- iew dep h es ima ion
we e e he eade o he Chap e 2 o his hesis ha speci ically add esses his p oblem. In
his Chap e we discuss he o iginal li e a u e conside o his esea ch.
Dep h can be es ima ed om a single iew using di e en image cues, o example ocus
(Ens and Law ence, 1993) o pe spec i e (S u m and Maybank, 1999). Lea ning-based
app oaches, as he one we use, basically disco e RGB pa e ns ha a e ele an o accu a e
dep h eg ession.
The pionee ing wo k o Saxena e al. (2009) ained a MRF o model dep h om a se o
global and local image ea u es. Be o e ha , Saxena e al. (2007) p esen ed an ea ly app oach
o dep h p edic ion om monocula and s e eo cues. Eigen e al. (2014) p esen ed a wo deep
con olu ional neu al ne wo k (CNN) s acked, one o p edic global dep h an he second one
ha e ines i locally. Build upon his me hod, Eigen and Fe gus (2015) ecen ly p esen ed a
h ee scale con olu ional ne wo k o es ima e dep h, su ace no mals and seman ic labeling.

3.3 Single and Mul i-View Dep h Fusion 45
High-G adien Low-G adien
Mul i-View 0.18 1.02
Single-View 0.36 0.42
Table 3.1 Median dep h e o [m] o single and mul i- iew dep h es ima ion, and high and
low-g adien pixels. This e alua ion has been done in he sequence li ing_ oom_0030a om
he NYU 2 da ase (one o he sequences wi h highe pa allax). The no malized h eshold
be ween high and low-g adien pixels is 0.35 (g ay scale).
Liu e al. (2015b) use a uni ied con inuous CRF-and-CNN amewo k o es ima e dep h. The
CNN is used o lea n he una y and pai wise po en ials ha he CRF uses o dep h p edic ion.
Based on Eigen and Fe gus (2015), Li e al. (2016) inco po a es mid-le el ea u es in
i s p edic ion using skip-laye s. I shows compe i i e esul s and a small ba ch-size aining
s a egy ha makes hei ne wo k as e o ain. Chak aba i e al. (2016) in oduces a
di e en me hod o p edic dep h om single- iew using deep neu al ne wo ks, showing
ha aining he ne wo k wi h a much iche ou pu imp o es he accu acy. Cao e al. (2016)
o mula es he dep h p edic ion as a classi ica ion p oblem and he ne ou pu is a pixel-
wise dis ibu ion o e a disc e e dep h ange. Finally, Goda d e al. (2017) p esen s an
unsupe ised ne wo k o dep h p edic ion using s e eo images.
3.3 Single and Mul i-View Dep h Fusion
S a e-o - he-a mul i- iew echniques ha e a s ong dependency on high-pa allax mo ion and
he e ogeneous- ex u e scenes. Only a educed se o salien pixels ha hold bo h cons ain s
has a small e o , and he e o o he majo i y o he poin s is la ge and unco ela ed. In
con as , single- iew me hods based on CNN ne wo ks achie e easonable e o s in all he
image bu hey a e locally co ela ed. Ou p oposal exploi s he bes p ope ies o hese wo
me hods. Speci ically, i uses a deep con olu ional ne wo k (CNN) o p oduce ough dep h
maps and uses hei s uc u e wi h he esul s o a semi-dense mul i- iew dep h me hod (Fig.
3.1).
Be o e del ing in o he echnical aspec s, we will mo i a e ou p oposal wi h some
illus a i e esul s. Table 3.1 shows he median dep h e o o he high-g adien and low-
g adien pixels o a mul i- iew and single iew econs uc ion using a medium/high-pa allax
sequence o he NYU 2 da ase . Fo he mul i- iew econs uc ion, he e o o he low-
g adien pixels inc eases by a ac o o 2. No ice ha he opposi e happens o he single- iew
econs uc ion: he e o o high-g adien pixels is he one inc easing by a ac o o 2. Fo
46 Combining Single-View Deep Lea ning Dep h wi h Mul i-View Dep h
2000
1000
00 1 23 4 E o (m)
0 1 23 4 0 1 23 4
Fig. 3.2 His og am o single- iew dep h e o [m] o h ee sample sequences. No ice he
mul iple modes, each one co esponding o a local image s uc u e, his can be seen in he
e o images in he op ow o he igu e.
his expe imen , he h eshold used o dis inguish be ween high and low-g adien pixels is
0.35 in g ay scale (whe e he maximum g adien would be 1).
Fu he mo e, he single- iew dep h e o usually has a s uc u e ha indica es he p esence
o local co ela ions. Fo ins ance, Fig. 3.2 shows he his og am o he single- iew dep h
es ima ion e o o h ee di e en sequences ( wo o he NYU 2 da ase and one o he
TUM da ase ). No ice ha he e o dis ibu ion is g ouped in di e en modes, each one
co esponding o an image segmen .
This e ec is caused by he use o he high-le el image ea u es o he la es laye s o
he CNN ne wo k, ha ex end o e dozens o pixels in he o iginal image and hence o e
homogeneous ex u e a eas. The di e en na u e o he e o s can be exploi ed o ou pe o m
bo h indi idual es ima ions. This usion, howe e , canno be naï ely implemen ed wi h a
simple global model as i equi es con en -based de o ma ions.
In he nex subsec ions we de ail he speci ic mul i and single- iew me hods ha we use
in his wo k and ou usion algo i hm.
3.3.1 Mul i- iew Dep h
Fo he es ima ion o he mul i- iew dep h we adop a di ec app oach (Engel e al., 2014),
ha allows us o es ima e a dense o semi-dense map in con as o he mo e spa se maps o
he ea u e-based app oaches. In o de o es ima e he dep h o a key ame
Ik
we i s selec
a se o
n
o e lapping ames
{I1,...,Io,...,In}
om he monocula sequence. A e ha ,
e e y pixel
xk
l
o he e e ence image
Ik
is i s backp ojec ed a an in e se dep h
ρ
and
3.3 Single and Mul i-View Dep h Fusion 47
p ojec ed again in e e y o e lapping image Io.
xo
l=Tko(xk
l,ρl) = KR⊤
ko 


K−1xk
l
||K−1xk
l||
ρl

− ko
,(3.1)
whe e
Tko
,
Rko
and
ko
a e espec i ely he ela i e ans o ma ion, o a ion and ansla ion be-
ween he key ame
Ik
and e e y o e lapping ame
Io
.
K
is he came a in e nal calib a ion
ma ix.
We de ine he o al pho o-me ic e o
C(ρ)
as he summa ion o e e y pho o-me ic e o
εl
be ween e e y pixel (o e e y high-g adien pixel i we wan a semi-dense map)
xk
l
in he
e e ence image
Ik
and i s co esponding one
xo
l
in e e y o he o e lapping image
Io
a an
hypo hesized in e se dep h ρl,
C(ρ) = 1
n
n
∑
o=1,o=k
∑
l=1
εl(Ik,Io,xk
l,ρl).(3.2)
The e o
εl(Ik,Io,xk
l,ρl)
o each indi idual pixel
xk
l
is he di e ence be ween he
pho ome ic alues o he pixel and i s co esponding one
εl(Ik,Io,xk
l,ρl) = Ik(xk
l)−Io(xo
l).(3.3)
The es ima ed dep h o e e y pixel
ˆ
ρ= ( ˆ
ρ1... ˆ
ρl... ˆ
ρ )⊤
is ob ained by he minimiza-
ion o he o al pho ome ic e o C(ρ):
ˆ
ρ=a gmin
ρ
C(ρ)(3.4)
3.3.2 Single- iew Dep h
Fo single- iew dep h es ima ion we use he Deep Con olu ional Neu al Ne wo k p esen ed
by Eigen and Fe gus (2015). This ne wo k uses h ee s acked CNN o p ocess he images
in h ee di e en scales. The inpu o he ne wo k is he RGB key ame
Ik
. As we use he
ne wo k s uc u e and pa ame e s eleased by he au ho s wi hou u he aining, ou inpu
image size is
320×240
. The ou pu o he ne wo k is he p edic ed dep h, ha we will deno e
as
s
. The size o he ou pu is
147×109
, ha we upsample in ou pipeline in o de o use i
wi h he mul i- iew dep h.
The i s scale CNN ex ac high-le el ea u es uned o dep h es ima ion. This CNN
p oduces
64
ea u e maps o size
19×14
ha a e he inpu , along wi h he RGB image, o
he second scale CNN. This second s acked CNN e ines he ou pu o he i s one wi h
48 Combining Single-View Deep Lea ning Dep h wi h Mul i-View Dep h
mid-le el ea u es o p oduce a i s coa se dep h map o size
74×55
. This dep h map is
upsampled and eeds a hi d s acked CNN ha does a local e inemen o he dep h. This
inal s ep is necessa y, as he con olu ion and pooling s eps o he p e ious laye s il e ou
he high- equency de ails.
The i s scale was ini ialized wi h wo di e en p e- ained ne wo ks: he AlexNe
(K izhe sky e al., 2012) and he Ox o d VGG (Simonyan and Zisse man, 2014). We use
he VGG e sion, he mos accu a e one as epo ed by he au ho s. This ne wo k has been
ained in indoo scenes wi h he NYUDep h 2 da ase (Na han Silbe man and Fe gus, 2012).
As hey used he o icial ain/ es spli s o he da ase , so do we. We decided o use his
neu al ne wo k because i was he bes -pe o ming dense single- iew me hod a he momen
we s a ed his wo k and s ill i is he one ha keeps be e ade o be ween quali y and
e iciency. We e e he eade o he o iginal wo k by Eigen and Fe gus (2015) o mo e
de ails on his pa o ou pipeline.
3.3.3 Dep h Fusion
As we men ioned be o e, he objec i e is o use he ou pu o each p e ious me hod while
keeping he bes p ope ies o each o hem: he single- iew eliable local s uc u e and he
accu a e, bu semi-dense mul i- iew dep h es ima ion. Le deno e
s
and
m
o he single- iew
dep h and he mul i- iew semi-dense dep h es ima ion, espec i ely.
s
is p edic ed as de ailed
in sec ion 3.3.2 and m=1
ρis he in e se o he in e se dep h es ima ed in sec ion 3.3.1.
The used dep h es ima ion
i j
o each pixel
(i,j)
o a key ame
Ik
is compu ed as a
weigh ed in e pola ion o dep hs o e he se o pixels in he mul i- iew dep h image
i j =∑
(u, )∈Ω
Wmu
si j (mu +(si j −su )),(3.5)
whe e
Ω
is he semi-dense se o pixels es ima ed by he mul i- iew algo i hm (e.g. in a high-
pa allax sequence, hey usually co espond wi h he high-g adien pixels). The in e pola ion
weigh s
Wmu
si j
model he likelihood o each pixel
(u, )∈Ω
belonging o he same local
s uc u e as pixel
(i,j)
. The in e pola ion can be in e p e ed in wo ways. Fi s , he dep h
g adien
(si j −su )
is added o each mul i- iew dep h
mu
, i.e. we c ea e dep h map o each
mu wi h he s uc u e o sand hen weigh hem wi h pixel based weigh s. Second, o each
dep h si j we modi y i acco ding o he weigh ed disc epancy be ween (mu −su ).
The key ing edien o his in e pola ion a e he weigh s
Wmu
si j
ha model a de o ma ion
based on he local image s uc u es. Each weigh is compu ed as he p oduc o ou di e en
3.3 Single and Mul i-View Dep h Fusion 49
ac o s. The i s ac o
˜
W1mu
si j =e−√(i−u)2+( j− )2))
σ1,(3.6)
simply measu es p oximi y based on he dis ance o he pixels
(i,j)
and
(u, )
. The pa ame e
σ1
con ols he adius o p oximi y o each poin . The emainde h ee ac o s depend on he
s uc u e o he single- iew p edic ion s. The second ac o
˜
W2mu
si j =1
|∇xsu −∇xsi j|+σ2
·1
|∇ysu −∇ysi j|+σ2
(3.7)
measu es he simila i y o dep h g adien s and assigns la ge weigh s o simila ones.
∇xsi j
and
∇ysi j
ep esen he dep h g adien in he
x
and
y
di ec ion espec i ely a he pixel
(i,j)
.
σ2
limi s he in luence o a poin o a oid ex emely high weigh s o e y simila o iden ical
g adien s. We se i o 0.1 in he expe imen s.
Finally, he ac o s
˜
W3mu
si j
and
˜
W4mu
si j
s eng hen he in luence be ween he poin s lying in
he same plane and a e de ined as
˜
W3mu
si j =e−|(si j+∇xsi j·(u−i))))−su |+σ3(3.8)
and
˜
W4mu
si j =e−|(si j+∇ysi j·( −j))))−su |+σ3,(3.9)
whe e
σ3
se s a minimum weigh o any poin in
Ω
. This is equi ed o a oid anishing
weigh s when hey a e combined wi h ˜
W1mu
si j and ˜
W2mu
si j .
The p oduc o his ou ac o makes a non-no malized weigh o each pixel in Ω
˜
Wsi j
mu =
4
∏
n=1
˜
Wn
si j
mu (3.10)
and ep esen s i s a ea o in luence. The pa ame e s
σ1
,
σ2
and
σ3
shape he a ea o
in luence and ha e o be selec ed o balance p oximi y, g adien and plana i y and o a oid
discon inui ies in he esul o he usion. This was done empi ically on a small se o h ee
images. The alues o he pa ame e s a e
15
,
0.1
and
1e−3
, espec i ely, and we kep hem
ixed o all ou expe imen s.
Fig. 3.3 shows his a ea o a poin on an image and how i is compu ed. No ice how
he in luence expands a ound he poin bu is kep inside he same local s uc u e ( he able).
Once all he ac o s has been compu ed, since all he pixels
(i,j)
a e in luenced by all he

50 Combining Single-View Deep Lea ning Dep h wi h Mul i-View Dep h
RGB image wi h he poin Weig h ac o s o he poin Non-no malized in luence
Fig. 3.3 Non-no malized in luence o he highligh ed ed poin in he image. Fi s column:
RGB inpu image wi h a ed poin o e he able, his poin ep esen one pixel es ima ed by
he mul i- iew algo i hm. Second column: each one o he weigh s calcula ed sepa a ely, he
hi d and ou h weigh s a e shown as a p oduc o a mo e in ui i e iew. Thi d column: Non-
no malized in luence o he highligh ed poin in he RGB image. No ice how i s in luence is
cu on he edge o he able. Figu e bes iewed in elec onic o ma .
pixels in
Ω
(see Eq. 3.5), we no malize he weigh s o each single- iew pixel so all he
weigh s o e a pixel (i,j)sum 1.
Wmu
si j =
˜
Wmu
si j −min(g,h)∈Ω˜
Wmgh
si j
∑(p,k)∈Ω˜
Wmpk
si j −min(g,h)∈Ω˜
Wmgh
si j
(3.11)
The no malized weigh s expand he local in luence o he whole image (see Fig. 3.4 and
Fig. 3.5 o a mo e de ailed iew). No ice how he in luence expands along planes e en i
he poin s in
Ω
do no each he end o he plane; and is sha ply educed when he local
s uc u e changes. Once hese in luence weigh s ha e been calcula ed and no malized, he
usion dep h es ima ion,
, o each poin
(i,j)
is a combina ion o all he selec ed poin s in
Ω, as p esen ed in Eq. 3.5.
3.3.4 Mul i- iew Low-E o Poin Selec ion
Up o now we ha e assumed ha all he poin s in he mul i- iew semi-dense dep h map
Ω
ha e low e o . This is easily achie able in high-pa allax sequences by using obus
es ima o s – obus cos unc ions o RANSAC. Howe e , i is p oblema ic o he degene a e
o quasi-degene a e low-pa allax geome ies ha we also a ge in his wo k. In his case,
mul i- iew dep hs may con ain la ge e o s ha will p opaga e o he used dep h map and
i is necessa y o il e hem ou . Unexpec edly, selec ing high g adien pixels was no
obus enough o emo e poin s wi h la ge dep h e o s and we ha e de eloped a wo s ep
3.3 Single and Mul i-View Dep h Fusion 51
Fig. 3.4 No malized in luence a ea o he poin s. No ice how i expands a ound local
s uc u e a eas gi en a se o poin s in
Ω
.Fi s column: RGB image wi h he poin s o
Ω
labeled wi h di e en colo s. Second column: in luence a eas compu ed by ou me hod.
No ice how his in luence expands in a eas wi h he same local s uc u e bu can be misled
in a eas whe e he e is a lack o poin s o whe e he es ima ion om he neu al ne is no
accu a e enough. Figu e bes iewed in colo .
algo i hm ha akes in o accoun pho ome ic and geome ic in o ma ion in he i s s ep and
he single- iew dep h map in he second one.
The i s s ep selec s a ixed pe cen age o he bes co espondence candida es – he bes
25%
in ou expe imen s– based on he p oduc o a pho ome ic and a geome ic sco es.
On one hand, he pho ome ic c i e ion ocuses on he quali y o he co espondences using
image in o ma ion. We apply a modi ied e sion o he second bes a io.We i s ex ac
he wo closes ma ches o a pixel (smalles pho ome ic e o s acco ding o Eq. 3.3). We
hen compu e he sco e as a unc ion o he a io be ween he dis ance o he wo desc ip o s
(a high a io sugges ing a good ma ch) and he g adien o he dis ance unc ion along he
epipola line (i.e., he e o unc ion p esen ing a dis inc V-shape a ound his ma ch and
52 Combining Single-View Deep Lea ning Dep h wi h Mul i-View Dep h
Fig. 3.5 De ail o he in luence a ea. No ice how i expands mainly in he a eas wi h same
local s uc u e. Figu e bes iewed in colo .
sugges ing spa ial accu acy). On he o he hand, he geome ic sco e simply backp opaga es
he image co espondence e o o he dep h es ima ion, esul ing in low sco es o low-
pa allax co espondences.
In a second s age we also use he s uc u e o he single- iew econs uc ion and apply
RANSAC o es ima e a spu ious- ee linea ans o ma ion be ween he mul i and single- iew
poin s using only he poin s p e- il e ed in he i s s age. We apply his linea model along
he en i e image, consensus wi h ou lie s is ound i small pa ches a e used. This educes
u he he numbe o spu ious dep h alues om he mul i- iew algo i hm. The esul is a
small se o low-e o poin s ha we use o he in e pola ion o he p e ious sec ion. As
men ioned be o e, in ou expe imen s his algo i hm beha es be e han a geome ic-only
compa ibili y es , especially in he low-pa allax sequences o he NYU 2 da ase .
3.4 Expe imen al Resul s
In his sec ion we e alua e he algo i hm and compa e i s pe o mance agains wo s a e-o -
he-a me hods: mul i- iew di ec mapping using TV egula iza ion (implemen ed ollowing
Newcombe e al. (2011), Handa e al. (2011)) and he single- iew dep h es ima ion using he
ne wo k o Eigen and Fe gus (2015). We ha e selec ed wo da ase s wi h di e en p ope ies.
The i s one is he NYU 2 Dep h Da ase (Na han Silbe man and Fe gus, 2012), a gene al
da ase aimed a image segmen a ion e alua ion and hence likely o con ain low-pa allax
and low- ex u e sequences. We analyze esul s in six sequences om he es se (i.e. he
single- iew ne had no been ained on hese sequences) selec ed jus o include di e en
3.4 Expe imen al Resul s 53
RMSE SCALE INVARIANT MEAN ERROR (m) MEAN ERROR (m)
Sequence TV Eigen Ou s(a) TV Eigen Ou s(a) TV Eigen Ou s(a) Ou s(m)
NYUDep h 2
ba h_0018 1.458 0.852 0.793 0.405 0.150 0.145 1.174 0.692 0.612 0.263
bed_0013 1.004 0.550 0.482 0.212 0.139 0.136 0.690 0.441 0.344 0.163
d _0032 2.212 0.710 0.694 0.416 0.209 0.204 1.797 0.581 0.554 0.318
ki _0032 3.599 1.621 1.572 0.812 0.592 0.583 2.920 1.222 1.183 0.805
l _0025 1.073 0.620 0.597 0.289 0.236 0.219 0.798 0.471 0.435 0.289
l _0030a 1.031 0.818 0.792 0.411 0.228 0.219 0.849 0.532 0.440 0.329
TUM
1_desk 1.581 0.433 0.410 0.255 0.121 0.103 1.211 0.317 0.294 0.154
1_ oom 1.467 0.323 0.301 0.167 0.092 0.081 1.163 0.231 0.207 0.102
Table 3.2 Le able: E o me ics o he NYU 2 and TUM da ase s. Fo each sequence
and me ic we compa e he TV- egula ized mul i- iew dep h, he single- iew dep h Eigen
and Fe gus (2015) and ou used dep h. Ou s(a) ep esen ou p oposal wi h he au omaic
selec ion o poin s. Righ able: Mean e o o he used dep h wi h manual mul i- iew
poin selec ion (Ou s(m)); selec ed poin s unde ce ain h eshold. (The e alua ion has been
pe o med in he i s 100 ames o each sequence)
ypes o ooms. The second one is he TUM RGB-D SLAM Da ase (S u m e al., 2012a), a
da ase o ien ed o isual SLAM and hen likely o p esen a bias bene i ing mul i- iew dep h.
In his case, we e alua ed wo sequences selec ed andomly.
We un ou algo i hm in a
320×240
subsampled e sion o he images, as his is he size
o he single- iew neu al ne wo k gi en by he au ho s. We also un ou mul i- iew dep h
es ima ion a his image size, and upsample he used dep h o
640×480
in o de o compa e
i agains he g ound u h D channel om he kinec came a.
As ou aim is o e alua e he accu acy o he dep h es ima ion, we will assume ha
came a poses a e known o he mul i- iew es ima ion. In he TUM RGB-D SLAM Da ase
(S u m e al., 2012a) we use he g ound u h came a poses. In he NYU 2 Dep h Da ase
sequences we es ima e hem using he RGB-D Dense Visual Odome y by Gu ié ez-Gómez
e al. (2015). These came a poses will emain ixed and used o c ea e he mul i- iew dep h
maps. As men ioned be o e, he pa ame e s o he usion algo i hm we e expe imen ally se
p io o he e alua ion on a small sepa a e se o images.
To e alua e he me hods, we compu ed h ee di e en me ics, he RMSE, he Mean
Absolu e E o in me e s and he scale in a ian e o p oposed in Eigen e al. (2014)
1
n∑id2
i−1
n2(∑idi)2
whe e
d
is
(log(y)−log(y∗))
,
y
and
y∗
a e he g ound u h dep h and
he es ima ed dep h espec i ely. The esul s a e summa ized in Table 3.2. Ou me hod
ou pe o ms he TV egula iza ion in bo h da ase s ob aining an a e age imp o emen o e
50% wi h espec o he mean o he e o in me e s. As expec ed, he TV egula iza ion
pe o ms be e in he TUM sequences and achie es lowe e o s, bu in e ms o imp o emen
he e seems no o be big di e ences be ween bo h da ase s. Ou usion o dep hs also
ou pe o ms he single- iew dep h econs uc ion, he imp o emen being
10%
on a e age.
60 Condi ion-In a ian Place Recogni ion
mul i- iew
desc ip o
-(n-1)
-n
1
desc ip o
da abase
CNN
CNN
CNN
Place Recogni ion
by desc ip o ma ching
Inpu
Sequence
Que y
Nea es Neighbo
Re ie ed Place:
Visi ed Places
Fig. 4.1 O e iew o ou p oposal. We ex ac desc ip o s (using deep ne wo ks) o small
sequences o
n
ames. We use such desc ip o s o ind he closes ma ch in a da abase o
al eady isi ed places.
challenging and esul in lowe pe o mance. Using mul iple ames in a sequence can
imp o e he obus ness o place ecogni ion agains such changes. Bu he sequence models
p oposed by he s a e o he a (Mil o d and Wye h, 2012, Nasee e al., 2018) a e handc a ed
o a ce ain se o assump ions (e.g. o e lapping ajec o ies, simila eloci y pa e ns), and
hei pe o mance su e s i hese a e no hold. Also, ypically, hey equi e a high numbe o
ames.
Desc ip o s di ec ly ex ac ed om CNNs ha e shown good gene aliza ion p ope ies
(Gomez-Ojeda e al., 2015), bu hey usually do no exploi mul i- iew in o ma ion. Im-
p o emen s usually come a he cos o la ge desc ip o s, wi h dimensionali y in he o de
o housands o hund eds o housands. The complexi y o all place ecogni ion algo i hms
depends on he size o he desc ip o and he numbe o images in he da abase, he la es
being ypically high. This limi s he applicabili y o hese echniques in obo ics and AR/VR
scena ios, in which he compu a ional budge is limi ed due o eal- ime cons ained loops
and limi ed on-boa d compu a ional powe .
In his hesis we a ge place ecogni ion in he p esence o challenging changes in he
condi ion o an en i onmen , ha e en ually happen in mos o he scenes as ime passes.

4.2 The Pa i ioned No dland Da ase 61
Fo example day/nigh illumina ion, seasonal and wea he changes, o objec s ha a e mo ed
(ca s, pe sons o u ni u e). We add ess his p oblem by gene a ing a global desc ip o o he
isual inpu , i.e., e e y inpu (one o se e al images) is encoded as a desc ip o and ma ched
e sus he es o he desc ip o s e ie ing he closes place.
Ou con ibu ion is he p oposal and e alua ion o h ee di e en deep ne wo k a chi-
ec u es ha exploi mul i- iew and empo al in o ma ion o place ecogni ion: 1) naï e
desc ip o g ouping, 2) lea ning he usion o single- iew desc ip o s, and 3) ecu en ne -
wo ks using LSTM (Long Sho Te m Memo y) laye s Hoch ei e and Schmidhube (1997).
Fig. 4.1 shows an o e iew o ou p oposal. Up o ou knowledge, ou s a e he i s models
ha use deep lea ning o combine mul iple iews o gene a e desc ip o s o place ecogni ion.
Encoding empo al in o ma ion allows us o model sho - ime ela ions be ween images (e.g.
sho sma phone ideos o li e pho os) wi hou he need o keeping a global map o long
image sequences o achie e a high accu acy.
We compa e ou models o s a e-o - he-a single- iew deep baselines and a non-deep
sequen ial one using wo s anda d da ase s: he Pa i ioned No dland (Olid e al., 2018) and
Alde ley (Mil o d and Wye h, 2012). The expe imen al esul s show ha he pe o mance o
ou h ee mul i- iew desc ip o s ou pe o ms single- iew ones. We also ou pe o m SeqS-
LAM, a s a e-o - he-a baseline o place ecogni ion om image sequences. Fu he mo e,
ou lea ned desc ip o s a e a leas one o de o magni ude smalle han hose o he s a e
o he a showing ha mul i- iew lea ning is able o ex ac ele an in o ma ion o place
ecogni ion.
The es o he chap e is o ganized as ollows. Sec ion 4.2 p esen s a pa i ion o he
No land da ase ha we ha e used in his chap e ha was published in i is a ailable in
ou p ojec websi e
1
. Sec ion 4.3 e e s he ela ed wo k. Sec ion 4.4 de ails ou ne wo k
a chi ec u es, and sec ion 4.5 de ails how hey a e ained. Finally, sec ion 4.6 p esen s he
expe imen al esul s and sec ion 4.7 he conclusions and lines o u u e wo k. Ou
code
and
a ideo showing ou esul s can be ound in ou p ojec websi e 2.
4.2 The Pa i ioned No dland Da ase
Fo his wo k, we ha e used he No dland ail oad ideos. In 2012, he No way b oadcas ing
company (NRK) made a documen a y abou he No dland Railway, a ailway line be ween
he ci ies o T ondheim and Bodø. They ilmed he
729
km jou ney wi h a came a in he on
1h ps://webdiis.uniza .es/ jm acil/p -no dland/
2h p://webdiis.uniza .es/~jm acil/cim p /
62 Condi ion-In a ian Place Recogni ion
aining se
es se
disca ded da a
Fig. 4.2 P oposed da ase pa i ion o he No dland da ase .
Top:
Geog aphical ep e-
sen a ion o he aining ( ed) and es (yellow) se s.
Bo om:
Index ep esen a ion o he
dis ibu ion, w. . . ame index in he ideos.
pa o he ain in win e , sp ing, all and summe . The leng h o each ideo is abou 10
hou s and each ame is imes amped wi h he GPS coo dina es.
This da ase has been used by o he esea ch g oups in place ecogni ion, o example
Gomez-Ojeda e al. (2015) and Low y and Mil o d (2016). Each g oup uses di e en
pa i ions o aining and es , making di icul o ep oduce he esul s. In his wo k we
4.3 Rela ed Wo k 63
p opose a speci ic pa i ion o he da ase and a baseline, o gua an ee a ai compa ison
be ween algo i hms. We eleased his da ase pa i ion wi h he publica ion o Olid e al.
(2018).
4.2.1 Da a P e-p ocessing
The i s s ep, c ea ing he da ase , was o ex ac he maximum numbe o images om each
ideo. Mo eo e , GPS da a co up ion was ixed and we also elimina ed unnels and s a ions.
A e hese s eps, g abbing one ame pe second, we ob ained
28,865
images pe ideo. We
used speed in o ma ion om he GPS da a o il e s a ions and a da kness h eshold o il e
unnels.
4.2.2 Da ase Pa i ions
Fig. 4.2 illus a es he pa i ion o he whole image se in he No dland da ase . We decided
o c ea e he es se wi h h ee di e en sequences o
1,150
images (a o al o
3,450
, yellow
in he igu e). The es o he images we e used o aining (
24,569
, ed in he igu e). By
using mul iple sec ions, he a ie y o places and appea ance changes con ained in he es se
inc eases. We also le a sepa a ion o a ew kilome e s be ween each es and ain sec ion
by disca ding some images in o de o gua an ee he di e ence be ween es and ain da a.
4.2.3 Place labels
Gi en he simila i y be ween consecu i e images, in his wo k we p opose o conside ha
wo images a e o he same place i empo ally hey a e sepa a ed by
3
images o less. We
applied a sliding window o
5
images o e he whole da ase in o de o g oup images aken
om i e consecu i e seconds. This p ocess can be seen in Fig. 4.3.
4.3 Rela ed Wo k
The e ha e been many wo ks add essing isual place ecogni ion and ela ed p oblems. Fo
a gene al o e iew, we e e he eade o wo su eys, Ga cia-Fidalgo and O iz (2015) on
opological mapping and Low y e al. (2016) exclusi ely o isual place ecogni ion. In his
sec ion we will ocus on he wo ks ha a e mos ele an o ou p oposal. We will e iew
i s he li e a u e on place ecogni ion desc ip o s, and la e e e o ull place ecogni ion
pipelines. No ice ha ou con ibu ion lies mainly on he o me , ha is, he p oposal o
no el mul i- iew desc ip o s.
64 Condi ion-In a ian Place Recogni ion
place 1
place 2
place 3
Fig. 4.3 A sliding window o i e images is conside ed in his wo k as he same place. No ice
he simila i y o consecu i e images. The igu e is bes iewed in elec onic o ma .
Desc ip o G ouping Desc ip o Fusion Recu en Desc ip o s
CNN
-(n-1)
-n
Inpu Sequence
CNN CNN
+
mul i- iew
desc ip o
(a)
-(n-1)
-n
Inpu Sequence
CNN CNN CNN
c
+
mul i- iew
desc ip o
(b)
-(n-1)
-n
Inpu Sequence
CNN CNN CNN
LSTM LSTM LSTM
mul i- iew
desc ip o
(c)
Fig. 4.4
Mul i- iew desc ip o s
p oposed in his hesis. F om le o igh : (a)
Desc ip o
G ouping
, whe e he desc ip o o a sequence is he conca ena ion o all he single image
desc ip o s. (b)
Desc ip o Fusion
, he ou pu o he CNNs se es as inpu o a ully-
connec ed laye ha combines he in o ma ion in o a single desc ip o . (c)
Recu en
Desc ip o s
, he ou pu o he CNNs se es as inpu o an LSTM ne wo k ha in eg a es
o e ime he single-image ea u es o c ea e a mul i-image desc ip o .
4.3.1 Desc ip o s
Local Desc ip o s
Techniques based on local desc ip o s add ess place ecogni ion by de ec ing a se o salien
keypoin s in an image and gene a ing desc ip o s o each one o hem. These desc ip o s
a e used o ind co espondences in o he images, ha would po en ially allow us o pe o m
isual place ecogni ion o e en 6DoF came a pose eco e y. De ec ion and desc ip ion a e
usually decoupled applying a me hod (Ha is e al., 1988, Lowe, 2004, Mikolajczyk e al.,
4.3 Rela ed Wo k 65
2005) ha de ec s salien poin s and hen gene a e e e y a desc ip o o e e y poin using
(Bay e al., 2006, Lowe, 2004, Rublee e al., 2011). Recen ly, some deep-lea ning app oaches
ha e add essed desc ip o gene a ion (Luo e al., 2018, 2019), keypoin de ec ion (Sa ino
e al., 2017, Ono e al., 2018) o bo h in an end- o-end manne (Re aud e al., 2019, Dusmanu
e al., 2019).
Global Desc ip o s
Despi e he ad an ages o local desc ip o s, global desc ip o s come handy when he asks
equi es la ge con ex . One example o his is in he p esence o la ge appea ance changes
like wea he o illumina ion condi ions. On hose occasions, when local pa e ns migh
change subs an ially, a global iew o he scene cap u ed by high le el ea u es (e.g. he
skyline o he ci y) may be mo e help ul.
T adi ional global codes include handc a ed holis ic image desc ip o s, like low- esolu ion
humbnails (Mil o d and Wye h, 2012) o GIST (Mu illo e al., 2013). None o hese a e
obus o appea ance changes due o scene dynamics, seasonal and wea he changes, o
ex eme iewpoin o ligh ing a ia ions. To add ess such cases, Low y and Mil o d (2016)
used PCA o educe he dimensionali y o desc ip o s elimina ing he dimensions ha a e
in luenced by condi ion changes. Chen e al. (2018) inco po a es a en ion in o de o ocus
on he mos ele an image ea u es o place ecogni ion. Desc ip o s based on CNNs ha e
shown a high deg ee o obus ness agains appea ance changes. Sünde hau e al. (2015a) and
Sünde hau e al. (2015b) showed ha CNNs ou pe o m o he models, especially o d as ic
appea ance changes. They used AlexNe (K izhe sky e al., 2012), p e ained on ImageNe
(Russako sky e al., 2015). The ea u es o AlexNe con ain seman ic in o ma ion abou he
whole scene, which imp o es he in a iance o ce ain appea ance changes. The ea e , may
o he wo ks ha e s udied CNNs as condi ion-in a ian ea u e ex ac o s (Gomez-Ojeda e al.,
2015, A andjelo ic e al., 2016, A oyo e al., 2016, Chen e al., 2017, Lopez-An eque a
e al., 2017, Olid e al., 2018). Gomez-Ojeda e al. (2015) we e he i s ha ained a ne wo k
as single-image ea u e ex ac o o isual place ecogni ion unde appea ance changes. In
Ne VLAD (A andjelo ic e al., 2016), hey p oposed a new ype o laye inspi ed in VLAD,
an image ep esen a ion commonly used in image e ie al. Chen e al. (2017) p oposed
a ne wo k ained o classi y he place he image was aken. Olid e al. (2018) p oposed a
model based on p e- ained VGG-16 Simonyan and Zisse man (2014) and ine- uned i o
he place ecogni ion ask in a T iple -Siamese a chi ec u e.

66 Condi ion-In a ian Place Recogni ion
4.3.2 Visual Place Re ie al
This g oups he app oaches used o ind e ie e he igh place o e e y image. Tha comes
down o he ma ching algo i hms uses o e ie e he igh image om he da abase o isi ed
places.
Single-View Place Recogni ion
These a e all hose app oaches which goal is o e ie e a single image om a single image;
igno ing he ime o pose ela ion be ween di e en ames in he que y o in he da abase.
The li e a u e add esses he nea es neighbo p oblem in place ecogni ion in many di e en
ways: b u e o ce (Olid e al., 2018, Mu illo e al., 2013), KD- ee (Mu illo e al., 2013) and
bag o wo ds (Gál ez-López and Ta dos, 2012). Low y and And easson (2018) p esen ed
a model using SURF de ec o and HOG ea u es and s udies he use o Bag o Wo ds and
Vec o s o Locally Agg ega ed Desc ip o s (VLAD) o place ma ching.
Mul i-View Place Recogni ion using Single-View Desc ip o s
Fo place ecogni ion, i can be assumed ha images in he da abase a e independen and all
ha can be done is 1- o-1 ma ching, o else ha you know he links be ween he images (in
he que y sequence and/o in he da abase) an use ha as a p io knowledge. Al hough he e
a e only a ew wo ks ha conside empo al and mul i- iew in o ma ion o place ecogni ion,
hey all ha e shown ha sequences p o ide use ul in o ma ion o place ecogni ion. Fo
ins ance, DBoW (Gál ez-López and Ta dos, 2012) and Bampis e al. (2016) inco po a e
a empo al consis ency cons ain . SeqSLAM (Mil o d and Wye h, 2012) and ollowing
wo ks (Peppe ell e al., 2014) use sequence ma ching, simila ly o Newman e al. (2006).
Di e en ly o ou app oach, hey assume linea empo al co ela ion o sequence ma ching.
Also, we use da a-d i en high-le el ea u es, while hey use downsampled images.
Mo e ecen ly, some wo ks ha e ex ended SeqSLAM in se e al aspec s. On he one hand,
Chen e al. (2014) e ines he empo al il e ing. On he o he hand, Nasee e al. (2018)
p oposes a g aph o single- iew desc ip o s (based on HOG and AlexNe ) o model and
ma ch image sequences. Thei app oach is simila o SeqSLAM, wi h wo main di e ences.
The mos s aigh o wa d one is ha hey use di e en desc ip o s. The second one, mo e
sub le, is ha hei sea ch o he bes -ma ching sequence does no assume a cons an speed
a ia ion be ween he sequences. SeqSLAM looks o s aigh lines in he simila i y ma ix,
while Nasee e al. (2018) uses a mo e sophis ica ed model. In any case, none o hem
model changes in he sequence di ec ion. Also, hey ypically ely on long- e m sequence
ma ching (i.e., que y and da abase sequences ha ing many consecu i e ma ching ames),
4.4 Ne wo k A chi ec u es 67
Me hod Desc ip o Size Accu acy a g
%
ou s (single- iew) 128 80.19%
ou s (single- iew) 256 80.62%
ou s (single- iew) 512 80.65%
ou s (single- iew) 1024 81.46%
he bigge he be e
Table 4.1 E alua ion o di e en desc ip o sizes on he
Pa i ioned No dland Da ase
(Olid e al., 2018) in oduced p e iously in his chap e .
1s column:
Me hod.
2nd column:
Desc ip o size in 32-bi s loa ing poin numbe s. 3 d column: A e age accu acy.
which limi s hei applicabili y o such case. Assuming ha consecu i e ames ha e simila
appea ance, Neube e al. (2015) combines CNN single- iew desc ip o s wi h a di ec ed
sea ch. Con inuing hei p e ious wo k Vyso ska and S achniss (2016) and Vyso ska and
S achniss (2017) p opose a combina ion o hei lazy da a associa ion wi h a hash educ ion
o single- iew CNN ea u es.
4.3.3 Mul i-View Place Recogni ion using Mul i-View Desc ip o
All he mul i- iew models desc ibed so a a e handc a ed. Up o ou knowledge, ou s a e
he i s ones ha lea n o combine mul iple single- iew ea u es maps in o a mul i- iew
desc ip o . We compa e h ee di e en app oaches o gene a e a desc ip o based on mul iple
images, simple conca ena ion as baseline, lea ning o use desc ip o s and using a ecu en
neu al ne wo k o accumula e he knowledge o e ime. The adi ional app oach o using
mul iple iews in place ecogni ion is adding ex a cons ain s when looking o he nea es
neighbo . In ou case, we do no add any cons ain s bu gene a e desc ip o s ha al eady
include empo al and/o spa ial in o ma ion.
4.4 Ne wo k A chi ec u es
In his sec ion, we discuss ou di e en models o place ecogni ion: A single- iew one,
based on ResNe -50, and he h ee mul i- iew ones p oposed in his chap e .
4.4.1 Single-View ResNe -50
Ou i s ne wo k is based on he model p esen ed in Olid e al. (2018). The main di e ence
is ha we s a om ResNe -50 (He e al., 2016) p e ained on ImageNe (Russako sky e al.,
2015) as ou backbone, ins ead o VGG-16 (Simonyan and Zisse man, 2014). Al hough i
68 Condi ion-In a ian Place Recogni ion
is common o di ec ly use he desc ip o s o di e en laye s (see Sec ion 4.6.1 o esul s
on his), in ou case we added and ained a ully connec ed laye a e ResNe -50 o lea n a
128
-dimensional desc ip o especially designa ed o he ask o isual place ecogni ion. We
chose a size o
128
expe imen ally (see Table 4.1) , as a easonable comp omise be ween
pe o mance and compaci y.
4.4.2 Desc ip o G ouping
In o de o include empo al in o ma ion in o he desc ip o s, ou i s app oach is he naï e
conca ena ion o he desc ip o s o consecu i e ames, see Fig. 4.4a. Thus, s a ing om
ou p e ious single- iew model, we i s choose a empo al window o ames (
n
), we hen
gene a e a
128
-dimensional desc ip o pe ame, and we inally conca ena e hem. The
desc ip o size is hen
128 ×n
. No ice ha his model is ained only om single- iew
samples. Hence, he ela ion be ween consecu i e ames is no lea ned and his model only
p o ides a il e ing e ec .
4.4.3 Desc ip o Fusion
Desc ip o G ouping, as he simples s a egy o conside se e al ames, is limi ed in i s
capabili y o weigh di e en ly ce ain ea u es (i.e. ea u es o some o he ames may be
mo e ep esen a i e o he place han o he s). I is also limi ed o cases whe e he sequences
(map/que y) a e aligned – meaning ha bo h sequences ollow he same ajec o y. Fo
ha eason, we designed a model ha lea ns how o use he in o ma ion o ou
n
- ames
window in o a mo e disc iminan –as well as smalle –
128
-dimensional desc ip o . Wi h his
Desc ip o Fusion s a egy, we add an ex a ully connec ed laye ha lea ns how o combine
he ou pu s o
n
ResNe -50 in o a single compac desc ip o . See Fig. 4.4b o an illus a ion
o his app oach. As his ne wo k is able o lea n how o weigh he ea u es om di e en
ames, i can model mo e complex cases. Fo example, when sequences a e eco ded in
e e se o de Desc ip o G ouping is limi ed, while Desc ip o Fusion has he capabili y o
lea ning a sui able usion.
4.4.4 Recu en Desc ip o s
Desc ip o Fusion does no explici ly exploi s he sequen ial na u e o he da a. Wi h
Recu en Desc ip o s, we upda e in an online manne he sequence coding as new ames
come, keeping he mos ele an p e ious in o ma ion. Wi h ha in en ion, we p opose a
Recu en Neu al Ne wo k (see Recu en Desc ip o s in Fig. 4.4c). In his model, e e y
4.5 T aining 69
que y-sequence 1
que y-sequence 2
que y-sequence M
place 1
place 2
place N
Fig. 4.5 Same place con en ion, illus a ed wi h an example whe e he que y-sequence has
a leng h o 3 ames. A place ep esen s a se o ames ha a e conside ed o be on he
same place. No ice ha a ame can be in mo e ha one place. A que y-sequence is an inpu
sequence o ou model. We wan o eco e he co esponding place o a que y-sequence.
ame is he inpu o a ResNe -50, and he op laye s se e as he inpu o a LSTM ne wo k
(Hoch ei e and Schmidhube , 1997), ha gene a es a
128
-dimensional desc ip o . LSTMs
keep an inne s a e, ha is upda ed wi h each inpu ame, and he ou pu depends on he
s a e and he inpu . Di e en ly o p e ious models, keeping a ecu en inne s a e allows
his ne wo k o p oduce a desc ip o om he i s ame, and upda e i sequen ially as mo e
ames a i e.
4.5 T aining
4.5.1 Con en ion o Same Place
Since ou desc ip o is gene a ed om a sequence o images (que y-sequence) ins ead o a
single image, we mus de ine when wo que y-sequence o
n
ames a e conside ed o be a
he same place ( he de ini ion o a place being da ase -dependen ). To illus a e his de ini ion
we will make use o Fig. 4.5. The igu e shows a sequence o ames and se e al examples
o que y-sequence, and also shows he se o ames ha we conside as he same place.
The e o e, du ing aining, we conside wo que y-sequence o be on he same place i hey
76 Condi ion-In a ian Place Recogni ion
Me hod T ained
on
Numbe
o
ames
Desc ip o
Size
D s N
# %
Olid e al. (2018) No land 1 128 0.15%
Olid e al. (2018) Alde ley 1 128 6.84%
ou s (single- iew) No land 1 128 1.65%
ou s (single- iew) Alde ley 1 128 6.8%
ou s (g ouping) Alde ley 3 384 9.05%
ou s (g ouping) Alde ley 6 768 11.48%
ou s ( usion) Alde ley 3 128 10.18%
ou s ( ecu en ) Alde ley 3 128 5.73%
bigges he bes
Table 4.3 Resul s on he
Alde ley Da ase
Mil o d and Wye h (2012),
1s column:
Me hod.
2nd column:
Da ase in which he model i has been ained wi h.
3 d column:
Numbe o
ames used o ecogni ion (e.g. 1 would imply o be single- iew).
4 h column:
Desc ip o
size in 32b loa ing poin numbe s.
5 h column:
Day s Nigh (D s N), que y wi h dayligh
image while e e ence da abase composed by nigh ime images.
Quali a i e esul s
. Fig. 4.9 shows se e al es samples. No ice he inc eased challenge
wi h espec o he No dland da ase , wi h he p esence o se e e illumina ion changes plus
inclusion o a i icial illumina ion and dynamic objec s.
4.6.3 Mul i-View E alua ion
Sequence Speed Changes:
Inspec ing he p e ious esul s (Table 4.2), Desc ip o G ouping
(
ou s (g ouping)
) ained only on single- iew and hen applied on mul i- iew by conca e-
na ion is he bes pe o ming. This is su p ising a i s sigh , as he o he wo models we e
ained on mul i- iew da a. We designed wo ex a expe imen s (
Re e se Gea
and
Random
Speed
) o illus a e why his is happening. Bo h expe imen s a e pe o med du ing es ime,
which means none o he ne wo ks has been e ained.
The
Re e se Gea
expe imen consis on changing he di ec ion o he ain mo ion on
one o he sequences a es ime (e.g. when es ing Win e s Fall, he sequence o Fall is
played in e e se o de , see Fig 4.10a). This expe imen will help o disce n how much he
model exploi s he mul i- iew in o ma ion a he han jus he sequence consis ency. Table
4.4 shows ha , as we expec ed, models ained wi h mul i- iew examples (
ou s ( usion)
and
ou s ( ecu en )
) ha e lea ned o exploi mul iple iews: I s pe o mance only deg ades by
6%
and
4%
espec i ely. On he o he side,
ou s (g ouping)
d ops se e ely i s pe o mance,
by 18%.

4.6 Expe imen al Resul s 77
Re e se Gea
3385 3390 3395 3400 3405 3410 3415
3415 3410 3405 3400 3395 3390 3385
Summe
Fall
(a)
Random Speed
130 131 133 136 140 145 151
130 135 140 144 148 151 153
Win e
Sp ing
(b)
Fig. 4.10
Expe imen se up de ails.
(a)
Re e se Gea
, in which he sequence is played in
e e se o de o one o he seasons (Fall in he igu e). (b)
Random Speed
, in which he
ehicle speed is modi ied o bo h seasons, e e ence (Win e ) and que y (Sp ing). In bo h
cases, we ma k wi h a dashed g een box he same-place h ee- ames sequences.
In he
Random Speed
expe imen we syn he ically modi ied he speed o he ain mo ion
on one o he sequences a es ime. Speci ically, we modi ied he ame a e along he
sequence simula ing changes on he ain eloci y, see Fig. 4.10b (in ou expe imen s he
eloci y was andomly mul iplied by
×1
,
×2
o
×3
a e e y momen o he sequence). The
“speed” is modi ied o he whole sequence, implying ha he one- o-one co espondence in
plain No dland does no hold. Table 4.4 p o es ha
Random Speed
is he mos challenging
se up o he
ou s (g ouping)
app oach, d opping i s accu acy o
36%
.
ou s ( usion)
and
78 Condi ion-In a ian Place Recogni ion
Me hod Numbe
o
ames
No mal
Tes
Re e se
Gea
Random
Speed Mean±
S d
# % % % %
SeqSLAM 3 33% 0.08% 9% 14.0±13.9
ou s (g ouping) 392% 74% 36% 67.3±23.3
ou s ( usion) 3 86% 80% 78% 81.33±3.4
ou s ( ecu en ) 3 86% 82% 84%84.0±1.6
bigges he bes
Table 4.4 Expe imen al esul s o
Speed Changes
in he
Pa i ioned No land Da ase
Olid e al. (2018). Compa ison o all ou mul i- iew me hods and SeqSLAM p esen ed by
Mil o d and Wye h (2012).
1s column:
Me hod.
2nd column:
Numbe o ames used o
ecogni ion.
3 d-5 h column:
Summe s Win e expe imen s: 3
d
column:
No mal Tes
co esponds o he one showed on Table 4.2. 4
h
column:
Re e se Gea
expe imen whe e
he que y ames a e all in e e sed o de , i.e. simula ing he ain has used a e e se gea .
5
h
column:
Random Speed
expe imen whe e he speed o he ain is simula ed o be
andom, which means some o he ames a e los ). The speed a ia ions a e independen
o he que y and he e e ence da abases, and his implies no mo e 1 o 1 co espondence.
These speed changes a e pe o m only on he es sequences, his implies ha o a ne wo k,
e.g.
ou s (g ouping)
, esul s on he expe imen s a e achie ed using he same se pa ame e s
han in Table 4.2 wi h no e aining.
ou s ( ecu en )
keep i s pe o mance a a e y simila le el han he s anda d No dland
se up (
78%
and
84%
espec i ely). No ice ha he goal o his expe imen is no c ea ing a
e y ealis ic se up, bu misaligning he que y sequence and he da abase ones. The aim is o
e alua e he gene aliza ion o e a plain ame- o- ame il e ing e ec .
We un he s a e-o - he-a mul i- iew baseline SeqSLAM (Mil o d and Wye h, 2012)
in bo h expe imen s,
Re e se Gea
and
Random Speed
, obse ing ha i s pe o mance
d ops in bo h. This should be expec ed, as SeqSLAM assumes a linea ela ion be ween he
eloci ies o he que y and he e e ence sequences (sequence consis ency).
The las column o Table 4.4 (
Mean/S d
) summa izes he conclusions o bo h expe -
imen s, epo ing he mean (bigges he bes ) and s anda d de ia ion (smalles he bes )
o all expe imen s (
No mal Tes
,
Re e se Gea
and
Random Speed
). Obse e ha
ou s
( ecu en )
is he bes pe o ming, p esen ing bo h he highes a e age accu acy and smalles
a ia ions. This con i ms ou hypo hesis: The sequence desc ip o s ha use lea ning (
ou s
( usion)
and (
ou s ( ecu en )
) a e mo e esilien han hose based on plain conca ena ion
(ou s (g ouping)) o handc a ed ela ions (SeqSLAM).
4.6 Expe imen al Resul s 79
Me hod Num
o
ames
Desc ip o
Size
Accu acy
W s S
No land
Accu acy
S s W
No land
Accu acy
D s N
Alde ley
# % %
SeqSLAM 3 6144 31% 33% 3.91%
SeqSLAM 10 20480 71% 70% 9.90%
SeqSLAM 100 204800 95% 94% -
ou s (g ouping) 3 384 92% 92% 9.05%
ou s (g ouping) 6 768 97% 97% 11.48%
ou s ( usion) 3128 87% 86% 10.18%
ou s ( ecu en ) 3128 85% 86% 5.73%
ou s ( ecu en ) 6 128 87% 88% -
he bigge he be e
Table 4.5 Resul s on he
Pa i ioned No dland Da ase
(Olid e al., 2018) and
Alde ley
adding mo e ames in he o he sea ch. Compa ison o all ou mul i- iew me hods and
SeqSLAM p esen ed by Mil o d and Wye h (2012).
1s column:
Me hod.
2nd column:
Numbe o ames used o ecogni ion (e.g., 1 s ands o single- iew).
3 d column:
De-
sc ip o size in 32-bi s loa ing poin numbe s.
4 h column:
Win e s Summe , que y aken
om win e and ma ched o summe da abase.
5 h column:
Summe s Win e , que y aken
om summe and ma ched o win e da abase.
6 h column:
Day s Nigh , que y aken om
nigh and ma ched o day da abase.
Las , in Table 4.5 we show a compa ison be ween ou me hods and SeqSLAM (Mil o d
and Wye h, 2012) on bo h No land and Alde ley. We also e alua e he e ec o conside ing
mo e ames in he di e en me hods.
Ou me hods ou pe o m SeqSLAM (Mil o d and Wye h, 2012), a s a e-o - he-a baseline
able o model in o ma ion om se e al ames, when bo h use he same numbe o ames
(speci ically,
3
). As all mul i- iew app oaches imp o e hei pe o mance when inc easing
he numbe o ames, we inc eased he numbe o ames used by SeqSLAM. No ice ha , in
o de o ou pe o m ou app oach in No land, he numbe o ames has o be inc eased up
o
100
and
10
in Alde ley. Simila ly occu s when inc easing he numbe o ames used by
ou s.
4.6.4 Execu ion ime
We compa ed he execu ion ime o ou models in he uppe pa o Table 4.6. The ou h
column (
Desc ip o Ex ac ion
) shows he ime needed o ex ac he desc ip o o a que y
3- ames sequence on a NVIDIA TITAN Xp. In his pa o ou pipeline Desc ip o G ouping
is he as es me hod, as is uses he simples ne wo k.
80 Condi ion-In a ian Place Recogni ion
Me hod Desc ip o
Size
Desc ip o
Ex ac ion
Sea ch
1 s 10K
ms ms
ou s ( usion) 128 17 3.86
ou s ( ecu en ) 128 22 3.86
ou s (g ouping) 384 15 10.70
- 1860 - 49.21
- 4096 - 111.44
- 6144 - 166.13
- 20480 - 688.24
- 204800 - 9279.62
smalles he bes
Table 4.6
Execu ion Time
o all ou models.
1s column:
Me hod.
2nd column:
Desc ip o
size.
3 d column:
Time in milliseconds needed o ex ac 1 desc ip o .
4 h column:
Gi en a
desc ip o and a e e ence da a base o
10K
desc ip o s, ime in milliseconds needed o ind
he bes ma ch.
Las column (
Sea ch
) shows he ime needed o ind he bes ma ch (Nea es Neighbo
(NN)) gi en a que y and a da abase o
10K
desc ip o s. No ice ha ou me hods Desc ip o
Fusion and Recu en Desc ip o s a e as e . This was expec ed, as hei desc ip o sizes
a e
n
(que y-sequence size) imes smalle (3 imes in ou expe imen s). Ou NN algo i hm
consis s on an exhaus i e sea ch h ough he da abase. We i e a e o e all he isi ed places,
compu e he dis ance be ween hei desc ip o s and he que y and keep he minimum-dis ance
one. Fo he dis ance unc ion we use he Squa ed Euclidean Dis ance (
d2(dq,di)
). The
compu a ional complexi y o his sea ch is
O(N)
whe e
N
is he numbe o elemen s in he
da abase and o he dis ance unc ion O(k)whe e kis he desc ip o size.
Addi ionally, we compu ed he sea ch ime co esponding o he sizes o some o he
o he desc ip o s used in Tables 4.2 and 4.3 (bo om pa o he able). As expec ed, he
ime inc eases wi h he desc ip o size. No ice ha high dimensional desc ip o s ule ou he
use o mo e e icien da a s uc u es, such as KD- ees, o speed up he sea ch echniques,
since i is no possible o ejec candida es by using he di e ence o a single coo dina e
(Ma imon and Shapi o, 1979). A di ec ed sea ch using he sequen iali y o da a (Neube
e al., 2015, Vyso ska and S achniss, 2016) would educe he numbe o compa isons.
Hashing (Vyso ska and S achniss, 2017) can also educe he compu a ional cos . In any
case, educing he dimensionali y o he desc ip o has a di ec in luence in he cos o all he
app oaches men ioned.
4.7 Conclusions 81
4.7 Conclusions
In his chap e we ha e in oduced h ee deep lea ning-based mul i- iew global desc ip o
models, ha ou pe o m exis ing baselines bo h in accu acy and compac ness. We analyzed
di e en app oaches o combine he in o ma ion o he ea u es om mul iple iews (G oup-
ing, Fusion and Recu en ), and we e alua ed hem on di e en expe imen al se ups in wo
public da ase s: Pa i ioned No land and Alde ley. Each model we p opose has i s own
s eng hs and weaknesses. On he one side, Desc ip o G ouping ensu es he sequen ial
consis ency o he ame in a sequence, achie ing he bes pe o mance in he s anda d
No dland/Alde ley benchma ks, whe e he in e - ame mo ion is simila in di e en uns. On
he o he side, Desc ip o Fusion and Recu en Desc ip o s a e able o lea n mo e complex
ela ions be ween ames and hence p o ed o be be e in cases whe e he eloci ies di e
o he ames o de ing is di e en . We also show he low compu a ional cos o all he
app oaches, demons a ing i s po en ial o obo ic applica ions.
We belie e ha ecu en models, in spi e o challenges associa ed o aining and
gene alizing o di e en ehicle dynamics, a e p omising. They a e obus o speed changes
and adap o di e en sequence leng hs wi hou inc easing he desc ip o and ne wo k sizes.
We ha e obse ed such challenges in p elimina y esul s, along wi h a sligh pe o mance
imp o emen wi h an inc ease o he sequence leng h (up o 6 ames). Encou aged by his,
ou u u e wo k will s udy ca e ully he use o ecu en models wi h la ge image sequences.
Howe e ha p esen s he challenge o he beginning and he end o he sequence being wo
place oo a apa .

Chap e 5
Co ne P edic ion o Layou
Recons uc ion
Ano he ele an cue o isual 3D econs uc ion is he ex ac ion o high-le el seman ic
in o ma ion co esponding o he s uc u e o he scene. Fo example, de ec ing and es ima ing
dep h and o ien a ion o plana s uc u es (Concha and Ci e a, 2015a). In his chap e , we
ocus on de ec ing he main s uc u e o indoo scenes. Also e e eed as layou eco e y,
ha is de ec ing he walls, ceiling and loo . We ha e no applied ou ad ances di ec ly o
isual localiza ion o mapping o sequences bu we e e he eade o Salas e al. (2015) as
an example o how his ad ances could be use o Visual SLAM.
The p oblem o 3D layou eco e y in indoo scenes has been a co e esea ch opic o
o e a decade. Howe e , he e a e s ill se e al majo challenges ha emain unsol ed. Among
he mos ele an ones, a majo pa o he s a e-o - he-a me hods make implici o explici
assump ions on he scenes –e.g. box-shaped o Manha an layou s. Also, cu en me hods a e
compu a ionally expensi e and no sui able o eal- ime applica ions like obo na iga ion
and AR/VR. In his wo k we p esen CFL (Co ne s o Layou ), he i s end- o-end model
ha p edic s layou co ne s o 3D layou eco e y on
360◦
images. Ou expe imen al esul s
show ha we ou pe o m he s a e o he a , making less assump ions on he scene han
o he wo ks, and wi h lowe cos . We also show ha ou model gene alizes be e o came a
posi ion a ia ions han con en ional app oaches by using EquiCon s, a con olu ion applied
di ec ly on he sphe ical p ojec ion and hence in a ian o he equi ec angula dis o ions.
84 Co ne P edic ion o Layou Recons uc ion
CFL: End- o-End
Layou Reco e y
Fig. 5.1
Co ne s o Layou
: The model end- o-end p edic s he layou co ne s om he
sphe ical image. Connec ing he co ne s and assuming ceiling- loo pa allelism, we can
di ec ly ob ain he 3D layou in a e y sho ime.
5.1 In oduc ion
Reco e ing he 3D layou o an indoo scene om a single iew has a ac ed he a en ion o
compu e ision and g aphics esea che s in he las decade. The idea is going beyond pu e
geome ical econs uc ions and p o ide highe -le el con ex ual in o ma ion abou he scene,
e en in he p esence o clu e . Layou es ima ion is a key echnology in se e al eme ging
applica ion ma ke s, such as augmen ed and i ual eali y and obo na iga ion (Salas e al.,
2015). Bu also o mo e adi ional ones, like eal es a e (Liu e al., 2015a).
Layou es ima ion, howe e , is no a i ial ask and he e a e se e al majo p oblems ha
s ill emain unsol ed. Fo example, mos exis ing me hods a e based on s ong assump ions
on he geome y (e.g. Manha an scenes) o he o e -simpli ica ion o he oom ypes (e.g.
box-shaped layou s), o en unde i ing he ichness o eal indoo spaces. The limi ed ield
o iew o con en ional came as leads o ambigui ies, which could be sol ed by conside ing
a wide con ex . Fo his eason i is ad an ageous o use wide ields o iew, like 360
◦
pano amas. In hese cases, howe e , he me hods o con en ional came as a e no sui able
due o he image dis o ions and new ones ha e o be de eloped (Pais e al., 2019).
In he las yea s, he main imp o emen s in layou eco e y om pano amas ha e come
om he applica ion o deep lea ning. The high-le el ea u es lea ned by deep ne wo ks ha e
p o en o be as use ul o his p oblem as o many o he s. Ne e heless, hese echniques
en ail o he p oblems such as he lack o da a o o e i ing. S a e-o - he-a me hods equi e
addi ional p e- and/o pos -p ocessing. As a consequence hey a e e y slow, and his is a
majo d awback conside ing he a o emen ioned applica ions o eal- ime layou eco e y.
In his wo k, we p esen Co ne s o Layou (CFL), he i s end- o-end neu al ne wo k
ha p edic s a map o he co ne s o he oom o di ec ly ob ain he 3D layou om a single
360◦
image (Figu e 5.1). This makes
CFL mo e han 100 imes as e
han he s a e o
he a , while s ill
ou pe o ming he accu acy o cu en app oaches
. Fu he mo e, ou
5.2 Rela ed Wo k 85
p oposal is no limi ed by ypical scene assump ions, meaning ha i can p edic complex
geome ies, such as ooms wi h mo e han ou walls o non s ic Manha an s uc u es.
Addi ionally, we p opose a no el implemen a ion o he con olu ion o
360◦
images (Ta eno
e al., 2018, Cohen e al., 2018) in he equi ec angula p ojec ion. We de o m he ke nel,using
he ad ances p esen ed by Dai e al. (2017b), o compensa e he dis o ion and make CFL
mo e
obus o came a o a ion and pose a ia ions
, gene alizing o unseen con igu a-
ions. Hence, i is equi alen o applying di ec ly a con olu ion ope a ion o he sphe ical
image, which is geome ically mo e cohe en han applying a s anda d con olu ion on he
equi ec angula pano ama. We ha e ex ensi ely e alua ed ou ne wo k in wo public da ase s
wi h se e al aining con igu a ions, including da a augmen a ion echniques o add ess
occlusions by en o cing he ne wo k o lea n om he con ex . We also p opose a
obus ness
analysis
o see he e ec o ex insic a ia ions in pano amas and da ase bias. Ou
code
and
labeled da ase can be ound he e: CFL webpage.
5.2 Rela ed Wo k
The layou o a oom p o ides a s ong p io o o he isual asks like single- iew (Eigen
and Fe gus, 2015) and mul i- iew dep h eco e y (Concha e al., 2014), ealis ic inse ions o
i ual objec s in o indoo images (Ka sch e al., 2011), indoo objec ecogni ion (Bao e al.,
2011, Song and Xiao, 2016), indoo place ecogni ion (Hussain e al., 2016) o human pose
es ima ion (Fouhey e al., 2014). A la ge a ie y o me hods ha e been de eloped o his
pu pose using mul iple inpu images (Tsai e al., 2011, Flin e al., 2011) o dep h senso s
(Zhang e al., 2013), which deli e high-quali y econs uc ion esul s. Fo he common case
when a single RGB image is a ailable, he p oblem becomes conside ably mo e challenging
and esea che s need e y o en o ely on s ong assump ions.
The seminal app oaches o layou p edic ion om a single iew we e (Delage e al., 2006,
Lee e al., 2009), ollowed by (Hedau e al., 2009a, Schwing e al., 2013). They basically
model he layou o he oom wi h a anishing-poin -aligned 3D box, being hence cons ained
o his pa icula oom geome y and unable o gene alize o o he s appea ing equen ly in
eal applica ions. Mos ecen app oaches exploi CNNs and hei excellen pe o mance in a
wide ange o applica ions such as image classi ica ion, segmen a ion and de ec ion. (Mallya
and Lazebnik, 2015, Ren e al., 2016, Zhang e al., 2017, Zhao e al., 2017), o example,
ocus on p edic ing he in o ma i e edges sepa a ing he geome ic classes (walls, loo and
ceiling). Al e na i ely, Dasgup a e al. (2016) p oposed a FCN o p edic labels o each o
he su aces o he oom. All hese me hods equi e ex a compu a ion added o he o wa d
p opaga ion o he ne wo k o e ie e he ac ual layou . In Lee e al. (2017), o example, an
92 Co ne P edic ion o Layou Recons uc ion
he ne wo k ou pu (
k=4
) and 3 in e media e laye s (
k={1,...,3}
). The o al loss is hen
he sum o e all pixels, he 4 esolu ions and bo h he edge and co ne maps
L=∑
k={1,...,4}
∑
m={e,c}
∑
i
Lm
i[k].(5.2)
5.3.3 F om Co ne Maps o 3D Layou
Cu en me hods (Zou e al., 2018, Fe nandez-Lab ado e al., 2018b, Zhang e al., 2014) use
p e-compu ed anishing poin s and pos e io op imiza ions, being cons ained o p oduce
s ic Manha an 3D layou s. Aiming o a as end- o-end simple model, CFL a oids ex a
compu a ion and adop a ep esen a ion usually e e ed as So /Weak Manha an (Fu lan
e al, 2013) o A lan a Wo ld (Joo e al, 2018). Following his, ho izon al di ec ions a e no
necessa ily o hogonal o each o he , hus elaxing he model assump ions. To his end, we
simply ollow a na u al ans o ma ion om co ne s coo dina es o 2D and 3D layou . The
2D co ne s coo dina es a e he maximum ac i a ions in he p obabili y map. Assuming ha
he co ne se is consis en , hey a e di ec ly joined, om le o igh , in he uni sphe e
space and e-p ojec ed o he equi ec angula image plane. The 3D layou is in e ed by only
assuming ceiling- loo pa allelism, lea ing he wall s uc u e uncons ained –i.e., we assume
ha he loo co ne s a e on he same plane and he op co ne s a e di ec ly abo e he loo
ones, bu we do no o ce he usual Manha an pe pendicula i y be ween walls. Co ne s a e
p ojec ed o loo and ceiling planes gi en a uni a y came a heigh ( i ial as esul s a e up o
scale). See Figu e 5.3.
He e we p o ide a u he explana ion o how he p ocess o go om 2D o 3D wo ks.
F om he p edic ed 2D co ne posi ions, we can di ec ly eco e he 3D layou by doing he
ollowing assump ions:
1. So Manha an o A lan a wo ld.
This is a elaxa ion o he Manha an Wo ld
assump ion whe eby ho izon al di ec ions a e no necessa ily o hogonal o each o he .
Tha is, walls can in e sec wi h each o he in any di ec ion.
2. Ceiling- loo pa allelism.
Co ne s can be classi ied depending on hei posi ion along
he e ical di ec ion (abo e o below he ho izon line, which in cen al pano amas is
a he middle ow) be ween ceiling and loo co ne s espec i ely. Floo co ne s a e on
he same loo plane and ceiling co ne s a e di ec ly abo e he loo ones. The e ical
di ec ion is he no mal di ec ion o bo h loo and ceiling planes.
3. Uni a y came a heigh .
This is i ial as esul s a e up o scale bu needed o p edic
he o al heigh o he oom.

5.3 Co ne s o Layou 93
Taking all o his in o accoun , we can de ine a plane as he se o all poin s
P= (x,y,z)
such ha
P·N+d=0
, whe e he no mal
N= (nx,ny,nz)
is a no malized ec o pe pendicula
o i s su ace and
d
is he dis ance ha sepa a es i om he o igin o coo dina es in he
di ec ion o he no mal. Due o assump ions b) and c),
N
o bo h he loo and ceiling planes
is equal and co esponds o he e ical di ec ion, and he dis ance
d
om he loo o he
came a is known. The dis ance o he ceiling is ye unknown.
Addi ionally, hanks o he na u e o sphe ical images, we can easily ob ain he 3D ay
R( ) = O+
V·
(pa ame ic ep esen a ion) going om he cen e o he sphe e
O= (ox,oy,oz)
h ough he co ne posi ion, wi h no malized di ec ion ec o

V= ( x, y, z)
. To ob ain he
no malized di ec ion ec o

V
, we need he co ne posi ion in he sphe e, hus we ans o m
he image coo dina es o he co ne s
(u, )
in o sphe ical coo dina es and hen o he Euclidean
3D space. Equa ions o his can be ound in Sec ion 4.1. o his chap e . In he i s place,
Eq (5.3) gi e us he angles ha de ine he poin (u, )in he sphe e.
φ= (u−W
2)2π
W;θ=−( −H
2)π
H(5.3)
Whe e W and H a e he wid h and heigh o he equi ec angula image. Second, once hese
o a ions a e known we can compu e he di ec ion o he ay. The e o e, using Eq (5.4) we
can calcula e 
V.

V=


−cos(θ)sin(φ)
sin(θ)
cos(θ)cos(φ)


(5.4)
The in e sec ion be ween he co ne ay and he co esponding loo o ceiling plane will
gi e us he ac ual 3D co ne poin
P= (x,y,z)
(up o scale), ie. he in e sec ion ep esen s
ha poin
P
on he su ace o he plane ha e i ies he ay equa ion:
(ox+ x· )nx+(oy+
y· )ny+(oz+ z· )nz+d=0
. The poin
P
o in e sec ion would simply be he esul o
e alua ing he calcula ed , Eq (5.5), in he ay equa ion R( ).
=−oxnx+oyny+oznz+d
xnx+ yny+ znz
(5.5)
Le ’s conside we ha e pe o med he ope a ions o compu e one co ne poin on he loo
plane,
PF= (xF,yF,zF)
. The co esponding poin on he ceiling plane (
PC
) will be on op o
i (ie.
xF=xC
and
yF=yC
). The e o e, we can use his o compu e
C
, Eq (5.6), and hus he
ceiling poin :
C=(xF−ox)
C
x
(5.6)
94 Co ne P edic ion o Layou Recons uc ion
Fig. 5.4
Sphe ical pa ame iza ion o EquiCon s
. The sphe ical ke nel, de ined by i s
angula size (
αw×αh
) and esolu ion (
w× h
), is con ol ed a ound he sphe e wi h angles
φand θ.
whe e

VC= ( C
x, C
y, C
z)
is compu ed as in (5.4) wi h he co esponding ceiling poin in he
image. No ice ha wi h PCwe ha e he in o ma ion we we e missing o eco e he ceiling
plane.
Limi a ions o CFL:
We di ec ly join co ne s om le o igh , meaning ha ou model
would no wo k i any wall is occluded because o he con exi y o he scene. In hose
pa icula cases, he joining p ocess should ollow a di e en o de . Fe nandez-Lab ado
e al. (2018b) p oposes a geome y-based pos -p ocessing ha could alle ia e his p oblem,
bu i s cos is high and i needs he Manha an Wo ld assump ion. The addi ion o his
pos -p ocessing in o ou wo k, in any case, could be done simila ly o Fe nandez-Lab ado
e al. (2018a).
5.4 Equi ec angula Con olu ions
Sphe ical images a e ecei ing an inc easing a en ion due o he g owing numbe o omnidi-
ec ional senso s in d ones, obo s and au onomous ca s. A naï e applica ion o con olu ional
ne wo ks o a equi ec angula p ojec ion, is no , in p inciple, a good choice due o he space-
a ying dis o ions in oduced by such p ojec ion.
5.4 Equi ec angula Con olu ions 95
In his sec ion we p esen a con olu ion ha we name EquiCon , which is de ined in he
sphe ical domain ins ead o he image domain and i is implici ly in a ian o equi ec angula
ep esen a ion dis o ions. The ke nel in EquiCon s is de ined as a sphe ical su ace pa ch
–see Figu e 5.4. We pa ame ize i s ecep i e ield by he angles
αw
and
αh
. Thus, we di ec ly
de ine a con olu ion o e he ield o iew. The ke nel is o a ed and applied along he sphe e
and i s posi ion is de ined by he sphe ical coo dina es (
φ
and
θ
in he igu e) o i s cen e .
Unlike s anda d ke nels, ha a e pa ame e ized by hei size
kw×kh
, wi h EquiCon s we
de ine he angula size (
αw×αh
) and esolu ion (
w× h
). In p ac ice, we keep he aspec
a io,
αw
w=αh
h
, and we use squa e ke nels, so we will e e he ield o iew as
α
(
αw=αh
)
and he esolu ion as
(
w= h
) espec i ely om now on. In his wo k, we choose alues o
esolu ion and ield o iew o be he same as he image.
5.4.1 EquiCon s De ails
In Dai e al. (2017b), hey in oduce de o mable con olu ions by lea ning addi ional o se s
om he p eceding ea u e maps. O se s a e added o he egula ke nel loca ions in he
S anda d Con olu ion enabling ee o m de o ma ion o he ke nel.
Inspi ed by his wo k, we de o m he shape o he ke nels acco ding o he geome ical
p io s o he equi ec angula image p ojec ion. To do ha , we gene a e o se s ha a e no
lea ned bu ixed gi en he sphe ical dis o ion model and cons an o e he same ho izon al
loca ions. He e, we desc ibe how o ob ain he dis o ed pixel loca ions om he o iginal
ones.
Le us de ine
(u0,0, 0,0)
as he pixel loca ion on he equi ec angula image whe e we
apply he con olu ion ope a ion (i.e. he image coo dina e whe e he cen e o he ke nel is
loca ed). Fi s , we de ine he coo dina es o e e y elemen in he ke nel and a e wa ds we
o a e hem o he poin o he sphe e whe e he ke nel is being applied. We de ine each poin
o he ke nel as
ˆpi j =


ˆxi j
ˆyi j
ˆzi j


=


i
j
d


,(5.7)
whe e
i
and
j
a e in ege s in he ange
[− −1
2, −1
2]
and
d
is he dis ance om he cen e o
he sphe e o he ke nel g id. In o de o co e he ield o iew α,
d=
2 an(α
2).(5.8)
96 Co ne P edic ion o Layou Recons uc ion
S anda d De o mable Equi ec angula
Fig. 5.5
E ec o o se s on a 3×3ke nel
. Le : Regula ke nel in S anda d Con olu ion.
Cen e : De o mable ke nel in Dai e al. (2017b). Righ : Sphe ical su ace pa ch in EquiCon s.
We p ojec each poin in o he sphe e su ace by no malizing he ec o s, and o a e hem
o align he ke nel cen e o he poin whe e he ke nel is applied.
pi j =


xi j
yi j
zi j


=Ry(φ0,0)Rx(θ0,0)ˆpi j
|ˆpi j|,(5.9)
whe e
Ra(β)
s ands o a o a ion ma ix o an angle
β
a ound he
a
axis.
φ0,0
and
θ0,0
a e
he sphe ical angles o he cen e o he ke nel –see Figu e 5.4, and a e de ined as
φ0,0= (u0,0−W
2)2π
W;θ0,0=−( 0,0−H
2)π
H,(5.10)
whe e
W
and
H
a e, espec i ely, he wid h and heigh o he equi ec angula image in pixels.
Finally, he es o elemen s a e back-p ojec ed o he equi ec angula image domain. Fi s ,
we con e he uni sphe e coo dina es o la i ude and longi ude angles:
φi j =a c an(xi j
zi j
);θi j =a csin(yi j).(5.11)
And hen, o he o iginal 2D equi ec angula image domain:
ui j = (φi j
2π+1
2)W; i j = (−θi j
π+1
2)H.(5.12)
In Figu e 5.5 we show how hese o se s a e applied o a egula ke nel; and in Figu e 5.6
h ee ke nel samples on he sphe ical and on he equi ec angula images.
5.4 Equi ec angula Con olu ions 97
Fig. 5.6
EquiCon s on sphe ical images.
We show h ee ke nel posi ions o highligh
he di e ences be ween he o se s. As we app oach o he poles (la ge
θ
angles) he
de o ma ion o he ke nel on he equi ec angula image is bigge , in o de o ep oduce a
egula ke nel on he sphe e su ace. Addi ionally, wi h EquiCon s, we do no use padding
when he ke nel is on he bo de o he image since o se s ake he poin s o hei co ec
posi ion on he o he side o he 360◦image.

98 Co ne P edic ion o Layou Recons uc ion
5.5 Expe imen s
We p esen a se o expe imen s o e alua e CFL using bo h S anda d Con olu ions (S dCon s)
and he p oposed Equi ec angula Con olu ions (EquiCon s). We do no only analyze he
co ne maps p edic ed by ou model, bu also he impac o each algo i hmic componen
h ough abla ion s udies. We epo he pe o mance o ou p oposal in wo di e en da ase s,
and show quali a i e 2D and 3D models o di e en indoo scenes.
5.5.1 Da ase s
We use wo public da ase s ha comp ise se e al indoo scenes, SUN360 (Xiao e al.,
2012) and S an o d (2D-3D-S) A meni e al. (2017) in equi ec angula p ojec ion (360
◦
).
The o me is used o abla ion s udies, and bo h a e used o compa ison agains se e al
s a e-o - he-a baselines.
SUN360 (Xiao e al., 2012)
: We use
∼
500 bed oom and li ing oom pano amas om his
da ase labeled by Zhang e al. (2014). We use hese labels bu , since all pano amas we e
labeled as box- ype ooms, we hand-label and subs i u e 35 pano amas ep esen ing mo e
ai h ully he ac ual shapes o he ooms. We spli he aw da ase in 85
%
aining scenes and
15
%
es scenes andomly by making su e ha he e we e ooms o mo e han 4 walls in bo h
pa i ions.
S an o d 2D-3D-S (A meni e al., 2017)
: This da ase con ains mo e challenging scena ios
like clu e ed labo a o ies o co ido s. In Zou e al. (2018), hey use a eas 1, 2, 4, 6 o
aining, and a ea 5 o es ing. Fo ou expe imen s we use same pa i ions and he g ound
u h p o ided by hem.
5.5.2 Implemen a ion de ails
The inpu o he ne wo k is a single pano amic RGB image o esolu ion
256×128
. The
ou pu s a e, on he one hand, he oom layou edge map and on he o he hand, he co ne
map, bo h o hem a esolu ion
128×64
. A widely used s a egy o imp o e gene aliza ion
o neu al ne wo ks is da a augmen a ion. We apply andom e asing, ho izon al mi o ing as
well as ho izon al o a ion om
0◦
o
360◦
o inpu images du ing aining. The weigh s a e
all ini ialized using ResNe -50 (He e al., 2016) ained on ImageNe (Russako sky e al.,
2015). Fo CFL EquiCon s we use he same ke nel esolu ions and ield o iews as in
ResNe -50. This means ha o a s anda d 3
×
3 ke nel applied o a W
×
H ea u e map,
=3
and
α= o
W
, whe e
o =360◦
o pano amas. We minimize he c oss-en opy loss using
Adam (Kingma and Ba, 2014), egula ized by penalizing he loss wi h he sum o he L2
5.5 Expe imen s 99
Co ne s
Con . IP EM IoU Acc P R F1
: 1 : 1 : 1 : 1 : 1
S dCon s - - 0.519 0.978 0.611 0.763 0.675
S dCon s -✓0.531 0.979 0.639 0.749 0.685
S dCon s ✓ ✓ 0.569 0.982 0.684 0.761 0.718
EquiCon s - - 0.485 0.972 0.551 0.786 0.642
EquiCon s -✓0.536 0.980 0.649 0.744 0.690
EquiCon s ✓ ✓ 0.580 0.983 0.697 0.762 0.726
bigge is be e
Table 5.3
Abla ion s udy on SUN360 da ase .
We show esul s o bo h S anda d Con-
olu ions (S dCon s) and ou p oposed Equi ec angula Con olu ions (EquiCon s) wi h
some modi ica ions: Using o no in e media e p edic ions (IP) in he decode and edge map
p edic ions (EM).
o all weigh s. The ini ial lea ning a e is
2.5e−4
and is exponen ially decayed by a a e o
0.995 e e y epoch. We apply a d opou a e o 0.3.
The ne wo k is implemen ed using Tenso Flow (Abadi e al., 2016) and ained and es ed
in a NVIDIA Ti an X. The aining ime o S dCon s is a ound
1
hou and he es ime is
0.31
seconds pe image. Fo EquiCon s, aining akes
3
hou s and es a ound
3.32
seconds
pe image.
5.5.3 Ne wo k’s ou pu e alua ion
We measu e he quali y o ou p edic ed p obabili y co ne maps using i e s anda d me ics:
in e sec ion o e union IoU, p ecision P, ecall R, F1 Sco e
F1
and accu acy Acc. Table 5.3
summa izes ou esul s and allows us o answe he ollowing ques ions:
Wha a e he e ec s o di e en con olu ions?
As one would expec , EquiCon s, awa e
o he dis o ion model, lea n in a non-dis o ed gene ic ea u e space achie ing accu a e
p edic ions, like S dCon s on con en ional images (Lee e al., 2017). Dis o ion unde -
s anding, addi ionally, gi es he ne wo k o he ad an ages. While S dCon s lea n s ong
bias co ela ion be ween ea u es and dis o ion pa e ns (e.g. ceiling line on he op o he
image o clu e in he mid-bo om), EquiCon s a e in a ian o ha . Fo his eason, he
pe o mance o EquiCon s does no deg ade when a ying he came a DOF pose – see
Sec ion 5.5.4. Addi ionally, EquiCon s allow o di ec ly le e age ne wo ks p e- ained on
con en ional images. Speci ically, his ansla es in o a as e con e gence, which is desi able
as, o da e,
360◦
da ase s con ain a less images han da ase s wi h con en ional images. In
100 Co ne P edic ion o Layou Recons uc ion
Fig. 5.7
EquiCon s show mo e consis en quali a i e esul s
whe eas S dCon s simply
do no unde s and ha he image w aps a ound he sphe e, losing he con inuous con ex ha
hese images p o ide.
omnidi ec ional images, he igh and he le edge a e he same spo in eali y so, ano he
s eng h o EquiCon s lie in he ac ha we can a oid padding when he ke nel eaches he
bo de o he image since o se s ake he poin s o hei co ec posi ion on he o he side o
he
360◦
image. This allows he model o unde s and he con inui y o he scene. S dCon s,
ins ead, simply do no unde s and ha he image w aps a ound he sphe e. As a consequence,
in mos cases when co ne s app oach he bo de s, S dCon s p edic hese co ne s wice, i.e.
a bo h ends, o he edges a one side would no coincide wi h he edges a he o he side.
This e ec is highligh ed in Figu e 5.7 and u he demons a ed in he supplemen a y ideo.
How can we e ine p edic ions?
The e a e some echniques ha we can use in o de o
ob ain mo e accu a e and e ined p edic ions. He e, we make py amid p elimina y p edic ions
in he decode and i e a i ely e ine hem, by eeding hem back o he ne wo k, un il he
inal p edic ion. Also, al hough we only use he co ne map o eco e he layou o he oom,
we ain he ne wo k o addi ionally p edic edge maps as an auxilia y ask. This is ano he
ep esen a ion o he same ask ha ensu es ha he ne wo k lea ns o exploi he ela ionship
be ween bo h ou pu s, i.e., he ne wo k lea ns how edges in e sec be ween hem gene a ing
he co ne s. The imp o emen is shown in he Table 5.3.
How can we deal wi h occlusions?
We do Random E asing Da a Augmen a ion. This
ope a ion andomly selec s ec angles in he aining images and emo es i s con en , gene -
a ing a ious le els o i ual occlusion. In his manne we simula e eal si ua ions whe e
objec s in he scene occlude he co ne s o he oom layou , and o ce he ne wo k o lea n
con ex -awa e ea u es o o e come his challenging si ua ion. Figu e 5.8 illus a es his
s a egy wi h an example.
Is i possible o elax he scene assump ions while keeping a good pe o mance?
By
a oiding cons ained Manha an 3D layou p edic ions we no only achie e be e esul s
5.5 Expe imen s 101
Inpu Pano ama Wi hou
andom e asing
Wi h
andom e asing
E asing example
Fig. 5.8
Augmen ing he da a wi h i ual occlusions.
Le : Image wi h e ased pixels.
Righ : Inpu pano ama and p edic ions wi hou and wi h pixel e asing. No ice he imp o e-
men by andom e asing.
compa ed wi h cu en a s, bu also we sa e in compu a ion. Addi ionally, ou model
o e comes he classic box- oom simpli ica ion ( ou -walls oom se ups), e en i we s ill ha e
a la gely unbalanced da ase a e labeling some pano amas mo e accu a ely o hei ac ual
shape. We add ess his p oblem by choosing a ba ch size o
16
and o cing i o always
include one non-box sample. This a o s he lea ning o mo e complex ooms despi e ha ing
ew examples.
F1Acc IoU
T ans S dCon s 55.32±8.23 95.46±1.3 39.135±7.82
EquiCon s 59.55 ±8.95 96.21 ±1.14 43.47 ±8.83
Ro x S dCon s 45.89±14.72 93.44±3.18 31.26±12.83
EquiCon s 46.2 ±15.1 94.43 ±2.18 31.625 ±13.41
Ro y S dCon s 72.28±2.7 98.21±0.21 57.54±3.25
EquiCon s 72.96 ±2.02 98.29 ±0.14 58.44 ±2.44
Table 5.4
Robus ness analysis
. Values ep esen he mean alue (bigge is be e )
±
s anda d
de ia ion (smalle is be e ) in %. We apply h ee ypes o ans o ma ions o he pano amas:
ansla ions
in
y
dependan on he oom heigh om
−0.3h
o
0.3h
,
o a ions
in
x
om
−30◦
o
+30◦
and
o a ions
in
y
om
0◦
o
360◦
. We do no use hese images o aining
bu jus o es ing in o de o show he gene aliza ion capabili ies o bo h models.
5.5.4 Robus ness analysis
We es ou model wi h p e iously unseen images whe e he came a iewpoin is di e en
om ha in he aining se . The dis o ion in equi ec angula p ojec ion is loca ion dependen ,
108 Co ne P edic ion o Layou Recons uc ion
Fig. 5.12 Layou p edic ions (ligh magen a) and g ound u h (da k magen a) o
complex
oom geome ies
on he SUN360 anno a ion da ase (Xiao e al., 2012). Bes iewed in
colo .

5.6 Conclusions 109
Fig. 5.13 Layou p edic ions (ligh magen a) and g ound u h (da k magen a) on he S an o d
2D-3D anno a ion da ase (A meni e al., 2017). Bes iewed in colo .
Chap e 6
Monocula and RGB-D SLAM on
Dynamic En i onmen s
The assump ion o scene igidi y is ypical in SLAM algo i hms. Such a s ong assump ion
limi s he use o mos isual SLAM sys ems in popula ed eal-wo ld en i onmen s, which
a e he a ge o se e al ele an applica ions like se ice obo ics o au onomous ehicles.
In his chap e we p esen DynaSLAM, a isual SLAM sys em ha , building on ORB-
SLAM2 (Mu -A al and Ta dós, 2017), adds he capabili ies o dynamic objec de ec ion and
backg ound inpain ing. DynaSLAM is obus in dynamic scena ios o monocula , s e eo and
RGB-D
con igu a ions. We a e capable o de ec ing he mo ing objec s ei he by mul i- iew
geome y, deep lea ning o bo h. Ha ing a s a ic map o he scene allows inpain ing he ame
backg ound ha has been occluded by such dynamic objec s.
We e alua e ou sys em in public monocula , s e eo and
RGB-D
da ase s. We s udy he
impac o se e al accu acy/speed ade-o s o assess he limi s o he p oposed me hodology.
DynaSLAM ou pe o ms he accu acy o s anda d isual SLAM baselines in highly dynamic
scena ios. And i also es ima es a map o he s a ic pa s o he scene, which is a mus o
long- e m applica ions in eal-wo ld en i onmen s.
6.1 In oduc ion
SLAM is a p e equisi e o many obo ic applica ions, o example collision-less na iga ion.
SLAM echniques es ima e join ly a map o an unknown en i onmen and he obo pose
wi hin such map, only om he da a s eams o i s on-boa d senso s. The map allows he
obo o con inually localize wi hin he same en i onmen wi hou accumula ing d i . This
112 Monocula and RGB-D SLAM on Dynamic En i onmen s
is in con as o odome y app oaches ha in eg a e he inc emen al mo ion es ima ed wi hin
a local window and a e unable o co ec he d i when e isi ing places.
Visual SLAM, whe e he main senso is a came a, has ecei ed a high deg ee o a en ion
and esea ch e o s o e he las yea s. The minimalis ic solu ion o a monocula came a
has p ac ical ad an ages wi h espec o size, powe and cos , bu also se e al challenges
such as he unobse abili y o he scale o s a e ini ializa ion. By using mo e complex se ups,
like s e eo o
RGB-D
came as, hese issues a e sol ed and he obus ness o isual SLAM
sys ems can be g ea ly imp o ed.
The esea ch communi y has add essed SLAM om many di e en angles. Howe e , he
as majo i y o he app oaches and da ase s assume a s a ic en i onmen . As a consequence,
hey can only manage small ac ions o dynamic con en by classi ying hem as ou lie s o
such s a ic model. Al hough he s a ic assump ion holds o some obo ic applica ions, i
limi s he applicabili y o isual SLAM in many ele an cases, such as in elligen au onomous
sys ems ope a ing in popula ed eal-wo ld en i onmen s o e long pe iods o ime.
Visual SLAM can be classi ied in o ea u e-based me hods (Klein and Mu ay, 2007,
Mu -A al e al., 2015), ha ely on salien poin s ma ching and can only es ima e a spa se
econs uc ion; and di ec me hods (S ühme e al., 2010, Newcombe e al., 2011, G abe
e al., 2011), which a e able o es ima e in p inciple a comple ely dense econs uc ion by he
di ec minimiza ion o he pho ome ic e o and TV egula iza ion. Some di ec me hods
ocus on he high-g adien a eas es ima ing semi-dense maps (Engel e al., 2014, 2017).
None o he abo e me hods, conside ed he s a e o he a , add ess he e y common
p oblem o dynamic objec s in he scene, e.g., people walking, bicycles o ca s. De ec ing and
dealing wi h dynamic objec s in isual SLAM e eals se e al challenges o bo h mapping
and acking, including:
1. How o de ec such dynamic objec s in he images o:
(a)
P e en he acking algo i hm om using ma ches ha belong o dynamic objec s.
(b)
P e en he mapping algo i hm om including mo ing objec s as pa o he 3D
map.
2.
How o comple e he pa o he 3D map ha is empo ally occluded by a mo ing
objec .
Many applica ions would g ea ly bene i om p og ess along hese lines. Among o he s,
augmen ed eali y, au onomous ehicles, and medical imaging. All o hem could o
ins ance sa ely euse maps om p e ious uns. De ec ing and dealing wi h dynamic objec s
is a equisi e o es ima e s able maps, use ul o long- e m applica ions. I he dynamic
6.1 In oduc ion 113
(a) Inpu RGB-D ames wi h dynamic con en .
(b) Ou pu
RGB-D
ames. Dynamic con en has been emo ed. Occluded backg ound has been econs uc ed
wi h in o ma ion om p e ious iews.
(c) Map o he s a ic pa o he scene, a e emo al o he dynamic objec s.
Fig. 6.1 O e iew o DynaSLAM esul s o he RGB-D case.
con en is no de ec ed, i becomes pa o he 3D map, complica ing i s usabili y o acking
o eloca ion pu poses.
In his wo k we p opose an on-line algo i hm o deal wi h dynamic objec s in
RGB-D
,
s e eo and monocula SLAM. This is done by adding a on -end s age o he s a e-o - he-a
ORB-SLAM2 sys em (Mu -A al and Ta dós, 2017), wi h he pu pose o ha ing a mo e
accu a e acking and a eusable map o he scene. In he monocula and s e eo cases ou
p oposal is o use a CNN o pixel-wise segmen he a p io i dynamic objec s in he ames
(e.g., people and ca s), so ha he SLAM algo i hm does no ex ac ea u es on hem. In he

114 Monocula and RGB-D SLAM on Dynamic En i onmen s
RGB-D case we p opose o combine mul i- iew geome y models and deep-lea ning-based
algo i hms o de ec ing dynamic objec s and, a e ha ing emo ed hem om he images,
inpain he occluded backg ound wi h he co ec in o ma ion o he scene (Fig. 6.1).
The es o he chap e is s uc u ed as ollows: sec ion 6.2 discusses ela ed wo k, sec ion
6.3 gi es he de ails o ou p oposal, sec ion 6.4 de ails he expe imen al esul s, and sec ion
6.5 p esen s he conclusions and lines o u u e wo k.
6.2 Rela ed Wo k
Dynamic objec s a e, in mos SLAM sys ems, classi ied as spu ious da a and he e o e
nei he included in he map no used o came a acking. The mos ypical ou lie ejec ion
algo i hms a e RANSAC (e.g., in ORB-SLAM (Mu -A al e al., 2015, Mu -A al and Ta dós,
2017)) and obus cos unc ions (e.g., in PTAM by Klein and Mu ay (2007)).
The e a e se e al SLAM sys ems ha add ess mo e speci ically he dynamic scene
con en . Wi hin ea u e-based SLAM me hods, some o he mos ele an on dealing wi h
dynamic scenes a e he ollowing. Tan e al. (2013) ha de ec changes ha ake place in
he scene by p ojec ing he map ea u es in o he cu en ame o appea ance and s uc u e
alida ion. Wangsi ipi ak and Mu ay (2009) ack known 3D dynamic objec s in he scene.
Simila ly, Riazuelo e al. (2017) deal wi h human ac i i y by de ec ing and acking people.
Mo e ecen ly, he wo k o Li and Lee (2017) uses dep h edges poin s, which ha e an
associa ed weigh indica ing i s p obabili y o belonging o a dynamic objec .
Di ec me hods a e, in gene al, mo e sensi i e o dynamic objec s in he scene. The mos
ele an wo ks speci ically designed o dynamic scenes a e men ioned bellow. Alcan a illa
e al. (2012) de ec mo ing objec s by means o a scene low ep esen a ion wi h s e eo cam-
e as. Wang and Huang (2014) segmen he dynamic objec s in he scene using RGB op ical
low. Kim and Kim (2016) p opose o ob ain he s a ic pa s o he scene by compu ing he
di e ence be ween consecu i e dep h images p ojec ed o e he same plane. Sun e al. (2017)
calcula e he di e ence in in ensi y be ween consecu i e RGB images. Pixel classi ica ion is
done wi h he segmen a ion o he quan ized dep h image.
All he me hods –bo h ea u e-based and di ec ones– ha map he s a ic scene pa s only
om he in o ma ion con ained in he sequence (Mu -A al and Ta dós, 2017, Mu -A al
e al., 2015, Tan e al., 2013, Li and Lee, 2017, Alcan a illa e al., 2012, Wang and Huang,
2014, Kim and Kim, 2016, Sun e al., 2017, Concha and Ci e a, 2015a), ail o es ima e
li elong models when an a p io i dynamic objec emains s a ic, e.g., pa ked ca s o people
si ing. On he o he hand, Wangsi ipi ak and Mu ay (2009), and Riazuelo e al. (2017)
would de ec hose a p io i dynamic objec s, bu would ail o de ec changes p oduced by
6.3 DynaSLAM Sys em Desc ip ion 115
Fig. 6.2 Block diag am o ou p oposal. In he s e eo and monocula pipeline (black con-
inuous line) he images pass h ough a Con olu ional Neu al Ne wo k (
Mask R-CNN
) o
compu ing he pixel-wise seman ic segmen a ion o he a p io i dynamic objec s be o e being
used o he mapping and acking. In he
RGB-D
case (black dashed line) a second app oach
based on mul i- iew geome y is added o a mo e accu a e mo ion segmen a ion, o which
we need a low-cos acking algo i hm. Once he posi ion o he came a is known (T acking
and Mapping ou pu ), we can inpain he backg ound occluded by dynamic objec s. The ed
do ed line ep esen s he da a low o he s o ed spa se map.
s a ic objec s, e.g., a chai a pe son is pushing, o a ball ha someone has h own. Tha is,
he o me app oach succeeds in de ec ing mo ing objec s, and he second one in de ec ing
se e al mo able objec s. Ou p oposal, DynaSLAM, combines mul i- iew geome y and
deep lea ning in o de o add ess bo h si ua ions. Simila ly, Amb us e al. (2016) segmen
dynamic objec s by combining a dynamic classi ie and mul i- iew geome y.
6.3 DynaSLAM Sys em Desc ip ion
Fig. 6.2 shows an o e iew o ou sys em. Fi s o all, he RGB channels pass h ough a
CNN ha segmen s ou pixel-wise all he a p io i dynamic con en , e.g., people o ehicles.
In he
RGB-D
case, we use mul i- iew geome y o imp o e he dynamic con en seg-
men a ion in wo ways. Fi s , we e ine he segmen a ion o he dynamic objec s p e iously
ob ained by he CNN. Second, we label as dynamic new objec ins ances ha a e s a ic mos
o he ime (i.e., de ec mo ing objec s ha we e no se o mo able in he CNN s age).
Fo ha pu pose, i is necessa y o know he came a pose, o which a low-cos acking
module has been implemen ed o localize he came a wi hin he al eady c ea ed scene map.
These segmen ed ames a e he ones which a e used o ob ain he came a ajec o y and
he map o he scene. No ice ha i he mo ing objec s in he scene a e no wi hin he
CNN classes, he mul i- iew geome y s age would s ill de ec he dynamic con en , bu he
accu acy migh dec ease.
Once his ull dynamic objec de ec ion and localiza ion o he came a ha e been done,
we aim o econs uc he occluded backg ound o he cu en ame wi h s a ic in o ma ion
116 Monocula and RGB-D SLAM on Dynamic En i onmen s
om p e ious iews. These syn he ic ames a e ele an o applica ions like augmen ed
and i ual eali y, and place ecogni ion in li elong mapping.
In he monocula and s e eo cases, he images a e segmen ed by he CNN so ha keypoin s
belonging o he a p io i dynamic objec s a e nei he acked no mapped.
All he di e en s ages a e desc ibed in dep h in he nex subsec ions (6.3.1 o 6.3.5).
6.3.1 Segmen a ion o Po en ially Dynamic Con en using a CNN
Fo de ec ing dynamic objec s we p opose o use a CNN ha ob ains a pixel-wise seman ic
segmen a ion o he images. In ou expe imen s we use
Mask R-CNN
(He e al., 2017),
which is he s a e o he a o objec ins ance segmen a ion.
Mask R-CNN
can ob ain bo h
pixel-wise seman ic segmen a ion and he ins ance labels. Fo his wo k we use he pixel-wise
seman ic segmen a ion in o ma ion, bu he ins ance labels could be use ul in u u e wo k
o he acking o he di e en mo ing objec s. We use he Tenso Flow implemen a ion by
Ma e po 1.
The inpu o
Mask R-CNN
is he RGB o iginal image. The idea is o segmen hose
classes ha a e po en ially dynamic o mo able (pe son, bicycle, ca , mo o cycle, ai plane,
bus, ain, uck, boa , bi d, ca , dog, ho se, sheep, cow, elephan , bea , zeb a and gi a e).
We conside ha , o mos en i onmen s, he dynamic objec s likely o appea a e included
wi hin his lis . I o he classes we e needed, he ne wo k, ained on
MS COCO
(Lin e al.,
2014), could be ine- uned wi h new aining da a.
The ou pu o he ne wo k, assuming ha he inpu is an RGB image o size
m×n×3
, is
a ma ix o size
m×n×l
, whe e
l
is he numbe o objec s in he image. Fo each ou pu
channel
i∈l
a bina y mask is ob ained. By combining all he channels in o one, we can
ob ain he segmen a ion o all dynamic objec s appea ing in one image o he scene.
6.3.2 Low-Cos T acking
A e he po en ially dynamic con en has been segmen ed, he pose o he came a is acked
using he s a ic pa o he image. Because he segmen con ou s usually become high-
g adien a eas, salien poin ea u es end o appea . We do no conside he ea u es in such
con ou a eas.
The acking implemen ed a his s age o he algo i hm is a simple and he e o e
compu a ionally ligh e e sion o he one in ORB-SLAM2 (Mu -A al and Ta dós, 2017). I
p ojec s he map ea u es in he image ame, sea ches o he co espondences in he s a ic
a eas o he image, and minimizes he ep ojec ion e o o op imize he came a pose.
1h ps://gi hub.com/ma e po /Mask_RCNN
6.3 DynaSLAM Sys em Desc ip ion 117
(a) Keypoin
x′
belongs o a s a ic objec (
z′=
zp o j).
(b) Keypoin
x′
belongs o a dynamic objec
(z′≪zp o j).
Fig. 6.3 Keypoin
x
om he Key F ame (KF) is p ojec ed in o he Cu en F ame (CF) using
i s dep h and came a pose, esul ing in poin
x′
wi h dep h
z′
. The p ojec ed dep h
zp o j
is
hen compu ed. A pixel is labeled as dynamic i he di e ence
∆z=zp o j −z′
is g ea e han
a h eshold τz.
6.3.3 Segmen a ion o Dynamic Con en using Mask R-CNN and Mul i-
iew Geome y
By using
Mask R-CNN
, mos o he dynamic objec s can be segmen ed and no used o
acking and mapping. Howe e , he e a e objec s ha canno be de ec ed by his app oach
because hey a e no a p io i dynamic, bu mo able. Examples o he la es a e a book ca ied
by someone, a chai ha someone is mo ing, o e en u ni u e changes in long- e m mapping.
The app oach u ilized o dealing wi h hese cases is de ailed in his sec ion.
Fo each inpu ame, we selec he p e ious key ames ha ha e he highes o e laps.
This is done by aking in o accoun bo h he dis ance and he o a ion be ween he new
ame and each o he key ames, simila ly o Tan e al. (2013). The numbe o o e lapping
key ames has been se o
5
in ou expe imen s, as a comp omise be ween compu a ional
cos and accu acy in he de ec ion o dynamic objec s.
We hen compu e he p ojec ion o each keypoin
x
om he p e ious key ames in o he
cu en ame, ob aining he keypoin s
x′
, as well as hei p ojec ed dep h
zp o j
, compu ed
om he came a mo ion. No ice ha he keypoin s
x
come om he ea u es ex ac o
algo i hm used in
ORB-SLAM2
. Fo each keypoin , whose co esponding 3D poin is
X
, we
124 Monocula and RGB-D SLAM on Dynamic En i onmen s
Sequence Dep h Edge Mo ion Segmen a ion DSLAM Mo ion Remo al DVO-SLAM DynaSLAM (N+G) (RGB-D)
SLAM
w/o Mo ion
De ec ion
w/ Mo ion
De ec ion
Imp o .
w/ MD
w/o Mo ion
De ec ion
w/ Mo ion
De ec ion
Imp o .
w/ MD
w/o Mo ion
De ec ion
w/ Mo ion
De ec ion
Imp o .
w/ MD
[m] [m] [m] [%] [m] [m] [%] [m] [m] [%]
w_hal 0.049 0.116 0.055 52.59% 0.529 0.125 76.32% 0.351 0.025 92.88%
w_xyz 0.060 0.202 0.040 80.20% 0.597 0.093 84.38% 0.459 0.015 96.73%
w_ py 0.179 0.515 0.076 85.24% 0.730 0.133 81.75% 0.662 0.035 94.71%
w_s a 0.026 0.470 0.024 94.89% 0.212 0.066 69.06% 0.090 0.006 93.33%
s_hal 0.043 - - - 0.062 0.047 23.70% 0.020 0.017 15.00%
s_xyz 0.040 - - - 0.051 0.048 4.55% 0.009 0.015 X
Dep h Edge SLAM by Li and Lee (2017).
Mo ion Segmen a ion DSLAM by Wang and Huang (2014)
Mo ion Remo al DVO-SLAM by Sun e al. (2017)
DynaSLAM w/o Mo ion De ec ion is ORB-SLAM2 by Mu -A al and Ta dós (2017)
Table 6.3 Absolu e ajec o y RMSE [m] o DynaSLAM agains s a e-o - he-a
RGB-D
SLAM sys ems in dynamic scenes. To e alua e he e ec i eness o he speci ic module
add essing dynamic con en , we epo he imp o emen wi h espec o he o iginal SLAM
sys ems (w/o Mo ion De ec ion). Ou esul s a e es ima ed using
Mask R-CNN
and mul i-
iew geome y.
used in e e y case. DynaSLAM signi ican ly ou pe o ms all o hem in all sequences (bo h
high and low dynamic ones). The e o is, in gene al, a ound 1-2 cm, simila o ha o he
s a e o he a in s a ic scenes. Ou mo ion de ec ion app oach also ou pe o ms he o he
me hods.
ORB-SLAM, he monocula e sion o ORB-SLAM2, is gene ally mo e accu a e han
he
RGB-D
one in dynamic scenes, due o hei di e en ini ializa ion algo i hms.
RGB-D
ORB-SLAM2 is ini ialized and s a s he acking om he e y i s ame, and hence
dynamic objec s can in oduce e o s. ORB-SLAM delays he ini ializa ion un il he e is
pa allax and consensus using he s a ici y assump ion. Hence, i does no ack he came a
o he ull sequence, some imes missing a subs an ial pa o i , o e en no ini ializing.
Table 6.4 shows he acking esul s and pe cen age o he acked ajec o y o ORB-
SLAM and DynaSLAM (monocula ) in he TUM da ase . The ini ializa ion in DynaSLAM is
always quicke han ha o ORB-SLAM. In ac , in highly dynamic sequences, ORB-SLAM
ini ializa ion only occu s when he mo ing objec s disappea om he scene. In conclusion,
al hough he accu acy o DynaSLAM is sligh ly lowe , i succeeds in boo s apping he
sys em wi h dynamic con en and p oducing a map wi hou such con en (see
Fig. 6.1
), o be
e-used o long- e m applica ions. The eason why DynaSLAM is sligh ly less accu a e is
ha he es ima ed ajec o y is longe , and he e is he e o e oom o accumula ing e o s.

6.4 Expe imen al Resul s 125
Sequence ORB-SLAM DynaSLAM
Mu -A al and Ta dós (2017) (Monocula )
ATE [m] % T aj ATE [m] % T aj
3/walking_hal sphe e 0.017 87.16 0.021 97.84
3/walking_xyz 0.012 57.63 0.014 87.37
2/desk_wi h_pe son 0.006 95.30 0.008 97.07
3/si ing_xyz 0.007 91.44 0.013 100.00
Table 6.4 Absolu e ajec o y RMSE [m] and pe cen age o success ully acked ajec o y
o bo h ORB-SLAM and DynaSLAM (monocula ).
Sequence ORB-SLAM2 (S e eo) DynaSLAM (S e eo)
Mu -A al and Ta dós (2017)
RPE RRE ATE RPE RRE ATE
[%] [◦/100m] [m] [%] [◦/100m] [m]
KITTI 00 0.70 0.25 1.3 0.74 0.26 1.4
KITTI 01 1.39 0.21 10.4 1.57 0.22 9.4
KITTI 02 0.76 0.23 5.7 0.80 0.24 6.7
KITTI 03 0.71 0.18 0.6 0.69 0.18 0.6
KITTI 04 0.48 0.13 0.2 0.45 0.09 0.2
KITTI 05 0.40 0.16 0.8 0.40 0.16 0.8
KITTI 06 0.51 0.15 0.8 0.50 0.17 0.8
KITTI 07 0.50 0.28 0.5 0.52 0.29 0.5
KITTI 08 1.05 0.32 3.6 1.05 0.32 3.5
KITTI 09 0.87 0.27 3.2 0.93 0.29 1.6
KITTI 10 0.60 0.27 1.0 0.67 0.32 1.2
Table 6.5 Compa ison o he RMSE o he ATE [m], he a e age o he RPE [
%
] and he
RRE [◦/100m] o DynaSLAM agains ORB-SLAM2 sys em o s e eo came as.
6.4.2 KITTI Da ase
The KITTI Da ase (Geige e al., 2013) con ains s e eo sequences eco ded om a ca in
u ban and highway en i onmen s. Table 6.5 shows ou esul s in he ele en aining sequences,
compa ed agains s e eo ORB-SLAM2. We use wo di e en me ics, he absolu e ajec o y
RMSE p oposed in S u m e al. (2012b), and he a e age ela i e ansla ion and o a ion
e o s, p oposed in Geige e al. (2013). Table 6.6 shows he esul s in he same sequences
o he monocula a ian s o ORB-SLAM and DynaSLAM.
126 Monocula and RGB-D SLAM on Dynamic En i onmen s
Sequence ORB-SLAM DynaSLAM (Monocula )
Mu -A al and Ta dós (2017)
KITTI 00 5.33 7.55
KITTI 02 21.28 26.29
KITTI 03 1.51 1.81
KITTI 04 1.62 0.97
KITTI 05 4.85 4.60
KITTI 06 12.34 14.74
KITTI 07 2.26 2.36
KITTI 08 46.68 40.28
KITTI 09 6.62 3.32
KITTI 10 8.80 6.78
Table 6.6 Absolu e ajec o y RMSE [m] o ORB-SLAM and DynaSLAM (monocula ).
No e ha he esul s a e simila in bo h he monocula and s e eo cases, bu he o me
is mo e sensi i e o dynamic objec s and he e o e o he addi ions in DynaSLAM. In some
sequences he accu acy o he acking is imp o ed when no using ea u es belonging o a
p io i dynamic objec s, i.e., ca s, bicycles, e c. An example o his would be he sequences
KITTI 01 and KITTI 04, in which all ehicles ha appea a e mo ing. In he sequences in
which mos o he eco ded ca s and ehicles a e pa ked (hence s a ic), he absolu e ajec o y
RMSE is usually bigge since he keypoin s used o acking a e mo e dis an and usually
belong o low- ex u e a eas (KITTI 00, KITTI 02, KITTI 06). Howe e , he loop closu e and
elocaliza ion algo i hms wo k mo e obus ly since he esul ing map only con ains s uc u al
objec s, i.e., he map can be e-used and wo k in long- e m applica ions.
As u u e wo k, i is in e es ing o make a dis inc ion be ween hose mo able and mo ing
objec s, by using only RGB in o ma ion. I a ca is de ec ed by he CNN (mo able) bu is
no cu en ly mo ing, i s co esponding keypoin s should be used o he local acking, bu
should no be in he map.
6.4.3 Timing Analysis
To comple e he e alua ion o ou p oposal, Table 6.7 shows he a e age compu a ional
ime o i s di e en s ages. No e ha DynaSLAM is no op imized o eal- ime ope a ion.
Howe e , i s capabili y o c ea ing li e-long maps o he s a ic scene con en a e also ele an
o unning on o line mode.
6.5 Conclusions 127
Sequence Low-Cos T acking [ms] Mul i- iew Geome y [ms] Backg ound Inpain ing [ms]
w_hal sphe e 1.69 333.68 208.09
w_ py 1.59 235.98 183.56
Table 6.7 DynaSLAM a e age compu a ional ime [ms].
Mu e al. show eal- ime esul s o and ORB-SLAM2 (Mu -A al and Ta dós, 2017).
He e al. (2017) epo ha
Mask R-CNN
uns a
195
ms pe image on a N idia Tesla M40
GPU.
The addi ion o he mul i- iew geome y s age is an addi ional slowdown, due mainly o
he egion g ow h algo i hm. The backg ound inpain ing also in oduces a delay, which is
ano he eason why i should be done a e he acking and mapping s age, as i has been
shown in Fig. 6.2.
6.5 Conclusions
We ha e p esen ed a isual SLAM sys em ha , building on ORB-SLAM, adds a mo ion
segmen a ion app oach ha makes i obus in dynamic en i onmen s o monocula , s e eo
and
RGB-D
came as, o e ing a solu ion o a e y well known Visual SLAM p oblem. Ou
sys em accu a ely acks he came a and c ea es a s a ic and he e o e eusable map o he
scene. In he
RGB-D
case, DynaSLAM is capable o ob aining he syn he ic RGB ames
wi h no dynamic con en and wi h he occluded backg ound inpain ed, as well as hei
co esponding syn hesized dep h ames, which migh be oge he e y use ul o i ual
eali y applica ions. We include a ideo showing he po en ial o DynaSLAM 2.
The compa ison agains he s a e o he a shows ha DynaSLAM achie es in mos cases
he highes accu acy.
In he ideos o he TUM da ase ha include Dynamic Objec s da ase , a he momen
o publishing his wo k DynaSLAM was he bes
RGB-D
SLAM solu ion. Cu en ly i is
s ill he mos accu a e, al hough wo ks like he one p esen ed in Dai e al. (2020) o e a
model ha does no equi e GPU sac i icing on pe o mance, he au ho s also claim ha
hei p oposal could be combined o some o he ideas p esen ed in his chap e . Simila
conclusions a e ound in Vincen e al. (2020). Being he compu a ion ime one o he bigges
d awback o DynaSLAM acco ding o hese ecen publica ion we hink is in e es ing o sha e
he wo k by Alonso e al. (2020) in oducing MiniNe , a eal- ime seman ic segmen a ion
CNN; ou p oposal in his chap e is no dependen on he CNN we a e cu en ly using bu
i can use a di e en less ime-consuming ne wo k as MiniNe . In he monocula case, ou
2h ps://you u.be/EabI_goFmQs
128 Monocula and RGB-D SLAM on Dynamic En i onmen s
accu acy is simila o ha o ORB-SLAM, ob aining howe e a s a ic map o he scene wi h
an ea lie ini ializa ion.
In he KITTI da ase DynaSLAM is sligh ly less accu a e han monocula and s e eo
ORB-SLAM, excep o hose cases in which dynamic objec s ep esen an impo an pa o
he scene. Howe e , ou es ima ed map only con ains s uc u al objec s and can he e o e be
e-used in long- e m applica ions.
Fu u e wo ks in his line o esea ch ha e looked, eal- ime pe o mance Dai e al. (2020),
Vincen e al. (2020), Alonso e al. (2020), an RGB-based mo ion de ec o , o a mo e ealis ic
appea ance o he syn hesized RGB ames by using a mo e elabo a e inpain ing echnique,
e.g., he one used by Pa hak e al. (2016) by he use o Gene a i e Ad e sa ial Ne wo ks
(GANs). This las idea was ca ied ou , pos e io o his wo k, by Bescos e al. (2019) whe e
hey use GANs o inpain al eady de ec ed dynamic objec s. A he momen i has only been
es ed in simula ion wi h g ound- u h segmen a ion.
Chap e 7
Conclusions
In gene al, 3D isual pe cep ion is a om being ully sol ed. The e a e signi ican esea ch
challenges ahead, in pa icula ela ed o scene unde s anding. In his hesis we ha e ad anced
he s a e o he a in se e al a eas o his exci ing and ele an opic.
The i s con ibu ion desc ibed in his hesis is on single- iew dep h es ima ion. In CAM-
Con s (Facil e al., 2019) we ha e p oposed a new ype o con olu ion and demons a ed he
ad an ages o accoun ing o he came a in insic pa ame e s in dep h es ima ion asks. We
ha e shown ha ou CAM-Con s allow us o ain and es wi h di e en came as; some hing
ha has been explo ed u he in López-An eque a e al. (2020). Fu u e wo k along his
line should explo e models ha le e age he came a in insics and, unlike CAM-Con s, do
no equi e o lea n how o use hem. A majo d awback o CAM-Con s is ha , as any
lea ning p ocedu e, hey a e s ongly da a-dependen . The e o e, a su icien ly well sampled
da ase o images aken by di e en came as is needed o a easonable pe o mance, and 1
o 2 came as migh no be enough. On his line, López-An eque a e al. (2020) has s a ed
making p og ess on hei p oposal o a canonical came a model, simila o he ocal leng h
no maliza ion we use in ou wo k.
We ha e demons a ed in Chap e 3 ha adi ional mul i- iew geome y and deep lea ning
can bene i om each o he , achie ing oghe he an accu acy ha ou pe o ms bo h o hem
sepa a ely. I is wo h ema king ha ou wo k was one o he i s add essing his idea. A e
us, many no el app oaches ha e been p oposed. In pa icula , I would highligh wo o hem
ha couple deep lea ning and mul i- iew geome y qui e igh ly, ins ead o hem being wo
di e en p ocedu es subsequen ly me ged. On he one side, CodeSLAM (Bloesch e al., 2018)
and i s ollowing wo k DeepFac o s (Cza nowski e al., 2020) success ully p opose a deep
neu al ne wo k ha de ines a mani old o each dep h map, in which adi ional mul i- iew
geome y op imiza ion inds he bes dep h maps acco ding o geome ic and pho o-me ic
e o s. On he o he hand, Zhou e al. (2018) p esen ed DeepTAM, con inuing hei wo k

130 Conclusions
DeMoN in Ummenho e e al. (2017), p oposing an deep neu al ne wo k ha i e a i ely
e ines i s p edic ions using mul i- iew geome y, achie ing an imp essi e accu acy.
On isual place ecogni ion, we ha e p esen ed h ee no el app oaches o mul i- iew
global desc ip o s ha a e obus o changes in he appea ance c ea ed by di e en condi ions.
We ha e es ed i o mul iple seasons and o di e en ligh condi ions. I is also wo h
ema king ha ou wo k is he i s one p oposing mul i- iew desc ip o s based on deep
lea ning, and ha ou desc ip o s we e u he explo ed in he da ase compiled in Wa bu g
e al. (2020) wi h simila conclusions as in ou wo k. Fo u u e esea ch, i would be
in e es ing o explo e condi ion-in a ian local desc ip o s ha would allow o eco e a
me ic pose and no only a opological one (Re aud e al., 2019).
In Chap e 5 we p esen ed CFL and EquiCon s. CFL is a ne wo k ha achie es s a e-o -
he-a esul s in indoo layou eco e y, and EquiCon s a e a special ype o con olu ions
ha adap is shape o he equi ec angula dis o ion. Bo h con ibu ions can ha e se e al
applica ions in Visual SLAM (Salas e al., 2015). EquiCon s is also a gene al model, om
which any deep ne wo k using pano amic images can bene i om.
Rega ding SLAM in dynamic en i onmen s, we ha e p oposed a pipeline o a oid
dynamic o mo able objec s o pe u b mapping and acking algo i hms assuming a igid
wo ld. Ou main con ibu ion is he design o he DynaSLAM pipeline and he inclusion o a
segmen a ion CNN embedded in a isual SLAM sys em. We demons a e ha ou esul s
a e e y compe i i e, and ha ou p oposal ou pe o ms s a e-o - he-a SLAM sys ems.
A easonable line o u u e wo k would be o ocus on acking he dynamic objec s and
possibly use i also in o i s ad an age. P elimina y esul s on his di ec ion can al eady be
seen in Balles e e al. (2020). A po en ial ad an age could be, o example: i an objec is
being acked and a some poin he came a is occluded by i , he objec mo ion es ima ion
would allow a easonable es ima ion o he came a mo ion o some ime.
We can d aw a gene al conclusion o his hesis by w i ing ha , on he one hand, we
ha e p oposed se e al no el me hods o use deep ne wo ks o 3D pe cep ion challenges.
And, on he o he hand, we ha e also made con ibu ions wi hin deep lea ning o his
pa icula domain. On he i s se o p oposals, we ha e de eloped no el me hods o use
mul i- iew and single- iew dep hs and o de ec and emo e dynamic objec s in isual
SLAM. On he second se o p oposals, we ha e de eloped no el mul i- iew embeddings o
place ecogni ion and wo no el con olu ion ypes, CAM-Con s and Equicon s, explici ly
including he came a in insics and demons a ing be e pe o mances o single- iew dep h
lea ning wi h mul iple came as and layou es ima ion om equi ec angula images.
7.1 Limi a ions and Fu u e Wo k 131
7.1 Limi a ions and Fu u e Wo k
Deep lea ning has supposed a g ea ad ance in 3D isual pe cep ion and i is making i s
way in o isual SLAM. The bigges limi a ion we ound while wo king on his hesis is he
dependency on da a, and mo e speci ically on good-quali y and di e se, su icien ly well
sampled da a. To exempli y his, look a he dense dep h es ima ion p oblem. The p og ess
achie ed by using deep lea ning has no p eceden s. Howe e , i is e y easy o all in o
small segmen s o he p oblem by e alua ing he models in a subse o he eal cases, e.g. a
ela i ely small da ase on a e y speci ic and biased domain. A common case is he aining
on di e en domains sepa a ely o , as we poin ed in Chap e 2, commonly used da ase s
only p o ide images aken by one ype o came a. This is di e en o adi ional 3D ision
algo i hms ha explici ly conside he came a model and do no make any assump ion (o
he smalles possible numbe o hem) in he ype o da a o domain a p io y. In his hesis
we always kep his in mind. CAM-Con s in oduce he came a model in o con olu ions o
he i s ime. We also made use o unbiased dep h es ima ion om a adi ional geome y-
based iangula ion o complemen he lea ned dep h p edic ion. We ha e adap ed s anda d
con olu ions o equi ec angula dis o ion in EquiCon s, again aking in o accoun he came a
model. Las ly, we ha e combined deep lea ning and adi ional me hods o dynamic objec
de ec ion in DynaSLAM. We ag ee wi h he gene al hough ha deep lea ning has a g ea
po en ial o 3D pe cep ion. Howe e , u u e esea ch needs o add ess he da a dependency,
c ea ing mo e comple e and gene al benchma ks (as López-An eque a e al. (2020)) and also
models ha accoun o he 3D- o-2D p ojec ion and he da a noise. We can ci e (Cza nowski
e al., 2020, Zhou e al., 2020) as examples o he o me , and Bayesian deep lea ning ( o
example, Gus a sson e al. (2020)) as a p omising line o wo k o he la e . Lea ning
om da a migh be he key o a comple e scene unde s anding, bu , in ou belie e, only
hose p oposals ha complemen machine lea ning wi h unce ain y, geome ic and physical
models will achie e he bes pe o mance.
Re e ences
Abadi, Ma ín, Paul Ba ham, Jianmin Chen, Zhi eng Chen, Andy Da is, Je ey Dean,
Ma hieu De in, Sanjay Ghemawa , Geo ey I ing, Michael Isa d, e al. (2016), “Tenso -
low: a sys em o la ge-scale machine lea ning.” In OSDI, olume 16, 265–283.
Alcan a illa, Pablo F, José J Yebes, Ja ie Almazán, and Luis M Be gasa (2012), “On
combining isual SLAM and dense scene low o inc ease he obus ness o localiza ion
and mapping in dynamic en i onmen s.” In ICRA.
Alonso, Inigo, Luis Riazuelo, and Ana C Mu illo (2020), “Minine : An e icien seman ic
segmen a ion con ne o eal- ime obo ic applica ions.” IEEE T ansac ions on Robo ics.
Amb us, Ra es, John Folkesson, and Pa ic Jens el (2016), “Unsupe ised objec segmen a-
ion h ough change de ec ion in a long e m au onomy scena io.” In Humanoid Robo s
(Humanoids), IEEE.
Amodei, Da io, Sunda am Anan hana ayanan, Rishi a Anubhai, Jingliang Bai, E ic Ba -
enbe g, Ca l Case, Ja ed Caspe , B yan Ca anza o, Qiang Cheng, Guoliang Chen, e al.
(2016), “Deep speech 2: End- o-end speech ecogni ion in english and manda in.” In
In e na ional con e ence on machine lea ning, 173–182.
A andjelo ic, Relja, Pe G ona , Akihiko To ii, Tomas Pajdla, and Jose Si ic (2016),
“Ne VLAD: CNN a chi ec u e o weakly supe ised place ecogni ion.” In P oceedings o
he IEEE Con e ence on Compu e Vision and Pa e n Recogni ion, 5297–5307.
A andjelo i´
c, Relja and And ew Zisse man (2014), “Disloca ion: Scalable desc ip o dis-
inc i eness o loca ion ecogni ion.” In Asian Con e ence on Compu e Vision, 188–204,
Sp inge .
A meni, I., A. Sax, A. R. Zami , and S. Sa a ese (2017), “Join 2D-3D-Seman ic Da a o
Indoo Scene Unde s anding.” A Xi .
A meni, I o, Sasha Sax, Ami R Zami , and Sil io Sa a ese (2017), “Join 2D-3D-seman ic
da a o indoo scene unde s anding.” a Xi p ep in a Xi :1702.01105.
A oyo, Robe o, Pablo F Alcan a illa, Luis M Be gasa, and Edua do Rome a (2016), “Fusion
and bina iza ion o CNN ea u es o obus opological localiza ion ac oss seasons.” In
In elligen Robo s and Sys ems (IROS), 2016 IEEE/RSJ In e na ional Con e ence on,
4656–4663, IEEE.
Balles e , I ene, Alejand o Fon an, Ja ie Ci e a, Klaus H S obl, and Rudolph T iebel
(2020), “Do : Dynamic objec acking o isual slam.” a Xi p ep in a Xi :2010.00052.