scieee Science in your language
[en] (orig)

The joint role of geometry and illumination on material recognition

Abstract

Observing and recognizing materials is a fundamental part of our daily life. Under typical viewing conditions, we are capable of effortlessly identifying the objects that surround us and recognizing the materials they are made of. Nevertheless, understanding the underlying perceptual processes that take place to accurately discern the visual properties of an object is a long-standing problem. In this work, we perform a comprehensive and systematic analysis of how the interplay of geometry, illumination, and their spatial frequencies affects human performance on material recognition tasks. We carry out large-scale behavioral experiments where participants are asked to recognize different reference materials among a pool of candidate samples. In the different experiments, we carefully sample the information in the frequency domain of the stimuli. From our analysis, we find significant first-order interactions between the geometry and the illumination, of both the reference and the candidates. In addition, we observe that simple image statistics and higher-order image histograms do not correlate with human performance. Therefore, we perform a high-level comparison of highly nonlinear statistics by training a deep neural network on material recognition tasks. Our results show that such models can accurately classify materials, which suggests that they are capable of defining a meaningful representation of material appearance from labeled proximal image data. Last, we find preliminary evidence that these highly nonlinear models and humans may use similar high-level factors for material recognition tasks. Lagunas, Manuel; Serrano, Ana; Gutiérrez, Diego; Masiá, Belén

Read accessible full text

The joint role of geometry and illumination on material recognition

Author: Lagunas, Manuel; Masiá, Belén; Gutiérrez, Diego; Serrano, Ana
Year: 2021
DOI: 10.1167/jov.21.2.2
Source: https://zaguan.unizar.es/record/100710/files/texto_completo.pdf
Jou nal o Vision (2021) 21(2):2, 1–18 1
The join ole o geome y and illumina ion on ma e ial
ecogni ion
Manuel Lagunas Uni e sidad de Za agoza, I3A, Za agoza, Spain
Ana Se ano Uni e sidad de Za agoza, I3A, Max Planck Ins i u e o
In o ma ics, Za agoza, Spain
Diego Gu ie ez Uni e sidad de Za agoza, I3A, Za agoza, Spain
Belen Masia Uni e sidad de Za agoza, I3A, Za agoza, Spain
Obse ing and ecognizing ma e ials is a undamen al
pa o ou daily li e. Unde ypical iewing condi ions,
we a e capable o effo lessly iden i ying he objec s
ha su ound us and ecognizing he ma e ials hey a e
made o . Ne e heless, unde s anding he unde lying
pe cep ual p ocesses ha ake place o accu a ely
disce n he isual p ope ies o an objec is a
long-s anding p oblem. In his wo k, we pe o m a
comp ehensi e and sys ema ic analysis o how he
in e play o geome y, illumina ion, and hei spa ial
equencies affec s human pe o mance on ma e ial
ecogni ion asks. We ca y ou la ge-scale beha io al
expe imen s whe e pa icipan s a e asked o ecognize
diffe en e e ence ma e ials among a pool o candida e
samples. In he diffe en expe imen s, we ca e ully
sample he in o ma ion in he equency domain o he
s imuli. F om ou analysis, we find significan fi s -o de
in e ac ions be ween he geome y and he
illumina ion, o bo h he e e ence and he candida es.
In addi ion, we obse e ha simple image s a is ics and
highe -o de image his og ams do no co ela e wi h
human pe o mance. The e o e, we pe o m a high-le el
compa ison o highly nonlinea s a is ics by aining a
deep neu al ne wo k on ma e ial ecogni ion asks. Ou
esul s show ha such models can accu a ely classi y
ma e ials, which sugges s ha hey a e capable o
defining a meaning ul ep esen a ion o ma e ial
appea ance om labeled p oximal image da a. Las , we
find p elimina y e idence ha hese highly nonlinea
models and humans may use simila high-le el ac o s
o ma e ial ecogni ion asks.
In oduc ion
Unde ypical iewing condi ions, humans a e
capable o e o lessly ecognizing ma e ials and
in e ing many o hei key physical p ope ies, jus
by b ie ly looking a hem. Al hough his is almos
an e o less p ocess, i is no a i ial ask. The
image ha is inpu o ou isual sys em esul s om
a complex combina ion o he su ace geome y, he
e lec ance o he ma e ial, he dis ibu ion o ligh s in
he en i onmen , and he obse e ’s poin o iew. To
ecognize he ma e ial o a su ace while being in a ian
o o he ac o s o he scene, ou isual sys em ca ies
ou an unde lying pe cep ual p ocess ha is no ye
ully unde s ood (Adelson, 2000;D o e al., 2001a;
Fleming e al., 2001).
So how does ou b ain ecognize ma e ials? We could
hink ha , simila o sol ing an in e se op ics p oblem,
ou b ain is es ima ing he physical p ope ies o each
ma e ial (Pizlo, 2001). This would imply knowledge
o many o he physical quan i ies abou he objec
and i s su ounding scene, om which ou b ain could
disen angle he e lec ance o he su ace. Howe e ,
we a ely ha e access o such p ecise in o ma ion,
so a ia ions based on Bayesian in e ence ha e been
p oposed (Ke s en e al., 2004).
O he app oaches a e based on image s a is ics,
and explain ma e ial ecogni ion as a p ocess whe e
ou b ain ex ac s image ea u es ha a e ele an o
desc ibe ma e ials. Then, i would y o ma ch hem
wi h p e iously acqui ed knowledge, o disce n he
ma e ial we a e obse ing. Conside ing his app oach
ou isual sys em would dis ega d he illumina ion,
mo ion, o o he ac o s in he scene and y o
ecognize ma e ials by ep esen ing hei ypical
appea ance in e ms o ea u es ins ead o explici ly
acqui ing an accu a e physical desc ip ion o each
ac o . This ype o image analysis can be ca ied ou
in he p ima y domain (Adelson, 2008;Fleming, 2014;
Geisle , 2008;Mo oyoshi e al., 2007;Nishida & Shinya,
1998), o in he equency domain (B ady & Oli a,
Ci a ion: Lagunas, M., Se ano, A., Gu ie ez, D., & Masia, B. (2021). The join ole o geome y and illumina ion on ma e ial
ecogni ion. Jou nal o Vision,21(2):2, 1–18, h ps://doi.o g/10.1167/jo .21.2.2.
h ps://doi.o g/10.1167/jo .21.2.2 Recei ed May 4, 2020; published Feb ua y 3, 2021 ISSN 1534-7362 Copy igh 2021 The Au ho s
This wo k is licensed unde a C ea i e Commons A ibu ion-NonComme cial-NoDe i a i es 4.0 In e na ional License.
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 2
Figu e 1. Two sphe es made o sil e , unde wo diffe en
illumina ions, leading o comple ely diffe en pixel-le el
s a is ics.
Figu e 2. Two objec s o diffe en geome ies bu made o he
same ma e ial, unde he same illumina ion. The objec on he
le seems o be made o a shinie ma e ial.
2012;Giesel & Zaidi, 2013;Oli a & To alba, 2001).
Howe e , i is a gued i ou isual sys em ac ually
de i es any aspec s o ma e ial pe cep ion om such
simple s a is ics (Ande son & Kim, 2009). Fo ins ance,
Fleming and S o s (2019) ha e ecen ly p oposed he
idea ha highly nonlinea encodings o he isual inpu
may be e explain he unde lying p ocesses o ma e ial
pe cep ion.
In his wo k, we ho oughly analyze how he
con ounding e ec s o illumina ion and geome y
in luence human pe o mance in ma e ial ecogni ion
asks. The same ma e ial can yield di e en appea ances
owing o changes in illumina ion and/o geome y
(Figu es 1 and 2), al hough i is possible o ha e wo
di e en ma e ials look he same by weaking he wo
pa ame e s (Vango p e al., 2007). We aim o u he
ou unde s anding o he complex in e play be ween
geome y and illumina ion in ma e ial ecogni ion.
We ha e ca ied ou la ge-scale, igo ous online
beha io al expe imen s whe e pa icipan s we e asked
o ecognize di e en ma e ials, gi en images o one
e e ence ma e ial and a pool o candida es. By using
pho o ealis ic compu e g aphics, we ob ain ca e ully
con olled s imuli, wi h a ying deg ees o in o ma ion
in he equency domain. In addi ion, we obse e
ha simple image s a is ics, image his og ams, and
his og ams o V1-like subband il e s do no co ela e
wi h human pe o mance in ma e ial ecogni ion asks.
Inspi ed by Fleming and S o s’ ecen wo k (2019),
we analyze highly nonlinea s a is ics by aining a
deep neu al ne wo k. We obse e ha such s a is ics
de ine a obus and accu a e ep esen a ion o ma e ial
appea ance and ind p elimina y e idence ha hese
models and humans may sha e simila high-le el ac o s
when ecognizing ma e ials.
Ma e ial ecogni ion
Recognizing ma e ials and in e ing hei key ea u es
by sigh is in aluable o many asks. Ou expe ience
sugges s ha humans a e able o co ec ly p edic a
wide a ie y o ough ma e ial ca ego ies like ex iles,
s ones, o me als (Fleming & Bül ho , 2005;Fleming,
2014;Ged e al., 2010;Li & F i z, 2012); o i ems
ha we would call “s u ” (Adelson, 2001)—like sand
o snow. Humans a e also capable o iden i ying he
ma e ials in a pho og aph by b ie ly looking a hem
(Sha an e al., 2009,2008) o o in e ing hei physical
p ope ies wi hou he need o ouch hem (Fleming
e al., 2013,2015a;Ja abo e al., 2014;Maloney &
B aina d, 2010;Nagai e al., 2015;Se ano e al.,
2016). This abili y is buil om expe ience, by ac ually
con i ming isual imp essions wi h o he senses. This
way, ma e ial pe cep ion becomes a cogni i e p ocess
(Palme , 1975) whose unde lying in icacies a e no
ully unde s ood ye (Ande son, 2011;Fleming e al.,
2015b;Thompson e al., 2011).
In e play o geome y and illumina ion
Ma e ial pe cep ion is a complex p ocess ha
in ol es a la ge numbe o dis inc dimensions (Mao e
al., 2019;Obein e al., 2004;Sè e, 1993) ha , some imes,
a e impossible o physically measu e (Hun e e al.,
1937). The illumina ion o a scene (Beck & P azdny,
1981;Bousseau e al., 2011;Zhang e al., 2015)and
he shape o a su ace, a e esponsible o he inal
appea ance o an objec (Nishida & Shinya, 1998;
Schlü e & Faul, 2019;Vango p e al., 2007) and,
he e o e, o ou pe cep ion o he ma e ials i is
made o (Olkkonen & B aina d, 2011). Humans a e
capable o es ima ing he e lec ance p ope ies o
a su ace (D o e al., 2001b) e en when he e is no
in o ma ion abou i s illumina ion (D o e al., 2001a;
Fleming e al., 2001), ye we pe o m be e unde
illumina ions ha ma ch eal-wo ld s a is ics (Fleming
e al., 2003). Indeed, geome y and illumina ion ha e a
join in e ac ion in ou pe cep ion o glossiness (Faul,
2019;Leloup e al., 2010;Ma low e al., 2012;Olkkonen
& B aina d, 2011)andcolo (Bloj e al., 1999). In his
wo k, we explo e he in e play o shape, illumina ion,
and hei spa ial equencies in ou pe o mance a
ecognizing ma e ials. To achie e ha , we launched
igo ous online beha io al expe imen s whe e we ely
on ealis ic compu e g aphics o gene a e he s imuli
and ca e ully a y hei in o ma ion in he equency
domain.
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 3
Image s a is ics and ma e ial pe cep ion
One o he goals in ma e ial pe cep ion esea ch is
o un angle he p ocesses ha happen on ou isual
sys em o comp ehend hei oles and know wha
in o ma ion hey ca y. The e is an ongoing discussion
on whe he ou isual sys em is sol ing an in e se
op ics p oblem (Kawa o e al., 1993;Pizlo, 2001)o i i
ma ches he s a is ics o he inpu o ou isual sys em
(Adelson, 2000;Mo oyoshi e al., 2007;Thompson
e al., 2016) o unde s and he wo ld ha su ounds us.
La e s udies ega ding ou isual sys em and how we
pe cei e ma e ials dismiss he in e se op ics app oach
and claim ha i is unlikely ha ou b ain es ima es
he pa ame e s o he e lec ance o a su ace, when,
o ins ance, we wan o measu e glossiness (Fleming,
2014;Geisle , 2008). Ins ead, hey sugges ha ou
isual sys em joins low and midle el s a is ics o make
judgmen s abou su ace p ope ies (Adelson, 2008).
On his hypo hesis, Mo oyoshi e al. (2007) sugges ha
he human isual sys em could be using some so o
measu e o his og am symme y o dis inguish glossy
su aces. O he wo ks ha e explo ed image s a is ics in
he equency domain (Hawken & Pa ke , 1987;Schille
e al., 1976), o ins ance, o cha ac e ize ma e ial
p ope ies (Giesel & Zaidi, 2013), o o disc imina e
ex u es (Julesz, 1962;Scha ali zky & Zisse man, 2001).
Howe e , i is a gued ha , i ou isual sys em ac ually
de i es any aspec s o ma e ial pe cep ion om simple
s a is ics (Ande son & Kim, 2009;Kim & Ande son,
2010;Olkkonen & B aina d, 2010). Ins ead, ecen
wo k by Fleming and S o s (2019) p oposes ha , o
in e he p ope ies o he scene, ou isual sys em
is doing an e icien and accu a e encoding o he
p oximal s imulus (image inpu o ou isual sys em).
Thus, highly nonlinea models, such as deep neu al
ne wo ks, may be e explain human pe cep ion. In line
wi h such obse a ions, Bell e al. (2015)showhow
deep neu al ne wo ks can be ained in a supe ised
ashion o accu a ely ecognize ma e ials, and Wang
e al. (2016) la e ex end i o also ecognize ma e ials
in ligh ields. Close o ou wo k, Lagunas e al. (2019)
de ise a deep lea ning-based ma e ial simila i y me ic
ha co ela es wi h human pe cep ion. They collec ed
judgemen s on pe cei ed ma e ial simila i y as a whole,
no explici ly aking in o accoun he in luence o
geome y o illumina ion, and build hei me ic upon
such judgemen s. In con as , we ocus on analyzing o
which ex en geome y and illumina ion do in e e e
wi h ou pe cep ion o ma e ial appea ance. We
launch se e al beha io al expe imen s wi h ca e ully
con olled s imuli, and ask pa icipan s o speci y
which ma e ials a e close o a e e ence. In addi ion,
aking inspi a ion om hese ecen wo ks, we explo e
how highly nonlinea models, such as deep neu al
ne wo ks, pe o m in ma e ial classi ica ion asks.
We ind ha such models a e capable o accu a ely
Figu e 3. G aphical use in e ace o he online beha io al
expe imen s. In pa icula , his sc eensho belongs o he TEST
SH. On he le , he use can see he e e ence ma e ial
oge he wi h he cu en selec ion. On he igh , she can
obse e all he candida e ma e ials. To selec one candida e
ma e ial, he use clicks on he co esponding image and i is
au oma ically added o he selec ion box on he le .
ecognizing ma e ials, and u he obse e ha deep
neu al ne wo ks may sha e simila high-le el ac o s o
humans when ecognizing ma e ials.
Me hods
We ca ied ou a se o online beha io al
expe imen s whe e we analyze he in luence o
geome y, illumina ion, and hei equencies in human
pe o mance o ma e ial ecogni ion asks. Pa icipan s
a e p esen ed wi h a e e ence ma e ial and hei main
ask is o pick i e ma e ials om a pool o candida es
ha hey hink a e close o he e e ence. A sc eensho
o he expe imen can be seen in Figu e 3.
S imuli
We ob ain ou s imuli om he da ase p oposed by
Lagunas e al. (2019). This da ase con ains images
c ea ed using pho o ealis ic compu e g aphics, wi h 15
di e en geome ies, 6 di e en eal-wo ld illumina ions
anging om indoo scena ios o u ban o na u al
landscapes, and 100 di e en ma e ials measu ed om
hei eal-wo ld coun e pa s which we e pooled om
Mi subishi Elec ic Resea ch Labo a o ies (MERL)
da abase (Ma usik e al., 2003). We sample he ollowing
ac o s o ou expe imen s:
Geome ies. Among he geome ies ha he da ase
con ains, we choose he sphe e and Ha an-2 geome y
(Ha an e al., 2016). These a e low and high spa ial
equency geome ies, espec i ely, sui able o es how
he spa ial equencies o he geome y a ec he inal
appea ance o he ma e ial and ou pe o mance a
ecognizing i .
•Sphe e: Rep esen ing a smoo h, and low spa ial
equency geome y, widely adop ed in p e ious
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 4
Figu e 4. Examples o he s imuli in each diffe en online
beha io al expe imen . On he le , we show an example o he
e e ence s imuli wi h one o he six illumina ions. On he igh ,
we show a small subse (6 o he 100 ma e ials) o he
candida e s imuli wi h S . Pe e s illumina ion.
beha io al expe imen s (Filip e al., 2008;Ja abo
e al., 2014;Ke & Pellacini, 2010;Sun e al., 2017).
•Ha an-2:1I is a geome y wi h high spa ial
equencies, and wi h high spa ial a ia ions ha
has been ob ained h ough op imiza ion echniques.
•Ha an-2: Su ace has had signi ican success
in ecen pe cep ual s udies and applica ions
(Gua ne a e al., 2018;Guo e al., 2018;Lagunas
e al., 2019;Vá a & Filip, 2016).
The s imuli in each di e en expe imen can be
obse ed in Figu e 4. The geome y in he e e ence
and candida e samples changes depending on he
expe imen , he de ails a e as ollows:
• Tes HH: Bo h he e e ence and he candida es
depic Ha an geome y.
• Tes HS: The e e ence depic s Ha an and he
candida es depic he sphe e.
• Tes SH: The e e ence depic s he sphe e while he
candida es depic Ha an.
• Tes SS: Bo h he e e ence and he candida es
depic he sphe e geome y.
Illumina ions. To p e en a pu e ma ching ask, we
choose di e en illumina ions be ween he e e ence
and candida e ma e ials o all beha io al expe imen s.
• The e e ence samples depic six di e en
illumina ions cap u ed om he eal wo ld. All
illumina ions can be obse ed in Figu e 5.To
ha e an in ui ion o he con en in he cap u ed
illumina ion, he inse s show he RGB in ensi y o
he ho izon al pu ple line. We use all illumina ions
in he da ase since hey con ain a mix o spa ial
equencies sui able o empi ically es how he
spa ial equencies o he illumina ion may a ec
human pe o mance on ma e ial ecogni ion asks.
The illumina ions G ace,Ennis,andU izi ha e
a b oad spa ial equency spec um, Pisa and
Doge mainly con ain medium and low-spa ial
equency con en , while Glacie mainly has
Figu e 5. Le : All illumina ions depic ed in he online beha io al
expe imen s. The inse co esponds o he pixel in ensi y o
he ho izon al pu ple line. Righ : Magni ude spec um o he
luminance o each illumina ion.
low-spa ial equency con en . To simpli y he
no a ion, we e e o hem h oughou he a icle
as high- equency, medium- equency, and
low- equency illumina ions, espec i ely.
• The candida e samples depic he S . Pe e s
illumina ion (excep in an addi ional expe imen
discussed in he Discussion whe e hey depic Doge
illumina ion). S . Pe e s is an illumina ion ha has
been used in he pas o se e al pe cep ual s udies
(Fleming e al., 2003;Se ano e al., 2016), and i
can be seen in Figu e 5. The inse shows he RGB
pixel in ensi y o he ho izon al pu ple line.
To quan i y he spa ial equencies o he
illumina ions, we ha e employed he high- equency
con en (HFC) measu e (B ossie e al., 2004). This
measu e cha ac e izes he equencies in a signal by
summing linea ly weigh ed alues o he spec al
magni ude, hus a oiding o a bi a ily choose a
sepa a ion be ween high and low equencies, o isually
assessing he slope o he 1/ ampli ude spec um. A
high HFC alue means highe equencies in he signal.
Figu e 6 shows he HFC o each illumina ion.
Ma e ials
We use all he ma e ials om he Lagunas e al.
da ase Lagunas e al. (2019). The e e ence ials a e
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 5
Figu e 6. HFC measu e compu ed o all he candida e and
e e ence illumina ions. We can obse e how high- equency
illumina ions (Uffizi,G ace,Ennis,S . Pe e s) also ha e a high
HFC alue, medium- equency illumina ions (Pisa,Doge)ha ea
lowe HFC alue, and, las , low- equency illumina ions (Glacie )
ha e he lowes HFC alue.
sampled uni o mly o co e all 100 ma e ial samples
in he da ase . Examples o he s imuli used in each
beha io al expe imen a e shown in Figu e 4,whe e he
image on he le shows he e e ence ma e ial and he
igh a ea shows a subse o he candida e ma e ials.
Pa icipan s
The online beha io al expe imen s we e designed
o wo k ac oss pla o ms on s anda d web b owse s,
and hey we e conduc ed h ough he Amazon
Mechanical Tu k (MTu k) pla o m. In o al, 847
unique use s ook pa in hem (368 use s belonging
o he expe imen s explained in Resul s, and 479
belonging o he addi ional expe imen s explained
in he Discussion), 44.61% o hem emale. Among
he pa icipan s, 62.47% claimed o be amilia wi h
compu e g aphics, 25.57% had no p e ious expe ience
and 9.96% decla ed hemsel es p o essionals. We also
sampled da a ega ding he de ices used du ing he
expe imen s: 94.10% used a moni o , 4.30% used a
able , and 1.60% used a mobile phone. In addi ion,
he mos common sc een size was 1366 ×728 pixels
(42.01% o pa icipan s), minimum sc een size was 640
×360 pixels ( wo people), and a maximum o 2560 ×
1414 pixels (one pe son). Use s we e no awa e o he
pu pose o he beha io al expe imen .
P ocedu e
Subjec s a e shown a e e ence sample and a g oup
o candida e ma e ial samples. Each expe imen , HIT
in MTu k e minology, consis s o 23 unique e e ence
ma e ial samples o ials, 36 o which a e sen inels used
o de ec malicious o lazy use s. Use s a e asked o
“selec i e ma e ial samples which you belie e a e close
o he one shown in he e e ence image.” Addi ionally,
we ins uc hem o make hei selec ion in dec easing
o de o con idence. We le he use s pick i e candida e
ma e ials because jus one answe would p o ide spa se
esul s. We launched 25 HITs o each expe imen
and each HIT was answe ed by six di e en use s.
This esul ed in a o al o 27.000 nonsen inel ials,
12.000 belonging o he ou expe imen s analyzed in
he Resul s, and 15.000 o hem belonging o he i e
addi ional expe imen s discussed in he Discussion
(a o al o nine di e en expe imen s wi h 25 HITs each,
each HIT answe ed by six use s and 20 nonsen inel
ials pe HIT). Use s we e no allowed o epea he
same HIT.
The se o ma e ials in he candida e samples does
no a y ac oss HITs; howe e , he posi ion o each
sample is andomized o each ial. This has a wo- old
pu pose: i p e en s he use om memo izing he
posi ion o he samples, and i p e en s hem om
selec ing only he candida e samples ha appea a he
op o hei sc een. The e e ence samples do no epea
ma e ials du ing a HIT and he e e ence ma e ial is
always p esen among he candida e samples. Du ing
he expe imen , s imuli keep a cons an display size o
300 ×300 pixels o he e e ence, and o 120 ×120
pixels o he candida e s imuli (excep o some o he
addi ional expe imen s explained in Discussion whe e
bo h e e ence and candida e s imuli a e displayed a
ei he 300 ×300 pixels o 120 ×120 pixels). Figu e 3
shows a sc eensho wi h he g aphical use in e ace
du ing he beha io al expe imen s. On he le -hand
side, we can obse e he selec ion panel wi h he cu en
ial and he selec ion o he cu en ma e ials. The
igh -hand side displays he se o candida e ma e ials
whe eo use s can pick hei selec ion. Use s we e no
able o go back and edo an al eady answe ed ial, bu
hey could edi hei cu en selec ion o i e ma e ials
un il hey we e sa is ied wi h hei choice. Addi ionally,
once he 23 ials o he HIT a e answe ed, o ha e an
in ui ion abou he main ea u es ha humans use o
ma e ial ecogni ion, we asked he use : “Which isual
cues did you conside o pe o m he es ?”
To minimize wo ke un eliabili y, he use pe o ms a
b ie aining be o e he eal es (Welinde e al., 2010).
To a oid gi ing he use u he in o ma ion abou he
es , we use a di e en geome y (Ha an-3 Ha an
e al., 2016) du ing he aining phase. In his phase,
he i ems o he in e ace a e explained and he use is
gi en guidance on how o pe o m he es using jus a
ew images (Ga ces e al., 2014;Lagunas e al., 2018;
Rubins ein e al., 2010).
Sen inels
Each sen inel shows a andomly selec ed image om
he pool o candida es as he e e ence sample. We
conside use answe s o he sen inel as alid i hey pick
Downloaded om jo .a ojou nals.o g on 03/15/2021

Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 6
he igh ma e ial wi hin hei i e selec ions, ega dless
o he o de . We ejec ed use s who did no co ec ly
answe a leas one o he h ee sen inel ques ions. To
ensu e ha use s’ answe s we e well hough and ha
hey we e paying a en ion o he expe imen , we also
ejec ed use s ha ook less han 5 seconds pe ial (on
a e age). In he end, we adop a conse a i e app oach
and ejec ed 19.8% o he pa icipan s, ga he ing 21.660
answe s (9.560 belonging o he beha io al expe imen s
explained in he Resul s and 12.100 belonging o he
addi ional expe imen s explained in he Discussion).
Resul s
We in es iga e which ac o s ha e a signi ican
in luence on use pe o mance and on he ime hey
ook o comple e each ial in he ou expe imen s: Tes
HH, Tes HS, Tes SH, and Tes SS. The ac o s we
include a e: he e e ence geome y G e , he candida e
geome y Gcand, and he illumina ion o he e e ence
sample I e , as well as hei i s -o de in e ac ions
( ecall ha he illumina ion o he candida e samples
emains cons an in hese beha io al expe imen s). We
also include he O de o appea ance o each ial.
We use a gene al linea mixed model wi h a binomial
dis ibu ion o he pe o mance since i is well-sui ed
o bina y dependen a iables like ou s, and a nega i e
binomial dis ibu ion o he ime, which p o ides
mo e accu a e models han he Poisson dis ibu ion
by allowing he mean and a iance o be di e en .
Because we canno assume ha ou obse a ions a e
independen , we model he po en ial e ec o each
pa icula subjec iewing he s imuli as a andom
e ec . Because we ha e ca ego ical a iables among
ou p edic o s, we e-code hem o dummy a iables
o he eg ession. In all ou es s, we ix a signi icance
alue (P- alue) o 0.05. Finally, o ac o s ha p esen
a signi ican in luence, we u he pe o m pai wise
compa isons o all hei le els (leas signi ican
di e ence pai wise mul iple compa ison es ).
Analysis o use pe o mance and ime
In ou online beha io al expe imen s, we ely on he
op i e accu acy o measu e use pe o mance. This
me ic conside s an answe as co ec i he e e ence is
among he i e candida e ma e ials ha he use picked
in he ial. Because pa icipan s picked i e ma e ials
anked in descending o de o con idence, he op one
accu acy could also be conside ed o ou analysis.
Howe e , he ask hey ha e o sol e is no easy and
use s ha e an o e all op one accu acy o 9.21% which
yields spa se esul s. A andom selec ion would yield a
op one accu acy o 1% and a op i e accu acy o 5%.
Figu e 7. Le : Top fi e accu acy o each o he ou beha io al
expe imen . Cen e : Top fi e accu acy o each e e ence
geome y G e .Righ : Top fi e accu acy o he candida e
geome y Gcand. We can see how use s seem o pe o m
be e when he candida e and e e ence a e a high- equency
geome y. All plo s ha e a 95% confidence in e al. The names
ma ked wi h ∗a e ound o ha e s a is ically significan
diffe ences.
Influence o he geome y
The e is a clea e ec in use pe o mance when he
he geome y changes, ega dless i ha change happens
in he candida e (Gcand,P=0.005) o he e e ence
geome y (G e ,P<0.001). This inding is expec ed,
because he geome y plays a key ole in how a su ace
e lec s he incoming ligh and, he e o e, will ha e an
impac on he inal appea ance o he ma e ial. Figu e 7
shows use pe o mance in e ms o op i e accu acy
wi h a 95% con idence in e al when he e e ence and
candida e geome y change join ly (le ) o indi idually
(cen e and igh ). Use s seem o pe o m be e when
hey ha e o ecognize he ma e ial in a high- equency
geome y compa ed wi h a low- equency one. Those
esul s also sugges ha changes in he equencies o
he e e ence geome y may ha e a bigge impac on
use pe o mance han changes in he equencies o
he candida e geome y (i.e., use s pe o m be e wi h
a high- equency e e ence geome y and low- equency
candida e geome y, compa ed o a low- equency
e e ence geome y and a high- equency candida e
geome y).
Influence o he e e ence illumina ion
We obse e ha he illumina ion o he e e ence
image has a signi ican e ec on use pe o mance
(I e ,P<0.001). This inding is expec ed because all
he ma e ials in a scene a e e lec ing he ligh ha
eaches hem; he e o e, he changes in illumina ion
can signi ican ly in luence he inal appea ance o a
ma e ial, and how we pe cei e i (Bousseau e al., 2011).
Figu e 8, le , shows he op i e accu acy o each
e e ence illumina ion and g oups o illumina ions wi h
s a is ically indis inguishable pe o mance. We can
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 7
Figu e 8. Le : Top fi e accu acy o each e e ence illumina ion (I e ). We can see how use s seem o pe o m be e wi h
high- equency illumina ions (Uffizi,G ace,Ennis), while hei pe o mance is wo se wi h a low- equency illumina ion (Glacie ).
Addi ionally, hey ha e an in e media e pe o mance o medium- equency illumina ions (Doge and Pisa). Cen e : Top fi e accu acy
o each e e ence illumina ion when he candida e geome y (Gcand) changes. We can obse e how use s seem o pe o m
significan ly be e wi h a high- equency geome y (Ha an) and illumina ion. On he o he hand, o low- equency illumina ions,
changes in he candida e geome y yield s a is ically indis inguishable pe o mance. Righ : Top fi e accu acy o each e e ence
illumina ion when he e e ence geome y (G e ) changes. We can obse e how use s seem o pe o m significan ly be e o all
high- equency illumina ions, excep o G ace. The ho izon al lines unde he x-axis ep esen g oups o s a is ically indis inguishable
pe o mance. We can obse e how he g oups usually clus e high-, medium- and low- equency illumina ions. The e e ence
illumina ions ma ked wi h ∗deno e significan diffe ences in use pe o mance be ween geome ies o ha illumina ion. The e o
ba s co espond o a 95% confidence in e al.
obse e how use s seem o ha e be e pe o mance
when he su ace hey a e e alua ing has been li wi h a
high- equency illumina ion (Ennis,G ace,andU izi),
whe eas use s seem o pe o m wo se in scenes wi h a
low- equency illumina ion (Glacie ); use s show an
in e media e pe o mance wi h a medium- equency
illumina ion (Doge and Pisa). Mo eo e , we pe o med
a leas signi ican di e ence pai wise mul iple
compa ison es o ob ain g oups o illumina ions
wi h s a is ically indis inguishable pe o mance. These
g oups can be obse ed in Figu e 8, unde he x-axis. I
we ocus on I e we can see how high- (g een), medium-
(blue), and low- equency ( ed) illumina ions yield
g oups o simila pe o mance. The e is an addi ional
g oup o s a is ically indis inguishable pe o mance
ep esen ed in pink.
Influence o ial o de
The o de o appea ance o he ials du ing he
expe imen does no ha e a signi ican in luence in use s
pe o mance (O de ,P=0.391).
Fi s o de in e ac ions
We ind ha he in e ac ion be ween he candida e
geome y and he e e ence illumina ion has a
signi ican e ec on use pe o mance (Gcand ∗I e ,
P<0.001). Use s seem o pe o m be e wi h a
high- equency geome y (compa ed wi h a low-
equency one) when he e e ence s imuli ea u es a
high- equency illumina ion (I e =U izi,P=0.019;
I e =[G ace, Ennis], P<0.001). On he o he
hand, he e seems o be no signi ican changes in
pe o mance be ween a high- and low- equency
candida e geome y when he e e ence s imuli has
a medium- o low- equency illumina ion (I e =
Doge,P=0.453; I e =Pisa,P=0.381; I e =
Glacie ,P=0.770). We a gue ha use pe o mance
is d i en by he e e ence sample. When he e e ence
ma e ial is li wi h a low- equency illumina ion,
use s seem o no be able o p ope ly ecognize i .
The e o e, changes in he candida e geome y a e
no ele an o use pe o mance. These esul s can
be seen in Figu e 8, cen e . Fu he mo e, unde he
x-axis, we can obse e he g oups wi h s a is ically
indis inguishable pe o mance whe e high-, medium-,
and low- equency illumina ions yield g oups o simila
pe o mance.
We also ound ou ha he in e ac ion be ween he
e e ence geome y and he e e ence illumina ion has
a signi ican impac in use pe o mance (G e ∗I e ,
P=0.012). Use s seem o show be e pe o mance
o all illumina ions wi h a high- equency e e ence
geome y (G e =Ha an,I e =U izi,P=0.002; I e
=[Ennis, Pisa, Doge, Glacie ], P<0.001), excep o
G ace illumina ion (P=0.176), whe e he di e ences in
humans pe o mance a e s a is ically indis inguishable.
These esul s, oge he wi h he g oups o s a is ically
indis inguishable pe o mance, can be seen in Figu e 8,
igh .
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 8
Figu e 9. Visualiza ions o use answe s o each o he ou online beha io al expe imen s (namely, TEST HH, TEST HS, TEST SH, and TEST
SS) using he -STE algo i hm (Van De Maa en & Weinbe ge , 2012). The inse shows he colo o each ma e ial based on he colo
classifica ion p oposed by Lagunas e al. (2019). We can see how, o all expe imen s, ma e ials wi h simila colo p ope ies a e
g ouped oge he . Fu he mo e, i we explo e he colo clus e s indi idually, we can see how he e is a second-le el a angemen by
eflec ance p ope ies. These obse a ions sugges ha use s may be pe o ming a wo-s ep p ocess while ecognizing ma e ials
whe e fi s , hey so hem ou by colo , and second, by eflec ance p ope ies.
In gene al, we canno conclude ha he e a e
signi ican changes in pe o mance due o he
in e ac ion be ween he candida e and e e ence
geome y (G e ∗Gcand,P=0.407). Ne e heless, wi h
a low- equency e e ence geome y (G e =sphe e),
use s seem o pe o m signi ican ly be e wi h a
high- equency candida e geome y (Gcand =Ha an,
P=0.009).
Analysis o he ime spen on each ial
To accoun o ime, we measu e he numbe o
milliseconds ha passed since he ial loaded in hei
sc een and un il hey picked all i e ma e ials and
p essed he “Con inue” bu on.
Influence o ial o de
We ind ha he o de o he ials has a signi ican
in luence on he a e age ime use s spend o answe
hem (P<0.001). Use s spend mo e ime in he i s
ques ions and ha a e ew ials he a e age ime
hey spend becomes s able a a ound 20 seconds
pe ial ( ecall ha he o de does no in luence
pe o mance). This phenomenon is expec ed as
use s ha e o amilia ize wi h he expe imen du ing
he i s i e a ions. As he es ad ances, hey lea n
how o in e ac wi h i and he ime hey spend
becomes s able. Addi ional igu es and esul s on he
ac o s ha in luence he spen ime can be ound in
Appendix A.
High-le el ac o s d i ing ma e ial ecogni ion
In addi ion o he analysis, we also y o gain
in ui ion on which high-le el ac o s d i e ma e ial
ecogni ion, in es iga e how simple image s a is ics
and image his og ams co ela e wi h human answe s,
and analyze highly nonlinea s a is ics in ma e ial
classi ica ion asks by aining a deep neu al ne wo k.
Visualizing use answe s
To gain in ui ion on which high-le el ac o s
humans migh use while ecognizing ma e ials, we
use a s ochas ic iple embedding me hod called he
( -S uden s ochas ic iple embedding ( -STE) (Van
De Maa en & Weinbe ge , 2012) di ec ly on use
answe s. This me hod maps use answe s om hei
o iginal non-nume ical domain in o a wo-dimensional
space ha can be easily isualized ( ind addi ional
de ails in he Appendix B). Figu e 9 shows he
wo-dimensional embeddings a e applying he -STE
algo i hm o he answe s o each online beha io al
expe imen . Each poin in he embedding ep esen s 1
o he 100 ma e ials om he Lagunas e al. da ase .
The inse s show he colo o each ma e ial based on
he colo classi ica ion p oposed by Lagunas e al. We
can obse e how ma e ials a e clus e ed by colo and,
i we ocus in a single colo , hey seem o be clus e ed
by e lec ance p ope ies (e.g., in Tes HH, ed colo
clus e , we can obse e how on he le he e a e specula
ma e ials while on he igh he e a e di use ma e ials).
This inding sugges s ha use s ha e ollowed a
wo-s ep s a egy o ecognize he ma e ials, and ha
he high-le el ac o s d i ing ma e ial ecogni ion migh
be colo i s , and he e lec ance p ope ies second. A
he end o he HIT, use s we e asked o w i e he main
isual ea u es hey used o ecognize ma e ials. Ou o
368 unique use s om he expe imen s analyzed in he
Resul s, 273 answe ed ha hey ha e used he colo s,
and 221 answe ed ha hey elied on he e lec ions.
Among hem, 157 answe ed bo h colo and e lec ions
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 9
as some o he isual cues hey ha e used o pe o m
he ask. This obse a ion, oge he wi h he -STE
isualiza ion, s eng hens he hypo hesis o a wo-s ep
s a egy.
Image s a is ics
P e ious s udies ocused on simple image s a is ics
as an a emp o u he unde s and ou isual sys em
(Adelson, 2008;Mo oyoshi e al., 2007). Ne e heless,
i is a gued whe he ou isual sys em ac ually
de i es any aspec s o ma e ial pe cep ion using such
simple s a is ics (Ande son & Kim, 2009;Kim &
Ande son, 2010;Olkkonen & B aina d, 2010). We
es ed he co ela ion be ween he i s ou s a is ical
momen s o he luminance (conside ed as he a io:
L=0.3086 ∗R+0.6094 ∗G+0.0820 ∗B), he pixel
in ensi y o each colo channel independen ly, and he
join RGB pixel in ensi y, di ec ly agains use s op i e
accu acy. To measu e co ela ion we employ a Pea son
Pand Spea man Sco ela ion es . We ound ou ha
he e is li le o no co ela ion, excep o he s anda d
de ia ion o he join RGB pixel in ensi y whe e
P2=0.43 (P<0.001) and S2=0.50 (P<0.001).
Addi ional in o ma ion can be ound in Appendix C.
Image his og ams
We also compu e he his og ams o he RGB pixel
in ensi y, o he luminance, o a Gaussian py amid (Lee
& Lee, 2016), o a Laplacian py amid (Bu & Adelson,
1983), and o log-Gabo il e s designed o simula e
he ecep i e ield o he simple cells o he P ima y
Visual Co ex (V1) (Fische e al., 2007). To see how
such his og ams would pe o m a classi ying ma e ials,
we ain a suppo ec o machine (SVM) ha akes he
image his og am as he inpu and classi ies he ma e ial
in ha image. We use a adial basis unc ion ke nel
(o Gaussian ke nel) in he SVM. We use all image
his og ams ha do no ea u e Ha an geome y as he
aining se and lea e he ones wi h Ha an as es se .
In he end, he bes pe o ming SVM uses he RGB
image his og am as he inpu and achie es a 24.17% op
i e accu acy in he es se .
In addi ion, we compa e he p edic ions o each
SVM di ec ly agains human answe s. Fo each
e e ence s imuli we compa e he i e selec ions o
he use agains he i e mos likely SVM ma e ial
p edic ions o ha s imuli. The bes SVM uses he
his og ams o V1-like subband il e s and ag ees wi h
humans 6.36% o he ime. Mo eo e , we compa e
his og am simila i ies agains human answe s using a
Χ2his og am dis ance (Pele & We man, 2010). Fo a
e e ence image s imuli we measu e i s simila i y agains
all possible candida e image s imuli and compa e he
closes i e agains pa icipan s answe s. The Gaussian
py amid his og am ob ained he bes esul , ag eeing
wi h humans 6.29% o he ime. These esul s show how
simple s a is ics, and highe -o de image his og ams
seem no o be capable o ully cap u ing human
beha io . We ha e added addi ional esul s on he
SVMs and human ag eemen in Appendix C.
Image equencies
To unde s and i humans’ pe o mance could be
explained by aking in o accoun he spa ial equency
o he e e ence s imuli, a hei iewed size, we ha e
added he HFC measu e, and he i s ou s a is ical
momen s o he e e ence s imuli magni ude spec um
o he ac o s analyzed in he Resul s. We ound ha he
Skewness (P<0.001) and Ku osis (P<0.001) o he
magni ude spec um seem o ha e a signi ican in luence
on humans pe o mance; howe e , hey p esen a e y
small e ec size.
Highly nonlinea models
Recen s udies sugges ha , o unde s and wha
su ounds us, ou isual sys em is doing an e icien
nonlinea encoding o he p oximal s imulus ( he
image inpu o ou isual sys em) and ha highly
nonlinea models migh be able o be e cap u e human
pe cep ion (Delanoy e al., 2020;Fleming&S o s,
2019). Inspi ed by his hypo hesis, we ha e ained a
deep neu al ne wo k called ResNe (He e al., 2016)
using a loss unc ion sui able o classi y he ma e ials
in he Lagunas e al. da ase . The images ea u e he
same illumina ions as he e e ence s imuli. We le
ou he images ende ed wi h Ha an geome ies o
alida ion and es ing pu poses, and use he es du ing
aining. To know which ma e ial he ne wo k classi ies
we add a so max laye a he end o he ne wo k. The
so max laye ou pu s he p obabili y o he inpu
image o belong o each ma e ial in he da ase . In
compa ison, he model used by Lagunas e al. does
no ha e he las ully connec ed and so max laye ,
and i is ained using a iple loss unc ion aiming o
simila i y ins ead o classi ica ion. A he end o he
aining, he model achie es a op-5 accu acy o 89.63%
on he es se , sugges ing ha such models a e ac ually
capable o ex ac ing meaning ul ea u es om labeled
p oximal image da a (addi ional de ails on he aining
can be ound in Appendix D). To gain in ui ion on how
he ne wo k has lea ned, we ha e used he Uni o m
Mani old App oxima ion and P ojec ion algo i hm
(McInnes & Healy, 2018). This algo i hm aims o
dec ease he dimensionali y o a se o ea u e ec o s
while main aining he global and local s uc u e o hei
o iginal mani old. Figu e 10 shows a wo-dimensional
isualiza ion o he es se ob ained using he 128
ea u es o he ully connec ed laye be o e so max. We
can obse e how ma e ials seem o be g ouped i s by
colo and hen by i s e lec ance p ope ies sugges ing
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 16
Lagunas, M., Ga ces, E., & Gu ie ez, D. (2018).
Lea ning icons appea ance simila i y. Mul imedia
Tools and Applica ions, 1–19.
Lagunas, M., Malpica, S., Se ano, A., Ga ces, E.,
Gu ie ez, D., & Masia, B. (2019). A Simila i y
Measu e o Ma e ial Appea ance. ACM
T ansac ions on G aphics (P oc. SIGGRAPH, 38(4).
Lee, S., & Lee, D. (2016). Fusion o IR and Visual
Images Based on Gaussian and Laplacian
Decomposi ion Using His og am Dis ibu ions
and Edge Selec ion. Ma hema ical P oblems in
Enginee ing, 2016.
Leloup, F. B., Poin e , M. R., Du é, P., & Hanselae ,
P. (2010). Geome y o illumina ion, luminance
con as , and gloss pe cep ion. JOSA, A27(9),
2046–2054.
Li, W., & F i z, M. (2012). Recognizing ma e ials
om i ual examples. In Eu opean Con e ence on
Compu e Vision (ECCV). Sp inge , 345–358.
Maloney, L. T., & B aina d, D. H. (2010). Colo and
ma e ial pe cep ion: Achie emen s and challenges.
Jou nal o Vision (JOV), 10(9), 19–19.
Mao,R.,Lagunas,M.,Masia,B.,&Gu ie ez,D.
(2019). The e ec o mo ion on he pe cep ion
o ma e ial appea ance. P oceedings o he ACM
symposium on applied pe cep ion (SAP). (p. 9).
ACM.
Ma low, P. J., Kim, J., & Ande son, B. L. (2012). The
pe cep ion and mispe cep ion o specula su ace
e lec ance. Cu en Biology, 22(20), 1909–1913.
Ma usik, W., P is e , H., B and, M., & McMillan, L.
(2003). A Da a-D i en Re lec ance Model. ACM
T ansac ions on G aphics (TOG), 22(3), 759–769.
McInnes, L., & Healy, J. (2018). Umap: Uni o m
mani old app oxima ion and p ojec ion
o dimension educ ion. a Xi p ep in
a Xi :1802.03426.
Mo oyoshi, I., Nishida, S., Sha an, L., & Adelson, E.
H. (2007). Image s a is ics and he pe cep ion o
su ace quali ies. Na u e, 447(7141), 206–209.
Nagai, T., Ma sushima, T., Koida, K., Tani, Y.,
Ki azaki, M., & Nakauchi, S. (2015). Tempo al
p ope ies o ma e ial ca ego iza ion and ma e ial
a ing: isual s non- isual ma e ial ea u es. Vision
Resea ch, 115, 259270. Pe cep ion o Ma e ial
P ope ies (Pa II).
Nishida, S., & Shinya, M. (1998). Use o image-based
in o ma ion in judgmen s o su ace- e lec ance
p ope ies. JOSA A15,12, 2951–2965.
Obein, G., Knoblauch, K., & Viéo , F. (2004).
Di e ence scaling o gloss: Nonlinea i y,
binocula i y, and cons ancy. Jou nal o Vision
(JOV), 4(9), 4–4.
Oli a, A., & To alba, A. (2001). Modeling he shape
o he scene: A holis ic ep esen a ion o he spa ial
en elope. In e na ional Jou nal o Compu e Vision
(IJCV), 42(3), 145–175.
Olkkonen, M., & B aina d, D. H. (2010). Pe cei ed
glossiness and ligh ness unde eal-wo ld
illumina ion. Jou nal o Vision (JOV), 10(9),
5–5.
Olkkonen, M., & Da id, H. B., (2011). Join e ec s
o illumina ion geome y and objec shape in he
pe cep ion o su ace e lec ance. i-Pe cep ion2,9,
1014–1034.
Palme , S. (1975). Visual pe cep ion and wo ld
knowledge: No es on a model o senso y-cogni i e
in e ac ion. Explo a ions in Cogni ion, 279–307.
Pele, O., & We man, M. (2010). The quad a ic-chi
his og am dis ance amily. In Eu opean con e ence
on compu e ision. Sp inge , 749–762.
Pizlo, Z. (2001). Pe cep ion iewed as an in e se
p oblem. Vision Resea ch41,24, 3145–3161.
Ramamoo hi, R., & Han ahan, P. (2001). An e icien
ep esen a ion o i adiance en i onmen maps. In
P oceedings o he Annual con e ence on Compu e
G aphics and In e ac i e Techniques. 497–500.
Rubins ein, M., Gu ie ez, D., So kine, O., & Shami , A.
(2010). A Compa a i e S udy o Image Re a ge ing.
ACM T ansac ions on G aphics (P oc. SIGGRAPH
Asia 2010), 29(6), 160:1–160:10.
Scha ali zky, F., & Zisse man, A. (2001). Viewpoin
in a ian ex u e ma ching and wide baseline
s e eo. In P oceedings o he IEEE In e na ional
Con e ence on Compu e Vision (ICCV),Vol. 2.
IEEE, 636–643.
Schille , P. H., Finlay, B. L., & Volman, S. F. (1976).
Quan i a i e s udies o single-cell p ope ies
in monkey s ia e co ex. I. Spa io empo al
o ganiza ion o ecep i e ields. Jou nal o
Neu ophysiology, 39(6), 1288–1319.
Schlü e , N., & Faul, F. (2019). Visual shape pe cep ion
in he case o anspa en objec s. Jou nal o Vision
(JOV), 19(4), 24–24.
Se ano, A., Gu ie ez, D., Myszkowski, K., Seidel, H.-
P., & Masia, B. (2016). An In ui i e Con ol Space
o Ma e ial Appea ance. ACM T ansac ions on
G aphics (TOG), 35(6), A icle 186 ( No .2016),
186:1–186:12 pages.
Sè e, R. (1993). P oblems connec ed wi h he concep
o gloss. Colo Resea ch & Applica ion, 18(4),
241–252.
Sha an, L., Rosenhol z, R., & Adelson, E. (2009).
Ma e ial pe cep ion: Wha can you see in a b ie
glance? Jou nal o Vision (JOV), 9(8), 784–
784.
Downloaded om jo .a ojou nals.o g on 03/15/2021

Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 17
Sha an, L., Rosenhol z, R., & Adelson, E. H. (2008).
Eye mo emen s o shape and ma e ial pe cep ion.
Jou nal o Vision (JOV), 8(6), 219–219.
Sun, T., Se ano, A., Gu ie ez, D., & Masia, B. (2017).
A ibu e-p ese ing gamu mapping o measu ed
BRDFs. In Compu e G aphics Fo um,Vol. 36.
Wiley Online Lib a y, 47–54.
Szegedy, C., Vanhoucke, V., Io e, S., Shlens, J.,
& Wojna, Z. (2015). Re hinking he incep ion
a chi ec u e o compu e ision. a Xi .
Thompson, W., Fleming, R., C eem-Regeh , S., &
S e anucci, J. K. (2011). Visual Pe cep ion om
a Compu e G aphics Pe spec i e (1s ed.).A.K.
Pe e s, L d., Na ick, MA, USA.
Thompson, W., Fleming, R., C eem-Regeh , S., &
S e anucci, J. K. (2016). Visual pe cep ion om a
compu e g aphics pe spec i e. AK Pe e s/CRC
P ess.
Van De Maa en, L., & Weinbe ge , K. (2012).
S ochas ic iple embedding. In Machine Lea ning
o Signal P ocessing (MLSP), 2012 IEEE
In e na ional Wo kshop on. IEEE, 16.
Vango p, P., Lau ijssen, J., & Du é, P. (2007). The
In luence o Shape on he Pe cep ion o Ma e ial
Re lec ance. ACM T ansac ions on G aphics
(TOG), 26(3), A icle 77 (July2007).
Ví a, R., & Filip, J. (2016). Minimal sampling o
e ec i e acquisi ion o aniso opic BRDFs. In
Compu e G aphics Fo um,Vol. 35. Wiley Online
Lib a y, 299–309.
Wang, T.-C., Zhu, J.-Y., Hi oaki, E., Chand ake ,
M., E os, A. A., & Ramamoo hi, R. (2016). A
4D ligh - ield da ase and CNN a chi ec u es o
ma e ial ecogni ion. In Eu opean Con e ence on
Compu e Vision. Sp inge , 121–138.
Welinde , P., B anson, S., Pe ona, P., & Belongie, S. J.
(2010). The mul idimensional wisdom o c owds. In
Ad ances in Neu al In o ma ion P ocessing Sys ems
(Neu IPS). 2424–2432.
Zhang, F., Ridde , H., & Pon , S. (2015). The in luence
o ligh ing on isual pe cep ion o ma e ial
quali ies. In Human Vision and Elec onic Imaging
XX, Vol. 9394. In e na ional Socie y o Op ics and
Pho onics, 93940Q.
Appendix A: Addi ional esul s on
he influence o ime
Addi ional de ails on he ime ha each pa icipan
spen doing he online beha io al expe imen . In
Figu e 16 we can see how he ime spen o answe
each ial becomes s able as he beha io al expe imen
ad ances.
In luence o e e ence illumina ion: The e e ence
illumina ion I e in luences he ime use s spend o
answe each ial (P=0.001). Use s spend mo e ime
when he s imuli a e li wi h Ennis illumina ion while
hey a e he as es when he illumina ion is Doge.
We did no ind a signi ican in luence o he
e e ence geome y G e o candida e geome y Gcand
in he a e age ime each use spen o answe each ial.
Fi s o de in e ac ions: We obse e ha use s ake
signi ican ly longe o answe he ials when bo h he
e e ence geome y and he candida e geome y change
(G e ∗Gcand,P=0.001). This happens in he case
whe e he e e ence geome y has mos ly low spa ial
equency con en and he candida e geome y changes
(Gcand =sphe e,P=0.002); and when he e e ence
has mos ly low spa ial equency (G e =sphe e,P=
0.001) and he candida e geome y changes.
Appendix B: Addi ional de ails on
he -STE algo i hm
The -STE algo i hm aims o ob ain an n-
dimensional embedding ha sa is ies as many
quali a i e compa isons o he ype “A is mo e simila
o B han C” as possible. In ou case, a wo-dimensional
embedding which is easie o isualize. Ne e heless, in
he use s udies, we ha e asked pa icipan s o selec i e
ma e ials om a pool o candida es and we do no ha e
such quali a i e compa isons. Howe e , we can assume
ha he selec ion o he use s will be close (mo e
simila ) o he e e ence han any o he ma e ial ha
was no selec ed. Based on his assump ion, we gene a e
iple s whe e he use selec ion is mo e simila o he
e e ence ma e ial han any o he andom ma e ial ha
is no wi hin he 5 selec ed ma e ials. We epea his
p ocess en imes o each o he 5 ma e ials selec ed by
Figu e 16. A e age ime he use s spen o each ial acco ding
o he o de o appea ance du ing he online beha io al
expe imen . We can obse e how, as he use p og esses
h ough he expe imen , he ime spen on each ial becomes
s able. The e o ba s co espond o a 95% confidence in e al.
Downloaded om jo .a ojou nals.o g on 03/15/2021
Jou nal o Vision (2021) 21(2):2, 1–18 Lagunas, Se ano, Gu ie ez, & Masia 18
he use making su e ha he new andomly sampled
ma e ial has no been andomly selec ed al eady no
ha i belongs o he pool o 5 selec ed ma e ials.
To un he -STE we se a lea ning a e o 1, and an
α=25 (deg ees o eedom o he S uden - ke nel).
Addi ionally, we apply a loga i hmic ans o ma ion o
he loss alue o he -STE. Those pa ame e s a e he
same o he answe s o he ou expe imen s.
Appendix C: Addi ional de ails on
image s a is ics
To measu e he co ela ion be ween image s a is ics
and use s pe o mance we employ a Pea son Pand
Spea man Sco ela ion es wi h a signi icance alue
(P- alue) o 0.05. The alue Pn ep esen s he Pea son
co ela ion o he n h s a is ical momen (same applies
o he Spea man Snco ela ion).
Luminance: We analyze i he momen s o he
luminance o each ma e ial image ha e a di ec
in luence on use s pe o mance. We ound ha he
momen s o he luminance a e no co ela ed wi h use s
pe o mance: P1=−0.14 (P=0.17),S1=−0.15
(P=0.15),P2=0.02 (P=0.83),S2=−0.03 (P=
0.78),P3=0.03 (P=0.77),S3=0.03 (P=0.78),
P4=0.01 (P=0.94),S4=0.05 (P=0.65).
RGB image: We analyze i he momen s o he
join RGB in ensi y o each ma e ial image ha e
a di ec in luence on use s pe o mance. We ound
ha he momen s o he join RGB in ensi y ha e
li le o no co ela ion wi h use s pe o mance
excep o he s anda d de ia ion: P1=−0.02
(P=0.79),S1=−0.06 (P=0.51),P2=0.43
(P<0.001),S2=0.50 (P<0.001),P3=0.16
(P=0.09),S3=0.22 (P=0.02),P4=−0.1
(P=0.30),S4=−0.06 (P=0.52).
We also es ed ou he co ela ion o each channel
and ound ou ha o all he channels he e is no
co ela ion o any o he i s 4 s a is ical momen s.
Red channel: On he ed channel he e seems o
be a sligh posi i e linea co ela ion be ween he
ou h momen (ku osis) and use s pe o mance. All
he o he s a is ics show no signi ican co ela ion:
P1=−0.10 (P=0.29), S1=−0.08 (P=0.42),
P2=0.03 (P=0.60), S2=−0.02 (P=0.87),
P3=0.07 (P=0.46), S3=0.07 (P=0.51), P4=0.20
(P=0.04), S4=0.15 (P=0.13).
G een channel: The e is no co ela ion be ween any
s a is ics on he g een channel: P1=−0.04 (P=0.66),
S1=−0.0.(P=0.74), P2=0.03 (P=0.55), S2=0.04
(P=0.67), P3=0.05 (P=0.64), S3=0.06 (P=0.53),
P4=0.05 (P=0.63), S4=0.01 (P=0.94).
Blue channel: Simila o he g een channel, he
blue does no show any co ela ion o he i s 4
s a is ical momen s: P1=0.03 (P=0.72), S1=−0.004
(P=0.93), P2=0.06 (P=0.52), S2=0.01 (P=0.95),
P3=0.13 (P=0.19), S3=0.10 (P=0.30), P4=0.16
(P=0.11), S4=−0.05 (P=0.61).
Addi ional esul s on he SVMs and his og am
simila i y
We ha e ained a o al o 6 SVM models, each
o hem using a di e en inpu : RGB pixel in ensi y,
luminance in ensi y, Gaussian py amid pixel in ensi y
(Lee & Lee, 2016), Laplacian py amid pixel in ensi y
(Bu & Adelson, 1983), joining he Gaussian
and Laplacian py amids, and using log-Gabo
il e s (Fische e al., 2007). Fo each o hem he
SVM achie ed a op-5 accu acy in he es se o :
24.17%, 15.16%, 22.50%, 6.33%, 7.52%, and 16.33%,
espec i ely. In addi ion, we ha e compa ed how he
SVM p edic ions ag eed wi h humans’ answe s om
he online beha io al expe imen s. Fo each SVM he
ag eemen is: 4.24%, 4.33%, 4.34%, 5.04%, 4.97%, and
6.36% espec i ely. Las , we ha e also compu ed he
his og am simila i y using a Χ2dis ance. Then, we
ha e aken he i e closes samples and compa ed ha
wi h human answe s. We do ha o each o he i e
di e en his og ams and each achie es an ag eemen
o : 5.95%, 5.45%, 6.29%, 4.97%, 5.04%, and 5.07%
espec i ely.
Appendix D: Addi ional de ails on
ResNe aining
To ain he 35 laye s ResNe (34 o he o iginal
model plus an addi ional ully connec ed) (He e al.,
2016) we ha e employed he da ase in oduced by
Lagunas e al. (Lagunas e al., 2019), which con ains
ende ings o ma e ials wi h di e en illumina ions and
geome ies. We keep he images ende ed wi h Ha an-3
geome y o alida ion pu poses and Ha an geome y
o es ing. All he o he images a e used o aining.
To ain he model o classi y ma e ials we use a so
c oss-en opy loss whe e samples ha do no belong
o he same class a e penalized (Szegedy e al., 2015).
The loss unc ion akes he p obabili ies ou pu o
he so max laye and penalizes when hey gi e a high
p obabili y o he ma e ials ha do no belong o he
inpu image. The images inpu o he model a e esized
o 224 ×224 pixels. The pa ame e s o he model a e
ini ialized using a p e ained e sion on ImageNe
da ase (Deng e al., 2009). We use he ADAM
algo i hm (Kingma & Ba, 2014) as he op imize . The
model has been ained du ing 50 i e a ions s a ing a
alea ning a eo 10
−3and decayed by a ac o o 10 a
he i e a ion 20, 35, and 45; he ba ch-size was se o
64 images. We use he PyTo ch amewo k and use an
N idia 2080Ti GPU.
Downloaded om jo .a ojou nals.o g on 03/15/2021