NLC: A Measu e Based on P ojec ions
Robe o Ruiz, Jos´e C. Riquelme, and Jes´us S. Aguila -Ruiz
Depa amen o de Lenguajes y Sis emas,
Uni e sidad de Se illa
A da. Reina Me cedes S/N.
41012 Se illa, Espa˜na
{ uiz, iquelme,aguila }@lsi.us.es
Abs ac . In his pape , we p opose a new ea u e selec ion c i e ion.
I is based on he p ojec ions o da a se elemen s on o each a ibu e.
The main ad an ages a e i s speed and simplici y in he e alua ion o
he a ibu es. The measu e allows ea u es o be so ed in ascending
o de o impo ance in he defini ion o he class. In o de o es he
ele ance o he new ea u e selec ion measu e, we compa e he esul s
induced by se e al classifie s be o e and a e applying he ea u e selec-
ion algo i hms.
1 In oduc ion
The selec ion o ele an ea u es is a cen al p oblem in machine lea ning. I
a ele an ea u e is emo ed, he measu e o he emaining ea u es will de-
e io a e. In o de o iden i y ele an a ibu es, we need o add ess wha a
good ea u e is o classifica ion. Wi hou defining he goodness o a ea u e o
ea u es, i does no make sense o alk abou bes o op imal ea u es. The
algo i hms e alua e he a ibu es based on gene al cha ac e is ics o he da a.
Fea u e Selec ion can be iewed as a sea ch p oblem, whe e each s a e in he
sea ch space specifies a subse o he possible ea u es. The need o e alua ion
is common o all sea ch s a egies.
In his pape , we p opose a new ea u e selec ion c i e ion no based on calcu-
la ed measu es be ween a ibu es, o complex and cos ly dis ance calcula ions.
This c i e ion is based on a unique alue called NLC. I ela es each a ibu e
wi h he label used o classifica ion. This alue is calcula ed by p ojec ing da a
se elemen s on o he espec i e axis o he a ibu e (o de ing he examples
by his a ibu e), hen c ossing he axis om he beginning o he g ea es
a ibu e alue, and coun ing he Numbe o Label Changes (NLC) p oduced.
This wo k has been suppo ed by he Spanish Resea ch Agency CICYT unde g an
TIC2001-1143-C03-02, and Jun a Andalucia unde coo dina ed ac ion ACC-1021-
TIC-2002.
V. Maˇ ´ık e al. (Eds.): DEXA 2003, LNCS 2736, pp. 907–916, 2003.
c
Sp inge -Ve lag Be lin Heidelbe g 2003
908 R. Ruiz, J.C. Riquelme, and J.S. Aguila -Ruiz
2 Rela ed Wo k
Fea u e selec ion algo i hms use diffe en e alua ion unc ions. Func ions a e
based in c i e ions o measu e he ele ance o he a ibu es. The e a e se e al
axonomies o hese e alua ion measu es in p e ious wo k, depending on diffe -
en c i e ions: Langley [9] g oup e alua ion unc ions in o wo ca ego ies: fil e
and w appe . Blum y Langley [3] p o ide a classifica ion o e alua ion unc ions
in o ou g oups, depending on he ela ion be ween he selec ion and he induc-
ion p ocess: embedded, fil e , w appe , weigh . Ano he diffe en classifica ion,
Doak [5] and Dash [4] p o ide a classifica ion o e alua ion measu e based on
hei gene al cha ac e is ics mo e hen in he ela ion wi h he induc ion p ocess.
The classifica ion ealized by Dash, sepa a e fi e diffe en ypes o measu es: dis-
ance, in o ma ion, dependence, consis ency y accu acy. Fea u e Selec ion can
be iewed as a sea ch p oblem, whe e each s a e in he sea ch space specifies a
subse o he possible ea u es. The need o e alua ion is common o all sea ch
s a egies. In gene al, a ibu e selec ion algo i hms pe o m a sea ch h ough he
space o ea u e subse s, and mus add ess ou basic issues affec ing he na u e
o he sea ch: 1) S a ing poin : o wa d and backwa d, acco ding o whe he i
began wi h no ea u es o wi h all ea u es. 2) Sea ch o ganiza ion: exhaus i e
o heu is ic sea ch. 3) E alua ion s a egy: w appe o fil e . 4) S opping c i e-
ion: a ea u e selec o mus decide when o s op sea ching h ough he space
o ea u e subse s. A p edefined numbe o ea u es a e selec ed, a p edefined
numbe o i e a ions eached. Whe he o no he addi ion o dele ion o any
ea u e p oduces a be e subse , we also s op he sea ch, i an op imal subse
acco ding o some e alua ion unc ion is ob ained.
3 Fea u e E alua ion
3.1 Obse a ions
To disco e main idea o he algo i hm we base on he da a se s IRIS and WINE,
because o he easy in e p e a ion o hei wo-dimensional p ojec ions.
In Figu e 1(a) i is possible o obse e ha i he p ojec ion o he examples
is made on he o dina e axis we can no ob ain in e als whe e any class is a
majo i y. Ne e heless, o he Pe alwid h a ibu e i is possible o app ecia e
some in e als whe e he class is unique: [0,0.6] o Se osa, [1.0,1.3] o Ve sicolo
and [1.8,2.5] o Vi ginica. This is because when p ojec ing he examples on his
a ibu e he numbe o label changes is minimum. Fo example, i is possible o
e i y ha o Pe alwid h he fi s label change akes place o alue 1 (se osa o
Ve sicolo ), he second in 1.3 (Ve sicolo o Vi ginica). The e a e o he changes
la e and he las one is in 1.8.
In Figu e 1(b) he same conclusion is eached wi h da a se WINE. We
analyze he p ojec ion o da a se elemen s on o C8 and C7 a ibu es. We
iden i y in e als whe e one class is a majo i y when c ossing he abscissas axis
om he beginning o he g ea es a ibu e alue: [0,1] o class 3, [1.5,2.3] o
NLC: A Measu e Based on P ojec ions 909
0
0,5
1
1,5
2
2,5
3
3,5
4
4,5
5
0 0,5 1 1,5 2 2,5 3
pe alwid h
sepalwid h
se osa e sicolo i ginica
(a)
0
0,1
0,2
0,3
0,4
0,5
0,6
0,7
0123456
C7
C8
class 1 class 2 class 3
(b)
Fig. 1. (a) IRIS. Rep esen a ion o A ibu es Sepalwid h-Pe alwid h (b) WINE. Rep-
esen a ion o A ibu es C8-C7
class 2 and [3,4] o class 1. Ne e heless, o C8 a ibu e on he o dina e axis
i is no possible o obse e any in e als whe e he class is unique.
We conclude ha i will be easie classi y by a ibu es wi h he smalles
numbe o label changes. I he a ibu es a e in ascending o de acco ding o
he NLC, we ob ain a anking lis wi h he be e a ibu es om he poin
o iew o he classifica ion, in I is his would be: Pe alwid h 16, Pe alleng h
19, Sepallen h 87 and Sepalwid h 120. This esul ag ees wi h wha is common
knowledge in da a mining, which s a es ha he wid h and leng h o pe als a e
mo e impo an han hose ela ed o sepals.
Classi ying IRIS wi h C4.5 by Sepalwid h only, we ob ain 59% accu acy and
by Pe alwi d h 95%. The a ibu es used in Figu e 1(b) a e he fi s and he las
on he anked lis , wi h a NLC alue o 43 and 139 espec i ely. Applying he
classifie C4.5, we ob ain 80.34% accu acy by C7 and 47.75% by C8.
3.2 Defini ions
Defini ion 1: An example e∈E is a uple o med by he Ca esian p oduc
o he alue se s o each a ibu e and he se C o labels. We define he
ope a ions a and lab o access he a ibu e and i s label (o class): a : E x
N→A and lab: E →C, whe e N is he se o na u al numbe s.
Defini ion 2: Le he uni e se U be a sequence o example om E. We will
say ha a da abase wi h n examples, each o hem wi h m a ibu es and
one class, o ms a pa icula uni e se. Then U=<u[1],...,u[n]>and as he
da abase is a sequence, he access o an example is achie ed by means o i s
posi ion. Likewise, he access o j- h a ibu e o he i- h example is made
by a (u[i],j), and o iden i ying i s label lab(u[i]).
Defini ion 3: An o de ed p ojec ed sequence is a sequence o med by he p o-
jec ion o he uni e se on o he i- h a ibu e. This sequence is so ed ou in
ascending o de .
Defini ion 4: A pa i ion in cons an subsequences is he se o subsequences
o med om he o de ed p ojec ed sequence o an a ibu e in such a way
as o main ain he p ojec ion o de . All he examples belonging o a sub-
sequence ha e he same class and e e y wo consecu i e subsequences a e
disjoin ed wi h espec o he class.
910 R. Ruiz, J.C. Riquelme, and J.S. Aguila -Ruiz
(a) (b)
7
5
212
10 6
O3 9
E48
O1 11
E 0-E
b
Oa
O
E
E8
7
5
10
6
9
3
4
2
O1
OEO
O
E
O
E
Fig. 2. Da a se wi h (a) en and (b) wel e elemen s and wo classes
Defini ion 5: Asubsequence o he same alue is he sequence composed o
he examples wi h iden ical alue om he i- h a ibu e wi hin he o de ed
p ojec ed sequence. This si ua ion can be o igina ed in con inuous a iables,
and i will be he way o deal wi h he disc e e a iables.
Defini ion 6: wo examples a e inconsis en i hey ma ch excep o he class
label.
3.3 Desc ip ion
The algo i hm is based on his basic p inciple: o coun he label changes o
examples p ojec ed on o each ea u e. I he a ibu es a e in ascending o de
acco ding o he NLC we will ha e a lis ha defines he p io i y o selec ion,
om g ea e o smalle impo ance.
Be o e o mally exposing he algo i hm, we will explain in mo e de ail he
main idea. Le us conside he si ua ion depic ed in Figu e 2(a), wi h en el-
emen s numbe ed and wo labels (O-odd numbe s and E-e en numbe s): he
p ojec ion o he examples on he abscissas axis p oduces h ee cons an subse-
quences {O,E,O}co esponding o he examples {[1,3,5][8,4,10,2,6][7,9]}. Iden i-
cally, wi h he p ojec ion on he o dina es axis we can ob ain six cons an subse-
quences {O,E,O,E,O,E} o med by he examples {[1][2,4][3,9][6,10][5,7][8]}.We
check ha he fi s a ibu e has wo label changes and he second one has fi e.
Applying ou hypo hesis, he fi s a ibu e is mo e ele an han he second
one, because i has a smalle NLC.
4 Algo i hm
The algo i hm is e y simple and as (Table 1). I has he capaci y o ope a e
wi h con inuous and disc e e a iables as well as wi h da abases which ha e
wo classes o mul iple classes. Fo each a ibu e, he aining-se is o de ed
(QuickSo [7], his algo i hm is O(n log n), on a e age and we coun he NLC
h oughou he o de ed p ojec ed sequence.
NLC: A Measu e Based on P ojec ions 911
Table 1. Main Algo i hm
Inpu : E aining (N examples, M a ibu es)
Ou pu : E educed (N examples, K a ibu es)
o each a ibu e Ai∈1..M
QuickSo (E,i)
NLCi←Numbe Changes(E,i)
NLC A ibu e Ranking
Selec he k i s
Table 2. Numbe Changes unc ion
Inpu : E aining (N examples, M a ibu es), i
Ou pu : numbe o label changes
o each example ej∈E wi h j in 1..N
i a (u[j],i) ∈subsequence o he same alue
changes = changes + ChangesSameValue()
else
i lab(u[j]) <> las Label)
changes = changes + 1
e u n(changes)
Applying he algo i hm o he example o he Figu e 2(b) we ob ain he
o de ed p ojec ed sequences:
{1,3,4,10,2,11,8,9,6,12,5,7}
{1,11,4,8,3,9,10,6,2,12,5,7}
and he pa i ions:
{[1,3][4,10,2][11,8,9,6,12,5,7]}
{[1,11][4,8][3,9][10,6,2,12][5,7]}
The elemen s’ p ojec ions on o he fi s a ibu e p oduce wo cons an sub-
sequences and one subsequence o he same alue wi h diffe en labels. The
elemen s’ p ojec ions on o he second a ibu e p oduces fi e cons an subse-
quences.
Numbe Changes conside s whe he we deal wi h diffe en alues om an a -
ibu e, o wi h a subsequence o he same alue ( his si ua ion can be o igina ed
in con inuous and disc e e a iables). In he fi s case, i compa es he p esen
label wi h he las one. Whe eas in he second case, whe e he subsequence is
o he same alue, i coun s he maximum possible changes by means o he
unc ion ChangesSameValue.
In he a ibu e ep esen ed on he o dina es axis (b) in Figu e 2(b), we
see se e al subsequences o he same alue wi h he same label, hen, we deal
wi h cons an subsequence, and he esul is ou label changes (NLC=4). In
he a ibu e on he abscissas axis (a), he fi s wo pa i ions a e cons an
912 R. Ruiz, J.C. Riquelme, and J.S. Aguila -Ruiz
subsequences, and he hi d is a subsequence o he same alue wi h wo labels.
The e o e, we conside he maximum possible NLC.
In he p e ious case, we ha e he subsequence [11,8,9,6,12,5,7] whe e ou
elemen s a e class O (odd) y h ee class E (e en). The e a e wo easons o
coun ing he maximum NLC: fi s , we wan o penalize he a ibu e in hese
inconsis ency si ua ions; and second, we wan o a oid ambigui ies ha could
be p oduced depending on he elemen s o de a e algo i hm QuicSo is ap-
plied. Fo example, in diffe en independen execu ions, we could ob ain hese
si ua ions: [E,E,E,O,O,O,O], [E,E,O,O,O,O,E], [O,O,E,E,O,O,E],. . . wi h NLC
equal o 1, 2 and 3 espec i ely. ChangesSameValue e u ns 5, he maximum.
The si ua ion is: [O,E,O,E,O,E,O]. This can be ob ained wi h low cos . I can be
deduced coun ing he class’ elemen s in he subsequence wi hou eso ing he
elemen s.
We conclude ha he a ibu e bwi h ou NLC is mo e ele an ha he
a ibu e awi h se en NLC.
5 Expe imen s
In his sec ion we compa e he quali y o selec ed a ibu es by he NCL mea-
su e wi h he selec ed a ibu es by he o he wo me hods: In o ma ion Gain
(IG) [10] and he Relie F me hod [8]. IG has been chosen because i is he mo e
popula concep and i is used mo e when you wan o e alua e he ele ance
o an a ibu e. And he Relie F me hod has been chosen because i is widely
e e enced in o he pape s. The Relie F me hod is a e sion o he Relie me hod
by Kononenko, wich pe mi s a ibu es wi h missing alues and mul iclass p ob-
lems. The quali y o each selec ed a ibu e was es ed by means o h ee clas-
sifie s: he Nai e Bayes [6], C4.5 [10] and 1-NN [1]. The implemen a ion o he
induc ion algo i hms and he o he s selec o s was done using he Weka lib a y1
and he compa ison was pe o med wi h eigh een da abases o he Uni e si y
om Cali o nia I ine [2]. The da a se s we e chosen wi h ew missing alues.
The p ocess ollowed o es he quali y o he a ibu es selec ed wi h he
NCL measu e was he ollowing. Fo all o iginal da a se s, we ob ained he
accu acy using he h ee classifie s and he size o he decision ees induced by
C4.5. We ob ained he same measu es a e applying each selec o algo i hm,
eco ding he numbe o a ibu es selec ed.
To asses he ob ained esul s, wo pai ed s a is ical es s wi h a confidence
le el o 95% we e ealized
In o de o es ablish he numbe o a ibu es in each case, we ob ain a anked
lis o ea u es wi h he h ee me hod and we use he lea ning cu e o obse e
he effec o added ea u es. S a ing wi h one ea u e ( he mos ele an one
fi s ) and g adually adding nex mos ele an ea u e one by one, we calcula e
i s accu acy a e. We selec he se o a ibu es wi h he bes accu acy. Applying
a diffe en classifie , we ob ain a diffe en se .
1h p://www.cs.waika o.ac.nz/ ml
NLC: A Measu e Based on P ojec ions 913
DB.da a
N={0,1,...9}
DB_N.da a DB_METHOD_N.da a
N={0,1,...9} METHOD +
Lea ning
Cu e
DB_N. es DB_METHOD_N. es
Same educ ion
DB. es
DB.da a 10 se s
10 se s
N={0,1,...9}
(a)
CLASSIFIER
METHOD={NLC,RLF,IG}
CLASSIFIER={C4.5,1NN,NB} N={0,1,...9}
N={0,1,...9}
(b) (c)
Fig. 3. C oss- alida ion p ocess
DB_N.da a
N={0,1,...9}
METHOD
METHOD={NLC,RLF,IG}
Ranking
Fea u e Lea ning
Cu e DB_METHOD_N.da a
Fig. 4. Reduc ion me hod: ea u e anking and lea ning cu e
Fo each da abase (DB), he measu es we e es ima ed aking he mean o
a en old c oss alida ion. A en- old c oss- alida ion is pe o med by di iding
he da a in o en blocks o cases ha ha e an app oxima ely simila size, and
o each block in u n, es ing he model cons uc ed om he emaining nine
blocks on he unseen cases in he hold-ou block (Figu e 3(a)). The same olds
we e used o each algo i hm aining-se s.
Each educing me hod was gi en a aining se (DB N.da a) consis ing o
90% o he a ailable da a, om which i e u ned a subse DB METHOD N
(Figu e 3(b)), whe e METHOD is one o NLC, RLF, IG, N is a alue in 0,1,...,9
and includes a classifie o ob ain he lea ning cu e (Figu e 4). We use he same
classifie ha we a e going o classi y he es se . Fo example, om DB 1.da a
we would ob ain DB IG 1.da a by applying he IG me hod. The emaining 10%
o he unseen da a (DB N. es ) was also educed (DB METHOD N. es ) (Fig-
u e 3(b)) and es ed on he ins ances o DB METHOD N.da a using a clas-
sifie (Figu e 3(c)). Fo example, we ob ain i is ig 1.da a by applying he IG
me hod o he i is 1.da a file gene a ed by he c oss alida ion. A e wa ds, we
use i is ig 1.da a o classi y i is ig 1. es by means o he nea es neighbo ech-
nique. When we deal wi h he lea ning cu e, we also apply 1NN (Figu e 4).
As a u he compa ison, ano he widely-used lea ne , C4.5 and NB, was
un on hese da a se s (Figu e 3(c)). Fo example, a e educing i is 3.da a
wi h NLC, i is nlc 3.da a was gene a ed, i was gi en as inpu o C4.5 and he
decision ee gene a ed was used o classi y he i is nlc 3. es file ( es files a e
educed oo).
I we conside all he possible esul s ha we ge using he o iginal da a, he
h ee selec ion me hods (NLC, RLF and IG) and he h ee classifie s (C4.5, 1NN
and NB) wi h eigh een da a se s aking he mean o a 10- old c oss alida ion,
914 R. Ruiz, J.C. Riquelme, and J.S. Aguila -Ruiz
Table 3. Accu acy ob ained wi h C4.5,1NN and nai e Bayes, selec ing he subse
wi h he bes accu acy
Da a C4.5 1NN NB
C4.5 NLC 123RLF IG 1NN NLC 123RLF IG NB NLC 123RLF IG
anne 98.6 98.4 98.4 98.0 99.3 99.0 98.9 98.9 86.3 89.2 ◦•90.0 92.4
bala 78.4 78.4 78.4 78.4 86.9 86.9 86.9 86.9 88.8 88.8 88.8 88.8
ge m 71.1 73.8 ◦◦ 70.5 74.5 72.4 69.7 70.8 71.0 74.8 75.5 73.4 74.7
diab 76.7 75.1 75.9 75.5 70.9 68.2 66.7 68.5 76.2 75.6 75.8 76.5
glas 69.2 70.1 71.0 67.3 70.5 71.9 •75.1 74.2 45.8 56.6 ◦◦ 47.7 51.0
gla2 77.8 79.6 79.2 79.0 79.2 78.3 •• 89.0 89.0 61.9 69.9 ◦◦ 64.4 69.9
h-s 78.5 83.7 ◦80.4 85.2 74.4 79.3 ◦•79.6 83.3 84.1 81.9 •85.2 84.4
iono 88.6 89.7 •92.6 91.2 86.6 87.2 •91.4 89.2 83.2 84.6 89.7 87.2
i is 94.0 92.7 94.0 92.7 95.3 93.3 93.3 93.3 95.3 92.7 •94.0 92.7
k - 99.5 99.3 99.5 99.5 96.5 97.4 ◦◦98.3 96.9 88.0 90.4 ◦• 94.0 90.4
lymp 78.3 74.9 77.7 76.3 79.7 77.0 83.0 77.8 83.8 83.8 81.8 83.8
segm 97.0 96.9 96.8 96.9 97.0 97.1 97.1 97.1 80.0 87.2 ◦87.2 87.2
sona 72.6 73.5 75.0 72.2 85.5 84.6 81.3 87.5 67.8 73.5 74.9 73.5
spli-2 94.4 94.0 94.0 94.2 73.9 90.0 ◦90.0 90.0 95.3 96.1 96.3 95.8
ehi 73.5 72.6 72.7 73.7 70.1 70.7 71.2 69.3 44.3 44.3 •48.5 44.3
owe 82.9 81.7 82.7 81.0 99.4 99.2 99.0 99.2 66.1 68.7 ◦69.2 69.1
wa e 76.6 77.6 78.1 77.4 73.8 79.1 ◦78.2 79.1 80.0 80.7 •81.4 80.7
zoo 93.1 94.1 93.1 92.1 96.0 97.0 95.0 95.0 95.0 95.0 91.0 92.1
we ge wo hund ed and eigh een esul s ((1+3) ×3×18 = 216). Now we a e
going o analyze hese esul s o ob ain some conclusions abou he pe o mance
o he diffe en me hods.
Table 3 shows a summa y o he esul s o he classifica ion using C4.5,
1NN and NB. Table shows how o en each me hod pe o ms significan ly be e
(deno ed by ◦) o wo se (deno ed by •) han da a wi hou educ ion (column
named 1), and be e o wo se han Relie F (RLF) and In o ma ion Gain (IG)
(column 2 and 3 espec i ely). Then, we ob ain fi y ou esul s compa ing NLC
wi h da a wi hou educ ion (18 da a se s ×3 classifie s = 54) and one hund ed
and eigh compa ing NLC wi h RLF and IG (18 da a se s ×3 classifie s ×2
= 108). NLC measu e is be e han da a wi hou educ ion in wel e o he
fi y ou cases, and in o y one a e equal and only in one is wo se han da a
wi hou educ ion. Fu he mo e, in ou o he one hund ed and eigh cases, he
se o a ibu es selec ed by he NLC measu e yields be e accu acy han he
wo o he me hods. In nine y h ee hey a e equal, and in ele en hey a e wo se
han he o he .
We selec he se o a ibu es wi h he bes accu acy. Applying each classi-
fie , we ob ain a diffe en se . The e o e, we ge h ee se o a ibu es o each
educ ion me hod o each da a se . We ob ain he pe cen age o he o iginal
ea u es e ained and we calcula e he a e age o e he eigh een da a se s (nine
esul s). All he me hods a e be ween 50% and 60% o he o iginal ea u es.
The expe imen s show ha by applying NLC, he knowledge a ained in
he o iginal aining file is conse ed in o he educed aining file, and he
NLC: A Measu e Based on P ojec ions 915
Table 4. Accu acy ob ained wi h C4.5,1NN and nai e Bayes, selec ing he h ee fi s
a ibu es o he anked lis
Da a C4.5 1NN NB
C4.5 NLC 12RLF IG 1NN NLC 12RLF IG NB NLC 12RLF IG
anneal 98.6 89.9 •92.5 90.6 99.3 91.1 91.6 91.6 86.3 84.0 •• 90.0 88.9
balance 78.4 69.4 69.4 69.4 86.9 68.0 69.6 67.7 88.8 74.2 73.6 73.3
gc edi 71.1 70.9 71.5 72.2 72.4 61.2 •• 70.6 70.6 74.8 71.6 71.3 73.9
diabe es 76.7 74.6 74.7 75.1 70.9 69.8 66.7 69.3 76.2 77.1 76.4 76.8
glass 69.2 71.9 67.7 65.4 70.5 66.3 66.3 64.5 45.8 53.8 ◦45.8 49.1
glass2 77.8 81.4 77.9 81.4 79.2 77.9 83.4 77.9 61.9 68.7 63.8 68.7
hea -s 78.5 72.6 •73.3 85.2 74.4 67.8 •• 73.3 84.8 84.1 74.8 •74.8 79.6
ionosphe 88.6 80.4 •• 88.3 90.0 86.6 84.3 85.7 88.6 83.2 78.3 •83.5 86.6
i is 94.0 94.0 94.0 94.0 95.3 95.3 95.3 95.3 95.3 95.3 95.3 95.3
k - s 99.5 90.4 90.4 90.4 96.5 90.4 90.4 90.4 88.0 90.4 90.4 90.4
lymph 78.3 77.7 79.7 77.7 79.7 75.7 83.7 75.0 83.8 75.0 80.3 72.3
segmen 97.0 90.5 ◦◦ 85.6 85.5 97.0 91.8 ◦◦ 85.3 88.5 80.0 77.0 ◦◦ 72.5 64.3
sona 72.6 68.8 70.2 70.7 85.5 73.1 66.8 70.2 67.8 73.1 70.6 70.6
splice-2 94.4 80.5 81.4 80.8 73.9 80.2 81.2 80.6 95.3 79.6 81.1 80.7
ehicle 73.5 53.5 •• 62.4 61.4 70.1 53.7 56.0 57.7 44.3 40.4 42.8 40.8
owel 82.9 69.1 70.6 72.1 99.4 79.6 80.1 82.8 66.1 57.9 56.0 58.7
wa e o m 76.6 66.2 64.5 65.3 73.8 57.0 56.3 56.4 80.0 65.1 66.1 64.9
zoo 93.1 84.2 ◦72.3 85.2 96.0 83.2 ◦71.3 87.2 95.0 84.2 ◦71.3 84.2
dimensionali y o da a is educed significan ly. We ob ain simila esul s wi h
he o he me hod, bu needing much mo e ime.
I is e y in e es ing o compa e he speed o a ibu e selec ion echniques.
We measu ed he ime aken in milliseconds o selec he anking o a ibu es.
NLC is an algo i hm wi h a e y sho compu a ion ime. NLC akes 792 mil-
liseconds in educing 18 da a se s whe eas Relie F akes 566 seconds and IG 2189
milliseconds. We ob ain he pe cen age o educ ion ime o each da a se s, and
we calcula e he a e age. NCL educe he compu a ional cos o he 99% o he
ime needed by Relie F and 50% o he ime needed by IG.
In o de o compa e he fi s a ibu es in he anking lis o each me hod,
we ob ain Table 4 whe e he da a se s a e educed o he fi s h ee a ibu es o
each anking lis . we obse e he accu acy o each educ ion me hod applying
he h ee classifie s. We ob ain simila esul s wi h he h ee me hods. Table
shows how o en each me hod pe o ms significan ly be e o wo se han Relie F
(RLF) and In o ma ion Gain (IG) (column 1 and 2 espec i ely). In en o he
one hund ed and eigh cases, he se o a ibu es selec ed by he NLC measu e
yields be e accu acy han he wo o he me hods. In eigh y hey a e equal, and
in ou een hey a e wo se han he o he .
6 Conclusions
In his pape we p esen a de e minis ic a ibu e selec ion c i e ion. The main
ad an ages a e i s speed and simplici y in he e alua ion o he a ibu es. The