scieee Open visual document viewer

A Comparison of PSO and GA Approaches for Gene Selection and Classification of Microarray Data

García Nieto, José Manuel; Alba, Enrique; Jourdan, Laetitia; Talbi, El-Ghazali

Full text

ACompa ison o PSO and GA App oaches o Gene Selec ion and Classi ica ion o Mic oa ay Da a José Ga cía-Nie o, En ique Alba Dep . de Leng. y Ciencias de la Compu ación Uni e si y o Málaga ETSI In o má ica, Málaga - 29071, Spain {jnie o,ea }@lcc.uma.es Lae i ia Jou dan, El-Ghazali Talbi LIFL-INRIA Fu u s B^ a M3, Ci é Scien i ique 59655 Villeneu e d’Ascq, F ance {jou dan, albi}@li l. Fea u e selec ion o gene exp ession analysis in cance p edic ion o en uses w appe classi ica ion me hods o dis- c imina e a ype o umo , o educe he numbe o genes o in es iga e in case o a new pa ien . By c ea ing clus- e s a big educ ion o he numbe o conside ed genes and an imp o emen o he classi ica ion accu acy can be inally achie ed. The de ini ion o he ea u e selec ion p oblem is his: gi en a se o ea u es F={ 1, ..., i, ..., n}, ind a subse F0⊆F ha maximizes a sco ing unc ion Θ : Γ →G such ha F0=a gmaxG⊂Γ{Θ(G)},(1) whe e Γ is he space o all possible ea u e subse s o F and Ga subse o Γ. The op imal ea u e selec ion p oblem has been shown o be NP-ha d. The e o e, only heu is ics app oaches a e able o deal wi h la ge size p oblems. In his wo k, we a e in e es ed in gene selec ion and classi- ica ion o DNA Mic oa ay da a in o de o dis inguish u- mo samples om no mal ones. Fo his pu pose, we p opose wo hyb id models ha use me aheu is ics and classi ica ion echniques. The i s one consis s o a Pa icle Swa m Op i- miza ion (PSO) combined wi h a SVM app oach as w appe me hod. The second model is based on he popula GA us- ing a specialized SSOCF [1] c osso e ope a o , ha will be also combined wi h SVM in ou app oach. A second impo an con ibu ion consis s in he ac ual disco e y o new and challenging esul s on six public da ase s iden i- ying signi ican in he de elopmen o a a ie y o cance s (leukemia, b eas , colon, o a ian, p os a e, and lung om he URL h p://sdmc.li .o g.sg/GEDa ase s/Da ase s. h ml). Fo ou PSOSV M app oach, a bina y e sion o PSO was implemen ed in C++ ollowing he skele on a chi ec u e o he MALLBA lib a y [2]. Fo he GASV M app oach he GA was implemen ed in C++ using he Pa adisEO F ame- wo k. Since he posi ion o a pa icle (ch omosome in GA) x ep esen s a gene subse , he e alua ion is ca ied ou by means o he SVM classi ie o assess he quali y o he ep- esen ed gene subse . The i ness o a pa icle/ch omosome xis calcula ed applying a Lea e One Ou C oss Valida ion (LOOCV) me hod o calcula e he a e o co ec classi i- ca ion (accu acy) o a SVM ained wi h his gene subse . The comple e i ness unc ion is desc ibed in Equa ion 2. i ness(x) = α·(100/accu acy) + β·# ea u es, (2) Copy igh is held by he au ho /owne (s). GECCO ’07, July 7-11, 2007, London, England, Uni ed Kingdom. ACM 978-1-59593-697-4/07/0007.. whe e αand βa e weigh alues se o 0.75 and 0.25 e- spec i ely. The objec i e he e consis s o maximizing he accu acy and minimizing he numbe o genes (# ea u es). Fo con enience (only minimiza ion o i ness) he i s ac- o is p esen ed as (100/accu acy). An special ini ializa ion me hod was adap ed o gene se- lec ion as ollows. The swa m/popula ion was di ided in o ou subse s o pa icles/ch omosomes ini ialized in di e en ways depending on he numbe o ea u es in each pa icle. Tha is, 10% o pa icles we e ini ialized wi h N(p e ixed alue) selec ed genes (1s) loca ed andomly. Ano he 20% o pa icles we e ini ialized wi h 2Ngenes, 30% wi h 3N genes and inally, he es o pa icles (40%) we e ini ialized andomly and 50% o he genes we e u ned on. Table 1: Subse s epo ed wi h 100% es accu acy Da ase Algo i hm Genes Leukemia PSOSV M 100(3) K01383 a , U03056 a , J04130 s a B eas PSOSV M 100(4) Con ig49744 RC, Con ig26884 RC Con ig25936 RC, Con ig13846 RC Colon PSOSV M 100(3) H64398, H73758 U27699 Lung GASV M 100(3) 33762 a , 34648 a 728 a , 829 s a O a ian GASV M 100(2) MZ1154.6306 MZ2653.8464 P os a e GASV M 100(3) 35935 a , 39801 a 40069 a In conclusion, bo h app oaches we e expe imen ally as- sessed on six well-known cance da ase s disco e ing new and challenging esul s, and iden i ying speci ic genes ha ou wo k sugges s as signi ican ones. In his sense, com- pa isons wi h se e al s a e o a me hods show compe i i e esul s acco ding o s anda d e alua ion. Resul s o 100% classi ica ion a e and ew genes pe subse (2, 3 and 4) a e ob ained in mos o ou execu ions (see Table 1). The use o an adap ed ini ializa ion me hod has shown a g ea in luence on he pe o mance o p oposed algo i hms, since i in o- duces an ea ly se o accep able solu ions in hei e olu ion p ocess. 1. REFERENCES [1] L. Jou dan, C. Dhaenens, and E.-G. Talbi. A gene ic algo i hm o ea u e selec ion in da a-mining o gene ics. In P oceedings o he 4 h Me aheu is ics In e na ional Con e encePo o (MIC’2001), pages 29–34, Po o, Po ugal, 2001. [2] E. Alba and M. g oup. Mallba: A Lib a y o Skele ons o Combina o ial Op imisa ion. In B. Monien and R. Feldmann, edi o s, P oceedings o he Eu o-Pa , olume LNCS 2400, pages 927–932, 2002. 427