scieee Open visual document viewer

A Robust Audio Fingerprinter Based on Pitch Class Histograms Applications for Ethnic Music Archives

Six, Joren; Cornelis, Olmo

Abstract

In this paper we present a new acoustic fingerprinting system, based on pitch class histograms. The aim of acoustic fingerprinting is to generate a small representation of an audio signal that can be used to identify identical, or recognize similar, audio snippets in a large audio set. A robust fingerprinting system generates similar fingerprints for perceptually similar audio signals. A piece of music with a noise added should generate an almost identical fingerprint as the original. The new system, presented here, has some interesting features which makes it a valuable tool to manage ethnic music archives: the fingerprints are rather robust against pitch shift, tempo changes, several synthetic audio effects, and reversal of the audio. When only part of the audio is used to generate a fingerprint, the system keeps working but retrieval performance degrades.

Full text

A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es 191 A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es JOREN SIX OLMO CORNELIS Royal Academy o Fine A s & Royal Conse a o y, Uni e si y College Ghen jo en.si[email p o ec ed] olmo.co ne[email p o ec ed] Abs ac In his pape we p esen a new acous ic inge p in ing sys em, based on pi ch class his og ams. The aim o acous ic inge p in ing is o gene a e a small ep esen a ion o an audio signal ha can be used o iden i y iden ical, o ecognize simila , audio snippe s in a la ge audio se . A obus inge p in ing sys em gene a es simila inge p in s o pe cep ually simila audio signals. A piece o music wi h a noise added should gene a e an almos iden ical inge p in as he o iginal. The new sys em, p esen ed he e, has some in e es ing ea u es which makes i a aluable ool o manage e hnic music a chi es: he inge p in s a e a he obus agains pi ch shi , empo changes, se e al syn he ic audio e ec s, and e e sal o he audio. When only pa o he audio is used o gene a e a inge p in , he sys em keeps wo king bu e ie al pe o mance deg ades. 1. In oduc ion In he p ocess o digi izing a la ge music collec ion i is possible ha he same music is p esen on di e en physical media, ei he as comple e copies o as copies o indi idual acks. Some imes i is ha d o keep ack o which physical media a e al eady digi ized and which a e s ill o p ocess. An abili y o sea ch o music based on he con en o he signal is a aluable ool o p e en duplica es en e ing he digi al e sion o he music a chi e. He e we p esen a sys em wi h hose capabili ies. Ano he use case o such sys em is o ( e)connec me a-da a o an audio agmen wi hou any in o ma ion, bu is p esen in he digi al connec ion. Fo la ge, his o ical collec ions o e hnic music he p oblems ske ched abo e a e almos ine i able. O en, indi idual collec ions o eco dings o discs –wi hou me a-da a– a e dona ed o museums. These collec ions usually a e e y di e se and se e al eco dings may al eady be p esen in he a chi e, which is whe e he need o con en based sea ch comes o play. Due o he na u e o he o iginal physical media - wax cylinde s, wi e eco dings, magne ic apes, g amophone eco ds - and he, o en abysmal, eco ding quali y a con en based sea ch sys em o e hnic music has o ha e special ea u es o obus ness. The goal o he sys em, as i is p esen ed he e, is o iden i y iden ical audio exce p s e en i hey we e p ocessed o digi ized in a di e en way. Ou esea ch is ocused on pi ch class his og ams which appea o be obus enough o he ask o acous ic inge p in ing, e en in he con ex o his o ical e hnic music collec ions. This pape is s uc u ed as ollows: we s a wi h an o e iew o he sys em and hen a gue why i shows po en ial. Then de ails abou he implemen a ion a e un eiled. The hi d sec ion desc ibes an expe imen wi h he sys em. The pape ends wi h a conclusion. 2. Sys em O e iew Figu e 1 shows a gene al acous ical inge p in ing sys em. Fea u es a e ex ac ed om audio and wi h hese ea u es a inge p in is cons uc ed. The inge p in is a small ep esen a ion o he audio. In he bes case, pe cep ually simila audio gene a es ela ed inge p in s, iden ical audio should gene a e iden ical inge p in s. Wi h he gene a ed inge p in and a lis o p e iously gene a ed inge p in s an unknown piece o audio is iden i ied. In an ideal sys em he inge p in s a e small bu unique o each piece o audio and sea ching h ough a la ge numbe o inge p in s is e icien . Al e na i e sys ems include he ones desc ibed by Hai sma & Kalke  Visi h p:// a sos.0110.be/ o mo e in o ma ion. A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es 192 (2002); Wang (2003); Allamanche (2001), he e is also a e iew a icle on audio inge p in ing by Cano e al. (2005). Figu e 1: A gene al audio inge p in e . The wo k low o ou sys em, which can be seen in Figu e 2, is exac ly he same as he gene al acous ic inge p in ing sys em bu i shows which ea u es a e ex ac ed and how a inge p in is c ea ed in ou sys em. The i s s ep is o ex ac ea u es om audio, in his case a pi ch ex ac ion algo i hm ex ac s pi ch om audio. The nex s ep is o c ea e a inge p in , he e o e we use a pi ch class his og am. A pi ch class his og am con ains how many imes any pi ch class has been anno a ed in a musical segmen o piece. A pi ch class is de ined he e as an in ege be ween 0 and 1200, o co espond wi h he cen uni in oduced by Helmhol z & Ellis (1912). I o example, he alue o 880Hz has been de ec ed, his equency in Hz can be con e ed o a cen alue c ela i e o a e e ence equency calcula ing c = 1200 × log2( / ). Wi h he s anda d =8:176Hz 1 his makes 8100 mod 1200=900 cen s. Fo his block o audio, one alue is added o he bin ep esen ing 900 cen s in he pi ch class his og am. I he nex block o audio 2 con ains o example 220Hz, he exac same hing happens. This is done o e and o e again o he en i e piece. Please no e ha his app oach comple ely igno es empo al in o ma ion, he nex sec ion explains he ad an age o doing his. In heo y also imb al in o ma ion is no included in a pi ch class his og am, in p ac ice i has an in luence. When he exac same melody is pe o med on piano and hen on lu e and subsequen ly bo h a e analysed, hey will gene a e sligh ly di e en pi ch class his og ams due o he cha ac e is ics o impe ec pi ch de ec ion algo i hms. E.g. some pi ch de ec ion algo i hms migh con use o e ones o undamen al equencies. The hi d s ep is o ma ch he cons uc ed inge p in wi h a lis o p e iously s o ed inge p in s. In ou sys em his en ails calcula ing a simila i y be ween pi ch class his og ams. Pi ch class his og ams a e essen ially p obabili y densi y unc ions, hey desc ibe how p obable i is a block o audio has a ce ain pi ch. The e a e di e en ways o calcula e simila i y be ween p obabili y densi y unc ions, o an o e iew, consul he a icle by Cha (2007). As he inal s ep in he p ocess, he iden i ied piece o audio is e u ned. 2.1 Pi ch Class His og ams as Acous ic Finge p in s The e has been a lo o esea ch abou pi ch class his og ams, o e y simila concep s unde some imes di e en names e.g. by Sundbe g & Tje nlund (1969); Moelan s e al. (2009); Gedik & Bozku (2010); Six & Co nelis (2011); Tzane akis e al. (2002), o name a ew. Especially he las a icle is in e es ing, in he u u e wo k sec ion o hey men ion he ollowing: Al hough mainly designed o gen e classi ica ion i is possible ha ea u es de i ed om Pi ch His og ams migh also be applicable o he p oblem o con en -based audio iden i ica ion o audio inge p in ing ( o an example o such a sys em see Allamanche [2001]). We a e planning o explo e his possibili y in he u u e. 1 The MIDI no e numbe s anda d le s no e numbe 0 co espond wi h a e e ence equency o 8:176Hz, which is C-1 wi h A4 uned o 440Hz. I he same e e ence equency is used o cen s, hen MIDI no e numbe s and cen s di e by a ac o 100. 2 The op imal size o a block o audio depends on he chosen pi ch de ec ion algo i hm. A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es 193 As a as we know Tzane akis e al. did no explo e he possibili y o using pi ch his og ams o audio inge p in ing any u he , his a icle can be seen as an elabo a ion on ha idea. Bo h Figu e 3 and Figu e 4 show why his app oach is easonable. The igu es show pi ch class his og ams o simila bu no equal e sions o an A ican song. A pen a onic scale appea s o be p esen . Figu e 3 illus a es ha pi ch class his og ams a e ela i ely obus agains se e e adap a ions o he unde lying audio: he his og am shape emains mo e o less he same. Figu e 4 shows he esul o audio e ec s which change he pi ch. Changing pi ch in audio shi s he his og am o e he ho izon al pi ch axis. When calcula ing a co ela ion be ween his og ams his needs o be aken in o accoun . Figu e 2: An acous ic inge p in ing scheme based on pi ch ea u es and pi ch class his og ams as inge p in s. Da a is p ocessed in an iden ical ashion as he gene al audio inge p in e in Figu e 1. A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es 194 Figu e 3: A pi ch class his og am o an A ican song. The his og am o he o iginal song is p esen , oge he wi h a his og am o a e e sed, a c opped and noisy ende ing o he song. I shows ha pi ch class his og ams a e ela i ely obus agains se e e mu ila ions o he unde lying audio. Figu e 4: A pi ch class his og am o an A ican song oge he wi h a his og am o a e sion played 5% as e and a pi ch shi ed e sion (wi hou a ec ing he du a ion). I is clea ha almos he same his og am is p esen h ee imes, only shi ed sligh ly o e he ho izon al pi ch axis. Figu e 4 also shows why igno ing empo al in o ma ion can be a good idea. Changing he playback speed o a song –wi h co esponding pi ch shi – esul s only in a ho izon al shi o he his og am, as can be seen in he illus a ion. In he con ex o analogue media his means ha magne ic ape digi ized on an inco ec speed can be ma ched wi h he same con en digi ized on he co ec speed. Then his og am o e lap o in e sec ion is used as a dis ance measu e because Gedik & Bozku (2010) show ha his measu e wo ks bes o pi ch class his og am e ie al asks. The o e lap c(h1, h2) be ween wo his og ams h1 and h2 wi h K classes is calcula ed wi h equa ion 1. To calcula e he co ela ion wi h a pi ch shi n equa ion 2 is used. To make su e ha he bin k emains wi hin he bounds o he his og am a mod K calcula ion is done. In ou applica ion his means ha he oc a e ela ion is espec ed, e.g. wi h n equal o 50 cen 3, he bin a 1170 cen 3 , he bin a 1170 cen o h1 is compa ed wi h he bin a (1170+50) mod 1200=20 cen o h2. Table 1: Simila i y be ween di e en pi ch class his og ams o se e al adap ed e sions o a song. I shows ha he his og am o he song wi h whi e noise added di e s he mos om he o iginal his og am (89%). To ind he pi ch shi n wi h maximum co ela ion, an exhaus i e sea ch is done by simply calcula ing he co ela ion o each possible shi . A possible signi ican pe o mance inc ease would be o de ec peaks on each his og am and hen compa e he his og ams on only hose posi ions (shi s), his is simila o ’ onic de ec ion’ in Gedik & Bozku (2010). 3 Hal a semi one, no o be con used wi h he Ame ican appe . A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es 195 Table 1 shows he co ela ion, as de ined by equa ion 2, be ween he di e en his og ams shown in Figu e 3 and 4, wi h op imal pi ch shi n. I shows ha he his og am based on he o iginal e sion is, o his song, e y much alike his og am based on he e e sed audio (96%). The e sion wi h added noise di e s he mos om he o iginal (89%). C opping one minu e om he song, which is 7 minu es and 20 seconds long, esul s in co ela ion o 94%. The 97% simila i y be ween he 5% as e and pi ch shi ed e sion can be explained by he ac ha a 5% speed inc ease ansla es o a pi ch shi o 84 cen s which is almos 100 cen s 4 . The only di e ence hen is he leng h o he song, i.e. he numbe o elemen s in he his og am, which can be no malized. Sec ion 3 shows i he desc ibed beha iou is unique o his one song o no . The implemen a ion o he sys em is done in Ja a and uses he pi ch es ima o desc ibed in McLeod (2009). Fo es ing pu poses, a pla o m independen e sion can be downloaded he e h p:// a sos.0110.be/ ag/FMA2012. The e you can also ind sc ip s and da a used o his pape . 3. Expe imen al Resul s To show ha a inge p in ing scheme based on pi ch class his og ams has po en ial, an expe imen was done on a da a se o 10272 songs om Cen al A ica (see appendix A o mo e in o on he da a se ). The expe imen was cons uc ed as ollows: om he da a se 50 andomly selec ed iles we e copied. A numbe o modi ica ions and e ec s –27 in o al– we e applied o hese 50 iles 5 , gene a ing 1350 modi ied songs. The goal o he expe imen was co ec ly ma ch hose 1350 songs o he o iginal in he da a se . C opping was done a he beginning o he ile, he a e age leng h o he 50 selec ed iles was abou 4 minu es, he sho es was one minu e in leng h. Table 2 shows he esul s o he expe imen . F om hose esul s some conclusions can be d awn. 1) Since he e ie al o he o iginal song always succeeded, i s ands o eason ha inge p in s o songs a e, a leas , unique wi hin his da a se . An impo an p ope y o a inge p in . 2) The e e sed audio is also e ie ed always, which shows ha he pi ch es ima o used gene a es almos iden ical es ima ions on e e sed audio. This is a good sani y check when using au oco ela ion based pi ch es ima o s. When using pi ch de ec o s based on ea -models his migh be less i ial. 3) Pi ch shi ing wo ks easonably well. 4) The pe o mance when lea ing ou he i s numbe o seconds deg ades quickly be ween 15 and 20 seconds. 5) The me hod does no handle whi e noise ha well. The 20%, 25% and 30% whi e noise e sions we e le ou o he able since no ma ches we e ound. 6) The da a se con ains monophonic and polyphonic music. The esul s show ha pi ch es ima o s which gene a e one pi ch pe block o audio migh su ice o his ask, e en wi h polyphonic music. 4 Since 2(84/1200) = 1.05 a shi o 84 cen s ansla es o a shi in equency (Hz) o i e pe cen . 5 SoX - Sound Exchange, a command line u ili y, was used o apply e ec s o he o iginal ile. Following command line ins uc ions we e used: pi ch, speed, e e se, im, and syn h whi enoise. Fo mo e in o ma ion on SoX, and he exac meaning o he e ec s, see h p://sox.s .ne . A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es 196 Table 2: The esul s o a e ie al ask on a da a se o 10272 iles. 27 e ec s we e applied o 50 songs, gene a ing 1350 modi ied e sions. The goal o he ask was o ind he o iginal e sion o he song. The pe cen ages show much o he modi ied e sions we e co ec ly iden i ied in he i s , i s wo, and i s h ee hi s. The o iginal and e e sed e sion a e e ie ed always. The pe o mance on pi ch shi and c opping is easonable, whi e noise and la ge empo changes a e p oblema ic. 4. Conclusion & Fu u e Wo k In his pape a new app oach o acous ic inge p in ing, based on pi ch class his og ams, was p esen ed. A e he in oduc ion, which ske ched he applica ions o he sys em, an o e iew o he wo king p inciples o acous ic inge p in ing in gene al and ou sys em in pa icula we e gi en. The second sec ion also explained why pi ch class his og ams can be used as inge p in s. Some de ails abou he implemen a ion a e also gi en. In sec ion h ee expe imen al e alua ion was done. This pape has shown ha an acous ic inge p in ing sys em based on pi ch class his og ams is a he obus and has po en ial bu a lo o ques ions emain open. The expe imen in his pape only discusses a e ie al ask o comple e songs and o a limi ed numbe o audio e ec s. Some u u e wo k includes: 1. Expand he e ie al ask o include mo e audio (Wes e n music) and apply mo e audio e ec s: echo, digi al analogue/analogue digi al con e sions, low bi a e encoding, band pass il e ing …Tes ing o he obus ness agains pi ch ins abili y, o en obse ed in old ape eco dings and es ing wi h ealis ic en i onmen al noise om a noise da abase - e.g. a noisy audience. 2. Documen he pe o mance dec ease o he sys em be e by using s anda d in o ma ion e ie al measu es (P ecision, Recall, ROC-cu es …). E.g. o do ailu e analysis when adding mo e and mo e noise. 3. In es iga e wha happens when pe cussi e songs - wi hou much pi ch in o ma ion - a e inge p in ed. 4. Compa ing his sys em wi h simila sys ems on he same da a se , using he same measu es. A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es 197 5. See i he sys em can be applied o iden i y small agmen s o music ins ead o comple e songs. How small is he minimum agmen ? Which adap a ions need o be done o b oadcas moni o ing, p ocessing s eams? 6. Expe imen wi h pi ch es ima o s o ch oma es ima ion algo i hms. I he pi ch es ima o is eplaced, is he e a signi ican impac on he esul s? 7. Handle scalabili y and pe o mance issues. Can he inge p in size be educed, wi hou loss o accu acy? Is i possible o speed up he ma ching s ep signi ican ly? The sys em desc ibed he e shows simila i ies wi h some co e song de ec ion sys ems, i is e y simila o he one by Se à & Gómez (2008). This is ema kable because he goal o bo h sys ems is di e en . He e we wan o iden i y he same audio, wi h some modi ica ions and in he o he sys em he goal is o iden i y in a ian musical ma e ial (co e songs), using simila ea u es. The simila i ies o he sys ems make sense i you look a iden ical audio, wi h some modi ica ions, as ’ he mos simila co e song’. This s a emen esul s in new ques ions: can co e song de ec ion sys ems be used o iden i y almos iden ical audio, o acous ic inge p in ing? O he in e se: how well does he inge p in ing sys em desc ibed he e pe o m a co e song de ec ion? Cu en ly, hose ques ions a e le as u u e wo k. The da a se es ed in sec ion 3 does no include di e en e sions - co e s - o he same song. As a inal ema k, we would like o no e ha his a icle is a he unique because i p esen s a gene ally applicable algo i hm ha is es ed on e hnic music i s . Only la e i will be applied o wes e n music. This is pa ly due o he ac ha we only ha e access o a la ge da a se wi h A ican music bu is also a philosophical s a emen : ins ead o adap ing echniques used on Wes e n music o applica ions wi h e hnic music, why no , o once, do i he o he way a ound? Re e ences Allamanche, E. (2001). Con en -based iden i ica ion o audio ma e ial using mpeg-7 low le el desc ip ion. In P oceedings o he 2nd in e na ional symposium on music in o ma ion e ie al (ISMIR 2001). Cano, P., Ba lle, E., Kalke , T., & Hai sma, J. (2005). A e iew o audio inge p in ing. The Jou nal o VLSI Signal P ocessing, 41, 271-284. Cha, S.-h. (2007). Comp ehensi e su ey on dis ance / simila i y measu es be ween p obabili y densi y unc ions. In e na ional Jou nal o Ma hema ical Models and Me hods in Applied Sciences, 1(4), 300–307. Gedik, A. C. & Bozku , B. (2010). Pi ch- equency his og am-based music in o ma ion e ie al o u kish music. Signal P ocessing, 90(4), 1049–1063. Hai sma, J. & Kalke , T. (2002). A highly obus audio inge p in ing sys em. In P oceedings o he 3 h In e na ional Symposium on Music In o ma ion Re ie al (ISMIR 2002). Helmhol z, H. on & Ellis, A. J. (1912). On he sensa ions o one as a physiological basis o he heo y o music ( ansla ed and expanded by Alexande J. Ellis, 2nd English d .) [Book]. Longmans, G een, London. McLeod, P. (2009). Fas , accu a e pi ch de ec ion ools o music analysis. Academisch p oe sch i , Uni e si y o O ago. Depa men o Compu e Science. Moelan s, D., Co nelis, O., & Leman, M. (2009). Explo ing a ican one scales. In P oceedings o he 10 h In e na ional Symposium on Music In o ma ion Re ie al (ISMIR 2009). Se à, J. & Gómez, E. (2008, 31/03/2008). Audio co e song iden i ica ion based on onal sequence alignmen . In Ieee in e na ional con e ence on acous ics, speech and signal p ocessing (icassp) (p. 61-64). Las Vegas, USA. A ailable om iles/publica ions/jse a ICASSP08.pd Six, J. & Co nelis, O. (2011). Ta sos - a Pla o m o Explo e Pi ch Scales in Non-Wes e n and Wes e n Music. In P oceedings o he 12 h In e na ional Symposium on Music In o ma ion Re ie al (ISMIR 2011). Sundbe g, J. & Tje nlund, P. (1969). Compu e measu emen s o he one scale in pe o med music by means o equency his og ams. STL-QPS, 10(2-3), 33-35. A Robus Audio Finge p in e Based on Pi ch Class His og ams Applica ions o E hnic Music A chi es 198 Tzane akis, G., E molinskyi, A., & Cook, P. (2002). Pi ch his og ams in audio and symbolic music in o ma ion e ie al. In P oceedings o he 3 h In e na ional Symposium on Music In o ma ion Re ie al (ISMIR 2002) (pp. 31–38). Wang, A. L. (2003). An Indus ial-S eng h Audio Sea ch Algo i hm. In P oceedings o he 4 h In e na ional Symposium on Music In o ma ion Re ie al (ISMIR 2003) (pp. 7–13). A Audio Ma e ial In sec oin 3 a subse o he music collec ion o he Royal Museum o Cen al A ica (RMCA, Te u en, Belgium) was used. The museum ocuses on he A ican cul u e and easu es all kinds o e hnog aphic objec s. The a chi e o he Depa men o E hnomusicology has a digi ized collec ion o abou 50.000 sound eco dings, wi h a o al o 3000 hou s o music, mos ly ield eco dings made in Cen al A ica o which he oldes da ing back o 1910. The audio a chi e is one o he bigges and bes documen ed 6 a chi es wo ldwide o he egion o Cen al A ica. On speci ic song, also om he RMCA collec ion, was used in sec ion 2.1. I has ape numbe MR.1954.1.18-4 and was eco ded in 1954 by missiona y Scohy-S ooban s in Bu undi. 6 The e is a websi e abou he audio da ase o he Royal Museum o Cen al A ica ea u ing comple e me a- da a and some audio agmen s. I can be ound a h p://music.a icamuseum.be