Untangling Attribution in Biodiversity Data Records
Abstract
The exactness, fitness-for-purpose (FFP) and reliability of primary biodiversity data can be enhanced by additional data beyond the basic taxon-location-date triad (Hill et al. 2010). Often, the only available data are the labels in legacy specimen collections. The digitization process is most efficient if all available information can be collected at once in a single event of specimen handling, rather than in separate phases. There is, however, a compromise between producing a faster catalogue for immediate use and housekeeping, and an accurate, wider-FFP database where all data have been thoroughly checked.Recognizing potential sources of error at digitization time may help making choices. During a dataset integration procedure, the quality and reliability of the data capture was analyzed. The dataset consisted of transcribed label data of over 58K pinned insects of agriculturally-relevant groups in XXth-century collections at six institutions in Spain, that resulted in almost 6000 collector strings. But collector names could beunidentified;misread;ambiguous,duplicated under variants; ormisplaced or misattributed to/from another entity, e.g. a location.This resulted in a high entropy level where one collector could be databased in multiple ways, artificially inflating the corresponding catalogues. The entropy was much higher in collections where collectors contributed few specimens, which is the case for university-based collections.By using simple indexing and cross-referencing techniques, the roster of names was significantly reduced, but full disambiguation of collectors required mining ancillary sources and consulting with people with long-standing knowledge of the collections. Overall, 42% of collector names were in error, resulting in excess entropy. METHODSSpecimen data were collated from the TETTRIS INC-STEP Project of Spanish pollinators complemented by some agriculturally-relevant groups in six collections deposited at five academic and research institutions in Madrid, Barcelona, Valencia and Pamplona (see Suppl. material 1 for full details). MCNB, MNCN, MUVHN were chiefly historical and created by researchers over long careers, while MZNA-R, MZNA-Z, UCME came mainly from academic coursework activities.Our procedures agreed with the overall strategy of Groom et al. (2022), while devising a specific workflow. Disambiguation and attribution included a number of steps (Suppl. material 1). First, names were normalized to facilitate grouping and sorting (7 steps). Then, ambiguous names were (whenever possible) attributed to actual persons by internal checks (clustering of names, matching localities and dates), consultation with external references (e.g. student lists), or consultation with collection curators (5 steps). Names were finally given a four-level identity qualification resulting from the disambiguation exercise (see Suppl. material 1). For further analyses, "very low" and "low" levels were considered a poor attribution, while "medium" and "full" levels were deemed good attribution.RESULTSThe disambiguation exercise reduced collector names from 5939 to 3473 (a 41.5% decrease). The total entropy, measured as Shannon's H', decreased by 8% while Simpson's dominance D increased by 29%, as expected (merging names created larger collections for some collectors). However, there were differences among collections. The three academic collections had both higher diversity and higher diversity reduction than the historical collections (Fig. 1). Historical collections tended to be much more concentrated (30 specimens per collector, 25% single-specimen collectors) than the academic collections (8 and 61%, respectively).These differences can be tracked to how the collections were formed. Historical collections tended to be well documented and created by researchers spanning longer careers, while academic collections were often linked to works created during coursework. Fig. 2 shows the career spans of collectors. The abrupt end of collection at the turn of the century for academic collections can be linked to the introduction of restrictive legislation in Spain (a mandatory reduction of course loads, and unaffordable permission requirements). Full disambiguation often requires manual, time-consuming verifications against a variety of external sources. Automated disambiguation procedures elevated good identification from an initial 21% to 47% in MZNA-R, but it was not possible to use manual checking against sources for this dataset. In contrast, the similar MZNA-Z collection could be checked against coursework lists, which resulted in a 80% good attribution rate (Fig. 3). Overall, 57% of the collector names across collections had a good attribution.CONCLUSIONDoing disambiguation exercises by people with contemporary knowledge of the collections appears to be critical to avoid losing much information from the collections. Curators have an invaluable knowledge of the collections, could locate relevant documents about them, and are the ones able to reduce collection entropy by at least 32%. Their early retirement would negate a significant way to disambiguate names that left little or no trace in the literature record.