scieee AI-readable full text Open interactive document viewer

How to Cite Datasets and Link to Publications

Ball, Alex; Duke, Monica

Abstract

The motivation to cite datasets arises from a recognition that data generated in the course of research are just as valuable to the ongoing academic discourse as papers and monographs. Scientific journals have traditionally supported research by disseminating knowledge in such detail that first, peer scientists could judge the strength of the conclusions based on the quality of the premises and research methods employed, and second, further investigations could be based upon it. In many disciplines, though, the paper alone is no longer sufficient for these purposes: the underlying data also need to be shared. As a medium, the journal paper owes its success in part to the control systems put in place around it: mechanisms allowing authors to be open about their research while still receiving due credit; metrics used to translate such attributions into rewards for authors and their institutions; and archives ensuring that the work is permanently available for reference and reuse. If datasets are to be regarded as first-class records of research, as they need to be, a similar set of control systems needs to be constructed around them. A major part of this work can be achieved using a robust citation mechanism for referencing datasets from within traditional publications. Provided the citation contains the name of a responsible agent, it can be used to assign due credit. By providing a globally unique identifier, it can be used to track the impact of a particular dataset. A citation is also an ideal place to provide the information needed to locate and access the dataset. In this way, datasets can take advantage of the infrastructure already in place to manage journal papers. The rise of electronic journals has led to new and valuable services being layered over the top of papers, among them the provision of forward links to papers citing the current one. Such links help the reader to gauge the impact of the paper, place it within the literature and in some cases gain awareness of flaws or issues discovered by others. Forward links from datasets to the papers that cite them provide all the same benefits, as well as ensuring that documentation for the dataset can be found. Ultimately, bibliographic links between datasets and papers are a necessary step if the culture of the scientific and research community as a whole is to shift towards data sharing, increasing the rapidity and transparency with which science advances.

Full text

A Digital Curation Centre ‘working level’ guide How to Cite Datasets and Link to Publications Alex Ball (DCC) and Monica Duke (DCC) Please cite as: Ball, A., & Duke, M. (2015). ‘How to Cite Datasets and Link to Publications’. DCC How-to Guides. Edinburgh: Digital Curation Centre. Available online: http://www.dcc.ac.uk/resources/how-guides Digital Curation Centre, 2015. Licensed under Creative Commons Attribution 4.0 International: http://creativecommons.org/licenses/by/4.0/ How to Cite Datasets and Link to Publications Introduction This guide will help you create links between your academic publications and the underlying datasets, so that anyone viewing the publication will be able to locate the dataset and vice versa. It provides a working knowledge of the issues and challenges involved, and of how current approaches seek to address them. This guide should interest researchers and principal investigators working on data-led research, as well as the data repositories with which they work. Why cite datasets and link them to publications? The motivation to cite datasets1arises from a recognition that data generated in the course of research are just as valuable to the ongoing academic discourse as papers and monographs. Scientific journals have traditionally supported research by disseminating knowledge in such detail that first, peer scientists could judge the strength of the conclusions based on the quality of the premises and research methods employed, and second, further investigations could be based upon it. In many disciplines, though, the paper alone is no longer sufficient for these purposes: the underlying data also need to be shared.2,3,4 As a medium, the journal paper owes its success in part to the control systems put in place around it: 1The term ‘dataset’ is used throughout this guide to mean a logically complete set of data; some systems or services prefer the terms ‘data product’ or ‘data package’. 2Stodden, V. (2009). Enabling reproducible research: Open licensing for scientific innovation. International Journal of Communications Law and Policy,13, 1–25. Retrieved 2 September 2010, from http: //www.ijclp.net/files/ijclp_web-doc_1-13-2009.pdf. 3Open to all? Case studies of openness in research. (2010, September). Research Information Network and National Endowment for Science, Technology and the Arts. Retrieved 1 May 2011, from http://www.rin.ac.uk/system/files/attachments/NESTA -RIN_Open_Science_V01_0.pdf. 4Lynch, C. (2009). Jim Gray’s fourth paradigm and the construction of the scientific record. In T. Hey, S. Tansley & K. Tolle (Eds.), The fourth paradigm: Data-intensive scientific discovery (pp. 177–183). Redmond, WA: Microsoft Research. Retrieved 14 July 2010, from http://research.microsoft.com/en-us/collaboration/ fourthparadigm/. mechanisms allowing authors to be open about their research while still receiving due credit; metrics used to translate such attributions into rewards for authors and their institutions; and archives ensuring that the work is permanently available for reference and reuse.5 If datasets are to be regarded as first-class records of research, as they need to be, a similar set of control systems needs to be constructed around them. A major part of this work can be achieved using a robust citation mechanism for referencing datasets from within traditional publications. Provided the citation contains the name of a responsible agent, it can be used to assign due credit. By providing a globally unique identifier, it can be used to track the impact of a particular dataset. A citation is also an ideal place to provide the information needed to locate and access the dataset. In this way, datasets can take advantage of the infrastructure already in place to manage journal papers. The rise of electronic journals has led to new and valuable services being layered over the top of papers, among them the provision of forward links to papers citing the current one. Such links help the reader to gauge the impact of the paper, place it within the literature and in some cases gain awareness of flaws or issues discovered by others. Forward links from datasets to the papers that cite them provide all the same benefits, as well as ensuring that documentation for the dataset can be found. Ultimately, bibliographic links between datasets 5Mackenzie Owen, J. (2007). The scientific article in the age of digitization (ch. 2). Information Science and Knowledge Management. Dordrecht: Springer. doi:10.1007/1-4020-5340-1. 2 and papers are a necessary step if the culture of the scientific and research community as a whole is to shift towards data sharing, increasing the rapidity and transparency with which science advances. Principles of data citation The FORCE11 Data Citation Synthesis Group – whose members include representatives of the Research Data Alliance, the ICSU World Data System, and a range of projects6– has published a set of data citation principles.7The principles build on earlier work in this area, most notably by CODATA,8the (US) National Academies of Sciences, Engineering, and Medicine,9 the DCC,10 and the Institute for Quantitative Social Science, Harvard University,11 and have been widely endorsed.12 Importance Data should be considered legitimate, citable products of research. Data citations should be accorded the same importance in the scholarly record as citations of other research objects, such as publications. Credit and Attribution Data citations should facilitate giving scholarly credit and normative and legal attribution to all contributors to the data, recognizing that a single style or mechanism of attribution may not be applicable to all data. Evidence In scholarly literature, whenever and wherever a claim 6FORCE11 Data Citation Synthesis Group, URL:https://www. force11.org/datacitation/workinggroup. 7FORCE11, Data Citation Synthesis Group. (2014). Joint declaration of data citation principles. Retrieved from https: //www.force11.org/datacitation. 8CODATA/ITSCI Task Force on Data Citation. (2013). Out of cite, out of mind: The current state of practice, policy and technology for data citation. Data Science Journal,12, CIDCR1–CIDCR75. doi:10.2481/dsj.OSOM13-043. 9Uhlir, P. F. (Ed.). (2012). For attribution – Developing data attribution and citation practices and standards: Summary of an international workshop. Washington, D.C.: National Academies Press. Retrieved from http://www.nap.edu/openbook.php ?record_id=13564. 10 Ball, A. & Duke, M. (2012). Data citation and linking. Edinburgh, UK: Digital Curation Centre. Retrieved from http://www.dcc.ac .uk/resources/briefing-papers/introduction-curation/ data-citation-and-linking. 11 Altman, M. & King, G. (2007). A proposed standard for the scholarly citation of quantitative data. D-Lib Magazine,13(3/4). doi:10.1045/march2007-altman 12 ‘Endorse the data citation principles’, URL:https://www. force11.org/datacitation/endorsements. relies upon data, the corresponding data should be cited. Unique Identification A data citation should include a persistent method for identification that is machine actionable, globally unique, and widely used by a community. Access Data citations should facilitate access to the data themselves and to such associated metadata, documentation, code, and other materials, as are necessary for both humans and machines to make informed use of the referenced data. Persistence Unique identifiers, and metadata describing the data, and its disposition, should persist – even beyond the lifespan of the data they describe. Specificity and Verifiability Data citations should facilitate identification of, access to, and verification of the specific data that support a claim. Citations or citation metadata should include information about provenance and fixity sufficient to facilitate verfiying that the specific timeslice, version and/or granular portion of data retrieved subsequently is the same as was originally cited. Interoperability and Flexibility Data citation methods should be sufficiently flexible to accommodate the variant practices among communities, but should not differ so much that they compromise interoperability of data citation practices across communities. Data citation for authors The first half of this guide is aimed predominantly at researchers. It discusses the practical business of citing datasets, such as how to construct a data citation and use it in a research paper. Ways of referencing data The usual way of referencing the data directly underlying a publication is by means of a data access statement. For open data, this statement should say what is available from which repository, and provide a URL, identifier or accession code to help access the data. For restricted data, the statement should indicate the legal or ethical reason for the restriction, and provide a 3 link to a permanent record explaining the conditions of access. While a simple statement of this sort fulfils the basic need to reference data, it falls short in several respects: •if there is a typographical error in the identifier or URL, there is no additional information to locate the data among the repository’s holdings; •authors may be tempted to give the URL of the repository, rather than one specific to the dataset; •it does not give due credit to the creators of the dataset – an especially important point if these are different from the authors of the publication; •it does not treat data as a first-class record of research. All of these issues may be resolved by enhancing the statement with a data citation. As with other citations, this involves providing an in-text pointer to an entry in the reference list. If the publisher is not willing to accept a data citation, it is sometimes possible to work around this by citing a data paper instead. This is a paper that describes the dataset and its collection without drawing any scientific conclusions from it. Such papers may be published in a special section of a regular journal, or in a dedicated data journal such as Earth System Science Data.13 The placement of the data access statement/in-text citation varies between journals. Some, including those published by PLoS and Pensoft, specify the use of a dedicated section, for example ‘Data resources’ or ‘Data access and terms of use’.14 Others encourage authors to place it at the end of the abstract. Where no specific advice is given, it is usually recommended that authors put the statement in the acknowledgements section; this is because, as with the acknowledgement of the grant, the statement is often a condition of funding and locating the two together simplifies the task of checking compliance.15 More detailed guidance on 13 Earth System Science Data,URL:http://www. earth-syst-sci-data.net/. 14 Penev, L., Mietchen, D., Chavan, V., Hagedorn, G., Remsen, D., Smith, V. & Shotton, D. (2011, May 26). Pensoft data publishing policies and guidelines for biodiversity data. Pensoft. Retrieved 4 July 2011, from http://www.pensoft.net/J_FILES/Pensoft_Data _Publishing_Policies_and_Guidelines.pdf. 15 Acknowledgement of funders in scholarly journal articles: Guidance for UK research funders, authors and publishers. (2008, February). Research Information Network. Retrieved 3 June 2011, from http: //www.rin.ac.uk/our-work/research-funding-policy-and -guidance/acknowledgement-funders-journal-articles. data access statements is given by some institutions.16 It is also possible to reference data from nontextual outputs, such as other datasets. Indeed, doing so may help satisfy the licensing conditions of the earlier datasets and encourage data sharing by supporting transitive credit models.17 One straightforward way of doing this is to include with the dataset a table that lists the source datasets and indicates the subset that each one contributed. An early example of this was published as part of the supplementary data of a 2011 paper on microattribution for genetic variation data.18 Another solution is to see if the repository holding the data will record the information among the metadata it holds for the dataset. Elements of a data citation The elements that would make up a complete citation are a matter of some debate. The following list is a superset taken from four different papers on the subject. 11,19,20,21 Author. The creator of the dataset.11,19,20,21 Publication date. Whichever is the later of: the date the dataset was made available,11 the date all quality assurance procedures were completed,19,20 and the date the embargo period (if applicable) expired.21 Title. As well as the name of the cited resource itself,11,21 this may also include the name of a facility19 and the titles of the top collection and main parent sub-collection (if any) of which the dataset is a part.20 16 See, for example, guidance given by the universities of Bath (http://www.bath.ac.uk/research/data/sharing-reuse/ data-access-statement.html) and Bristol (http://data.bris. ac.uk/research/using-data/). 17 Katz, D. S. (2014). Transitive credit as a means to address social and technological concerns stemming from citation and attribution of digital products. Journal of Open Research Software,2(1), e20. doi:10.5334/jors.be. 18 Giardine, B., Borg, J., Higgs, D. R., Peterson, K. R., Philipsen, S., Maglott, D., . . . Patrinos, G. P. (2011). Systematic documentation and analysis of human genetic variation in hemoglobinopathies using the microattribution approach. Nature Genetics,43, 295–301. doi:10.1038/ng.785. 19 Lawrence, B., Jones, C., Matthews, B., Pepler, S. & Callaghan, S. (2011). Citation and peer review of data: Moving towards formal data publication. International Journal of Digital Curation,6(2), 4–37. doi:10.2218/ijdc.v6i2.205 20 Green, T. (2010, February). We need publishing standards for datasets and data tables. OECD Publishing. doi:10.1787/ 787355886123 21 Starr, J. & Gastl, A. (2011). isCitedBy: A metadata scheme for DataCite. D-Lib Magazine,17(1/2). doi:10.1045/january2011 -starr 4 Edition. The level or stage of processing of the data, indicating how raw or refined the dataset is.19 Version. A number increased when the data changes, as the result of adding more data points or re-running a derivation process, for example.21 Feature name and URI. The name of an ISO 19101:200222 ‘feature’ (e.g. GridSeries, ProfileSeries) and the URI identifying its standard definition, used to pick out a subset of the data.19 Resource type. Examples: ‘database’,20 ‘dataset’.21 Publisher. The organisation either hosting the data21 or performing quality assurance.19 Unique numeric fingerprint (UNF). A cryptographic hash of the data, used to ensure no changes have occurred since the citation.11 Identifier. An identifier for the data, according to a persistent scheme.11,19,20,21 Location. A persistent URL from which the dataset is available. Some identifier schemes provide these via an identifier resolver service.11,19,20,21 The most important of these elements – the ones that should be present in any citation – are the author, the title and date, the location, and the publisher. These give due credit, allow the reader to judge the relevance of the dataset, permit access to it, and give reassurances about its quality or persistence, respectively. In theory, they should between them uniquely identify the dataset; in practice, a formal identifier is often needed. The most efficient solution is to give a location that consists of a resolver service and an identifier (for an example, see Figure 3 below). Note that the way in which these elements would be styled and combined together in the finished citation depends on the style in use for citations of textual publications. Figure 1 provides example data citations drawn from commonly used style manuals,23,24,25,26 22 ISO 19101. (2002). Geographic information – Reference model. 1st ed. International Organization for Standardization. 23 Publication Manual of the American Psychological Association (6th ed., p. 211). (2010). Washington, DC: American Psychological Association. 24 Chicago Manual of Style (16th ed., p. 764). (2010). Chicago, IL: University of Chicago Press. 25 Gibaldi, J. (2008). MLA style manual and guide to scholarly publishing (3rd ed., pp. 213-214, 238-239). New York: Modern Language Association of America. 26 Ritter, R. M. (2002). Oxford Manual of Style (p. 551). Oxford, UK: Oxford University Press. Waddingham, A. (Ed.). (2014). New Hart’s rules: The Oxford style guide (2nd ed., pp. 368-376). Oxford, UK: Oxford University Press. while Figure 2 shows the citation formats suggested by three data repositories. APA Cool, H. E. M., & Bell, M. (2011). Excavations at St Peter’s Church, Barton-upon-Humber [Data set]. doi:10.5284/1000389 Chicago (Footnote) H. E. M. Cool and Mark Bell, Excavations at St Peter’s Church, Barton-upon-Humber (accessed May 1, 2011), doi:10.5284/1000389. (Bibliography) Cool, H. E. M., and Mark Bell. Excavations at St Peter’s Church, Barton-upon-Humber (accessed May 1, 2011). doi:10.5284/1000389. MLA Cool, H. E. M., and Mark Bell. “Excavations at St Peter’s Church, Barton-upon-Humber.” Archaeology Data Service, 2001. Web. 1 May 2011. 〈http: //dx.doi.org/10.5284/1000389〉. Oxford Cool, H. E. M. and Bell, M. (2011), Excavations at St Peter’s Church, Barton-upon-Humber [dataset] (York: Archaeology Data Service), doi: 10.5284/1000389 Figure 1: Data citations in common styles PANGAEA Willmes, S et al. (2009): Onset dates of annual snowmelt on Antarctic sea ice in 2007/2008. doi:10.1594/PANGAEA.701380 Dryad Kingsolver JG, Hoekstra HE, Hoekstra JM, Berrigan D, Vignieri SN, Hill CE, Hoang A, Gibert P, Beerli P (2001) Data from: The strength of phenotypic selection in natural populations. Dryad Digital Repository. doi:10.5061/dryad.166 Dataverse Frederico Girosi; Gary King, 2006, ‘Cause of Death Data’, http://hdl.handle.net/1902.1/UOVMCPSWOL UNF:3:9JU+SmVyHgwRhAKclQ85Cg== IQSS Dataverse Network [Distributor] V3 [Version]. Figure 2: Data citation formats suggested by repositories Digital Object Identifiers There are several types of persistent identifier that could be used to identify datasets: examples include Handles, Archival Resource Keys (ARKs) and Persistent URLs (PURLs), all of which can be resolved to an 5 Internet location. The scheme that is gaining most traction is the Digital Object Identifier (DOI). The DOI System is an identifier scheme administered by the International DOI Foundation.27 It is built on the Handle System but has its own conventions and an independent business model. The identifiers themselves have the standard Handle structure of prefix, slash, suffix (see Figure 3). All DOI prefixes begin with ‘10.’ to mark them as such; the prefix may be further subdivided with dots, but otherwise the characters in a DOI have no special significance. http://dx.doi.org/ | {z } 10.5284 | {z } /1000389 | {z } resolver service prefix suffix (assigning (resource) body) Figure 3: Anatomy of a DOI While there are several services available that can resolve a DOI to an Internet location,28 the preferred one is http://dx.doi.org/. Appending a DOI to this URL creates a further URL that can be used to access the associated resource. Authors are encouraged to use the URL version of the DOI wherever possible, though some publishers prefer to print the bare DOI and embed the URL form as a hyperlink in digital versions. Individuals wishing to register a DOI for their dataset would normally do so via their disciplinary data archive, institutional data repository, or a data sharing service such as figshare29 or Synapse.30 Contributor identifiers If contributors have a common name, or move between many different institutions, giving them an unambiguous credit is somewhat problematic. A possible solution is for each contributor to be given a unique identifier, to be used in connection with all their publications, data contributions, and so on. While several identifier schemes are already well established, most are arguably unsatisfactory because they are either too narrowly scoped, proprietary or focused on authentication rather than attribution. There are 27 DOI System, UR L:http://www.doi.org/. 28 Some publishers provide resolvers for their own DOIs, while the Handle resolver http://hdl.handle.net/ can be used for any DOI. 29 Figshare, URL: http://figshare.com/. 30 Synapse, URL:https://www.synapse.org/. however two schemes being developed specifically for attribution. The Open Researcher and Contributor Identifier (ORCID) is a scheme specifically aimed at academic authors.31 It has gained support from over 300 organisations, including major academic publishers, and been integrated into numerous research systems. Researchers can associate with their ORCID profiles a list of works to which they have contributed, as well as grants received and their educational and employment history. ORCID profiles can also be linked to identifiers and profiles from other schemes such as Thomson Reuters’ ResearcherID,32 Scopus,33 Scholar Universe,34 and RePEc.35 The International Standard Name Identifier (ISNI) scheme is an ISO standard for registering ‘Public Identities’: people, pseudonyms, personas and legal entities involved in the creation or distribution of intellectual property.36 It is thus a broader scheme than ORCID, allowing organisations to be identified as well as individuals. ISNIs take the form of a 16-digit number (though the last digit may be ‘X’); each identifier is supported by a metadata record containing details such as name(s), date of birth, fields of endeavour and roles within them, titles of creations and a URI for further information. As the primary utility for such identifiers is to support software tools, they are better placed in machine-readable metadata than written out for human inspection. It is therefore recommended that authors do not attempt to include ORCIDs or similar in their reference lists, but rather ensure they supply their own ORCID to publishers and repositories at the point of making their submission. Granularity With print publications, the issue of citing at different levels of granularity is relatively straightforward. The documents listed within a bibliography or reference list represent intellectual wholes: single-author monographs are referenced as whole books, but with journal issues, conference proceedings and edited collections, the relevant papers are referenced individually. More granular references (to sections, pages, etc.) are made 31 ORCID, URL:http://orcid.org/. 32 ResearcherID, URL: http://www.researcherid.com/. 33 Scopus, URL: http://www.scopus.com/. 34 Scholar Universe, UR L:http://www.scholaruniverse.com/. 35 RePEc Author Service, URL:http://authors.repec.org/. 36 ISO 27729. (2012). Information and documentation – International standard name identifier (ISNI). International Organization for Standardization. 6 at the point of citation in the text, rather than in the reference list. Datasets are a little more complicated. A dataset may form part of a collection and be made up of several files, each containing several tables, each containing many data points. There are also more abstract subsets that can be used, such as features and parameters. At the other end of the scale, it is not always obvious what would constitute an intellectual whole: it can be argued, for example, that investigations should be the primary units of citation rather than individual datasets.37 For authors, the pragmatic solution is to list datasets at whatever level of granularity has been chosen by the host repository for assigning identifiers. If a finer level of granularity is required, the in-text citation should provide the reader with the information needed to find the subset. As conventions for doing this have yet to be established, if the repository provides identifiers at several levels of granularity, the finest-grained level that meets the need of the citation should be used in the reference list, to minimise the additional information needed. Citing unreleased data If citing a dataset that is not yet released, the rule of thumb is to provide in the reference as much information about it as is already known. At a miniumum, this should include the creator and title of the dataset. If the dataset has not yet been deposited, the date of collection should be included. If the dataset has been deposited but an online record is not yet available, the date can be given as ‘in press’ and the repository given in the publisher position. Once an online record is available, the full citation can be given. The full details of the status of the dataset – whether deposited, embargoed, restricted or openly available – should be explained in the data access statement. As with references to manuscripts that have not yet been published, authors should revisit references to unreleased data prior to publication to ensure the information is as up to date as possible. Citing physical data There is no difference in principle between how one should cite physical data, such as samples or materials, and digital data. Physical data is often less reproducible 37 Lawrence, B. (2011, January 7). Citation, Digital Object Identifiers, persistence, correction and metadata [Blog post]. Retrieved 12 May 2011, from http://home.badc.rl.ac.uk/ lawrence/blog/2011/01/07/citation , _digital _object _identifiers,_persistence,_correction_and_metadata. or shareable than digital data, but the majority of issues that apply to it also apply to digital data that are too sensitive or voluminous to be transported over the Internet. In practice, the issue most likely to cause confusion is how and whether to provide a URL for the physical data. If the physical data has an identifier in a scheme with a resolver service, this should be used as the URL. For example, the International Geo Sample Number (IGSN) is associated with a catalogue whose records can be accessed by appending the number to the resolver service URL: http://www.geosamples.org/profile?igsn=. If the physical data has an identifier that cannot be resolved, it should be quoted elsewhere in the reference. The URL, meanwhile, should point to a page explaining how to gain access to the data, if applicable. Summary for researchers •If you have generated/collected data to be used as evidence in an academic publication, you should deposit it with a suitable data archive or repository as soon as you are able. If they do not provide you with a persistent identifier or URL for your data, encourage them to do so. •When citing a dataset in a paper, use the citation style required by the editor/publisher. If no form is suggested for datasets, take a standard data citation style and adapt it to match the style for textual publications. •Give dataset identifiers in the form of a URL wherever possible, unless otherwise directed. •Include data citations alongside those for textual publications. Some reference management packages now include support for datasets, which should make this easier. •Cite datasets at the finest-grained level available that meets your need. If that is not fine enough, provide details of the subset of data you are using at the point in the text where you make the citation. •If a dataset exists in several versions, be sure to cite the exact version you used. •When you publish a paper that cites a dataset, notify the repository that holds the dataset, so it can add a link from that dataset to your paper. 7 Data citation for repositories The remainder of this guide is aimed at data repositories. It looks at the underlying infrastructure that supports data citation, and suggests ways in which repositories might participate in and build on existing activity. Tools and services This section provides an overview of some of the technologies available to support data citation. DataCite DOIs The task of managing DOI registers is delegated to registration agencies that each specialise in a type of resource. For research datasets, the registration agency is the DataCite Consortium.38 The consortium is made up of libraries and data centres from across the globe, led by the German National Library of Science and Technology (TIB). Among the services it provides are human and machine interfaces for simple end-user administration of DOI registrations. DataCite also collects metadata about each dataset it registers.39 These metadata may be searched through a Web interface40 or harvested using OAI-PMH.41 Any repository wishing to register DOIs needs to obtain a username and password from DataCite to gain access to the registration service. Alternatively, the organisation can manage its DOIs through a third-party service such as EZID.42 The username and password are not needed for the metadata search or OAI-PMH services. While best practice has yet to emerge on some matters, certain conventions are already becoming established. •When organisations register a DOI for a resource, they should not introduce semantic elements into the suffix, especially not metadata that might change over time (e.g. publisher, archive, owner). •As DOIs are used to cite data as evidence, the dataset to which a DOI points should also remain unchanged, with any new version receiving a new DOI. 38 DataCite, URL:http://www.datacite.org/. 39 DataCite Metadata Schema Repository, URL:http://schema. datacite.org/. 40 DataCite Metadata Search service, URL:http://search. datacite.org/. 41 DataCite OAI-PMH service, URL:http://oai.datacite.org/. 42 EZID, URL:http://ezid.cdlib.org/. Various organisations have shared their experiences of working with DataCite: •Archaeology Data Service; University of Southampton;43 •Australian Antarctic Division; Australian National University; Dryad;44 •ForestPlots.net, University of Leeds;45 •Griffith University;46 •UK Data Archive;47 •University of Bristol.48 Notification Services The CLADDIER Project developed a prototype Citation Notification Service for use by digital object repositories, based on the TrackBack protocol.49 The TrackBack protocol is one of a family of linkback protocols that allow a blog article to list and link to later articles that mention or comment on it, allowing the reader to follow a debate across many blogs.50 CLADDIER extended the protocol to permit richer metadata to be communicated each way between the citing and cited systems,51 allow previous TrackBacks to be updated or deleted, and reduce the likelihood of spam 43 British Library. (2013). Working with the British Library and DataCite: Institutional case studies. Retrieved from http://www.bl.uk/ aboutus/stratpolprog/digi/datasets/DataCiteCaseStudies _2013.pdf. 44 Australian National Data Service. (n.d.). Data citation [YouTube playlist]. Retrieved from https://www.youtube.com/playlist ?list=PLG25fMbdLRa4peWpeZslW0cLSPYNjcbc1. 45 British Library. (2015). Datacite case study: ForestPlots.net at the University of Leeds. Retrieved from http://www.bl.uk/aboutus/ stratpolprog/digi/datasets/ForestPlot_CaseStudy_ForBL .pdf. 46 Simons, N., Visser, K. & Searle, S. (2013). Growing institutional support for data citation: Results of a partnership between Griffith University and the Australian National Data Service. D-Lib Magazine,19(11/12). doi:10.1045/november2013-simons. 47 Jisc. (2012, April). Data identifiers: How to ensure your data is properly cited [Webinar]. Retrieved from https://www.jisc.ac .uk/events/data-identifiers-how-to-ensure-your-data-is -properly-cited-11-apr-2012. 48 Duke, M. & Gray, S. (2014). Assigning Digital Object Identifiers to research data at the University of Bristol. Edinburgh, UK: Digital Curation Centre. Retrieved from http://www .dcc .ac .uk/ resources/persistent-identifiers. 49 CLADDIER Project page, URL:http://www.jisc.ac.uk/ whatwedo/programmes/digitalrepositories2005/claddier. 50 Six Apart. (2007). TrackBack manual. Retrieved 18 October 2011, from http://www.movabletype.org/documentation/ trackback_manual.html. 51 Matthews, B., Portwin, K., Jones, C. & Lawrence, B. (2007, November 30). Recommendations for data/publication linkage (CLADDIER Project Report No. 3). STFC. retrieved 20 June 2012, from http://ie-repository.jisc.ac.uk/221/. 8 TrackBacks.52 As a demonstration, CLADDIER implemented the Citation Notification System in STFC’s ePub repository and the BADC repository. The follow-on project StoreLink implemented the system as plugins for EPrints, DSpace and Fedora repository software.53 StoreLink was itself followed by the Webtracks Project, which generalised the system to form the InterRepository Communication (InteRCom) protocol and extend its usage beyond e-print repositories to STFC’s ICAT data catalogue, open electronic notebooks and scientific publishers.54 Example Knowledge Blog55 was developed as an alternative scholarly publication platform based on WordPress56 blogging software. It makes heavy use of linkbacks, for example as the mechanism for linking an article with its reviews, and could therefore be used together with the Citation Notification Service to provide bidirectional links to datasets. Its KCite plugin allows for the automatic generation of citations from just a DOI (using the metadata lookup API from the CrossRef and DataCite registration agencies) or a PubMed identifier.57 The US-based SHARE initiative is implementing a different notification system that, while not directly related to citation, may assist with setting up links between systems.58 The SHARE Notify service collects metadata from publishers and repositories about events such as data or pre-prints being deposited, or papers being published. This metadata is indexed in a database and made available through JSON and Atom feeds. If a repository is aware that a manuscript related to a dataset it holds is about to be published, it could monitor the feeds to discover when publication occurs. Similarly, by contributing to the SHARE Notify service, the repository could enable a publisher to discover 52 Matthews, B., Duncan, A., Jones, C., Neylon, C., Borkum, M., Coles, S. & Hunter, P. (2009, December). A protocol for exchanging scientific citations. Fifth IEEE International Conference on e-Science (e-Science 2009) (pp. 171–177). Los Alamitos, CA: IEEE Computer Society. doi:10.1109/e-Science.2009.32. 53 StoreLink Project summary Web page, URL:http://www.jisc. ac.uk/whatwedo/programmes/digitalrepositories2007/ storelink.aspx. 54 Webtracks Project blog, URL:http://webtracks. jiscinvolve.org/. 55 Knowledge Blog, URL:http://knowledgeblog.org/. 56 WordPress, URL: https://wordpress.org/. 57 KCite WordPress plugin, URL:https://wordpress.org/ plugins/kcite/. 58 SHARE initiative, UR L:http://www.share-research.org/. when the dataset underlying a manuscript has been released. Citation tracking services One of the benefits of using formal data citations is that it should make it easier to assemble evidence that a dataset has had impact. As of mid-2015 it is quite hard to do this due to the variety of ways in which datasets are referenced in the literature. Nevertheless there are some services available that index these references. The Thomson Reuters Data Citation Index was launched in October 2012.59 It tracks citations and less structured references to data at four levels of granularity: nanopublications (see below), datasets, research studies, and data repositories. It relies not only on access to the full text of publications but also on an index of available datasets, the information for which is drawn from data repositories, data discovery services and DataCite. Europe PubMed Central routinely text-mines its archive of full text articles for data citations. The resulting information is used by the PLoS Article Level Metrics service,60 and is also available through the Europe PubMed Central RESTful Web service.61 For more information about tracking the impact of datasets, please see the DCC guide ‘How to Track the Impact of Research Data with Metrics’.62 Network tracking services The DLI (Data Literature Interlinking) Service is a result of a collaboration between the ICSU-WDS/RDA Data Publishing Services Working Group and the OpenAIRE initiative.63 The aim of the service is to create a centrally curated graph of links between publications and datasets. Instead of relying on many bilateral arrangements between organisations, the idea is that publishers contribute to the graph links relating to articles they have published, while repositories contribute links relating to datasets they hold. Any 59 Thomson Reuters Data Citation Index product page, URL:http: //wokinfo.com/products_tools/multidisciplinary/dci/. 60 Lin, J. & Fenner, M. (2013, December 5). Research findings: Going deeper than the article [Web log post]. Retrieved from PLoS Tech Blog: http://blogs.plos.org/tech/research-findings -going-deeper-than-the-article/. 61 Europe PubMed Central RESTful Web service, URL:http: //europepmc.org/RestfulWebService. 62 Ball, A. & Duke, M. (2015). How to track the impact of research data with metrics. Edinburgh, UK: Digital Curation Centre. Retrieved from http://www.dcc.ac.uk/resources/how -guides/track-data-impact-metrics. 63 DLI Service, URL:http://dliservice. research-infrastructures.eu/. 9