Full text
A Digital Curation Centre ‘working level’ guide How to Track the Impact of Research Data with Metrics Alex Ball (DCC) and Monica Duke (DCC) Please cite as: Ball, A., & Duke, M. (2015). ‘How to Track the Impact of Research Data with Metrics’. DCC How-to Guides. Edinburgh: Digital Curation Centre. Available online: http://www.dcc.ac.uk/resources/how-guides Digital Curation Centre, 2015. Licensed under Creative Commons Attribution 4.0 International: http://creativecommons.org/licenses/by/4.0/
How to Track the Impact of Research Data with Metrics Introduction This guide will help you to track and measure the impact of research data, whether your own or that of your department/institution. It provides an overview of the key impact measurement concepts and the services and tools available for measuring impact. After discussing some of the current issues and challenges, it provides some tips on increasing the impact of your own data. This guide should interest researchers and principal investigators working on data-led research, administrators working with research quality assessment submissions, librarians and others helping to track the impact of data within institutions. Why measure the impact of research data? A key measure of the worth of research is the impact it has or, to put it another way, the difference it is making both within the academic community and beyond. In recent years funding bodies have placed increasing emphasis on monitoring the potential and actual impact of the research projects they fund, as distinct from evaluating the intrinsic academic quality and value of research outputs as judged solely by other academics. Since 1997, the NSF has judged the merit of research proposals on their intellectual merit and their broader impact.1In the UK, impact plans became part of the bidding process for all Research Councils in 2009, though in 2010 the purpose of the plans was clarified by reformulating them as Pathways to Impact.2 In this part of their proposals, researchers are asked to consider how they might maximise the academic, societal and economic impact of their research. At the other end of the research lifecycle, the 1National Science Board. (2011, December 11). National Science Foundation’s merit review criteria: Review and revisions (Report No. NSB/MR-11-22). National Science Foundation. Retrieved from http://www.nsf.gov/nsb/publications/2011/ meritreviewcriteria.pdf. 2Hodgson, S. & Porter, L. (2010, October 13). Pathways to impact. Presentation given to the University of Cambridge Research Operations Office. Retrieved from http://www.admin.cam.ac .uk/offices/research/documents/local/presentations/ 2010_10_13_NERC_Impact.pdf. 2014 Research Excellence Framework (REF) in the UK included impact as an explicit element alongside outputs and environment.3Submissions were in the form of case studies of social, economic and cultural benefits and impacts arising from research activity. Furthermore, when the Higher Education Funding Council for England (HEFCE) – one of the agencies responsible for the REF – undertook a review of the role of metrics in research assessment in 2014-2015, it considered how they might be used to assess both the quality of academic research and its broader impact.4 There are many reasons underlying this emphasis on impact. For one, it provides tangible evidence of benefit to weigh against the costs of research. For another, it provides an engaging way of comparing peer research programmes across the globe, albeit through the lens of proxy indicators, when undertaking strategic decision-making or benchmarking. It is not, however, ideal for making comparisons across disciplines as each one has its own pattern of impact, operating over a different timescale. In order to accommodate these differences as far as possible, funders tend to take into account a wide variety of ways in which research can be influential. This means going beyond 3Research Excellence Framework. (2011, March). Decisions on assessing research impact (Report No. REF 01.2011). HEFCE et al. Retrieved from http://www.ref.ac.uk/media/ref/content/ pub/decisionsonassessingresearchimpact/01_11.pdf. 4Higher Education Funding Council for England. (2014). Independent review of the role of metrics in research assessment. Retrieved from http://www.hefce.ac.uk/rsrch/metrics/. 2
a traditional bibliometric analysis of academic outputs to consider how wider societal needs have been met by research efforts. Research can have impact by influencing practice or policy, generating wealth, driving industrial innovations, tackling pressing societal questions or problems, or meeting the needs of a particular community. It is therefore in the interests of researchers and institutions to track the impact of their research. An obvious place to start is with the impact of research outputs, including datasets. Admittedly, the prospect of using quantitative measures for assessing impact is not without controversy. Social and political concerns include the encroachment on academic freedom and creativity, and the effects on the well-being of researchers of working in a culture of measurement.5The limitations of what is being measured must also be recognised: awareness is needed of the specific types of impact being recorded, which may not be comprehensive especially given how broad the consideration of impact could be. The output from tools for tracking impact must also be carefully considered. As discussed below, due to the immaturity of the area some of the measurements may not be comparable. A note of caution must be sounded if using derived data for decision making: knowledge of the strengths and weaknesses of different metrics must be taken into account. Researchers can, however, start to use some metrics as indicators of impact and follow them up as potential leads that could become the basis of demonstrating impact in a case study. Furthermore, by monitoring usage of their shared datasets, researchers can get to know which forms of data preparation and data publication work best, and adjust their practices accordingly. By tracking who is reusing their data, researchers may uncover opportunities for collaboration from among their peers, and may identify communities who, even though they were not the original intended audience (e.g. the public), have an interest in the data. Institutions can benefit from data usage monitoring when they come to •plan for and monitor the success of the infrastructure providing access to the data, in particular to gauge capacity requirements in storage, archival and network systems; 5Kansa, E. (2014, January 27). It’s the neoliberalism, stupid: Why instrumentalist arguments for Open Access, Open Data, and Open Science are not enough [Web log post]. Retrieved from London School of Economics, The Impact of Social Sciences web log: http://blogs.lse.ac.uk/impactofsocialsciences/2014/01/ 27/its-the-neoliberalism-stupid-kansa/. •instigate promotional activities and celebrate data sharing and re-use successes by researchers at the institution; •create special collections around popular datasets; •meet funder requirements to safeguard data for the appropriate length of time since last use.6 As an example of the latter points, in October 2014, the Engineering and Physical Sciences Research Council (EPSRC) set out a clarification of its expectations of the research organisations it funds.7It explained that organisations are expected to log requests to access the data they hold, and recommended they also do likewise for data their researchers have deposited elsewhere (Expectation III). Not only are such logs ‘a valuable indicator of impact’, they along with other measures of interest8can inform decisions about data retention. In particular, EPSRC stated it did not expect a dataset to be retained if no interest has been shown in it for a period of 10 years (Expectation VII). These considerations are all important in the wider movement to improve quality and transparency, increase efficiency, and widen the opportunities for academic research through data sharing. Narrative accounts of high-impact data sharing and re-use can be used to advocate cultural change. Meanwhile, data metrics may be used to incentivise data sharing within a framework of professional advancement and reward that recognizes data as a fundamental research output.9 Impact measurement concepts Impact is, in its figurative sense, the effect or influence that one agent, event or resource has on another. It is distinct from, but related to, concepts such as attention (how many people know about the entity) and dissemination (how widely a resource has been 6Jones, C. (2014, December 2). RDMF12: Notes from breakout group 4 (tracking uses) [Web log post]. Retrieved from http: //www.dcc.ac.uk/blog/rdmf12-notes-breakout-group-4. 7Engineering and Physical Sciences Research Council. (2014, October 9). Clarifications of EPSRC expectations on research data management. Retrieved from http : / / www .epsrc .ac .uk / files / aboutus / standards / clarificationsofexpectationsresearchdatamanagement/. 8The document states it is ‘reasonable to use data citations, or any other metric based on reliable sources of evidence and widely accepted at the time, to evaluate if interest has been shown in a dataset’ (Expectation VII). 9Costas, R., Meijer, I., Zahedi, Z. & Wouters, P. (2013, April). The value of research data: Metrics for datasets from a cultural and technical point of view. Retrieved from Knowledge Exchange website: http://www.knowledge-exchange.info/datametrics. 3
distributed). When considering proposed metrics, it is therefore important to consider what exactly is being measured and the strength of the evidence it provides for the impact of the entity in which one is interested.10 For example, citation counts are often used as a measure of the influence that a paper has on subsequent literature in a discipline. They are not a true measure, as papers may be cited for reasons other than acknowledging influence (e.g. as part of a refutation, or to acknowledge an unused line of enquiry), but serve as a useful proxy measure. In aggregrate, citation counts can be used as a proxy measure for the impact of other entities – the authors, the funding body, and so on – but at a weaker level of confidence. The h-index, for example, is a measure of researcher impact and productivity derived from the citation counts of papers.11 Researchers have an index hif exactly hof their published papers have been cited hor more times. This way of aggregating the citation counts means that researchers have to produce highly cited papers in quantity in order to score highly; a large quantity of poorly cited papers or a one-off influential paper are not enough. Another measure derived from citation counts is the Journal Impact Factor (JIF), which gauges the impact of a journal in a given year. It is defined as the mean number of citations received by the papers published by the journal in the preceding two years from papers published (in any journal) in the year in question. The official JIFs are calculated by Thomson Reuters from the papers indexed by Web of Knowledge, and published in the Journal Citation Reports (JCR) product.12 Despite being a measure of the impact of a journal, not that of its constituent papers, it is often used as a proxy measure for the prestige of the journal, and thereby (controversially) of the impact of the authors whose papers are published in that journal. There are compelling arguments against using the JIF in so simplistic a manner. It is no longer the case that it is prohibitively time consuming to apply metrics to articles and authors on an individual basis. Furthermore, measuring the impact of an entity through indirect means, as happens with both the JIF and the h-index, provides only an incomplete picture. 10 Hicks, D., Wouters, P., Waltman, L., de Rijcke, S. & Rafols, I. (2015, April 23). Bibliometrics: The Leiden Manifesto for research metrics. Nature,520, 429–431. doi:10.1038/520429a. 11 Hirsch, J. E. (2005). An index to quantify an individual’s scientific research output. Proceedings of the National Academy of Sciences of the United States of America,102, 16569–16572. doi:10.1073/pnas.0507655102. 12 Variations on the JIF are possible; for example, the JCR also provides a Five-Year Impact Factor which may be more relevant for disciplines with longer publication cycles. Traditionally, we have attempted to measure the impact of the journal in which that research was published as a proxy for the impact of the research itself. However, this method is becoming increasingly problematic as more research is created and disseminated digitally and in forms beyond the traditional journal article. – Andy Wesolek13 This is one reason why funding bodies have so far been reluctant to use metrics directly, preferring narrative case studies that can explore the full range of possible impact. For example, Figure 1 shows the types of impact recognised by BBSRC in its Policy on Maximising the Impact of Research.14 Specifically data-related impact researchers can have includes •reuse of data they have shared to derive new knowledge; •incorporation of their data into larger datasets or data products; •widespread use of software or workflows they have written. No one metric can hope to represent fairly all these possibilities, so it is worth exploring the variety of metrics that can be used. As noted above, there are risks and concerns about reading too much into any given statistic, but metrics do provide an accessible Scientific advancement Knowledge Jobs Equipment Skills Training Schools New companies Knowledge economy Processes Products Inward investment Wealth creation Communication Public Engagement Societal issues International development Public health Policy Excellent people Excellent research Figure 1: The variety of impact recognised by the UK Biotechnology and Biological Sciences Research Council 13 Wesolek, A. (2014). Metrics: Understanding your impact. Retrieved from Clemson University Libraries website: http: //libguides.clemson.edu/metrics. 14 Biotechnology and Biological Sciences Research Council. (2012). Bbsrc policy on maximising the impact of research. Retrieved from http://www.bbsrc.ac.uk/documents/bbsrc-impact-policy -pdf/. 4
way of uncovering evidence that might be suitable for use in an impact case study. Citations of data The most mature emerging model for measuring the impact of data is one that is analogous to the publication and citation of literature.15 Going beyond mere data sharing, where data is simply made available (e.g. as files on a website), data publication implies that the data has entered a framework for checking its quality, ensuring it is fit for reuse, making it searchable and discoverable, and guaranteeing its long-term accessibility. The resulting dataset is given stable bibliographic information so that it can be reliably cited by other scholarly outputs. Such citations can be counted in the usual manner to provide evidence of the impact of the dataset, subject to the limitation described above. The analogy can be taken further, as datasets can meaningfully make references as well as receive citations. One suggested formal mechanism for this is to package the data within a research object rich in metadata.16 If this idea gains traction, derived data products could then cite the source data, with referencing between data products (or research objects) providing a complementary citing network alongside that of publications. While direct citation of and between datasets is far from widespread, and may be considered an aspirational end goal, some disciplines are taking a transitional approach where citations are made instead to a data paper. This is a paper that describes the dataset and its collection without drawing any scientific conclusions from it. Such papers may be published in a special section of a regular journal, or in a dedicated data journal such as the Journal of Open Archaeology Data.17 Citations of data papers may be interpreted as citations of the underlying dataset for the purposes of assembling evidence of impact. In many disciplines, however, the dominant approach is the traditional one of citing the first paper to make use of the data, relying on that paper to indicate if and how the data has been shared. It is not usually possible, at least not without significant manual effort, to identify whether citations to such papers should 15 Costas, R., Meijer, I., Zahedi, Z. & Wouters, P. (2013, April). The value of research data: Metrics for datasets from a cultural and technical point of view. Retrieved from Knowledge Exchange website: http://www.knowledge-exchange.info/datametrics. 16 Research Objects, URL:http://www.researchobject.org. 17 Journal of Open Archaeology Data,URL:http:// openarchaeologydata.metajnl.com/. count towards the impact of the argumentation and conclusions of the paper, the underlying dataset, or both. In such disciplines, therefore, citation counts are of little help as an indicator of data impact, so alternative indicators must be found. Resolutions Many datasets have been given persistent, unique identifiers to assist with unambiguous referencing, and many of the schemes in use are resolvable. In other words, there are bridge services that map the identifiers to one or more Internet locations. One such scheme is the Digital Object Identifier (DOI), for which DataCite is the main Registration Agency for research datasets. Among the services it provides is DataCite Statistics,18 which shows on a monthly basis the number of times the top ten DOIs for each prefix have been resolved to a URL. Account holders have access to resolution data for all the DOIs they manage. These statistics give an indication of how often references to the dataset have been followed. Page views Web servers log each interaction they have with a client, so by analysing the logs it is possible to count approximately how many times a webpage has been opened by a browser. Some sites additionally embed in their pages JavaScript code that notifies an analysis application each time a page is viewed. Either way, this statistic can be used to infer the level of interest in that page. When datasets are made available online, best practice is to provide a corresponding webpage displaying a catalogue record for the dataset. At a minimum one would expect the page to display the dataset title, a statement of responsibility, a short description, and a download link (or instructions on how to gain access). The number of times the dataset catalogue page has been viewed gives an indication of the level of interest in the dataset, and the level of awareness of its existence. Downloads Web server logs can also be used to count the number of times a data file has been downloaded. This indicates a stronger level of interest in the data than can be inferred from a count of catalogue page views, since it implies a desire to look at the actual data, but 18 DataCite Statistics, URL:http://stats.datacite.org/. 5
the statistic alone does not reveal the use to which the downloaded data might be put. At the time of writing the number of repositories and data archives that make download statistics openly available is quite small – one is VectorBase19 – but several others are known to collect and use them internally.20 For example, the UK Data Service Discover catalogue can put search results in order from most to least downloaded. Social media links Perhaps the alternative metrics closest to references in journal articles are those that measure the topicality of the dataset on social media platforms. If people are moved to share or discuss a dataset with friends, colleagues and the wider world, there is a likelihood it has affected them in some way, meaning it is worth looking closer for evidence of impact. Twitter21 is a social networking tool that enables users to send short messages known as ‘tweets’ to their followers. As they are limited to 140 characters, tweets lend themselves to immediate reaction and brief sentiments. A tweet referring to a piece of research might contain a link to, say, a research output, a project website or a blog post that discusses it, accompanied by a comment on it. Detecting tweets that relate to a dataset can be tricky, but a possible search strategy is to look for mentions of the dataset’s identifier or links to its catalogue page. Once a tweet has been found, for an accurate picture of the impact the dataset is having it is useful to consider whether the tone of the tweet is positive, negative or neutral, just as it is when analysing traditional citations. It can also be informative to track any ensuing conversation – replies or forwards (‘retweets’) – as these can indicate how far others agree with the sentiments of the original tweet. It should be noted that there is a cultural dimension to tweeting; for example, in some communities it is not considered a professional activity for academics. Social bookmarking and bibliographic services are not used quite so widely as Twitter, but are much less ‘noisy’ as a source of evidence of interest in scholarly outputs. Services such as Mendeley, 19 VectorBase, URL:https://www.vectorbase.org/. 20 Costas, R., Meijer, I., Zahedi, Z. & Wouters, P. (2013, April). The value of research data: Metrics for datasets from a cultural and technical point of view (pp. 39–42). Retrieved from Knowledge Exchange website: http://www.knowledge-exchange.info/ datametrics. 21 Twitter, URL:http://twitter.com/. 22 Beal, E. (2014, September 26) [Twitter status]. Retrieved from https://twitter.com/bumblebeal/status/5154607345167278 08. Applying altmetrics to data could be even more useful than for papers; data not well cited, but much interest in how they are used #1amconf – Eleanor Beal (via Twitter)22 CiteULike, BibSonomy and Delicious allow users to record online resources for their own reference or to recommend them to others.23 While the functionality offered by each service differs, one can typically discover how many users have bookmarked a particular link or resource. Some, notably Digg and Reddit,24 also provide information on how many users have up-voted or down-voted the resource; this may give an impression on whether the resource is having a positive impact or not. Several blogging platforms use a system of linkbacks to track conversations between blogs. When a new post is written to Blog A about an old post in Blog B, Blog A sends a notification to Blog B. Blog B can use this information to incorporate an extract of the new post as a comment on the old one. If a data repository or archive is set up to receive linkbacks, it can use them to monitor where particular datasets are being mentioned among publishing platforms supporting the protocol. Post-publication peer review One proposed innovation in the field of scholarly communications is the use of post-publication peer review as a method of quality control. While it is has yet to establish itself as a practice, there are several places where such reviews may be found, such as Faculty of 1000 (both as an integral part of its own publications and as a service reviewing other literature) and PubPeer.25 While the emphasis has to date been on reviewing journal papers, there are moves to apply the principle to datasets as well. The data papers submitted to data journal Earth System Science Data, for example, are given a brief internal review before being published in its companion title Earth System Science Data Discussions,26 where anyone is able to 23 Mendeley, URL:http://www.mendeley.com/. CiteULike, URL: http://www.citeulike.org/. BibSonomy, URL:http://www. bibsonomy.org/. Delicious, URL:https://delicious.com/. 24 Digg, URL:http://digg.com/. Reddit, URL:http://www. reddit.com/. 25 Faculty of 1000, URL:http://f1000.com/. PubPeer, URL: https://pubpeer.com/. 26 Earth System Science Data Discussions,URL:http: //www.earth-syst-sci-data-discuss.net/papers_in_open_ discussion.html. 6
submit a review. Only once the paper has satisfactorily passed this public review phase can it proceed for publication in the main journal. While post-publication peer review is concerned with quality rather than impact, the text or nature of such reviews may reveal evidence of reuse. There are suggestions that data archives might invite those who have downloaded and reused their data to leave feedback on the dataset’s landing page.27 Not only would this provide a scalable source of peer review and insights into what makes data reusable, it would also provide confirmation that other researchers had attempted to use the data and, in the case of successful reuse, that the data has made an impact. Impact measurement services Thomson Reuters Data Citation Index Data Citation Index at a glance •Counts formal and informal citations of datasets by papers •For: researchers, librarians, funders •Pricing: institutional subscription, price on application •http://wokinfo.com/products_tools/ multidisciplinary/dci/ In October 2012, Thomson Reuters launched the Data Citation Index (DCI) as part of its Web of Knowledge service.28 It provides records at four levels of granularity: nanopublication, dataset, data study (a research activity producing one or more datasets) and repository. The records can be searched and filtered in various ways, in the same way as (and indeed in combination with) the other indices in Web of Knowledge. The records are linked so that, for example, from a repository record one can view records for the data studies and datasets held by that repository. Sample citations are also provided. On each record, the DCI displays the number of times the entity has been cited in Web of Knowledge. Recognising the variety of ways in which datasets 27 Marjan Grootveld, J. v. E. (2012). Peer-reviewed open research data: Results of a pilot. International Journal of Digital Curation,7(2), 81–91. doi:10.2218/ijdc.v7i2.231. 28 Herther, N. K. (2012, October 29). Thomson Reuters tackles open access datasets with Data Citation Index. NewsBreaks. Retrieved from http://newsbreaks.infotoday.com/NewsBreaks/ Thomson-Reuters-Tackles-Open-Access-Datasets-With-Data -Citation-Index-85849.asp. and repositories can be cited, the DCI counts not only entries in the reference list but also less formal citations that occur elsewhere in scholarly papers (for example, in the abstract or acknowledgements). Selection for inclusion in the DCI is at the level of whole repositories rather than individual datasets or studies. The criteria used for selecting repositories include longevity, sustainability, activity (in terms of new data being deposited), metadata held for the data (ideally in English, with links to associated literature, funding information, and so on), and quality assurance procedures.29 ImpactStory ImpactStory at a glance •Collects altmetric statistics for a portfolio of scholarly outputs •For: researchers •Pricing: US$60 per annum •https://impactstory.org/ ImpactStory allows researchers to build a profile to showcase their various academic activities.30 After registration, users add to their profile their various scholarly outputs, such as articles, presentation slides, videos, data, or software. This can be done by quoting their respective URLs or identifiers including PubMed IDs and DOIs. ImpactStory then uses various external services to track metrics relevant to the impact of those resources. Some of the metrics used are specific to the service used to host the resource. For example, ImpactStory tracks the number of times videos on Vimeo and YouTube have been viewed and ‘liked’, the number of times a GitHub repository has been forked, and the number of times resources in the Dryad data repository, Figshare, PLoS journals and SlideShare have been downloaded. Other metrics track interest in a resource independent of where it is hosted. The service can look up citation counts in Scopus, bookmark counts in Mendeley and CiteULike, and mentions on Facebook, Google+, Twitter, Wikipedia, and in blog posts. The metrics are used to compile reports on the interest shown in the user’s portfolio of 29 Thomson Reuters. (2012). Repository evaluation, selection, and coverage policies for the Data Citation Index within Thomson Reuters Web of Knowledge. Retrieved from http://wokinfo .com//products_tools/multidisciplinary/dci/selection _essay/. 30 ImpactStory, URL:https://impactstory.org/. 7
outputs, highlighting the most popular resources and providing some aggregate statistics. These reports are typically emailed to the user on a weekly basis. ImpactStory operates as a non-profit organisation registered in the USA. It grew out of a hackathon project, ‘total-impact’, that was developed at the 2011 Beyond Impact workshop. Since 2012 it has received funding from the Open Knowledge Foundation, the National Science Foundation, Jisc and the Sloan Foundation, and is also supported by registration fees of $60 per annum (at the time of writing in 2015). The data collected by the service is made open, unless restricted by third parties, and may be exported for single items or whole profiles at any time. The code and governance of the service are also open. Ideas for further development of the service are invited through a feedback forum, and users can vote for the ideas they would like the developers to prioritise. Example ‘Mousing over’ elements of an ImpactStory profile reveals more information. The pop-up for Bik et al.’s 2012 dataset on benthic microbial eukaryote communities reveals it has received more Dryad views than 86% of the datasets from that year tracked by the service.31 PLoS Article-Level Metrics Article-Level Metrics at a glance •Collects altmetric statistics for PLoS articles; other publishers/repositories may use the same software •For: researchers, institutions, funders •Pricing: Free •http://article-level-metrics.plos.org/ alm-info/ 31 Holly Bik’s ImpactStory profile – datasets, URL:https:// impactstory.org/HollyBik/products/datasets In 2009, the Public Library of Science (PLoS) launched its Article-Level Metrics (ALM) service.32 This compiles a set of impact indicators from PLoS’s own systems and various other services, and makes them available in both a visual way and via an application programming interface (API). The metrics compiled include •usage statistics (views and downloads) from PLoS and Pubmed Central; •interactions (comments, notes, ratings) on the PLoS website; •citations identified by Scopus, Web of Science and others; •references made in social networks like Twitter and Facebook, on various blogging platforms, or on Wikipedia. The metrics are displayed on the landing pages for PLoS articles, and can also be compiled into custom reports.33 PLoS released the source code for the ALM application in 2011.34 It was used as the basis of the (now discontinued) ScienceCard service, which provided an author-centric view on the same data.35 It was also taken up by other publishers and service providers, most significantly by CrossRef Labs, meaning statistics are available for many non-PLoS papers as well.36 While the implementations of the software so far have concentrated on papers, the software itself is resource-type agnostic, so could be applied to datasets. PlumX PlumX is the main product of Plum Analytics,37 a company formed in 2011 and acquired by EBSCO at the start of 2014.38 It aims to provide a more comprehensive picture of research impact than citations alone, 32 PLoS ALM website, URL:http://article-level-metrics. plos.org/alm-info/ 33 ALM Reports website, URL:http://almreports.plos.org/ 34 Lagotto (Article-Level Metrics) source code repository, URL: https://github.com/mfenner/lagotto 35 Fenner, M. (2011, September 28). Announcing ScienceCard [Web log post]. Retrieved from http://blogs .plos .org/ mfenner/2011/09/28/announcing-sciencecard/. 36 Lin, J. & Fenner, M. (2014, February 24). One step closer to article-level metrics openly available for all scholarly content [Web log post]. Retrieved from http://articlemetrics.github.io/ blog/2014/02/24/alms/. 37 PlumX website, URL:http://plu.mx/. 38 Harris, S. (2014, October). Acquisition opens up altmetrics options. Research Information. Retrieved from http://www .researchinformation.info/features/feature.php?feature _id=490. 8
PlumX at a glance •Collects altmetric statistics for an organisation’s scholarly outputs •For: institutions, funders, publishers •Pricing: institutional subscription, price on application •http://plu.mx/ and in particular to give insight into the impact that resources have in the period before the first citations are counted. The PlumX product is aimed at organisations rather than individuals, so Plum Analytics counts among its customers universities, corporations, publishers and funders, and reports rapid growth since the acquisition. PlumX aggregates information from a wide range of external sources about the impact of research outputs, including datasets and source code as well as more traditional publications. The metrics are grouped into five categories: •usage: the number of times the resource has been viewed or downloaded, the number of times a link to it (from Twitter or Facebook) has been clicked, the number of users contributing to it (on GitHub), the number of libraries that hold a copy; •captures: the number of times the resource has been marked as being of interest (e.g. bookmarked on Delicious; added to a Mendeley library; followed, forked or watched on GitHub); •mentions: the number of blog posts written about it, the number of comments made about it (on Facebook, Slideshare, YouTube, etc.), the number of reviews received (on Amazon or Goodreads); •social media: the number of times the resource has been recommended (e.g. by means of ‘likes’ on Facebook, ‘+1s’ on Google+, net upvotes on Reddit, tweets); •citations: the number of citations the resource has received according to Scopus, CrossRef, and various other sources. The totals for these metrics are displayed in a dashboard; bar charts or sunburst diagrams supplement a tabular view, and data is not only available at the level of individual resources, but also aggregated for individual researchers, resource types, and various levels of organisational unit (programmes, departments, whole organisations).39 A ‘plum print’ summary is 39 Example PlumX profile for the University of Pittsburgh, URL: https://plu.mx/pitt/. available for embedding into other sites, such as an institutional repository. The information shown in the embedded widget is customisable and links to the original data source are made available. Researchers can help seed the information available by linking their PlumX profiles to accounts they have in other systems (e.g. Slideshare, GitHub).40 Example When Jason Colditz wrote a blog post on Open Access publishing, he illustrated it with a flowchart of the publication process that he had deposited in figshare.41 His institution’s PlumX profile tracks interest in the figure.42 Altmetric Altmetric at a glance •Collects altmetric statistics for an organisation’s scholarly outputs •For: researchers, institutions, publishers •Pricing: free for researchers, price on application for commercial/institutional licenses •http://www.altmetric.com/ Altmetric is an article-centred service which monitors various sources for mentions of scholarly articles. 40 Michalek, A. (2014, July 31). Plum Analytics and our approach to altmetrics [Webinar recording]. Retrieved from http://bit.ly/ PlumW923. 41 Colditz, J. (2012). Publication process with oa [Figure]. doi: 10.6084/m9.figshare.91458. 42 PlumX profile for the figure ‘Publication Process with OA’, URL:https://plu.mx/a/ 22ogwhQ4i9naHqqVVtbKuB8m1MoJ9sfe83lmONsd_u0 9
•Konkiel, S., Dalmau, M. & Scherer, D. (2015). Altmetrics and analytics for digital special collections and institutional repositories.doi:10 .6084/m9 .figshare.1392140 •MacGillivray, M. (2012, December 12). Metrics for repository impact [Webinar recording]. Retrieved from http://www.rsp .ac .uk/events/impact -metrics-for-repositories/ •Neylon, C. (2014, October 3). Altmetrics: What are they good for? [Web log post]. Retrieved from PloS Opens web log: http://blogs.plos.org/opens/ 2014/10/03/altmetrics-what-are-they-good -for/ •National Information Standards Organization. (2012, November). NISO Webinar: Beyond publish or perish: alternative metrics for scholarship. Retrieved from http://www.niso.org/news/events/2012/ nisowebinars/alternative_metrics/ •National Information Standards Organization. (2014b, June). NISO Virtual Conference: Transforming assessment: alternative metrics and other trends. Retrieved from http://www .niso .org/news/ events/2014/virtual/assessment/ •Penfield, T., Baker, M. J., Scoble, R. & Wykes, M. C. (2014). Assessment, evaluations, and definitions of research impact: A review. Research Evaluation,23, 21–32. doi:10.1093/reseval/rvt021 •Piwowar, H. A. & Vision, T. J. (2013). Data reuse and the open data citation advantage. PeerJ,1, e175. doi:10.7717/peerj.175 •Public Libary of Science. (2012). Altmetrics collection. Retrieved from http://www .ploscollections .org/altmetrics •Strasser, C. (2013, October 15). Universities can improve academic services through wider recognition of altmetrics and alt-products [Web log post]. Retrieved from London School of Economics, The Impact of Social Sciences web log: http://blogs.lse .ac.uk/impactofsocialsciences/2013/10/15/ universities-can-improve-academic-services -through-altmetrics/ •Tattersall, A. & Beecroft, C. (2014, September). Altmetrics in the academy: Strategies for better academic engagement, dissemination and measurement. Presentation given at the Social Media for Researchers symposium, Sheffield Hallam University. Retrieved from http://t.co/2jPWTLeJJO •UK Data Service. (n.d.). Data impact blog. Retrieved from http://blog.ukdataservice.ac.uk/ •Utrecht University Library. (2015). Research impact and visibility: Traditional and altmetrics. Retrieved from http://libguides.library.uu .nl/researchimpact The Digital Curation Centre (DCC) is a consortium of the Universities of Edinburgh, Glasgow and Bath, and receives funding from Jisc. Follow the DCC on Twitter: @digitalcuration, #ukdcc Published: 29 June 2015 16