scieee AI-readable full text Open interactive document viewer

Supporting appendices for: The hunt for research data: development of an open-source workflow for tracking institutionally affiliated research data publications

Gee, Bryan M.

Abstract

These are the supporting appendices for the manuscript titled "The hunt for research data: development of an open-source workflow for tracking institutionally affiliated research data publications," which has been accepted for publication in the Journal of eScience Librarianship.

Full text

1 Supporting appendices for: “The hunt for research data: development of an opensource workflow for tracking institutionally affiliated research data publications” Bryan M. Gee University of Texas Libraries, The University of Texas at Austin, Austin, TX, United States of America; [email protected]; 0000-0003-4517-329 These are the supporting appendices for the journal article, “The hunt for research data: development of an open-source workflow for tracking institutionally affiliated research data publications,” which is accepted for publication in the Journal of eScience Librarianship at the time of this deposit. The article will be available at the following DOI: https://doi.org/10.7191/jeslib.1170. These appendices are also available on the publisher’s website. 2 Appendix A. Case studies on metadata discrepancies and plasticity As mentioned in the main text, the workflow’s development was shaped by discovery of numerous metadata deficiencies and discrepancies at the repository level. Understanding the full history is not essential for re-using the workflow, but case studies involving four repositories (Dryad, Figshare, Mendeley Data, Texas Data Repository) are described here to provide context for anyone who may be interested in specific examples that informed this workflow. Texas Data Repository The first case study pertains to historic suboptimal formatting of affiliation metadata in the Texas Data Repository (TDR), UT Austin’s institutional data repository; at the onset of this project, TDR deposits were not detected via RORor single-permutation-string DataCite queries. Crossvalidation revealed that this was because affiliations for TDR datasets were being crosswalked to DataCite as a string enclosed in parentheses: ‘(University of Texas at Austin).’ This is an unexpected permutation of the institutional name (this formatting occurs in a few other repositories) and one that was not intentional per communications with TDR staff. This formatting explains, in part, the initial inability to discover TDR datasets through the DataCite API (the absence of a leading “The” being the other explanator), and because of the relative volume of TDR deposits, ‘(University of Texas at Austin)’ became the most commonly identified permutation across all affiliated datasets once the workflow was modified to retrieve them (see archived conference presentations by Shensky and Gee 2024; Gee 2025c). This crosswalk issue was raised with the TDR development team and resolved, with DataCite records for all existing datasets subsequently batch-updated prior to the most recent presentation in April 2025 (Gee and Shensky 2025). TDR deposits are now listed with ‘University of Texas at Austin’ (which 3 is still not the official permutation), and ‘(University of Texas at Austin)’ is now the 7th most frequently recovered permutation (Fig. 2). As discussed in the main text, it is hoped that the recently implemented functionality to add ROR identifiers in TDR will further improve the metadata quality. Dryad The second case study pertains to unexpected and erroneous changes to data metadata in Dryad. During this project, comparisons were made between repositories with respect to their annual publication volume of UT Austin-affiliated datasets through the use of the publicationYear field in DataCite. Early in development, Dryad exhibited an essentially flat pattern over the past decade, oscillating between a narrow range of about 35 to 45 annual datasets with at least one UT Austin researcher listed (Shensky and Gee 2024; Fig. 10). Prior to the most recent presentation in April 2025 (Gee and Shensky 2025), the workflow returned a peculiar pattern for Dryad, with a sharp drop-off to single-digit volume in 2024, but a 2025 volume that was already at the level of previous years (Fig. A1), despite being only a third of the way into the calendar year. This peculiar pattern was not detected for any other high-use repositories that were identified in the workflow. 4 Figure A1. Comparison of the number of Dryad datasets (for UT Austin-affiliated datasets and datasets across all affiliations) with a timestamp in a given year across three timestamps. availableYear and publicationYear (DataCite) are fields derived from the DataCite API (the former is extracted from a full timestamp; the latter is verbatim); publicationYear (Dryad) is derived from a full date field in the Dryad API. These dates should be congruent (similar to the availableYear and publicationYear (Dryad) fields), but there is a marked discrepancy in 2024 and 2025, with smaller discrepancies in 2019 and 2020. This dataset (Gee 2025b) omits file-level DOIs, which are retrieved from the DataCite API but not the Dryad API. Data as of July 1, 2025. 5 Examination of specific datasets in the DataCite API, the Dryad API, and DataCite Commons (web interface), as well as aggregate data for all Dryad datasets in both APIs, revealed that for most Dryad datasets published in 2024, the publicationYear field was, at some point, updated to list ‘2025.’ At least for UT Austin-affiliated datasets, other DataCite fields such as date.Issued and registered displayed full timestamps with the correct year. This appears unrelated to the updated field, which almost always lists a date in the current year, likely as part of the Make Data Count process (which requires regular updates to incorporate new metric data), and also appears unrelated to versioning since most other datasets published prior to 2024 have congruent timestamps in other fields and on the landing page (i.e. if publicationYear was affected by versioning, there should be a more even distribution of discordant dates across many years). Additionally, when examining the full corpus of Dryad datasets, there remain several hundred with a publicationYear of 2024 (Fig. A1). Across all Dryad datasets, nearly 13,000 had discordant publication year fields between the DataCite and Dryad APIs (Gee 2025b). There are two periods in which the annual count based on DataCite metadata is much higher than that based on Dryad metadata (2019–2020, 2025) and one period in which the inverse is true (2024), indicating that the discordance goes both ways (some datasets appear younger than they are, while others appear older than they are; Fig. A1). The publicationYear field is convenient for annual summaries since it does not require extracting the year from a full timestamp and is used by DataCite to create automatic summaries in its REST API and as a primary filter in DataCite Commons. In order to ensure accurate analysis, this workflow pivoted to use a different DataCite field from which year had to be extracted (registered), which is suboptimal because these fields should, in theory, be distinct 6 timestamps that may not be equivalent at finer timescales (e.g., month, day). This issue was brought to the attention of Dryad staff and resolved several months later in mid-July, but it potentially impacted other processes that involved retrieval of static Dryad records from DataCite in recent months. Discordance in date metadata is not exclusive to Dryad; there appears to be a common issue in old Figshare datasets that have been versioned in which one creation date listed for the parent DOI is the most recent date of publication (dates.date where dates.dateType=’Created’) and another (created) is the earliest date of publication (e.g., for Prokop 2013, with 65 versions between 2013 and 2023, both years appear in different creation date fields; dates for updated are also discordant), and likely similar patterns likely exist in other repositories but were not identified here. Some discordance over time is to be expected as repositories change their practices (although metadata discordance in general can sometimes be remedied by re-curation), whereas the issue with Dryad metadata was systemic but without a clear connection to a changed practice. Mendeley Data The third case study pertains to changes to affiliation recording and crosswalking in Mendeley Data. Based on examination of metadata for deposits affiliated with UT Austin, which is not, and has never been, an institutional member, Mendeley Data historically only crosswalked affiliation metadata to DataCite for institutional members. However, affiliation metadata was (sometimes) still recorded for authors at non-member institutions on dataset landing pages. Mendeley Data then appears to have started crosswalking all affiliation metadata to DataCite, regardless of institutional membership, in 2023; the timing suggests a connection to involvement in GREI. However, UT Austin-affiliated datasets published prior to 2023, of which there appear to be at 7 least several dozen to potentially several hundred based on a web interface search, still do not have affiliation metadata in DataCite, (e.g., Jayadev et al. 2020; House and Boehm 2021; Solomon-Lane et al. 2022). The data that can be obtained from the DataCite API will thus underestimate the total volume of affiliated Mendeley Data deposits and, if taken at face value, give a misleading impression that this repository was only used by affiliated researchers within the past few years. This case study is also an excellent demonstration of the importance of recuration for improving quality of already-published datasets’ metadata. Figure A2. Comparison of annual volume of affiliated Mendeley Data datasets as retrieved from the DataCite API. Figshare 8 The fourth case study pertains to affiliation metadata in Figshare and is multi-faceted, touching not only on metadata plasticity but also on re-curation, the value of data sharing, and a more common use case for data sharing (reproducibility). As part of the development of the secondary workflow to identify affiliated Figshare datasets that lack affiliation metadata, this study examined the dataset from the RADS study (Mohr and Narlock 2024), specifically the CSV versions (rather than the RData versions) of the files. Initially, the focus was on identifying which journals and publishers were associated with affiliated Figshare deposits (Table 5; Gee 2025b), and because the CSV files flatten any nested fields from the JSON response into ‘NA’ entries, this examination re-queried these deposits’ DataCite records through the DataCite API. Incidentally, when examining the new output, a significant number of deposits were detected that lack affiliation metadata matching any one of the six RADS institutions (Table A1); many included no affiliation metadata at all. Table A1. Summary of whether Figshare datasets previously reported with a RADS institution affiliation still include this linking metadata. Counts represent the retained records after removal of DOIs ending in ‘.v*’ (i.e. retaining only the ‘parent’ DOI). Matching was done using the institutional strings used in the RADS DataCite query. ‘Percent matched’ is essentially the percent of datasets that ‘should’ have been retrieved, assuming that all current DataCite affiliation metadata is accurate. *This count includes some duplication when a dataset was originally listed with multiple RADS institutions; it is counted for each institution. Institution Total count* Unmatched count Percent matched Cornell 298 136 54.3% Duke 91 58 36.2% 9 Michigan 185 123 33.5% Minnesota 112 56 50.0% Virginia Tech 141 129 8.51% Washington U 268 224 16.4% TOTAL 1,096 726 33.7% Multiple hypotheses and their resultant predictions were developed to explain the widespread lack of RADS affiliation metadata among purportedly affiliated Figshare deposits. 1. Hypothesis: The metadata discrepancies are related to the methodology of Johnston et al. (2024), either to the process of searching for DOIs in DataCite or in processing a DataCite output. a. Prediction: If a methodological flaw (e.g., capture of large numbers of ‘false positives’) is the cause, these discrepancies should also appear for other repositories at similar frequencies. 2. Hypothesis: The metadata discrepancies are related to Figshare, potentially related to the automation of dataset and metadata creation at scale through publisher partners. a. Prediction: If a Figshare-specific process is the cause, these discrepancies should be primarily restricted to Figshare. To this end, 4,000 dataset DOIs from the RADS dataset were randomly selected and queried for current affiliation metadata through the DataCite API. The code for this resampling uses a set seed [random_state] and thus can be independently reproduced (Gee 2025a), but the dataset is also included in the data publication (Gee 2025b). This workflow then looked for exact 16 drastically different summaries of annual publication volume (Fig. A1; Gee 2025b). The code for validation processes such as this (or the reanalysis of the RADS dataset with specific emphasis on purportedly affiliated Figshare deposits; Gee 2025a) is provided mainly for reproducing the specific results of this study. Another example is the provision of the script to calculate what subset of affiliated datasets can be recovered with either a ROR-based search in a single metadata field or a single-string-based search in a single metadata field (Gee 2025a); this process was known to be incomplete in retrieval from the beginning but was used to establish a baseline for increasing retrieval coverage. These processes often involved examination and comparison of specific datasets’ metadata in different platforms (e.g., comparison of publication date in the Dryad web interface, the Dryad API, the DataCite web interface [DataCite Commons], and the DataCite API); in most instances, these were randomly selected from datasets that were retrieved through the primary workflow, although in some instances, personal knowledge of a specific dataset (not necessarily affiliated with UT Austin) in a certain platform led to examination of other datasets. Having published several datasets in repositories myself, I was also able to examine how metadata that I had (or had not) provided with my deposits was reflected in the DataCite API. Similarly, for processes involving retrieval of article metadata through Crossref or OpenAlex, singular articles were targeted for examination of different metadata fields, some of which are my own publications. Finally, much of the exploration of Figshare metadata was drawn from the author’s personal experience as an active scientific researcher with several dozen peer-reviewed publications. For example, the integration between Figshare and some scholarly publishers like Springer Nature, Taylor & Francis, and PLOS in which files uploaded as ‘supplemental information’ are automatically published on Figshare upon publication of the associated article is 17 not well-known among researchers and even to editors, as no Figshare account is necessary for the depositor. My knowledge that this process exists and at least some aspects of how it functions (e.g., sometimes assigning each file a separate DOI) is based on first-hand experience as a researcher – I personally “published” eight Figshare deposits, mostly through Taylor & Francis titles (e.g., Gee et al. 2021, 2023), before discovering that this is how supplemental information is hosted by the publisher. In these titles in which I have published, there is no indication that this integration exists in the journal instructions, submission portal, or any part of the production process. I have never logged into Figshare, and none of my auto-created deposits record even basic author metadata such as my ORCID or affiliation at the time, so it seems unlikely that I would be able to claim, edit, or version any of these datasets to improve their quality (although I have noticed at least one instance in which either the journal or Figshare versioned the dataset without my knowledge let alone any request from me; compare Gee et al. 2019 and Gee et al. 2020). 18 Appendix C. Additional details on the cross-validation process As noted in the main text, the primary purpose of the cross-validation process is identification of additional permutations of a focal institution’s name that can be added to a multi-permutation query (excluding highly granular ones that are likely one-offs). This process thus mainly focuses on datasets that were discovered through a specific repository API but not from the DataCite API in order to improve the affiliation-based DataCite query (e.g., adding a new permutation of ‘UT Austin’). However, the inverse (datasets only found in DataCite’s API but not in a repository’s API) also occurs, albeit very rarely. The explanations for these edge cases are likely repositoryspecific and are often unclear, but they are noted here with possible explanations. The first example results from a lack of updated or crosswalked metadata for old deposits. For example, some dated deposits in Dryad (e.g., McTavish et al. 2015) and Zenodo (e.g., Dixon 2014) have affiliation metadata recorded on the landing page and in the repository API, but, as of this work, this metadata has not been crosswalked to DataCite. This is not unlike historic (pre-2023) datasets in Mendeley Data that lack a researcher from a member institution and for which no affiliation metadata is recorded. The second example is when a DOI appears to have been published but was then taken down, without any public-facing metadata label or explanation; rather than redirecting to the “DOI not found” page on doi.org (which typically indicates a lack of DOI activation), these DOIs redirect to a repository page indicating that the dataset was not found. This was observed for Dryad (e.g., Nandakumar et al. 2020a; Larter and Ryan 2023) and TDR (Hodges 2019a). For Dryad, it is presumed that these datasets were deaccessioned; this repository does not currently have a tombstone process. Curiously, there appears to be a dataset with nearly identical metadata and a public DOI (Nandakumar et 19 al. 2020b) for one of the non-functional Dryad DOIs (Nandakumar et al. 2020a). For TDR, the explanation is less clear (Dataverse installations have a tombstone process), but the unavailable dataset also appears to have a duplicate, publicly available dataset with nearly identical metadata in TDR (Hodges 2019b). The final example is a TDR dataset that is published and publicly accessible, but the DOI has not been activated (Bernard 2017). All instances of non-functional DOIs have been reported to their respective repositories and demonstrate the added utility of this workflow for identifying datasets in need of metadata re-curation. 20 Appendix D. Tested approaches to programmatic retrieval of mediated Figshare deposits with Crossref DOIs: the PLOS case study As noted in the main text, the sheer volume of mediated Figshare deposits with Crossref DOIs (millions) render a process like that developed for mediated Figshare deposits with DataCite DOIs impractical because of the time and API conditions (e.g., rate limiting allowances, server stability) that would be necessary just to retrieve all such objects (it could be more tractable with the public data file). Three different programmatic approaches were tested with PLOS specifically (not the only partner to mediate Figshare deposits with Crossref DOIs), but they are either very time-intensive, or they are not readily transferable to other publisher partners. The code for all three is provided as a proof of concept (Gee 2025a). The first approach involves identifying a list of affiliated articles published by a partner (PLOS), constructing a hypothetical SI file based on observed DOI construction patterns for that publisher (adding a ‘.s00*’ or ‘.t00*’ suffix to the article DOI), and pinging the URL to see if it exists. This process is rather time-consuming because it requires individual queries to test the existence of each hypothetical DOI, and even if a DOI does exist, no additional information is returned. The testing could be done with the Crossref API in order to obtain some metadata, but this is not recommended given the volume of API calls that would be necessary. All of the files associated with one article would have to be retrieved for collective assessment (e.g., SI files 1–19 could be accessory non-data files and file S20 could be a raw data file), and a means of assessing the content (‘nature’) of files from the very limited Crossref metadata remains elusive. The approach in particular is discouraged for use because of its time cost and is provided in the codebase only to demonstrate that this was explored. 21 The second approach is to use text-scraping (through the beautifulsoup4 module; Richardson 2025) to target the HTML blocks that contain information about SI files. For PLOS articles, this provides more information than the Crossref API because the HTML blocks contain additional, valuable information like the title of the deposit, which includes a value from a semicontrolled vocabulary (e.g., Figure, Table, Text); a description of the file; and the file format (e.g., DOCX, XLSX, PDF). Additional fields could either be scraped from the page or retrieved from the Crossref records for the article (e.g., author information) and collectively could be used, for example, to assess whether a given file or at least one file from a multi-file set for one article qualifies as ‘data.’ This process is more time-efficient than the first approach, but it is intrinsically sensitive to any future modifications to the HTML code and would require a custom approach for each publisher partner’s distinct HTML formatting. Whether other publisher partners who use Crossref for mediated Figshare deposits provide the same level of metadata for SI content in the web version of articles was not assessed. In early stages of testing, there was also some variation in how metadata were entered between PLOS titles. The third approach is extremely quick but very coarse-grained and involves the use of the PLOS Open Science Indicators (OSI) dataset (Public Library of Science 2024). This dataset, which is on version 10 (the workflow was tested on version 9) and under continuous development, was created through analysis of the entire corpus of PLOS articles and includes a categorization of where data were shared. The secondary workflow using this dataset retrieves a list of affiliated PLOS articles and then filters the PLOS OSI dataset for any affiliated article that has data shared in part or in whole through SI, which is in actuality the mediated Figshare process. This method assumes that the PLOS algorithms for identifying ‘data’ are philosophically aligned with those of the investigator and provides no additional information on the count, nature (e.g., file formats), or DOIs of these mediated deposits. It also relies on periodic versioning of the core dataset to be useful in the long-term and will quickly become slightly outdated after release. This approach mainly proved useful in early testing and exploration because it runs extremely quickly (< 60 seconds), and it could be sufficient for certain use cases (e.g., a quick approximation). Appendix E. Additional cleaning processes 22 As mentioned in the main text, there are likely to be additional cleaning steps for smaller-scale repositories, but these will likely be specific to each institution since they mostly involve specialist repositories with more heterogenous usage (or lack thereof) across disciplines and research institutions. Some examples for UT Austin are noted here, mainly to contextualize specific lines of code for these granular processes (Gee 2025a), some of which target individual datasets after manual examination of the outputs during workflow development. For the Environmental Molecular Science Laboratory (EMSL), it appears that individual but related samples are deposited with separate DOIs, as there are sets of datasets with identical metadata (authors, publication date, etc.), separate (but numerically sequential) DOIs, and the same type of data in each deposit (e.g., 15 deposits titled ‘Data for EMSL Project 47414 from March 2020’; Gao et al. 2020a, 2020b, 2020c, are cited here as examples). These were deduplicated based on a combination of metadata fields (dataset title, first author, relationType, relatedIdentifier, containerIdentifier). For DesignSafe, affiliation metadata is sub-optimally entered in the contributor.name field, rather than in an affiliation field, so it is not apparent which affiliation corresponds to which author(s). Additionally, because the repository is hosted at the Texas Advanced Computing Center (TACC) at UT Austin, all datasets list ‘University of Texas at Austin’ as a ‘contributor’ regardless of whether a UT Austin researcher was actually involved. Examination of random deposits suggested that when a UT Austin researcher was legitimately involved, the affiliation will specifically be ‘University of Texas at Austin (utexas.edu).’ This distinction was used to remove DesignSafe entries without this particular form of the affiliation. 23 Finally, listed repository names were corrected and standardized. For example, some Zenodo datasets do not list Zenodo as the publisher; as this is a freeform editable field in Zenodo (in contrast to many other repositories in which the repository is locked as the value in this field), any value can be entered in this field (e.g., Aziz Zanjani [2024] lists ‘The Seismic Record’ as the publisher). There are probably appropriate scenarios for entering a value other than the hosting repository, but for the purposes of this work, any deposit with ‘zenodo’ in the DOI string was changed to list Zenodo as the publisher. Reusers of this workflow will need to manually check their own outputs, as additional outliers will likely be unique to an institution (the ability for authors to edit the publisher field exists in other platforms; the Global Biodiversity Information Facility [GBIF] and Finnish FairData are other examples encountered here). Longer-tenured repositories are more likely to also have different permutations of their name listed in metadata records (e.g., ‘Texas Data Repository’ vs. ‘Texas Research Data Repository’), which were standardized here for identified repositories. The workflow also contains an optional additional processing of Dataverse deposits (these could be Harvard Dataverse for any institution or an institutional data repository on Dataverse software like TDR). I have observed that some authors create multiple ‘datasets’ in TDR, each labeled as ‘dataset’ in the Dataverse parlance (a DOI-backed container labeled as ‘dataset’ in DataCite), that all pertain to a single manuscript/article and that are nested under a single ‘dataverse’ (a non-DOI-backed container). Examples include five datasets (code, reconstruction parameters, system parameters, results, measurements) for one study on zebrafish (Yang 2024a, 2024b, 2024c, 2024d, 2024e); and four datasets (two types of data, experimental conditions, README) for one study on mechano-lysis in blood clots (Rausch 2025a, 2025b, 24 2025c, 2025d). Authors seem to separate materials based on format or an internal organizational delimiter (e.g., samples, specimens), rather than along metadata (e.g., different authorship lists, different licensing), and in contrast to the mediated Figshare process, they are not splitting each file into its own deposit (a standalone deposit for a README file may be an exception). Nonetheless, this practice can be considered to partially inflate the number of Dataverse ‘datasets’ because one could argue that all of the files would be in a single DOI-backed dataset in another repository that lacks an overarching container. The division of materials is probably the result of researchers taking advantage of the ‘dataverse’ container, which has no equivalent in most other repositories. This optional processing step consolidates objects with the same publication date, creators, and rights into a single entry. Although it is possible to retrieve dataverse metadata through the Dataverse API, some researchers create more expansive dataverses, beyond the level of one manuscript, that house multiple sets of deposits, and this step requires an additional API call. The resultant counts from this process can be drastically lower than the pre-consolidation count: if applied to the UT Austin datasets, the original 1,487 DOIs are consolidated into 949 (63%) entries. Because the observed splitting of materials in these instances is both logical and manual, and I have occasionally observed similar intentional splitting in other repositories, the results described in the main text do not incorporate this step. Finally, while developing rules for cleaning and consolidating deposits, duplicate publishing of the same dataset in different repositories was identified. Presently, three pairs affiliated with UT Austin are identified: TDR + Zenodo (Samineni 2022; Samineni and Kumar 2022); Dryad + ScienceDB (Wang et al. 2023, 2023b); and Harvard 25 Dataverse + Zenodo (Ishikawa et al. 2025a, 2025b). The reasons for this duplication are unclear, but with their identification, the researchers could be contacted to inquire about this practice. This duplication is different than paired Dryad-Zenodo deposits that share the same title and most other metadata fields and that will be retrieved if a workflow user includes resource types other than ‘dataset’ in the query. These related deposits result from a partnership in which a researcher submitting to Dryad has the option to create a linked Zenodo deposit for supplemental information and/or software (Lowenberg 2021). Linked Zenodo deposits will contain separate files, have different resource type labels, and be licensed differently, but other metadata attributes like title and authorship are shared with the Dryad deposit. 32 References Bernard, Rachel. 2017. “JGR BernardBehr2017 EBSD Data.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/SB9TOY. Dixon, Groves. 2014. “CpGoe_package.” Dataset. Zenodo. https://doi.org/10.5281/zenodo.12626. Gao, Fei, Ljiljana Pasa-Tolic, Donald Baer, et al. 2020a. Data for EMSL Project 47414 from March 2020. Dataset. Environmental Molecular Sciences Laboratory. https://doi.org/10.25582/data.2020-03.1970181/2390842. Gao, Fei, Ljiljana Pasa-Tolic, Donald Baer, et al. 2020b. Data for EMSL Project 47414 from March 2020. Dataset. Environmental Molecular Sciences Laboratory. https://doi.org/10.25582/data.2020-03.1970180/2390841. Gao, Fei, Ljiljana Pasa-Tolic, Donald Baer, et al. 2020c. Data for EMSL Project 47414 from March 2020. Dataset. Environmental Molecular Sciences Laboratory. https://doi.org/10.25582/data.2020-03.1970179/2390843. Gee, Bryan M. 2025a. “Code for: The Hunt for Research Data: Development of an Open-Source Workflow for Tracking Institutionally Affiliated Research Data Publications.” Software. Zenodo. https://doi.org/10.5281/zenodo.18037530. Gee, Bryan M. 2025b. "Data for: The Hunt for Research Data: Development of an Open-Source Workflow for Tracking Institutionally Affiliated Research Data Publications.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/R9MSCP. Gee, Bryan. 2025c. “Development of an Open, Automated Workflow for Gaining Bibliometric Insights into University Research Data Publishing.” Presentation. Research Data Access and 33 Preservation (RDAP) Annual Meeting, Virtual. OSF, March 13. https://doi.org/10.17605/OSF.IO/T659S. Gee, Bryan, and Michael Shensky. 2025. “Completing the Picture of Institutional Research Output and Impact: Automating Discovery and Assessment of Research Data and Software.” Presentation. Coalition for Networked Information (CNI) Spring Meeting, Milwaukee, WI. April 7. https://www.cni.org/topics/information-accessretrieval/completing-the-picture-of-institutional-research-output-and-impact-automatingdiscovery-and-assessment-of-research-data-and-software. Gee, Bryan M., William G. Parker, and Adam D. Marsh. 2019. “Redescription of Anaschisma (Temnospondyli: Metoposauridae) from the Late Triassic of Wyoming and the Phylogeny of the Metoposauridae.” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.8192291.v1. Gee, Bryan M., William G. Parker, and Adam D. Marsh. 2020. “Redescription of Anaschisma (Temnospondyli: Metoposauridae) from the Late Triassic of Wyoming and the Phylogeny of the Metoposauridae.” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.8192291.v2. Gee, Bryan M., David S. Berman, Amy C. Henrici, Jason D. Pardo, and Adam K. Huttenlocker.2021. “New Information on the Dissorophid Conjunctio (Temnospondyli) Based on a Specimen from the Cutler Formation of Colorado, U.S.A.” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.14184953.Gee, Bryan M. Charles V. Beightol, and Christian A. Sidor. 2023. “A New Lapillopsid from Antarctica and a Reappraisal of the Phylogenetic Relationships of Early Diverging Stereospondyls.” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.23591802. 34 Hendricks, Ginny, Dominika Tkaczyk, Jennifer Lin, and Patricia Feeney. 2020. “Crossref: The Sustainable Source of Community-Owned Scholarly Metadata.” Quantitative Science Studies 1 (1): 414–27. https://doi.org/10.1162/qss_a_00022. Hodges, Ben R. 2019a. “Outtxt_bump_0065_baseline.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/0UTCS3. Hodges, Ben R. 2019b. “Outtxt_bump_0065_baseline.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/B7GU5P. House, Christopher, and Christoph Boehm. 2021. “Data for: Optimal Taylor Rules When Targets Are Uncertain.” Dataset. Mendeley Data, April 26. https://doi.org/10.17632/8yp2dmc5kk.1. Ishikawa, Hiroyuki Tako. 2025. “Luhman 16AB IGRINS Spectral Atlas.” Dataset. Zenodo. https://doi.org/10.5281/zenodo.15001025. Ishikawa, Hiroyuki Tako, Stanimir Metchev, Megan E. Tannock, et al. 2025. “Luhman 16AB IGRINS Spectral Atlas.” Dataset. Harvard Dataverse. https://doi.org/10.7910/DVN/TKY3KC. Jayadev, Gopika, Benjamin Leibowicz, and Erhan Kutanoglu. 2020. “Data for: U.S. Electricity Infrastructure of the Future: Generation and Transmission Pathways through 2050.” Dataset. Mendeley Data. https://doi.org/10.17632/cjcfw8b4xf.1. Johnston, Lisa R., Alicia Hofelich Mohr, Joel Herndon, et al. 2024. “Seek and You May (Not) Find: A Multi-Institutional Analysis of Where Research Data Are Shared.” PLOS ONE 19 (4): e0302426. https://doi.org/10.1371/journal.pone.0302426. Larter, Luke, and Michael Ryan. 2023. “Female Preferences for More Elaborate Signals Are an Emergent Outcome of Male Chorusing Interactions in Túngara Frogs.” Dataset. Dryad. https://doi.org/10.5061/dryad.7d7wm37zs. 35 Li, Jia. 2020. “Visualization1-FALCON.Avi.” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.11956734.v1. Lowenberg, Daniella. 2021. “Doing It Right: A Better Approach for Software & Data.” Blog. Dryad News, February 8. https://blog.datadryad.org/2021/02/08/doing-it-right-a-betterapproach-for-software-amp-data/. McTavish, Emily Jane, Jared E. Decker, Robert D. Schnabel, Jeremy F. Taylor, and David M. Hillis. 2013. “Data from: New World Cattle Show Ancestry from Multiple Independent Domestication Events.” Dataset. Dryad. https://doi.org/10.5061/DRYAD.42TR0. Mohr, Alicia Hofelich, and Mikala Narlock. 2024. “DataCurationNetwork/Rads-Metadata: Article Acceptance.” Dataset. Zenodo. https://doi.org/10.5281/zenodo.11073357. Nandakumar, Nagaraja, John Forder, Steven Warach, and Jose Merino. 2020a. “Reversible Diffusion-Weighted Imaging Lesions in Acute Ischemic Stroke: A Systematic Review.” Dataset. Dryad. https://doi.org/10.5061/dryad.qv9s4mwb1. Nandakumar, Nagaraja, John Forder, Steven Warach, and Jose Merino. 2020b. “Reversible Diffusion Weighted Imaging Lesion in Acute Ischemic Stroke – A Systematic Review.” Dataset. Dryad. https://doi.org/10.5061/dryad.mpg4f4qvp. Prokop, Andreas. 2013. “2nd Year Drosophila Developmental Genetics Practical.” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.156395.v65. Public Library of Science. 2024. “PLOS Open Science Indicators.” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.21687686.v9. Rausch, Manuel. 2025a. “Experimental Conditions.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/POACZN. 36 Rausch, Manuel. 2025b. “Measures of Lysis.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/WBBBND. Rausch, Manuel. 2025c. “Mechanical Data.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/R7XZMJ. Rausch, Manuel. 2025d. “ReadMe.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/OQXGVP. Samineni, Laxmicharan. 2022. “Data Repository Moringa Filber Filters npJ Clean Water.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/BKRUCG. Samineni, Laxmicharan, and Manish Kumar. 2022. “Data Repository Moringa Filber Filters npJ Clean Water.” Dataset. Zenodo. https://doi.org/10.5281/zenodo.6607397. Shensky, Michael, and Bryan Gee. 2024. Generating Institutional Level Open Science Metrics with Scripted Processes for Tracking Publication of Research Data and Software. Poster. Southeast Data Librarian Symposium (SEDLS), Virtual. OSF, November 7. https://osf.io/9e4v3. Solomon-Lane, Tessa, Hans Hofmann, and Rebecca Butler. 2022. “Vasopressin Mediates Nonapeptide and Glucocorticoid Signaling and Social Dynamics in Juvenile Dominance Hierarchies of a Highly Social Cichlid Fish.” Dataset. Mendeley Data. https://doi.org/10.17632/9rhgvznb87.1. Strecker, Dorothea. 2025. “How Permanent Are Metadata for Research Data? Understanding Changes in DataCite DOI Metadata.” Preprint, arXiv, December 6. https://doi.org/10.48550/arXiv.2412.05128. Teague, Richard. 2019a. “HD 163296 Rotation Maps.” Dataset. Harvard Dataverse. https://doi.org/10.7910/DVN/C2ZUNO. 37 Teague, Richard. 2019b. “TW Hydrae Rotation Maps.” Dataset. Harvard Dataverse. https://doi.org/10.7910/DVN/KXELJL. Wang, Xuezhao, Yunyun He, Brian E. Sedio, et al. 2023a. “Phytochemical Diversity Impacts Herbivory in a Tropical Rainforest Tree Community.” Dataset. Science DB. https://doi.org/10.57760/sciencedb.10798. Wang, Xuezhao, Yunyun He, Brian E. Sedio, et al. 2023b. “Phytochemical Diversity Impacts Herbivory in a Tropical Rainforest Tree Community.” Dataset. Dryad. https://doi.org/10.5061/DRYAD.N2Z34TN32. Yang, Siqi. 2024a. “Reconstruction Parameters.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/BAM64U. Yang, Siqi. 2024b. “System Parameters.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/ZNGIZR. Yang, Siqi. 2025a. “Code.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/LFYAIO. Yang, Siqi. 2025b. “Measurements.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/KJW2UG. Yang, Siqi. 2025c. “Reconstruction Results.” Dataset. Texas Data Repository. https://doi.org/10.18738/T8/M26D5E. Zhang, Xuan, Jing Li, Bang-Zhen Pan, et al. 2021. “Additional File 11 of Extended Mining of the Oil Biosynthesis Pathway in Biofuel Plant Jatropha Curcas by Combined Analysis of Transcriptome and Gene Interactome Data.” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.15253711.v1. 38 Zheng, Guoqiang, Xiaoyun Dong, Jiaping Wei, et al. 2022. “Additional File 3 of Integrated Methylome and Transcriptome Analysis Unravel the Cold Tolerance Mechanism in Winter Rapeseed(Brassica Napus L.).” Dataset. Figshare. https://doi.org/10.6084/m9.figshare.20652829.v1.