Code and Data for: Reproducibility and Open Science in Economics
Full text
Reproducibility and Open Science in Economics Lars Vilhuber a Drawing on my experience as the American Economic Association’s data editor, I examine the current state of open science in economics as facilitated by and related to reproducibility. I touch on the tension between accessibility, sharing, and preservation. The guiding theme is the accessibility of the key ingredients for scholarship : manuscripts, data, software, and the necessary technology to combine the latter two in order to produce knowledge. I analyze how economic research balances openness with necessary restrictions, particularly regarding administrative and confidential data. I argue that a large degree of openness is nevertheless present, with extensive networks that include thousands of researchers supporting collaborative science. I argue that resource constraints, such as software licensing costs and computational resource requirements, pose similar challenges. I illustrate concrete benefits of open science in the economics literature, using recent articles. I wrap up by discussing the state of access to scientific articles in economics. Mots clés : sciences ouvertes; reproductibilité; accès à resources Keywords: open science; reproducibility; ressource access Classification JEL : A11; A13; B40. a. Cornell University, [email protected] 1
Reproducibility and Open Science in Economics 1. INTRODUCTION As a graduate student in economics at Université de Montréal, reading the economics literature was easy. While the main university library had all the relevant subscriptions, our department librarian, Fethy Mili, would populate the library of the economics department with multi-hued rows of working papers. Mili was also one of the key creators of what was initially known as Working Papers in Economics ( WoPEc ) and BibEc (for printed working papers) (Cruz et al. [2000]; Krichel [1997]; Krichel et Zimmermann [2009]), populating the latter since 1993 (Bátiz‐Lazo et al. [2012], p. 450). The overall network, known as Research Papers in Economics ( RePEc ), was born contemporaneously with the more widely known arXiv (Ginsparg [2011]) and the more centralized Social Science Research Network ( SSRN ) (Social Science Research Network [2025]). While electronic working papers were mostly free in those days (there was no way to pay for them), Mili’s work consisted of sending out postage-paid envelopes to all of the various economics departments that were publishing the working papers, and then cataloging the incoming printed materials electronically, for public consumption. In the end, information about the existence of the working papers was freely available, but access to (printed) working papers still required a small fee to cover the cost of shipping. I also experienced the openness of code sharing, with code samples by prominent authors being available to graduate students, though discovery was much more difficult at the time. The Statistical Software Components ( SSC ), primarily but not exclusively for STATA packages, appeared in 1998 (Cox [2010]; Cox et Jenkins [2022]), providing a convenient and open way to catalog, distribute, and provide open access to additional Stata functionality. 1 Data sharing was harder, of course, with lots of floppies 2 being exchanged, but also the use of departmental FTP 3 servers (David Card’s collection of data, or the NBER’s), or even the replication archive of the Journal of Applied Econometrics, instantiated at Queens University in 1994 under the long-running guidance of James McKinnon. 4 On the other hand, administrative data, such as the French administrative data used in AKM required travel to Paris, sitting in a room without windows at an assigned time, and typing the code into the system that had access. 5 When data were available, computing was straightforward : You logged on to the university’s big computer, running some variant of Unix, and used whatever 1 . This kind of functionality was inspired by similar functionality available for other software, like CPAN for Perl and CRAN for R, but not for most statistical software used by economists. 2. https://en.wikipedia.org/wiki/Floppy_disk 3. https://en.wikipedia.org/wiki/File_Transfer_Protocol 4 . The JAE archive was migrated to the ZBW’s archives in 2022 and can now be found at https: //journaldata.zbw.eu/journals/jae , but legacy files are still visible as of 2024 at http: //qed.econ.queensu.ca/jae/legacy.html. 5 . For non-local authors, this meant traveling for extended time periods as a Ph.D. student (Margolis) or spending a sabbatical in Paris (Abowd), neither of which is a cheap endeavor. Both were, and still are, enjoyable, though. 2Revue économique – vol. , no, , p. 1-33
Lars Vilhuber software was available. Software licenses were paid by the university, as was the computing hardware itself. Laptops powerful (and light!) enough to do actual work were only then emerging. In 2025, there are concerned discussions about the cost of publishing academic articles, of accessing those same academic articles, of the ever increasing use of administrative data (Card et al. [2010a,b]; Chetty [2012]; Einav et al. [2014]) that would appear to be hidden behind insurmountable access restrictions, the use of “proprietary software”, and the increasing use of large computing infrastructure, all of which would seem to be restricting access to the basic elements of conducting research in economics. In this article, I will draw on my experience working on many aspects of increasing access to data and materials of all kinds, in particular my recent experience as the inaugural data editor of the American Economic Association ( AEA ) (Duflo et al. [2018]), to paint a picture of economics in an era of open science. How different are matters in practice now compared to that early view of the field of economics, back when I was a graduate student? In this article, I will discuss the current state of open science in economics as facilitated by and related to reproducibility. I will touch on the tension between accessibility, sharing, and preservation, and some of the approaches that are being implemented, sometimes tentatively, in economics, and sometimes elsewhere. My view will be biased - I am an active participant in this space, primarily via my current appointment as data (and reproducibility) editor of the American Economic Association, but also as a past participant in networks that have and foster access, and a researcher and editor in the space of disclosure limitation. The guiding theme will be the accessibility of the key ingredients for scholarship : manuscripts (or more generally, documents), data, software, and the necessary technology to combine the latter two in order to produce knowledge as published in manuscripts. My focus will be on the latter three, though I will provide some observations about scholarly publishing in the last section. In the conclusion, I will identify a few areas where there is (continued) movement towards greater openness. 2. CONCEPTS In order to write about “Open Science,” a definition is needed. Open science is a surprisingly difficult term to define precisely, and multiple overlapping definition are commonly referenced. UNESCO [2022] sees four components to open science : open scientific knowledge (publications, data, code, and teaching materials “openly available, accessible and reusable for everyone”), open science infrastructures (which encompasses both physical infrastructure such as instruments and laboratories, as well as virtual components such as open access publication platforms), science communication (knowledge translation), and broad engagement beyond the boundaries of the academy. It also recognizes the limitations of such access in a caveat : Revue économique – vol. , no, , p. 1-33 3
Reproducibility and Open Science in Economics ... human rights, security, personal privacy, ... In such cases, it may still be possible to share the existence of such information or share it among certain users who meet defined access criteria. The Open Knowledge Foundation (Open Knowledge Foundation [2024], OKF) defines “open” as (my emphasis) ... anyone can freely access, use, modify, and share for any purpose (subject, at most, to requirements that preserve provenance and openness). Less broadly, Vicente-Saez et al. [2018] identify a consensus that defines open science as “transparent and accessible knowledge that is shared and developed through collaborative networks.” In this article, I will focus on the what UNESCO 2022 calls open science “knowledge” and will briefly discuss “infrastructures.” I will highlight how some elements have been quite widespread in economics for some time. I will try to identify limits to fully open accessibility, some of which are intrinsic to the nature of the research conducted in economics, and describe how widespread such limitations may be. In particular, I will highlight how those access restrictions are not, as many think, an impediment to open science, in the sense that aforementioned “collaborative networks” can still access these resources. A key ambiguity will arise in how big such networks need to be in order to be considered “open.” Consider the realized size of several relevant networks in economics. The National Bureau of Economic Research ( NBER ) defines its affiliated scholars as a network : 𝑛 = 1804 as of January 2025, primarily in North America (National Bureau of Economic Research [2025]). However, in 2024, a total of 2,966 authors published 1,223 NBER working papers. J-PAL has approximately 𝑛 = 1725 affiliates at 120 universities on all populated continents (Abdul Latif Jameel Poverty Action Lab [2025]). Between 2001 and June 2024, there had been 𝑛 = 2023 researchers on projects that used confidential U.S. Census Bureau in the Federal Statistical Research Data Centers ( FSRDC ) (U.S. Census Bureau [2024]). In an average year (2013-2023), 𝑛 = 1238 students graduate from a U.S. university with a Ph.D. in economics (National Science Foundation [2024], Table 1-5). In the approximately 30 years since inception of RePEc , 𝑛 = 526 authors from 52 institutions in Norway have published a paper listed on RePEc (presumably in economics) (IDEAS//RePEc [2025]). All of these are measured across different spatio-temporal dimensions. Are they large? Context and purpose matter. Some may intersect. How many U.S. graduate students in the past 10 years have published an NBER working paper and are now at a Norwegian institution? It is harder to measure the actual key criterion : How many people can potentially enter the network, become part of it, or, as will be the key use of this concept later, how many people can potentially access data of certain types, accessible via some sort of network. The actual size, relative to the number of potential entrants, is only a proxy for that, and often, a detailed analysis of entry criteria is necessary. For instance, what does it take to enter the network of research data centers (in the US, in France, in Canada, etc.) in order to be able to access the same data as others. Finally, it may be relevant to measure how diverse these networks are. 4Revue économique – vol. , no, , p. 1-33
Lars Vilhuber This matters for the ability of insiders of a certain network to criticize each other. For this, size is not a good measure — even 𝑛 = 2 may be sufficient. Journals play a key role in this space, and will be an important background to my discussion and experiences, possibly also my biases. Most top economics journals have a data (and possibly code) availability policy. 6 The AEA’s policy was first implemented in 2005 (American Economic Association [2005]; Bernanke [2004]). While the focus of early availability policies was on the data, the code often came along for the ride, albeit not always in its most complete form. I am the AEA ’s inaugural data editor, appointed in 2018 (Duflo et al. [2018]). The AEA implemented a new policy in 2019 (American Economic Association [2019]). Many other economics journals appointed data editors around that time, and multiple journals coordinated on a common policy core called Data and Code Availability Standard ( DCAS ) (Koren et al. [2022]), and revised their policies to align with DCAS , see American Economic Association [2024] for the AEA. 7 A key part of these newer policies was increased pre-publication monitoring of the content of replication packages (Christian et al. [2018]; Duflo et al. [2018]). One way to start to move away from a binary perspective of access is to consider time as summary metric that captures what is needed in order to access generic resources, whether data, manuscripts, or computing resources. Time might be needed to write an application to access data, or time might be needed in order to obtain access to large-scale computing resources. Time might be needed in order to obtain grant funding that allows to purchase such resources. I choose time, rather than money, as the metric, since it might appear to be slightly more egalitarian, given that much of science has (theoretical) access to subsidies and grants. In the other dimension, the number of people who have some probability of accessing the resource (the size of the network) can be taken as an approximate measure of openness, regardless of how interesting or valuable they might consider the data to be. Figure 1, taken from Vilhuber [2023], serves to illustrate this idea, for access to data, with various institutions that facilitate that access mapped out into the space of time vs. size of network. I will return to this throughout the discussion. 3. DATA ACCESS One subcomponent of open science, and locus of much attention throughout the literature in the social sciences, are “open data.” This in principle easy - why should the data used in research not be open? However, the various caveats that 6 . For a review of the history of data and code availability policies in economics, see Vlaeminck [2021]. 7 . These initiatives are not restricted to economics, of course. Political science (Basile et al. [2023]; Data Access & Research Transparency (DA-RT) [2014]; Jacoby [2015]), sociology (Sociological Science [2018]; Weeden [2023]), and general initiatives like the Transparency and Openness Promotion (TOP) guidelines (Nosek et al. [2015]). Revue économique – vol. , no, , p. 1-33 5
Reproducibility and Open Science in Economics Figure 1 – Conceptual trade-off between number of individuals accessing data, and time required to do so. Figure first published in Vilhuber [2023]. policies and principles include are important to recognize. (Open Knowledge Foundation [2024], OKF) mentions “requirements that preserve provenance and openness,” which does not take into account privacy. UNESCO [2022] does note “human rights, security, personal privacy.” On the other hand, even much data that is available to almost anybody on the Internet may not actually be “open.” Consider the S&P 500, viewed in newspaper and many websites (e.g. S&P Dow Jones Indices LLC [2025]), is not “open data” because it does not allow for free re-use. OKF defines “open data” as requiring machine readability, absence of licensing charges, and free re-use, but does not mandate availability via download on the internet, absence of all fees, nor absence of any technical measures, such as a requirement to register and agree to abide by these rules (Open Knowledge Foundation [2024]). I will discuss two sub-areas within this space : Secondary data use, and primary data generation. Much of economic research uses data collected by others, such as survey organizations and national statistical offices, but also private company data and various administrative data sources (“organic data”, Groves [2011a,b]). Primary data generation is more frequent in behavioral and development economics. Secondary data : The data produced by the United States government are in the public domain (i.e., without any restrictions on usage or attribution), which makes it “open” in the above sense (Copyright Act of 1976, Wikipedia [2025]). Many countries have switched their government data to default to open data (Statistics Canada [2012]; UK Government [2014]). However, many well known “public-use” data are not always “open” : IPUMS has a redistribution restriction 6Revue économique – vol. , no, , p. 1-33
Lars Vilhuber (encapsulated in “Terms of Use”, not a license), and some geographic data by international statistical offices remain under more stringent licensing requirements in other countries (e.g., United Kingdom). Many well-known survey organizations impose redistribution and usage restrictions that are not consistent with the OKF definition. Many such redistribution restrictions apply to datasets with more detailed personal information collected through surveys, and are meant to ensure continued compliance with ethical rules of behavior, often with informed consent agreements by survey participants and local privacy laws. Notable examples include PSID (Institute for Social Research [2024]), World Value Survey (Haerpfer et al. [2024]), Demographic and Health Surveys ( DHS ) Program (DHS Program [2024]), German Socio-Economic Panel (Goebel, Grabka, Liebig, Kroh et al. [2019]; Goebel, Grabka, Liebig, Schröder et al. [2024]), all of which have broad (cost-free) usage, conditional on registration and compliance with usage restrictions. 8Table 1 shows a few examples. Tableau 1 – Example restrictions Dataset Variant Usage.restrictions.and.justification DHS No redistribution and access by unauthorized individuals. To ensure that provider meets host-country agreements, such as compliance with usage by real people with legitimate research purposes. WVS Except for joint WVS-EVS files The data can be used freely for non-commercial purposes such as research, publication, teaching. Data redistribution is prohibited : publication of the original WVS datasets at other online platforms is against the WVSA Constitution. PSID Public data Use the data in the PSID data sets only for scientific research and aggregate statistical reporting; Make no attempts to identify study participants; PSID Restricted data In order to safeguard the confidentiality of respondents at the highest level, some data are provided only under conditions of a restricted use contract, including human subjects review and data security plan. SOEP European version No redistribution and access by unauthorized individuals. Compliance with German Federal Data Protection Act. SOEP International version Same as European version, but 95% sample, and institutional agreement required. The steps needed to obtain access range from click-through agreements to acknowledge compliance with CC-BY licenses (consistent with open access definitions) to the need to write a paragraph about the purpose on how the data is to be used (maybe consistent), to various commitments to not redistribute the data, all while being cost-free. The first three rows in Table 1 apply some variant of the latter : While anybody can obtain access, the restriction on redistribution is not consistent with definitions of open access. Yet there are no impediments to actually using the data in research. The last three rows of Table 1, however, go further. Researchers using SOEP data must satisfy certain geographical requirements, such as presence of the researcher in the countries covered by General Data Protection Regulation ( GDPR ), 8. I come back to the case of the DHS Program in Conclusion. Revue économique – vol. , no, , p. 1-33 7
Reproducibility and Open Science in Economics otherwise they can only access a modified version of the data. Similar restrictions apply to the restricted-use data of many other surveys, for instance, the US-based National Longitudinal Survey of Youth ( NLSY ). Users of restricted PSID data must comply with even more stringent rules : secure approval from their institution’s ethics board, and use of secure computing environments. Clearly, these restrictions will inhibit broader use of the data per se. Yet these restrictions are not globally very restrictive : There are in Economics alone 1238 new US-based researchers being given the ability to request access to geo-restricted NLSY data every year : newly minted Ph.Ds, as noted earlier. A similar number in the EU likely obtain the right to access the geo-restricted SOEP data as well. Despite the restrictions on the use of SOEP data, there are currently 6,443 papers that in some fashion have used the data. 9 The PSID bibliography 10 lists over 1,300 dissertations and over 5,300 articles. In the space depicted in Figure 1, access is very much towards the left of the figure. 11. Organic data : However, most economists, when asked about data subject to restrictions, will think of “proprietary data,” a term often applied to any data that may be subject to restrictions of use and re-use. Access to most administrative data is typically not “open” in the sense of OKF, but are they open enough, given the privacy concerns that are attached to these data? Many have argued that access is not broad enough (Card et al. [2010b]; Einav et al. [2014]), while acknowledging the difficulty of addressing privacy and security of the data at scale. J. M. Abowd et I. M. Schmutte [2018] discuss the challenge of making the choice of between accuracy of (public) statistics and data, and the privacy loss inherent in doing so. Nagaraj, Shears et al. [2020] argue that more openness improves scientific progress, and Nagaraj et Tranchero [2023] study this in the context of the US system for providing access ( FSRDC ). They note that 4% of US-based empirical authors have had some access to the FSRDC system. I am not aware of similar studies for other countries, such as France (the equivalent system is the Centre d’accès sécurisé aux données ( CASD ), Gadouche [2019]) or Canada (Currie et al. [2015]). In absolute terms, these networks host several hundred researchers every year. For the 781 projects using Census Bureau data in the FSRDC , 2,084 researchers have had access to confidential data between 1998 9 . Source : https://www.diw.de/en/diw_01.c.789503.en/publications_based_on_ soep_data__soeplit.html accessed on 2025-02-08. 10 . Source : https://psidonline.isr.umich.edu/publications/Bibliography/ search.aspx accessed on 2025-02-08 11 . In fact, authors sometimes forget to abide by the rules for these datasets. As AEA Data Editor, I get notified via “take-down requests” from data providers 12-15 times per year, including in 2024 from the PSID, and have posted information on how to achieve compliance, at least in some cases, at https: //aeadataeditor.github.io/posts/2024-11-01-psid-requests . Most of the cases affect papers published prior to my tenure, because I do alert authors to data use agreement violations that I am aware of. Ultimately, however, it is the authors’ obligation to remain compliant with such data use agreements. 8Revue économique – vol. , no, , p. 1-33
Lars Vilhuber and 2024. 12 Nagaraj et Tranchero [2023] mention 861 papers in scientific journals. Similar numbers can be obtained for the French (6841 researchers from 1109 institutions on 1797 projects with 417 publications) 13 and Canadian (2201 active researchers from 42 universities as of January 2025, with 3245 papers published between 2000 and May 2024) 14 networks. 15 The number of publications from access to these networks is smaller in absolute terms then those from PSID and SOEP, though likely higher in impact (Nagaraj et Tranchero [2023]). 16 Primary data collection : The discussion of choices made by survey organizations should in principle be applicable when smaller teams of economists, not entire survey organizations, do the primary data collection. Similar to survey organizations, such teams have to balance the privacy of their respondents with the benefits of open science, in particular the broader knowledge to be gained from open access to the data. Many, so it would seem, provide much of the data in replication packages, subject to de-identification (see Bjarkefur et al. [2021]; Kopper et al. [2020], for examples), though typically not with stronger disclosure avoidance measures similar to those employed by statistical agencies and larger survey institutions (for a brief discussion of the issues and one possible solution, see Mukherjee et al. [2023]). For research teams, ethics boards and institutional review boards ( IRB s) have a role to play (Grant et al. [2019]), with some arguing very strongly that greater availability to others (though not blind publication of all data) is required in order to maximize the societal benefits that are the quid pro quo for the respondents’ consent to their privacy being invaded (Grant et al. [2019]; Meyer [2018]). Making such data as broadly available, while respecting the privacy of respondents, is precisely what open access to such data promises, modulo appropriate access restrictions or data use agreements similar to those outlined in Table 1. In general, however, primary data collections do not have access to robust third-party systems that would allow for access similar to the access required by PSID and similar organizations, situated between no access and fully public access. Thus, while access may be requested in ad-hoc fashion via the original authors, this is known to be fraught with problems (Gabelica et al. [2022]; Watson [2022]). An ideal scenario would see researchers deposit the data they collected in third-party repositories, which then handle issues such as verifying ethics approval and secure access mechanisms. Some full-service repositories, such as Inter-university Consortium for Political and Social Research ( ICPSR ) 12 . Own calculations based on Census Bureau data, see https://labordynamicsinstitute. github.io/fsrdc-external-census-projects/. 13 . From https://www.casd.eu/ and https://www.casd.eu/ toutes-les-publications/liste/0/20/ as of May 2025. 14. Provided by Grant Gibson, CRDCN, on February 10, 2025. 15 . The Nagaraj et Tranchero [2023] number only includes papers published by economists. Other numbers are counts of researchers and publications in all disciplines, in non-peer-reviewed publications, and include non-economists. 16 . For an analysis of code availability over time for SOEP-based publications, see Fink et al. [2025]. Revue économique – vol. , no, , p. 1-33 9
Reproducibility and Open Science in Economics institutions, including in lowand middle-income countriess ( LMIC s), may well not have the funds to purchase proprietary software, but access to computers may be equally constraining. The template README requests information on the type of computer that was used by the original researcher, to provide a benchmark to future re-users. Acquiring access to sufficient memory (random access memory ( RAM )), storage, and use over time of those resources can be expensive, even when renting such resources in cloud environments (which very few researchers appear to be doing). Traditionally, that access may be embedded within a single purchased computer, which may have (in 2024) around 32GB of RAM, 1-2 TB of storage, and have 4-12 compute cores available exclusively to the owner. More complex analyses may require access to shared compute clusters (using hundreds or thousands of compute cores), very large storage arrays (measured in the twoto three-digit TB range), and may require up to 1024 GB of RAM. Cutting edge analyses may require specialized chips, such as one or more graphical processing units ( GPU s), or even a cluster of GPU s. I have observed analyses that may run data cleaning or data acquisition processes for months at a time. The vast majority of articles published in economics journals usually require no more computing resources than a modern laptop provides, in all the dimensions enumerated in the previous paragraph. In fact, a formal quantitative measurement of resource usage in economics articles is surprisingly hard to obtain, as most researchers are not very good at reporting the resources they have used to conduct their research. In part, this is because measuring such usage is non-trivial, but to a larger extent, I postulate that this is because most research institutions provide such resources to their researchers in a “convenient way,” and researchers conduct research within those constraints. More importantly, however, it suggests an important constraint on how “open” access can be for some if not all economics research. Some newer research requires vastly different types of resources. Studies using raw satellite data may require more than 10TB of data storage (Khachiyan, Thomas, Zhou, Hanson, Cloninger, Rosing et A. Khandelwal [2022]; Khachiyan, Thomas, Zhou, Hanson, Cloninger, Rosing et A. K. Khandelwal [2022]), may need more than 20,000 compute hours on a cloud provider (Rudik [2020a,b]), or the use of one (Dell [2024, 2025]) or dozens (Khachiyan, Thomas, Zhou, Hanson, Cloninger, Rosing et A. Khandelwal [2022]) GPU s. 24 Access to the code and data for the papers mentioned is open : The AEA-related replication packages for these articles are licensed under a standard Creative Commons Attribution ( CC-BY ) license. Some of the data not included in the Dell [2025] replication package is on Huggingface, also under a CC-BY license (Silcock et al. [2024]). Open access to satellite data is one of the canonical examples of the benefits of open access (Nagaraj, Shears et al. [2020]). In these cases, the computational resources may restrict the benefits of the open access of data and code. 24 . As of January 2025, the type of GPU used by Dell [2025], costs between USD 4500 and USD 7700, or between 2 and 7 times as much as a standard laptop. 16 Revue économique – vol. , no, , p. 1-33
Lars Vilhuber Are such computational constraints a problem? No consistent analysis exists that correlates resource requirements to academic outcomes such as citations, primarily because it is very hard to measure consistently the resource requirements of economic articles. The very small sample in the previous paragraph may serve to illustrate this, but without controls for scientific merit, is purely an indicator. Silcock et al. [2024] had been downloaded 98 times in December 2024, six months after the arXiv paper associated with it was published (Silcock et al. [2024]). Rudik [2020a] has had 1432 views, 124 downloads for replication package, as the manuscript (Rudik [2020b]) has 15 citations. The replication package Khachiyan, Thomas, Zhou, Hanson, Cloninger, Rosing et A. Khandelwal [2022] has 2124 views, 175 downloads, while the manuscript (Khachiyan, Thomas, Zhou, Hanson, Cloninger, Rosing et A. K. Khandelwal [2022]) has 5 citations (all as of January 2025). For comparison, the average article in one of the AEA ’s journals has 908 views and 106 downloads (Vilhuber [2025], Table 4). 6. THE BENEFITS OF OPEN SCIENCE IN ECONOMICS Tthe discussion about benefits of open science often centers around data availability. The World Bank, in its annual World Development Report, identifies data availability — for research, for commerce, for education — as a key contributor, and highlights that many LMIC continue to have impediments to the reliable provision of open access data (World Bank [2021], pg. 62). The (theoretical) optimal level of data availability intersects with privacy, making the optimal level of data availability a non-trivial balance between public and private benefits, and private costs (J. M. Abowd et I. M. Schmutte [2018]; J. M. Abowd, I. M. Schmutte et al. [2019]; Acquisti et al. [2016]; Duch-Brown et al. [2017]). The more recent discussions (and court cases) surrounding the use of data in the training of large language models have only re-emphasized this tension (Panettieri [2025]). A different thread in the discussion brings up normative reasons for transparency, for instance around Mertonian (Merton [1942]) norms of openness (see Miguel [2021], for an overview). Openness at all stages of research may act as a moderator for publication bias (Brodeur et al. [2016]; Miguel [2021]) Some of the benefits of open science have been measured indirectly, through increased citations, say. Some recent studies find some advantages for studies with linked (openly available) data (Christensen et al. [2019]; Colavizza et al. [2020]; Piwowar et al. [2013]). 25 The economics literature has emphasized the benefits of (balanced) open science. Nagaraj, Shears et al. [2020] discuss the canonical example of improved data access through (free) public-use data in the context of satellite imagery. Patents can be usefully investigated, since they are both openly viewable but also access-limiting by their very nature. Economists 25 . Some of these studies need to be taken with a grain of salt, since they typically measure whether data is referenced, not whether it is actually openly available. Revue économique – vol. , no, , p. 1-33 17
Reproducibility and Open Science in Economics Figure 5 – Extracts from Ferguson et al. [2023, Figures 1 and 7], reused under CC-BY 4.0. have looked at how restrictiveness, duration, and type of patents affect scientific progress. In general, restrictions reduce social welfare (Murray et al. [2016]; Williams [2013]), even enabling anti-competitive behavior that directly identify welfare loss (Xie et al. [2020]). The much broader applicability of open science practices is also much more recent, and as of yet, hard to measure. Nevertheless, it appears to be widely accepted in economics, as evidenced by very strong positive attitudes documented in Ferguson et al. [2023], as selectively depicted in Figure 5. The top left part of Figure 5 shows behavior (black) and opinions (color scale) in regards to data sharing among economists, with more than 50% of economists having shared data, but over 90% being very much or moderately (the two green colors) in favor of sharing data. The right panel illustrates the evolution over time, depicting the proportion of social scientists in the four disciplines (economics, political science, psychology, and sociology) who had adopted an open science practice as of a particular date, showing again the rapid increase in active data sharing amongst economists from around 60% in 2011 to the 2020 number of over 90%. I want to add to this discussion of benefits three concrete case studies that advance science in economics, while relying heavily on the openness of prior research. The three articles in question leverage the open availability of code and data, with licenses that allow for re-use, to improve econometric methods for future researchers. The first paper relies on the empirical recomputation of prior papers to assess a theoretically ambiguous potential bias in inference. The second paper also focuses on inference, and uses actual data from previous papers to simulate the relevance of the impact. Both then provide new software (R and Stata packages), under open source licenses, to “fix” the problem for future studies. 18 Revue économique – vol. , no, , p. 1-33
Lars Vilhuber The third study selects studies again where data are available, and leverages the centralized availability of such replication packages to select the studies in their sample. Roth [2022] uses 12 previously published papers to assess whether the usual tests for pre-existing differences in trends when using difference-in-differences methods are properly controlling for power, and the empirical impact on subsequent inference. The theoretical bias is ambiguous, so an empirical evaluation is necessary. The study both leverages the open availability of materials in economic journals, but also illustrates the limitations imposed by imperfect adherence to openness. To wit, Roth writes that he searched for “the phrase “event study” in papers published in the American Economic Review, American Economic Journal : Applied Economics, and American Economic Journal : Economic Policy between 2014 and June 2018 ... The search returned 70 total papers that include a figure that the authors describe as an event-study plot.” but continues to then be limited by lack of data in the majority of cases : “I exclude 43 papers for which data to replicate the main event-study plot were unavailable. (Roth [2022], pg. 307)” 26 While it remains unclear whether the excluded papers are non-compliant with the AEA’s policy at the time, or whether they have legitimate reasons not to provide the data (Roth does not provide the raw result of his search), the paper is able to make an important methodological point (723 / 415 citations as of January 2025, per Google Scholar/ OpenAlex) because it is able to fully recompute the results in previous papers, apply new tests and methodologies, and come to meaningful recommendations and tools — Roth provides an (open source) R package to implement his methodology. Chaisemartin et Ramirez-Cuellar [2024] use open access information on RCTs (AEA registry) to find 15 RCTs of a particular type (clustered paired or small strata), of which 4 have publicly available data (and reproducible artifacts). They then provide results both on simulations using these data, and in particular, reestimate the regressions used in those studies and apply their proposed solution, showing that the number of significant effects is reduced by one-third. In other work (Chaisemartin et D’Haultfœuille [2020, 2024]), the results from various other papers are also recomputed to empirically demonstrate the relevance of the proposed methods, and software packages (e.g. Chaisemartin, D’Haultfoeuille et al. [2025]) are developed and made openly available. 27 Chaisemartin et D’Haultfœuille [2020] has been cited between 2,600 and 4,600 times (OA, GS). Finally, Goldsmith-Pinkham et al. [2024] investigate contamination bias (“each treatment’s effect are contaminated by nonconvex averages of the effects of other 26. He also excludes another 15 papers for reasons not related to data availability. 27 . Note that as of January 2025, many of the packages do not have an explicit open source license — or any license — applied, a common feature of economists working in the open source world. My presumption is that they simply assume that everybody knows that the code is openly available. Revue économique – vol. , no, , p. 1-33 19
Reproducibility and Open Science in Economics treatments”). To do so, they search the centralized repository of AEA replication packages 28 to identify packages that, crucially, contain data. They then reproduce one of each original paper’s specifications, conduct several tests, and conclude that “economically and statistically meaningful contamination bias [is present] in two of the three observational studies while showing no evidence for bias in any of the experimental studies.” (pg. 4046). Goldsmith-Pinkham et al. [2024] has been cited between 65 and 118 (OA, GS) times. 29 In each of these examples, the ability to access prior data, code, and information is critical to improving future scientific progress. In some cases, it remains limited by both historical and unavoidable limitations on openness, as well as scaling limitation based on absence of “push button” reproducibility. 7. OPEN INFRASTRUCTURE : PUBLICATIONS The challenges of openly accessible written scholarship are manifold, with the current focus on “Plan S”, master publication agreements, and in the US, similar efforts under the moniker of the “Nelson memo” (Brainard [2024]; Brainard et Kaiser [2022]). I note that the economics profession has a very long history of making much of the written knowledge available at very low cost via working papers (Vilhuber [2020b]), with the first working papers at the reputable NBER working paper series going back to 1973 (Welch [1973]). Over the past eight years, I have or have had three editorial appointments. I am the Data Editor for the journals of the American Economic Association ( AEA ) (Duflo et al. [2018]), a column editor for the open access Harvard Data Science Review ( HDSR ) (Vilhuber, I. Schmutte et al. [2023]), and until January 2025, the joint executive editor for the open access and multi-disciplinary Journal of Privacy and Confidentiality ( JPC ), for which I continue to manage the publication infrastructure (J. Abowd et al. [2025]). I will use each of these to highlight a particular pattern in broadening access to publications, without any claim to generality. The AEA is a not-for-profit organization, as are many other learned societies. It self-publishes eight journals, plus the proceedings of the annual conference, without relying on a commercial publishing house. Depending on the measurement, three of these publications are in the top ten journals in economics (Mogstad et al. [2022]). Its publication costs account for about half of its overall operating expenses, and are only partially offset by directly attributable subscription and membership fees (Cherry Bekaert, LLP [2024]). In fact, 6 of the top 10 journals in economics (Mogstad et al. [2022]) are published by societies (JEL, JEP, 28 . “These studies were identified by a systematic search of papers in the AEA Data and Code Repository.” [pg. 4043] 29 . Google Scholar citations are directly reported from a view of the article’s page on Google Scholar, and may include citations to multiple versions. OpenAlex citations are the sum of citations to all recorded works on OA with the same title and by the same authors. 20 Revue économique – vol. , no, , p. 1-33
Lars Vilhuber Econometrica, AER, Restud, JOLE), some of which have as sole or primary purpose the publication of the journal. 30 A further three journals are primarily associated with economics departments (QJE, JPE, RESTAT), which arguably may not be driven by pure profit. The sole outlier in the top ten is the Journal of Financial Economics ( JFE ), which is owned by Elsevier, a big commercial publisher. It should be noted that the European Economics Association ( EEA ) severed its relationship with Elsevier in 2003 for its official journal, creating a journal that is fully owned by the association, adding to the list of society-owned journals in economics (Tirole et al. [2003]). Access to these journals is generally still on a subscription basis (JEP is the exception, being free to read), but given the primarily not-for-profit organization of its owners, personal subscriptions (often via society membership) are quite low, compared to journals in many other sciences. For instance, as of 2024, a personal subscription to the Review of Economic Studies is $156 or €141 per annum; a yearly membership to the AEA , providing access to the seven subscription journals and the proceedings is $25 for students and researchers in low-income countries, and $150 at the highest personal income tier. As outlined earlier for the AEA, these subscription fees cover only a small fraction of the production costs. Nevertheless, even these (arguably low) costs do not satisfy “Plan S” or “Nelson memo” requirements, which require no access cost to the end consumer, and in the case of “Plan S”, also require a liberal license allowing for re-use. 31 Interestingly in the context of the previous sections, all of the society-owned journals in the previous paragraph have appointed data (reproducibility) editors. 32 From 2018 to January 2025, I was the executive editor of the JPC , an open access multi-disciplinary journal, having taken over the journal from Stephen Fienberg (Vilhuber [2018]). 33 The journal does not charge submission fees, but in 2025, to cover costs, it started charging publication fees. It remains free to read. Articles default to a Creative Commons Attribution-NonCommercial-NoDerivatives (CCBY-NC-ND) license, though authors are allowed to choose a more liberal license, for instance to comply with “Plan S” (which does not allow for the “no-derivatives” part). As executive editor, I was responsible for all aspects of running the journal, not just finding referees for the articles that I am responsible for. The journal is made available through open-source software called Open Journal System, hosted by its creators at Simon Fraser University’s Public Knowledge Project, preserved via industry-standard mechanisms (CLOCKSS, a non-profit) in case 30 . JOLE is a bit of an outlier, in that one becomes a member of the Society of Labor Economists by subscribing to the journal, rather than the other way around. 31 . “The author(s) [...] grant(s) [...] a free, irrevocable, worldwide, right of access to, and a license to copy, use, distribute, transmit and display the work publicly and to make and distribute derivative works, [...] subject to proper attribution of authorship” (Max-Planck-Gesellschaft [2023]). 32 . Two additional societies, not previously mentioned, also employ data editors : the Canadian Economics Association (CEA) and the Western Economics Association International (WEAI). 33 . Fienberg, together with Cynthia Dwork and Alan Karr, founded the journal in 2009 (J. M. Abowd, Nissim et al. [2009]). Fienberg passed away in 2016 (Slavković et al. [2018]). In 2023, I recruited Rachel Cummings to jointly manage the journal. Revue économique – vol. , no, , p. 1-33 21
Reproducibility and Open Science in Economics the journal ever needs to shut down, indexed in a variety of academic indexes, including via assignment of DOI . Copy-editing is done through a mixture of professional copy-editors and volunteer work by editors and board members. Until 2024, all editors, including myself, were unpaid, and referees are, like in much of the publishing industry, unpaid volunteers. Yet I did pay bills, for each of the above components of a properly managed, indexed, and preserved academic journal — and professional copy-editors and university staff do not work for free. I am thus quite aware of the absolute minimum cost of running a (small) journal. Over the years, funding has come from a variety of chaired professorships at Carnegie Mellon (Fienberg), Cornell (Abowd), and Harvard (Dwork). In order to make such funding more robust, a non-profit society was created to better and more robustly structure the funding situation (J. M. Abowd, Dwork et al. [2024]), and publication fees were implemented – a step back, to some extent, on the prior “complete” open access model. Time will tell if this will stabilize the funding situation, while maintaining the foundational commitment to open access. Others, in particular Sociological Science, have shown that it is feasible to sustainably publish high-quality research 8. CONCLUSION I have described in this article how open data and code are in the academic literature in economics, and how access to software and hardware resources can be limiting factors. Almost all code is openly accessible in top journals. The vast majority of data is accessible with little to no effort, and a large proportion of the remaining restricted data can be accessed by networks that include thousands of researchers. I provide a few concrete examples where the openness of the data and code available allows others to directly build on prior results. Broader assessments of the benefits of open access are more difficult to measure, in part because the right controls are hard to construct, in part because the community is still only starting to learn how to technically leverage the openness in a large-scale fashion. Nevertheless, access to networks of data access and financing remains one of the key worries. I am a regular participant in discussions within three research networks that provide access to restriced data ( FSRDC , Canadian Research Data Center Network ( CRDCN ), and CASD ). Core discussions center around equity and access, and how to balance those criteria while preserving the privacy of the respondents for which these networks act as curators. One under-appreciated aspect of open access is that it enables persistence. For any data that is subject to a gatekeeping mechanism, however objective, impartial, and lightweight it may be at the present time, such mechanisms can and do disappear. Every one of the networks that I mention above has a mandate to work within budget constraints, and those budgets are determined fundamentally by external forces, typically government-based funding agencies. A particular striking example, as of the writing of this article, is access to data from the USAIDfunded DHS program. Data access to DHS data was classified as ‘moderately easy to obtain’ (as per Table 2) until February 2025, as it took only a day or two 22 Revue économique – vol. , no, , p. 1-33
Lars Vilhuber to obtain access subject to a lightweight data use agreement. My team at the AEA regularly went through the process to obtain data used by other researchers, a process that was easy to navigate even for an undergraduate researcher on my team. In February 2025, the second Trump administration shut down USAID and “paused” the DHS program, with very little notice, and no recourse. While the DHS program system was still accessible to those with prior access permission in late February 2025, no further access requests were accepted. That turns it into a ‘very difficult to obtain’ dataset. 34 These events are a note of caution that any kind of redistribution restriction may very well negatively impact future availability to the research community. Whereas journals can subscribe to mechanisms that allow past issues to remain accessible even when the journal is shuttered, 35 and open access software can be preserved by communities with an interest in continued use, 36 no such mechanism exists for most restricted-access data. While open code may ensure we can recompute, and advances in computational infrastructure will bring the current cutting edge into the space of everyday-accessibility, data that are not open can and will disappear. That is concerning. Annexe Tableau A1 – Access categories and whether data can be shared privately Access restrictions No Yes Total Percent No n/a 50 Very Easy to Obtain 11 7 18 (38.89%) Moderately Easy to Obtain 3 5 8 (62.50%) Moderately Difficult to Obtain 16 9 25 (36.00%) Very Difficult to Obtain 9 5 14 (35.71%) Any restriction 39 26 65 (40.00%) *Percentages are calculated as the number of ’Yes’ divided by the total number of responses. An article can have multiple categories of data; the sum of responses is therefore higher than the number of articles. 34 . As of March 2025, efforts are underway to preserve the DHS program data, for instance via IPUMS, and to resuscitate both the access to historical data, as well continue funding new surveys. 35. See CLOCKSS. 36. See f.i. Data Rescue Project. Revue économique – vol. , no, , p. 1-33 23
Reproducibility and Open Science in Economics Inferring Stata and R usage Stata usage is inferred from downloads of Stata packages from the SSC web server, which is the sole official location to obtain these packages. Private mirrors may exist, and not all Stata packages are installed from SSC - both the Stata Journal and Github are likely to be significant sources. Data are obtained from log files from the SSC web server, provided by Kit Baum. R usage is inferred from downloads of R binaries (for Windows and MacOS). While this is likely to be less frequent than Stata packages, relative patterns are of interest. There is no easy way to obtain the full list of downloads for all packages, other than cycling through several thousand such packages. The data stem from one of dozens of Comprehensive R Archive Network ( CRAN ) mirrors, though this one, managed by Posit PBC, is the first one listed in the list of mirrors that are offered to users. Data in both cases is for February 2025. As a first step, these downloads were mapped to countries, and then aggregated by regions (Table A2 for Stata, Table A3 for R). Mapped onto a world map, custom_region countries Regional Downloads Fraction 1 China 2 5438811 50.79 2 North America 3 3509194 32.77 3 Europe 40 912383 8.52 4 Rest of Asia 45 521557 4.87 5 Africa 51 168663 1.57 6 Latin America & Caribbean 28 127078 1.19 7 Australia 1 26703 0.25 8 Rest of Oceania 5 4763 0.04 Tableau A2 – Regional Statistics for Stata Downloads custom_region countries Regional Downloads Fraction 1 North America 3 438562 66.62 2 Europe 50 83367 12.66 3 Rest of Asia 48 72724 11.05 4 Latin America & Caribbean 42 25316 3.85 5 China 1 18641 2.83 6 Africa 51 10124 1.54 7 Australia 1 8166 1.24 8 Rest of Oceania 11 1406 0.21 9 Other 1 9 0.00 Tableau A3 – Regional Statistics for R Downloads this provides a pretty picture, though not very informative (Figure A1). The data suggest that downloads from China are not fully captured by the Posit-managed mirror, possibly because of peculiarities of the Chinese internet infrastructure. While this is possible for other countries as well, it is not fully detectable. I have 24 Revue économique – vol. , no, , p. 1-33
Lars Vilhuber Figure A1 – Downloads of Stata packages and R software, by region, without China. Source : published data by Kit Baum and Posit PBC, own calculations. Revue économique – vol. , no, , p. 1-33 25
Reproducibility and Open Science in Economics Tirole J. et al. [mars 2003], « Editorial », Journal of the European Economic Association, 1 (1), p. iii-iv. U.S. Census Bureau [2024], uscensusbureau/fsrdc-external-census-projects at 2625c2169f2d1b22cac262b90f7af87b7e969d6b, Github. UK Government [2014], Open Government Licence for public sector information V3. UNESCO [nov. 2022], Understanding open science, UNESCO. Vicente-Saez R. et Martinez-Fuentes C. [juill. 2018], « Open Science now : A systematic literature review for an integrated definition », en, Journal of Business Research, 88, p. 428-436. Vilhuber L. [déc. 2018], « Relaunching the Journal of Privacy and Confidentiality », en, Journal of Privacy and Confidentiality, 8 (1), Number : 1. – [mai 2020a], Migrating historical AEA supplements, Cornell University. – [déc. 2020b], « Reproducibility and Replicability in Economics », Harvard Data Science Review, 2 (4). – [juin 2023], « Reproducibility and transparency versus privacy and confidentiality : Reflections from a data editor », en, Journal of Econometrics, 235 (2), Published online, p. 2285-2294. – [mai 2025], « Report of the AEA Data Editor », en, AEA Papers and Proceedings. Vilhuber L., Connolly M. et al. [nov. 2022], A template README for social science replication packages, Publisher : Zenodo, Zenodo. Vilhuber L., Schmutte I., et al. [July 2023], “Reinforcing Reproducibility and Replicability: An Introduction”, Harvard Data Science Review, 5 (3). Vilhuber L., Turitto J. et Welch K. [mai 2020], « Report by the AEA Data Editor », en, AEA Papers and Proceedings, 110, p. 764-75. Vlaeminck S. [août 2021], « Dawning of a new age? Economics journals’ data policies on the test bench », en, LIBER Quarterly : The Journal of the Association of European Research Libraries, 31 (1), p. 1-29. Watson C. [juin 2022], « Many researchers say they’ll share data — but don’t », en, Nature, 606 (7916), Bandiera_abtest : a Cg_type : News Publisher : Nature Publishing Group Subject_term : Funding, Ethics, p. 853-853. Weeden K. A. [oct. 2023], « Crisis? What Crisis? Sociology’s Slow Progress Toward Scientific Transparency », en, Harvard Data Science Review, 5 (4). Welch F. [juin 1973], Education, Information, and Efficiency, National Bureau of Economic Research. Wikipedia [jan. 2025], Copyright status of works by the federal government of the United States, en, Page Version ID : 1270780906. Williams H. L. [fév. 2013], « Intellectual Property Rights and Innovation : Evidence from the Human Genome », en, Journal of Political Economy, 121 (1), p. 1-27. World Bank [2021], World Development Report 2021 : Data for Better Lives, EN, World Bank. Xie J. et Gerakos J. [mai 2020], « The Anticompetitive Effects of Common Ownership : The Case of Paragraph IV Generic Entry », en, AEA Papers and Proceedings, 110, p. 569-572. 32 Revue économique – vol. , no, , p. 1-33
Zillow [2021], Zillow’s Assessor and Real Estate Database, en-US.