scieee AI-readable full text Open interactive document viewer

Refractive datasets as a sensemaking methodology in closed data ecosystems

Beers, Anna; Gildersleve, Patrick; Aragón, Pablo; Tripodi, Francesca

Abstract

Data, code, and supplement for:Beers, A., Ito, V., Orozco, A., Gildersleve, P., Aragón, P., & Tripodi, F. (2025). Refractive datasets as a sensemaking methodology in closed data ecosystems. Big Data & Society, 12(4). https://doi.org/10.1177/20539517251406193 (Original work published 2025). The following files are uploaded:1. RefractiveDatasets_Code.py. This is commented Python code for reproducing analyses in this paper. It assumes that any data downloaded will be present in a folder titled Data located one level up from where you run the script ("../Data"). Some data can be regenerated via this code; other data is provided with an explanation of how to independently retrieve it in the comments. 2. Figure2Data.csv. Time series data for Figure 2, representing Wikipedia traffic across different Wikipedia projects. 3. Figure3Data.csv. Time series data for Figure 4, representing the median of five data pulls of Google Trends traffic for a given keyword. 4. GoogleTrends_DataSamples.zip. Zipped CSV files corresponding to five unique data downloads from Google Trends used to construct Figure3Data.csv. 5. GoogleTrendsAnalysisSupplement.pdf. A supplemental analysis of the robustness of Google Trends data portrayed in Figure 3. 6. Figure45_HurricaneClusterArticles.csv. A list of articles found in the "Atlantic Hurricanes" cluster identified in Figures 4 and 5. 7. Figure45_ClusteredData.zip. A compressed GEXF network file representing clustered correlation relationships between articles portrayed in Figures 4 and 5. 8. Figure45.gephi. A Gephi visualization file for the network visualization portrayed in Figures 4 and 5.

Full text

Beers et al. 1 Supplement: Refractive Datasets as a Sensemaking Methodology in Closed Data Ecosystems Journal Title XX(X):2–6 ©The Author(s) 2025 Reprints and permission: sagepub.co.uk/journalsPermissions.nav DOI: 10.1177/ToBeAssigned www.sagepub.com/ SAGE Anna Beers,1Viviane Ito,1Agustin Orozco,1Patrick Gildersleve,2 Pablo Aragón,3and Francesca Tripodi1 Prepared using sagej.cls Abstract As digital platforms restrict their APIs, researchers face diminishing options for studying social phenomena in digital environments. During what has been called the post-API era, researchers have found themselves looking for reliable data sources in an unreliable and frequently changing platform data ecosystem. In this context, we propose analyzing refractive datasets as a methodology for researchers to understand the dynamics of closed data platforms. Refractive datasets come from platforms with relatively more open data policies, and their analysis sheds light on platforms with more restrictive data policies. Like a prism, refractive datasets reflect but also transform data-based phenomena unfolding on closed platforms. Using refractive datasets from Wikipedia and Google Trends, we present three studies to demonstrate our methodology. We first show how refractive data from Wikipedia’s multiple language editions can be used to understand a fractured global platform ecosystem in a case study of hydroxychloroquine, a purported COVID-19 medicine. Second, we use Google Trends to show how similar refractive analyses can be used to recover information lost to platform deletion, in a profile of an online panic over the drug brand Galaxy Gas. Finally, we show how Wikipedia data can be used as a grounding point for a refractive analysis of how new generative algorithms reproduce and distort data across the social web. We discuss how refractive datasets can be a way for researchers to “sensemake” in increasingly opaque big data environments, enabling interpretivist analyses which aim to generate new hypotheses rather than verify existing claims. Keywords computational social science, data access, wikipedia, cross-platform dynamics, mixed methods, transnational information dynamics 1University of North Carolina at Chapel Hill, School of Information and Library Science, USA 2University of Exeter, Department of Communications, UK 3Universitat Pompeu Fabra, Department of Information and Communication, ES Corresponding author: Anna Beers Prepared using sagej.cls [Version: 2017/01/17 v1.20] Beers et al. 3 Supplement Validating Google Trends Data The academic use of Google Trends data has been a matter of some controversy (Lazer et al.,2014;Hölzl et al.,2025). At least two factors fuel this controversy. First, Google Trends data is a sample of Google’s search traffic data, and that sample is re-generated on each day that Google Trends data is queried. Second, various idiosyncrasies of Google Trends’ data presentation and history, such as its presentation in scaled rather than absolute traffic values, can lead to errors in analysis. Prior research has shown that studies that use Google Trends data rarely provide sufficient information to allow for reproducibility, or even consider in their analysis that, e.g., Google Trends data is sampled anew potentially with each query. To respond to these critiques, we follow recommendations from the "checklist" offered in Hölzl et al. (2025) to mitigate (but not eliminate) risk in the analysis of Google Trends data. We first define the rationale for our keyword selection and the corresponding construct we wish to measure. In Case Study 2, we wish to measure the salience of the Galaxy Gas phenomena to internet users in the United States during a time period in which external reporting claims it was popular (Holtermann,2024). To do this, we use Google Trends to measure the popularity of the Google search term "galaxy gas" in Google’s "United States" region from July 1st to November 1st 2024. The Google Trends interface suggests many other "related" queries to "galaxy gas," such as question formats pertaining to Galaxy Gas ("what is galaxy gas"), more general phenomena pertaining to nitrous oxide ("nitrous oxide," "whippets"), and internet memes associated with Galaxy Gas ("lil t"). In addition, Google provides two automatically-generated search categories relating to Galaxy Gas, one referring to the Galaxy Gas "Topic," and another to the Galaxy Gas "Company." These categories combine multiple search queries in a non-transparent way, supposedly related to the automatically-generated label. We find that simple question-formatted queries relating to Galaxy Gas follow the same qualitative pattern observed in Case Study 2 as the simple "galaxy gas" search term (lowercase, as shown), making their use redundant. More general terms, such as "nitrous oxide" or "whippets," show a similar qualitative pattern, but appear to be confounded at certain moments by unrelated stories of celebrities abusing nitrous oxide. Queries related to memes associated with Galaxy Gas, such as "lil t," show idiosyncratic behavior, with greater prevalence in the earlier part of the case study, when we predict Prepared using sagej.cls 4Journal Title XX(X) Figure 1. A line chart depicting daily Wikipedia traffic data from July 1st to November 1st, 2024 for several pages relating to Galaxy Gas. Point A refers to unrelated traffic peaks relating to celebrity use of nitrous oxide (but not Galaxy Gas). Point B refers to the publication date of a New York Times article documenting the Galaxy Gas phenomenon. that the Galaxy Gas phenomenon was spreading widely on social media platforms via influencers. Finally, the automatically-generated Galaxy Gas "Topic" follows the simple "galaxy gas" query closely, while the "Company" category has no observable traffic, suggesting that the latter category is non-functional. Given these comparisons, we find that the simple "galaxy gas" query is a reasonable search term to use within Google Trends. To be more confident that our metric, Google Trends data on the "galaxy gas" search query, corresponds to our construct of interest, we can also compare our data with another dataset used in this study: Wikipedia page traffic data (Figure 1). Wikipedia page traffic data is an inexact comparison, as the page for "Galaxy Gas" was only created near the end of September 2024, when the brand was receiving widespread mainstream media attention. Reassuringly, traffic to the "Galaxy Gas" Wikipedia page tails off in the same way that our Google Trends data for "galaxy gas" does once entering October 2024. We can also compare traffic data for related Wikipedia pages, "Nitrous oxide," "Recreational use of nitrous oxide," and "Whipped-cream charger," to our Google Trends data. We find that this data replicates broadly the slow increase of traffic into late September that we find in our Google Trends data, but is marked by early differences that we attribute to certain popular news stories about celebrities using (non-Galaxy Gas) nitrous oxide in August 2024. Prepared using sagej.cls Beers et al. 5 Figure 2. A line chart depicting daily Google Trends traffic data from July 1st to November 1st, 2024 for the query "galaxy gas." Different color lines represent different samples of the same query on different dates, per the legend. To be sure that we are not missing confounding factors in our data, we compare the "galaxy gas" search query to two other queries representing control cases: one for the search query "cocaine," another illicit substance, and one for the search query "lebanon," which was a region heavily reported on by the New York Times in September 2024. We find that neither query resembles the traffic pattern found in "galaxy gas," providing supporting evidence that the query is not simply representing increased interests in illicit substances broadly, or following underlying seasonal trends in media consumption. We replicate other findings that Google Trends data changes when queried from day to day. This can make analyses of particularly low-volume search queries problematic, as trends or "spikes" in a given dataset can turn out to be illusory. To reassure ourselves that such sampling variation is not affecting our analysis, we sample Google Trends data at 5 different dates from two locations (the United States and Germany) without a VPN while logged into personal accounts. In our findings, we use the median value found from these five samples. The trends and peaks observed in this data do not substantially change depending on which sample we use. Figure 2shows these five samples overlaid on top of one another. The maximum difference on Google Trends’ normalized 0-100 scale for any given time point between these five samples is 3. Given that our interpretations do not change depending on our sample, and that the overall variation displayed is low, we are reassured that we are not interpreting noise. The final recommendation from Hölzl et al. (2025) is to reflect on the generalizability and specificity of this dataset. We do not claim in Case Study 2 that we have identified Prepared using sagej.cls 6Journal Title XX(X) a strict pattern of information transmission that generalizes to all internet-mediated phenomena. That said, there are some specificities to Google Trends data that may limit this analysis’s applicability. For example, it may be that users encountering Galaxy Gas via social media platforms are more likely to use the search engines on those platforms, rather than Google, and thus that the portion of search traffic owed to social media engagement in this dataset is suppressed. We encourage more research into how refractive search engine data may distort comparisons of volume between different platform populations flattened into one dataset. References Holtermann, C. (2024). What Is Galaxy Gas, and Why Are Young People Inhaling It? The New York Times. Hölzl, J., Keusch, F., and Sajons, C. (2025). The (mis)use of Google Trends data in the social sciences - A systematic review, critique, and recommendations. Social Science Research, 126. Lazer, D., Kennedy, R., King, G., and Vespignani, A. (2014). The Parable of Google Flu: Traps in Big Data Analysis. Science, 343(6176). Prepared using sagej.cls