Full text
Social Scientific Data Quality and Reproducibility in the AI Era: Challenges and Pathways Stefan Dietze1,2 Session moderator: Anna Jacyszyn3 Interdisciplinary Colloquium on Digitalisation of Research, DiTraRe, 4 December 2025 (1) GESIS Leibniz Institute for the Social Sciences + (2) Heinrich Heine University Düsseldorf (3) FIZ Karlsruhe - Leibniz Institute for Information Infrastructure
Photos and recording Pixabay, ste_phania 2 Interdisciplinary Colloquium on Digitalisation of Research, Stefan Dietze, 4 December 2025 www.youtube.com/@DiTraRe
3 Interdisciplinary Colloquium on Digitalisation of Research, Stefan Dietze, 4 December 2025
Social scientific data quality and reproducibility in the AI era: challenges and pathways DiTraRe Interdisciplinary Colloquium on Digitalisation of Research, 4 December 2025 Stefan Dietze
Social science research is changing ▪Emergence of large volumes of behavioral data (e.g. from social media) has introduced new research field (CSS), methods and data 2
Behavioral web data for the social sciences ▪Online discourse (e.g. in social media, online news) ▪Social web activity streams (posts, shares, likes, follows etc) ▪Web search behaviour, e.g. browsing, navigation or search engine interactions ▪Low-level behavioral traces (scrolling, mouse movements, gaze behavior etc) ▪General characteristics oClose to users & their personal (potentially sensitive) information oLarge and heterogeneous 3
Web data tends to be „big“ Source: Domo via PCMag 4
New kinds of data require new kinds of methods Methods widely used (e.g. for social media analysis) : ▪Time series analysis (auto-regressive models, ARIMA etc) ▪Network/graph analysis ▪Dictionary-based methods (e.g. for sentiment analysis) ▪Tailored machine learning models (trained from scratch) ▪Pretrained open source language models (e.g. BERT) ▪Pretrained proprietary LLMs (like GPT/ChatGPT) Substantial differences with respect to: ▪Scalability (ability to handle larger volumes of data) ▪Robustness (ability to handle noisy or biased data) ▪Efficiency (compute/resource requirements) ▪Transparency & interpretability ▪Reproducibility „AI“ 5
Beyond basic use of AI for data analysis: LLMs for simulating human behavior Santurkar, S., et al., Whose Opinions Do Language Models Reflect?, International Conference on Machine Learning (ICML2023) 6
Beyond just reproducibility Pineau et al., Improving reproducibility in machine learning research, Journal of Machine Learning Research 22 (2021) 1-20. 17
Beyond reproducibility: do benchmarks assess generalisable learnings? Example: Twitter bot detection Chris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk, and Philipp Zimmer. 2023. Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot Detection. ACM WebConf2023 18 „Shortcuts“ in the data
Beyond reproducibility: do benchmarks assess generalisable learnings? Example: Twitter bot detection Chris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk, and Philipp Zimmer. 2023. Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot Detection. ACM WebConf2023 Take-aways ▪AI benchmark data does not represent real-world data/problems but contains shortcuts ▪Shortcut learning [Geirhos2020] is widespread and leads to poor generalisability ▪Reproducible results ≠generalisable results ▪Benchmarking, i.e. understanding what is state-of-the-art in AI/NLP is hard Geirhos, R., Jacobsen, JH., Michaelis, C. et al. Shortcut learning in deep neural networks. Nature Machine Intelligence 2, 665–673 (2020). 19 „Shortcuts“ in the data
Addressing reproducibility & generalisability in CSS/AI research? 1. Empowering researchers to find state-of-the-art methods (“benchmarking / state-of-the-art crisis”) 2. Improving the interpretability of scholarly reporting (“reporting problem”) 3. Ensuring data availability & access (“access problem”) Reproducibility Replicability Robustness Generalisability 22
Overview 1. Empowering researchers to find state-of-the-art methods (“benchmarking / state-of-the-art crisis”) 2. Improving the interpretability of scholarly reporting (“reporting problem”) 3. Ensuring data availability & access (“access problem”) 23
Key challenge: how to identify high quality methods? How to find SotA methods for given task (e.g. stance detection on specific tweet sample)? •Review literature: labor-intensive, methods often poorly cited / not traceable •Code/model repositories (e.g. HuggingFace, GitHub): lack context (e.g. related research, comparisons with other methods etc) •Ad-hoc choices („I use what I know“) Benchmarking of AI/CS methods •Use of standard evaluation corpora & metrics to compare method performance / quality •In theory: benchmarks assess whether a published method is good/bad/state-of-the-art •In practice: benchmarks and benchmarking practices (eg baseline choices) are flawed, e.g. do not evaluate generalisability 24
Finding AI methods for the social sciences: GESIS Methods Hub Released in Q3 2025 Integrated into GESIS Search, MyBinder, Jupyter4NFDI •Platform for finding, sharing & using/executing data science & AI methods •Empowering social scientists with & without technical expertise to use complex state-of-the-art methods & LLMs •GESIS-curated and community-based methods and tutorials •Focus on reproducibility, quality, citability (DOIs), benchmarking, provenance https://methodshub.gesis.org 25
Benchmarking: evaluating generalisability of NLP models Feger, M., Boland, K., Dietze, S., Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments, In ACL2025. 26 Example case: argument mining in tweets/social media posts as established NLP task
Feger, M., Boland, K., Dietze, S., Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments, In ACL2025. Do models actually generalise? •Train-on-one-test-on-another (dataset) experiments on 17 AM datasets •Using state-of-the-art Transformer-based language models (BERT, RoBERTa, WRAP) •Models do not generalise („do not learn to detect arguments“): performance degrades when models are tested on OOD data 27 Benchmarking: evaluating generalisability of NLP models
Feger, M., Boland, K., Dietze, S., Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments, In ACL2025. •Leave-one-out cross validation: models trained on all datasets but the target dataset (rows) •Performance degradation significant (despite more diverse training data) •Performance drop particularly for datasets that seemed „easy“ to learn 28 Realistic benchmarking: evaluating generalisability of NLP models
Detecting model, task and dataset mentions: model performance Otto, W., Zloch, M., Gan, L., Karmakar, S., Dietze, S. (2023). GSAP-NER: A Novel Task, Corpus, and Baseline for Scholarly Entity Extraction Focused on Machine Learning Models and Datasets. In Findings of the Association for Computational Linguistics: EMNLP 2023 36
Understanding methods and data in CSS (AAAI ICWSM publications) Tasks Methods 37
Understanding methods and data in CSS (AAAI ICWSM publications) Citations of ML models over time Citations of data sources over time 38
MethodMiner: a tool for mining task, dataset & model mentions 39 Otto, W., Upadhyaya, S., Gan, L., Silva, K. (2025), Track Machine Learning in Your Research Domain. In 2nd Conference on Research Data Infrastructure (CoRDI)
Shared AI task @ ACL2025: mining data, model, software mentions https://sdproc.org/2025/somd25.html 44
Overview 1. Empowering researchers to find state-of-the-art methods (“benchmarking / state-of-the-art crisis”) 2. Improving the interpretability of scholarly reporting (“reporting problem”) 3. Ensuring data availability & access (“access problem”) 45
46 Challenge: dependencies on 3rd party gatekeepers Behavioral data is not distributed as the web but tied to platforms/gatekeepers
Challenge: volatility & decay of web data •Data is not persistent •Example: deletion ratio of tweets between 25-29 % •Differs between different samples Khan, M.T., Dimitrov, D., Dietze, S., Characterization of Tweet Deletion Patterns in the Context of COVID-19 Discourse and Polarization, ACM Hypertext 2025 47
Challenge: data evolution impacts methods (quality/reproducibility) ▪Vocabulary evolves: e.g. vocabulary shift, over- /underrepresentation of topics/vocabulary in particular time periods (e.g. Twitter COVID19discourse 2020 vs prior periods) ▪PLMs/LLMs require frequent training and updates (and continuous access to data) Source: Hombaiah et al., “Dynamic Language Models for continuously evolving Content”, SIGKDD2021 48
Responsible social media archiving @ GESIS: examples X/Twitter (https://data.gesis.org/tweetskb) ▪Sampling: 1% - random sample ▪Dataset size: > 14 billion tweets ▪Time period: Feb 2013 - June 2023 Telegram (https://data.gesis.org/telescope) ▪Sampling: seed lists + snowball sampling ▪Dataset: ~120M messages from ~71K public channels and metadata for ~500K channels ▪Time period: Feb 2024 and running Fact-checked claims (https://data.gesis.org/claimskg) ▪Sampling method: 13 factchecking websites ▪Dataset: 74066 claims and 72128 claim reviews ▪Time period: claims published between 1996 –2023 4Chan ▪Sampling method: all boards ▪Dataset size: 4,676,378 threads, 264,898,231 posts ▪Time period: Nov 2023 and running ▪In preparation: BlueSky, YouTube, … https://www.gesis.org/gesis-web-data 50
https://stefandietze.net https://gesis.org/en/kts Thank you!
6
www.ditrare.de/en Thank you for joining! Stay connected ■DiTraRe ○Website: www.ditrare.de/en ○Email: ditrar[email protected] ○LinkedIn: www.linkedin.com/company/ditrare ○Mastodon: social.kit.edu/@DiTraRe ○YouTube: www.youtube.com/@DiTraRe ○Zenodo: zenodo.org/communities/ditrare ■Discussion forum: www.ditrare.de/en/forum ■Newsletter: www.ditrare.de/en/newsletter 7