scieee AI-readable full text Open interactive document viewer

OSSGameBench: A Large-Scale Dataset of Contributor Activities in the Open-Source Video Games

Marsad, Faiz; Weeraddana, Nimmi

Abstract

Video games have evolved beyond entertainment to support education, training, and well-being through deeply immersive and interactive experiences. Yet, unlike other software systems where standardized benchmark datasets have accelerated research on downstream tasks such as defect prioritization and localization, the game development domain lacks comparable resources that reflect its distinctive characteristics. For instance, defects in games often arise from the complex interplay between code and non-code assets, whereas in other software they do not. Therefore, specialized datasets are required to capture these multidisciplinary demands unique to game development.This gap limits the development and evaluation of data-driven approaches for tasks such as defect prioritization, defect localization, and automated repair. To fill this void, we introduce OSSGameBench, a curated benchmark dataset derived from 970 open-source game repositories in GitHub, which contains issues, comments, commits, pull requests, and review comments, with explicit mappings among these entities to enable reproducible and comprehensive defect analysis tasks in game development. In this Zenodo repository, we will include the OSSGameBench dataset and the replication package, which contains the scripts used to collect and analyze the data. Please check the README.md file and our Online Appendix.pdf for detailed instructions on how to use our dataset and/or execute these scripts. Please cite our paper: @inproceedings{marsad2026msrdata, Author = {Marsad, Faiz and Weeraddana, Nimmi}, Title = {{OSSGameBench: A Large-Scale Dataset of Contributor Activities in the Open-Source Video Games}}, Year = {2026}, Booktitle = {Proc. of the International Conference on Mining Software Repositories (MSR)} }

Full text

OSSGameBench: A Large-Scale Dataset of Contributor Activities in Open-Source Video Games Online Appendix 1. Abstract 3 2. Dataset Schema 4 Projects 5 Entity name 5 Attributes 5 Relative path to the data 5 Open the data in Python 5 Issue reports 6 Entity name 6 Relative path to the data 6 Open the data in Python 6 Issue comments 7 Entity name 7 Relative path to the data 7 Open the data in Python 7 Pull request 8 Entity name 8 Relative path to the data 8 Open the data in Python 8 PR review comments 9 Entity name 9 Relative path to the data 9 Open the data in Python 9 Issue-PR mapping 10 Entity name 10 1 Relative path to the data 10 Open the data in Python 10 Commits metadata 11 Entity name 11 Relative path to the data 11 Open the data in Python 11 Commits 12 Entity name 12 Relative path to the data 12 Open the data in Python 12 PR - commit mapping 13 Entity name 13 Relative path to the data 13 Open the data in Python 13 3. Analysis: Game genres 15 Introduction 15 Details of the coding process 16 Evaluation of the coding agreement 16 Results 17 References 18 2 1. Abstract Video games have evolved beyond entertainment to support education, training, and well-being through deeply immersive and interactive experiences. Yet, unlike other software systems where standardized benchmark datasets have accelerated research on downstream tasks such as defect prioritization and localization, the game development domain lacks comparable resources that reflect its distinctive characteristics. For instance, defects in games often arise from the complex interplay between code and non-code assets, whereas in other software they do not. Therefore, specialized datasets are required to capture these multidisciplinary demands unique to game development. This gap limits the development and evaluation of data-driven approaches for tasks such as defect prioritization, defect localization, and automated repair. To fill this void, we introduce OSSGameBench, a curated benchmark dataset derived from 970 open-source game repositories in GitHub, which contains issues, comments, commits, pull requests, and review comments, with explicit mappings among these entities to enable reproducible and comprehensive defect analysis tasks in game development. In this document, we provide extra information about our dataset for anyone who is interested in using it. 3 2. Dataset Schema 4 Projects Entity name project Attributes Attribute name Description project The full name of the project repository in the format owner/repo (e.g., AlmasB/FXGL). is_fork Indicates whether the repository is a fork of another project (true or false). languages A dictionary mapping programming languages to the number of bytes of code written in each language (e.g., {"Java": 515464, "Shell": 1702, "Rich Text Format": 928}). topics A list of topics or tags associated with the repository on GitHub (e.g., ["game", "java", "2d-games"]). created_at The timestamp representing when the repository was created. readme_content Content of the README file, if available Relative path to the data OSSGameBench/project.parquet Open the data in Python import pandas as pd projects = pd.read_parquet(“OSSGameBench/project.parquet") 5 Issue reports Entity name issue Attributes Attribute name Description project The full name of the project repository in the format owner/repo (e.g., AlmasB/FXGL). number Issue number (int) contributor Issue author’s login/display name. title Issue title. body Issue description text (may be empty). labels List of label names (strings). state open / closed created_at Creation timestamp closed_at Close timestamp (nullable) Note that one repo can contain multiple issues. Issue entity can be linked to project entity by <repo> attribute. Relative path to the data OSSGameBench/issue.parquet Open the data in Python import pandas as pd issues = pd.read_parquet(“OSSGameBench/issue.parquet") 6 Issue comments Entity name issue_comment Attributes Attribute name Description project owner/name issue_number Issue number (int). contributor Comment author. body Comment text. created_at Comment timestamp Note that one issue can have multiple comments, and comments can be ordered by the created_at time. Issue_comment entity can be linked to issue entity by <repo,issue_number>. attributes. Relative path to the data OSSGameBench/issue_comment.parquet Open the data in Python import pandas as pd issue_comment_mapping = pd.read_parquet(“OSSGameBench/issue_comment.parquet") 7 Pull request Entity name pr Attributes Attribute name Description project The full name of the project repository in the format owner/repo (e.g., AlmasB/FXGL). number PR number (int) contributor PR author’s login/display name. title PR title. body PR description text (may be empty). labels List of label names (strings). state open / closed created_at Creation timestamp closed_at Close timestamp (nullable) merged_at PR merge timestamp (nullable) Note that one repo can contain multiple PRs. pr entity can be linked to the project entity by <repo> attribute. Relative path to the data OSSGameBench/pr.parquet Open the data in Python import pandas as pd prs = pd.read_parquet(“OSSGameBench/pr.parquet") 8 PR review comments Entity name pr_comment Attributes Attribute name Description project owner/name pr_number PR number (int). contributor Comment author. body Comment text. created_at Comment timestamp Note that one PR can have multiple review comments, and comments can be ordered by the created_at time. pr_comment entity can be linked to pr entity by <repo,pr_number>. attributes. Relative path to the data OSSGameBench/pr_comment.parquet Open the data in Python import pandas as pd pr_review_comments = pd.read_parquet(“OSSGameBench/pr_comment.parquet") 9 Details of the coding process In fact, two coders classify the README files of 100 randomly sampled game projects; one coder is a human (the second author), and the other is an LLM (OpenAI GPT-4.1).2 The human coder reviewed Table A.1 to understand how each genre is described. Then, they investigated the text of each README file in our sample and categorized it accordingly. The second coder, i.e., OpenAI GPT-4.1-mini model, was provided the following prompt for each README file content. You are a video game developer. Determine whether the README text discusses the genre of the game Use ONLY the genres defined in the JSON below. If a subtype/example is mentioned, map it to its parent 'genre'. You can guess the genre, but it should be one of the genres in the JSON below. Return strictly one of the following: - A comma-separated list of genre names exactly as they appear in the 'genre' field, OR - 'Not found' if no genre is mentioned (or information is not sufficient). GENRE: <CONTENT OF TABLE A.1> README TEXT: <TEXT OF THE README FILE> Evaluation of the coding agreement After coding, we compute the inter-coder reliability using Krippendorff’s α,3 a widely used reliability coefficient that corrects for chance agreement and generalizes to multiple coders, missing data, and different levels of measurement. This metric is particularly suitable for assessing agreement in multi-label coding tasks, such as genre classification, where a single project may belong to more than one genre. Specifically, we calculate α separately for each genre and report both the macro-averaged and prevalence-weighted average coefficients to capture the overall agreement between the human and LLM coders. 3 https://repository.upenn.edu/handle/20.500.14332/2089 2 https://openai.com/index/gpt-4-1/ 16 Results Our evaluation resulted in an average coefficient value of 0.67 across all predefined genres, indicating moderate-to-substantial agreement between the coders. With this agreement, we extend the LLM script to code the README files of all the projects in our dataset. This information is now available in our dataset for validation and replication purposes. We then visualize the distribution of genres in our dataset; see Figure 5 in our paper. Accordingly, the most frequent genres were role-playing (136 projects), action (129), and simulation (122). A close inspection of this data further reveals that 300 projects lack genre information in their README files, 66% of which are game development tools. For example, FXGL11 is a game engine for Java and Kotlin projects. This absence of genre metadata is also expected when repositories serve tooling or infrastructural purposes rather than representing games themselves. For instance, XStreaming4 is a mobile streaming client that connects to Xbox consoles, enabling remote gameplay on mobile devices. This variety highlights the dataset’s coverage, capturing not only playable games but also engines and tools that represent the broader open-source game development ecosystem. 4 https://github.com/Geocld/XStreaming 17 References Humble, Áine M. 2009. “Technique Triangulation for Validation in Directed Content Analysis.” International Journal of Qualitative Methods 8 (3): 34–51. Swacha, Jakub. 2025. “The Relative Popularity of Video Game Genres in the ScienTific Literature: A Bibliographic Survey.” Multimodal Technologies and Interaction 9. 18