scieee AI-readable full text Open interactive document viewer

Assessing the Prevalence and Security Risks of Credential Leakage in GitHub Repositories of Undergraduate Open-Source Projects

Daniel O. Cabasa; Andrei C. Co; Derick James M. Espinosa; Aminah C. Malic; Joan C. Mag-isa

Abstract

This cross-sectional prevalence study measured the occurrence of credential leakage in undergraduate open-source projects hosted on public GitHub repositories collected on September 27, 2025. The main goal was to determine the rate at which the repositories showed sensitive credentials in the projects started on or after January 1, 2023. Candidate repositories were identified using tailored search queries combining undergraduate project indicators with backend and database keywords, and a minimum size threshold of 3MB was applied to filter the results. Fifty repositories that fulfilled specific inclusion criteria were selected. The information and metadata were accessed through the GitHub REST API v4 and the repositories were shallow-cloned and inspected locally. Gitleaks v8.21.2 and TruffleHog v3.82.6 were used to detect secrets. Provenance was stored in metadata and raw API outputs in the form of JSON files. A rubric based on CVSS generated repository-level risk scores between 0 and 6. Outputs were hashed, and secret strings were redacted. Where necessary, descriptive statistics, Wilson 95 percent confidence intervals, and chi-square or Fisher exact tests were used. Findings showed that 31 out of 50 repositories, or 62.0%, 95% confidence interval: 48.15 -74.14, had one or more leaked secrets. The analysis identified 325 cases of secret leakage; the majority of them were database-related credentials (n= 209, 64.3 percent), then the source-control keys (n= 65, 20.0 percent), and the cloud/API tokens (n= 51, 15.7 percent). The risk level of the repositories was categorized into High (n=15), Medium (n=16), or Low (n=19). The leakage was not significantly related to programming language (χ² = 4.27, df = 7, p = 0.749). The results show that credential leakage is prevalent in student repositories and usually includes high-impact secrets. It is suggested to implement curricular and operational interventions, such as the implementation of secure-by-default assignment templates, secret scanning, environment-based configuration, and institutional remediation and rotation processes.

Full text

cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 270 Assessing the Prevalence and Security Risks of Credential Leakage in GitHub Repositories of Undergraduate Open-Source Projects Daniel O. Cabasa; Andrei C. Co; Derick James M. Espinosa; Aminah C. Malic; Joan C. Mag-isa Bachelor of Science in Information Technology, Technological University of the Philippines, Taguig, Philippines [email protected]; an[email protected]; [email protected]; [email protected]; joan_m[email protected] DOI: 10.47760/cognizance.2025.v05i10.025 Abstract: This cross-sectional prevalence study measured the occurrence of credential leakage in undergraduate open-source projects hosted on public GitHub repositories collected on September 27, 2025. The main goal was to determine the rate at which the repositories showed sensitive credentials in the projects started on or after January 1, 2023. Candidate repositories were identified using tailored search queries combining undergraduate project indicators with backend and database keywords, and a minimum size threshold of 3MB was applied to filter the results. Fifty repositories that fulfilled specific inclusion criteria were selected. The information and metadata were accessed through the GitHub REST API v4 and the repositories were shallow-cloned and inspected locally. Gitleaks v8.21.2 and TruffleHog v3.82.6 were used to detect secrets. Provenance was stored in metadata and raw API outputs in the form of JSON files. A rubric based on CVSS generated repository-level risk scores between 0 and 6. Outputs were hashed, and secret strings were redacted. Where necessary, descriptive statistics, Wilson 95 percent confidence intervals, and chi-square or Fisher exact tests were used. Findings showed that 31 out of 50 repositories, or 62.0%, 95% confidence interval: 48.15 -74.14, had one or more leaked secrets. The analysis identified 325 cases of secret leakage; the majority of them were database-related credentials (n= 209, 64.3 percent), then the source-control keys (n= 65, 20.0 percent), and the cloud/API tokens (n= 51, 15.7 percent). The risk level of the repositories was categorized into High (n=15), Medium (n=16), or Low (n=19). The leakage was not significantly related to programming language (χ² = 4.27, df = 7, p = 0.749). The results show that credential leakage is prevalent in student repositories and usually includes highimpact secrets. It is suggested to implement curricular and operational interventions, such as the implementation of secure-by-default assignment templates, secret scanning, environment-based configuration, and institutional remediation and rotation processes. Keywords: credential leakage, GitHub, undergraduate open-source projects, prevalence study, secret scanning, security risk assessment cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 271 I. INTRODUCTION Platforms such as GitHub have become essential tools of collaboration, version control, and sharing of code among software developers working at both professional and student levels in modern software development. These platforms may be extremely helpful as far as productivity and learning are concerned, but also a big security threat when any of the much sensitive credentials including API keys, database connection strings, service tokens or private keys are spilled, on to an open repo. Acquiring these credentials is simple once issued and is then quickly harvested and consumed by bad actors to gain unauthorized access, financial gain or cause enormous security breaches. Empirical studies are used to draw the severity of such a phenomenon. Meli et al. (2019) found out that millions of files on GitHub contained sensitive data on hundreds of thousands of repositories and that millions of secrets were live and could be leveraged once they had been discovered [1]. According to the State of Secrets Sprawl Report of the GitGuardian which was published in 2024, the number of secrets that have been leaked into public repositories has already increased over 23 million considering that this is 25% more than the volume that was leaked the previous year (GitGuardian, 2024) [2]. Krause et al (2022) researched the processes of prevention and remediation; they found out that among the developers of a sample, more than one out of every three had accidentally leaked secrets at some point in their career writing programs, whether through lack of program knowledge on how to do this or lack of support in tools [3]. The risks are especially concerning among the student developers. Research indicates that although students understand the need to use secure coding, the main problem is that they do not effectively apply it. According to Mutanga (2022), undergraduates in the field of developers have shown a tendency of not having the necessary knowledge and experience to implement security concepts in practice [4]. Similarly, Saeed et al. (2019) have systematically observed failures of the students in the field of applying secure coding rules to academic tasks, especially at the beginning of the curriculum [5]. The majority of current studies address professional developers or major open-source projects, and the aspects of credential leakage by student-written repositories have not been empirically studied. However, there is still some worry about how many times student repository credential leaks occur, what kind of secrets are revealed and whether students are aware of countermeasures like .gitignore policies and secret-management behavior. In this paper, we will discuss how common credential leakage is in publicly available GitHub repositories created by student developers. Specifically, it will provide a response to the following sub-questions (1) What percentage of the sampled repositories have at least one exposed secret? (2) what sorts and extent of secrets can be disclosed? (3) What repository level properties relate to the presence of credential leakage? II. METHODOLOGY A. Study Design This study used a cross-sectional prevalence design to study the prevalence of credential leakage in publicly accessible repositories of undergraduate open-source projects hosted on GitHub [6]. The main aim was to determine the occurrence of credential leakage. Specifically, the study identified the proportion of sampled repositories containing at least one exposed secret, characterized the types and severity of the exposed credentials, and examined repository-level factors associated with leakage. The scope was limited to repositories created from January 2023 onward, with 50 repositories selected per predefined criteria. The sample size was based on power analysis to detect a prevalence of about 30% with a 95% confidence interval of acceptable width [7]. Following recommendations on structuring methods sections, the study design is explicitly presented to ensure clarity, reproducibility, and adherence to established reporting guidelines. [8] B. Data Collection The data and metadata were collected from GitHub on September 27, 2025, using the REST API v4 [9], each data or repository was cloned locally for inspection. The repositories were identified through search cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 272 queries used to filter undergraduate open-source projects (e.g., “student,” “university”) with backend or database keywords and a minimum size of greater than or equal to 3 Megabytes, limited to repositories created on or after January 1, 2023. For provenance, all API responses and raw metadata (repository name, owner, creation date, size, primary language, etc.) were archived in JSON. For reproducibility, all scripts and analysis outputs are archived in a public GitHub repository [12] Fig 1. PRISMA Flow Diagram of GitHub Repository Selection and Analysis Fig. 1 outlines the selection process, where repositories are first filtered using specific GitHub search queries, then filtered for duplicates, then filtered out as irrelevant or too small. Eligibility is determined based on implementing a risk-score-system that includes language, keywords, database / backend indicators, size of repository, activity, and evidence of student input. Finally, approximately fifty repositories are selected as the group of analysis. C. Data Processing and Analysis Phase 1: Secret Detection and Analysis Each selected repository was shallow-cloned locally (git clone --depth 1) for surface-level inspection. Two complementary tools were then applied: Gitleaks v8.21.2 [10] scanned the working tree locally, while TruffleHog v3.82.6 [11] analyzed the repository’s remote history for historical secrets without requiring a full re-clone. Both detectors produced structured JSON reports, which were archived in the raw reports directory for subsequent anonymization and statistical analysis. cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 273 Fig 2. Flowchart of Phase 1: Secret Detection and Analysis Fig. 2 illustrates the sequential process applied to each repository: shallow cloning, parallel scanning with Gitleaks and TruffleHog, and generation of structured JSON outputs forwarded to anonymization and analysis. Phase 2: Risk assessment framework A CVSS-inspired rubric [13] was applied to each repository, scoring findings on three dimensions: Exploitability, Impact, and Exposure. Scores were aggregated to a total risk score (0–6) and mapped to risk levels. TABLE I RISK SCORING RUBRIC AND CLASSIFICATION. cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 274 Table I summarizes the rubric used in Phase 1 in the evaluation of the detected credentials. Exploitability is used to describe how easily a secret can be exploited. Impact measures potential damage depending on the occurrence and type of findings (an extra weight on a database/backend indicator), and Exposure is the degree of availability. Since only the public repositories were examined, Exposure was set to 2. The combination of the three dimensions resulted in a total risk score of 0 to 6, and it was then charted to categorical risk levels: Low (02), Medium (34), and High (56). Phase 3: Data processing and anonymization Outputs from Gitleaks [10] and TruffleHog [11] were parsed into a unified JSON/CSV summary (repository metadata, language, size, counts of findings, and risk score). To preserve privacy, owner identifiers were anonymized with SHA-256 (8-character hash) and all secret strings were redacted. Raw tool reports were stored in a restricted directory (reports/private/), while anonymized summaries were stored separately for analysis. Phase 4: Statistical analysis The descriptive statistics were summarized, including the frequency of repositories with a ≥1 secret and Wilson 95% confidence intervals [14]. The frequency tables were used to summarise the distributions based on programming language, size category, and risk level. The relationships between the existence of a secret and categorical variables were tested by chi-square or Fisher's exact tests, as needed [15]. D. Ethical Considerations The study investigated only publicly available GitHub repositories and had no direct engagement with the owners of the repositories. Any personally identifiable data was anonymised by hashing of usernames and redacting the identification of any secrets that were detected before analysis; any raw non-redacted data was stored in a directory with controlled access. By the principles of responsible disclosure, no particular vulnerabilities or confidential material were publicized, and the results are reported only in aggregate. III. RESULTS AND DISCUSSION A. Repository and Secret Collection TABLE I Summary of Credential Leakage Prevalence in Student GitHub Repositories (n = 50) The analysis in Table I shows that the number of repositories analyzed in GitHub is 50, and only 31 of the repositories (62 percent) had at least one of the following secrets exposed. According to it, more than half of all student developers of the dataset quietly put sensitive data in their repositories, such as API keys, database logins, or configuration tokens. The existence of such a vast amount of prevalence underscores the fact that the issue of credential leakage is not a single mistake, and it is an inclination prevalent in a large percentage of student developers. Analysis Value Total repositories analyzed 50 Repositories with ≥1 secret 31 Prevalence of leakage (%) 62.0 95% Confidence Interval (%) 48.15 – 74.14 cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 275 B. Assessment and Categorization TABLE II Distribution of Exposed Secrets by Category, Top Types, and Typical Risk (N = 325 findings) The analysis in Table II demonstrates that out of 325 instances of secrets found in the examined repositories, three major groups, such as Databases/ Data-store credentials and Cloud and API tokens, as well as Source Control and Keys, appeared. Both categories present certain yet serious risks to the state of student projects and the feasibility of damaging them that can be done by malicious individuals. The distribution has encouraging news: there are and have been few exposures of the lower type of risk (high-quality assurances). With 64.3% database-related leaks or 20.0% source-control exposures, most of the findings could translate to the resultant giving easy access to a system or be used as a direct corruptionpromoting system. Cloud and API tokens are uncommon, yet they can present some serious threats, as the possibility of breaking the services and stealing the quota exists. TABLE III Repository-level Correlates of Credential Leakage Table III presents the results of a chi-square test to show how the programming language used in the student GitHub project repositories (JavaScript, Java, PHP, C etc.) depends on whether or not there is credential leakage. The chi-square test value is: kh2 = 4.27 and 7 degrees of freedom (df = 7). P-value = 0.749, which is larger than the customary level of significance (0.05 or 0.01). This indicates that the correlation that exists between the programming language that is in the code, as well as the existence of uncovered secrets, is not statistically significant. Leakage of credentials is not languagedependent, that is, it would occur in repositories regardless of whether the project was created using JavaScript, Java, PHP, C#, or any other language. Category Top Types (counts) Total (n, %) Typical Risk Databases / Data-store Sqlserver (72), MongoDB (66), Jdbc (45), URI (23), Postgres (3) 209 (64.3%) High (direct DB compromise) Cloud & API To kens API Key (9), Swell (9), Tomorrowio (6), TravisCI (6), Miro (6) 51 (15.7%) Medium–High (service abuse, quota theft) Source-control & Keys Private Key (52), Githuboauth2 (4), Gitlab (4), JWT Token (2), Gitlab-Pat (1) 65 (20.0%) High (account takeover, code repo access) Variable Language (JS, Java, PHP, C#, Other) Comparison/Statistic With vs. without secrets (χ²) Value χ² = 4.27, df = 7 Test Chi-square test p-value 0.749 cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 276 Fig 1: Severity of affected repositories The table presents the result of a chi-square test in order to demonstrate how the programming language used in the student GitHub project repositories (JavaScript, Java, PHP, C, etc.) changes in the presence or absence of a leakage of the credentials. IB. ki2 =7 df=4: chi square test value = kh2 = 4.27. P = -0.749, higher than the standard level of significance (.05 or 01). Fig 2. Prevalence of Exposed Secrets: a donut chart showing that 31 of 50 repositories (62.0%, 95% CI: 48.2–74.1) contained at least one exposed secret. Figure 2 shows that only 38 percent of the 50 student repositories studied based on the donut chart had zero exposed secrets, and 62 percent of there were one or more types of secrets. This indicates that the issue of credential leakage is a trending and widely prevalent issue, and suggests that the majority of student developers cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 277 are not following secure code such as the use of .gitignore or secret management systems. It is high, meaning it should be prevented by being conscious of it. C. Data and Statistical Analysis Fig 3. Distribution of Secret Categories Figure 3 demonstrates how frequently and in what proportion the different categories of secrets that are being revealed in the repositories under consideration appear. The credentials associated with the databases are more prevalent in the exposures, the SQL server (72, 22.1%), MongoDB (66, 21.4%), JDBC (45, 14.6%), and the URI strings (23, 7.5) are filled with the majority of the leaked information according to the results. This implies that students usually refer to or misuse their database connection parameters and that subject projects to greatly exposed to being attacked directly into a database. As many as a given number of these projects may be at the academic level, inclusion of placing live database credentials puts the act of such irresponsible habits that would most probably be incorporated into any work experience. Overall, and as the distribution of Figure 3 suggests, the issue lies not in its scope but is very widespread, spanning databases, token API keys, and source-control credentials. The abundance of database exposures (over 70 percent) and private key exposures (over 70 percent) indicates that the overwhelming and most significant fraction of leaks concerns high-impact secrets, and allows one to directly use them to compromise data, as opposed to low-level misconfigurations. This further confirms that a comprehensive teaching of safe codes is in most, which is an automatic secret scanner, and a high level of adherence to the best practices of version control. IV. CONCLUSION AND RECOMMENDATION In this study, it is shown that credit leakage is common to the student-created open GitHub repositories, with 62.0 (31/50, 95% CI 48.15-74.14) of each containing at least one exposed secret. High-risk (15), mediumrisk (16) and low-risk (19) distribution of severity shows a serious potential to circumstances in which unauthorized entry and data breach are possible. The leading category of leaked credentials was database credentials and private keys (MongoDB, SQL server, JDBC) and API keys and cloud service credentials were systemic because they posed a threat to the confidentiality and integrity of data. The exposure of coding to cognizancejournal.com Daniel O. Cabasa et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 270-279 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 278 statistics and credential exposure were not significantly related (kh2 = 4.27,df = 7, p = 0.749) in meaning that leakage will happen across language ecosystems but was the result of insecure development habits rather than specific technical environments. The consequences of this finding are that they need multi-level interventions. The teachers are expected to work with assignment templates created with the use of secure-by-default templates, prepared environment variables, pre-written gitignore templates, require pre-submission secrets scanning and learn secure workflows in Git and credential rotation. The student developers are not to hard code the credentials, scan sparsely to storage, scan in-place the local scan before committing, and even upgrade leaked secrets after committing. Temporary scoped credentials should be offered, mass rotation facilities need to be implemented, and leaks should be scanned at scale across the institution. To further improve this study, a number of proposals are implied to be made by future research. To enhance the external validity and discussion of the results, first, it would be desirable to increase the sample of repositories that would be representative of a broader number of institutions, fields, and geographical areas. Longitudinal would also be useful to monitor whether credential leakage would decline as months go by, especially as effective measures ( institutional interventions ) or greater student awareness are put in place. Also, including qualitative data with surveys or interviews may provide more knowledge on why students still hardcode credentials with identification of knowledge gaps, familiarization with the tools, or practice of workflow, that statistical analysis may fail to yield. The other strength enhancement needed would be comparative research exploring professional or sizable open-source, as well as student projects, to better put into perspective how much the exposure of credentials is a student problem, or a symptom of widely spread software development practices. Finally, adequacy of various tools and institutional mechanisms against detecting secrets and mitigating through various mechanisms should be assessed by the future studies in order to be able to offer more tangible recommendations of best practices and tooling effectiveness. REFERENCES 1. Meli, M., McNiece, M. R., & Reaves, B. (2019). How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories. Proceedings 2019 Network and Distributed System Security Symposium. https://doi.org/10.14722/ndss.2019.23418 2. GitGuardian. (2024). The State of Secrets Sprawl Report 2024. GitGuardian. https://www.gitguardian.com 3. Krause, A., Klemmer, J. H., Huaman, N., Wermke, D., Acar, Y., & Fahl, S. (2022). Committed by Accident: Studying Prevention and Remediation Strategies Against Secret Leakage in Source Code Repositories. ArXiv.org. https://arxiv.org/abs/2211.06213 4. Mutanga, M. (2022). Secure software development awareness: A case study of undergraduate developers. Journal of Research in Engineering and Applied Sciences, 7(3), 39–46. https://doi.org/10.52326/jes.utm.2022.29(2).08 5. Saeed Al-Haj, Naeem Seliya, & Kemner, C. L. (2019, June 15). Pedagogical Assessment of Secure Coding in Student Programs. Asee.org. https://peer.asee.org/pedagogical-assessment-of-secure-coding-in-student-programs? 6. Setia, M. S. (2016). Methodology Series Module 3: Cross-sectional Studies. Indian Journal of Dermatology, 61(3), 261–264. “Cross-sectional designs characterize the prevalence of outcomes in a population.” https://pmc.ncbi.nlm.nih.gov/articles/PMC9536510 7. Naing, L., et al. (2022). Sample size calculation for prevalence studies using Scalex. BMC Medical Research Methodology. “Parameters include confidence level, expected prevalence, and precision (acceptable width)” https://bmcmedresmethodol.biomedcentral.com/articles/10.1186/s12874-022-01694-7 8. Eldawlatly, A. A., & Meo, S. A. (2019). Writing the methods section: Basic elements of methods section in scientific papers. Saudi Journal of Anaesthesia. PMCID: PMC6398288. Retrieved from https://pmc.ncbi.nlm.nih.gov/articles/PMC6398288/ 9. GitHub, Inc. (2023). GitHub REST API documentation. GitHub. https://docs.github.com/en/rest 10. Zricethezav. (2023). Gitleaks: Detect and prevent secrets in git repos. GitHub. https://github.com/zricethezav/gitleaks