scieee AI-readable full text Open interactive document viewer

CSRFing the SSO Waves: Security Testing of SSO-Based Account Linking Process

Bisegna, Andrea

Abstract

Single Sign-On (SSO) account linking allows users to connect their Service Provider (SP) accounts with their Identity Provider (IdP) identities, and is commonly implemented through OAuth 2.0, OpenID Connect (OIDC), or SAML 2.0 federation flows. However, the security of this process is often overlooked. In this work, we present a systematic analysis of a critical Account Hijacking attack targeting SSO-based account linking (SSOLinking). The attack exploits two Cross-Site Request Forgery (CSRF) vulnerabilities: an Authentication CSRF (Login CSRF) at the IdP and a CSRF on the SP button that triggers the linking process. We introduce an automated security testing approach, implemented as the SSOLinking Checker extension of the Micro-Id-Gym penetration testing framework, to detect these vulnerabilities. Our large-scale empirical study evaluates 648 popular websites supporting SSO account linking with major IdPs. Of the 48 eligible SPs, 21 were found vulnerable (43.7%), including high-profile services such as Goodreads, Naver, and Workable. Our findings demonstrate that CSRF-based Account Hijacking in OAuth-, OIDC-, and SAML-backed SSO account linking is widespread and dangerous.

Full text

CSRF-ing the SSO waves: security testing of SSO-based account linking process Andrea Bisegna Center for Cybersecurity, Fondazione Bruno Kessler Trento, Italy a.bise[email protected] Matteo Bitussi Center for Cybersecurity, Fondazione Bruno Kessler Trento, Italy [email protected] Roberto Carbone Center for Cybersecurity, Fondazione Bruno Kessler Trento, Italy [email protected] Luca Compagna SAP Security Research France [email protected] Silvio Ranise Center for Cybersecurity, Fondazione Bruno Kessler and Department of Mathematics, University of Trento Trento, Italy [email protected] Avinash Sudhodanan Independent Researcher CA, USA [email protected] Abstract—The Single Sign-On based account linking process (SSOLinking in short) allows users to link their accounts at Service Provider (SP) websites to their Identity Providers (IdP) accounts. We focus on a serious (and overlooked) attack, namely an Account Hijack targeting the SSOLinking and relying on two CSRF vulnerabilities, one affecting the IdP and the other the SP. The former is an Authentication CSRF (also known as Login CSRF) and the latter is a CSRF on the button triggering the SSOLinking. We propose a security testing approach to help testers automatically detect such attacks. We implemented our testing technique as an extension (namely SSOLinking Checker) to the open-source penetration testing tool Micro-Id-Gym. To demonstrate the effectiveness of our approach and the pervasiveness of the SSOLinking Account Hijack, we conducted an experimental analysis against a selection of popular SPs that offer the SSOLinking with major IdPs. The results of our experiments are alarming: out of the 648 web sites we considered, 48 qualified for conducting our experiments and 21 of these suffered from SSOLinking vulnerability (i.e. 43.7%). Our findings (we responsibly disclosed to the affected vendors) include severe vulnerabilities among the web sites of Goodreads, Naver, Workable, etc. 1. Introduction More and more websites enrich their standard authentication processes with Single Sign-On (SSO) to smoothly sign-in users via popular Identity Providers (IdPs) like Facebook and Google. Ensuring the security of these SSO processes is thus paramount. A little mistake in their implementation may introduce a vulnerability and jeopardize the entire authentication, sometimes even enabling an attacker to take complete control of a victim’s account on a website. For instance, Cross-Site Request Forgery (CSRF) [25] is a vulnerability that enables a malicious website to forge state-changing HTTP requests from a victim’s web browser. CSRF has been known for more than 20 years and has been demoted and removed from the OWASP Top 10 list in 2017 [26], mainly because of the rollout of framework supported defenses [17]. More recently, browser vendors rolled out SameSite cookies [24] and Fetch metadata headers [23] to further eradicate CSRF attacks. However, most of these defenses are built to ensure session binding within the workflows executed in a single website and may not be effective in SSO processes that are cross-site by construction. For instance, the recent SameSite cookies solution, enforced by major browsers like Chrome since 2020, aims to prevent certain cookies from being sent along with cross-site requests so to invalidate these requests. But SSO processes strongly rely on cross-site requests and thus such a defense, if not disabled for the specific SSO related cookies, would simply break the execution of the SSO processes. Protecting an SSO process against CSRF requires a careful combination of specific CSRF defenses built for the SSO process itself (e.g., the state parameter within the OAuth 2.0 protocol [21]) together with standard CSRF defenses preventing the SSO process from being unintentionally executed. In this paper, we demonstrate that CSRF is far from being a solved problem for cross-site scenarios such as those implemented with SSO processes. On the contrary, we advocate the importance of (i) raising more awareness for CSRF issues in these scenarios as well as (ii) designing dynamic security testing techniques to support testers in detecting these issues. In our study, we focus on the SSObased account linking process (SSOLinking, in short) and on an overlooked CSRF attack vector for SSO which was introduced by Rich Lundeen at the BlackHat conference in 2013 [22]. SSOLinking allows users to link their accounts at Service Provider (SP) websites, to their IdP account. This enables these users to authenticate at SPs through an IdP account, thereby eliminating the need to maintain a separate set of credentials for each SP. This added advantage encourages SPs to support the SSOLinking process. In fact, our experiments indicate that around 13% of the top 200 SP websites implementing the standard SSO login also support SSOLinking. The CSRF attack vector we focus on in this paper relies on two different CSRF vulnerabilities, one affecting the IdP and the other the SP. The first vulnerability is an Authentication CSRF [34]—whose well-known instance is Login CSRF [5]—at the IdP. This vulnerability enables an attacker to authenticate a victim into an attacker-controlled account at the IdP. The second one is a CSRF on the action initiating the SSOLinking process at the SP (hereinafter referred to as SSOLinking-Init-CSRF). By combining both vulnerabilities, an attacker can craft an exploit web page that authenticates the victim to an attacker-controlled IdP account and initiates the SSOLinking process to connect the attacker’s IdP account to the victim’s SP account. Upon successful exploitation, the attacker will be able to authenticate to the victim’s SP account by executing the SSO login at the SP through the attacker’s IdP account. By doing so, the attack leads to a complete account hijacking, granting the attacker full control over the victim’s SP account. Our analysis shows that 43.7% of the SPs implementing the SSOLinking feature are vulnerable to SSOLinkingInit-CSRF. By combining the two vulnerabilities, an attacker can execute an account hijack of the victim at the SP. In this paper, as mentioned earlier, we will refer to this attack as SSOLinking Account Hijack. Although this attack was discussed by Lundeen et al. in 2013 [22], it was not considered in past research that assessed the implications of CSRF attacks on the SSOLinking process (e.g., [36], [19], [37], [34]). This may have led to the lack of awareness of this attack. We inspect the two CSRF vulnerabilities and design a testing strategy to detect them. Considering that the number of IdPs is significantly less than that of SPs, that a handful of very popular IdPs is serving the majority of SPs [13], and that IdPs are quite reluctant to fix Authentication CSRF, we believe it is not worth to invest on implementing an automated testing technique for IdPs. We rather opted to test manually the four most popular IdPs identified in [13] against Authentication CSRF and found out that three of them are vulnerable. Under the assumption that IdPs may be vulnerable to Authentication CSRF, we developed an automated testing technique for SSOLinking-Init-CSRF. The technique relies on existing technologies [8] that we extended to support our testing strategy. The goal is to empower a tester at the SP, the party that is impacted by the account hijack, with a security testing technique to detect the issue at the SP side, without requiring strong security expertise. Indeed, the tester shall only provide a Selenium script that executes the SSOLinking process with a specific IdP. Writing this kind of scripts is common practice for web applications as they act as simple integration tests, validating that the process executes as expected. Our security testing technique, which we implemented into the SSOLinking Checker prototype, runs and observes the execution of this script and performs additional automatically-generated tests in the background to check whether the process is vulnerable. When a vulnerability is detected, the SSOLinking Checker also generates an HTML Proof-of-Concept (HTML POC) webpage that allows the tester to easily reproduce the issue. This webpage is used both by SSOLinking Checker to perform the attack and as a proof for the SP to test its service. For each major IdP, we select 12 popular SPs featuring the SSOLinking process and run SSOLinking Checker against them. The results are quite alarming: 21 out of 48 selected unique pairs SP-IdP are vulnerable to SSOLinking-Init-CSRF. We reported our findings to all the affected vendors. We are still interacting with some vendors to clarify the issues. Some of them already confirmed (and patched) the vulnerability we reported and we received some monetary rewards in the context of bug bounty programs. These results confirm our conjecture concerning the lack of awareness and the need of security testing techniques, like those implemented in the SSOLinking Checker, to support testers. To summarize, the contributions of this paper are as follows: 1) We propose an approach to assist a tester at the SP to automatically detect vulnerabilities enabling the SSOLinking Account Hijack. We have implemented our approach into the SSOLinking Checker prototype, based on the Burp Suite [12]. 2) We demonstrate the effectiveness of our approach and the pervasiveness of the SSOLinking Account Hijack by performing an experimental analysis against a selection of popular websites that offer the SSObased account linking process with major IdPs (48 unique pairs SP-IdP). 3) We detected and reported many CSRF vulnerabilities on prominent SPs and IdPs, enabling the SSOLinking Account Hijack: •3 major IdPs (namely Google, Facebook and LinkedIn) are vulnerable to Authentication CSRF; •21 unique pairs SP-IdP (out of 48 analyzed) are affected by SSOLinking-Init-CSRF, including Goodreads, Naver and Workable; •We responsibly disclosed our findings and received some monetary rewards. 4) We noticed that the large majority 19/21 of the vulnerable SPs had protected their SSOLinking flow from the popular CSRF attack involving the absence (or incorrect validation) of the OAuth 2.0 state parameter. This finding indicates the importance of studies like ours to spread awareness of the SSOLinking Account Hijack. Structure of the paper. In Section 2, we introduce some background. In Section 3 we discuss the SSOLinking process and the SSOLinking Account Hijack. Then, in Section 4, we present our approach to support SP testers. In Section 5, we describe the experimental analysis on popular IdPs and SPs, by detailing the dataset selection procedure, the methodology we followed for the analysis and the results. In Section 6, we delve into various mitigations for the prevention and detection of the SSOLinking Account Hijack. In Section 7 and Section 8, we present the responsible disclosure process and some limitations of our work, respectively. Section 9 presents some related work. We conclude and overview future work in Section 10. 2. Background This section provides background knowledge on various types of CSRF attacks, widely-used defenses to prevent them, the testing framework we used, and obstacles to test automation. 2.1. CSRF Attacks The Cross-Site-Request-Forgery (CSRF) attack is also known as sea surf or session riding attacks, taking advantage of the inherent statelessness of the web to simulate user actions from one website to another. Typically, CSRF is used to perform actions on behalf of the attacker using the victim’s authenticated session. If a victim is logged into a website, an attacker can compel the victim’s browser to execute actions on the same site by sending a forged HTTP request. In the context of this work, our focus is on cross-sites, as it involves two entities. A successful CSRF attack can be detrimental to both businesses and users, potentially leading to unauthorized fund transfers, password changes, identity theft, data theft, and the theft of session cookies. In the past, CSRF attacks were thought to only affect state-changing actions caused by authenticated users due to the assumption that only authenticated users can perform high-impact actions, such as making purchases from one account to another. In [5], Barth et al. introduced the concept of Login CSRF, where an attacker tricks a victim’s browser into sending an HTTP request to the authentication endpoint of a website with the attacker’s credentials. As a result, the victim gets authenticated as the attacker. In this scenario, the attacker can monitor the actions performed by the victim on the vulnerable website, allowing the attacker to steal sensitive information from the victim. The authors provided examples from Google and PayPal to illustrate the impact of this attack. More recently, several variants of the Login CSRF attack have been reported in the Single Sign-On (SSO) domain [9], [3]. In [34], the authors classified CSRF attacks into two categories: (i) Non-Authenticated CSRF attacks, which do not require the victim to have an authenticated session with the vulnerable website, and (ii) Authenticated CSRF attacks, which are variants that necessitate the victim to have a valid authenticated session. 2.2. Defenses for CSRF While there are many ways to defend websites against CSRF attacks [20], four primary approaches stand out: token-based, fetch metadata-based, SameSite cookies, and user interaction-based. Token-Based. It helps a web site to maintain session integrity by using a secret token. The token is a unique, non-guessable value, cryptographically bound to the session identifier (e.g., using the Set-Cookie HTTP header), that is generated by the server-side application and transmitted to the browser in such a way that it is included in all the HTTP request made by the browser. In case the HTTP request contains a not valid token, the server-side application rejects the request. The usage of CSRF tokens can prevent CSRF attacks by making it impossible for an attacker to build a fully valid HTTP request suitable for feeding to a victim user. Since the attacker cannot determine or predict the value of a user’s CSRF token, the attacker cannot build a request with all the parameters that are necessary for the application. In the case of the SSO scenario and more specifically in OAuth 2.0, the usage of the state parameter if seeded with a secure random, should avoid CSRF attacks. As reported in the OAuth 2.0 standard [21], the state parameter is defined as “an opaque value used by the client to maintain state between the request and callback. The authorization server includes this value when redirecting the user-agent back to the client”. Fetch Metadata-Based. The Fetch Metadata request headers stand as a cutting-edge advancement in web platform security, designed to equip servers with potent defenses against cross-origin attacks. These headers, denoted as Sec-Fetch-*, provide vital contextual data about an HTTP request, granting the receiving server the ability to proactively enforce security measures before handling the request. This feature enables developers to make informed decisions about whether to approve or decline a request, taking into consideration its origin and intended scope. This methodology guarantees that only legitimate requests originating from their own application receive responses, thereby bolstering the overall security posture. There exist other header-based CSRF defenses including Referer and Origin header validation (interested readers can refer to [29], [5], [16], [34]). SameSite Cookies-Based. SameSite is a cookie attribute defined in [14] that enables servers to specify whether a cookie should be included in cross-site requests. It offers three options: Lax,Strict, or None. When set to Strict, the cookie is withheld from all cross-site requests, even those initiated by regular links. On the other hand, the default Lax mode permits cookies for regular links from external sources but restricts them for CSRF-prone request methods like POST. Lax mode exclusively allows crosssite requests through top-level navigation and safe HTTP methods. Additionally, when set to None, it specifies that the cookie should be sent in all contexts. User Interaction. As proposed in [3], this defense should recognize whether a request intentionally comes from the user or not. To do so, it requires involving the user in additional steps before accepting the request. These steps could include asking the user to perform authentication again, solving a captcha, or requesting consent at the IdP side for every SSO request. This type of defense can effectively prevent CSRF attacks but may have a high impact on the user experience. 2.3. Micro-Id-Gym Micro-Id-Gym1(MIG) is an open source tool [8] used to assist system administrators and testers in the deployment and pentesting of Identity Management protocol deployment. In particular, the tool provides a plugin to support pentesting activities on SSO implementations. The pentesting tool is based-on Burp2web proxy and provides a good set of APIs to interact with. However, as we delve into fully automated testing, we encounter 1. https://st.fbk.eu/tools/Micro-Id-Gym.html 2. https://portswigger.net/burp specific challenges, including dealing with obstacles such as captchas. 2.4. Obstacles to test automation In web applications operating in production environments where captchas are enforced, automated UI testing can present significant challenges. This is particularly evident when captchas disrupt automated penetration testing activities. The concept of captchas inherently clashes with automation, as it is primarily designed to prevent automated bots from executing actions within the application. Similarly, multi-factor authentication (MFA) often requires user involvement, adding an extra layer of complexity to automated testing. These are common obstacles faced in automated testing. While these challenges may exist within production systems, it is less likely that they are present in the testing environments utilized by software testers at the SP. Otherwise, these systems would be not very testable. To work around these issues, we manually address them, assuming that testing systems on these websites do not feature the same automation hurdles. 3. SSO-based Linking Account The SSO-based account linking process (SSOLinking, in short) allows users to link their accounts on a website (SP) to an SSO account they own at an IdP. As a consequence, users can log in on SP by leveraging popular IdPs for the authentication, while keeping their existing profiles on the SP. This ensures a SSO experience, besides the traditional form-based login. Four entities take part in the protocol: the user U, controlling a web Browser, the SP and IdP. The main steps of the protocol are reported in Figure 1 and detailed below. (1-3) At first, U logs in on the SP website, obtaining an authenticated session identifier on the SP website (say cookie(U,SP)). The big arrows in Figure 1 represent the multiple steps required for the login process. (4) U opens the SSOLinking page on SP, selects the preferred IdP and clicks on the account linking button (5) Browser sends an HTTP redirection request at SP with cookie(U,SP) (6-7) SP redirects Browser to the IdP (8) IdP provides the cookie policy and asks for the credentials (9-10) U accepts the cookie policy and types the credentials (11) IdP asks for the U’s consent of sharing information with SP to link the U account at SP with the U account at IdP (e.g., the one reported in Figure 2) (12-13) U provides the consent (14-15) IdP redirects Browser to the SP making the user’s browser send authentication data of the user on IdP (say code(U,IdP)) to the SP web site. (16) SP confirms the linking of the U account at IdP, identified by code(U,IdP), with the U account at SP, identified by cookie(U,SP) At the end of the process, SP links the user’s IdP account with the existing user’s SP account. This enables the user to login (via SSO) to the SP account using the IdP account. Two variants of the protocol are possible. In case Browser has an authenticated session identifier on the IdP web site (say cookie(U,IdP)), in Step 7, cookie(U,IdP) is sent from Browser to IdP in the header of the request. As a consequence, steps 8-10 (marked with a red dashed rectangle in Figure 1) are not performed. Similarly, in case the consent has been already provided by the user in a previous execution, then steps 11-13 (marked with a blue dotted rectangle in Figure 1) are not performed. Finally, what is presented here is the classical process, but there may be variants in the implementations on SPs. For instance, some SPs may use captcha and similar mechanisms to further protect the authentication phase. Moreover, as we will detail in Section 5.3, in some SPs, the account linking button does not trigger a redirection request to SP (i.e. step 5 of Figure 1), but a direct request to IdP, generated by using JavaScript from the IdP library at the SP side. 3.1. SSOLinking Account Hijack In this paper, we focus on a severe account hijack (hereafter SSOLinking Account Hijack), reported in [22], that can be performed in the SSO-based account linking depicted in Figure 1. We consider a classical web attacker, which (i) owns/controls a (malicious) website (attack website), (ii) is capable to forge HTTP requests from the victim’s web browser (e.g., by making the victim visit a link), and (iii) creates accounts at the target SP and the target IdP. In addition, we assume that the victim user has an active session on SP (obtained by authenticating themselves at the target SP through the traditional form-based login). Then, we consider the following assumptions on SP and IdP: (A1) IdP suffers from an Authentication CSRF vulnerability (a.k.a. login-CSRF, i.e., due to a missing CSRF protection, an attacker can force-login the victim to an attacker-controlled account on the IdP) and, (A2) IdP-seamless: IdP has the SSO flow seamless (i.e., no IdP interface is shown) if the user has already logged in at the IdP. This means that steps 8-13 inside the red and blue rectangles in Figure 1 are not performed. (A3) SP suffers from a SSOLinking-Init-CSRF vulnerability, namely the account linking button at the SP is vulnerable to CSRF. (A4) SP enables U to login (via SSO) using the IdP account. The steps to reproduce the attack are as follows. The attacker aplays the role of U and executes SSOLinking (by using its own credentials). In particular, it provides the consent to link an account at SP with its own account at IdP (step 12). Then the attacker unlinks his IdP account on his SP account. The victim v—playing the role of U and having an active session with the SP (i.e. Browser owns cookie(v,SP))—visits the malicious website owned by the attacker and clicks on: Figure 1: The SSO-based account linking. •a link (exploiting the Authentication CSRF vulnerability at IdP and IdP-seamless) to be (transparently) authenticated in the IdP as the attacker a(thus, setting cookie(a,IdP) on Browser of v); •a link (exploiting SSOLinking-Init-CSRF vulnerability at SP) to send an HTTP redirection request to SP. As a consequence, Browser automatically includes cookie(v,SP) (step 5, Figure 1). Then, SP redirects Browser to IdP (step 7) and Browser sends cookie(a,IdP) to IdP (variant of steps 8-10) As a result, the attacker aaccount at IdP, identified by code(a,IdP), is linked with the victim vaccount at SP, identified by cookie(v,SP). In case (A4) holds, namely SP enables U to login via SSO using the IdP account (that is the attacker account), the impact of this attack is severe because the attacker’s level of access to SP is identical to the legitimate user’s. The attack results in a complete account hijack, meaning the attacker gains full control over the victim’s account. In addition, the attack remains “invisible”: while preparing the links, the attacker should properly set the values of the target attribute (e.g., _blank), telling the browser to open the link on new (tiny) windows, which may easily go unnoticed by the user. Also, as revealed by our analysis, usually SPs neither notify users about other devices or active sessions nor allow users to see all the active sessions. 3.2. Comparison with another SSOLinking Account Hijack It is interesting to note that there is a similar CSRF attack vector on the SSOLinking, leading an attacker to hijack the account of the victim. This attack is much more studied compared to the previous one (for instance described in Attack #10 in [34] and in [3]. The steps to reproduce this attack are as follows. The attacker a plays the role of U and executes SSOLinking (by providing its own credentials) until step 14 of Figure 1. Then, aintercepts the Authorization code response with code(a,IdP) and forces a victim user v(when the victim visits an attacker-controlled website) to send it to SP. Assuming that vhas an active session with SP, Browser sends cookie(v,SP) to SP as well (cf. step 15). As a result, the account of aat IdP is linked with the account of the victim vat SP. This is exactly the same result of the attack reported in Section 3, but this attack exploits an incorrect CSRF protection in SSO callback endpoints (cf. Sections 2.2 and 9), while the previous one exploits a missing CSRF protection in the SSO initiation request (besides the Authentication CSRF affecting the IdP). The need of CSRF protection in SSO callback endpoints is well-known and SSO protocols already offer solutions for that (e.g., the state parameter in OAuth 2.0). On the contrary, missing CSRF protection in the initial request of SSOLinking, leading to SSOLinking Account Hijack, is understudied, as also pointed out by our study. 4. Our Approach Exploiting the SSOLinking Account Hijack presented in the previous section relies thus on two vulnerabilities: an Authentication CSRF on the IdP side and an SSOLinking-Init-CSRF on the SSO-based account linking process on the SP side. The Authentication CSRF has been studied in various works [5], [34], [6] and many of such vulnerabilities have been reported, also to IdPs. However, in most of the cases the IdP opted to not fix such a vulnerability to allow users to make use of the oneclick login feature. If Authentication CSRF vulnerabilities are not fixed, then investing further effort in developing testing techniques for this is not so worth. We validated Figure 2: Consent dialog shown by an IdP (Google) for an SP (Goodreads). these two hypothesis—(I1) IdPs tend to be vulnerable to Authentication CSRF and (I2) IdPs tend to be reluctant to fix these vulnerabilities—in a specific experiment over the four major IdPs identified in [13]. The results of this experiment are detailed in Section 5.3.1 and confirm our hypothesis. The CSRF vulnerabilities underlying our attack vector, when analysed in isolation within one party’s (either SP’s or IdP’s) boundary, may be judged as less-critical. For instance, an IdP may consider Authentication CSRF vulnerabilities as low priority as the victim will likely notice that they are logged in at the wrong account. Similarly, the SP may consider SSOLinking-Init-CSRF as non-critical, as the IdP will show a consent dialog (similar to the one shown in Figure 2) and prevent the SSO flow from proceeding (for previously-unlinked IdP accounts). The underestimation of these vulnerabilities has led to an interesting and dangerous scenario where IdPs have been introducing features that aid Authentication CSRF and SPs have not been patching SSOLinking-InitCSRF vulnerabilities. For instance, popular IdPs including Facebook, Instagram, and LinkedIn have introduced the one-click login feature that allows users to login to their IdP account by visiting a URL. These URLs are ideal payloads for Authentication CSRF attacks. In our approach we thus assume that websites and, more specifically, IdPs can be vulnerable to Authentication CSRF. Under this assumption we design a testing technique to detect the SSOLinking Account Hijack on the SP side. Manual testing for this attack requires security expertise, is error-prone, and cannot be efficiently repeated on regular basis. In-line with CI/CD best-practices, we believe website developers conduct regular functional regression tests, anytime changes are implemented. Our vision is thus to empower developers with automated security tests upon the functional ones: developers create a functional test for SSOLinking, and SSOLinking Checker can then execute the security tests and accurately report vulnerabilities. Figure 3 presents our approach. The Tester at the SP provides as input a Selenium script (s1) whose user actions (e.g., click on a button) are run to execute SSOLinking. The HTTP messages (requests and responses) generated while executing the user actions are collected by the Proxy and retrieved by our checker via the Proxy API. Next, the “Attack steps inference” module infers, from the execution of (s1) and from the retrieved HTTP messages, the attack steps to be run to detect the SSOLinking Account Hijack at this specific SP. These attack steps are saved as a new Selenium script (s2). Among the actions in (s2) some relate with the IdP used for SSOLinking. These are created by the “Attack steps inference” module by means of the “IdPs metadata”, a little database populated offline by us with the metadata of the major IdPs and easily extendible for other IdPs. Also, one of the action in (s2) requires an HTML form to send a specific cross-site HTTP request to the SP. This form is generated by the “HTML POC Generator” module. If the attack steps in (s2) are all successfully executed, then the SP is reported vulnerable to the Tester. Besides presenting the verdict for the attack (successful or not), the output also comprises a proof-ofconcept of the attack (if the attack was successful), and a log file with the entire HTTP traffic. This information enables the Tester to drill-down into the details of the test as well as to reproduce the attack. We provide hereafter more details about the key parts of our approach and its implementation within our SSOLinking Checker. 4.1. Input Selenium Script The input Selenium script (s1) is just a UI integration test that testers are used to create for web applications [32]. Indeed, (s1) just comprises the user actions to functionally execute with the Browser the entire SSOLinking process in an automated way at anytime, so to be certain the process is working as expected. We can decompose (s1) in 4 main parts: Login at SP, Link accounts, Assert, and Unlink accounts. Let us discuss each one of them, also presenting some real examples. Login at SP. This comprises the user actions to login a testing user at the SP (cf. step 1 in Figure 1). We illustrate an example for the Daily Mail website in Listings 1. The Daily Mail login page is opened and the cookie policy accepted. Then email and password are filled-in and the login button clicked. 1open | https://www.dailymail.co.uk/registration/login.html |//opens login page 2click | xpath=/html/.../div/button[2] | //accepts the cookie policy 3click | xpath=/html/.../div[2]/input | //clicks on email text field 4type | xpath=/html/.../div[2]/input | [email protected] //types the email 5click | xpath=/html/.../div[3]/input | //clicks on password text field 6type | xpath=/html/.../div[3]/input | 12345678 //types the password 7click | xpath=/html/.../div[5]/button | //clicks the login button Listing 1: Login at SP - Daily Mail example. Link accounts. This includes the user actions to link the testing user account as authenticated at the SP with the one at the IdP (cf. steps 4, 9, and 12 in Figure 1). We continue the example for the Daily Mail website (SP) when linking accounts with Twitter (IdP) in Listings 2. The Selenium commands are identical to the ones in the previous listing and the comments describe each user action. Note, however, that the credentials of a testing user at the IdP are provided within the actions and easily retrievable in our approach. We use <email-idp> and <password-idp> to refer to these retrieved credentials. 1open | https://www.dailymail.co.uk/registration/profile/edit.html | //opens account page 2click | xpath=/html/.../li[1]/a | //clicks on Twitter link account button (a) Architecture (b) Process (c) Collection of IdPs metadata. Figure 3: High level view of our approach. 3click | xpath=/html/.../button[2] | //accepts Twitter cookies 4click | id=email | //clicks on email text field 5type | id=email | [email protected] //types the email of user on IdP, <email-idp> 6click | id=pass | //clicks on password text field 7type | id=pass | abcdefgh //types the password on IdP, <password-idp> 8click | id=loginbutton | //clicks the login button 9click | xpath=/html/.../div | //clicks to give consent to link accounts Listing 2: Link accounts - Daily Mail with Twitter. Assert. Now that the linking between accounts should be done, the Tester wants to be sure this was completed successfully. Assertions provide an easy way to set the expectations for the UI test. Listing 3 presents the assertions for our example on linking accounts between Daily Mail (SP) and Twitter (IdP). The idea is very simple: the account profile is accessed and the fact that the linking button with Twitter is not there allows to derive that the linking was already done. 1open | https://www.dailymail.co.uk/registration/profile/edit.html | //re-opens account page for @assertion purposes 2assert not clickable | xpath=/html/.../div/a | //verifies the not availability of the account linking button with Twitter Listing 3: Assert - Daily Mail with Twitter. Unlink accounts. The Input Selenium Script shall be repeatable, so to enable the Tester to test the process anytime. To do so (s1) is completed with a reset process to unlink the accounts. Indeed, if the accounts stay linked, then the previous parts of the script would simply fail. Listing 4 completes our script example and illustrates the user actions to unlink the user account at Daily Mail from the account on Twitter. 1open | https://www.dailymail.co.uk/registration/profile/edit.html | //re-opens account page 2click | xpath=/html/.../div/a | //clicks unlink account button for Twitter 3click | xpath=/html/.../span/a[2] | //confirm to unlink the accounts Listing 4: Unlink accounts - Daily Mail with Twitter. 4.2. IdPs metadata Our approach relies on a few IdP-related metadata. This is an offline activity that requires little effort and can be shared among few researchers to collect metadata for many IdPs (cf. Figure 3-(c)). 1host: twitter.com 2redirect: [/o/oauth2/auth] 3login: 4open | https://www.twitter.com | //access Twitter 5click | xpath=/html/body/div[3]/.../button[2] | //clicks to authenticate 6click | id=email | //clicks on email text field 7type | id=email | <email-idp> //types the email 8click | id=pass | //clicks on password text field 9type | id=pass | <password-idp> //types the password 10 click | id=loginbutton | //clicks the button Listing 5: IdP metadata - Twitter. The procedure to support new IdPs is straightforward and amount to append in the IdPs metadata database three IdP information. These information are described hereafter and a complete example is presented in Listing 5 for the Twitter IdP. (1) IdP Host (host). Simply the host of the IdP. (2) IdP Redirection Signature (redirect). This is the list of pattern signatures that our approach uses to detect the redirection response that the SP generates to reach a specific IdP host. In most of the cases, this metadata entry will be something like dialog/oauth. It is referred as macro <IdP_redirection> in lines 3 and 5 of Listing 6. (3) IdP Login (login). This comprises the Selenium actions to login at the IdP. It is referred as macro <IdP_login_act> in Listing 6. The user credentials are not hardcoded, but they are rather specified as macros. We recall that our approach automatically extracts them from the “Link Account” fragment of (s1). 4.3. Infer the attack steps While the Selenium script (s1) is executed, the “Attack steps inference” module infers the attack steps to be run to detect the SSOLinking Account Hijack at a specific SP and saves them in a new Selenium script (s2). The inference relies on a new component MIG-L that we introduced to extend the open-source platform MIG [7]. MIG-L defines a specification language we extended to execute tests in MIG where, e.g., new Selenium scripts can be created and run on-the-fly by processing the execution of other Selenium scripts. 1run($s1) //runs Selenium script (s1) 2// Inference begin 3mark($L,<IdP_redirection>) 4save_act([1,index($L)),$SP_login_act) 5save_msg(<IdP_redirection>,$SP_acc) 6save_host($SP_acc, $IdP_host) 7save_poc($SP_acc,$POC) 8mark($A,contains{assert, @assertion}) 9save_act($A,$SP_assert) 10 add($s2,$SP_login_act) 11 add($s2,get_idp_login($IdP_host, <IdP_login_act>)) 12 add($s2,"open | http://localhost/"+$POC) 13 add($s2,"click | id=SSOLinking") 14 add($s2,"wait | 2000") 15 add($s2,$SP_assert) 16 // Inference end 17 run($s2) //runs Selenium script (s2) 18 result($s2) //compute the results: vulnerable if (s2) successfully executed Listing 6: Test procedure for SSOLinking Account Hijack with inference procedure. For instance, Listing 6 is the procedure written in MIG-L to test the entire SSOLinking Account Hijack at any SP (commands are presented in red, macros in purple, and comments in blue). The first run command just indicates that the Selenium script (s1) is executed. Then, the inference procedure starts and proceeds as following line by line: Line 3. The inference procedure observes the HTTP messages and marks with Lthe user action of (s1) whose corresponding HTTP messages comprise the value of the macro <IdP_redirection>. The macro is extracted from the IdPs metadata database. Line 4. All the user actions in (s1) from the first till the one marked with Lare saved into the variable $SP_login_act. Note that closing round bracket indicates that the action marked with Lis excluded. All these actions capture the SP login activity. Lines 5-7. The first HTTP message request whose HTTP response comprises the macro <IdP_redirection> is saved into variable $SP_acc. The host to which the response is redirected is saved into $IdP_host and then an HTML POC is generated by the “HTML POC Generator” module. This HTML POC, whose filepath is saved into variable $POC, enables later on the sending of a cross-site request mimicking the original one (more details in Section 4.4). Lines 8-9. All user actions in (s1) that contain any of the keywords assert or @assertion as a comment are marked with Aand then saved in variable $SP_assert. Lines 10-15. The new Selenium script (s2) can now be built. First, all the user actions related to SP login (saved into $SP_login_act) are added. Then, the user actions related to IdP login are added. These are simply extracted from the IdP metadata by selecting the right IdP host $IdP_host. We refer to these actions as <IdP_login_act>. Then three Selenium commands are also appended into (s2): the first will open the HTML POC local page (notice that the path depends on $POC), the second will run the cross-site request, and the third will just add some waiting time to ensure the request is properly processed. Last, but not least, the assertions collected in $SP_assert are added to (s2). Now that the inference is completed and (s2) is created, the test procedure simply proceeds in running (s2) in line 17 and checking its result in line 18. If all the Selenium actions in (s2) are successfully executed, then the SP is vulnerable (more details in Section 4.5). 4.4. Generate the HTML POC This module generates a proof-of-concept form able to craft the proper CSRF exploit to probe the SSOLinkingInit-CSRF. In details, it creates a HTML page with a form to send the HTTP request in $SP_acc properly changed. Indeed, according to the HTTP method of $SP_acc, the module builds the page and the HTML form. In case of a GET message, the form makes a simple request to the specific URL, while for a POST message the content-type and body are added. 4.5. Output The output of our testing approach comprises three elements. The first element is the test result that provides the verdict whether the attack was successful or not. If all the user actions in (s1) and those created for (s2) are successfully executed, then the attack is successful. If some user actions of (s2) were unsuccessful (e.g., an unexpected page is displayed, an element which should be clicked is not present in the page), then the attack is unsuccessful. The second element of the output is the HTML POC useful to reproduce the attack (when the attack was reported). The last element is a log file with all the HTTP messages saved during the execution of the test. The key to avoiding false positives and false negatives is by correctly defining the Selenium assertions that determines successful SSOLinking (see Section 4.1). Having said that, some tests may still fail because of network problems, CAPTCHA challenges, and other reasons, but such occurrences are detectable within our approach. 4.6. Implementation We implemented our approach in SSOLinking Checker. It is as an extension of Micro-Id-Gym [7] programmed in Java and uses the API of the widelyused penetration testing tool Burp to perform standard proxy engine operations such as collecting HTTP traffic to search for specific HTTP message, setting proxy rules to alter the HTTP traffic, etc. To specify the procedure to infer the attack steps, we used a JSON version of MIGL, which is an equivalent variant currently supported by MIG of the one described in Section 4.3. The source code, installation guide and tutorial of our prototype are publicly available in [1]. 5. Experimental analysis We demonstrated the effectiveness of our approach and the pervasiveness of the SSOLinking Account Hijack against a selection of popular websites featuring the SSObased account linking process with major IdPs (48 unique pairs SP-IdP). In doing so, our approach discovered 21 vulnerable unique pairs SP-IdP out of the 48 analyzed. In the rest of this section, we detail the selection of the major IdPs and popular websites (see Section 5.1) that we used in our experiments, the methodology used for our experiments (see Section 5.2), and the results obtained (see Section 5.3). 5.1. Dataset selection We selected 4 major IdPs and for each of them 12 SPs supporting the SSO-based account linking process and login via SSO using the IdP account. Our selection leverages the dataset released by [13] in 2018, where the authors crawled the most popular one million websites to infer whether they integrate the SSO login process. This is very valuable information for our study as websites featuring the SSO-based account linking process are clearly a subset of those having SSO login. Even if some of the data from [13] is clearly outdated, this did not impact our study and we were able to select our candidates easily. Listing 7 reports an example of an entry of the dataset from [13]. Each entry specifies in the field sso which IdPs that website is using (if any). For this example, workable.com is inferred to have SSO login with Google and LinkedIn IdPs. 1{ "rank": "4824" 2"url": ["https://www.workable.com/signin", 3"http://workable.com", 4"https://www.workable.com/signup", 5"https://www.workable.com/"], 6"sso": ["google", "linkedin"], 7"redir_from": [] } Listing 7: Entry (excerpt) from [13]. For each IdP mentioned in the dataset we counted how many websites were making usage of them in the 2018 dataset and selected the 4 IdPs with the highest number of occurrences. Table 1 presents the IdPs candidates and the number of occurrences reported for each of them. The selected IdPs for our experimental analysis are then Facebook, Google, Twitter and Linkedin. For each one of the selected IdPs, we then proceeded in selecting 12 SPs having SSO-based account linking with the IdP. We considered one by one the website entries in their ranking order from the 2018 dataset for which one of our selected IdP occurred. We manually checked whether the SSO-based account linking process was implemented by that website with the IdP. As shown in Table 2, we manually analyzed a corpus of 648 SPs to discover 67 websites supporting the SSO-based account linking process with the selected IdPs. Among the 67 SPs, we discarded 19 websites due to web issues that occurred in the execution of the process (14) and automation detection issues (5). With web issues, we refer to issues in the SP website that leads to e.g., broken pages and links during the account linking or unlinking processes. For automation detection issues we mean those SPs that e.g., identify the use of Selenium libraries to automate browser actions as a bot issue, which denies access to the website. The SPs selection procedure is successfully finished by obtaining 12 fully working SPs for each selected IdP. Some of these SPs are used in our dataset for more than one IdP. For instance, sp4 occurs in our dataset as paired with Facebook, Google, and Twitter IdPs. In total, we tested 48 SSOLinking processes (i.e. unique pairs SP-IdP) including 37 different SP websites. 5.2. Methodology Here we describe the methodology we followed to perform the experiments on our selected dataset. TABLE 1: IdP candidates. IdP Name Occurrences Facebook 43,333 Google 26,186 Twitter 12,465 LinkedIn 3,939 Amazon 1,459 Yahoo 1,924 TABLE 2: SPs selection per IdP. SPs per IdP SPs with SSOLinking SPs with issues Web Autom. Facebook 61 18 4 2 Google 128 12 0 0 LinkedIn 371 24 9 3 Twitter 88 13 1 0 Total 648 67 14 5 First of all, we focused our attention to validate the two hypothesis (I1) and (I2) we made in our approach for IdPs (cf. Section 4). We manually analysed the four IdPs in our dataset against the Authentication CSRF and discovered that three of them are vulnerable and that they do not plan to fix the problem. More details are given in Section 5.3.1. While analysing the IdPs we also collected the few IdPs metadata required for our approach (cf. Listing 5) which provides a concrete example of the IdP metadata for Facebook. We make available the other IdPs metadata for Google, Twitter and LinkedIn in [1]. The last step in our methodology aims to evaluate the pervasiveness of SSOLinking Account Hijack. In this respect, we run our SSOLinking Checker against the 48 pairs of SPs and IdPs in our dataset. For each pair of SP and IdP, we played the role of a tester at the SP and we created the Selenium script to execute the SSO-based account linking process. In doing so we manually performed the registration of users on the SP websites and we followed the steps described in our approach (see Section 4). For the purpose of the experimental analysis, namely to simplify the manual effort being able to increase the number of analysed SPs, we considered the challenging option to leverage state-of-the-art automation tools to automatically perform the registration of the users and the execution of the SSOLinking for generating the selenium script. Initially, we explored Shepherd [15], a tool for basic website authentication. However, it lacked support for automated registration and relied on credentials available or leaked on the web, rendering it unsuitable for our experiments. Ultimately, we experimented with the xdriver-open [11], a tool also designed to assist in website authentication but the tool was not successful in completing registrations on most websites in our dataset [13]. All the scripts are available in [1]. Our SSOLinking Checker detected the account hijack in almost 50% of our dataset. More details are given in Section 5.3.2, while responsible disclosures and confirmations from vulnerable SPs are discussed in Section 7. All in all, our results indicate that this attack may really be overlooked by the web community. 5.3. Results We report here the results obtained by following the experiments described in our methodology. 5.3.1. Manual testing of the IdPs. We analysed manually the four IdPs in our dataset to validate our two hypothesis: (I1) IdPs tend to be vulnerable to Authentication CSRF; [24] Mozilla. SameSite cookies. https://developer.mozilla.org/en-US/ docs/Web/HTTP/Headers/Set-Cookie/SameSite. [25] OWASP. Cross Site Request Forgery (CSRF). https://owasp.org/ www-community/attacks/csrf. [26] OWASP. What changed from 2013 to 2017? https://owasp.org/ www-project-top-ten/2017/Release Notes. [27] Giancarlo Pellegrino and Davide Balzarotti. Toward black-box detection of logic flaws in web applications. In NDSS, volume 14, pages 23–26, 2014. [28] Victor Le Pochat, Tom Van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczy´ nski, and Wouter Joosen. Tranco: A researchoriented top sites ranking hardened against manipulation. arXiv preprint arXiv:1806.01156, 2018. [29] Open Web Application Security Project. Cross-Site Request Forgery Prevention Cheat Sheet. https://cheatsheetseries.owasp. org/cheatsheets/Cross-Site Request Forgery Prevention Cheat Sheet.html. [30] Ethan Shernan, Henry Carter, Dave Tian, Patrick Traynor, and Kevin Butler. More Guidelines Than Rules: CSRF Vulnerabilities from Noncompliant OAuth 2.0 Implementations. In Proceedings of the 12th International Conference on Detection of Intrusions and Malware, and Vulnerability Assessment - Volume 9148, DIMVA 2015, page 239–260, Berlin, Heidelberg, 2015. Springer-Verlag. [31] Marco Squarcina, Pedro Ad˜ ao, Lorenzo Veronese, and Matteo Maffei. Cookie Crumbles: Breaking and Fixing Web Session Integrity. In 32nd USENIX Security Symposium (USENIX Security 23), pages 5539–5556, 2023. [32] Sristee. Getting started with Selenium: Guide to automated UI testing. https://www.divami.com/blog/ selenium-guide-to-automated-ui-testing/. [33] Avinash Sudhodanan, Alessandro Armando, Roberto Carbone, and Luca Compagna. Attack Patterns for Black-Box Security Testing of Multi-Party Web Applications. In NDSS, 2016. [34] Avinash Sudhodanan, Roberto Carbone, Luca Compagna, Nicolas Dolgin, Alessandro Armando, and Umberto Morelli. Large-scale analysis & detection of authentication cross-site request forgeries. In 2017 IEEE European symposium on security and privacy (EuroS&P), pages 350–365. IEEE, 2017. [35] San-Tsai Sun and Konstantin Beznosov. The devil is in the (implementation) details: an empirical analysis of oauth sso systems. In Proceedings of the 2012 ACM conference on Computer and communications security, pages 378–390, 2012. [36] Rui Wang, Yuchen Zhou, Shuo Chen, Shaz Qadeer, David Evans, and Yuri Gurevich. Explicating SDKs: Uncovering assumptions underlying secure authentication and authorization. In 22nd USENIX Security Symposium (USENIX Security 13), pages 399– 314, Washington, D.C., August 2013. USENIX Association. [37] Ronghai Yang, Guanchen Li, Wing Cheong Lau, Kehuan Zhang, and Pili Hu. Model-based security testing: An empirical study on oauth 2.0 implementations. In Proceedings of the 11th ACM on Asia Conference on Computer and Communications Security, ASIA CCS ’16, page 651–662, New York, NY, USA, 2016. Association for Computing Machinery.