Full text
CERN openlab Report // 2023 PROJECT SPECIFICATION In the rapidly evolving landscape of data management and security, the integration of authentication and authorization mechanisms is crucial for ensuring the confidentiality, integrity, and availability of sensitive information. This project aims to assess the feasibility and effectiveness of integrating Apache Knox with the Single Sign-On (SSO) at CERN. The primary objective of this project is to evaluate the integration of Apache Knox, an open-source gateway for securing and centralizing access to multiple Hadoop clusters, with the CERN Single Sign-On infrastructure. The goal is to enhance the security and streamline the authentication process for accessing distributed data from Hadoop and HBase clusters hosted centrally at CERN. The successful integration of Apache Knox with CERN SSO can provide CERN with a robust and secure solution for managing access to distributed data, thereby improving user data security. As an integral part of this project, Apache Knox will be deployed and evaluated within a controlled development environment. This evaluation seeks to ascertain its compatibility with various components of the Hadoop Ecosystem at CERN, including HDFS, Yarn, HBase, MapReduce, and Spark History Server. Special emphasis will be placed on configuring the SSO gateway, integrating with CERN SSO (e.g., with Keycloak and application portal entities), and testing the high availability mode, along with authorization-based API calls to services like the webHDFS REST API. 2
CERN openlab Report // 2023 ABSTRACT This project explores the implementation of Apache Knox to secure various Hadoop web services, specifically HBaseUI, YARNUI, HDFSUI, and the Spark History Server, within the context of a summer internship at CERN. The primary goal was to transition from the existing SPNEGO-based security framework to an Apache Knox-based solution, addressing the limitations of SPNEGO while enhancing the security and user experience for accessing Hadoop services. SPNEGO, which relies on Kerberos tickets for authentication, presents several challenges, including complex browser configurations, inconsistent user experiences, and frequent access issues across different browsers. These issues often lead to a high volume of support requests, burdening the IT support team. Apache Knox, by contrast, offers a single point of authentication and access for Hadoop services, simplifying the security architecture and improving accessibility. Knox's topology-based configuration and customizable rewrite rules allow for flexible protection of services, providing a more consistent and user-friendly experience. However, the project also addresses the complexities of the initial setup, the performance overhead introduced by Knox, and the importance of ensuring high availability of the gateway. The project involved configuring Knox, setting up secure access through CERN's OIDC-based SSO, and protecting key Hadoop services. Additionally, it explored overcoming challenges in securing the Spark History Server, which required custom proxy setups due to CORS-related issues. The implementation was further enhanced by automating the configuration and deployment process using Puppet, and by building RPM packages to streamline future updates and maintenance. Comprehensive monitoring was established to ensure the reliability and performance of the secured services. Lastly, the project concluded with the creation of a detailed poster summarizing the work, which highlights the practical benefits of transitioning to Apache Knox. The results demonstrate that while Knox introduces some setup complexity initially, it significantly streamlines security management and access control for Hadoop services, ultimately offering a more reliable and user-friendly solution than SPNEGO. 3
CERN openlab Report // 2023 TABLE OF CONTENTS INTRODUCTION.........................................................................................................................................5 Introduction.............................................................................................................................................. 5 SPNEGO current state and issues.................................................................................................... 5 Apache Knox as a solution.......................................................................................................................6 Pros and cons of Knox vs SPNEGO........................................................................................................ 8 Pros of Knox:...............................................................................................................................8 Cons of Knox:............................................................................................................................. 8 Pros of SPNEGO:........................................................................................................................9 Cons of SPNEGO:.......................................................................................................................9 Basic setup of Apache knox..................................................................................................................... 9 Binary download................................................................................................................................9 Installation..........................................................................................................................................9 Create the master secret.............................................................................................................10 keystore files..............................................................................................................................10 Creation of the __gateway-credentials.jceks file................................................................10 Creation of the gateway.jks file.................................................................................................12 Running the gateway........................................................................................................................14 Making our first topology file..........................................................................................................14 Setting up the keycloak client (Application portal)...................................................................15 Setting up the knoxsso.xml file................................................................................................. 15 Setting up the gateway-site.xml file..........................................................................................16 Writing our first default.xml file............................................................................................... 16 Remove the SPNEGO conf from core-site.xml and hdfs-site.xml............................................17 Protection of the hadoop cluster services using Knox............................................................................17 HDFS WebUI...................................................................................................................................18 YARN WebUI.................................................................................................................................. 19 Spark history.................................................................................................................................... 20 Spark needs a proxy to work..................................................................................................... 21 4
CERN openlab Report // 2023 HBase...............................................................................................................................................22 Map reduce history.......................................................................................................................... 24 REST APIs.......................................................................................................................................24 Curling the endpoints.......................................................................................................................24 CURL examples........................................................................................................................ 25 Knox rewrite rules............................................................................................................................26 Knox HA..........................................................................................................................................27 ACLs................................................................................................................................................29 Auditing........................................................................................................................................... 29 Metrics............................................................................................................................................. 30 Current Knox architecture...................................................................................................................... 31 Usefull ressources...................................................................................................................................31 JIRA boards..................................................................................................................................... 31 Mailing lists..................................................................................................................................... 32 Cloudera forum................................................................................................................................ 32 INTRODUCTION This documentation serves as a guide to set up Apache Knox to protect HBaseUI YARNUI HDFSUI and Spark history Server in the context of the Summer Internship project of Thomas Mauran. Gitlab repo rpm: here Gitlab repo IT Module branch knox_sso_setup: here Introduction In this documentation, we will explore the setup of Apache Knox to secure HBaseUI, YARNUI, HDFSUI, and the Spark History Server within the context of my summer internship project at CERN. The primary objective is to transition from the existing SPNEGO-based security framework to an Apache Knox-based solution. This transition aims to address the shortcomings of SPNEGO, streamline authentication processes, and enhance the user experience for accessing Hadoop services. 5
CERN openlab Report // 2023 SPNEGO current state and issues Currently the security is handled by SPNEGO (Simple and Protected GSSAPI Negotiation Mechanism) which uses Keberos tickets to make sure users are logged in and have the right to access both WEBUIsand APIs. This system works in theory, but has multiple flaws in practice. Iin the actual way of setting up SPNEGO, in reality every user must modify some policy in its own browser to make SPNEGO work. Those instructions are available here and even though it doesn't seem to be easy at first it's not that easy to do. The major problem being that even if you correctly configured your browser, it might not even work on some services. For example, on Firefox a lot of pages on almost all the services returns errors making them unavailable. To solve this, you basically have to change your browser, which is not handy at all. And as computer scientist these ideas to test it in another browser might be easy to get but for a regular user who took already a lot of time to configure it's browser it seems impossible to figure out what is the source root of the problem, inevitably leading in the creation of support tickets which in the end overloads the managing team with problems they shouldn't have to deal with. Apache Knox as a solution Apache Knox Gateway is a system that provides a single point of authentication and access for Apache Hadoop services in a cluster. It provides a single point of access for all REST and HTTP interactions with Apache Hadoop clusters. 6
CERN openlab Report // 2023 Knox uses a system of topology files allowing you to protect different services. This topology like architecture is pretty easy to maintain and user-friendly after getting a bit of experience with it. In the same direction, one of Apache Knox's great strengths is its customizability. If you have to add another custom REST service you can simply write rewrite rules for it, and it will work out of the box. I had to tweak some of those rewrite rules on some services for the purpose of this deployment and once you understand the process it's easy to debug and customize. 7
CERN openlab Report // 2023 Here is a typical Knox architecture, showcasing the token verification and services behind a proxy Pros and cons of Knox vs SPNEGO Pros of Knox: 1. Single Point of Access: Knox provides a centralized gateway for all Hadoop services, simplifying access management and reducing the complexity of dealing with multiple endpoints. 2. Ease of Use: Once set up, Knox requires minimal configuration on the client-side, eliminating the need for users to modify their browser settings. 3. Customizability: Knox's topology files and rewrite rules allow for flexible and easy customization to protect various services, including custom REST services. 4. Consistent User Experience: Knox ensures a more consistent and reliable user experience across different browsers and services using CERN SSO, reducing the likelihood of access issues. Cons of Knox: 1. Initial Setup Complexity: Setting up Knox for the first time can be complex and time-consuming, requiring a good understanding of the system and its configuration. 2. Performance Overhead: Introducing an additional gateway layer may introduce some performance overhead, potentially impacting the response time of protected services. 8
CERN openlab Report // 2023 3. Dependency on Gateway: All traffic to Hadoop services passes through the Knox Gateway, creating a single point of failure. Ensuring high availability of the gateway is crucial. Pros of SPNEGO: 1. Integrated with Kerberos: SPNEGO leverages Kerberos authentication, which is already widely used and trusted for secure authentication in many enterprise environments. 2. Direct Service Access: Users can directly access services without needing an additional gateway layer, which can simplify the architecture and reduce potential bottlenecks. 3. Low Overhead: Without the need for an intermediary gateway, SPNEGO can potentially offer lower latency and better performance for direct service access. Cons of SPNEGO: 1. Browser Configuration: Users must configure their browsers to support SPNEGO, which can be complex and inconsistent across different browsers and versions. 2. Inconsistent User Experience: Even with proper configuration, SPNEGO may not work consistently across all services and browsers, leading to frequent access issues. 3. Support Burden: The need for user-side configuration and troubleshooting can lead to a high volume of support tickets, placing a burden on the IT support team. 4. Limited Customization: SPNEGO does not offer the same level of flexibility and customizability for protecting and managing access to various services as Knox does. Basic setup of Apache Knox Prerequisites - Make sure you are running on Java 8or 11. - The following examples will use 123456 as a password, obviously do not use that as a password Binary download First, we need to download the binary files available here as Gateway Server Binary 2.0.0. Make sure you then unzip the file and work from inside of it. Installation Once downloaded (or built) you will have access to a ./bin folder containing your binaries. Including: 9
CERN openlab Report // 2023 After signing in to the CERN SSO you should access the following page: If that is not the case make sure your knoxsso.xml file is filled with the correct information and check the logs/ files. Setting up the gateway-site.xml file Here is the working gateway-site.xml that you need to use instead of the one in conf/ gateway-site.xml Writing our first default.xml file We are going to write a new file named default.xml from which we will add all our services later on. Here is an example of default.xml file: default.xml Make sure you change the following values: 16
CERN openlab Report // 2023 - sso.authentication.provider.url -> https://<hostname>:8443/gateway/knoxsso/api/v1/websso - WEBHDFS -> https://<namenode host>:50070/webhdfs Make sure you stop and start the gateway again. To test our configuration we are going to add a folder and a file in our hdfs, in my case /user/tmauran/hello.txt You shall now have access to the following url and get your file downloaded after loging in CERN SSO page: https://<hostname>:8443/gateway/default/webhdfs/v1/<your>/<path>?op=OPE N If that is not the case make sure your default.xml file is filled with the correct information and check the logs/ files. Remove the SPNEGO conf from core-site.xml and hdfs-site.xml To disable SPNEGO we had to comment parts related to it in both /etc/hadoop/conf/core-site.xml and /etc/hadoop/conf/hdfs-site.xml. Here is an example of working core-site.xml In /etc/hadoop/conf/hdfs-site.xml we will comment the following properties: - dfs.web.authentication.kerberos.principal - dfs.web.authentication.kerberos.keytab - dfs.journalnode.kerberos.internal.spnego.principal Protection of the hadoop cluster services using Knox Apache knox got a set of already defined rewrite and services available in data/services. Those are then usable in the topology files as follows 17
CERN openlab Report // 2023 HDFS WebUI <service> <role>HDFSUI</role> <version>2.7.0</version> <url>https://ithdpdev-ekleszcz01.cern.ch:50070</url> </service> To protect each service we need to enable multiple things in the core-site.xml config. First of all make sure you have access to the following homepage UI: https://<hostname>:8443/gateway/homepage/home/ If you get prompted a login page make sure you started the demo LDAP server The default credentials are username: guest password: guest-password You should end up on something looking like that: 18
CERN openlab Report // 2023 You must now download the PEM TLS certificate that we are gonna need in our core-site configuration. You can directly get the certificate using the following command: /usr/hdp/knox/bin/knoxcli.sh export-cert --type pem In core-site.xml add the following things: <!-- Knox Security --> <property> <name>hadoop.http.authentication.type</name> <value>org.apache.hadoop.security.authentication.server.JWTRedirectAuthentic ationHandler</value> </property> <property> <name>hadoop.http.authentication.authentication.provider.url</name> <value>https://ithdpdev-ekleszcz01.cern.ch:8443/gateway/knoxsso/api/v1/webss o</value> </property> <property> <name>hadoop.http.authentication.public.key.pem</name> <value><Your public token from the .pem file></value> </property> You should now not be able to directly access hdfs without getting logged in YARN WebUI <service> <role>YARNUI</role> <url>https://ithdpdev-ekleszcz01.cern.ch:8088</url> </service> To enable YARNUI make sure you defined the YARN service. 19
CERN openlab Report // 2023 To fix the CSS problem you might encounter on the :8443/gateway/default/yarn you need to change this line in data/services/yarnui/2.7.0/rewrite.xml What this does is directly go on the server URL to get the static files. Yarn UI will by default be protected if you applied the HDFS configuration mentioned above in core-site.xml Use the following rewrite rules: rules Spark history <service> <role>SPARK3HISTORYUI</role> <url>https://ithdpdev-ekleszcz01.cern.ch:18080</url> </service> Use the following rewrite rules: rules To secure enable Spark UI redirection to Knox when trying to directly access it you will have to add the following properties in /etc/spark/conf/spark-defaults.conf: spark-default.conf spark.org.apache.hadoop.security.authentication.server.AuthenticationFilter. param.kerberos.principal=<Host principal> spark.org.apache.hadoop.security.authentication.server.AuthenticationFilter. param.kerberos.keytab=<Path to the keytab> spark.ui.filters=org.apache.spark.deploy.yarn.YarnProxyRedirectFilter,org.ap ache.hadoop.security.authentication.server.AuthenticationFilter spark.org.apache.hadoop.security.authentication.server.AuthenticationFilter. param.type org.apache.hadoop.security.authentication.server.JWTRedirectAuthenticationHa ndler spark.org.apache.hadoop.security.authentication.server.AuthenticationFilter. param.authentication.provider.url https://<Knox host>:8443/gateway/knoxsso/api/v1/websso 20
CERN openlab Report // 2023 spark.org.apache.hadoop.security.authentication.server.AuthenticationFilter. param.public.key.pem=<Your pem key> The PEM key is the one you retrieved in the HDFS part. Make sure you remove every \n otherwise you will get an error saying that your PEM key is corrupted. We also need to use the new hadoop-client-api-3.3.6.jar instead of 3.3.4, to do so cp /usr/hdp/hadoop/share/hadoop/common/lib/hadoop-auth-3.3.6.jar /usr/hdp/spark/jars/ #then remove the 3.3.4 one rm -rf /usr/hdp/spark/jars/hadoop-auth-3.3.4.jar rm -rf /usr/hdp/spark/jars/hadoop-client-api-3.3.4.jar Spark needs a proxy to work As Spark History UI was causing some issues when being proxied by Knox (302 CORS error on the jquery api call) we had to put in place 2 small proxies in front of each spark instance. The goal of those are to redirect to either spark instance using the capability of the HTML redirect tag. This mechanism is exactly the one featured in HDFS, and this is why we automatically get redirected to the namenode when clicking on HDFS in the homepage. # Install nginx from yum yum install nginx # From the gitlab content do the following # Copy the nginx.conf in /etc/nginx/nginx.conf # Copy the index.html in /usr/share/nginx/html/index.html # Add the hostnames in /usr/share/nginx/html/spark-hostnames # We save the hostname of the server on the /hostname route for the script 21
CERN openlab Report // 2023 to work hostname > /usr/share/nginx/html/hostname sudo mkdir -p /etc/pki/nginx # Here we could use the custom certificate generated as for the other services but for a matter of time saving we just another one sudo openssl req -new -x509 -nodes -days 3650 -newkey rsa:2048 -keyout /etc/pki/nginx/server.key -out /etc/pki/nginx/server.crt -subj "/C=US/ST=California/L=San Francisco/O=My Company/CN=your_domain.com/emailAddress[email protected]" # Start nginx using nginx So basically the proxy is a simple NGINX server with one index.html containing the following tag: <meta http-equiv="REFRESH" content="0;url=https://hdp-ekleszcz-dev-gateway.cern.ch:18080" /> This tag allows us to force the redirection to the Spark UI, it was inspired of what is done in the Hadoop source code where a similar thing is hapening. Code available here We don't know why spark blocks CORS doing this weird 302 error and here are the forum / mailing list follow-ups for this issue: Cloudera forum -Error with Pac4j Query param -Spark history CORS header access control allow origin -Spark history UI for apache Knox Mailing list threads -Force redirect to a service instead of proxying -Error with Pac4j provider on query param of original URL -Error with Sparkhistory UI HBase Make sure you remove the following properties to disable simply pass the Kerberos UI to simple 22
CERN openlab Report // 2023 <property> <name>hbase.security.authentication</name> <value>simple</value> </property> <property> <name>hbase.security.authentication.ui</name> <value>simple</value> </property> <property> <name>hbase.security.authentication</name> <value>simple</value> </property> Currently, HBase doesn't provide any mechanism to redirect to a JWT provider, this is currently a work in progress but meaning to protect HBase we need to protect it behind the firewall and make it inaccessible from outside. This way, the only working path is to go through Knox proxy. JIRA ticket Make sure you remove port 60010 from the whitelist NFT ports to make it inaccessible from outside. We are gonna replace this 60010 rule with one allowing only the IP of the other nodes to that port. That way the port will be closed online but when Knox tries to proxy request from the master node to the backup node it will be able to do so and display the corresponding UI. To do so we are gonna use this command: nft add rule inet filter default_in ip saddr @cluster_hg_ipset4 tcp dport 60010 accept nft add rule inet filter default_in ip6 saddr @cluster_hg_ipset6 tcp dport 60010 accept Make sure you do it on both master nodes. Then do the same for the region servers. Remove the port 16030 and add those rules: 23
CERN openlab Report // 2023 nft add rule inet filter default_in ip saddr @cluster_hg_ipset4 tcp dport 16030 accept nft add rule inet filter default_in ip6 saddr @cluster_hg_ipset6 tcp dport 16030 accept To add the new service in the topology we need to declare a service with the HBASEUI role as following: <service> <role>HBASEUI</role> <url>https://ithdpdev-ekleszcz01.cern.ch:60010</url> </service> Then the service will be accessible with the following URL, make sure you replace the host and port with the master one, this is currently not very handy but could be reworked in the end. https://<knox-host>:8443/gateway/default/hbase/webui/master?host=ithdpd ev-ekleszcz01.cern.ch&port=60010 In the end HBase rewrite rules will be the trickiest one to maintain, since we not only need to maintain the first page after the quick link access but also all the sub pages of the whole UI. Map reduce history Add this service in the default topology <service> <role>JOBHISTORYUI</role> <url>https://ithdpdev-ekleszcz01.cern.ch:19888</url> </service> As for the other services make sure you update the whitelist port in the gateway-site.xml file to allow 19888. You will also need to change the rewrite files to the one on GitLab as some rules are not correct in the current version REST APIs Curling the endpoints To CURL the different services, we need to use the CERN auth-get-sso-cookie tool. 24
CERN openlab Report // 2023 ⚠ auth-get-sso-cookie doesn't work with account who have 2FA activated, to use it you will need to be Apache Knox adds a cookie called hadoop-jwt when you connect to sso, one thing is that Kerberized Hadoop environment also add another cookie called hadoop-auth that is mandatory to be able to CURL those endpoints. To acquire both of those cookies we are gonna be using the auth-get-sso-cookie on the namenode endpoint which will provide us those 2 cookies auth-get-sso-cookie -vv -u "https://ithdpdev-ekleszcz01.cern.ch:50070/" -o auth here we are storing the cookie in the auth file you can cat this file to make sure you have both hadoop-jwt and Hadoop-auth, if it' s not the case make sure you are using the right Kerberos ticket. kdestroy and kinit a new one on an account which doesn' t have 2FA (hbuilder for example). One you have the auth file you can CURL each of the endpoint using the -Lb flag (-L allows relocation following, and is mandatory while -b allows you to pass a cookie file). CURL examples # Namenode curl -Lb auth "https://ithdpdev-ekleszcz01.cern.ch:50070/dfshealth.html#tab-overview" # Yarn curl -Lb auth "https://ithdpdev-ekleszcz01.cern.ch:8088" curl -Lb auth "https://ithdpdev-ekleszcz01.cern.ch:8443/gateway/default/yarn" # HBase curl -Lb auth "https://ithdpdev-ekleszcz01.cern.ch:8443/gateway/default/hbase/webui/master ??&host=ithdpdev-ekleszcz01.cern.ch&port=60010" # Mapreduce history 25
CERN openlab Report // 2023 Conclusion In conclusion, the project successfully implemented Apache Knox to secure various Hadoop web services at CERN, replacing the existing SPNEGO-based security framework. The transition to Knox addressed several challenges associated with SPNEGO, including complex browser configurations and inconsistent user experiences. By providing a centralized authentication and access point, Knox not only enhanced the security of the Hadoop ecosystem but also significantly improved usability for end-users. The project also demonstrated the importance of a flexible security architecture, leveraging Knox's topology-based configuration and rewrite rules to protect key services such as HBaseUI, YARNUI, HDFSUI, and the Spark History Server. Despite initial setup complexities and the need for custom solutions to handle specific challenges, such as CORS issues with the Spark History Server, Knox proved to be a robust and scalable solution. Furthermore, the automation of the deployment process using Puppet, coupled with the creation of a custom RPM repository via the Koji system, ensured that the Knox-based security setup could be easily maintained and replicated. This automation not only streamlined the deployment but also reduced the potential for errors and configuration inconsistencies, contributing to the long-term sustainability of the solution. On a personal note, this project was an invaluable learning experience. Working in such a dynamic and innovative environment at CERN allowed me to gain deep insights into both security architecture and automation practices. I would like to express my gratitude to my supervisor, Emil Kleszcz, for his guidance and support throughout the project. His mentorship played a crucial role in the successful completion of this work, and I am deeply appreciative of the opportunity to learn and grow under his supervision. Overall, the project achieved its goals of enhancing security, improving user experience, and providing a maintainable and scalable solution for securing Hadoop services at CERN. The successful implementation and automation of Apache Knox mark a significant step forward in the security management of the organization's Hadoop infrastructure. 32
CERN openlab Report // 2023 Useful resources JIRA boards Hadoop JIRA board Knox JIRA board Spark JIRA board HBase JIRA board Mailing lists Archives Cloudera forum Forum 33