Exploring OpenSearch deployment in Kubernetes
Abstract
The OpenSearch service at CERN has been operating since 2016 on puppet-managed servers. Currently managing 122 OpenSearch and OpenDistro clusters powered by a pool of 156 powerful physical machines running AlmaLinux 9. In the current architecture, multiple clusters live within the same host, resources isolation (e.g., cpu, ram) is achieved on the process level and disk space is managed with LVM. This deviates from the standard OpenSearch deployment. The objective of this project is to get familiar and explore a containerized deployment of OpenSearch using the standard Docker + Kubernetes images on the CERN IT Kubernetes Service.
Full text
August 2024 Exploring OpenSearch deployment in Kubernetes AUTHOR(S): Cristobal Gimenez Devis SUPERVISOR(S): Sokratis Papadopoulos Luis Pigueiras
Exploring OpenSearch deployment in Kubernetes 2 CERN openlab Report // 2024 PROJECT SPECIFICATION We are currently deploying OpenSearch using puppet on bare metal and have developed an ecosystem of code and operations around it. In this project we want to explore the standard way of deploying and operating OpenSearch: Kubernetes.' Firstly, we will create Kubernetes clusters using the CERN IT Kubernetes service and then explore two ways of deploying OpenSearch there: with and without the OpenSearch Kubernetes operator. We will then perform all kinds of operations in our Kubernetes-hosted OpenSearch clusters, like cluster upgrades, scaling and failures and compare the experience of using the operator or not. Another part of the project refers to setting up GitOps using ArgoCD and performing a monitoring mockup using Prometheus and Grafana. Finally, benchmarking has to be performed so that the actual performance of the OpenSearch clusters using different kinds of volume types is measured and compared against the current OpenSearch clusters living in bare metal. As an outcome, a summary of the learning points, should be produced.
Exploring OpenSearch deployment in Kubernetes 3 CERN openlab Report // 2024 ABSTRACT The OpenSearch service at CERN has been operating since 2016 on puppet-managed servers. Currently managing 122 OpenSearch and OpenDistro clusters powered by a pool of 156 powerful physical machines running AlmaLinux 9. In the current architecture, multiple clusters live within the same host, resources isolation (e.g., cpu, ram) is achieved on the process level and disk space is managed with LVM. This deviates from the standard OpenSearch deployment. The objective of this project is to get familiar and explore a containerized deployment of OpenSearch using the standard Docker + Kubernetes images on the CERN IT Kubernetes Service.
Exploring OpenSearch deployment in Kubernetes 4 CERN openlab Report // 2024 TABLE OF CONTENTS INTRODUCTION 01 CLUSTER CREATION 02 DEPLOYMENT 03 GITOPS 04 NODE FAILURES 05 MONITORING 06 BENCHMARK TESTS 07 SCALING AND UPGRADING 08 CONCLUSION 09
Exploring OpenSearch deployment in Kubernetes 5 CERN openlab Report // 2024 1. Introduction In this project, we want to create a prototype deployment of the OpenSearch service in Kubernetes clusters, and explore the different features of such deployment, like performance, flexibility and simplicity of operations. For that, we will need to explore different things': 1. Cluster creation2: First, we will see what do we need to do to create a Kubernetes cluster, and how to operate with it. 2. Deployment2: Next, once we have our Kubernetes clusters, we have to deploy OpenSearch, we will explore two options, with the help of the community-made OpenSearch K8s operator, and with only pure Helm charts. 3. GitOps2: Everytime we need to change something about the cluster, we would need to manually operate with it, We will explore GitOps integration to use CI/CD pipelines in order to synchronize changes automatically. 4. Failures2: In this section, we will simulate node failures', explore what happens to the OpenSearch cluster and how does it manage itself in order to be available again. 5. Monitoring2: After the development part, we have to monitor the cluster, in order to be on the lookout for failures and metrics. Here, we will explore how monitoring can be integrated. 6. Benchmark tests2: We will also need to make some tests in order to check the performance of the new Kubernetes clusters and compare them to the current deployment. In this section, we will make Benchmark tests to measure performance. 7. Scaling and upgrading2: Here, we will explore how to scale and upgrade nodes. 8. Conclusion2: Finally, we will end the report with a summary and conclusion. 2. Cluster creation To create Kubernetes clusters, we will use Magnum, which is an OpenStack service provided by the CERN IT Kubernetes team to easily create Kubernetes clusters on demand. In our project, we want to test three different types of Kubernetes clusters': Single-master cluster (all VMs)2: A simple Kubernetes cluster with one master node and three data nodes. This cluster is only a prototype and will not be used in production, since the single master is a single point of failure. Multi-master cluster (all VMs)2: A Kubernetes cluster with three master nodes and three data nodes, this type of cluster is a production-like deployment, since it is more robust to failures. Single-master cluster (physical machine)2: A Kubernetes cluster with one master node and three data nodes, the difference in this cluster is that the data nodes are bare metal machines rather than virtual machines. However, this was dropped early in the project since the creation of a cluster with physical machines requires extra steps, operating with them is also not as easy and some extra issues that we will see in this section. To create a Kubernetes cluster we will use OpenStack in the CLI': This is the basic command to create a Kubernetes cluster. However, we also need to specify the cluster architecture and some extra configuration, we must also add some labels': $ openstack coe cluster create <name>
Exploring OpenSearch deployment in Kubernetes 6 CERN openlab Report // 2024 --cluster_template <template> : The Kubernetes version which will run in the cluster, in our case we will use version 1.30.1-1 for single-master clusters and 1.30.1-1-multi for multi-master clusters. --keypair <keypair> : This flag associates an RSA keypair in order to communicate with SSH to the cluster. --node_count <count> : The number of worker nodes. --master_count <count> : The number of master nodes. --flavor <flavor> : The flavor (specifications) of all the worker nodes. All the possible flavors can be seen in OpenStack, in our case we will use the m2.large flavor. --master_flavor : The flavor (specifications) of all the master nodes. All the possible flavors can be seen in OpenStack, in our case we will use the m2.large flavor. --merge_labels : This flag makes the cluster use the labels we specify in the CLI and also the default labels in Magnum. --labels cinder_csi_enabled=<enabled> : This flag is required in order to allow the use CinderCSI persistance volumes. --labels monitoring_enabled=<enabled> : This flag is required to enable monitoring in our cluster (grafana, prometheus). --labels grafana_admin_passwd=<pass> : Initial admin password for grafana dashboard (in case monitoring is enabled) --labels ingress_controller=<controller> : Type of ingress class controller, usually nginx. Before we proceed we must note': The cluster template used is 1.30.1-1 or 1.30.1-1-multi The keypair used is named crgimene CinderCSI and monitoring are enabled, the initial grafana password is "admin" and the default ingress controller is nginx a. Creation of single-master cluster (VMs) To create a single-master cluster we need one master node and three worker nodes (the recommended amount is three worker nodes). For that, we use this command to create a single-master cluster': b. Creation of multi-master cluster (VMs) To create a multi-master cluster we need three master nodes and three worker nodes (the recommended amounts are three master nodes in different availability zones and three worker nodes). For that, we use this command to create a multi-master cluster': $ openstack coe cluster create <name> --cluster_template 1.30.1-1 --keypair crgimene --node_count 3 --master_count 1 --flavor m2.large --master_flavor m2.large –merge_labels --labels cinder_csi_enabled=true --labels monitoring_enabled=true --labels grafana_admin_passwd=admin --labels ingress_controller=nginx $ openstack coe cluster create <name> --cluster_template 1.30.1-1-mutli --keypair crgimene --node_count 3 --master_count 3 --flavor m2.large --master_flavor m2.large –merge_labels --labels cinder_csi_enabled=true --labels monitoring_enabled=true --labels grafana_admin_passwd=admin --labels ingress_controller=nginx
Exploring OpenSearch deployment in Kubernetes 7 CERN openlab Report // 2024 c. Creation of single-master cluster (PHYSICAL MACHINES) To create a single-master cluster with physical machine we need to first, create a cluster with virtual machines, and then add the physical machines after the creation': After creating the cluster, we have one master and one worker node (in this example), now, we have to add the physical machines, in order to do that, we create a new nodegroup in the cluster with the physical machine flavor. Here, the flag --role defines the role of the node inside the cluster and p1.dl8822089.S513-V-IP413 is the flavor of the physical machines. Once this is done, we have our physical machines in the cluster, however, there is one more issue (at least, currently)', the physical machines are not assigned an internal IP, thus making operating with the machine hard, while some things (e.g. ingress) are practically not supported. Due to these issues, working with physical machines was dropped, and the focus became only VMs. 3. Deployment Once we have our cluster setup, we must now deploy OpenSearch on our cluster, for that we will explore two options for the deployment': Using the community-made OpenSearch Kubernetes operator which simplifies a lot of common operations, and adds features for scaling, updates, draining, etc. Using Helm Charts only, in order to not have one extra dependency and for better default GitOps monitoring. a. Deployment with operator In order to deploy OpenSearch with the Kubernetes operator, first, we need to install the operator in our cluster. For that, we will install and deploy the Helm chart of the operator. First, we install the repository containing the chart': Then, we install the operator: $ openstack coe cluster create <name> --cluster_template 1.30.1-1 --keypair crgimene --node_count 1 --master_count 1 --flavor m2.large --master_flavor m2.large --merge_labels --labels cinder_csi_enabled=true --labels monitoring_enabled=true --labels grafana_admin_passwd=admin --labels ingress_controller=nginx $ openstack coe nodegroup create <name_of_cluster> <name_of_nodegroup> --node_count 1 --flavor p1.dl8822089.S513-V-IP413 --role <role> --merge_labels --labels cinder_csi_enabled=true --labels heat_container_agent_tag=trainstable-6 $ helm repo add opensearch-operator https://opensearchproject.github.io/opensearch-k8s-operator $ helm install opensearch-operator opensearch-operator/opensearch-operator
Exploring OpenSearch deployment in Kubernetes 8 CERN openlab Report // 2024 Once the operator is deployed, we must deploy the clusters. The operator defines a CustomResourceDefinition in Kubernetes in order to deploy the clusters. For that, we must make a .yaml file which will contain the specification of both the OpenSearch nodes and the OpenSearch Dashboards. In that file, we must specify several attributes: version: The version of both the OpenSearch Docker image, and the version of the OpenSearch Dashboards. Memory and cpu requests for each pod. replicas: The number of OpenSearch nodes diskSize: Disk space for each of the OpenSearch nodes. roles: The roles of the cluster nodes, can be “cluster_manager”, or “data”. storageClass: The PVC storage class, if CinderCSI is enabled it can be used. An example of this file can be seen in gitlab. b. Deployment with Helm To deploy OpenSearch with only Helm charts, we must install the repositories of both the OpenSearch Helm chart and the OpenSearch Dashboards chart This installs the repositories of the both the clusters and the Dashboards, now what is left is to deploy them': The values.yaml file used when deploying the cluster is a Helm value file, just like when deploying the clusters with the operator we must specify': version: The version of both the OpenSearch Docker image, and the version of the OpenSearch Dashboards. Memory and cpu requests for each pod. replicas: The number of OpenSearch nodes diskSize: Disk space for each of the OpenSearch nodes. roles: The roles of the cluster nodes, can be “cluster_manager”, or “data”. storageClass: The PVC storage class, if CinderCSI is enabled it can be used. Note : Starting from OpenSearch version 2.12.0 forward, it is required to setup an initial admin password, as an environment variable, if we use that version on our clusters we must also specify a password like in this example value file in gitlab. In the case of the operator, there is currently a GitHub issue open because it’s not possible to set this password, which means that we can only use versions lower than 2.12.0. $ helm repo add opensearch https://opensearch-project.github.io/helm-charts $ helm install opensearch-clusters opensearch/opensearch -f values.yaml $ helm install opensearch-dashboards opensearch/opensearch-Dashboards
Exploring OpenSearch deployment in Kubernetes 9 CERN openlab Report // 2024 c. Access OpenSearch Dashboards Once we have our OpenSearch cluster deployed and running we want to access the OpenSearch Dashboards in order to view metrics, use the developer tools, etc. We have two ways of accessing the Dashboards. Port-forward The OpenSearch Dashboards pod exposes the port 5601 to communicate with it, so, that is the port we must port-forward. Then, we can either port-forward the pod directly, or the service which redirects to the pod. Then, we can access https://localhost:5601 to view the Dashboards. Create an Ingress Another way of viewing the Dashboards is creating an Ingress to it. An Ingress is a Kubernetes object that manages external access to the services in a cluster, by exposing a HTTP and HTTPS port So, we can make an Ingress which redirects to the service Dashboards which will communicate with the pod running the OpenSearch Dashboards. To deploy an Ingress we must: Change the Magnum values to allow a nginx ingress. Tag one of the worker nodes as the Ingress. Deploy the Ingress which points to the service. In order to create the Ingress, we need to update the namespace kube-system with new Helm values. We then run: Inside the magnum-values.yaml we add and delete the following: Delete the first line “USER-SUPPLIED VALUES”. Then, in the section ingress-nginx we add: $ kubectl port-forward opensearch-dashboards 5601 $ kubectl port-forward svc/opensearch-service-dashboards 5601 $ helm get values cern-magnum -n kube-system > magnum-values.yaml