Full text
CERN openlab Report // 2024 PROJECT SPECIFICATION This project aims at providing an improved log collection, processing, and storage pipeline for the internal logging events occurring inside all of the OKD clusters, such as PaaS, WebEOS, Drupal and App-Catalogue, where scientists, researchers, and employees can request access to the internal cluster and get resources provisioned to them. Specifically, the project will be divided into three stages: ● Research and evaluation Investigate and document existing solutions and technologies for logging pipelines. Research how other CERN projects approached the problem. Document in detail each one of the shortcomings, advantages and use cases from the previous findings for each of the technologies. Investigate the feasibility of having a solution working for both for OKD and Kubernetes Magnum. ● Small scale implementation Deploy a complete pipeline in a simple but real project, in order to gain knowledge about the potential implementation obstacles, considerations or limitations and getting to know the technologies. Since the deployment will occur in a production environment, it will already be a benefit to the organization. ● Proof of concept Configure a proof of concept logging pipeline solution for OKD deployments in our clusters in addition to Kubernetes Magnum, in case of technical feasibility. 2
CERN openlab Report // 2024 ABSTRACT To address future processing demands, it's essential to enhance logging capabilities by implementing a more robust, efficient, and maintainable system, in replacement of our current Fluentd system. In addition, OKD maintainer RedHat already deprecated Fluentd log collector Operator and plans to phase out and remove it in a future release, providing only bug fixes and support in the meantime. Their official alternative is Vector for the internal log collector Operator, which is our chosen option for researching and proof of concept. Fluentd is an open source data collector and processor used as a general purpose logging tool written in Ruby, under the CNCF. Fluentbit serves the same purpose, also under the CNCF, but is written in C, which makes it faster and more memory efficient. Vector is also a logging collector and processor that serves the same purpose, but is written in Rust and is developed by DataDog. This report details the research and modifications to the internal OKD logging pipeline out of Fluentd towards FluentBit and Vector in order to match the current system requirements, make it more robust and efficient, and also be compatible with future releases of OKD. Also, we will explore ways to simplify the logging pipeline itself, merging different components into one aggregator. Our OKD telemetry pipeline collects logs from containers and Kubernetes itself, parses them, and sends them to OpenSearch, through the Central Monitoring system. Our Kubernetes Magnum clusters implement almost the same pipeline, in addition to enriching the data with metrics about the logs themselves, in order to send them to Prometheus. Therefore, we will try to make our solution cross-compatible for OKD and Kubernetes Magnum. 3
CERN openlab Report // 2024 TABLE OF CONTENTS INTRODUCTION AND CONTEXT 01 RESEARCH AND EVALUATIONS 02 RESEARCH RESULTS 06 PROOF OF CONCEPT 06 CONCLUSION 06 4
CERN openlab Report // 2024 1. Introduction and context a. Telemetry To understand logging, we need to first understand telemetry: it refers to the automated collection, transmission and storage of data from various systems and applications to monitor and analyze their performance, behavior, and health. It is used to gain insights, for compliance, traceability, and system status monitoring for correct IT administration. Over the big umbrella of telemetry, we classify three components: logs, metrics and traces. Firstly, metrics are a measurement of a service captured at runtime. Each metric consists of not only the measurement itself, but also the time at which it was captured. They act as important indicators of availability and performance. Then, traces are a collection of structured logs with context, correlation, and hierarchy. They are useful to give us the big picture of what happens and its path when a request is made to an application. Finally, logs are a time stamped text record, usually structured, from information, debug information, warning and errors. b. Logging In this project, we will specifically focus on logging, given the shorter project timespan limitations, although most of the technologies we analyzed also support traces and / or metrics. Logs often contain detailed debugging or diagnostic info, such as inputs to an operation, the result of the operation, and any support metadata to the operation. The act of logging consists in emitting a log record to any source, being over network, through a file, stdout or any other mechanism. c. How logging tools work All common logging technologies work with the sources →transformers → sinks pattern. It is also called input →filter → output, depending on the logging processor. Figure 1: logging tools generic workflow 5
CERN openlab Report // 2024 In Figure 1, we can see the general pattern of many sources, to one processor to many sinks. The first step is data collection: it can come in different forms and depends on each processor which it supports. In general, they all strongly support logging in different ways such as file tailing, stdout, HTTP, Kubernetes and others. In general, the log aggregators have first-party or plugin support for different types of inputs, in order to connect to the source and parse the logs seamlessly in a standardized output. The second step is first processing the logs. The system can filter logs and discard them based on specific rules. It can also tag the logs in order to redirect them to a specific processing pipeline, while also modifying the fields, transforming the data or even enriching it with further data. It can also aggregate multiple sources into one log pipeline, while also buffering the output in time, size or quantity to sinks. The last step of the logging pipeline is the output or sinking process. Its point is to send the already filtered and processed logs to another service, or to save them. It is up to the output process to decide what to do with the received logs. Some output services include ElasticSearch, Prometheus, message queues, topics, to a compressed file, or even any network exposed service. In general, logging is the most supported element of telemetry inside the aggregators. Some of them support metrics in different stages of development, such as first party support, through plugins, or manually. Traces are in general not supported except if you configure them manually. d. The golden standard The historical golden standard of logging pipeline is known as the ELK stack. ElasticSearch, LogStash and Kibana. LogStash is one of the first tools to collect, process and save logs in a unified way. LogStash generally collects and processes the logs, and then sends them to ElasticSearch. The latter acts as a database and search engine, where Kibana acts as a user-friendly way of analyzing, observing and filtering them. We at CERN already use Kibana and OpenSearch (ElasticSearch alternative), provided as a service by the OpenSearch team. Despite this, we do not use LogStash, since the Kubernetes space started using new tools such as Fluentd instead, because of its modular design and it being community driven through the CNCF. 6
CERN openlab Report // 2024 e. How we currently use Fluentd Figure 2: architecture with Fluentd aggregator We currently use a Kubernetes Operator called OpenShift Logging Operator. It is designed to deploy a full-stack configuration using Fluentd, OpenSearch and Grafana together. Since we at CERN already have a central service both for OpenSearch and Grafana, we only end up using the operator to deploy Fluentd. What the Logging Operator deploys using Fluentd is not enough for our requirements. It is capable of getting system logs from both nodes and container logs from the pods, but we also need some processing on top of this gathered data, in addition to data enrichment with events and proxy accesses. Given this consideration, we deployed another Fluentd system, in this case a StatefulSet to collect the logs from the various sources. Another problem faced in the past is that Fluentd is not able to collect Kubernetes events, therefore we were forced to deploy FluentBit too, which supports said collection. In Figure 2, you can observe the current architecture for the Fluentd collector. It gets logs both from the Fluentd daemonset forwarder and from HAProxy, transforms them and outputs them to the central monitoring system that uses OpenSearch. f. Why do we need an upgrade? Log collection, processing, storage and retrieval is becoming an important issue to address in large organizations with enormous web traffic, data processing and cloud computing requirements like CERN. It enables real-time monitoring and alerting, which are crucial for maintaining the health and correct operation of IT infrastructure. For organizations with complex architectures such as CERN, it also provides visibility into the operations, helping to identify and troubleshoot any occurring issues. Secondly, logs are essential for security and compliance purposes, since they record access to our systems, assisting in threat detection and usage analysis. Given the previous points, a fully working logging pipeline is vital to CERN. Therefore, we need to ensure the correct operation of it in our internal systems. 7
CERN openlab Report // 2024 Fluentd is a community driven open source project created by the company Treasure Data. The company does not actively contribute nor support the tool anymore, and is pushing users towards FluentBit. While the community still contributes, there is no active development anymore. Given that, the OpenShift Logging Operator that we use inside OKD will discontinue its support for Fluentd in favor of Vector. While our current implementation of Fluentd is working as expected and no logs are lost, it can still happen that it crashes due to its memory inefficiency. Thanks to the nature of Kubernetes and its self-healing mechanism, in contrast to VMs, even if containers crash they will still recover automatically. Therefore, it would be beneficial to migrate away from Fluentd to more modern, scalable, efficient and supported solutions. Finally, we will also explore further ways of improving the logging pipeline, specifically through intelligent anomaly detection. We will dive into the “Anomaly Detection” OpenSearch feature. It promises to generate alerts on unusual behavior changes in our logging data, providing insights into our data. It uses Random Cut Forests (RCF), an unsupervised machine learning algorithm that models the incoming data and assigns an anomaly grade and confidence score for each incoming input data point. 8
CERN openlab Report // 2024 2. Research and evaluation of Vector a. Decision: rework the logging system We decided to rebuild the logging system into a unified one, with a new tool. We will consider Vector and FluentBit as alternatives. The goal is to achieve a daemonset collector, and only one central aggregator and processor. As requirements, we need the new system to be functionally identical to the previous system. Also, if technically possible, it would be ideal to have compatibility not only with OKD, but also with Kubernetes Magnum. b. Starting point In order to act on our choice and make a more informed decision, we need to gain more experience and knowledge on these two new tools being Vector and FluentBit. We wanted to obtain real world experience with these tools before using them in a big project. Therefore, I contributed to the legacy-web-redirector project, implementing logging capabilities for it. Given that Vector is the new choice for the upstream OKD operator, we chose Vector as a starting point. c. Available benchmarks We analyzed the performance benchmarks published by Vector and realized that they were irregular, unkempt, deprecated and biased. In general, the benchmarking project is not maintained anymore, and was developed between 2019 and 2021. Also, the test suite uses old versions of both Vector and FluentBit, with a negative bias towards FluentBit by using an older version at the time period by almost 3 years. Further research and benchmarks need to be done with versions of FluentBit > 3.0 and Vector > 0.3. Therefore, the presented official benchmarks are not to be trusted. Here are the versions and dates used for the benchmarks: Fluentbit = 1.1.0 May 20, 2019, https://fluentbit.io/announcements/v1.1.0/ Vector = 0.20.0 February 8, 2022, https://vector.dev/releases/0.20.0/ Vector test harness, 2019 - 2021 https://github.com/vectordotdev/vector-test-harness/ d. Inner workings The first step of the project consisted of the research phase of current technologies that solve the organization problems and adjust to their needs. We first came into contact with Vector, from DataDog. It is described as a high-performance observability data pipeline, capable of collecting, transforming, and routing logs, metrics, and traces to any supported vendors. It works in a sources →transformers →sinks architecture, as all common logging technologies explained before. 9
CERN openlab Report // 2024 3. Research and evaluations - OpenSearch feature a. Research We decided to research an OpenSearch feature, specifically the “Anomaly Detection” for logs. In order to use anomaly detection, you need to ensure your analyzed data is in a numerical integer format, since it is a requirement for the internal detection algorithm. Also, each feature to analyze has a maximum of 5 inputs. Therefore, you are forced to do either: ● One hot encoding: ○ For example, for HTTP verbs, we can easily encode them in 5 categories. GET, DEL, POST, PATCH / PUT, OTHERS ○ For HTTP return codes, since we have around 10 for each starting with 2, 4 and 5, we need to bin them. Therefore, aggregate all return codes starting with 2 in a 2xx field, and the same for 4xx and 5xx, in order to not go over the 5 input maximum. ● Raw numerical data: ○ For data such as response time or frequencies you can use the numerical value directly. Also, we explain which aggregation techniques should be used for each use-case: ●average(): For detecting anomalies in the central tendency of your data. Website or Service response time in ms. Low response time could mean network error, and high response time could mean performance degradation or service saturation. ●count(): For detecting anomalies in the frequency of events. Tracking 2xx, 4xx, 5xx HTTP response codes. Sudden spike in 4xx and 5xx in addition to drop in 2xx could relate to a service problem. ●sum(), min(), max(): self explanatory 16
CERN openlab Report // 2024 Figure 5: Example anomaly detection for status codes In Figure 5, you can observe an anomaly detection inside the OpenSearch dashboard. It works based on probabilistic models, where if in the specified time window the probability distribution changes dramatically, an anomaly occurrence is cataloged. These anomalies can be connected to messaging systems such as Email and Slack as an alert. c. Results We also decided to not do further experiments with OpenSearch Anomaly Detection for the following reasons: i. Manual categorization Ideally, we would want to input the entire log string and have a ML algorithm analyze it and detect anomalies using a language model, instead of parsing the observable features by hand. Also, this gives low flexibility for different kinds of logs, and forces you to manually create a new pipeline for each traceback, such as Python, Java, Nginx, etc. ii. Numeric dependency As said before, the probabilistic features need to be in a numerical format, one-hot encoded which scales in size exponentially or binned which gives poor traceability. iii. Conclusion We conclude that OpenSearch anomaly detection current state is not where we need it to be. Data has to manually be encoded and binned for it to be useful, which adds unnecessary complexity to the architecture. Also, features and anomaly detection has to be manually set up, which makes it unmaintainable for the future. We decided that everything metrics related is best done by Prometheus, while keeping OpenSearch only as a log aggregator, short term storage and search engine for users. 17
CERN openlab Report // 2024 4. Proof of concept Overview The provided FluentBit configuration implements a comprehensive logging solution for OKD clusters. It's deployed as a DaemonSet, ensuring that a FluentBit instance runs on each node in the cluster. The configuration covers various aspects of log collection, processing, and forwarding, tailored specifically for OKD environments. Log Sources The configuration collects logs from multiple sources: 1. Container logs (/var/log/pods/*/*/*log) 2. Linux system journal (/var/log/journal) 3. Audit logs a. Linux: i. /var/log/audit/audit.log b. Kubernetes: i. /var/log/kube-apiserver/audit.log c. OpenShift audit logs: i. /var/log/oauth-apiserver/audit.log ii. /var/log/openshift-apiserver/audit.log iii. /var/log/oauth-server/audit.log d. Open Virtual Network (OVN) audit logs: i. /var/log/ovn/acl-audit-log.log Processing and Filtering The configuration applies several filters and modifications to the collected logs: 1. Record Modification: ○ Adds a UUID to each log entry ○ Renames fields for consistency (e.g., log to message) 2. Kubernetes Metadata: ○ Enriches container logs with Kubernetes metadata 3. Parsing: ○ Parses JSON messages from the event router ○ Parses OVN audit logs ○ Handles audit log timestamps 4. Tagging and Rewriting: ○ Retags journal logs based on their source (e.g., Kibana, event router, system containers) 5. Custom Processing: 18
CERN openlab Report // 2024 ○ Uses Lua scripts for advanced processing Output The configuration sets up two outputs: 1. Standard Output: For debugging purposes 2. Forward: Sends logs to a Fluentd aggregator service in the openshift-logging namespace Custom Lua processing ○ Fixing timestamp formats: needed when the log does not have a correctly set timestamp, which is required in Fluentd in a specific format function fix_timestamp(tag, timestamp, record) if record["@timestamp"]and record["@timestamp"] ~= "" then return 0, timestamp, record end -- get the decimal part of the timestamp nsec = timestamp % 1 n_sec_str = tostring(nsec) -- get the first 9 digits of the decimal part in a string nsec_last_digits = n_sec_str:match("%.(%d+)"):sub(1,9) -- in UTC (!), format the main timestamp formatted_time = os.date("!%Y-%m-%dT%H:%M:%S", timestamp) -- add the timezone offset, which is +00:00 for UTC tz_offset = "+00:00" -- concatenate the main timestamp, the decimal part and the timezone offset full_timestamp = string.format("%s.%s%s", formatted_time, nsec_last_digits, tz_offset) record["@timestamp"] = full_timestamp return 1, timestamp, record end ○ Adjusting systemd fields to fit Fluentd requirements function fix_systemd_fields(tag, timestamp, record) record["message"] = record["MESSAGE"] record["hostname"] = record["_HOSTNAME"] record["systemd"]={ ["u"] = { 19
CERN openlab Report // 2024 ["SYSLOG_FACILITY"] = record["SYSLOG_FACILITY"], ["SYSLOG_IDENTIFIER"] = record["SYSLOG_IDENTIFIER"] }, ["t"] = { ["SYSTEMD_UNIT"] = record["_SYSTEMD_UNIT"], ["PID"] = record["_PID"] } } return 1, timestamp, record end ○ Modifying Kubernetes fields to fit Fluentd requirements function fix_kubernetes_fields(tag, timestamp, record) record["hostname"] = record["kubernetes"]["host"] record["host"]=nil return 1, timestamp, record end Custom Parsers Several custom parsers are defined to handle specific log formats: ● OVN log parser: specifically used for OVN logs, where the log format is different and we need to extract the first part of it as the log timestamp: [PARSER] Name ovn_parser Format regex Regex ^(?<log_timestamp>[^|]+).*$ Time_Key log_timestamp Time_Format %Y-%m-%dT%H:%M:%S.%Lz ●JSON message parser: needed to parse only a specific key of the log as a JSON: [FILTER] Name parser Match kubernetes.var.log.pods.*eventrouter* Key_Name message Parser json_message_parser Preserve_Key true Reserve_Data true 20
CERN openlab Report // 2024 ... [PARSER] Name json_message_parser Format json ●Custom “receivedTimestamp”parser, needed to parse the data in JSON but also parsing a specific key with a specific time format as the log time: [PARSER] Name requestReceivedTimestamp Format json Time_Key requestReceivedTimestamp Time_Keep On Time_Format %Y-%m-%dT%H:%M:%S.%N%z ●Kubernetes pod log parser: extract from the currently parsed log file name part of the necessary Kubernetes metadata, such as namespace name, pod name, docker id and container name: [PARSER] Name kube-pod-log Format regex Reserve_Data true Regex (?<namespace_name>[a-z0-9](?:[-a-z0-9]*[a-z0-9])?(?:_[a-z0-9]([-a-z0-9]*[a-z0-9 ])?)*)_(?<pod_name>[^_]+)_(?<docker_id>.*)\.(?<container_name>.*)\..*\.log$ Key Features 1. Pods Log Handling: Supports logs from containers using CRI and Docker parsers 2. Efficient Log Tailing: Uses database files to track log file positions 3. Flexible Filtering: Implements various filtering mechanisms to process logs based on their source and content 4. Kubernetes Integration: Deep integration with Kubernetes, including metadata enrichment and namespace-aware processing 5. Audit Log Support: Comprehensive handling of various audit log sources in the OKD environment 6. Custom Processing: Uses Lua scripts for advanced log manipulation, allowing for complex transformations 21
CERN openlab Report // 2024 Results The proof of concept yielded several significant findings: 1. Pod Log Collection: functions as expected, but requires custom Regex parsers for accurate collection, in combination with Lua functions for modifications. 2. Systemd journal Log Collection: functions as expected, but requires custom Lua functions for collection and modifications. 3. Lua Functions: essential for complex parsing tasks, including: ○ Utilizing the internal FluentBit timestamp ○ Accessing or modifying deeply nested structures ○ Performing time parsing in various ways and operating on strings 4. Container ID Collection: feasible, but only when collecting pod logs from /var/log/containers, a location being deprecated by Kubernetes. 5. Audit Integration: does not function correctly due to insufficient plugin support in FluentBit. Future work While the proof of concept successfully replaces most of Fluentd's functionality with FluentBit, several areas require further development: 1. Multiline Log Handling: Comprehensive testing is needed to ensure proper handling of multiline logs. 2. Message Tracking: Implement a hashing mechanism for consistent message tracking, replacing the current UUID system. 3. Limitations in Metadata Collection: container_image_id and namespace_id cannot be collected in FluentBit due to the lack of custom Kubernetes API integration, which is available in Fluentd. 4. Audit Log Handling: Migrate this functionality to Fluentd due to lack of Openshift-specific plugins 5. Custom Plugin Development: Explore the possibility of developing custom plugins for FluentBit to address specific OKD logging requirements that are currently not met. 6. Integration Testing: Perform extensive integration testing with the existing logging ecosystem to ensure seamless compatibility and data consistency. Conclusion The FluentBit-based proof of concept demonstrates a viable replacement for the existing Fluentd system in OKD clusters. Key conclusions include: 1. Viability: The solution provides a robust and flexible logging framework capable of handling the complexities of OKD environments. 2. Performance and Maintainability: FluentBit offers improved performance and maintainability compared to the current Fluentd system. 22
CERN openlab Report // 2024 3. Functionality Preservation: Critical logging functionality is maintained, ensuring continuity in log collection and processing. 4. Adaptability: The use of Lua scripts showcases the solution's adaptability in addressing OKD-specific requirements and overcoming some of FluentBit's limitations. 5. Kubernetes Integration: The proof of concept demonstrates deep integration with Kubernetes, including metadata enrichment and namespace-aware processing. 6. Challenges: While largely successful, the implementation revealed some limitations, particularly in audit log handling and certain metadata collection capabilities. 7. Future-Ready: The proof of concept lays a strong foundation for future enhancements and optimizations in OKD logging infrastructure. Having said that, FluentBit configuration provides a robust and flexible logging solution for OKD clusters. It addresses the complexities of collecting logs from various sources in a Kubernetes environment, enriching them with relevant metadata, and preparing them for further processing and analysis. The use of Lua scripts for custom processing demonstrates the adaptability of the solution to specific requirements of the OKD environment. Therefore, our FluentBit-based proof of concept demonstrates a viable replacement for the existing Fluentd system, offering improved performance and maintainability while preserving critical logging functionality. The use of Lua scripts provided the flexibility needed to overcome FluentBit's limitations and maintain compatibility with our existing logging ecosystem. 23