Full text
Online and Adaptive PID Tuning for Thermal Zone Temperature Control with Bayesian Optimization Christos D. Tsaknakis1,3∗, Christos D. Korkas1,2, Dimitrios G. Vamvakas1,3, Yiannis Boutalis1,3and Elias B. Kosmatopoulos1,3 Abstract— This paper presents a Bayesian Optimization (BO)-based method for online tuning of a Proportional-IntegralDerivative (PID) controller in a real-world thermal zone HVAC application. Unlike conventional PID tuning methods that require manual intervention and offline calibration, the proposed approach optimizes PID gains autonomously and continuously, adapting daily to external disturbances (e.g., weather fluctuations, occupancy changes) without requiring a system model. A novel transition error mechanism accounts for occupancy patterns, enabling the controller to anticipate heating/cooling demands by preemptively adjusting set-points, thereby improving thermal comfort and energy efficiency. Through a 90-day real-world experiment under varying winter conditions (January–March), the BO-optimized PID controller is evaluated against a traditional on-off dead-band rule-based controller (RBC), demonstrating 5–7% energy savings despite lacking direct energy consumption feedback during tuning, a 45% reduction in temperature deviations from set-points, and faster convergence to target temperatures during occupancy transitions. These findings validate BO as a robust, model-free solution for real-time PID tuning in dynamic environments, balancing energy efficiency with occupant preferences. I. INTRODUCTION Modern building management systems aim to optimize energy efficiency, reduce operational costs, and enhance user comfort. In developed countries, residential buildings consume almost 40% of the total electricity production [1], with heating, ventilation, and air conditioning (HVAC) systems accounting for approximately 20–40% of this energy demand [2]. Given the substantial energy footprint of HVAC systems, improving their efficiency is essential not only for economic and environmental benefits but also for addressing energy poverty challenges while maintaining indoor thermal comfort. One of the most widely used control strategies in HVAC systems is the Proportional-IntegralDerivative (PID) controller, which has been the dominant choice in industrial automation for decades, with more than 90% of industrial controllers based on PID architecture [3]. In many applications, PID controllers regulate room temperature in single-zone control loops. However, their performance depends highly on proper parameter tuning. The *Corresponding author 1Christos D. Tsaknakis, Christos D. Korkas, Dimitrios G. Vamvakas and Elias B. Kosmatopoulos are with the Informatics & Telematics Institute (I.T.I.), Center for Research and Technology Hellas (CERTH), Greece. [email protected] 2Christos D. Korkas is also with the Electrical and Computer Engineering Dept. (ECE), University of Western Macedonia (UOWM), Greece. 3Christos D. Tsaknakis,Dimitrios G. Vamvakas, Yiannis Boutalis and Elias B. Kosmatopoulos are also with the Electrical and Computer Engineering Dept. (ECE), Democritus University of Thrace (DUTH), Greece. stability of PIDs is also important, and many techniques have been implemented, such as local model networks [4]. Traditional tuning methods which have been utilized in modern problems, such as Ziegler-Nichols [5] and CohenCoon [6], involve empirical rules derived from open-loop step response data, while methods for optimal tuning of PID for closed-loop systems have also been proposed [7]. Although these methods offer a starting point, they often require manual adjustments by control engineers or rely on preset configurations that may not adapt well to dynamic environmental conditions. As a result, PID controllers often operate suboptimally, leading to energy inefficiencies and increased operational costs. To enhance PID performance in many cases, various advanced tuning approaches have been explored, including fuzzy logic [8], [9], neural networks [10], adaptive control [11], decoupling control [12], and BO [13]. Among these, BO stands out as a data-driven, modelfree technique that balances exploration and exploitation. It enables automatic PID tuning without requiring an explicit mathematical model of the system dynamics. In this work, we propose an online BO framework for real-world HVAC control, in which PID parameters are continuously updated to adapt to daily variations in external weather conditions and occupancy patterns. Unlike static tuning approaches, our method dynamically optimizes PID gains in an automated, model-free manner, ensuring efficient temperature regulation while adhering to user comfort preferences. By integrating transition error mechanisms, our approach anticipates occupancy patterns, enabling preheating/precooling strategies that further improve energy efficiency and, most importantly, comfort and user preferences. A. Related Work BO has emerged as an active learning approach to efficiently learn and optimize cost functions directly from data [14]. In this work, we employ BO with Gaussian Process Regression (GPR), a widely used method due to its versatility. While GPR is inherently a stochastic process, one of its most well-known applications is as a regression model. Here, we leverage BO with GPR under safety constraints, ensuring that the algorithm respects predefined limits while minimizing the cost function [15]. The incorporation of safety constraints is particularly crucial in environments such as thermal zones and buildings, where human interaction is a key factor. BO has also been widely adopted as a data-efficient machine learning technique for tuning various controllers, par2025 33rd Mediterranean Conference on Control and Automation (MED) June 10 - 13, 2025. Tangier,Morocco 979-8-3315-7719-3/25/$31.00 ©2025 IEEE 411 2025 33rd Mediterranean Conference on Control and Automation (MED) | 979-8-3315-7719-3/25/$31.00 ©2025 IEEE | DOI: 10.1109/MED64031.2025.11073464 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:30 UTC from IEEE Xplore. Restrictions apply.
ticularly in black-box optimization problems. Recent studies have applied BO as an active learning approach for directly optimizing cost functions [14]. Several works [16], [17] have incorporated safety constraints to restrict exploration in applications ranging from synthetic data algorithms and safe movie recommendations to spinal cord therapy and robotics. Notably, Gelbart et al. [18] introduced an acquisition function modification to confine exploration within a desirable space, specifically in the meta-optimization of machine learning algorithms and sampling techniques. BO has demonstrated success across numerous domains, including control systems and engineering applications. It has been effectively applied to inverted pendulums [19], multi-armed bandit problems [20], robotics control [21], engine calibration [22], and heat pump applications with PI controller tuning [23]. For instance, Marco et al. [24] used BO to jointly tune process and measurement noise covariances and to learn the gains of a state feedback controller, which was ultimately refined using an LQR approach. Additionally, Abdelrahman et al. [25] showed the potential of BO to maximize energy generation in photovoltaic power plants. Further research highlights its utility in controller tuning. Mohammad Khosravi et al. [26] presented an automated, model-free, and data-driven method for PID cascade tuning, along with BO-based tuning of a heat pump controller [27]. The method has also been applied in tuning the parameters of a bipedal walker [28], optimizing a PD quadrotor controller, and automating control in a bulk tailings treatment plant [29]. Additionally, Qihang He et al. [30] explored Parallel BO for optimizing an industrial bio-oil processing unit, while Kim et al. [31] successfully applied BO to aircraft maneuver control, yielding promising results. Despite its extensive adoption across various domains, the integration of BO into complex building energy management systems and thermal zone temperature control remains relatively underexplored. The literature review we conducted revealed that, while studies such as [32] have made significant contributions,their approach did not fully capture the complexities of modern buildings or account for scheduling preferences. Given the potential benefits of BO in optimizing energy systems, further exploration in this domain is both necessary and promising. B. Main Contributions This work presents a novel online application of BO for optimizing the operation of complex building systems, continuously adapting from January to March to ensure efficient and intelligent thermal management. Unlike conventional approaches, this method operates with minimal information requirements, relying solely on thermostat data without access to direct energy consumption measurements. Using the probabilistic nature of BO, the system dynamically fine-tunes PID controller parameters in response to changing weather conditions, fluctuating occupancy patterns, and thermostat readings. A key innovation is the transition error mechanism, which penalizes deviations from expected occupancy schedules, enabling the system to efficiently track usage patterns and proactively implement preheating/precooling strategies. This ensures that indoor temperatures are precisely adjusted in anticipation of occupancy changes, minimizing energy waste while maintaining thermal comfort. The online and adaptive nature of this approach allows continuous optimization in real-time, providing sustained efficiency throughout the heating season. Furthermore, the on/off thermostat control strategy ensures seamless integration into existing building management systems, making the method not only computationally efficient but also highly practical for realworld deployment, delivering improved comfort and energy savings without requiring extensive sensor infrastructure or direct energy data. II. SYSTEM MODEL AND PROBLEM FORMULATION In this section, we will briefly describe the system model we used to test and apply our method. In order to simulate the thermal behavior and the dynamics of a building we used a resistor-capacitor (RC) model firstly introduced in [33]. RC models are used widely for building simulations because of the analogy between electricity and the thermal physics [34]–[36]. The 5R1C model we used is shown in Fig. 1 . It is presumed that only one glass surface is in contact with the outside environment. The zone’s other surfaces are in contact with the building’s thermal zones and are modeled as adiabatic. The system dynamics are described by the following circuit equations: Cm dTm dt +Tm(Htr3+Hem) = ϕmtot (1) where Tmis the thermal mass’ temperature, Htr3and Hem are the thermal conductances and ϕmtot is an equivalent thermal heat flux. In order to solve the differential equation numerically, it is mandatory to be discretised, so we apply the Crank-Nicolson method: Tmk+1 =ϕmtot +Tmk(Cm ∆t−0.5(Htr3+Hem)) Cm ∆t+ 0.5(Htr3+Hem)(2) where krepresents the timestep of length ∆t. It is worth to mention that the heating or cooling demand for the system is calculated in order to ensure that the desired-target temperature is reached. More information on the system’s dynamics can be found on GitHub. For the simulation, an energy plus weather file for the Zurich-Kloten, Switzerland region in 2013 was used. In this work, our goal is to regulate the temperature of a room optimally. The temperature needs to be driven at the proper set-point as soon as the user desires it, while also be maintained there as stable as possible. This criterion is expressed as follows: J= t=N X t=0 Wt(3) Wtexpresses the temparature deviation between the target set-point defined by the user and the room’s temperature for each timestep t. This is defined as the cost and needs 412 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:30 UTC from IEEE Xplore. Restrictions apply.
Fig. 1: RC model . to be minimized. The occupants’ scheduling preferences are displayed in Table I, for which we integrate a PID controller in order to achieve the target temperature for the proper time intervals. The standard algorithm exhibited in (4) minimizes the temperature error (e), between target temperature and measured indoor temperature, with the aid of the proportional (Kp), integral (Ki) and (Kd) parameters. PIDerror =Kpe+KiZe dt +Kd de dt (4) This classic PID formulation served as a baseline for comparison with our approach. Our approach defined a slightly alternated PID error where the error (e) is defined as: etotal(t) = ecurrent(t) + w∗etransition(t)(5) and the PID error as: PIDerror =Kpetotal +KiZetotal dt +Kd detotal dt (6) where ecurrent(t)is the standard temperature error, etransition(t)poses as the error between the target temperature set-point of the next time interval and the current target set-point and wis a weight factor. By defining the PID error in this way, we achieve not only temperature stability at the target set-point but also anticipation of the next interval’s set-point. By doing so, we ensure smooth transitions between temperature changes, minimizing abrupt fluctuations and enchancing thermal comfort and energy efficiency. In order to find the optimal look-ahead time horizon, we introduce a parameter Twhich defines the number of future steps considered. This parameter also needs to be optimized, therefore is a parameter of the BO as well. For the PID controller, we also added basic anti-windup capability usually found in practice where the integral part Kiis limited in its value and resets to zero when set-point changes affect the system, in order to avoid overshoots of the temperature [37], [38]. Hour-Intervals Target Temperature 00:00 - 06:00 21°C 06:00 - 08:30 22°C 08:30 - 17:30 19°C 17:30 - 20:00 21°C 20:00 - 21:30 22°C 21:30 - 00:00 21°C TABLE I: Target temperature settings for different hour intervals III. PROPOSED METHODOLOGY Our methodology, as mentioned before, controls the on/off function of the thermostat and is aiming at limiting the temperature fluctuations around the temperature set-point. To solve this optimization problem, we use the Gaussian Process Bayesian Optimization algorithm. Our algorithm uses Gaussian processes to make predictions about fbased on noisy evaluations, and uses their predictive uncertainty to guide exploration. The goal is to approximate a non linear map f(a) : A−→ R. A key assumption in Gaussian Process modeling is that the function values at different points in Aare jointly Gaussian-distributed random variables. A Gaussian Process ins fully characterized by its prior mean function m(a)and a covariance function k(a, a′), often referred to as kernel. The choice of the kernel has a crucial role in the correlation structure between different function values, capturing the similarity between inputs aand a′. In the context of learning mapping from the controller parameters and environmental conditions (such diverse weather conditions) to cost function values, Gaussian Process provide a probabilistic approach.Suppose we have tobservations of the cost function at different controller parameters and conditions( in our case Kp, Ki, Kd, T and the look ahead period), represented as ˜a= (a, z), with corresponding function evaluations yt= [ ˆ Ji(˜ ai, ..., ˆ Ji( ˜at)]. All these observations are affected by Gaussian noise such as ˆ Ji( ˜an) = Ji( ˜an)+ωnwhere ωn∼N(0, σ2).Given these observations , the Gaussian Process predicts the function Ji(.)at new input (a, z)using Gaussian distribution with variance: σ2 t(˜a;Ji) = k(˜a, ˜a)−kt(˜a)(Kt+Itσ2 ω)−1kT t(˜a)(7) and mean: µt(˜a;Ji) = kt(˜a)(Kt+Itσ2 ω)−1yt(8) µt(˜a;Ji) = kt(˜a)(Kt+Itσ2 ω)−1ytwhere It∈Rtxt is the identity matrix, kt(a;z)=[k(˜a, ˜a1), ..., k(˜a, ˜at)∈Rtxt is the kernel matrix with entries [Kt]m,n =k( ˜am,˜an), m, n ∈ 1, ..., t. It is essential to note that the cost function in (3) follows a Gaussian Process [39]. For the optimization variables in this work, we use the Matern 5/2 kernel, whereas the squared exponential kernel is used for context. This kernel encodes twice differentiable functions. while the squared exponential kernel models smooth, infinitely differentiable functions. 413 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:30 UTC from IEEE Xplore. Restrictions apply.
To ensure safety, we maintain an increasing sequence of subsets [40]. St⊆ bounded Dis established as safe by using the Gaussian Process posterior. In that way, the algorithm always chooses a sample inside the Stand his goals are, the expansion of the safe region and the localization of the high-reward regions inside St. It maintains a set Mt⊆DStof decisions that are potential maximizers of the cost function, which in our work represents the total error between the target temperature at each time-step and the air temperature in the thermal zone. So in each step it picks a decision x, that is the one with the largest predictive variance among Gt⊆DMt. IV. SIMULATION RESULTS As stated, we use the 5R1C simulation model to evaluate the performance both of the proposed algorithm and the baselines method as well. We test these methods on one thermal zone for a heating case through winter. For each set of optimization variables Kp, Ki, Kd, T , the controller’s performance is evaluated based on the current day’s performace. Then the training process continues in subsequent days. This approach enables online experiment, allowing the controller to be trained in an adaptive way through diverse winter conditions. For our simulations, historical weather data from Zurich-Kloten for the year of 2013 were acquired from an EnergyPlus file. The proposed online PID controller is trained over a simulation period of 3 months(90 days), starting from January to March. This period simulates a wide range of winter conditions, including varying temperatures, enabling the PID to learn and adapt effectively to diverse operational scenarios. To evaluate the adaptability and robustness of the method, the controller is tested on a different day from the training set. We chose 1st of December, as it still is a winter day with low temperatures, however unseen conditions from the PID, so is suitable for testing. A. Baseline Controllers 1) Traditional PID Controller: Our main comparison study will be conducted by utilizing the traditional PID controller for the regulation of the indoor temperature. This controller’s variables are optimized through BO with Safe Constraints as well, however it is using the Equation (4) and not the Equation (6), which results in a sub-optimal performance, trying to reduce the frequency oscillations of temperature around the target set-point for each time interval. 2) Rule Based Controller: The Rule Based Controller is implementing a human-based but very common control strategy. This strategy is responsible for the thermostat control. This strategy, provides user defined but far for optimal setpoints and offers acceptable performance. So RBC follows a simple strategy, where the user sets the thermostat to the desired set-point according to his schedule, so the thermostat is either on operation trying to achieve the target temperature or off operation as long as the target set-point is achieved. Thermostat checks the indoor air temperature every 5 minutes and we are adopting a ±0.5oChysterisis in order to simulate a real-life thermostat, and also to prevent noise provoked by nervous switching of the heater near set point value. B. Cost Comparison with Baseline Methods To evaluate the performance of the proposed method, we conduct a comparative analysis against baseline methods. This comparison assesses performance on a day excluded from the training phase, following the schedule specified by the user. We use two key indicators: total energy consumption and the deviation error between the target temperature set-point and the actual indoor air temperature throughout the day. In addition to the baseline controllers, we include several PID-based controllers for a comprehensive evaluation. These include: A PID controller with the worst parameters observed during the online training period, to highlight the robustness of online BO. A PID controller with the best optimized parameters derived from the same training process. A basic PID controller using a simplified online training approach based on Equation (3). An RBC controller for further comparison. This analysis reveals the effectiveness of the proposed method in achieving optimal performance under varying conditions, without proposing a worst than RBC, PID controller, during the training period. As illustrated in Table II, the Optimized PID, the worsttuned PID and the traditional PID controllers all significantly outperform RBC in terms of temperature regulation, as they exhibit reduced temperature deviations from the target setpoint compared to the RBC. The optimized PID achieves 45% improvement, the worst-tuned 32 % and the traditional 38 % improvement. Furthermore, even without explicitly utilizing energy consumption into the cost function during the training of PIDs, the controllers still achieve 4-7 % less energy consumption relative to the RBC. As illustrated in Figure 2a the optimized PID, by adopting preheating/precooling strategies (oval shaped note), maintains temperature closer to the target set-point compared to the traditional PID. As shown in Figure 2b the worst PID outperforms the RBC also in temperature deviations, validating the robustness and the adaptability of the method. V. CONCLUSION AND FUTURE WORK In this work, we tuned a PID controller for the on/off operation of a smart thermostat using BO with safe constraints, ensuring both efficiency and user comfort. Through an online experiment under varying weather conditions, we showed that the controller could be dynamically adjusted while respecting user scheduling preferences by incorporating a transition error strategy. This feature helps the system follow occupancy patterns more effectively, allowing it to preheat or pre-cool spaces ahead of time, ensuring the desired temperature is reached when needed while also reducing unnecessary energy use. We compared our approach to both a simple Rule-Based Control (RBC) strategy and a traditional PID controller. The results show that our method achieves better temperature regulation while using less energy, making 414 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:30 UTC from IEEE Xplore. Restrictions apply.
(a) (b) Fig. 2: Temperature Deviation Comparison. (a) Optimized and Traditional PID; (b) Worst tuned PID and RBC . TABLE II: Performance Comparison Between Proposed Method and Baselines Worst PID Optimized PID Traditional PID RBC Temperature Deviation Improvement With Respect to RBC 32% 45% 38% - Energy Consumption [kWh] 8.05 7.70 7.83 8.31 it a strong alternative to conventional thermostat control strategies. In future work, we plan to integrate this PID controller into a Reinforcement Learning (RL) framework for setpoint regulation to further improve energy efficiency. This hybrid approach will allow the system to dynamically adjust temperature targets based on occupancy, weather forecasts, and energy pricing, leading to even smarter energy use. We also aim to explore its application in more complex building systems, incorporating factors like photovoltaic (PV) generation, internal thermal gains from occupants and equipment, and unpredictable disturbances such as open windows and doors. Additionally, we will analyze how different building materials and thermal capacitance affect performance. We also aim to extend testing to year-round simulations and evaluate the method across heating and cooling seasons to assess its effectiveness in diverse real-world conditions. Last but not least, we aim at encapsulating energy consumption as an input parameter into the training process of the PIDs by incorporating it into the cost function and optimizing both energy efficiency and thermal comfort. ACKNOWLEDGMENT We acknowledge partial support of this work by the European Commission Horizon Europe - REHOUSE : Renovation packagEs for HOlistic improvement of EU’s bUildingS Efficiency, maximizing RES generation and cost-effectiveness (Grant agreement ID: 101079951) and Horizon Europe - 415 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:30 UTC from IEEE Xplore. Restrictions apply.
Harmonise : Hierarchical and Agile Resource Management Optimization for Networks in Smart Energy Communities ( Grant agreement ID: 101138595) REFERENCES [1] X. Cao, X. Dai, and J. Liu, “Building energy-consumption status worldwide and the state-of-the-art technologies for zero-energy buildings during the past decade,” Energy and buildings, vol. 128, pp. 198– 213, 2016. [2] L. P´ erez-Lombard, J. Ortiz, and C. Pout, “A review on buildings energy consumption information,” Energy and buildings, vol. 40, no. 3, pp. 394–398, 2008. [3] K. Nouman, Z. Asim, and K. Qasim, “Comprehensive study on performance of pid controller and its applications,” in 2018 2nd IEEE Advanced Information Management, Communicates, Electronic and Automation Control Conference (IMCEC). IEEE, 2018, pp. 1574– 1579. [4] C. H. Mayr, C. Hametner, M. Kozek, and S. Jakubek, “Nonlinear stable pid controller design using local model networks,” in 2012 20th Mediterranean Conference on Control & Automation (MED). IEEE, 2012, pp. 842–847. [5] V. V. Patel, “Ziegler-nichols tuning method: Understanding the pid controller,” Resonance, vol. 25, no. 10, pp. 1385–1397, 2020. [6] E. Joseph and O. Olaiya, “Cohen-coon pid tuning method; a better option to ziegler nichols-pid tuning method,” ENginerring Research, vol. 2, no. 11, pp. 141–145, 2017. [7] K. G. Papadopoulos, E. N. Papastefanaki, and N. I. Margaris, “Optimal tuning of pid controllers for type-iii control loops,” in 2011 19th Mediterranean Conference on Control Automation (MED), 2011, pp. 1295–1300. [8] M. Elnour and W. I. M. Taha, “Pid and fuzzy logic in temperature control system,” in 2013 International Conference on Computing, Electrical and Electronic Engineering (ICCEEE). IEEE, 2013, pp. 172–177. [9] D. Babunski, J. Berisha, E. Zaev, and X. Bajrami, “Application of fuzzy logic and pid controller for mobile robot navigation,” in 2020 9th Mediterranean Conference on Embedded Computing (MECO), 2020, pp. 1–4. [10] G.-Q. Zeng, X.-Q. Xie, M.-R. Chen, and J. Weng, “Adaptive population extremal optimization-based pid neural network for multivariable nonlinear control systems,” Swarm and evolutionary computation, vol. 44, pp. 320–334, 2019. [11] X. Zuo, J.-w. Liu, X. Wang, and H.-q. Liang, “Adaptive pid and model reference adaptive control switch controller for nonlinear hydraulic actuator,” Mathematical Problems in Engineering, vol. 2017, no. 1, p. 6970146, 2017. [12] J. Garrido, F. V´ azquez, and F. Morilla, “Multivariable pid control by decoupling,” International Journal of Systems Science, vol. 47, no. 5, pp. 1054–1072, 2016. [13] B. Boulkroune, X. Jordens, B. Mrak, J. Verhelst, B. Depraetere, J. Meskens, and P. Bovijn, “Enhancing pi tuning in plant commissioning through bayesian optimization,” in 2024 European Control Conference (ECC). IEEE, 2024, pp. 3220–3225. [14] J. Snoek, H. Larochelle, and R. P. Adams, “Practical bayesian optimization of machine learning algorithms,” Advances in neural information processing systems, vol. 25, 2012. [15] R. Marin-Perez, I. T. Michailidis, D. Garcia-Carrillo, C. D. Korkas, E. B. Kosmatopoulos, and A. Skarmeta, “Plug-n-harvest architecture for secure and intelligent management of near-zero energy buildings,” Sensors, vol. 19, no. 4, p. 843, 2019. [16] Y. Sui, A. Gotovos, J. Burdick, and A. Krause, “Safe exploration for optimization with gaussian processes,” in International conference on machine learning. PMLR, 2015, pp. 997–1005. [17] I. T. Michailidis, A. C. Kapoutsis, C. D. Korkas, P. T. Michailidis, K. A. Alexandridou, C. Ravanis, and E. B. Kosmatopoulos, “Embedding autonomy in large-scale iot ecosystems using cao and l4g-cao,” Discover Internet of Things, vol. 1, pp. 1–22, 2021. [18] M. A. Gelbart, J. Snoek, and R. P. Adams, “Bayesian optimization with unknown constraints,” arXiv preprint arXiv:1403.5607, 2014. [19] J. Park, C. Lee, D. Kim, and S. Han, “Efficient lqr parameter tuning for a flying inverted pendulum via bayesian optimization,” in 2024 24th International Conference on Control, Automation and Systems (ICCAS). IEEE, 2024, pp. 691–696. [20] A. Nandy, C. Kumar, D. Mewada, and S. Sharma, “Bayesian optimization–multi-armed bandit problem,” arXiv preprint arXiv:2012.07885, 2020. [21] R. Martinez-Cantin, “Bayesian optimization with adaptive kernels for robot control,” in 2017 IEEE international conference on robotics and automation (ICRA). IEEE, 2017, pp. 3350–3356. [22] A. Pal, L. Zhu, Y. Wang, and G. G. Zhu, “Multi-objective stochastic bayesian optimization for iterative engine calibration,” in 2020 American Control Conference (ACC). IEEE, 2020, pp. 4893–4898. [23] M. Khosravi, A. Eichler, N. Schmid, R. S. Smith, and P. Heer, “Controller tuning by bayesian optimization an application to a heat pump,” in 2019 18th European Control Conference (ECC), 2019, pp. 1467–1472. [24] A. Marco, F. Berkenkamp, P. Hennig, A. P. Schoellig, A. Krause, S. Schaal, and S. Trimpe, “Virtual vs. real: Trading off simulations and physical experiments in reinforcement learning with bayesian optimization,” in 2017 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2017, pp. 1557–1563. [25] H. Abdelrahman, F. Berkenkamp, J. Poland, and A. Krause, “Bayesian optimization for maximum power point tracking in photovoltaic power plants,” in 2016 European Control Conference (ECC). IEEE, 2016, pp. 2078–2083. [26] M. Khosravi, V. N. Behrunani, P. Myszkorowski, R. S. Smith, A. Rupenyan, and J. Lygeros, “Performance-driven cascade controller tuning with bayesian optimization,” IEEE Transactions on Industrial Electronics, vol. 69, no. 1, pp. 1032–1042, 2021. [27] M. Khosravi, A. Eichler, N. Schmid, R. S. Smith, and P. Heer, “Controller tuning by bayesian optimization an application to a heat pump,” in 2019 18th European Control Conference (ECC). IEEE, 2019, pp. 1467–1472. [28] R. Calandra, A. Seyfarth, J. Peters, and M. P. Deisenroth, “An experimental comparison of bayesian optimization for bipedal locomotion,” in 2014 IEEE international conference on robotics and automation (ICRA). IEEE, 2014, pp. 1951–1958. [29] J. Van Niekerk, J. D. Le Roux, and I. K. Craig, “On-line automatic controller tuning using bayesian optimisation-a bulk tailings treatment plant case study,” IFAC-PapersOnLine, vol. 55, no. 21, pp. 126–131, 2022. [30] Q. He, Q. Liu, Y. Liang, W. Lyu, D. Huang, and C. Shang, “Riskaverse pid tuning based on scenario programming and parallel bayesian optimization,” Industrial & Engineering Chemistry Research, 2024. [31] D. Kim, H.-S. Oh, and I.-C. Moon, “Black-box modeling for aircraft maneuver control with bayesian optimization,” International Journal of Control, Automation and Systems, vol. 17, pp. 1558–1568, 2019. [32] M. Fiducioso, S. Curi, B. Schumacher, M. Gwerder, and A. Krause, “Safe contextual bayesian optimization for sustainable room temperature pid control tuning,” arXiv preprint arXiv:1906.12086, 2019. [33] P. Jayathissa, M. Luzzatto, J. Schmidli, J. Hofer, Z. Nagy, and A. Schlueter, “Optimising building net energy demand with dynamic bipv shading,” Applied Energy, vol. 202, pp. 726–735, 2017. [34] P. Bacher and H. Madsen, “Identifying suitable models for the heat dynamics of buildings,” Energy and buildings, vol. 43, no. 7, pp. 1511– 1522, 2011. [35] R. Sonderegger, “Diagnostic tests determining the thermal response of a house,” 1977. [36] H. Madsen and J. Holst, “Estimation of continuous-time models for the heat dynamics of a building,” Energy and buildings, vol. 22, no. 1, pp. 67–79, 1995. [37] R. McDowall, Fundamentals of HVAC Control Systems: SI Edition Hardbound Book. Elsevier, 2009. [38] A. Visioli, “Modified anti-windup scheme for pid controllers,” IEE Proceedings-Control Theory and Applications, vol. 150, no. 1, pp. 49–54, 2003. [39] C. E. Rasmussen, “Gaussian processes in machine learning,” in Summer school on machine learning. Springer, 2003, pp. 63–71. [40] Y. Sui, V. Zhuang, J. Burdick, and Y. Yue, “Stagewise safe bayesian optimization with gaussian processes,” in International conference on machine learning. PMLR, 2018, pp. 4781–4789. 416 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:30 UTC from IEEE Xplore. Restrictions apply.