scieee AI-readable full text Open interactive document viewer

A hybrid algorithm for iterative adaptation of feedforward controllers: An application on electromechanical switches

Serrano-Seco, Eloy; Moya-Lasheras, Eduardo; Ramirez-Laboreo, Edgar

Abstract

Electromechanical switching devices such as relays, solenoid valves, and contactors offer several technical and economic advantages that make them widely used in industry. However, uncontrolled operations result in undesirable impact-related phenomena at the end of the stroke. As a solution, different soft-landing controls have been proposed. Among them, feedforward control with iterative techniques that adapt its parameters is a solution when real-time feedback is not available. However, these techniques typically require a large number of operations to converge or are computationally intensive, which limits a real implementation. In this paper, we present a new algorithm for the iterative adaptation that is able to eventually adapt the search coordinate system and to reduce the search dimensional size in order to accelerate convergence. Moreover, it automatically toggles between a derivative-free and a gradient-based method to balance exploration and exploitation. To demonstrate the high potential of the proposal, each novel part of the algorithm is compared with a state-of-the-art approach via simulation.

Full text

Contents lists available at ScienceDirect European Journal of Control journal homepage: www.sciencedirect.com/journal/european-journal-of-control A hybrid algorithm for iterative adaptation of feedforward controllers: An application on electromechanical switches Eloy Serrano-Seco ∗, Eduardo Moya-Lasheras , Edgar Ramirez-Laboreo Departamento de Informatica e Ingenieria de Sistemas (DIIS) and Instituto de Investigacion en Ingenieria de Aragon (I3A), Universidad de Zaragoza, Zaragoza, 50018, Spain A R T I C L E I N F O Recommended by T. Parisini Keywords: Mechatronics Adaptive control Iterative methods Run-to-run control Optimization algorithms Adaptive coordinate descent A B S T R A C T Electromechanical switching devices such as relays, solenoid valves, and contactors offer several technical and economic advantages that make them widely used in industry. However, uncontrolled operations result in undesirable impact-related phenomena at the end of the stroke. As a solution, different soft-landing controls have been proposed. Among them, feedforward control with iterative techniques that adapt its parameters is a solution when real-time feedback is not available. However, these techniques typically require a large number of operations to converge or are computationally intensive, which limits a real implementation. In this paper, we present a new algorithm for the iterative adaptation that is able to eventually adapt the search coordinate system and to reduce the search dimensional size in order to accelerate convergence. Moreover, it automatically toggles between a derivative-free and a gradient-based method to balance exploration and exploitation. To demonstrate the high potential of the proposal, each novel part of the algorithm is compared with a state-of-the-art approach via simulation. 1. Introduction Today, solenoid valves (Angadi & Jackson, 2022) and electromechanical relays (Gergič & Hercog, 2019) are used in virtually all industries, ranging from household appliances and automotive applications to robotics and medical devices. In general, the basic operating principle of all electromechanical switching devices is similar: when electrical energy is applied, a magnetic force accelerates a moving component to the end of the stroke. This causes undesirable phenomena, including bouncing and violent impacts, which result in premature device wear and acoustic noise. In an effort to solve or reduce these phenomena, several control strategies have been proposed, generally with the same objective: to reach the final position with zero velocity. Among these strategies are those based on backstepping control (Al Saaideh et al., 2022), sliding-mode control (Deschaux et al., 2018), extremum-seeking adaptive control (Benosman & Atınç, 2015), or iterative learning control (Moya-Lasheras & Sagues, 2024). Like other authors (Braun et al., 2018), some of our previous works employ a feedforward controller for two main reasons. First, a dynamic property of these systems, differential flatness, allows us to easily design the controller by model inversion. Secondly, feedforward control provides immediate responses to reference changes and is able to compensate for known disturbances. Despite its advantages, it alone is not robust to design errors, modeling errors, or system changes. ∗Corresponding author. E-mail address: [email protected] (E. Serrano-Seco). To address these limitations, various complementary strategies exist, including conventional feedback controllers with observers (Schroedter et al., 2018), learning algorithms (Grotjahn & Heimann, 2002), and parameter adjustments based on measurable variables (Yeh & Hsu, 1999). Nevertheless, all the previously mentioned controllers are dependent on feedback of the variable to be controlled. However, in some cases, these variables cannot be measured or observed due to economic or technical constraints. Therefore, we explored an alternative (Moya-Lasheras et al., 2023) based on run-to-run controllers. The main idea is to iteratively update the feedforward controller parameters from measurable variables that, even though they cannot be directly used to observe and control the variable of interest, they provide a performance index for each iteration (i.e., switching operation). In a first approach, the iterative adaptation law was implemented using a Pattern Search (Lewis & Torczon, 2000) algorithm. Although the method is computationally light and accurate, it requires too many evaluations to converge. Specifically, it needs 2𝑞+1 evaluations (where 𝑞 is the search space dimension) to determine whether to move to a new point. Our latest work Ramirez-Laboreo et al. (2024) demonstrated that convergence speed can be improved through sensitivity-based parameter reduction. However, it did not offer an automated approach for determining the number of parameters to be reduced for a given problem, among other limitations. https://doi.org/10.1016/j.ejcon.2025.101305 Received 6 June 2025; Accepted 9 July 2025 European Journal of Control xxx (xxxx) xxx Available online 23 July 2025 0947-3580/© 2025 The Authors. Published by Elsevier Ltd on behalf of European Control Association. This is an open access article under the CC BY-NC-ND license ( http://creativecommons.org/licenses/by-nc-nd/4.0/ ). Please cite this article as: Eloy Serrano-Seco et al., European Journal of Control, https://doi.org/10.1016/j.ejcon.2025.101305 E. Serrano-Seco et al. In this line, Loshchilov et al. (2011) presents an Adaptive Coordinate Descent algorithm. The strategy involves periodically updating the coordinate system by a Covariance Matrix Adaptation Evolution Strategy (CMA-ES) and Adaptive Encoding to decompose the problem into as many one-dimensional problems as dimensions in the general problem. Although it is an interesting idea, as concluded by the authors in Hansen and Ostermeier (2001), the function evaluations needed are about 10 𝑞, 30 𝑞 for a real-world search problem, and 100 𝑞2 for complete adaptation. Given that each evaluation involves a switching operation, the number of switching operations with unsatisfactory performance would be excessively high. In terms of one-dimensional search, the authors of Loshchilov et al. (2011) suggest derivative-free methods or the use of gradients. Gradient methods are a powerful tool, especially if the objective function is known. A similar alternative are subgradient methods, since they can work with approximations based on the value of the function to be optimized across the search space. Another option are the sign gradient descent methods, first introduced in the RProp (Resilient Propagation) algorithm (Moulay et al., 2019). Despite being technically a gradientbased method, the RProp algorithm has a low computational load, as it only needs to calculate the sign of the gradient, not the gradient itself. Nevertheless, adjusting the hyperparameters of these algorithms can be a challenging task. In contrast, some gradient descent methods implement an adaptive step size without the need for hyperparameters. This paper presents a new Run-to-Run controller based on an Adaptive Coordinates algorithm (R2R-AC) in order to automate the improvement process described in Ramirez-Laboreo et al. (2024) and to enhance the performance of the iterative adaptation law. This new algorithm leverages the controller sensitivity with respect to its parameters to calculate an alternative search basis that decomposes the initial 𝑞-dimensional problem into 𝑞 one-dimensional problems to optimize on the descending coordinate with the highest improvement potential. We analyze and compare three versions of the algorithm that differ in how the step size is computed: one based on derivative-free methods, another based on gradient-based algorithms, and a hybrid one that toggles between the other two to enhance the exploration–exploitation tradeoff. The paper is organized as follows. Section 2 provides a concise overview of the dynamic and control model that has prompted the development of the proposed algorithm. Section 3 develops the iterative adaptation law of the R2R-AC strategy and discusses the different versions previously mentioned. Section 4 contains simulation results that demonstrate the functionality of our proposals and the comparison with a state-of-the-art feedforward run-to-run controller. Finally, the conclusions are discussed in Section 5. 2. Background of the control system In this section, we briefly describe the system dynamics and control where the need for the proposed new algorithm has arisen. For a more detailed explanation, readers are referred to our works (Moya-Lasheras et al., 2023; Ramirez-Laboreo et al., 2024). 2.1. System dynamics The system is modeled as a single-coil reluctance actuator, subject to two types of forces: passive elastic forces—typically modeled as ideal springs—and a magnetic force. The magnetic force is generated when current flows through the coil, causing an inner fixed core to become magnetized and attract the movable core. The typical method of supplying the actuator with power is by providing a voltage. We describe the dynamics of the system using a state-space model, where the voltage 𝑢 is the input to our system, and the position 𝑧, velocity 𝑣 and magnetic flux linkage 𝜆 are the state variables. The state equations are defined as 𝑧 =𝑣, (1) 𝑣 =1 𝑚(−𝑘s(𝑧−𝑧s)−1 2𝜆2𝜕 𝜕𝑧 ),(2)  𝜆= −𝑅 𝜆 (𝑧, 𝜆) + 𝑢, (3) where 𝑚, 𝑘s, 𝑧𝑠, 𝑅, and  are the moving mass, the spring stiffness, the spring resting position, the coil resistance, and an auxiliary function based on the magnetic reluctance concept, respectively. This auxiliary function considers the magnetic saturation and flux fringing phenomena, (𝑧, 𝜆) = 𝜅1 1 − |𝜆|∕𝜅2 +𝜅3+𝜅4𝑧 1 + 𝜅5𝑧log(𝜅6∕𝑧),(4) where 𝜅1, 𝜅2, 𝜅3, 𝜅4, 𝜅5, and 𝜅6 are positive constants. Overall, the system dynamics depends on 𝑞= 9 uncertain parameters, which can be grouped in the parameter vector 𝑝. 𝑝=[𝑘s𝑧s𝑚 𝜅1𝜅2𝜅3𝜅4𝜅5𝜅6]⊺.(5) Note that the resistance 𝑅 is treated independently as a parameter without uncertainty, as it can be easily measured. 2.2. Control The control structure used in this study is schematized in Fig. 1. This is an iterative control design for a real-world scenario with two particularities: •The variable to be controlled, the position 𝑧, cannot be fed back for several reasons. Firstly, a position sensor is more expensive than the switching devices. Secondly, a protective housing impedes access to the component whose position needs to be known. In addition, a real-time estimate is also unavailable. •Errors in the model parameters are not negligible. These devices are produced at a low cost, with relaxed manufacturing tolerances, causing variability in the value of the parameters. Due to the first point, the control strategy is focused on a feedforward controller. It is designed by model inversion, taking advantage of a structural property shared by such devices: differential flatness (Lévine, 2011). Considering that the objective is soft-landing control, as in our previous works, 𝑧(𝑡) is the desired trajectory, designed as a 5th-degree polynomial from 𝑧0, the initial mechanical limit of motion of the moving component to be controlled, to 𝑧f, the final limit. The trajectory is defined over the time interval [𝑡0, 𝑡f], subject to boundary conditions of zero initial and final velocities and accelerations. In short, the feedforward control term, 𝑢𝑓𝑓 , is defined as a function of 𝑧(𝑡), its derivatives, and the parameter vector 𝑝. Although this applies to a specific case, it can be generalized as 𝑢𝑓𝑓 =𝑢𝑓𝑓 (𝑡, 𝜃) for parametric controllers where 𝜃 represents any vector of control parameters. In our case 𝜃 is a normalized version of 𝑝. Due to the second point, including a feedback loop to adapt the control parameters 𝜃 is essential. Some previous works propose associating a measurement related to the control objective to a cost value 𝐽. In a real-world application, due to measurement difficulties, 𝐽 may be computed based on indirect measurements associated with the impacts (Moya-Lasheras et al., 2023; Winkel et al., 2023). In simulation, however, the impact velocity 𝑣c could be directly used as feedback, 𝐽=||𝑣c||.(6) At this point, it is reasonable to assume that the process which relates 𝜃 to 𝐽 is unknown or difficult to work with analytically. In response to this, the proposal is to use a black-box optimizer as the iterative adaptation law to minimize 𝐽 along switchings. Since we are focusing on a real-world application, the optimizer must be not only accurate, European Journal of Control xxx (xxxx) xxx 2 E. Serrano-Seco et al. Fig. 1. General control diagram. The subscript 𝑘 denotes the variables of the 𝑘th evaluation of the run-to-run adaptation law. The feedforward block computes 𝑢𝑓𝑓 from the parameter vector 𝜃 and the desired reference signal 𝑟. The adaptation law updates the feedforward parameters 𝜃 using the cost 𝐽, which is derived from the measurable output 𝑦. but also computationally light and able to quickly discard poorly performing evaluations to minimize unsatisfactory switching operations. Therefore, to speed up convergence, a state-of-the-art feedforward runto-run controller based on a Pattern Search algorithm (Ramirez-Laboreo et al., 2024) shows a strategy to reduce the number of search dimension, without relying on further cost function evaluations. Assuming that larger changes in the control action translate into larger changes in the cost value, this strategy is based on a local sensitivity analysis of smooth (differentiable) controllers. The Fisher matrix, (𝜃), which can be computed from the sensitivity of the controller with respect to 𝜃, (𝜃) = ∫𝑡f 𝑡0[(𝜕𝑢𝑓𝑓 (𝑡, 𝜃) 𝜕𝜃 )⊺(𝜕𝑢𝑓𝑓 (𝑡, 𝜃) 𝜕𝜃 )]d𝜏, (7) is evaluated at a nominal point 𝜃nom. Since (𝜃nom) is symmetric and positive semidefinite, the eigendecomposition and the singular value decomposition coincide. That is, the Fisher matrix can be expressed as (𝜃nom) = 𝑉 𝛬 𝑉 ⊺, where 𝑉∈R𝑞×𝑞 is the orthonormal matrix with the eigenvectors (or singular vectors) of (𝜃nom) as columns and 𝛬∈R𝑞×𝑞 is the diagonal matrix of the corresponding eigenvalues (or singular values). Finally, a transformation of the known vector 𝜃 to a new vector 𝑋∈R𝑞 is defined in terms of the basis change matrix 𝑉 as 𝜃=𝜃nom +𝑉(𝑋−𝑋nom)⟺𝑋=𝑋nom +𝑉⊺(𝜃−𝜃nom),(8) where 𝑋nom is the nominal value of 𝑋, which can be chosen arbitrarily. Thanks to this procedure, the controller is parametrized in such a way that the correlation between the sensitivities with respect to the new parameters in vector 𝑋 is low. In addition, 𝛬 provides information about these sensitivities, enabling the exclusion of coordinates with negligible sensitivity from the optimization process. 3. New algorithm This section is divided into two parts. The first one, following the idea presented in Loshchilov et al. (2011), introduces a new algorithm which solves the open questions posed in Ramirez-Laboreo et al. (2024), i.e., what is the appropriate number of dimensions to optimize in each situation and, since the analysis is local, how often the alternative coordinate system should be recalculated. The second part presents a classical derivative-free method, a gradient-based method, and a hybrid methodology that enables toggling between the two options to calculate the step size of the previous algorithm. 3.1. Iterative adaptation law of the R2R-AC The behavior of the algorithm is presented in Algorithm 1. The new algorithm must fulfill two requirements: it must eventually upgrade the alternative coordinate system and it must select automatically the number of search dimensions. We propose to convert the complete optimization of the 𝜃-problem into successive 𝑋-optimization problems. Each 𝑋-optimization problem optimizes 𝑋 in a new canonical basis initialized at 𝑋nom =𝟎𝑞. This decision simplifies the transformation of 𝑋 into 𝜃 as 𝜃=𝜃nom +𝑉 𝑋, (9) with the only eventual needs to update 𝜃nom as the best evaluated 𝜃 (Algorithm 1, line 18) and update the transformation matrix 𝑉 (line 2) as shown in the previous section (according to Ramirez-Laboreo et al. (2024)). From now on, for clarity, since we work mainly in the 𝑋-optimization space, with a slight abuse of notation, we denote the cost — obtained for a value 𝜃 that depends on 𝑋, (9) — as a direct function of 𝑋: 𝐽(𝑋) = 𝐽. Once the strategy for updating the alternative coordinate system has been selected, the remaining tasks are to determine the timing of the update event and the search problem dimensions. The proposed solution is to first perform an exploration process to find the coordinate with the greatest descent and then to exploit this coordinate. For the exploration process, 𝑉 should be ordered such that the associated eigenvalues are arranged from largest to smallest, i.e., 𝛬(1,1) ≥𝛬(2,2) ≥⋯≥𝛬(𝑞,𝑞),(10) where the subscripts in parentheses represent the position in the matrix. This step permits the ordering of the coordinates from the highest to the lowest sensitivity of the control action to the parameters 𝑋. The proposal for selecting the coordinate of greatest descent is based on pattern search methods: the origin of the coordinate system (line 13) and two points along each coordinate (line 7), 𝑋+ and 𝑋−, symmetrically located at a distance 𝛿 from the origin coordinate, are evaluated, 𝑋+←𝑋b+𝛿⋅𝑒𝑑,(11) 𝑋−←𝑋b−𝛿⋅𝑒𝑑,(12) where 𝑋b is the point associated with the lowest cost value, whose value is 𝟎𝑞 at the start of each 𝑋-optimization problem, 𝛿∈R is the step size, and 𝑒𝑑∈R𝑞 is the unit vector that defines the direction of the 𝑑∈ [1, 𝑞] coordinate of the canonical basis. This pattern is evaluated sequentially coordinate by coordinate (lines 4–10) until a cost-improving coordinate has been found. At this moment the exploration process ends and the exploitation process begins by a line-search (lines 11–16). The algorithm continues to look for lower cost points along the corresponding direction and orientation by a method that embraces the philosophy of sign gradient descent algorithms, 𝑋next ←𝑋b+sgn (𝑋b (𝑑))𝛿⋅𝑒𝑑,(13) where 𝑋next is the next point to be evaluated, sgn(⋅) is the sign operator, and 𝑋b (𝑑) is the 𝑑-th component of 𝑋b. Note that (13) only applies after a better point has been found, so in this case, 𝑋b (𝑑)≠0. When the evaluation of 𝑋next does not improve the cost, the present 𝑋-optimization problem ends, 𝜃nom and 𝑉 are updated and the 𝑋optimization problem is reset (line 18). Thus, it is not necessary to complete the pattern to shift it, and the number of dimensions of the problem is automatically reduced to the minimum allowed to obtain improvements. An exception applies to this sequence, when the algorithm has moved on the first coordinate (line 17), 𝜃 is not updated, so neither is 𝑉, and the evaluation of the second coordinate continues around the best point found on the first coordinate. The main reason is not to transform the algorithm into a single gradient search method in which the descending coordinate is calculated through the coordinate that further modifies 𝑢𝑓𝑓 , because the convergence speed may decrease due to limited information and reduced opportunities to directly identify a new best point. To complete the algorithm, it only remains to define how to upgrade the step size (lines 9 and 15). This is discussed in the following subsection. European Journal of Control xxx (xxxx) xxx 3 E. Serrano-Seco et al. Algorithm 1 Iterative adaptation law of the R2R-AC Initialize: 𝛿, 𝜃nom, 𝑑←0, 𝑋b←𝟎𝑞, 𝐽(𝑋next )←∞ 1: while true do ⊳ Alternative coordinate system 2: 𝑉← sorted eigenvectors of (𝜃nom) ⊳ Best descent coordinate exploration 3: Evaluate 𝐽(𝑋b) 4: while 𝐽(𝑋next )> 𝐽(𝑋b) do 5: 𝑑←(𝑑mod 𝑞) + 1 ⊳ Next 𝑑 6: 𝑋+←𝑋b+𝛿⋅𝑒𝑑; 𝑋−←𝑋b−𝛿⋅𝑒𝑑 7: Evaluate 𝐽(𝑋+) and 𝐽(𝑋−) 8: 𝑋next ←arg min𝑋∈{𝑋+, 𝑋−}𝐽(𝑋) 9: Update 𝛿 ⊳ Algorithm 2 10: end while ⊳ Descent coordinate exploitation: Line-search 11: while 𝐽(𝑋next )< 𝐽(𝑋b) do 12: 𝑋b←𝑋next 13: 𝑋next ←𝑋b+sgn (𝑋b (𝑑))𝛿⋅𝑒𝑑, 14: Evaluate 𝐽(𝑋next ) 15: Update 𝛿 ⊳ Algorithm 2 16: end while ⊳ Reset the 𝑋 optimization problem 17: if 𝑑≠ 1 then 18: 𝜃nom ←𝜃nom +𝑉 𝑋b; 𝑑←0; 𝑋b←𝟎𝑞 19: end if 20: end while 3.2. Step-size update: hybrid method The new algorithm enables the 𝑞-dimensional problem to be reduced to the behavior of a one-dimensional problem. For the sake of simplicity, this subsection assumes that a descending shift always occurs. For clarity, we work with 𝑒𝑑, the unit vector of the desired direction and orientation. The fact that there is no analytical information on the relationship between the cost value 𝐽 and the parameters to optimize 𝑋, together with the need of a computationally light and fast convergence process, has led us to the use of direct search methods based on updating a step size 𝛿. We explore two approaches: derivative-free methods and adaptive methods based on the objective function. In the basic derivative method, the step size decreases as the number of evaluations increases; however, other methods allow for expanding the step size if certain conditions are met. The latter methods are used in algorithms as Pattern Search and RProp, among others. When it appears the process is progressing in the right direction, i.e., the cost improves, the step size should be increased (expansion) to reach the possibly distant optimum point more quickly. Conversely, when the process has reached a minimum, i.e., the cost does not improve, the step size should be decreased (contraction) to allow for an approach to the minimum cost and reduce fluctuations. In short, the next point to evaluate 𝑋next can be calculated as 𝑋next =𝑋b+𝛿DF ⋅𝑒𝑑,(14) where 𝑋b is the point considered best, 𝛿DF is the derivative-free step size, which is eventually updated as 𝛿DF ←{max(𝛼con 𝛿DF, 𝛿DF min) if 𝐽(𝑋next )> 𝐽(𝑋b) min(𝛼exp 𝛿DF, 𝛿DF max) if 𝐽(𝑋next )≤𝐽(𝑋b) ,(15) where 𝛼con and 𝛼exp are the contraction and expansion constants such that 0< 𝛼con <1< 𝛼exp, and 𝛿DF min and 𝛿DF max are the minimum and maximum allowed step sizes, respectively. Unfortunately, the tuning of 𝛼 values suffers from a trade-off between faster convergence and the Fig. 2. Geometric interpretation of the first-order method adopted. likelihood of reaching a local minima, which tends to be higher as the dimensionality of the problem increases. In contrast, adaptive methods based on the objective function are able to dynamically fit the step size, often utilizing gradient information in first-order methods. A typical geometric interpretation of these methods is illustrated in Fig. 2 for an ideal situation with a convex objective function. The aim is to reach a point 𝑋∗ with a lower cost 𝐽∗ from a known point 𝑋b, for our algorithm, the best evaluated point. To this end, the step size is approximated as the ratio of the cost difference 𝐽(𝑋b) − 𝐽∗ to the gradient 𝑔 at 𝑋b. To determine 𝐽∗, common approaches include treating it as a constant equal to 𝐽opt , the cost of the optimal point 𝑋opt , i.e., the minimum value of the objective function. Even if the real value is unknown, 𝐽∗= 0 is adequate on many situations, but it is usually a strong assumption. Alternatively, 𝐽∗ can vary by iteration (Algorithm 2, line 5), which is often more practical, as large step sizes can be more detrimental than an advantage, leading to divergence. As shown in Fig. 2, larger step sizes introduce greater approximation errors, potentially affecting convergence. For our problem, our non-static search coordinate system complicates the collection of information on 𝑔, as a reminder, the gradient with respect to 𝑋. Thus, we propose a reinterpretation of this method by working with 𝑔, an average slope of the trajectory followed by the algorithm that represents the average improvement capacity. In addiction, given the possible non-smoothed decreasing cost, we work with  𝐽, an average value of the cost. These are computed as  𝐽←𝛽 𝐽+ (1 − 𝛽)𝐽(𝑋𝑘),(16) 𝑔 ←√𝛽 𝑔2+ (1 − 𝛽)(𝐽(𝑋𝑘) − 𝐽(𝑋𝑘−1) ‖𝑋𝑘−𝑋𝑘−1‖)2 ,(17) where 𝛽 < 1 is a positive constant that acts as a decay factor, and the subscript 𝑘 refers to the evaluation number. With these adjustments, the next point to be evaluated and the value of the gradient-based step size 𝛿GB are calculated, respectively, as 𝑋next =𝑋b+𝛿GB ⋅𝑒𝑑,(18) 𝛿GB =( 𝐽−𝐽∗) 𝑔 .(19) Analyzing the equation, an additional advantage of slope filtering is that it mitigates the problem of excessively oscillating step sizes. Algorithm 2 Process to update the step size 𝛿 Initialize:  𝐽, 𝑔, 𝛽, 𝛿DF, 𝛼con, 𝛼exp, 𝛿DF min, 𝛿DF max, 𝐽∗ 1: for 𝑘←1 to num. evaluations do 2: Update  𝐽 and 𝑔 ⊳ Eq. (16) and Eq. (17) 3: if 𝛿 must be updated then 4: Update 𝛿DF and 𝛿GB ⊳ Eq. (15) and Eq. (19) 5: Update 𝐽∗ 6: 𝛿←{𝛿DF if  𝐽≤𝐽∗ 𝛿GB if  𝐽 > 𝐽 ∗ 7: end if 8: end for European Journal of Control xxx (xxxx) xxx 4 E. Serrano-Seco et al. Table 1 Nominal parameter values. 𝑘s55 N∕m 𝜅51320 m−1 𝑧s0.015 m 𝜅69.73 ⋅10−3 m 𝑚1.6⋅10−3 kg 𝑅50 Ω 𝜅11.35 H−1 𝑧010−3 m 𝜅20.0229 Wb 𝑧f0 𝜅33.88 H−1 𝑡00 𝜅47.67 ⋅104H−1∕m 𝑡f3.5⋅10−3 s Table 2 Initial hyperparameters of the control strategies. Control 𝛿DF 𝛿DF min 𝛿DF max 𝛼con 𝛼exp 𝛽 R2R-PS+ 0.2 2⋅10−10 2 0.5 2 – R2R-AC 0.2 2⋅10−10 2 0.7 1.1 0.8 This concept is also used in several stochastic gradient descent algorithms, including RMSProp (Ma et al., 2022). Despite the adaptability of gradient-based methods, they are highly susceptible to the geometry of the objective function and the initial evaluations. Our hybrid version for calculating 𝛿 tries to take advantage of the good performances of both strategies and avoid their drawbacks by toggling between (15) and (19) based on whether  𝐽 is less than or greater than 𝐽∗, respectively (line 6). In this form the first resource of the algorithm is (15), but, when convergence slows significantly or the evaluated point is far from the optimal point, the strategy toggles to (19). The update of the step size 𝛿 is summarized in Algorithm 2. 4. Simulated results In this section, to illustrate the benefits of the new algorithm, we analyze the improvements achieved over our previous work (RamirezLaboreo et al., 2024). To assess the improvements introduced by R2R-AC and the hybrid step size, the following control strategies are evaluated: •R2R-PS+: the best solution in Ramirez-Laboreo et al. (2024); it uses for the iterative adaptation law a Pattern Search algorithm, with a single initial basis change (9) and four dimensions. •R2R-AC (DF): proposed R2R-AC in which 𝛿←𝛿DF. •R2R-AC (GB): proposed R2R-AC in which 𝛿←𝛿GB. •R2R-AC (Hy): proposed R2R-AC in which 𝛿 is computed using the Algorithm 2, the complete proposal. These strategies have been tested through simulation on the problem presented in Section 2. Recall that the detailed methodology underlying the design of the feedforward controller, as well as the formulation of the desired trajectory, can be found in Ramirez-Laboreo et al. (2024). It is assumed that the model acting in the role of the real system and the feedforward controller are based on exactly the same equations. However, it is reasonable to expect discrepancies between the model parameters identification and the nominal (initial) feedforward controller parameters (see Table 1). To emulate these errors, two Monte Carlo analyses of 10 000 trials, and 300 switching operations in each trial are performed for each run-to-run strategy. For each Monte Carlo analysis, a test set with a different 𝑝-vector, (5), for each trial is generated. The first test set is associated with a situation where the errors are small, and the second with a situation where the errors are larger. To ensure fair comparisons, these test sets are common for all algorithms. The initial hyperparameters required of each algorithm (see Table 2) have been adjusted to optimize the results of the test where the errors are small. The same hyperparameters have been applied to the other test. 4.1. Small errors in the initial parameters For this situation, each component of each 𝑝-vector of the real system model is randomly and independently perturbed up to 5%, i.e., the parameters of the real device under consideration vary with a uniform probability distribution between 95% and 105% of the values in Table 1. Fig. 3 shows the results of the first analysis. The graphs represent the evolution of the cost, 𝐽, with respect to each evaluation or switching operation. Due to the large number of simulations required to capture the variability of the parameters across devices, the results are presented by the median (𝑃50) and the 10th and 90th percentiles (𝑃10 and 𝑃90, respectively) of the distribution of values obtained for the 10 000 simulated experiments. For reference, the cost of a switching operation without control, namely with a 30 V constant activation, is also plotted. To demonstrate the improvement introduced by R2R-AC, Fig. 3(a) (the best solution in Ramirez-Laboreo et al. (2024)) and 3(b) must be compared, as they differ only in the search strategy; both methods compute the step size using the same derivative-free approach. As can be seen, while the results at the end are quite similar, the new R2R-AC (DF) strategy shows a notable improvement in the convergence speed. While our previous feedforward run-to-run controller requires 300 switching operations for 90% of the trials to converge, R2R-AC (DF) requires only approximately 70. The same applies to 50% and 10% of the trials, for which the required number of switching operations is reduced by approximately half. The other two strategies require the definition of 𝐽∗. For R2R-AC (GB), since gradient-based methods are sensitive to excessively large step sizes, 𝐽∗ is variable and is calculated as the minimum evaluated cost, 𝐽min, multiplied by a positive constant 𝛾 < 1. With this strategy, the convergence speed also increases, but the final results are worse in this case. For R2R-AC (Hy), the expression of 𝐽∗ (see Fig. 3(d)) is derived from a smoothed 𝑃90 behavior of R2R-AC (DF) under the constraint that the initial value matches the cost obtained from an evaluation without control. In this way, processes below the 90th percentile should remain unchanged, as can be seen when Figs. 3(b) and 3(d) are compared, while those above the 90th percentile should improve their performance when the hybrid method is applied to calculate the step size. For this reason, the 97.5th percentil (𝑃97.5) is also represented in these figures. Comparing this index, with R2R-AC (Hy) the target value is almost reached at the end of the trials, while with R2R-AC (DF) the convergence continues but progresses at a slow pace. Figs. 4(a) and 4(b) illustrate the effect of the hybrid strategy on the 𝐽 evolution of two individual processes, one with slow convergence and another that converges to an unacceptable cost, respectively. 4.2. Larger errors in the initial parameters Analogous to the previous subsection, the 𝑝 test set has been generated with perturbations up to 25% instead of 5%. Fig. 5 shows the results of the second analysis. As in the previous case, 𝑃10, 𝑃50 and 𝑃90 of the distribution of values obtained for the 10000 simulated experiments are shown. Each control strategy, i.e., R2R-PS+, R2R-AC (DF), R2R-AC (GB), and R2R-AC (Hy), uses the same hyperparameters as those employed in the previous subsection. This includes the definition of 𝐽∗ across evaluations: for R2R-AC (GB), 𝐽∗←𝐽min ⋅𝛾, and for R2R-AC (Hy), 𝐽∗ is the smoothed behavior of 𝑃90 with the set of nominal parameters perturbed up to 5%. As can be seen, the improvement of R2R-AC, compared to RamirezLaboreo et al. (2024) is considerable. In contrast to cases with initial low 𝑝 error, R2R-AC (GB) is able to offer better performance than R2R-AC (DF). However, our complete proposal, R2R-AC (Hy), achieves the best results regardless of the error level. As a summary, Table 3 compares 𝑃90 of 𝐽 of the 10000 simulated experiments at the 300th switching operation in both situations. European Journal of Control xxx (xxxx) xxx 5 E. Serrano-Seco et al. Fig. 3. Cost values with respect to the number of switching operations when parameter perturbations set to 5%. Each graph shows the median (𝑃50) and the 10th and 90th percentiles (𝑃10 and 𝑃90, respectively) of the distribution of values obtained for the 10 000 simulated experiments. The cost without control is also represented. The 97.5th percentile (𝑃97.5) is also represented in (b) and (d) to show the hybrid method improvement. Fig. 4. Effect of the hybrid strategy versus derivative-free strategy. Evolution of 𝐽 in two specific processes selected as representative. (a) Process with slow convergence. (b) Process with convergence to an unacceptable cost. 5. Conclusions In this work, we have presented R2R-AC (Hy), a run-to-run control scheme with a new algorithm for iteratively adapting the parameters of a feedforward controller from indirect measurements. However, as outlined, the new algorithm could be used for other differentiably parametrized controllers. The improvement over the previous approach has been achieved both for small initial parameter errors and for larger errors, where the previous technique is not effective. The improvements have been obtained by integrating three concepts into the algorithm: a scheduled basis change based on the sensitivity of the feedforward law, a continuous process of exploration and exploitation, and a hybrid method that toggles between a derivative-free and a gradient-based strategy to calculate the step size. Likewise, this new algorithm automates two aspects of our previous work: the update of the coordinate system and the number of search dimensions to reduce, as the new Fig. 5. Cost values with respect to the number of switching operations when parameter perturbations set to 25%. Each graph shows the median (𝑃50) and the 10th and 90th percentiles (𝑃10 and 𝑃90, respectively) of the distribution of values obtained for the 10000 simulated experiments. The cost without control is also represented. Table 3 Comparison of 𝑃90 of 𝐽(m∕𝐬) in the 300th switching operation. Perturbation R2R-PS+ R2R-AC (DF) R2R-AC (GB) R2R-AC (Hy) 5% 0.2268 0.1777 0.3178 0.1772 25% 1.2381 0.5541 0.3750 0.2645 algorithm selects the minimum number of dimensions to improve its feedforward controller behavior. As future work, we would like to address the possibility of an estimation technique for the objective function 𝐽∗ or to analyze and improve the gradient-based step-size calculation in other scenarios, such as stochastic processes or noisy measurements, by considering alternative approaches from the literature. In addition, we also intend to perform real laboratory tests on different systems to verify that the experimental results agree with those observed in simulation and the generality of the method. CRediT authorship contribution statement Eloy Serrano-Seco: Data curation, Writing – original draft, Investigation, Conceptualization, Formal analysis, Software, Methodology, Visualization, Validation. Eduardo Moya-Lasheras: Supervision, Writing – review & editing, Conceptualization. Edgar Ramirez-Laboreo: Supervision, Funding acquisition, Writing – review & editing, Project administration, Conceptualization. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. European Journal of Control xxx (xxxx) xxx 6 E. Serrano-Seco et al. Acknowledgments This work was supported in part via grants CPP2021-008938, PID2021-124137OB-I00, and TED2021-130224B-I00, funded by MCIN/ AEI/ 10.13039/501100011033, by the European Union NextGenerationEU/PRTR, and by ERDF A way of making Europe, in part by the Government of Aragón - EU, under grant T45_23R, in part by the ‘‘Programa Investigo’’ funded by the European Union - Next Generation EU, and in part by Fundación Ibercaja and the University of Zaragoza, Spain, via project JIUZ2023-IA-07. References Al Saaideh, M., Boker, A. M., & Al Janaideh, M. (2022). Output-feedback control of electromagnetic actuated micropositioning system with uncertain nonlinearities and unknown gap variation. In IEEE conf. decision and control (pp. 2481–2486). IEEE. Angadi, S. V., & Jackson, R. L. (2022). A critical review on the solenoid valve reliability, performance and remaining useful life including its industrial applications. Eng. Fail. Anal., 136, Article 106231. Benosman, M., & Atınç, G. M. (2015). Extremum seeking-based adaptive control for electromagnetic actuators. International Journal of Control, 88(3), 517–530. Braun, T., Reuter, J., & Rudolph, J. (2018). Observer design for self-sensing of solenoid actuators with application to soft landing. IEEE Transactions on Control Systems Technology, 27(4), 1720–1727. Deschaux, F., Gouaisbaut, F., & Ariba, Y. (2018). Nonlinear control for an uncertain electromagnetic actuator. In IEEE conf. decision and control (pp. 2316–2321). IEEE. Gergič, B., & Hercog, D. (2019). Design and implementation of a measurement system for high-speed testing of electromechanical relays. Measurement, 135, 112–121. Grotjahn, M., & Heimann, B. (2002). Model-based feedforward control in industrial robotics. International Journal of Robotics Research, 21(1), 45–60. Hansen, N., & Ostermeier, A. (2001). Completely derandomized self-adaptation in evolution strategies. Evol. Comput., 9(2), 159–195. Lévine, J. (2011). On necessary and sufficient conditions for differential flatness. Appl. Algebra Eng., Commun. Comput., 22(1), 47–90. Lewis, R. M., & Torczon, V. (2000). Pattern search methods for linearly constrained minimization. SIAM J. Optimization, 10(3), 917–941. Loshchilov, I., Schoenauer, M., & Sebag, M. (2011). Adaptive coordinate descent. In Prod. 13th GECCO (pp. 885–892). Ma, C., Wu, L., & Weinan, E. (2022). A qualitative study of the dynamic behavior for adaptive gradient algorithms. In Math. sci. mach. learning (pp. 671–692). PMLR. Moulay, E., Léchappé, V., & Plestan, F. (2019). Properties of the sign gradient descent algorithms. Inf. Sci., 492, 29–39. Moya-Lasheras, E., Ramirez-Laboreo, E., & Serrano-Seco, E. (2023). Run-to-Run Adaptive Nonlinear Feedforward Control of Electromechanical Switching Devices. IFAC-PapersOnLine, 56(2), 5358–5363, 22nd IFAC World Congr. Moya-Lasheras, E., & Sagues, C. (2024). Iterative-only learning control for soft landing of short-stroke reluctance actuators. Control Engineering Practice, 152, Article 106067. Ramirez-Laboreo, E., Moya-Lasheras, E., & Serrano-Seco, E. (2024). Faster run-to-run feedforward control of electromechanical switching devices: A sensitivity-based approach. In 2024 European control conf. (pp. 1321–1326). Schroedter, R., Roth, M., Janschek, K., & Sandner, T. (2018). Flatness-based openloop and closed-loop control for electrostatic quasi-static microscanners using jerk-limited trajectory design. Mechatronics, 56, 318–331. Winkel, F., Scholz, P., Wallscheid, O., & Böcker, J. (2023). Reducing contact bouncing of a relay by optimizing the switch signal during run-time. IEEE Trans. Autom. Sci. Eng.. Yeh, S.-S., & Hsu, P.-L. (1999). An optimal and adaptive design of the feedforward motion controller. IEEE/ASME Transactions on Mechatronics, 4(4), 428–439. European Journal of Control xxx (xxxx) xxx 7