scieee AI-readable full text Open interactive document viewer

On various multi-layer perceptron and radial basis function based artificial neural networks in the process of a hot flow curve description

Opěla, Petr

Abstract

In recent years, the study of the hot deformation behavior of various materials is significantly marked by an increasing utilization of artificial neural networks, which are frequently employed for a hot flow curve description. This specific kind of description is commonly solved via a Feed-Forward Multi-Layer Perceptron architecture and rarely via a Radial Basis architecture. Both network architectures are compared to assess their suitability in the process of a hot flow curve description under a wide range of thermomechanical conditions. The performed survey is also aimed on the eventual utilization of corresponding modifications of both studied networks, namely on a Cascade-Forward Multi-Layer Perceptron and Generalized Regression network. The main results have shown that the Feed-Forward Multi-Layer Perceptron architecture represents a good choice if very high accuracy is a crucial goal. However, in the case of this architecture, finding the proper parameters can be time-consuming and the hardware burdensome. On the contrary, for the flow curve description the almost unused Radial Basis network offers a very easy training procedure and significantly shorter computing time under acceptable accuracy. The results of the submitted research should then serve as a background for the selection and following application of a suitable network architecture in the process of solving future flow curve description tasks.

Full text

Original Article On various multi-layer perceptron and radial basis function based artificial neural networks in the process of a hot flow curve description Petr Op ela * , Ivo Schindler, Petr Kawulok, Rostislav Kawulok, Stanislav Rusz, Horymı ´r Navr atil Faculty of Materials Science and Technology, VSBeTechnical University of Ostrava, 17. Listopadu 2172/15, 70800 OstravaePoruba, Czech Republic article info Article history: Received 3 May 2021 Accepted 21 July 2021 Available online 24 July 2021 Keywords: Hot deformation behavior Hot flow curve description Multi-layer feed-forward network Multi-layer cascade-forward network Radial basis network Generalized regression network abstract In recent years, the study of the hot deformation behavior of various materials is significantly marked by an increasing utilization of artificial neural networks, which are frequently employed for a hot flow curve description. This specific kind of description is commonly solved via a Feed-Forward Multi-Layer Perceptron architecture and rarely via a Radial Basis architecture. Both network architectures are compared to assess their suitability in the process of a hot flow curve description under a wide range of thermomechanical conditions. The performed survey is also aimed on the eventual utilization of corresponding modifications of both studied networks, namely on a Cascade-Forward Multi-Layer Perceptron and Generalized Regression network. The main results have shown that the Feed-Forward Multi-Layer Perceptron architecture represents a good choice if very high accuracy is a crucial goal. However, in the case of this architecture, finding the proper parameters can be time-consuming and the hardware burdensome. On the contrary, for the flow curve description the almost unused Radial Basis network offers a very easy training procedure and significantly shorter computing time under acceptable accuracy. The results of the submitted research should then serve as a background for the selection and following application of a suitable network architecture in the process of solving future flow curve description tasks. ©2021 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). 1. Introduction Efforts to mathematically express a flow stress evolution of by-deformation loaded material can be dated back to the beginning of the 20th century [1]. Since the beginning of these endeavors, various mathematical or physical-mathematical based approaches have been employed to deal with an approximation of experimentally acquired flow curve datasets in the pursuit of assembly dependable prediction models *Corresponding author. E-mail address: [email protected] (P. Op ela). Available online at www.sciencedirect.com journal homepage: www.elsevier.com/locate/jmrt journal of materials research and technology 2021;14:1837e1847 https://doi.org/10.1016/j.jmrt.2021.07.100 2238-7854/©2021 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http:// creativecommons.org/licenses/by-nc-nd/4.0/). [1e4]. This gave rise to a number of different and still-used flow stress models [1,2,5e14]efrom the simpler to the more sophisticated and aimed primarily on the description under hot deformation conditions (Fig. 1), see e.g., Cingara & McQueen's relation [7] describing the flow curve increasing part or Ebrahimi's equation [8] considering a rapid flow stress decrease. Time has shown that the utilization of some of these models has become more popular, see the by-straincompensated Garofalo's inverse-hyperbolic-sine equation, for instance [15,16]. Several models have also undergone later modifications [17e24] in order to increase their reliability or the coverage of deformation conditions. Nevertheless, despite a variety of proposed models, the accuracy of flow curve descriptions still has the potential for improvement eespecially under hot deformation circumstances, since the well-known competition between work hardening and softening phenomena (dynamic recrystallization and dynamic recovery) makes a hot flow curve description more difficult [1e4]. Recent years have been significantly marked by an increasing interest in the practical application of artificial neural networks (ANN), see the overview in [25] or some specific applications, e.g., predicting the occurrence of disinfection by-products in tap water [26e28] or utilization in the field of wastewater treatment [29,30]. So, it is not a big surprise that the occurrence of this phenomenon came to be inevitable even in the field of flow curve descriptions. As it is known, the flow curve description issue embodies a nonlinear regression task ehighly noticeable especially under hot deformation circumstances [1e4]. It is understandable that different ANN architectures, as well as their various learning (training) techniques [31], making it possible to solve various tasks, have been designed since the first mention of an ANN came into the world [32]. It can be said that the flagship position, in the case of the discussed flow curve regression, is occupied by the socalled feed-forward multi-layer perceptron networks [33,34]. This ANN type was, for instance, utilized to describe the flow behavior of the BT25 titanium alloy [35], 6061 aluminum alloy [36], as-extruded 7075 aluminum alloy [37], Al-6.2Zn-0.70Mg0.30Mn-0.17Zr alloy [38] and inconel 718 superalloy [39]. It is worth mentioning that the same type of network architecture has been, in some cases, incorporated inside the abovementioned flow stress models eresulting in hybrid approaches [4,40e42]. A related cascade-forward variation was then applied to predict the flow stress curves of near btitanium alloy [43]. Another network type, which has also been considered to solve the flow curve description issue, nevertheless to a somewhat lesser extent, is relatively younger and operating in a different way esee so-called radial basis neural networks [44,45] and the corresponding flow stress descriptions of Al-base metalematrix composites [46] and Ni 3 Albased superalloy JG4246A [47]. Ongoing research in the ANN field then brought attempts to improve neural network training methods. With respect to the above-mentioned multi-layer perceptrons, a pretraining phase is sometimes applied since the training of these networks can be a hard nut to crack, especially when operating with a higher layer number. This pretraining phase has then also been utilized in some research dealing with the flow curve description. Specifically, the pretraining has been in these works carried out via individual restricted Boltzmann machines [48] stacked into one network, see the deep belief networks [49] and their applications in the case of the AleZneMgeCu alloy [50] and Ni-based superalloy [51]. It is worth noting that besides the aforementioned ANN approaches, there is another machine learning approach also suitable for a regression analysis and growing in popularity in the field of the hot flow curve description ethe talk is about a support vector machine [52], see applications in [53e56]. The conducted research then proved that both artificial neural networks and support vector machines have higher description accuracy than the abovementioned flow stress models. In the framework of the submitted research, an experimental hot-compression flow curve dataset of chromiummolybdenum steel is descripted by means of four different artificial neural network architectures. The research is aimed at those ANN architectures which are used for the flow curve description both frequently, i.e., Feed-Forward Multi-Layer Perceptron (FF-MLP), and rarely, like Radial Basis Neural Network (RBNN). In addition, slightly different versions of these architectures have been selected in order to assess their potential suitability for flow curve description purposes e namely a Cascade-Forward Multi-Layer Perceptron (CF-MLP) and Generalized Regression Neural Network (GRNN) [57]. The main goal is to evaluate all proposed architectures in relation to their flow curve description possibilities and assess in this regard primarily the description accuracy and even their advantages and eventual disadvantages. 2. Experimental hot flow curve dataset A hot flow curve dataset of a chromium-molybdenum steel (delivered by T RINECK E ZELEZ ARNY, a.s. with a composition in wt.%: 0.29 C, 0.79 Cr, 0.21 Mo, 1.20 Mn, 0.27 Si) has been employed in order to demonstrate the description possibilities of the above-mentioned ANN architectures. The dataset acquirement was realized via a series of hot uniaxial compression tests conducted by means of a Gleeble 3800. Cylindrical hot-compression samples with a diameter of 10 mm and a length of 15 mm were always treated by a preheating regime consisting of a direct heating up to a Fig. 1 eInfluence of thermomechanical conditions and related softening mechanisms on the flow stress course under hot deformation conditions. journal of materials research and technology 2021;14:1837e18471838 deformation temperature by a heating rate of 5 K s 1 with a following dwell-time of 300 s. A uniaxial compression deformation up to a true (logarithmic) strain of 1.0 has then been applied under strain rates of 0.02, 0.2, 2 and 20 s 1 in combination with deformation temperatures of 1043, 1113, 1203, 1303, 1413, and 1553 K. A detailed experimental description is then available in [58]. The conducted experiment has yielded in a set of twenty-four flow curves covering combinations of six deformation temperatures and four strain rates. 3. Flow curve description via artificial neural networks 3.1. Data-preparation phase As outlined in the section 1., the flow curve description issue is going to be solved via four ANN architectures: FF-MLP, its modification CF-MLP, RBNN and its modified form GRNN. Each selected network provides a functional relationship between input and output variables on the basis of individual computational units (artificial neurons) communicating via synaptic weights and clustered into layers (i.e., forming an artificial neural network) [31,59]. In the current research, the network input and output variables are represented by a deformation temperature, T(K), strain rate, ε_(s 1 ), true (logarithmic) strain, ε(), and true flow stress, s(MPa), respectively. Although an inner calculation mechanism and training methodism can be more or less different, all the mentioned architectures can undergo an identical data-preparing phase as stated below. In order to deal with the ANN overfitting issue [30], the very first step of the preparation phase is to divide the flow curve dataset in two groups ea training set (directly involved in network training) and a testing set (intended for an evaluation of the prediction capability after the training step) [60]. Previous research in the field of flow curve descriptions via the ANN approach [38,50,51] utilizes a training part range from 70% to 80% of the total data. Based on these experiences, the studied dataset has been divided as follows: 75% (900 ordered quadruples, T-ε_-ε-s) and 25% (300 ordered quadruples) of the entire dataset are chosen for the training and testing part, respectively. Both the training and testing data are selected to cover the entire field of thermomechanical conditions (see Table 1). As a second step, in order to increase the accuracy and convergence rate, the data-preparing phase consists of data normalization since the input data are often distributed in various ranges and dimensions. To do so, the vectors of the individual input variables have been normalized to have zero mean values and unity standard deviations esee the detailed procedure in [61]. 3.2. Essence of used networks and corresponding training methodism After the preparation phase introduced above is done, it is possible to proceed to the network training procedure ethe basis of the network adaptation. However, despite the fact that the preparation phase was performed mutually for all the considered networks, the network training will have to be realized individually because of their different architecture. Only one training feature was set as mutual for all ea network performance function in the form of a mean squared error [62]. 3.2.1. Feed-forward and Cascade-Forward Multi-Layer Perceptron As mentioned in the section 1., flow curve description tasks are solved mainly via the FF-MLP architecture ea basic block diagram (Fig. 2) demonstrates the FF-MLP utilized for the purpose of this research. As indicated, the input matrix (consisting of three vectors) is processed through artificial neurons clustered into one or more hidden layers which are together with one output layer stacked to form a single multilayer network. Each neuron computes a vector sum of all previous-layer neuron outputs which are multiplied elementby-element by individual synaptic weights. This sum of products is added to a specific bias value, and the resulting weighted sum is processed through an activation function to gain a final neuron output vector [59,63]. The neuron activation functions have been selected on the basis of previous experiences [4,42,61] to be a hyperbolic tangent sigmoid and pure linear [59] as regards to the hidden and output layers, respectively. The training procedure (finding of optimal weights and biases, i.e., minimizing the network performance function) of the introduced FF-MLP network has been realized via the wellknown LevenbergeMarquardt (LM) gradient-based iterative optimization algorithm [64e66] under the back propagation of a network error [67]. Note, a network generalization capability was treated by introducing the Bayesian regularization [68,69] into the utilized LM algorithm. The same training procedure has been applied also in the case of the CF-MLP architecture. Compared to the FF-MLP, the CF-MLP variant has one essential modification eeach layer (no matter if hidden or output) is individually linked with the input matrix and also with all previous layers [59]esee a simplified block diagram in Fig. 3. Note, finding the proper number of hidden layers and their neurons is, in the case of both above-introduced MLP versions, solved via the adaptation phase eas will be discussed later. 3.2.2. Radial basis and Generalized Regression Neural Networks The conception of the RBNN architecture [44,45] is, when compared with the MLP introduced above, somewhat different esee a block diagram in Fig. 4 representing an RBNN utilized in the current research. At first, the input matrix is processed through radial basis (RB) neurons arranged into one Table 1 eDistribution of the dataset into a training and testing part. ε_(s 1 )/T(K) 1043 1113 1203 1303 1413 1553 0.02 train test train train test train 0.2 train train test train train train 2 test train train test train train 20 train train train train train test journal of materials research and technology 2021;14:1837e1847 1839 RB layer. The number of RB neurons can theoretically be equal to the number of T-ε_-ε-straining quadruples (i.e., 900, see the data split introduced above). Each neuron of the RB layer then accepts all three input vectors and also three specific weight values. These specific weights are for each neuron equal to a different input T-ε_-εtriplet, specifically, the first neuron accepts a weight triplet equal to the first input triplet, etc. [59]. In the frame of a specific RB neuron, Euclidean distances [70,71] between each input triplet and a specific-neuron-weight triplet are computed, and the resulting vector of Euclidean distances is then multiplied element-by-element by a neuron bias value. The weighted input obtained in this way is processed through the Gaussian radial basis activation function (G.rb) [59,72] to get a neuron output vector. Note, all bias values of the RB layer neurons are equal to a quotient between a value of 0.8326 and a specific G.rb spread. The final networkoutput vector is then calculated as a linear combination of the RB-layer outputs added to a linear layer bias value [59]. The nature of the RB layer means that an RBNN training is limited on finding the optimal weight and bias values only in the last (linear) layer [59]ewhich is easily feasible via the linear least square method [73]. As mentioned above, the number of RB neurons is theoretically equal to the number of T-ε_-ε-straining quadruples. However, one or more RB-layer outputs will become redundant if the weights connecting these outputs with the linear neuron are equal to zero ethe collinearity issue [74]. So, a real number of RB neurons can be reduced by a k-number of redundant neurons. Note, the knumber is closely related to the above-mentioned G.rb spread. Finding the proper G.rb spread is then a matter of the adaptation phase [59]. It should be noted that the described RBNN regression approach is in some ways similar to the support vector regression (SVR) mentioned in the introduction, specifically if the SVR employs a Gaussian kernel etheir similarities were studied e.g., in [75]. The GRNN architecture [57](Fig. 5) is almost the same as the RBNN architecture introduced above. The RB layer is practically identical, except for the fact that the RB-layer neuron number is directly equal to the number of T-ε_-ε-straining quadruples. The last, special linear (SL) layer is then constructed differently. The neuron of this SL layer accepts output vectors of all RB-layer neuros, and these are multiplied by specific weights equal to Fig. 2 eBlock diagram of the utilized Feed-Forward Multi-Layer Perceptron (FF-MLP) architecture. Fig. 3 eBlock diagram of the utilized Cascade-Forward Multi-Layer Perceptron (CF-MLP) architecture. journal of materials research and technology 2021;14:1837e18471840 the values of the output vector (s-singletons) ei.e., the first RBlayer neuron output is multiplied element-by-element by the first s-value, etc. The vectors of these individual products are thenaddedtogether,andtheresultingvector(inotherwordsthe vector sum of the weighted RB-layer neuron outputs) is then divided element-by-element by a vector sum of non-weighted RB-layer neuron outputs ea normalized weight function [76]. The resulting quotient is then considered to be the searched network output. Note, there is no bias value in the SL layer [59]. The firmly fixed and known weight values (taken from training inputeoutput vectors) have one essential consequence enetwork training is not required. As in the case of the RBNN, the adaptation phase is aimed only at searching for an appropriate value of the G.rb spread [59]. 4. Results and discussion of ANN adaptation phase As is known, the ANN adaptation phase is about finding the proper settings (hyperparameters) of specific network architecture. For the purposes of this research, appropriate hyperparameters of all the above-introduced ANN architectures have been established on the basis of a trial-error methodology under evaluation via the Root Mean Squared Error [62], RMSE (MPa): RMSE ¼ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 I,X I i¼1 ðTiAiÞ2 v u u t(1) In this equation, T i (MPa) and A i (MPa) symbolize the i-th target (experimental) and approximated true flow stress value, respectively. i¼[1, I]3ℕ, where Iis the number of Tε_-ε-squadruples in the training (I¼900) or testing (I¼300) dataset. The adaptation phase of the FF-MLP and CF-MLP architectures has two typical under-change hyperparameters ethe number of hidden layers, m, and its neurons, n, where the m and nvalues were changed in an interval of [1,3]3ℕand [1,20] 3ℕ, respectively (i.e., 60 network configuration). Based on the above-mentioned gradient methodology, each configuration has been trained under thirty pseudorandom initializations of weights and biases (i.e., 1800 examined combinations in total), and the corresponding sixty best RMSE values are then graphically expressed in Fig. 6. With regard to the FF-MLP architecture with one hidden layer (Fig. 6 (a)), it is evident that the RMSE-value of the training set is gradually declining with the increasing number of hidden neurons while the testing RMSE remains more or less unchanged. The observed slowlyopening scissors indicates an increasing network overtraining and thus reveals a range of unfavorable numbers of hidden neurons. The gap between the training and testing RMSE values is then also growing with the increasing number of hidden layers, and the range of unsuitable numbers of hidden neurons becomes wider. It is also noticeable that while the trend of the testing RMSE values remains practically unaffected by the number of hidden layers, the trend of the training RMSE is not linear anymore. Almost the same situation can then be observed in the case of the CF-MLP Fig. 4 eBlock diagram of the utilized Radial Basis Neural Network (RBNN) architecture. Fig. 5 eBlock diagram of the utilized Generalized Regression Neural Network (GRNN) architecture. journal of materials research and technology 2021;14:1837e1847 1841 architecture (Fig. 6 (def)). Based on the performed statistical observation while considering an RMSE value and overtraining issue (generalization capability), it was possible to select suitable settings as three hidden layers with three neurons and two layers with two neurons with respect to the FF-MLP and CF-MLP architecture, respectively. As mentioned above, the adaptation phase of the RBNN and GRNN architectures has only one hyperparameter to change ethe G.rb spread. The spread value was for both architectures changed in an interval of [0.01, 0.99] 3ℚwith a step of 0.01 and in an interval of [1, 100] 3ℚwith a step of 0.1 (i.e., 1090 examined combinations in total). The influence of the spread value on the RMSE function is graphically expressed in Fig. 7. With regard to the RBNN architecture (Fig. 7 (a)), it is visible that the training and testing RMSE functions are almost mirror images of each other. The training RMSE gradually increases with the growing spread and becomes saturated around a spread of 20, and even if the evolution of the testing RMSE under lower spread values seems to be not so unambiguous, the overall trend can be considered to be a mirror image of the training one. It is then clear that the spread values higher than 40 have practically no more effect on the change of RBNN accuracy. The situation around the GRNN architecture is then quite different. While an area of a significant RMSE change is given by the spread values from 0.01 to 40 with respect to the RBNN architecture (Fig. 7 (a)), a substantial RMSE change of the GRNN architecture takes place in a narrow range from 0.01 to ca. 1.00 (Fig. 7 (b)). In this range, the training RMSE is rapidly growing and then taking a saturation state. The testing set then practically follows the same trend, although its RMSE values are higher in the first stage. Taking into account both the training and testing RMSE, the ideal G.rb spread has been selected to be 45 and 0.52 with regards to the RBNN and GRNN architectures, respectively. As with the above-mentioned collinearity issue, the utilized G.rb spread is, in the case of the RBNN architecture, closely related to the number of redundant neurons, k, in the radial basis layer. In connection with the spread of 45, the k-value was determined to be 875 eso the number of neurons that really participated in the RB layer of the RBNN architecture is thus only equal to 25. The ideal hyperparameters of the individual ANN architectures are then summarized in Table 2. This table also contains the corresponding training and testing RMSE values as well as their differences. It is evident that the lowest RMSE values were achieved in the case of the FF-MLP architecture although the RMSE values of the CF-MLP and RBNN architectures are also feasible. The lowest difference between the training and testing RMSE values, which indicates the best results with respect to the network generalization, is then associated with the RBNN architecture. In all respects, the worst results are coupled with the GRNN architecture eit was practically not possible to find a G.rb spread leading to lower RMSE values and ensuring a better network generalization capability. The clustered histograms in Fig. 8 then offer a more detailed statistical evaluation of the ANN hyperparameters selected above (Table 2). These histograms capture the distribution of the true flow stress residues, D i (MPa), and include the corresponding mean values, m(MPa) [77], and standard Fig. 6 eEvaluation of the adaptation phase of the Feed-Forward Multi-Layer Perceptron (FF-MLP) and Cascade-Forward Multi-Layer Perceptron (CF-MLP) architectures. Fig. 7 eEvaluation of the adaptation phase of the Radial Basis Neural Network (RBNN) and Generalized Regression Neural Network (GRNN) architectures. journal of materials research and technology 2021;14:1837e18471842 deviations, d(MPa) [78] (Note, the T i and A i values have the same meaning as in Eq. (1)): Di¼TiAi(2) m¼1 I,X I i¼1 Di(3) d¼ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 I,X I i¼1 ðDimÞ2 v u u t(4) At first glance, the obvious phenomenon is the difference between the MLP-based and radial-basis-function-based architectures. While both MLP architectures have training and testing residues ranging between 10 and 10 MPa, the RBNN and especially the GRNN are burdened by a bigger variance. Approximately 91% of FF-MLP and 78% of CF-MLP training residues are then ranging in an even narrower range (from 4 to 4 MPa). This statistically-revealed accuracy of the MLP architectures is then confirmed by the graphical comparison of experimental and calculated flow curves (see the good fit between boxes and solid or dash lines in Fig. 9). The RBNN residues range from 50 to 20 MPa and from 40 to 20 MPa with regards to the training and testing set, respectively. However, residues lower than 30 MPa comprise no more than ca 0.33%, which practically corresponds with three training values and one testing value (see the absolute frequency in Fig. 8 subcharts (a) and (b)). Further analysis has revealed an interesting fact eall RBNN residues lower than 15 MPa or higher than 15 MPa are practically linked only with a true strain value of 0.02 and in two cases also with 0.04, i.e., with the beginning of deformation. So, purely statistically speaking, the RBNN description has comparable accuracy to the MLP-based architectures if the first datapoint of each flow curve is omitted. Nevertheless, despite this statistically favorable accuracy, the shape of the RBNN curves does not always correspond with the target curves, which is especially noticeable under a strain rate of 0.02 s 1 (see the dash-dot lines in Fig. 9). The residues of the GRNN description are then ranging in a wide interval efrom 70 to 40 MPa and from 50 to 50 MPa as regards to the training and testing set, respectively. This fact is, of course, reflected by the corresponding standard deviations which are, in comparison with the other three Fig. 8 eDistribution of the true flow stress residues. Table 2 eIdeal hyperparameters and related statistics of the examined ANN architectures. ANN Architecture Hidden Layers Hidden Neurons RB Neurons G.rb Spread RMSE (MPa) Train Test difference FF-MLP 3 3 ee2.434 4.082 1.648 CF-MLP 2 2 ee3.218 4.781 1.563 RBNN ee25 45 5.507 6.761 1.254 GRNN ee900 0.52 13.271 20.034 6.763 journal of materials research and technology 2021;14:1837e1847 1843 architectures, unduly high. It is also noticeable that the fraction of excessively large residues, especially those higher than 30 MPa, is not negligible. These statistical conclusions can then be linked to the observations in Fig. 9 esome dot curves are excessively out of the target flow stress level, which is getting worse mainly with the decreasing temperature. It can be stated that the observations performed above have practically confirmed the leadership position of the MLPbased architectures, specifically the FF-MLP architecture. Both the FF-MLP and CF-MLP architectures have reached almost identical accuracy. However, the FF-MLP is more accurate, and its adaptation phase took a shorter duration (about ½) due to the lower number of weight connections. It is clear that neither the RBNN nor the GRNN architecture reached the accuracy of the MLP-based architectures. Nevertheless, the accuracy of the RBNN architecture is not so extremely different as in the case of the GRNN architecture. In addition, the RBNN architecture has two significant advantages eits training procedure (and thus also the adaptation phase) is both very fast and easy to perform due to the linear character of the corresponding regression calculations. As mentioned above, the RBNN was examined under 1090 G.rb spread values and the total computing time was equal only to 2 min. For comparison, the MLP-based architectures, as mentioned above, have been both examined under 60 configurations with 30 pseudorandom weight initializations. Considering only one initialization, a computing time for all 60 configurations was approximately equal to 10 and 26 min for the FF-MLP and CFMLP architectures, respectively. Finding the proper hyperparameters of the RBNN architecture is thus much less burdensome for the computer hardware. Note, an acceleration of the MLP-based training would require a utilization of pretraining procedures [50,51]. However, the simplicity of the RBNN training makes it very easy to perform it e.g., in common spreadsheet applications which are highly impractical for the MLP-based networks. 5. Conclusion The submitted research brings a comparison among four artificial neural network architectures, namely a FeedForward Multi-Layer Perceptron (FF-MLP), Cascade-Forward Multi-Layer Perceptron (CF-MLP), Radial Basis Neural Network (RBNN) and Generalized Regression Neural Network (GRNN), in the process of a hot flow curve description. The research practically confirmed the leadership position of the FF-MLP architecture, since its achieved description accuracy is practically the best. Almost identical accuracy was then achieved only in the case of the CF-MLP architecture. Nevertheless, both MLP-based architectures are disadvantaged by a time-consuming and hardware-burdensome adaptation phase. The accuracy of the RBNN architecture is statistically very close to the MLP-based architectures and, at the same time, the RBNN architecture offers a significantly faster adaptation phase. However, the relatively favorable statistical accuracy, easy and fast training advantage are in the case of the studied material somewhat blemished by the flow curve shape not always corresponding with the target one. The GRNN architecture can then be considered entirely inappropriate, at least in the case of the studied material, for a flow curve description. Fig. 9 eComparison of target (experimental) and calculated flow curves. journal of materials research and technology 2021;14:1837e18471844 Declaration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Acknowledgments This research was funded by project no. CZ.02.1.01/0.0/0.0/ 17_049/0008399 from the EU and CR financial funds provided by the Operational Programme “Research, Development and Education”and in within the frame of the Student Grant Competition SP2021/73 funded by Ministry of Education, Youth and Sports of the Czech Republic. references [1] Gronostajski Z. The constitutive equations for FEM analysis. J Mater Process Technol 2000;106(1e3):40e4. https://doi.org/ 10.1016/S0924-0136(00)00635-X. [2] Lin YC, Chen XM. A critical review of experimental results and constitutive descriptions for metals and alloys in hot working. Mater Des 2011;32(4):1733e59. https://doi.org/ 10.1016/j.matdes.2010.11.048. [3] Ebrahimi R, Shafiei E. Mathematical modeling of single peak dynamic recrystallization flow stress curves in metallic alloys. In: Sztwiertnia K, editor. Recrystallization. Rijeka: InTech; 2012. p. 207e25. https://doi.org/10.5772/34445. [4] Op ela P, Kawulok P, Schindler I, Kawulok R, Rusz S, Navr atil H. On the zenerehollomon parameter, multi-layer perceptron and multivariate polynomials in the struggle for the peak and steady-state description. Metals 2020;10(11):1413. https://doi.org/10.3390/met10111413. [5] Fields DS, Backofen WA. Determination of strain hardening characteristics by torsion testing. Proc Am Soc Test Mater 1957;57:1259e72. [6] Johnson GR, Cook WH. A constitutive model and data for metals subjected to large strains, high strain rates and high temperatures. In: American defense preparedness association, organizer. Proceedings of the 7th international symposium on ballistics; 1983 Apr 19-21. p. 541e7. The Hague, Netherlands. [7] Cingara A, McQueen HJ. New formula for calculating flow curves from high temperature constitutive data for 300 austenitic steels. J Mater Process Technol 1992;36(1):31e42. https://doi.org/10.1016/0924-0136(92)90236-L. [8] Ebrahimi R, Zahiri SH, Najafizadeh A. Mathematical modelling of the stressestrain curves of Ti-IF steel at high temperature. J Mater Process Technol 2006;171(2):301e5. https://doi.org/10.1016/j.jmatprotec.2005.06.072. [9] Lin YC, Chen MS, Zhong J. Constitutive modeling for elevated temperature flow behavior of 42CrMo steel. Comput Mater Sci 2008;42(3):470e7. https://doi.org/10.1016/ j.commatsci.2007.08.011. [10] Momeni A, Dehghani K, Ebrahimi GR, Keshmiri H. Modeling the flow curve characteristics of 410 martensitic stainless steel under hot working condition. Metall Mater Trans 2010;41:2898e904. https://doi.org/10.1007/s11661-010-0350-z. [11] Solhjoo S. Determination of flow stress under hot deformation conditions. Mater Sci Eng, A 2012;552:566e8. https://doi.org/10.1016/j.msea.2012.05.057. [12] Hensel A, Spittel T. Kraftund Arbeitsbedarf bildsamer Formgebungsverfahren. 1st ed. Leipzig: Deutscher Verlag fu ¨r Grundstoffindustrie; 1978. [13] Schindler I, Kawulok P, O cen a sek V, Op ela P, Kawulok R, Rusz S. Flow stress and hot deformation activation energy of 6082 aluminium alloy influenced by initial structural state. Metals 2019;9(12):1248. https://doi.org/10.3390/met9121248. [14] Razali MK, Irani M, Joun MS. General modeling of flow stress curves of alloys at elevated temperatures using Bi-linearly interpolated or closed-form functions for material parameters. J Mater Res Technol 2019;8(3):2710e20. https:// doi.org/10.1016/j.jmrt.2019.04.007. [15] Wang Z, Wang A, Xie J, Liu P. Hot deformation behavior and strain-compensated constitutive equation of nano-sized SiC particle-reinforced Al-Si matrix composites. Materials 2020;13(8):1812. https://doi.org/10.3390/ma13081812. [16] Chen R, Zhang S, Liu X, Feng F. A flow stress model of 300M steel for isothermal tension. Materials 2021;14(2):252. https:// doi.org/10.3390/ma14020252. [17] Quan G, Tong Y, Luo G, Zhou J. A characterization for the flow behavior of 42CrMo steel. Comput Mater Sci 2010;50(1):167e71. https://doi.org/10.1016/ j.commatsci.2010.07.021. [18] Shen J, Hu L, Sun Y, Wan Z, Feng X, Ning Y. A comparative study on artificial neural network, phenomenological-based constitutive and modified fieldsebackofen models to predict flow stress in Ti-4Al-3V-2Mo-2Fe alloy. J Mater Eng Perform 2019;28:4302e15. https://doi.org/10.1007/s11665-019-04174-0. [19] Nayak KC, Date PP. Development of constitutive relationship for thermomechanical processing of Al-SiC composite eliminating deformation heating. J Mater Eng Perform 2019;28:5323e43. https://doi.org/10.1007/s11665-019-04277-8. [20] Akbari Z, Mirzadeh H, Cabrera JM. A simple constitutive model for predicting flow stress of medium carbon microalloyed steel during hot deformation. Mater Des 2015;77:126e31. https://doi.org/10.1016/j.matdes.2015.04.005. [21] Mohamadizadeh A, Zarei-Hanzaki A, Abedi HR. Modified constitutive analysis and activation energy evolution of a low-density steel considering the effects of deformation parameters. Mech Mater 2016;95:60e70. https://doi.org/ 10.1016/j.mechmat.2016.01.001. [22] Liu L, Wu YX, Gong H, Wang K. Modification of constitutive model and evolution of activation energy on 2219 aluminum alloy during warm deformation process. Trans Nonferrous Metals Soc China 2019;29(3):448e59. https://doi.org/10.1016/ S1003-6326(19)64954-X. [23] Wang F, Shen J, Zhang Y, Ning Y. A modified constitutive model for the description of the flow behavior of the Ti-10V2Fe-3Al alloy during hot plastic deformation. Metals 2019;9(8):844. https://doi.org/10.3390/met9080844. [24] Spigarelli S, El Mehtedi M. A new constitutive model for the plastic flow of metals at elevated temperatures. J Mater Eng Perform 2014;23:658e65. https://doi.org/10.1007/s11665-0130779-5. [25] Rabunal JR, Dorado J. Artificial neural networks in real-life applications. 2nd. London: Idea Group Publishing; 2006. [26] Deng Y, Zhou X, Shen J, Xiao G, Hong H, Lin H, et al. New methods based on Back Propagation (BP) and Radial Basis Function (RBF) Artificial Neural Networks (ANNs) for predicting the occurrence of haloketones in tap water. Sci Total Environ 2021;772:145534. https://doi.org/10.1016/ j.scitotenv.2021.145534. [27] Lin H, Dai Q, Zheng L, Hong H, Deng W, Wu F. Radial basis function artificial neural network able to accurately predict disinfection by-product levels in tap water: taking haloacetic acids as a case study. Chemosphere 2020;248:125999. https:// doi.org/10.1016/j.chemosphere.2020.125999. journal of materials research and technology 2021;14:1837e1847 1845