Full text
Eur. Phys. J. Spec. Top. https://doi.org/10.1140/epjs/s11734-025-01760-3 THE EUROPEAN PHYSICAL JOURNAL SPECIAL TOPICS Regular Article Impact of amplitude and phase damping noise on quantum reinforcement learning: challenges and opportunities Mar´ıa Laura Olivera-Atencio1,a, Lucas Lamata2,b,andJes´us Casado-Pascual1,3,c 1F´ısica Te´orica, Universidad de Sevilla, Apartado de Correos 1065, 41080 Seville, Spain 2Departamento de F´ısica At´omica, Molecular y Nuclear, Universidad de Sevilla, 41080 Seville, Spain 3Multidisciplinary Unit for Energy Science, Universidad de Sevilla, 41080 Seville, Spain Received 31 March 2025 / Accepted 23 June 2025 ©The Author(s) 2025 Abstract Quantum machine learning (QML) is an emerging field with significant potential, yet it remains highly susceptible to noise, which poses a major challenge to its practical implementation. While various noise mitigation strategies have been proposed to enhance algorithmic performance, the impact of noise is not fully understood. In this work, we investigate the effects of amplitude and phase damping noise on a quantum reinforcement learning algorithm. Through analytical and numerical analysis, we assess how these noise sources influence the learning process and overall performance. Our findings contribute to a deeper understanding of the role of noise in quantum learning algorithms and suggest that, rather than being purely detrimental, unavoidable noise may present opportunities to enhance QML processes. 1 Introduction Quantum machine learning (QML) is a rapidly growing field within quantum technologies that seeks to perform machine learning tasks more efficiently than classical supercomputers in terms of time, space, and energy resources [1–5]. Leveraging quantum superposition and entanglement, the goal is to enable a more scalable implementation of various machine learning algorithms using quantum computers [6]. A major challenge in quantum computing, which also affects QML, is the fragility of highly entangled manybody quantum states. These states are susceptible to interactions with unintended quantum systems, leading to decoherence and the loss of computational properties necessary to solve a given problem. However, an emerging perspective in QML explores decoherence and dissipation not only as an obstacle but also as a potential resource for enhancing quantum learning [7–13]. This approach is motivated by the fact that effective learning, both classical and quantum, often requires some form of nonlinearity. Since isolated quantum systems evolve linearly (i.e., unitarily), some form of coupling—whether through quantum measurement (projective, weak, etc.) or dissipative and/or dephasing processes governed by a master equation (Markovian or non-Markovian)—may play a crucial role in enabling richer learning dynamics. In this context, recent work has shown that carefully tuned amplitude, phase, and depolarizing noise can improve the performance of variational quantum algorithms, further supporting the idea that noise can be harnessed as a useful feature in QML [14]. In previous works, we have explored the role of thermal dissipation in QML protocols [10]. Our findings indicate that, rather than being purely detrimental, thermal dissipation can sometimes enhance the learning process. In this paper, we build upon this research by specifically analyzing phase damping noise (PDN) and amplitude damping noise (ADN) in a protocol of quantum reinforcement learning. By investigating these types of noise, we aim to deepen the understanding of when and how these types of noise can be beneficial in quantum learning protocols. The remainder of this work is structured as follows. In Sect. 2, we provide a brief description of the problem under study, the types of noise considered, and the algorithm to which they are applied. In Sect. 3,wepresentthe ae-mail: moliv[email protected] (corresponding author) be-mail: [email protected] ce-mail: [email protected] (corresponding author) 0123456789().: V,-vol 123
Eur. Phys. J. Spec. Top. numerical results and analyze the conditions under which noise can enhance the algorithm performance. Finally, in Sect. 4, we summarize our findings and conclusions. 2 Problem statement In the reinforcement learning algorithm presented in Ref. [15], the agent A is a known and controllable quantum system described by a state vector |φ, or equivalently, by the associated density operator ρ=|φφ|. For the sake of clarity, we will restrict ourselves to the simplest case, where the system is a single qubit with computational basis {|0,|1}. The interaction of the environment E with the agent A over a time interval τis characterized by the unitary time evolution operator U(τ)=e−iHτ/, where His an unknown Hamiltonian. This Hamiltonian, written in terms of its unknown excited state |eand ground state |g, takes the form H=ω 2(|ee|−|gg|), where ωis a characteristic frequency of the system. The goal of the algorithm is to “learn” how to construct, at least approximately, either of the stationary states |eor |g, despite the challenge posed by the fact that these stationary states are unknown. To this end, the algorithm exploits the fact that, for τ=τn=2nπ/ω with n being a natural number, the stationary states |ee|and |gg|are the only pure states invariant under the unitary time evolution, i.e., the only pure states satisfying the equation U(τ)|φφ|U†(τ)=|φφ|. Consequently, actions compatible with this property are rewarded, while those that are incompatible are penalized. The case τ=τnis an exception since U(τn)=(−1)nI, with Ibeing the identity operator, making all pure states satisfy the condition U(τn)|φφ|U†(τn)=|φφ|. For a detailed explanation of why this algorithm can be classified as a quantum reinforcement learning algorithm, see Refs. [10,15]. In the presence of noise, the time evolution that governs the interaction between the environment and the agent is no longer unitary. Specifically, for PDN and ADN—the cases analyzed in this work—the time evolution over an interval τtakes the form Eτ(ρ)=U(τ)E0(τ)ρE† 0(τ)+E1(τ)ρE† 1(τ)U†(τ), (1) where {E0(τ), E1(τ)}are Kraus operators [16] whose form depends on the type of noise considered [6]. Specifically, in the case of PDN, the Kraus operators and the corresponding time evolution take the form E0(τ)= |gg|+e−τ/TD|ee|,E1(τ)=√1−e−2τ/TD|ee|,and Eτ(ρ)=ρee|ee|+ρgg|gg|+e−τ/TDe−iωτ ρeg|eg|+eiωτ ρge|ge|, (2) where ραβ =α|ρ|β, with α,β∈{e,g},andTDis a parameter known as the decoherence time [17,18]. For ADN, the Kraus operator E0(τ) remains the same as for PDN, while E1(τ)=√1−e−2τ/TD|ge|, leading to the time evolution Eτ(ρ)=|gg|+ρeee−2τ/TD(|ee|−|gg|) +e−τ/TDe−iωτ ρeg|eg|+eiωτ ρge|ge|.(3) In this case, in addition to the suppression of the off-diagonal elements of the density operator in the eigenstate basis of H(decoherence), there is also a decay of the excited state |eto the ground state |gwith a mean decay time of TD/2. To some extent, this type of noise can be regarded as a specific instance of the thermal dissipation analyzed in Ref. [10], but at absolute temperature equal to zero. Note that, as in the unitary time evolution without noise, the stationary states |ee|and |gg|are the only pure states invariant under the non-unitary evolution induced by PDN [Eq. (2)]. In contrast, under ADN [Eq. (3)], only the ground state |gg|remains invariant. This asymmetry will be crucial to the algorithm performance in the presence of ADN, as discussed later. Next, we provide a detailed description of how the algorithm works in the presence of these types of noise. The algorithm consists of a large number of iterations, indexed by a natural number k. The unitary transformation generated in the kth iteration is denoted by Dk. Moreover, in each iteration, a numerical value within the interval [0, 1] is assigned to a parameter wk, called the exploration parameter. The goal is for Dk|φφ|D† k, with |φφ|being a given state, to gradually approach either of the target states |ee|or |gg|as the number of iterations increases. Starting with the initial values D0=Iand w0= 1, the values of Dk+1 and wk+1 are iteratively updated from the previous values Dkand wkaccording to the following steps: 1. The unitary transformation Dkis applied to one of the computational basis state, say, the state |00|,to construct the state ρk=Dk|00|D† k. 123
Eur. Phys. J. Spec. Top. 2. The resulting system evolves for a time τ, yielding the transformed state ρ k=Eτ(ρk), where Eτis given by Eq. (2)or(3), depending on the type of noise. 3. The initial unitary transformation is then reversed, yielding ρ k=D† kρ kDk. Note that, in the presence of PDN, if ρkhad reached one of the target states |ee|or |gg|, then, after the time evolution, ρ kwould remain in that state and, consequently, ρ kwould be equal to |00|. In the presence of ADN, this would only hold if ρk had reached the ground state |gg|. 4. A measurement in the computational basis is performed on the system obtained in the previous step, yielding a result mk∈{0, 1}. As deduced from the previous discussion, the measurement outcome mk= 0 is compatible with ρkhaving reached one of the target states, while the outcome mk= 1 is incompatible with this. 5. Depending on the outcome of mk, the following procedure is applied: •If mk= 0, since the outcome is compatible with having reached one of the target states, a reward is granted by reducing the exploration parameter according to wk+1 =rwk, where r∈(0, 1) is a parameter known as the reward rate. Additionally, the unitary transformation Dkremains unchanged, i.e., Dk+1 =Dk, and the process returns to step 1. •If mk= 1, since the outcome is incompatible with having reached one of the target states, a punishment is applied by increasing the exploration parameter to wk+1 = min(pwk, 1), where p>1 is a parameter known as the punishment rate and the min function ensures that wk+1 does not exceed 1. Additionally, three pseudorandom numbers αk,βk,andγkare generated, uniformly distributed within the exploration interval [−wkπ, wkπ], and used to construct the pseudo-random rotation Rk=Dke−iβkY/2e−iγkZ/2e−iαkX/2D† k, (4) where X=|01|+|10|,Y=−i(|01|−|10|), and Z=|00|−|11|are the Pauli operators. Finally, the unitary transformation Dkis updated to Dk+1 =RkDk=Dke−iβkY/2e−iγkZ/2e−iαkX/2, the qubit is restored to its original state |00|by applying the unitary transformation X, and the process returns to step 1. The previously described algorithm is considered to converge if the exploration parameter wkapproaches zero as kincreases. In this case, the pseudo-random rotations Rktend to the identity operator, and consequently, the unitary transformations Dkconverge to a constant value. Moreover, the faster wkapproaches zero, the faster the algorithm converges. 3Results To examine the impact of the previously discussed noise sources on the algorithm from the preceding section, we have implemented it using a Hamiltonian of the form H=ω 4(√3X−Z), (5) which corresponds to setting |e=(|0+√3|1)/2and|g=(−√3|0+|1)/2 in the Hamiltonian expression introduced earlier. To work with dimensionless quantities, we define the parameters ˜τ=ωτ and ˜ TD=ωTD.The measurement process described in item 4 of the previous section is simulated as follows: at each iteration, we calculate the probability of obtaining 0 as the measurement outcome using the expression Pk(0) = Tr(|00|ρ k)= Tr[|00|D† kEτ(ρk)Dk]=Tr[ρkEτ(ρk)], with Eτgiven by Eq. (2)or(3), depending on the type of noise considered. Then, a pseudo-random number χkis generated, uniformly distributed in the interval [0, 1]. If χk≤Pk(0), the measurement outcome is mk= 0; otherwise, it is mk=1. Thanks to the fact that the excited and ground states are known in the considered example, we can use this knowledge to assess the accuracy of the algorithm described in the previous section. To this end, at each iteration, we compute the square root fidelity between the state ρkand the stationary states |ee|and |gg|, given by f(e) k=|e|Dk|0| and f(g) k=|g|Dk|0|, respectively. Since, a priori, it is not known which of the two stationary states is closer to ρk, it is also convenient to consider the highest fidelity between f(e) kand f(g) k, i.e., fk= max(f(e) k, f(g) k). The closer the value of fkapproaches 1 as kincreases, the more accurate the algorithm estimation of one of the stationary states will be. It is worth mentioning that the quantities wk,f(e) k,f(g) k,andfkare random variables, whose values will vary from one realization of the algorithm to another. The randomness of these variables arises from two factors: first, the inherent stochastic nature of the measurement outcome mk, and second, the pseudo-random selection of the angles αk,βk,andγk. For this reason, it is convenient to perform a large number Nof realizations (in our calculations, 123
Eur. Phys. J. Spec. Top. Fig. 1 Mean fidelity Fkas a function of the number of iterations kin the presence of PDN (left panels) and ADN (right panels). Results are shown for different dimensionless decoherence times, namely, ˜ TD=1(red dotted lines), ˜ TD= 10 (blue dashed lines), ˜ TD= 100 (green dashed-dotted lines), and ˜ TD=∞(black solid lines), and for two values of the dimensionless evolution time, specifically, ˜τ= 1 (top panels) and ˜τ=2π(bottom panels). The case ˜ TD=∞ corresponds to the scenario with no noise, where the evolution is unitary we use N= 1000) and consider the arithmetic mean of the values for wk,f(e) k,f(g) k,andfkobtained in each realization, which will be denoted as Wk,F(e) k,F(g) k,andFk, respectively. In Fig. 1, we depict the mean fidelity Fkas a function of the number of iterations kin the presence of PDN (left panels) and ADN (right panels). The results are shown for several dimensionless decoherence times, namely, ˜ TD= 1 (red dotted lines), ˜ TD= 10 (blue dashed lines), ˜ TD= 100 (green dashed-dotted lines), and ˜ TD=∞(black solid lines), and for two values of the dimensionless evolution time, specifically, ˜τ= 1 (top panels) and ˜τ=2π (bottom panels). The decoherence time ˜ TD=∞corresponds to the case without noise, where the evolution is unitary. For all parameter values considered, we have verified the convergence of the algorithm by checking that the mean exploration parameter Wkdecreases to zero as the number of iterations kincreases sufficiently, although these results are not shown in the figure. As observed in the top left panel, for ˜τ= 1, the effect of PDN on the algorithm is minimal, allowing it to perform well even when the decoherence time is comparable to the system’s characteristic timescales. In fact, for ˜ TD= 1, the algorithm performs slightly better than in the absence of noise. In contrast, for ˜τ=2π, the presence of PDN considerably enhances the algorithm performance for dimensionless decoherence times that are not too large, specifically for ˜ TD=1and ˜ TD= 10, with better performance for smaller decoherence times (see bottom left panel). This occurs because, as mentioned in Sect. 2, the unitary evolution operator U(τ) becomes equal to minus the identity operator for τ=2π/ω, making all states invariant under the unitary evolution over a time 2π/ω. As a result, for ˜τ=2πand in the absence of noise, the algorithm is unable to distinguish between stationary and non-stationary states, causing it to fail. This is not the case in the presence of PDN, in which the stationary states |ee|and |gg|are the only pure states that remain invariant under the non-unitary evolution described by Eq. (2), even for ˜τ=2π. This allows the algorithm to distinguish between stationary and non-stationary states, resulting in a significant improvement in its performance with respect to the noise-free case. In the presence of ADN (right panels), the behavior is consistent with that observed for PDN, and the explanation for the improved performance in the noisy case compared to the noise-free case in the bottom right panel remains applicable. However, in this case, only the stationary state |gg|remains invariant under the non-unitary time evolution given by Eq. (3), unlike in the PDN case, where both stationary states were invariant. To analyze how this difference is reflected in the algorithm performance, Fig. 2shows the mean fidelities associated with the ground state, F(g) k(top panels), and the excited state, F(e) k(bottom panels), as a function of the number of iterations kfor the PDN case (left panels) and the ADN case (right panels). As seen in this figure, for PDN, the results for F(g) kand F(e) kshow little dependence on the dimensionless decoherence time ˜ TDand remain close to those obtained in the noise-free case, i.e., for ˜ TD=∞. In contrast, in the presence of ADN, the fidelity F(g) kincreases significantly as ˜ TDdecreases, reaching values close to 1 for ˜ TD=1and ˜ TD= 10, while F(e) kexhibits a substantial decline. This asymmetry in the behavior of F(g) kand F(e) k, which is observed for ADN but absent in 123
Eur. Phys. J. Spec. Top. Fig. 2 Mean fidelities associated with the ground state, F(g) k(top panels), and the excited state, F(e) k (bottom panels), as a function of the number of iterations kfor phase damping noise (left panels) and amplitude damping noise (right panels). The values of the dimensionless decoherence times ˜ TDare thesameasinFig.1,and the dimensionless evolution time is ˜τ=1 the PDN case, stems from the distinct properties of the stationary states under each type of noise. As previously mentioned, under PDN, both stationary states, |ee|and |gg|, remain invariant under the non-unitary time evolution given by Eq. (2). As a result, the algorithm converges to either of these states with similar probability, leading to the comparable values of F(g) kand F(e) kobserved in the left panels of Fig. 2. In contrast, under ADN [Eq. (3)], only the ground state |gg|remains invariant. Consequently, the algorithm preferentially converges to this state, leading to a significant enhancement of F(g) kwhile F(e) kdecreases accordingly. As the dimensionless decoherence time ˜ TDincreases, the non-unitary time evolution (3) gradually approaches the unitary case ˜ TD=∞, where both |ee|and |gg|are once again invariant. This progressively reduces the asymmetry observed for lower values of ˜ TD. According to the previous discussion, one might conclude that the presence of ADN would be advantageous primarily when aiming to construct the ground state |gg|, while it would be detrimental if the goal were to construct the excited state |ee|. However, the unitary transformations Dkobtained through the previously presented algorithm also allow for an approximate preparation of the excited state by applying them to the computational basis state |11|instead of |00|. Indeed, due to the unitary nature of the operators Dk, the state vectors Dk|1and Dk|0are orthogonal. Consequently, if Dk|00|D† kis close to the ground state |gg|, then Dk|11|D† kwill be close to the excited state |ee|. To confirm this, Fig. 3displays the mean fidelities associated with the ground state (top panels) and the excited state (bottom panels) as functions of the number of iterations k. The left panels correspond to the states Dk|00|D† k, while the right panels correspond to the states Dk|11|D† k. The values of ˜ TDand ˜τare the same as in Fig. 3. As seen in the figure, while the states Dk|00|D† kgradually approach the ground state |gg|as kincreases for ˜ TD=1and ˜ TD= 10 (top left panel), the states Dk|11|D† k similarly converge to the excited state |ee|(bottom right panel). In summary, the unitary transformations Dk allow for the calculation of both the ground and excited states by simply applying them to different computational basis states. 4 Conclusions In this work, we have analyzed the impact of two common types of noise—PDN and ADN—on the reinforcement learning quantum algorithm proposed in Ref. [15]. Through the study of specific examples, we have shown that the presence of noise does not necessarily hinder the algorithm performance; in some cases, it can even have a beneficial effect. 123
Eur. Phys. J. Spec. Top. Fig. 3 Mean fidelities associated with the ground state, F(g) k(top panels), and the excited state, F(e) k (bottom panels), as a function of the number of iterations k. The left panels correspond to the states Dk|00|D† k, while the right panels correspond to the states Dk|11|D† k.The values of ˜ TDand ˜τare the same as in Fig. 2 In particular, we have demonstrated that for certain values of the evolution time τ, the presence of noise has little impact when the algorithm accuracy is assessed using the mean fidelity Fk. However, for other values of τ, noise can significantly enhance the algorithm performance. Furthermore, we have shown that the two types of noise affect the stationary-state fidelities, F(g) kand F(e) k, in markedly different ways. While PDN influences both fidelities symmetrically, ADN introduces an asymmetry, favoring convergence to the ground state over the excited state. We have explained this difference by analyzing the pure states that remain invariant under the non-unitary evolution associated with each type of noise. Although this asymmetry might suggest that ADN enhances the preparation of the system in the ground state compared to the noise-free case, we have also demonstrated that the unitary transformation generated by the algorithm also enables the preparation of the excited state simply by applying it to a different computational basis state. In this work, for the sake of clarity and simplicity in the description, we have restricted ourselves to the case of a single qubit. However, in future research, we aim to analyze how the results described here generalize when the number of qubits is increased. Acknowledgements The authors acknowledge project PID2022-136228NB-C22 funded by MCIN/AEI/ 10.13039/501100011033 and by “ERDF A way of making Europe”, EU. Furthermore, this work has been partially financially supported by the Ministry of Economic Affairs and Digital Transformation of the Spanish Government through the QUANTUM ENIA project call—Quantum Spain project, and by the European Union through the Recovery, Transformation and Resilience Plan—NextGenerationEU within the framework of the “Digital Spain 2026 Agenda”. It has also been co-financed by EU, Ministerio de Hacienda y Funci´on P´ublica, FEDER and Junta de Andaluc´ıa (project SOL2024-31833). Funding Funding for open access publishing: Universidad de Sevilla/CBUA. Data availability All simulation scripts used to produce the results presented in this paper are available at https://gi thub.com/MLOA25/YQS. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. 123
Eur. Phys. J. Spec. Top. References 1. J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, S. Lloyd, Quantum machine learning. Nature 549(7671), 195–202 (2017). https://doi.org/10.1038/nature23474 2. M. Schuld, F. Petruccione, Machine Learning with Quantum Computers. Quantum Science and Technology, 2nd edn. (Springer, Cham, 2021) 3. A. Melnikov, M. Kordzanganeh, A. Alodjants, R.-K. Lee, Quantum machine learning: from physics to software engineering. Adv. Phys. X 8(1), 2165452 (2023). https://doi.org/10.1080/23746149.2023.2165452 4. L. Lamata, Quantum machine learning implementations: proposals and experiments. Adv. Quantum Technol. 6(7), 2300059 (2023). https://doi.org/10.1002/qute.202300059 5. Y. Wang, J. Liu, A comprehensive review of quantum machine learning: from NISQ to fault tolerance. Rep. Prog. Phys. 87(11), 116402 (2024). https://doi.org/10.1088/1361-6633/ad7f69 6. M.A. Nielsen, I.L. Chuang, Quantum Computing and Quantum Information (Cambridge University Press, Cambridge, 2000) 7. E. Ghasemian, M.K. Tovassoly, Generation of Werner-like states via a two-qubit system plunged in a thermal reservoir and their application in solving binary classification problems. Sci. Rep. 11, 3554 (2021). https://doi.org/10.1038/s4 1598-021-82880-32 8. E. Ghasemian, M.K. Tavassoly, Hybrid classical-quantum machine learning based on dissipative two-qubit channels. Sci. Rep. 12(1), 20440 (2022). https://doi.org/10.1038/s41598-022-24346-8 9. E. Ghasemian, Stationary states of a dissipative two-qubit quantum channel and their applications for quantum machine learning. Quantum Mach. Intell. 5, 13 (2023). https://doi.org/10.1007/s42484-023-00096-2 10. M.L. Olivera-Atencio, L. Lamata, M. Morillo, J. Casado-Pascual, Quantum reinforcement learning in the presence of thermal dissipation. Phys. Rev. E 108, 014128 (2023). https://doi.org/10.1103/PhysRevE.108.014128 11. L. Domingo, G. Carlo, F. Borondo, Taking advantage of noise in quantum reservoir computing. Sci. Rep. 13(1), 8790 (2023). https://doi.org/10.1038/s41598-023-35461-5 12. A. Sannia, R. Mart´ınez-Pe˜na, M.C. Soriano, G.L. Giorgi, R. Zambrini, Dissipation as a resource for quantum reservoir computing. Quantum 8, 1291 (2024). https://doi.org/10.22331/q-2024-03-20-1291 13. M.L. Olivera-Atencio, L. Lamata, J. Casado-Pascual, Benefits of open quantum systems for quantum machine learning. Adv. Quantum Technol. (2023). https://doi.org/10.1002/qute.202300247 14. W. Somogyi, E. Pankovets, V. Kuzmin, A. Melnikov, Method for noise-induced regularization in quantum neural networks (2024). arXiv:2410.19921 15. F. Albarr´an-Arriagada, J.C. Retamal, E. Solano, L. Lamata, Reinforcement learning for semi-autonomous approximate quantum Eigensolver. Mach. Learn. Sci. Technol. 1(1), 015002 (2020). https://doi.org/10.1088/2632-2153/ab43b4 16. K. Kraus, States, Effects, and Operations. Lecture Notes in Physics, 1983rd edn. (Springer, Berlin, 1983) 17. W.H. Zurek, Decoherence, einselection, and the quantum origins of the classical. Rev. Mod. Phys. 75, 715–775 (2003). https://doi.org/10.1103/RevModPhys.75.715 18. H.-P. Breuer, F. Petruccione, Theory of Open Quantum Systems (Oxford University Press, Oxford, 2003) 123