Please cite this article as: O. Marinov, M. Jamal Deen, Juan A. Jiménez-Tejada, Low-frequency noise in downscaled silicon transistors: Trends, theory and practice, Physics Reports 990 (2022) 1–179 ©2023. This manuscript version is made available under the CC-BY-NCND 4.0 license http://creativecommons.org/licenses/by-nc-nd/4.0/ Digital Object Identifier: 10.1016/j.physrep.2022.06.005 Source: https://www.sciencedirect.com/science/article/abs/pii/S0370157322002502?via%3Dihub
1 of 286 Low-Frequency Noise in Downscaled Silicon Transistors: Trends, Theory and Practice O. Marinov 1 , M. Jamal Deen 1 and Juan A. Jiménez-Tejada 2 1 Electrical & Computer Engineering, McMaster University, Hamilton, ON L8S 4K1, Canada (
[email protected]) 2 Departamento de Electrónica y Tecnología de Computadores, Universidad de Granada, 18071 Granada, Spain Abstract By the continuing downscaling of sub-micron transistors in the range of few to one deca-nanometers, we focus on the increasing relative level of the low-frequency noise in these devices. Large amount of published data and models are reviewed and summarized, in order to capture the state-of-the-art, and to observe that the 1/area scaling of low-frequency noise holds even for carbon nanotube devices, but the noise becomes too large in order to have fully deterministic devices with area less than 10nm×10nm. The low-frequency noise models are discussed from the point of view that the noise can be both intrinsic and coupled to the charge transport in the devices, which provided a coherent picture, and more interestingly, showed that the models converge each to other, despite the many issues that one can find for the physical origin of each model. Several derivations are made to explain crossovers in noise spectra, variable random telegraph amplitudes, duality between energy and distance of charge traps, behaviors and trends for figures of merit by device downscaling, practical constraints for micropower amplifiers and dependence of phase noise on the harmonics in the oscillation signal, uncertainty and techniques of averaging by noise characterization. We have also shown how the unavoidable statistical variations by fabrication is embedded in the devices as a spatial “frozen noise”, which also follows 1/area scaling law and limits the production yield, from one side, and from other side, the “frozen noise” contributes generically to temporal 1/f noise by randomly probing the embedded variations during device operation, owing to the purely statistical accumulation of variance that follows from cause-consequence principle, and irrespectively of the actual physical process. The accumulation of variance is known as statistics of “innovation variance”, which explains the nearly log-normal distributions in the values for low-frequency noise parameters gathered from different devices, bias and other conditions, thus, the origin of geometric averaging in lowfrequency noise characterizations. At present, the many models generally coincide each with other, and what makes the difference, are the values, which, however, scatter prominently in nanodevices. Perhaps, one should make some changes in the approach to the low-frequency noise in electronic devices, to emphasize the “statistics behind the numbers”, because the general physical assumptions in each model always fail at some point by the device downscaling, but irrespectively of that, the statistics works, since the low-frequency noise scales consistently with the 1/area law.
2 of 286 Contents I. Introduction II. Intrinsic (uncorrelated, uncoupled) and coupled (correlated) behavior in a fluctuating system (behind and beyond ∆µ – ∆n controversy) II.1. Coupled behavior II.2. Intrinsic behavior II.3. Combination of coupled and intrinsic fluctuations III. Noise in BJT III.1. Differences in measurement setups III.2. Differences in BJT fabrication, IFO III.3. Crossover between different noise sources in BJT III.3.1. Intrinsic noise for the current flow III.3.2. Base and collector currents are strongly correlated III.3.3. Lorentzian noise superimposes 1/f noise III.4. Measurement and characterization uncertainty – experimental accuracy, fitting and averaging III.4.1. Accuracy and the smoothness of the measured spectra III.4.2. Deviations from the ideal 1/f slope in the spectra III.4.3. Data processing IV. Noise in MOS transistors IV.1. Models and predictability IV.1.1. Intrinsic noise (mobility fluctuation) IV.1.2. Coupled noise component (number fluctuation) IV.1.3. Empirical factors that impact the scattering of data in noise measurements IV.1.4. Modeling factors to interpret the scattering of data in noise measurements IV.2. Charge trap profiling of gate dielectrics IV.2.1. Spatial profiling of trap density IV.2.2. Energy profiling of trap density IV.3. RTS noise in MOS transistors IV.3.1. Time constants of RTS noise IV.3.2. Amplitude of RTS noise IV.4. Figures of merit for MOS transistors IV.4.1. Definition IV.4.2. Input and output referred noise – scalability of normalized noise IV.4.3. Noise factor, noise resistance, noise temperature IV.4.4. Physical figures – trap density, Hooge parameter, scattering parameter IV.4.5. Performance figures – RF to LFN V. Noise in advanced Si-based transistors V.1. Noise in SiGe, stacked gate and strained transistors – higher performance, but higher noise too V.1.1. Observations in SiGe HBTs V.1.2. Observations in SiGe MOS transistors V.1.3. Modeling the 1/f noise in SiGe MOS transistors V.2. Forward body bias in MOS transistors – not a panacea, but it helps VI. Noise in advanced transistor structures VI.1. From SOI toward gate all around VI.1.1. Partially depleted SOI MOS transistors VI.1.2. Fully depleted SOI MOS transistors VI.1.3. Effects of oxide traps in the back gate insulator
3 of 286 VI.1.4. Capacitive coupling in the semiconductor body VI.1.5. Two-MOS two-junction gate transistor VI.2. Nanotubes and nanowires – 1D seems too noisy VI.3. Between 3D and 1D – the graphene and transition metal dichalcogenide 2D transistors VII. Impact of LFN in circuits VII.1. Trading power for noise VII.2. Upand cross-conversion in phase noise VII.2.1. Definitions VII.2.2. Phase-noise in oscillators VII.2.3. Suggestions for the design of oscillators VII.2.4. Summary VII.3. Noise in sensors VII.3.1. Noise in electrochemical sensors VII.3.2. Noise in photo detectors and imaging arrays sensors VIII. Outlook for the LFN VIII.1. Trend for the LFN level and variations VIII.2. Common assumptions for noise VIII.3. The fabrication “frozen noise” – from spatial variations to yield problems VIII.3.1. Length uncertainty variances VIII.3.2. Consequences from the length uncertainty variances VIII.3.3. Fabrication frozen noise. Definition VIII.3.4. Investigation of fabrication frozen noise VIII.4. Statistical accumulation of variance “innovates” 1/f noise VIII.5. Statistical origin of the temporal 1/f noise in relation with the spatial “frozen noise” VIII.6. Consequences from statistical nature of LFN – distributions in spectra, techniques of averaging, data volume and coordinates, instrumentation IX. Conclusion
4 of 286 I. Introduction The above aphorism is the famous summary of the ancient Greek philosopher Plato, said approximately 2500 years ago, regarding the thoughts of Heraclitus (the Ephesus) [1]. Being not loaded with the many interpretations of this aphorism, we can simply rephrase that everything varies, including the random variation, which one usually calls noise; it just needs time this to happen. Later, in section VIII, we will show that this can be origin of 1/f noise – the most difficult for physical interpretation noise in the nature, which always “snakes out” when attempting to describe it absolutely in finite values, but at infinite limits both in time and frequency. The low-frequency noise (LFN) is always present in electronic devices. However, obviously, the LFN is not “appreciated” due to the fact that it is assumed as undesirable effect, and small enough not to bother much, as compared to other more important and definitely physically better sound effects and useful for the practice properties of the electronic devices, such as gain, high frequency of operation, versatility in making of functions, etc. We can cite again the “present-day assessment” from [2] made in 1981, by re-quoting the Mac Donald’s text in “Noise and Fluctuations” from 1962 that “It is probably fair comment to say that to many physicists the subject of fluctuations (or “noise” to put it bluntly) appears rather esoteric and, perhaps, pointless; spontaneous fluctuations seem nothing, but an unwanted evil, which only an unwise experimenter would encounter!” However, the low-frequency noise became prominently large, especially in small devices, and it “snakes in” in many cases as a limiting factor for applications, such as high resolution sensors, precise and stable oscillators and other; and considerable interest is given to the “slow” (as compared to the operating frequency of the devices) noise, with either 1/f power spectrum density (PSD) in frequency domain, or with bistable random telegraph signal (RTS) behavior in time. Therefore, in this work we focus on the achievements related to lowfrequency noise in electronic devices, mainly in transistors, in order to identify issues related to low-frequency noise in the aggressive device downscaling nowadays, and also to attempt giving an outlook for the evolution of the issues in the near future. Before we approach to the discussions in this work, we first briefly introduce the types of noise in respect to their spectrum. The main types of electronic noise are summarized in Table 1 in terms of current noise PSD. The thermal noise is due to random motion of charge carriers driven thermodynamically in the devices by uncorrelated scattering, and, therefore, it has uniform, or “white” spectrum, in analogy with the spectrum of the white light. The shot noise originates from the fact that the minimum charge of the carriers is the charge of the electron, q≈1.6×10 −19 C, and when many of these discrete charges overcome randomly an emission barrier and traverse the device quickly, then there are current “shots”, each of short time, nearly Dirac pulses, and accordingly, each of which having uniform spectrum. Since the “shots” are randomly occurring and uncorrelated, then the spectrum of the shot noise is also uniform “white” spectrum. The white noise, either thermal or shot noise, or both, is broadband, and it occurs from low frequencies to the maximum frequency at which the device can operate, by the assumption that the device is ideally uniform, and there is nothing else, except for charge carrier motion. In the real devices, however, there can be many other random processes, such as generation-recombination of carriers, charge trapping, phonon scattering, etc, normally with lower “speed” or “repetition rate”, which, therefore, cause increase in noise spectra at the low-
5 of 286 frequency end. If the random process has a characteristic time constant τ, or equivalently, a rate 1/τ, then the process is random at time scale t>τ, or equivalently, at frequency f<1/(2πτ), and the noise spectrum is uniform for these low frequencies. In contrary, at short time scales t<τ, or equivalently, at frequency f>1/(2πτ), the variability of the noise signal is less, e.g. an RTS “spends” some time in “on” or “off” state before doing a transition to the other state. Consequently, the power spectrum density decays as 1/f² at high frequencies f>1/(2πτ), and the overall spectrum of random process with characteristic time constant τ is the Lorentzian spectrum, as depicted in Table 1. The Lorentzian spectrum is found to originate usually to bistable processes, such as generation-recombination, but more precisely, the requirement for this spectrum is exponential decay in the autocorrelation function of the random process x(t), that is, x(t)x(t±Δt)dt∝exp(−Δt/τ), thus, the characteristic time constant τ is the correlation time in x(t). In the last row of Table 1, the so-called “flicker” noise is given. The power spectrum density of the flicker noise is inversely proportional to the frequency in the entire frequency range, with a slope normally close 1/f, and again by analogy to the light spectrum, it is also called “pink” noise. Several concepts for the origin of the flicker noise are suggested in the literature. One of the most popular concept is superposition of Lorentzian processes with time constants distributed as 1/τ, by assumption of particular distributions of traps in the device structure. Another approach is to “stretch” the above exponential function for the autocorrelation, say τ=τ o +Δτ, where Δτ is distributed in some way, e.g. normally. Other approaches are to define ½ differential operation, random rate perturbations that decay as square root of time, etc. Overall, a unique explanation for the 1/f noise is not available at present (and most probably, it will be never available after the many suggestions made so far), although the 1/f spectrum occurs in virtually any system, from electronics, through the level or river Nile, to biology, music and finances, and there are many reasonable models and consistent explanations for 1/f noise in particular cases. From above, the low-frequency noise in small devices is relatively increased, and explanations and models that predict this noise are available. Therefore, we review the models and their predictions extensively in this work, both numerically and by keeping the link to the physical assumptions behind the models. In order to derive a common point, we have also intentionally suppressed the controversy, which has accompanied the subject of low-frequency noise for the origin of 1/f noise, because we have observed that the predictions of the different models coincide, when the word is for numerical values and behaviors in respect to bias, temperature and device sizes; and all of the popular low-frequency noise models are somewhat mesoscopic, or better to say compact models, with the microscopic effects averaged after one or several mathematical integrations, which does not really allow to fully and precisely inspect the assumed microscopic physical origin of the fluctuations from the scattered data obtained after measurement of the noise. Indeed, after the analyses in section VIII, one may not always need to mandatory assume microscopic origin for the low-frequency noise. To approach to the review, we first provide in section II details and the generic models related to intrinsic and coupled noise in the forms, in which the low-frequency noise is lumped in compact models, that are widely used at present. Then, in section III, we review the state of the art low-frequency noise in bipolar junction transistors (BJTs), analyzing the different factors related to the issues with low-frequency noise, such as crossover between bulk, surface and barrier noise, fabrication, superposition of Lorentzian noise, and decomposition to individual noise components in small-area BJT, problems with averaging techniques, and also, we have provided some derivations that explain details in the bias behavior of the low-frequency noise in BJT, using the generic models
6 of 286 from the previous section II. Having the observations and results for BJTs, we have pursued in section IV a detailed review for the level and models of the low-frequency noise in MOS transistors, since these transistors are at the frontier of the device downscaling nowadays. From these, we found that the different models converge each to other, including for ranges (e.g. of gate oxide thickness and trap densities), at which the assumptions for the models are actually violated. Interestingly, we have observed, that the models for the noise in BJT also converge to the models for MOS transistors at some instances, such as by noise from interfacial oxides. Also, we showed that models based on superposition of distributed traps cannot discriminate clearly whether the noise is due to distribution of energy or depth of the trap in the gate oxide, since both are always together in the models. We have also provided some extensions in gaps found in the literature, such as for unequal amplitudes in RTS noise in nominally identical MOS transistors, relations between figures of merit for low and high frequency noise, and for frequency performance of the MOS transistors. Since the downscaling of MOS transistors required modifications in the classical MOS structure (one gate MOS with uniformly doped silicon body), as well as, germanium is used widely at present to improve the mobility in npn BJTs, the low-frequency noise in the modified SiGe heterostructural bipolar transistors (HBTs) and MOS transistors with multiple gates is addressed in section V. Overall, the observation is that the low-frequency noise increases, when mixing materials or adding new layers in the transistor structures, which is different from claims made last decade, and thus, it is a potential issue for the future. Nevertheless, forward body biasing and multiple gates, seem, improve the low-frequency noise performance of MOS transistors, but cannot compensate for the increase of the scattering in the noise levels in advanced silicon based transistors, as discussed in section VI, in which it is shown that the ultimate downscaling, e.g. in carbon nanotube devices, leads to stochastic behavior of the device, although, surprisingly, the data for these device still match with the general trend (1/area). The consequences of the increased low-frequency noise for the practice are discussed in section VII for two cases – the tradeoff between low noise and low power in low-power amplifiers and for phase noise in oscillators, providing some derivations that help to quickly estimate the noise performance of the amplifier from its consumption, and the phase noise of the oscillator from the harmonics and symmetry of the generated RF signal. Finally, in section VIII, we provide extrapolation of the results gathered during almost one century, to identify that the low-frequency noise in nanodevices can impact the reliability even of digital circuits. Analyses in this section showed that the variations during fabrication, called as “frozen noise”, also contribute to the temporal noise in the devices; that the values for noise parameters are corresponding to the range of the Heisenberg uncertainty for the free electrons in semiconductors; and that the purely statistical accumulation of variance with time can be a mechanism generically creating 1/f noise in the nature. This accumulation mechanism is known as “innovation variance”, and it seems is overlooked for electronic devices as potential source of background 1/f noise, while it provides a reason, which explains why the geometric averaging of scattered noise data should be preferable. Summarizing our review and analyses in section IX, we conclude that the low-frequency noise in electronic devices follows consistently the phenomenological law (1/area) in relative units for the noise power in ratio to DC power, which basically sets a barrier for downscaling of deterministically behaving devices at sizes below about 10×10nm², which is nothing, but the size range of the viruses – the smallest structures, in which the nature was able to embed reproducibility and functionality over a period in the range of about 10 9 years. The different
7 of 286 models for low-frequency noise appear to coincide in their predictions for noise magnitudes, which actually implies that the statistics is more important for the noise than the physical phenomena through which it causes noise in the devices. Thus, the numbers might be more important and a change in the coordinate system for noise perhaps will be helpful to describe the low-frequency noise even better than the descriptions are at present, without arguing for and against that strongly, as it happened in the near past whether the 1/f noise is due to mobility or number fluctuation. It is due to both, and in addition, perhaps due to other fluctuations, which stay hidden from us at present. It might be reasonable to state that we can observe only a “microscopic” window, if not smaller, by probing the variations as 1/f noise in electronic devices, which accumulate variations from the continuously changing matter. Therefore, we have the aphorism above, behind which is the unity of continuity and variation, and with the information available to us, let our discussion begin “flowing” with the first topic on intrinsic and coupled noise in device currents, for which, it seems, we have agreement in principle at present. II. Intrinsic (uncorrelated, uncoupled) and coupled (correlated) behavior in a fluctuating system (behind and beyond ∆µ – ∆n controversy) In this work, we use a conceptual approach of “intrinsic and coupled noise” that helps to identify the noise in electronic devices. The approach is based on the general assumption that the fluctuations in one quantity can originate intrinsically from its nature, or the fluctuations can be induced, thus, coupled from fluctuations of another quantity. II.1. Coupled behavior For the noise S I in a quantity I, when S I is coupled to another fluctuation S V , one has V 2 DC 2 DC I S I g K I S = , (1) where K is a parameter (usually taken K=1, if no suppression or enhancement in S I due to correlation to other process exists), I DC is the average value of I, and g=∂I/∂V is the coupling coefficient between I and V, by assuming also that I and V are immediately, instantly and fully correlated each to other. Obviously for an electronic device (e.g. resistor, diode or transistor), I, V and g are electrical current, voltage and conductance, respectively. So, the units for the power spectrum densities (PSDs) become A 2 /Hz for S I and V 2 /Hz for S V . We illustrate eq. (1) for the voltage noise in bipolar junction transistors (BJTs) and MOS transistors, using the data from a past report of the International Technology Roadmap for Semiconductors (ITRS 2006) [3], which extended its predictions until 2020. ITRS provides values for the input referred voltage S V@1Hz at 1Hz normalized (multiplied) with device active area A of 1µm 2 for npn BJTs and nMOS transistors, as shown in Figure 1a. Let us express S I /I DC2 with the simplest SPICE model for 1/f noise, and use the transconductance of the transistors g m in eq. (1), also taking K=1 and multiplying the equation with the device area A. That is ( ) ( ) f SA I g f KA I S A Hz1@V 2 DC mF 2 DC I × = × = . (2) Here, the parameter K F is a measure for the ratio noise/DC in the output current of the transistors. The ratio g m /I C ≈1/ϕ t ≈38.5 V -1 in BJT, where ϕ t =kT/q≈0.026V is the thermal voltage at room temperature T=300K. The
8 of 286 ratio g m /I D varies with the bias of the MOS transistor, but at low gate overdrive of (V G −V T )=0.1V, g m /I D ~13 V -1 . So, we write ( ) ( ) ( ) ( ) ( ) =−× ×≈ϕ× ≈× 0.1VVVat MOSfor ,13SAHz1 BJTfor ,5.38SAHz1SAf KA TG 2 Hz1@V 2 Hz1@V 2 tf@V F . (3) The results for (A×K F ) calculated by eq. (3) are shown in Figure 1b, when the values for (A×S V ) from Figure 1a are used. Interestingly, the input referred voltage noise in BJT is about 2 orders of magnitude less than that in MOS transistors, but the difference is smaller in the output current noise; and the difference varies further at higher gate overdrive of MOS transistor, because g m /I D ∝1/(V G −V T ). Now, we discuss eq. (1) in more details. Assume that the noise is caused by a fluctuation S Q of trapping charges, and these charges change the voltage V on a capacitance C. Then, the coupling Q=CV is given by the capacitance C, and 2 Q 2 DC V 2 DC 2 DC I C S I g KS I g K I S = = (4) Consequently, the charge Q=qN is coupled to the number of charges N by the elementary charge of the electron q (1.6×10 -19 C), and one writes the general form of the equation for the so-called “number fluctuation” (Δn) N 2 2 DC 2 Q 2 DC V 2 DC 2 DC I S C q I g K C S I g KS I g K I S = = = , (5) which is widely used to study the effect of charge trapping in interfaces on the current in a device, e.g. charge trapping in the gate insulator of a MOS transistor and its drain current. This equation is for number fluctuation, since it assumes that the fluctuation S N in the number of trapped charges causes the noise in the current I. Obviously, the unit is Hz -1 for the power spectrum density S N , and S I is also PSD in units A 2 /Hz. The different Δn-models for different devices (at particular set of physically based assumptions) derive different expressions for S N and coupling parameters (C, g, K), which can be physics-, bias-, processand designdependent. Overall, all complex derivations based on the assumption for charge trapping end with a form similar to eq. (5). These will be presented along with the discussions on the specific devices. However, one important note should be made. The fluctuation in the number N of trapped charges is assumed in the origin of the noise in Δn-models, and this fluctuation is indirectly transferred to a fluctuation in the number n of charge carriers that provide the current I in the device, by electrostatic (Coulomb) balancing the charge using the Gauss law. So, one can mistakenly assume from the expression for drift current with a density J=qnμE that the Δn-model is for the number n of charge carriers, by neglecting the variations in charge carrier mobility μ and in the electric field E, caused by the trapped (and thus immobile) charge. For example, the trapped charge can cause change in the mobility, so that dg∝μ∂n/∂N+n∂μ/∂N≠μ∂n/∂N and the estimate for g obtained from electrostatic balance ∂n/∂N=1 and μ=constant will be inaccurate. To resolve these problems, a correlated (to the number of charges) mobility model (Δn-Δμ) is used widely in nMOS transistors, and a scattering parameter is introduced. Furthermore, both (Δn-Δμ) and (Δn) models assume that the current flow is continuous and free of inherent fluctuations. Certainly, this is a good approximation when analyzing the average I DC , but it is not exactly true,
15 of 286 It is interesting to observe in eq. (21) that the prefactor 38 µm -1 =1/(26nm) ±3.4dB for the trend in Figure 4 corresponds to a diameter 2πr th , which has been regarded as the effective diameter, in which the electrostatic band bending due to a single excess (trapped) electron charge disturbs the silicon properties [46]. At room temperature, r th ~5nm and corresponds to the distance in which the electrostatic band bending due to a single charge is equal kT=26meV. Also, with ε being permittivity, a common term is q 2 /(A E ε/t IFO ) 2 in the expressions in all models for the noise from IFO, which implies that the charge capture at IFO causes coupled noise in the BJT current, according to eq. (5). The differences between the models come from the different assumptions for the actual physical mechanism, which is behind the charge fluctuation S N in eq. (5). These models lead to a variety of figures of merit for parameters related to the noise from IFO, e.g. tanδ~0.3-3 for direct tunneling [10]. According to [34], S N ∝ λN t /C mo2 or S N ∝D it /C mo2 for two step tunneling and random walk models, respectively, where C mo is surface capacitance of IFO (may differ significantly from ε SiO2 /t IFO ), and the figures of merit are N t /C mo2 ~5×10 29 cmF -2 eV -1 and D it /C mo2 ~10 22 cm 2 F -2 eV -1 . These values for tanδ, N t /C mo2 and D it /C mo2 are solely introduced to fit data from noise measurements, and they are not examined from other experiments, to the best of our knowledge. (iv) Non-uniform IFO model. There is some uncertainty in the models for the noise from IFO. The models predict power-law dependence between noise level and t IFO , and do not suggest exponential dependence, observed in some of the experiments, especially from the samples with thicker IFO. An approach to implement exponential dependences was taken in [20] for samples of different widths W of the emitter. The results are shown with solid squares in Figure 2 and Figure 4 as function of emitter area and average IFO thickness, respectively. It is suggested in [20] that t IFO =0.55nm in the middle of the emitter for samples with wide emitter, while IFO is ∆t IFO =0.25nm thicker at the periphery of the emitter. Then, the IFO thickness along the emitter width is assumed to decrease from t IFO +∆t IFO at the edge of the emitter toward t IFO in the middle of the emitter, by an empirically assumed exponential function, given by ( ) ( ) oIFOIFOIFO w/wexpttwt −∆+= , w≤W/2, (24) where w is the distance from the edge (w=0) toward the middle of the emitter (w≤W/2), and w o ≈80nm is a characteristic width of the thicker peripheral IFO. Schematically, the non-uniform IFO layer in the emitter opening is illustrated in Figure 5. (For convenience during integrations, the authors of [20] have chosen hyperbolic functions instead of exponential function, but they have mentioned that the choice is not unique, and other function can be used.) Then, as explained in [20] in details, the effective recombination rate is RR={a 1 +a 2 ×exp[a 3 ×t IFO (w)]} -1 , with a 1 , a 2 and a 3 being constants; and the current densities of the DC and noise currents are obtained by integration along the width of the emitter, using J(w)∝RR(w) and j(w)∝t IFO (w) 3 ×J(w) 2 , respectively. The above model for noise from non-uniform IFO is shown in [20] to have a good agreement with the experimental data. We also observe that the data are very close to the trend in Figure 4, when re-calculating the average thickness of IFO from the information in [20]. However, despite that exponential functions of t IFO are used in the above model for several quantities, the overall dependence of noise level vs. average t IFO remains a power law and we note that the exponential dependence of noise on t IFO , although observed in [14, 33, 34] experimentally, is not explained theoretically yet. Nevertheless, the approach in [20] provides a way to analyze noise from non-uniform IFO in BJT, and can help to identify different contributions to the total low-frequency
16 of 286 noise in deep sub-micron BJT, in which a crossover between several noise sources is obvious; and this will be discussed next. III.3. Crossover between different noise sources in BJT Another reason for the scattering in the data for the normalized noise in Figure 2 is that several noise sources in BJT can contribute in different proportions. The proportions between the contributions can vary with the biasing level, and with the layout and fabrication of BJT. Note that the relation S I /I DC2 ∝1/A in eq. (14) is valid only if strictly one noise source dominates, which is usual case for a particular device at particular bias, but not always the case, when the bias, layout, fabrication and interconnection vary, especially for sub-micron area devices. For example, the generation-recombination (GR) currents dominate the base current at low bias, while the injection through the emitter junction of BJT is low, but the noise due to GR currents modulates the injection barrier, and the GR noise is coupled to the collector current by the transconductance, resulting in a quadratic (or nearly) dependence S IC ∝I C2 . However, at higher biasing, the injection current dominates, and the intrinsic (Hooge) noise in the emitter resistance or injection current takes over the GR noise, resulting in a linear (or nearly) dependence S IC ∝I C and eventually a decrease of normalized noise S norm =S IC /I C2 can be observed at high bias of BJT. Such cross-over between the noise sources in BJT is reported in [21], for example, when the base current was varied over 2-3 decades, and the cross-over causes a bias-dependent variation for K F , as illustrated earlier in Figure 2 with the solid diamonds labeled as “S IB ∝I B1.2 ”. A portion of these data is presented later in Figure 6 as function of the base DC current. One explanation for such cross-over could be that the increased carrier density screens the coupling from trapped charges, but this is not elaborated for BJT, to the best of our knowledge, although explored for MOS transistors [47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59]. We will show that the evolution of the peripheral noise, as uncoupled to the emitter area in BJT, can lead to reduction of the slope in the bias dependence of the noise, while the noise from IFO has opposite evolution. Typical types of cross-over in the noise in BJT are between areal and peripheral noise sources in the emitter region [14, 21, 35, 36, 42, 60], noise from tunneling through IFO and series resistances (carrier number fluctuation, or coupled noise to the injection process, precisely) and diffusion noise (mobility fluctuation, or intrinsic noise of the current flow in emitter and base regions) [10, 19, 21, 32, 42, 61, 62], surface and bulk noise sources [14, 21, 24, 36, 37, 42, 60], and for very small-area BJTs (A E ≤0.1µm 2 ), the extension length of the base to the contact via becomes an issue for the low-frequency noise [36, 37], since that extension is with large area compared to A E , and it is vulnerable to the surface noise from the oxide on the top of the structure. Also, in small-area BJTs, the crossover between generation-recombination (GR), random telegraph signal (RTS) and 1/f noise becomes apparent [11, 12, 17, 18, 30, 35, 63]. The identification of noise sources uses many techniques, such as noise partitioning (or decomposition the total noise in several components), superposition of noise components, evolution of noise components and correlation between them with bias, area and perimeter of the emitter, fitting to physical models and equivalent circuits. The variety of techniques is large and a fully systematic approach in reviewing these is not possible. Nevertheless, there are several useful relations that are accumulated during the years. These are reviewed in [31] and some of them are also discussed below. III.3.1. Intrinsic noise for the current flow The intrinsic noise for the current flow is assumed to be due to mobility fluctuation, which follows the Hooge eq. (6) for the carriers in the emitter and base along the direction of the current flow in BJT, while the concentration
17 of 286 of carriers varies, as given by the theory for operation of BJT. Then, since the BJT theory is based on diffusion, the fluctuation in the mobility is transferred to fluctuation in the diffusion coefficient D m for the minority carriers in the emitter and base regions outside the depletion region of the base-emitter junction, using the Einstein relation D=µφ t . The resulting expressions for the intrinsic noise of the current flow in the base and collector currents are in the form of I f S and I f S C ' C IB ' B I CB α = α = , for the base and collector current noise, respectively, (25) where the quantities α’ are in unit Ampere and given by [62] mm m m m H 2 m D d 1 )d(n )0(n with , )d(n )0(n ln d qD '×ν +≈ α=α , (26) Respectively for S IB and S IC , d=d E or d=d B are the thicknesses of the emitter and base regions (along the coordinate of the current flow), D m corresponds to the diffusion coefficients of the minority carriers in the emitter and base, α H corresponds to Hooge parameter for minority carriers in the emitter and base, n m (0) are minority carrier concentrations at the emitter-base junction in the emitter and base sides, n m (d) are minority carrier concentrations at the other sides of the emitter and base (metal contact for the emitter and collector-base junction for the base), and ν are recombination velocity (~10 5 cm/s) of minority carriers at the emitter metal contact or carrier saturation velocity (~10 7 cm/s) that the minority carriers in the base usually reach at the basecollector junction. Slightly different expressions for α’ can be found in [61]. Eq. (25) shows that the intrinsic noise for the current flow (due to mobility or diffusion fluctuation) in the base and collector currents is a linear function of the DC currents, and it corresponds the right-hand term of eq. (11). This linear dependence is usually used to identify the diffusion noise in BJT [19, 32, 42] after splitting the total noise into linear and quadratic functions of the DC current, because other noise sources result in non-linear dependence between noise and DC currents. For example, S I ∝I DC2 for recombination noise at the base surface or at the surface of the emitter-base space-charge region [42]. The expressions for the diffusion S IB are modified in [10], when IFO is present in the emitter, but the linear dependence remains. Worth mentioning, a linear dependence between noise power spectrum density and DC current in BJT is rarely observed and we will show later that such dependence can be derived from peripheral noise uncoupled to the emitter area. III.3.2. Base and collector currents are strongly correlated It is well established that the low-frequency noise in the base and collector currents are strongly correlated. Their normalized cross-correlation spectrum, called often coherence (and corresponding to correlation coefficient for random quantities), is close to one [19, 32, 37, 60]. The coherence between two noise spectra is given by ( ) yx 2 xy yx SS S s,sCoherence = , (27) where S x and S y are the individual power spectrum densities (PSDs) of the two noise spectra s x and s y , and |S xy |=|s x s y* | is the magnitude of the cross-correlation PSD of these noise spectra s x and s y . If the coherence is close to 1, then s x and s y are strongly correlated and most probably originate from the same noise source, since the different noise sources are independent each from other, either by assumption, or because they contribute in
18 of 286 different proportion at different terminals in BJT. Simply, a coherence=1 for base and collector noise currents means that S IB and S IC are coupled in BJT; and we can use eq. (11) to investigate the coupling. The reason to address the coupling is that in the literature the correlation between base and collector currents is usually analyzed in terms of feedback on the emitter resistance r e (see Figure 3a) and current gain β, while both the feedback and β are secondary effects of the primary diffusion process that explains the operation of BJT. The diffusion process depends on the diffusion properties of the base-emitter junction and the voltage applied on this junction. For example, generation-recombination (GR) currents in base current do not affect the DC collector current, if r e =0, as seen from Gummel plots for DC currents in BJT, whereas the GR noise in the base current occurs in the collector current well correlated. So, in a simplified form for diffusion current in the base of BJT, we can write ϕ = t BE 0BDIF,B V expII , (28) where I B0 ∝D m and exp(V BE /φ t )∝n m (0), as defined after eq. (26). Assume that D m ∝ µ fluctuates as given by eq. (25), V BE is coupled to this fluctuation, but other parameters, e.g. φ t , ν, d in eq. (26), do not fluctuate. By taking the logarithm of eq. (28), one writes ( ) ( ) ( ) ( ) ( ) t BE m m t BE 0B 0B DIF,B DIF,B V D DV I I I I ϕ ∂ + ∂ = ϕ ∂ + ∂ = ∂ . (29) Provided that the square of the differentials corresponds to power spectrum densities and the variation of V BE is coupled to the variation of the base current (correlation coefficient = ±1), using eq. (25) we get ( ) DIF,B ' B 2 DIF,B I 2 t IV I 1 f I S S DIFF,B DIF,BBE α == ϕ , (30) where S VBE | IB,DIFF is the noise in V BE coupled from the diffusion noise in I B . Since the collector current is also coupled to V BE by a relation similar to eq. (28), then S VBE | IB,DIFF will add a noise component in the collector current I C , which is coupled to the diffusion noise in the base current, and the total noise in I C is given by ( ) C ' C 2 t V 2 C I 2 t IV 2 C I I 1 f S I S S I S BE DIFF,C DIF,BBE C α + ϕ =+ ϕ = = coupled + intrinsic, (31) where S VBE = S VBE | IB,DIFF + S VBE | IFO +… is the total noise in V BE coupled from different noise sources, i.e. I D,DIF , IFO and other, and S IC,DIF is the intrinsic noise in I C due to mobility fluctuation in the base of the transistor, as given by eq. (25). Since the concentration of minority carriers m n depends on V BE , then S VBE can be regarded as a coupled number fluctuation in BJT. The above discussion is for unilateral case for noise propagation from input to output, for which the noise sources associated to the base and emitter of BJT couple noise into the collector current, and the intrinsic noise in the collector current is not coupled back to the base. This a reasonable approximation, because of two reasons (at least). First, the contribution of the intrinsic noise S IC,DIF is usually negligible, and second, the backward isolation from collector to base and emitter is high as compared to the forward transmission, conservatively in the bandwidth of the low-frequency noise. Also, the first example below
19 of 286 demonstrates that the diffusion noise in I B and I C might not be discriminated each from other in experiments and cumulative experience suggests that the noise in BJT can be successfully referred to the input as a noise in the base current. Now, let us see how eq. (31) helps to deal with noise partitioning and superposition in BJT. Example. Diffusion dominates DC and noise currents The first example is by an assumption that the diffusion dominates both the DC and noise currents in BJT and the current gain β=I C /I B =constant. Using eqs. (30) and (31) for the noise in I C we write C ' C B ' B C ' C 2 t IV 2 C I I 1 fI 1 fI 1 f S I S DIF,BBE C α + α = α + ϕ = . (32) Similarly, including the coupled noise from I C into I B , for the noise in I B we get B 'B C ' C B ' B 2 t IV 2 B I I 1 fI 1 fI 1 f S I S DIF,CBE B α + α = α + ϕ = , (33) which is the same as eq. (32) for the noise in I C . Since β=constant, then the ratio of the noise in I C and I B is 2 C ' BB ' C B ' CC ' B 2 B ' B B ' C C ' CC 'B B ' B B C B ' C C ' C C B C ' B I I II II I I II I f I I I f I f I I I f S S B C β≡ α+α α+α β= α+ β α α+βα = α + α α + α = =constant, (34) and from DC and noise measurements, one cannot separate the parameters α’ for the diffusion noise in I B and I C . One can estimate approximate values from eq. (26), but the final values for the Hooge parameter α H will remain in an arbitrary ratio. A good guess for initial values of α H can be obtained by attributing the noise first only to I B and then only to I C , and then check which number makes sense, but this is a guess – the measurement cannot reliably separate the two parameters from one device, in which the diffusion dominates both the DC and noise currents; and it is reasonable to estimate only one value of α H for the dominant noise contribution. The approach of finding the noise source with dominant contribution is presented in [19] in convenient form. This approach is based on choosing an appropriate equivalent circuit of the device (and measurement setup), with several, i.e. 3, possible noise sources. Then, the measured noise is attributed to each noise source, and one of them usually fits the data well, which is consistent with the assumption that one noise source practically dominates in the total noise. The approach of the dominant noise source is used several times and it identified that the dominant noise source can be presented as current noise source in the base current [19, 32, 37, 61], although the actual physical process can be in the emitter, e.g. IFO, or at the surface above base-emitter junction. Nevertheless, once the dominant noise source is known, the physical identification of the noise is much easier and focused on finding which physical model can describe the behavior of dominant noise source. This is discussed below with the second example for noise partitioning based on eq. (31) for the coupled and intrinsic noise in BJT. Example. Noise partitioning. Non-linear dependence of current noise on diffusion noise The second example for the use of eq. (31) is for cases when the current noise in BJT has non-linear bias
20 of 286 dependence with a slope different from one of the diffusion noise, and the slope changes with the bias current too, indicating cross-over between different physical origins for the noise. According to the discussions above for dominant current noise in the base of BJT, first, we rewrite eq. (31) with the term for the intrinsic noise in I C omitted, but replaced with the diffusion noise in the base current. That is 2 C DIF,B ' B 2 C 2 t V I I I 1 f I S S BE C α + ϕ = . (35) Here, S VBE denotes all noise sources coupled to V BE , except for the diffusion noise in the base current, which is explicitly shown. Also, we assume that I C is a diffusion current, while I B =I B,DIF +I B,GR consists of diffusion current I B,DIF , but may have generation-recombination DC component I B,GR due to either bulk or surface recombination. Second, we refer the noise to the base as the current noise, using AC and DC current gain, β AC =∂I C /∂I B and β DC =I C /I B , respectively. The equation for the base-referred current noise is ( ) ( ) 2 AC DC GR,BDIF,B DIF,B GR,B ' B 2 GR,BDIF,B 2 t V 2 AC I I II I I 1 f II S S S BE C B β β + + α ++ ϕ = β = . (36) Owing to the presence of I B,GR , both β AC and β DC vary with bias, even if one assumes that the ratio β=I C /I B,DIF between the diffusion currents is constant (valid assumption until high bias is applied, that causes high level of injection of majority carriers in the base, and consequently, D m and β decrease. The high-injection regime is usually not used in low-frequency noise measurement, but it can be reached in submicron-area BJT – see [35]). One can observe many possible cases that follow from eq. (36), since the evolution of S IB with I B is dependent on several ratios, e.g. β DC /β AC , I B,GR /I B,DIF , which are bias dependent. Let us look at several of these cases. Case 1 in second example: The diffusion is dominant in the DC currents, that is I B,GR <<I B,DIF =I B and β DC =β AC =β=constant. In this case, eq. (36) reduces to ( ) ( ) B ' B 2 B 2 t V I I f I S S BE B α + ϕ = , if diffusion dominates DC (37) One situation is that no noise is coupled to the diffusion process in BJT. In this situation, S IB ∝I B , and this is the situation of dominant diffusion (intrinsic) noise discussed above – see eq. (33). A second situation is when the coupled noise has a constant value as a voltage noise, S VBE =constant, and it is dominant. In the second situation, S IB ∝(I B ) 2 . Constant S VBE is following from bias-independent fluctuations that can be regarded as resistance fluctuations, e.g. noise form IFO in the emitter [10], or charge trapping and recombination that modulate potential of the base, emitter or emitter-base junction [42]. The normalized noise of these fluctuations is nearly constant [42], that is S I /I DC2 =constant. Let us take the noise form IFO. It is shown in [10] that the noise of IFO is a fluctuation in the resistance of the emitter region (see r e in Figure 3a). f " S I S I S r S r S IFO 2 t IFOV 2 B I 2 E I 2 e r 2 IFO r BEBE eIFO α = ϕ ===≈ = constant in respect to bias, (38) as far as β=constant is assumed for the case when the diffusion dominates the DC currents. Here, we have recalled the left-hand equality of eq. (30), which describes the coupling between voltage and current noise on a
21 of 286 pn junction. So, adding the coupled noise from IFO in eq. (37), one gets ( ) ( ) B ' B 2 B IFO I I f I f " S B α + α = , if diffusion dominates DC. (39) There is a critical value I Bcr =α’ B /α” IFO in eq. (39) for crossover between diffusion noise and noise from IFO. If I B <I Bcr , then S IB ∝I B because the diffusion noise dominates. In contrary, if I B >I Bcr , then S IB ∝(I B ) 2 , because the noise from IFO dominates. The crossover is observed for the large-area samples in [10], as shown in Figure 6 with squares, and similar data are reported in [19]. Based on the different slopes of S IB as function of I B for diffusion noise and noise form IFO, these noise sources are separated and accordingly characterized [10, 19, 32], respectively, in terms of Hooge parameter α H for the diffusion noise, since α’ B ∝α H in eq. (26), and in terms of K F =f×S IB /I B2 =α” IFO for the noise from IFO, as follows from eq. (38). The data suggest cumulatively that the diffusion noise is important for BJT with larger emitter area and thin IFO, while the noise from IFO dominates BJTs with emitter area less than 10µm 2 and t IFO >0.5nm. The nowadays BJTs usually are in the latter category. Case 2 in second example (peripheral-areal noise): The diffusion is not always dominant in the DC currents. For example, at low bias, a generation-recombination (GR) process can contribute to I B a component, given by ϕη = tGR BE 0GRGR,B V expII , (40) where η GR is non-ideality factor with values usually between 1.5 and 2. Following the procedure after eq. (28), the GR process couples voltage noise in the V BE , given by f I S S " GR 2GR,B I 2 GR 2 t IV GR GR,BBE α =η= ϕ ≈ constant in respect to bias, (41) since S IGR /(I GR ) 2 ≈constant for charge trapping [42], as mentioned above. In the presence of I B,GR , the ratio β DC /β AC is ( ) ( ) η + + = ∂×β +∂ + ×β = +∂ ∂ + = β β DIF,BGR GR,B GR,BDIF,B DIF,B DIF,B GR,BDIF,B GR,BDIF,B DIF,B GR,BDIF,B C GR,BDIF,B C AC DC I I 1 II I I II II I II I II I (42) ( ) << η >> = +η +η = β β GR,BDIF,B GR GR,BDIF,B GR,BDIF,BGR GR,BDIF,BGR AC DC II if , 1 II if ,1 II II (43) From eq. (36), in the case of existing GR component in I B , the base current noise becomes 2 GR GR,B DIF,B DIF,B ' B " GR I I I I 1 ff S B η + α + α = (44) At high bias I B ≈I B,DIF >>I B,GR , and one observes S IB ∝(I B ) 2 , since also I B,DIF >I Bcr =α’ B /α” GR . At low bias, however,
22 of 286 when I B,DIF << I B,GR ≈I B , several situations are possible. If I B,DIF >I Bcr =α’ B /α” GR , then S IB ∝(I B ) 2 , but slightly attenuated with (η GR ) 2 , and one may observe small step in the S IB ∝(I B ) 2 dependence at the crossover between I B ≈I B,GR and I B ≈I B,DIF . In a situation when the coupled noise is low and I Bcr is larger than the DC current for crossover between I B ≈I B,GR and I B ≈I B,DIF , one can find a dependence that differs from S IB ∝(I B ) 2 and S IB ∝(I B ). For this situation, we rewrite eq. (44) with α” GR omitted, as η + η + α = GR GR,B DIF,B DIF,BGR GR,B ' B I I I I I 1 f S B , for diffusion (intrinsic) noise. (45) At the higher bias end, where I B,GR < I B,DIF ≈I B , then S IB ∝(I B ). In contrary, at the lower bias end, where η GR I B,DIF < I B,GR ≈I B , then =η∝ =η = − ηϕ η α = η α =1 @ IS to 2 @constant from 1 2 V exp I I 1 f I I f S GRBI GR GRt BE 0B 20GR 2 GR ' B DIF,B 2 GR 2GR,B ' B I B B . (46) This equation implies that the intrinsic diffusion noise in BJT levels off when the base DC current is dominated by generation-recombination currents. This effect is observed several times when the BJT was stressed electrically or after proton irradiation, and the generation-recombination currents are increased. An example from [21] is shown in Figure 6. Before stress, circles in Figure 6, a slope close S IB ∝(I B ) is observed, indicating low level of coupled noise owing to screening of surface noise from emitter-base junction, achieved by a high doping concentration at the surface of the base, using superficial base doping (SBD). Also, the step in the S IB ∝(I B ) dependence is observed in this sample at the transition region from low I B ≈I B,GR to higher I B ≈I B,DIF . After the stress, triangles in Figure 6, the S IB ∝(I B ) dependence levels off due to increased I B,GR . The departure from S IB ∝(I B ) and S IB ∝(I B ) 2 is attributed to generation-recombination currents at the surface above the base-emitter junction [14, 36, 37, 42, 60], because it is found that apart from 1/A E dependence for the normalized noise (K F =f ×S IB /I B2 ∝1/A E ) in the base current, both the DC and noise currents in the base scale with the perimeter P E of the emitter [11, 14, 36, 37, 60], rather than only with the emitter area A E . So, partitioning between areal and peripheral noise in BJT takes place, by using equations in the form of 2GR,B P,F 2DIF,B A,F 2 P,B P,F 2 A,B A,F I I f K I f K I f K I f K S E E E E B +≈+= , (47) where K F,A (≡K F,AE ) is the normalized areal noise, with K F,A ∝1/A E , and K F,P (≡K F,PE ) is the normalized peripheral noise, with K F,P ∝1/P E . The peripheral noise is usually analyzed as a surface noise attributed to an area A P =P E ×W P at the spacer oxide [36, 37, 42], where the characteristic width W P is taken as the width of the depletion region of base-emitter junction [37]. The fluctuation of charge trapping, or tunneling associated with the oxide at this surface A P , is regarded as the origin for GR current and corresponding surface noise. This illustrated in Figure 7. The fluctuation of charge at/in the oxide modulates the potential in the base-emitter junction in vicinity of the trapped charge causing fluctuation in the carrier concentrations in this vicinity, and thus, current noise in BJT due to number fluctuation of carriers. The assumption that the traps affect the depletion region of width W P of base-emitter junction is reasonable, but the value of W P ~15-20nm in [37] is
23 of 286 small as compared to values 50-100nm for peripheral effects. Our estimate [46] is that the potential bending around a single trapped charge results in a strong and almost constant effect in a distance of about 15nm, and the effect gradually decreases in a distance up to 50nm. This is shown with dashed lines in Figure 7. So, a “remote” coupling of the fluctuation of trapped charges is expected, and the width, as well as the depth, of coupling of surface noise is normally in the range of 50-80nm [20, 37]. Nevertheless, this range of distances is smaller than the width of the emitter, and it implies a peripheral noise that scales with the perimeter of the emitter, rather than with the area of the emitter. The surface noise can couple noise in the resistance of the base too, since the length of base extension to the metal contact can be large as compared to the emitter width, but this effect is not dominant noise source in BJT, because the doping of the base and the extension is high, and they are not depleted from carriers in principle. Therefore, outside the vicinity of depletion region of the base-emitter junction, the large number of carriers screens the fluctuation of trapped charges, although some additional noise is observed in minimum-sized BJTs [37]. To separate the peripheral noise from areal noise in BJT during experiments, one writes eq. (47) in two equivalent forms, given by ( ) ( ) 2DIF,B 2GR,B E E P,FEA,FE 2DIF,B IE I I P A KPKA I SA f B ×+×= × , ≈ areal noise (A E ×K F,A ) at high I B , (48) ( ) ( ) P,FE 2GR,B 2DIF,B E E A,FE 2GR,B IE KP I I A P KA I SP f B ×+×= × , ≈ peripheral noise (P E ×K F,P ) at low I B . (49) Then, the DC currents and low-frequency noise of samples with known A E and P E and different ratio A E /P E are measured in several decades for I B . A check for quadratic dependence S IB ∝(I B ) 2 is helpful to verify that the noise is coupled (from IFO and surface) and not intrinsic (from diffusion) in order to guarantee the validity of eq. (47). Next, the diffusion and surface DC currents are separated from Gummel plots for each sample in a manner so that I B,DIF ∝A E and I B,GR ∝P E in order to verify that the assumptions for scaling with emitter area and perimeter are valid in the experiment, e.g. I B,GR might not be a peripheral current in BJT. Finally, the quantities in the left-hand sides of eqs. (48) and (49) are calculated and plotted against the quantity [A E ×(I B,GR ) 2 ]/[P E ×(I B,DIF ) 2 ] and its reciprocal, respectively. If the assumptions for areal and peripheral noise apply for the samples, the intersection of the linear fit in the first plot with the vertical axis should match with the slope of the linear fit in the other plot (and vice versa). If so, then one can use (A E ×K F,A ) as the measure for the areal noise in BJT and (P E ×K F,P ) as the measure for the peripheral noise; and proceed to physical identification of noise sources. Certainly, the procedure for separation of areal and peripheral noise is quite demanding, and we observe only portions of this procedure reported in the literature [11, 36, 37, 60]. Also, some peripheral effects in BJT, such as thicker IFO at emitter edges [20, 36] and current crowding at emitter periphery [35], may result in higher noise, but not necessary cause higher base DC current I B,GR . In these cases, the method above may result in unrealistic values from physical point of view [37] and it has to be modified to converge the approaches in [36] to that in [31], the latter discussed earlier by the help of eq. (24) for non-uniform IFO. The modifications are the following. From DC characteristics of small and large area BJT with different ratio A E /P E , one has to determine the characteristic width W P of the peripheral regions and the areal and peripheral DC current density pre-factors J 0A and J 0P , respectively, using the equation
24 of 286 ( ) ϕη + ϕη + ϕη −= tGR BE GR0PE tP BE P0PE tA BE A0PEEB V expJWP V expJWP V expJWPAI , (50) so that one set of values for W P , J 0A , J 0P , J 0GR , η A ≈η P ≈1, and η GR ≈2 fits the measured curves. An initial value for W P can be between 50nm [36] and 80nm [31]. The last term in the equation is for surface currents, and at high bias can be neglected, where also an approximation η=η A =η P ≈1 holds, and the equation can be reduced to P0P E E A0P E E t BE E B JW A P JW A P 1 V exp A I+ −= ηϕ − . (51) The right-hand side of this equation reduces to J OA for large-area, nearly square-shaped BJTs, and thus, J OA is easy to find. With the initial value for W P , one can find initial value for J OP , using data from smaller-area, rectangular-shaped transistors; but the next step will require optimization procedure, first, to fit the data from DC measurements at high bias, varying W P and J OP in eq. (51), and then, another optimization procedure to fit all data, both at high and low bias, using eq. (50). This is not convenient. Therefore, another approach is taken in [36]. It is estimated that P E ×W P /A E <<1, which reduces eq. (51) to P0P E E A0 t BE E B JW A P J V exp A I+= ηϕ − . (52) The quantity in the left-hand of the equation is the current density prefactor J B0 for the base current in the samples, and J B0 is obtained for each sample from DC measurements. The values of J B0 are plotted against the ratio P E /A E , and from the slope of this plot, the values for the product W P ×J OP are obtained and directly used, since, actually, W P ×J OP is needed in eq. (50), in order to analyze the noise form peripheral IFO; and the GR currents can be easily removed from the measured data too. Having W P ×J OP estimated, one can obtain the peripheral I P and areal I A emitter currents by ϕ ≈ ϕη = t BE P0PE tP BE P0PEP V expJWP V expJWPI , since η=η A =η P ≈1 is assumed, (53) and from measured base current I B one can get the areal emitter current I A , using P B A III −= ( –I B,GR , if necessary). (54) Care should be taken so that I A is greater than zero. By making substitutions I P ↔I B,GR and I A ↔I B,DIF , eqs. (48) and (49) can be rewritten as ( ) ( ) 2 A 2 P E E P,FEA,FE 2 A IE I I P A KPKA I SA f B ×+×= × , (55) ( ) ( ) P,FE 2 P 2 A E E A,FE 2 P IE KP I I A P KA I SP f B ×+×= × , (56) by also neglecting the noise from surface currents. At this point, all modifications for the noise partitioning between areal and peripheral noise from IFO are made, and the procedure continues as described above for surface noise. One note should be made. We have assumed that P E ×W P /A E <<1, which might be not precise, if
31 of 286 this current is possible, resulting in decreased sensitivity of the setup in some cases, especially at high DC currents (>1mA). Overall, even after careful preparation of the measurement setup, the error associated with the preamplifier and its front-end is in the range 3%-5%, which adds another 0.4dB uncertainty for the normalized noise. The error from the spectrum analyzers is usually low, within 1%, but it requires that the measurement range is chosen accordingly. However, in some cases such as RTS noise with small or high duty cycle and occasional spikes, the range has to be increased to avoid saturation. Practically, the error from spectrum analyzer is about 3%, which is another 0.25dB uncertainty for the normalized noise. Adding all the errors, the instruments usually cause 2dB uncertainty in low-frequency noise experiments, in which 3-6 decades for the DC biasing currents are considered. The smoothness of the measured spectra is also a problem in low-frequency noise experiments. At the lowfrequency end of the spectrum, one usually desires a resolution of 1Hz or better. This, in turn, results in time window for signal capturing of 1 s or more. Then, since the noise is a random signal, the Fourier transformation of single captured record is with large scatter around the average, e.g. 10 dB or more. Therefore, many records, e.g. 100, need to be captured and averaged, which increases the measurement time to 10 and more minutes per spectrum, in order to obtain spectrum with scattering about 2dB. So, the measurement time becomes an issue for low-frequency noise, and it causes about 30% uncertainty for the normalized noise, if no additional method for averaging is used. One example for insufficient averaging during measurement is shown in Figure 10 at the tradeoff with measurement time. The scattering in this figure is about ½ decade and one will meet with difficulties to examine the normalized noise from this figure with accuracy better than 1/4 decade. III.4.2. Deviations from the ideal 1/f slope in the spectra The second problem related to the experimental accuracy for the normalized noise is the variability of the slope in the 1/f-like spectra. The variations are typically two: the steepness of the slope is different from 1/f; and the slope of the spectrum varies with the frequency, having humps owing to large Lorentzian components in the spectrum. There are physically based models with variable slope in noise spectrum, e.g. analytical in [2, 73], numerical in [68] and the aforementioned for superposition of Lorentzian components [17, 66, 67]. These models are based on deviation from uniform trap density or gaps in the uniform distribution, but a mature characterization technique based on any of those models is not available. There is no standard technique that resolves the problem of the variability of the slope in the 1/f-like spectra. However, some approaches are popular. These are choice of frequency of interest, fitting of formal, semi-empirical or Monte Carlo model, and post-measurement averaging. Choice of frequency of interest The approach of choosing a frequency of interest from the whole spectrum is suitable for volume noise characterizations, e.g. industrial tests. This approach minimizes the time of measurement and compresses the data into a small volume, so that comparisons between many samples and biasing regimes become feasible. The approach is particularly suitable when the goal is to examine the evolution of the noise level with bias or to obtain a statistics for the noise levels among devices on one or from several wafers. Due to its simplicity, the approach of choosing a frequency of interest is widely used, almost in every single publication, and it allowed verifying that the variation of the noise is a reciprocal function of the device area, as shown with one example from [74] in Figure 11 for MOS transistors. However, the approach of choosing a single frequency from the spectrum is vulnerable to large uncertainty owing to deviation of noise spectrum from 1/f slope, since the
32 of 286 approach inherently assumes 1/f slope when estimating the normalized noise at 1Hz, as required for the noise parameter K F . A deviation δ=10-20% from 1/f slope causes error 10dB×δ also multiplied by the number of the frequency decades of extrapolation when referring the noise to 1Hz. A frequency 10Hz is typically chosen in experiments to minimize the measurement time, and therefore, 10-20% (0.4-0.8dB) error is easily introduced for the value of K F . In addition, the randomness of the corner frequency f c of Lorentzian noise in sub-micrometer area devices causes variation from high values for estimated K F , if f c ≈10Hz, to low values, if f c is one-two frequency decades apart from 10Hz. The uncertainty for the value of K F in this case is very large – one and more decades, as one can deduce from Figure 11, and one can estimate meaningless value for K F , if only one frequency point from the noise spectrum is used for the purpose. Therefore, one has to do statistical analysis of many data obtained by the approach of choosing a frequency of interest in order to obtain representative value for the normalized noise and the noise parameter K F . Fitting of formal, semi-empirical or Monte Carlo model The other approach of fitting of formal, semi-empirical or Monte Carlo model to the measured noise spectra is more reliable and less vulnerable to variations in the slope of the spectrum, and therefore is used often in the research on low-frequency noise. The fitting model is chosen usually in the form of eq. (62), and it is enhanced with two more parameters, resulting in ( ) τπ+ τ ++= i2 i ii A B F whiteI f21 B I f K SS F F , (79) where B F =0.8…1.2 reflects the deviation of the flicker noise spectrum from the ideal 1/f slope, and A F =1…2 reflects the bias dependence of the flicker noise. The approach of fitting a model, however, has several drawbacks. The first drawback is that it cannot be formalized into algorithm, because the selection of values for A F , B F and number of Lorentzian components is not unique, and the selection is also dependent on the preselected optimization criterion, the latter also chosen by preferences of the individual researcher. The second drawback, which follows from the first, is that the fitting procedure requires intensive human assistance in an interactive manner, which makes the approach very slow and vulnerable to human errors and individual preferences and skills. The last, but not the least, drawback is that the value and the unit for K F varies with values of B F and A F . As the consequence, the quantitative comparison for the level of the flicker noise from different measurements and samples is almost impossible. Nevertheless, the first two terms of eq. (79) are implemented in the device models for circuit simulations and the evaluation of the corresponding parameters K F , A F and B F helps in the design practice. Therefore, the approach of fitting a formal noise model for BJTs is taken in [32, 34, 37]. Interestingly, persons from industry are co-authoring these publications. III.4.3. Data processing The third problem related to the experimental accuracy for the normalized noise is the data processing from which K F is evaluated. As mentioned above, the scattering in the measured noise spectra is about 2dB and the variations between individual values from different measurements and devices can be larger than 1 decade – see again Figure 11. In such situation, one has to perform post-measurement averaging in order to evaluate the noise and to obtain a representative value for the normalized noise and K F . The simplest way to perform averaging of a noise spectrum is to plot the spectrum in a log-log scale and to draw a line where the density of measured points is the highest. Form the intersect of the extrapolated line with the
33 of 286 axis for noise level at 1Hz, one gets the power spectrum density S(1Hz) at frequency 1 Hz, then the normalized noise and K F are calculated. This approach is used very often, due to its simplicity, but it is vulnerable to human errors. An example that demonstrates the range of errors by manual fitting of data from noise measurements is given in Figure 12. The squares in this figure are the data reported by the authors in a numerical form in a table. The diamonds are the same data, but reported in a graphical form. The discrepancy between numerical and graphical data is apparent, although the manual fit of the data and the fitting of the two data series using least mean square method yield equations, which result in almost overlapping lines when plotted together. Looking closer at the numbers of the fitting equations, the values of the parameters vary about 5% both for the prefactor and exponential coefficient. So, one expects at least 0.2dB uncertainty for the normalized noise when the data from noise measurements are processed manually in a graphical form. This uncertainty could be much higher, if the data scatter, as in Figure 10 or in Figure 11. Another approach that is equivalent to averaging is to analyze the noise in a bandwidth, rather than at individual frequencies. This is a standard procedure for resistors [75, 76, 77, 78], but rarely used for electronic devices, perhaps because the noise in a bandwidth is an integral measure, the frequency slope of the noise is essentially inaccessible from the noise in a bandwidth, and also the noise in a bandwidth is not implemented in device models, since it is expected to be a result from simulations. An attempt for modeling the noise in a bandwidth is presented in [72]. Numerical methods for averaging The numerical methods for averaging of noise spectra can be divided into two groups – arithmetic (precisely, root-mean-square RMS) and geometric averaging of power spectrum density (PSD) (or derivates of PSD, such as f×PSD×Area/I DC2 ). For several power spectrum densities S i , with i=1, 2 …i max , the arithmetic (RMS) averaging uses = max i i i max avg S i 1 S , for arithmetic (RMS) mean, (80) and ( ) − − =σ max i i 2 avgi max SS 1i 1 , for arithmetic (RMS) standard deviation. (81) The geometric averaging [28, 29, 34, 36, 79] uses logarithm of the power spectrum densities, given by ( ) ii,dB SlogdB10S ×= , PSD in dB. (82) Then, = max i i i,dB max avg,dB S i 1 S , for geometric mean, (83) and ( ) − − =σ max i i 2 avg,dBi,dB max dB SS 1i 1 , for geometric standard deviation. (84)
34 of 286 When the variations between individual spectra S i are small, e.g. in large-area BJT, then both methods yield similar results, that is ( ) avgavg,dB SlogdB10S ×≈ and ( ) σ+×≈σ+ avgdBavg,dB SlogdB10S . (85) However, if the variations between individual spectra S i are spread over one decade or more, then the arithmetic averaging results in variation larger than the average [72], σ>S avg , it becomes impractical (because the noise is attributed to σ rather than to S avg ), and the geometric averaging is more suitable for sub-micrometer area devices, since the distribution of noise variation tends to log-normal distribution [28, 29, 79, 80]. Therefore, many authors use geometric averaging [17, 28, 29, 34, 36, 66, 67, 79, 80, 81, 82]. Worth mentioning, it is empirically observed that the noise variations are better described by log-normal distribution, and substantial work is expected to explain the origin of this empirical observation [79, 80]. A reason is given later in section VIII.6. “Consequences from statistical nature of LFN – distributions in spectra, techniques of averaging, data volume and coordinates, instrumentation”. Interestingly, assuming Poisson distribution of threading dislocations in strained silicon wafers, the geometric model of noise variation explains the data in [82]. However, one should be careful when using the equations for geometric averaging and modeling of noise, in order to preserve the physical consistence of figures of merit (FOM). For example, defining FOM for relative variation (in respect to mean level of noise), the arithmetic (RMS) averaging suggests [72, 79] avg ari S FOM σ = , for relative variation from arithmetic (RMS) averaging, (86) The FOM for relative variation from geometric averaging is actually σ dB [28, 79, 80, 81], because it follows from eq. (85) that [17, 66, 67] ( ) dB10 avg dBgeo dB 101 S FOM σ =+ σ σ≡ , for relative variation from geometric averaging, (87) whereas the FOM=σ dB /|S dB,avg | as attempted in [33, 34, 36] is just a number that will violate the rules for physical units, if one tries to convert it to ratio of noise level. In the above equations, note again that σ is standard deviation and S avg is average of noise S, the latter being a variance (usually per unit frequency) in principle. To provide impression for the noise variation in BJTs, we present in Figure 13 several views of the few data available from literature. These data have been also obtained using geometric standard deviation according eq. (84). The shaded areas in the figure represent cases when the variation of the noise is larger than the average level, that is σ>K F . The top-left plot in the figure shows the data vs. emitter area, as originally reported. One can see that the noise variation increases in small-area BJT and in sub-micron area BJT the variation is dominating over the average. The slope in this plot is somehow low, only -2.5dB/dec, but the bottom-left plot, obtained using eq. (85) or eq. (87), clearly shows that the absolute variations are a strong function of the emitter area, σ∝(A E ) -1.5 , and taking into account that K F ∝(A E ) -1 , then one observes the aforementioned increase of (A E ) -0.5 in noise variation as function of decreasing device area, as predicted by Monte Carlo simulations [17, 66, 67] and analytically in [72]. More interesting, and to the best of our knowledge not explored explicitly yet, are the observations that one can make in the right-hand plots in Figure 13. The main observation is that the noise variation scales with the average noise. This is clear in the bottom-right plot, where the slope 3:2 suggests a trend σ∝(K F ) 1.5 for the
35 of 286 absolute variation σ of the noise as function of the average noise K F , and reassembles the corresponding slope - 3:2 for the dependence σ∝(A E ) -1.5 . Owing to the limited number of experimental data points, it is difficult to identify properly the factors that cause the relation between noise variation and average noise. For example, there is no unique line that can be drawn in the top-right plot. The overall behavior in Figure 13 suggests that 5.1 EEF AAK − ∝∝σ , but the limited amount of experimental data does not allow parameterizing the dependence reliably. Nevertheless, for BJTs with emitter area A E <0.3µm 2 and K F >3×10 -8 , the noise variation appears to be more important than the average noise, since the noise variation affects the repeatability in the noise performance of the devices. Consequently, large discrepancy between simulations based only on average noise and the real noise performance of the individual device is expected for nowadays BJT. More discussions on averaging techniques are given in section VIII. “ Outlook for the LFN ” after eq.(469). Summing all factors discussed above we estimate that the uncertainty of each data point in Figure 2 is about 3dB, owing to different measurement setups and techniques for model fitting and averaging. This uncertainty, perhaps, contributes significantly to the value of σ dB that is shown in the figure. Nevertheless, we use the data for npn BJTs from Figure 2 as a benchmark in comparisons to other devices. When comparing the values for K F , we use the data which are within and close to the interval A E ×K F ≈5.6×10 -9 µm 2 ±2σ dB . When comparing the values for input voltage noise S V , the same reference data are referred as the base voltage noise using the relation f×A E ×S VB (f)=A E ×K F ×(φ t ) 2 from eq.(3), which applies when the noise in BJT is coupled from the emitter area to the diffusion currents, according to the discussions on eq. (36). Comparative study of noise in npn and pnp BJTs Past publications , for example [83], conclude after comparative studies on devices from complementary technologies (BiCMOS) that the noise in pnp BJTs is lower than the noise in their npn counterparts. We have searched for published data for low-frequency noise in silicon pnp BJT. The data are collected from [12, 32, 83], stored in numerical form in [84], and shown in Figure 14 together with data and trend for the noise in silicon npn BJTs. The information for noise in silicon pnp BJTs is less available in the last two decades, as compared to the data for silicon npn BJTs and the more recent interest on SiGe HBTs. In contrary to the above comparative studies, we did not found data for the noise in pnp BJT, which could confirm the conclusion of the comparative studies. This is illustrated in Figure 14. The open circles and gray lines are from Figure 2 and represent data and trend for noise coefficient K F in npn BJTs. The solid symbols and black lines in Figure 14 represent data and trend for noise coefficient K F in pnp BJTs, and they are above the trend for npn BJTs. Also, studies [34] on noise in complementary SiGe HBTs indicate that the noise in pnp transistors is higher than the noise in npn transistors. Interestingly, the data for pnp BJTs tends also to log-normal distribution, as shown in the insert of Figure 14. Nevertheless, the difference in noise coefficients K F between npn and pnp transistors is within the range of scattering of the data for low-frequency noise, and it does not lend much confidence to make a general conclusion whether the noise is lower in npn or in pnp BJTs. Some differences in the bias dependence of the noise and for the impact of IFO in npn and pnp transistors are reported in [32], which makes the fair comparison of the noise in these transistors difficult. IV. Noise in MOS transistors The low-frequency noise in MOS transistors is widely studied because of several reasons. First, a physical model has been proposed in [85] long time ago and before any model for noise in BJT was accepted. Second, the MOS
36 of 286 technologies are the driving force in the downscaling of electronic devices in the last decades. Third, the lowfrequency noise in MOS transistors is higher than the noise in BJTs, as one can see in Figure 15. In this figure, the data are for 134 nMOS transistors from [22, 47, 48, 49, 50, 51, 72, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123], and for 53 pMOS transistors from [47, 52, 86, 88, 95, 96, 97, 100, 102, 104, 105, 109, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132]. Many data points overlap in Figure 15, the lowest levels for noise in MOS transistors is higher than the highest levels of noise in BJT, and the data scatter about 3.5 decades for MOS transistors, while the data scatter less than 2 decades for BJT. Owing to these observations in the figure, it is clear that the control of the noise in MOS transistors is more difficult than it is in BJT. The main reason is that the operation of the MOS transistor is governed by a surface current transport, which is exposed to interface phenomena at the semiconductor-gate dielectric interface, while in BJT the current transport is mostly in the bulk of the semiconductor and it is “less” sensitive to the surface of the semiconductor. The consequence is that the current research on low-frequency noise in MOS transistors is focused on the fabrication of the gate stack at the semiconductor surface and although important and studied, the variation of the noise with the channel size is not dominant in the scope of many publications. Therefore, many researchers use samples of “standard” gate area, e.g. around 2-3µm 2 , or 10-20 µm 2 , but different gate stacks, and unique trend in Figure 15 cannot be placed accurately. Nevertheless, the product W×L×S VG for MOS transistors tends to log-normal distribution, as shown in the insert of Figure 15, and the mean in this distribution is about 2.5 decades above the average noise in BJT, which is in agreement with past ITRS predictions about the difference in the noise between MOS and bipolar transistors [3] – see again Figure 1a. Interestingly, when looking at the distributions for the product W×L×S VG separately for nMOS and pMOS transistors in Figure 16, the average noise in pMOS transistors is about 3dB (2 times) lower than the noise in nMOS transistors, but at the same time, the distribution of the noise values is broader for pMOS transistors, which practically does not lend much confidence to conclude which type of MOS transistors is with lower levels of noise. Comparisons for noise in pMOS and nMOS transistors report opposite observations – noise in pMOS transistors is lower [95] and higher [83, 98] than in nMOS transistors. Having the above observations in mind, and the data for noise available, we begin the discussion on the physics and models underneath these observations and the scattered data of low-frequency noise in MOS transistors. IV.1. Models and predictability It is widely accepted that the low-frequency noise in the drain current of the MOS transistor is coupled by the charge and mobility fluctuations related to the interface between semiconductor and gate dielectric. The approach of intrinsic and coupled noise allows comparing the low-frequency noise of MOS transistors to other devices, e.g. bipolar junction transistors, as it is discussed in the previous section. So, with the obvious for MOS transistor notations, we rewrite eqs. (11) and (12) as fn S I g K I S eff H V 2 D m 2 D I G D α + = , (88) for the purpose of comparing input referred voltages, and fn S C q I g K I S eff H N 2 ox 2 D m 2 D I t D α + = , (89)
37 of 286 for the purpose of analyses of charge trapping at interface between semiconductor and gate dielectric. Here, I D [A] is the DC drain current, S ID [A 2 /Hz] is the power spectrum density (PSD) of the output drain current noise. For the coupled portion of noise, popular as noise from number fluctuation ∆n, the relevant quantities are transconductance g m =∂I D /∂V G [A/V=S], PSD of input referred (gate) voltage noise S VG [V 2 /Hz], gate capacitance per unit area C ox [F/cm 2 ], PSD of trapped charge per unit area S Nt [cm -4 /Hz], and K is coupling parameter which depends on many factors, including bias, according to different physical models, but practically it is taken K≈1. IV.1.1. Intrinsic noise (mobility fluctuation) For the intrinsic portion of noise, popular as noise from mobility fluctuation ∆µ, α H is the Hooge parameter and n eff is effective number of carriers in the channel [9]. The effective number n eff ≤n takes into account the nonuniform distribution of the total number n of carriers in the channel, when the drain voltage is not zero, and can be deduced by using eq. (13), but the procedure of evaluation is complicated, and one usually takes n eff ≈n when investigating the low-frequency noise in MOS transistors in strong inversion regime at gate biasing V GS >V T above the threshold voltage V T , especially when the transistor operates in linear (ohmic) mode at low drain bias V DS ~50mV<V GS -V T . Evidently, even before assuming any physics for the noise, quite a lot approximations and assumptions can be taken differently by different researchers in the equations for the low-frequency noise in MOS transistor, and this causes large spread in the values reported for the parameters related to the noise, and sometimes even to controversial conclusions. We illustrate this with one example for the intrinsic component of the noise. Assume that V GS -V T ≥0.1V and 0≤V DS ≤ V GS -V T -0.05V. Such biasing of MOS transistor is usually regarded as operation in ohmic mode, and one takes n=WLC ox (V GS -V T )/q, neglecting that the charge concentration is lower at the drain side of the channel. Assume also that the intrinsic (Hooge) low-frequency noise is dominant (e.g. pMOS transistor), so that the Hooge parameter is estimated from eq. (88) as ( ) q VVWLC I S fn I S fn@ TGSox 2 D I 2 D I H DD − ==α , at approximation n=WLC ox (V GS -V T )/q. (90) Now, let us increase the “accuracy” of the calculation of the total number of carriers, assuming gradual approximation for charge concentration. The average number of carriers n avg is the mean value of the number of carriers at source and drain sides, n avg =(n source +n drain )/2, and n avg becomes ( ) ( ) −−= −− + − =2 V VV q WLC q VVVWLC q VVWLC 2 1 n DS TGS oxDSTGSoxTGSox avg . (91) Accordingly, the Hooge parameter is estimated as −−==α 2 V VV q WLC I S fn I S fn@ DS TGS ox 2 D I avg 2 D I avgH DD , at n avg =(n source +n drain )/2. (92) Let us increase further the “accuracy” of the calculation using the integral for current crowding given by eq. (13), rewritten for 1D case, by assuming constant current density J, charge sheet approximation (the thickness of inversion layer T≡Dirac function, yet not accurate for sub-100nm MOS transistors) and thus WTJ=constant.
38 of 286 ( ) ( ) α = ⋅ α == L 0 2 H 2 L 0 2 L 0 4 H 2 D I norm )x('n dx fL dxWTJ dxWTJ f)x('n I S S D , (93) where the coordinate x=0…L is along the channel length, and ( ) −−= ∂ ∂ =L x VVV q WC x n x'n DSTGS ox . (94) Then, −− − = −− = DSTGS TGS DSox L 0DSTGS ox L 0 VVV VV ln VWC qL L x VVV q WC dx )x('n dx (95) So, one gets the following expression for the Hooge parameter −− − =α DSTGS TGS DSox 2 D I effH VVV VV lnq VWLC I S fn@ D , with −− − = DSTGS TGS DSox eff VVV VV lnq VWLC n , (96) where the effective number of carriers in the channel, n eff [9], is lower than the above approximations n=WLC ox (V GS −V T )/q and n avg =WLC ox (V GS −V T −V DS /2)/q. Provided that one processes the same data from one sample, but using different assumptions for number of carriers, Figure 17 illustrates possible discrepancies in the values estimated for Hooge parameter, since the value of α H is proportional to the estimated number of carriers in eqs. (90), (92) and (96). One can see from Figure 17 that 20% to 50% variations (1dB to 3dB) in the estimated values for α H can be easily introduced by changing only the characterization model. Similar vulnerability is discussed in [133] for the coupled part of the low-frequency noise in MOS transistors (popular also as number fluctuation ∆n) when the drain bias is not low, and brings the transistor in saturation mode, resulting in doubling the noise levels (3dB increase). Thus, ½ decade in the scattering of the data in Figure 15 we attribute to the use of different characterization models and procedures, since no standard method is currently accepted for MOS transistors. IV.1.2. Coupled noise component (number fluctuation) The coupled noise component in MOS transistors is widely studied in terms of trapping at semiconductor-oxide interface states [134] or charge trapping in gate oxide, as originally suggested in [85]. These approaches are known as random walk and tunneling ∆n models for 1/f noise in MOS transistors, respectively. They are illustrated in Figure 18 and discussed below. Another possibility for charge trapping and associated coupled noise component is when the traps are in the semiconductor [135], either in the inversion or/and in the depletion layer of the MOS transistor channel. The exploration of this possibility is feasible, when Lorentzian noise spectra are present, but random-telegraph signal (RTS) noise waveforms are not observed, and the noise is attributed to traps with particular energy, in contrary
39 of 286 to the expectation for uniformly distributed in energy and space traps in gate dielectrics. This possibility of trapping in the semiconductor of the MOS transistors is rarely addressed in the literature, since the interface effects are dominant in the MOS transistors, and not discussed below. Random walk model The random walk model [134] assumes that a charge carrier is captured at an interface state and after a time τ it moves to a neighboring interface state. The mean free path of the charges “walking” at the semiconductorinsulator interface is taken as atomic distance l≈0.2nm and the cross-section of charge trapping is taken σ=10 -16 cm 2 =(0.1nm) 2 . For charge sheet (2D) approximation of the inversion layer in MOS channel, the distribution of time constants is ( ) τ ≈ τ π σ =τ 1.01 2 g l . (97) For the random walk, it is derived in [134] from superposition of Lorentzian spectra (see eqs. (73), (74) and (75) earlier) that the gate referred voltage noise in eq. (88) is it 2 ox it 2 ox V D1.0 WL kT C q f 1 kTD 2WL 1 C q f 1 S G ≈ π σ =l , for random walk, (98) where the interface states are assumed to be distributed uniformly with density D it [cm − 2 eV − 1 ]=constant. Tunneling model The tunneling ∆n model for 1/f noise in MOS transistors, which was originally suggested in [85], has been followed up widely [47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 72, 74, 80, 82, 83, 87, 89, 91, 95, 101, 103, 106, 107, 110, 124, 126, 127, 128, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145]. It assumes that some charge carriers are trapped in the depth of the gate oxide. The tunneling probability is exponentially decaying function of the distance x ti from semiconductor-insulator interface to the position of the trap in the oxide. Therefore, the rate 1/τ of tunneling events is λ − τ = τ ti o x exp 11 , (99) where 1/τ o is attempt rate (frequency) taken by different assumptions in the range of from 10 7 s − 1 [134] to 10 10 s − 1 [49, 55, 140], and λ is tunneling (electron wave) attenuation distance, given by eq. (23). The values for λ range between 0.06 nm and 0.22 nm for different semiconductor-insulator pairs, electrons and holes – see Table 2, because λ is function of the product of effective mass of the carrier in the insulator and the energy offset between the bands in semiconductor and insulator, according to Wentzel-Kramer-Brillouin (WKB) approximation. For Si-SiO 2 and electrons, one usually takes λ=0.1nm – see [146], for example. At an assumption for uniformly distributed traps in the oxide depth, ∂N t /∂x ti =constant, as shown in many publications, for example in [70], it follows from eq. (99) that ln(τ)∝x ti /λ∂τ/τ=∂x ti /λ, and the distribution of the traps vs. their time constants becomes ( ) τ = τ λ = τ∂ ∂ ∂ ∂ = τ∂ ∂ =τ b constant x x NN g ti ti tt . (100)
40 of 286 Then, the superposition of Lorentzian spectra of the individual traps in the oxide produces 1/f noise – see again the procedure by eqs. (72), (73), (74) and (75). In this way (see [28] for example), the gate referred voltage noise in eq. (88) becomes as t 2 ox V N WL kT C q f 1 S G λ = , for tunneling in gate insulator, (101) where the traps in the oxide are assumed to be distributed uniformly with density N t [cm -3 eV -1 ]=constant. Comparison of random walk and tunneling models Comparing eqs. (98) and (101), one sees that the random walk and tunneling models converge, although the physical assumptions behind these models are different. The reason is that D it in the random walk model and N t in the tunneling model are assumed uniformly distributed both in space and energy, in order to obtain 1/τ distribution for the time constants of the traps, and from this – 1/f noise. This is discussed in more detail in [28, 147], where also is observed that D it from the low-frequency noise technique converges with charge pumping technique in the frequency range 10kHz to 100kHz that is accessible by both techniques. The deduction in [147, 148] is that the 1/τ distribution arises from traps’ effective cross-section σ eff apparent at semiconductor-dielectric interface, and σ eff is exponential function from both activation energy E B of interface states and distance x ti from semiconductor-insulator interface, given by λ − −∝σ∝ τ ti B eff x exp kT E exp 1 . (102) Accordingly, ( ) [ ] λ ∂ + ∂ = τ τ∂ =τ∂ ti B x kT E ln , (103) and following the procedure in eq. (100), one sees that the 1/τ distribution can be achieved either or both assuming uniform distribution for the activation energy at the interface, ∂D it /∂E B =constant, or uniform spatial distribution of the traps in the depth of the gate insulator, ∂N t /∂x ti =constant. Figure 19 illustrates how the uniform distributions in trap energy E B =0.2...0.6eV and tunneling distance x ti =0.8…2.3 nm produce 1/f noise. Therefore, one can quite arbitrary attribute the 1/f noise to interface states or oxide traps, and this demonstrates the convergence between random walk and tunneling models for the gate referred 1/f noise voltage S VG in MOS transistors. In fact, both trapping mechanisms can superimpose, one can combine tunneling and random walk models, and we write for S VG in eq. (88) that ( ) itt 2 ox V D1.0N WL kT C q f 1 S G +λ = . (104) Rewritten in terms of charge fluctuation, S Nt in eq. (89) is then given by ( ) ittN D1.0N WL kT f 1 S t +λ= . (105) One can assume even a two-step process: first carrier capture in interface state and then random walk and tunneling. Perhaps, this is the reason why the range for τ o in eq. (99) is taken with values in several decades by
47 of 286 2 m D 2 m D o g I 1 g I 1K θ+≈ µ µ θ+= , for MOS transistors with correlated ∆n-∆µ fluctuation, (121) where the mobility degradation parameter θ may vary in sign, magnitude and with bias, depending on type of trapping and scattering mechanism in MOS transistor, e.g. Coulomb, remote, surface, phonon scattering and screening of charge of the inversion layer [47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 72, 89, 110, 124, 137], but θ is usually taken as a constant in one transistor, because of the following reasons. At low gate bias, V G - V T <0.1V, the contribution of the correlated mobility fluctuation is low, because the ratio I D /g m <0.1, and since θ<1V -1 , then K≈1. At high gate bias, I D /g m ∝(V G −V T )>0.5V, K becomes a nearly quadratic function of (V G −V T ), but the product θ(V G −V T ) is usually less than 1, and variations in θ are difficult to inspect reliably from noise measurements data with an experimental uncertainty 1-3dB – see for example fig.4 in [89]. So, the experimental characterizations assume θ=constant; and the model of eq. (120) is very attractive since all quantities (S ID , I D , g m ) can be estimated directly from the measured data without a need of device modeling. For example, the product (S ID /I D2 )(I D /g m ) 2 at low gate bias (V G −V T )<0.1V directly estimates S FB , and then at higher gate bias (V G −V T )>0.5V, the value for θ can be obtain from the slope of the graph G V S vs. V G , see for example [89], again using (S ID /I D2 )(I D /g m ) 2 =S VG , since I D /g m ∝(V G −V T ). The direct parameter extraction by using the model of eq. (120) is a routine approach in noise characterization and most of the results presented here were obtained using this approach. Actually, it is argued in [133] that assuming ∆n noise being regarded as noise S VFB ≈constant in flat band voltage, there is inherent error in the model that underestimates the noise level when the transistor is in saturation regime. The error is about 3dB, it is small, it is in the range of experimental inaccuracy, and using approximations for charge concentrations would cause similar uncertainty. So the model of eq. (120) is very attractive in experimental characterizations and it is widely used. The situation, however, has changed when high-k dielectrics and high doping of the channel are used in order to reduce the channel length of the transistors below 100nm. The electric fields and scattering increased in these devices and the correlated mobility fluctuation becomes an issue not only because the contribution from the scattering reduces significantly the effective mobility in MOS transistors, but also because the noise models are developed at an assumption of dominant phonon scattering, while Coulomb screening by the inversion layer and roughness scattering [157, 158] also take place. Therefore, the alternative forms of the model with correlated ∆n- ∆µ fluctuations for MOS transistors are discussed below, because they are derived at physical assumptions for charge transport in MOS transistors, while the model of eq. (120) above was introduced formally by using the concept for coupled noise at the assumption that the oxide charge trapping can be referred as noise in the flat band voltage. These models are known as “unified” model for 1/f noise in MOS transistors and use charge fluctuation due to trapping in gate oxide accompanied with scattering due to Coulomb interaction around the trapped charge. Unified model. The starting point in deriving the unified model is that in event of charge trapping in the oxide at distance x ti from semiconductor interface at coordinate (w,l) in the channel region, in which the carrier areal concentration is n’=∂ 2 n/∂w∂l, the drain current changes as [54, 55, 56]
48 of 286 ( ) µ µ ±−=⇔µ+µ∝ d 'n 'dn I dI dwdld'qdn'dndI D D D , (122) where dn’ and dµ are functions of trap occupancy, given by ( ) ,N 11 with dNdN N d and ,CCC q n* , *nn' n' R with dNRdN N 'n 'dn ts o tst t itdox t tt t α+ µ = µ µα= ∂ µ∂ = µ µ ++ ϕ = + == ∂ ∂ = (123) where C d and C it are the capacitance per unit gate area due to depletion under the conductive channel in the MOS transistor and fixed interface states at the semiconductor-insulator interface, respectively, and α s is a scattering coefficient, which takes into account for Coulomb interaction between oxide trapped charge and thus changing the mobility of the carriers in the inversion layer. In order to solve the integrals, several assumptions are made: the oxide trapping dN t is negligible as compared to n’, N t , α s and µ can be replaced with constants, and d(N t α s µ)=−d(n’α s µ/R). In this way, the general equation for the unified 1/f noise model for MOS transistor is written in several equivalent forms, one of which is = µλ = D S BGD V 0V 2 2 D V and Vconstant atI dv 'n R *N fL kTqI S , (124) where V G , V B , V S and V D (on which the areal carrier concentration n’ in the channel of MOS transistor depends on) are the potentials of the gate, body, source and drain terminals, respectively, all other quantities are related to n’, and also by using the approximation ( ) ( ) 2 st 2 R/'n1N'nC'BnA*N µα±=++= , (125) where the parameters A=NOIA, B=NOIB and C=NOIC (being assumed constant fitting parameters in BSIM3 model) are approximately corresponding to A=N t , B=2α s µN t /R≈2α s µN t and C=(α s µ/R) 2 ≈(α s µ) 2 . The sign ± is chosen by whether the trap is repulsive or attracting for the carriers in the channel, as mentioned in [48]. In the next step, the unified model is split into three integrals, which always have analytical solutions at A, B and C constant. These solutions are appropriate for compact modeling, since they capture the coupling of noise in the MOS channel by the non-uniform carrier concentration at various drain biasing, depending only on n’ at drain and source sides, where the overdrive (V G -V T -V) with V=V S or V D , are known from the biasing. The model, however is slightly inconvenient for experimental characterizations, because it requires additional estimates for capacitances, the equations are long, see [48, 54, 55], and cannot be rewritten in a form so that A, B and C do not depend on each other. Nevertheless, the unified model converges to the flat band model of eq. (120) above in ohmic and sub-threshold operation of MOS transistor, and the derivations established for linear regime that ( ) ( ) ( ) ( ) ( ) tTG TGox s itdoxtTGox ssTG 3...2VV if , q VVC q CCCVVC R/'n)VV( ϕ>− − µα≈ ++ϕ+− µα=µα=−θ (126) Therefore, since θ is directly obtained from measurements as discussed above, α s can be easily evaluated from
49 of 286 ox s C q µ θ =α ~10 -15 Vs (127) taking typical values for 180nm nMOS (C ox ~0.8 µF/cm 2 , θ~1 V -1 , µ~200 cm 2 /Vs). Effects of different scattering mechanisms. The issue with correlated mobility is now addressed. It is due to the fact that neither α s or µ are constants. This is because changing the gate bias, the electric field changes and there are several crossovers between different scattering mechanisms [157, 158]. These are illustrated in Figure 24. The mobility is given by the Matthiessen rule, as bitrph 11111 µ + µ + µ + µ = µ , (128) where µ ph is due to phonon scattering, µ r is due to surface roughness scattering, µ it is due to Coulomb scattering caused by interface states and oxide traps and µ b is due to scattering with ionized impurities in the semiconductor. The investigations in [157, 158] have established that the different components have different dependences with the biasing of the MOS transistor. - The phonon scattering mobility is given by 3.0 E dpl 3.0 Si E 75.1 ph 3.0 G e ph ph 'n N q TaETa 1 t + η ε η ≈= µ , (129) with T being the absolute temperature, a ph , e t , η E =0.5….0.3 being constants, ε Si being the permittivity of semiconductor (~ 1.04×10 -12 F/cm for silicon), E G being the gate electric field, E G =q(N dpl +η E n’)/ε Si , and N dpl being the depletion charge per unit area, according to ϕ ε = i sub tsub Si dpl n N lnN q4 N , (130) where N sub is the volume impurity concentration in semiconductor and n i is the free charge concentration of intrinsic semiconductor (~10 10 cm -3 for silicon at room temperature). The line labeled with “phonon scattering” in Figure 24 illustrates µ ph . At fixed temperature, one can combine the constants into the phonon scattering parameter α ph and rewrite eq. (129) as ( ) 3.0 ph 3.0 E dpl ph ph 'nn'n N 1+α= + η α= µ η , (131) where n η is a scaled version of N dpl . - The surface roughness scattering mobility is given by ( ) 12 E dpl e Si E r e G r r 'n N q aEa 1 r r ± + η ε η ≈= µ , (132) where a r and e r are approximately constants. The line labeled with “surface roughness scattering” in Figure 24 illustrates µ r , showing that this mobility degradation is observed when the electric field is high, e.g. above
50 of 286 0.5MV/cm. Again combining the constants into a scattering parameter α r , one writes ( ) ( ) 12 r e E dpl r r 'nn'n N 1 r ± η +α= + η α= µ . (133) - The Coulomb scattering is prominent when the inversion layer is with low concentration of carriers. The inversion layer screens the Coulomb scattering at higher gate bias. For the Coulomb scattering due to ionized impurities in the semiconductor, the mobility µ b is given by 'n'n N a 1 bsub b b α ≈= µ , (134) where α b is the corresponding scattering parameter. The curve labeled with “impurity screening” in Figure 24 illustrates µ b , showing that this mobility degradation is observed when the electric field is low, e.g. below 0.2MV/cm. At a little higher gate biasing, as shown with the curve “interface screening” in Figure 24, the Coulomb scattering caused by interface states and oxide traps takes place. Respectively, the mobility µ it is 'n'n NbD a 1 ittitit it it α ≈ + = µ , (135) where α it is the corresponding scattering parameter and b it refers the oxide traps as apparent areal density at the semiconductor-oxide interface. - Combining all scattering mechanisms, one gets ( ) ( ) ( ) ts o 12 r 3.0 ph itb N)'n( 1 'nn'nn 'n 'n 1×α+ µ =+α++α+ α + α = µ ± ηη , (136) where the terms are written in a sequence as they are dominant from low to high gate biasing, resulting in a bellshaped curve for µ, labeled with “effective mobility” in Figure 24. The last equation implies that the effective scattering parameter α s varies with the gate biasing, and α s is high at low and high areal carrier density n’ in the inversion layer in MOS transistor channel, as shown with circles in Figure 24. Despite these variations, the term ( ) q VVC 1R/'n1)VV(1 TGox ssTG − µα+≈µα+=−θ+ (137) for correlated mobility fluctuation does not change very much, as shown with squares in Figure 24, because the product α s µ is constant, if one scattering mechanism is dominant. This can be seen from eq. (136) by neglecting 1/µ o . Usually, the term [1+θ(V G −V T )]≈[1+θI D /g m ] can be fitted approximately with a linear function within the experimental inaccuracy, especially in silicon MOS transistors with not very thin (>4nm) and well processed SiO 2 for gate insulator. In these transistors, since the transistor channel is L>0.3 µm, then the phonon scattering usually dominates, because N sub <3×10 16 cm -3 and E G <0.6 MV/cm. A solution at dominant Coulomb scattering due to interface states and oxide traps is deduced in [159], proposing a modification of eqs. (124) and (125). The modified equation is 2 t itt 2 0C t 2 D I 'nN 'n 1 fWL NkT 'n 'n 1 fWL NkT I S D µα + λ ≡ µ µ + λ = , (138)
51 of 286 where the parameter µ C0 ≡N t /α it reflects the parameter for screened Coulomb scattering at interface states and oxide traps in eq. (135). This scattering mechanism would be pronounced in sub-100 nm MOS transistors, since N sub ≈5×10 17 cm -3 , and E G <0.8MV/cm. Assuming µ≈µ it in eq. (136), then α it µ it (n’) -0.5 ≈constant≈1/N t , which is a paradox of cancelling the mobility when the other scattering mechanisms are neglected. In fact, the identification of scattering coefficients for MOS transistors from sub-100nm MOS is difficult, because N sub >7×10 16 cm -3 in these transistors in order to compensate for DIBL (drain induced barrier lowering that affects the threshold voltage V T ), resulting in E G >0.6 MV/cm necessary to invert the channel conductance and control the inversion layer. These are accompanied with crossover between different scattering mechanisms, and the linear approximation is not precise anymore. Two examples for crossover are shown in Figure 25. In both examples, the correlated mobility noise due to Coulomb scattering is screened and ceases with increasing the gate overdrive voltage. In the second example in Figure 25b, the noise of phonon or roughens scattering causes a rising correlated mobility noise. These data have been explained in other manner in [57], and the crossover between different scattering mechanisms is not analyzed in details for the case of low-frequency noise, although some suggestions are available in [58, 159]. Issues with condition iii. This condition states that the charge exchange between channel, interface states and oxide traps occurs only at Fermi level and the charge trapping is a Shockley–Read–Hall (SRH) process. The assumption allows using Fermi-Dirac statistics in the superposition of the fluctuations of individual traps, by integration over energy. For example, the gate referred noise voltage is obtained in [54, 55, 56] from ( ) ( ) dxdE)E(f1)E(f 1 N4 WLC q S ox G t 02 t 2 ox 2 V − ωτ+ τ = ∞+ ∞− , (139) using the probability function f(E) of Fermi-Dirac statistics, given by kT EE exp1 1 )E(f F − + = , (140) where E F is the quasi Fermi level in the inversion layer of MOS transistor channel. So, the factor f(1-f) provides the term τ e τ c /(τ e +τ c ) 2 of the emission and capture time constants of charge trapping in the limit of Shockley– Read–Hall process. The term τ e τ c /(τ e +τ c ) 2 participates in the expression for Lorentzian spectrum of a generationrecombination process. The factor f(1-f) is sharply peaking function of E, which allows the inner integral of eq. (139) to be solved assuming all other quantities unchanged. Since, ( ) kTdE)E(f1)E(f =− +∞ ∞ − , (141) then the absolute temperature T occurs in the final expressions for the noise, as given by eqs. (104), (105), (109), (124), (138) and other derived from them. These equations suggest that the 1/f noise in MOS transistors should be proportional to the absolute temperature even assuming tunneling in gate oxides. The issue is that this proportionality is not observed experimentally [160], rather, the 1/f noise is found to be temperature independent in nMOS transistors [91] (with a small variation in the slope of the spectrum), while the tunneling model for
52 of 286 noise should be valid for nMOS transistors. Thus, the use of Shockley–Read–Hall process might be incorrect for the tunneling noise from oxide traps, since it is questionable whether the traps are in equilibrium. On the other hand, the capture and emission time constants in RTS noise in MOS transistors are found to follow Shockley– Read–Hall process [68, 74, 107, 112, 148, 161, 162]. Other issues related to the assumptions for superposition of tunneling events in gate oxide in creating 1/f noise are discussed in [4]. Issues with condition iv. This condition for ∆n models of MOS transistors states that the trapping centers are uniformly distributed both in space and energy, the populations of traps and carriers are large enough to be assumed continuous and approximated with averages, and the mobility and electric field do not fluctuate. Let us inspect the numbers for a sub-100 µm MOS transistor from the L=65nm technology. Assume W=3L=195nm, WL=2.1×10 -10 cm 2 , EOT=2nm, C ox =1.8µF/cm 2 , (V G -V T )=0.5V, inversion layer carrier density n’=C ox (V G -V T )/q=5.5×10 12 cm − 2 , number of carriers WLn’≈700. Using ITRS past predictions for the 65nm technology node [3], we would have WLS VG (at 1 Hz)=FOM SVG =1.6×10 − 10 µm 2 V 2 /Hz, which would result in N t =3.68×10 17 cm − 3 eV − 1 , assuming tunneling attenuation distance λ=0.2nm close to that of HfO 2 . The traps only within ±3kT around quasi Fermi level fluctuate, the other are either occupied or empty. For a noise measurement in 3 frequency decades, the traps are located in a slice of oxide thickness ∆t ox =λln(10 3 )=1.38nm. So, the number of traps observed in the measurement would be 01.1WLNtkT6 tox ≈∆ , for W/L=195/65 nm transistor in 3 decades of frequency. (142) Evidently, the population of traps is small in sub-100nm MOS transistors and the assumption in the ∆n models for 1/f noise in MOS transistors for continuous distributions approximated with averages is not valid. One should (and do) observe Lorentzian spectra and RTS noise, instead of 1/f noise, in these devices. However, 1/f noise is still present in sub-100nm devices, and it is perhaps from the intrinsic (Hooge or mobility) noise, because the number of carriers is in the range of several hundred carriers, which are at least 10 carriers per frequency decade. So, the assumption, as made in the ∆n models for 1/f noise in MOS transistors, that the mobility (electric field or other quantity related to the intrinsic properties of carrier transport) can be neglected, is not valid for sub-100nm MOS transistors. Nevertheless, the ∆n model converges well with the observed RTS in small-area MOS transistors, as it will be shown later in section IV.3 , and the model should not be “retired”; rather, the model will be properly analyzed and extended to describe the noise variation when the populations of traps and carriers are small. In fact, the statistical modeling of noise based on the ∆n model has begun, both empirically, analytically and by means of simulations. In this paper, we introduced the problem of statistical noise modeling in the section for BJT – please see again eqs. (80), (81), (82), (83), (84), (85) and the discussion after eq. (85). More discussions on averaging techniques are given in section VIII. “ Outlook for the LFN ” after eq.(469). Issues with condition v. This condition for ∆n models of 1/f noise in MOS transistors states that the barrier for tunneling in the oxide (or for trapping in random walk model) is constant and not affected by the electric field. Also, the barrier is at the semiconductor-dielectric interface and the channel carriers are at this interface too. Certainly, these assumptions were good for gate oxides with thickness t ox >10nm until the gate electric field was
53 of 286 less than E G ≈(V G -V T )/t ox <1V/10nm=1MV/cm, c.f. nodes with minimum gate length L min >0.35µm, which was the case when the ∆n models were developed. In a thin insulator and at high fields, however, the approximations with constant parameters become rough. Consider Figure 26 for a MOS transistor with gate insulator stack, such as HfO 2 with interfacial layer of SiO 2 . The steep slope of the potential in the semiconductor creates a potential well near the dielectric. Owing to quantum effects, the energy levels of electrons increase with approximately ∆Φ q ~0.2eV at electric field greater than 0.5MV/cm [51, 163]. One effect is that the tunneling barrier effectively decreases with ∆Φ q , which is also accompanied with a departure from rectangular barrier, leading to another barrier lowering with ∆Φ e , especially when the oxide is thinner than 1.5 nm [164]. Consequently, the tunneling attenuation distance λ increases according to from Wentzel-Kramer-Brillouin (WKB) approximation of eq. (23). The barrier lowering was modeled in [51] using the Schottky limit for emission over barrier, given by ox ox 3 4 Eq πε =∆Φ , (143) Other quantum effect at high electric field is that the centroid of the inversion layer is moved 1.5-2.5 nm in the depth of semiconductor, as shown with ∆x q in Figure 26, which is 2-3 times deeper than for the case of using Poisson equation alone, without considering quantum effects. Similar effect due to depletion of polysilicon gates occurs at the gate side [151], shown with ∆x d ~2mn in Figure 26. There are consequences for ∆n model of 1/f noise in MOS transistors. The first consequence is that ∆x q and ∆x d result in potential drops. For ∆x q the increase of the surface potential is V1 cm103 'n x'n q 213 Si q s × ≈ ε ∆ =ψ∆ ~0.1-0.3V in strong inversion. (144) The order of magnitudes is similar for ∆x d . Thus, the voltage across gate dielectric is approximately V)5.0... 2.0(VV~V1 cm102 'n VV x'n q5.1VV~ 5.1VVVVV TG 213 TG Si q TG sTGxdsTGox −− × −−≈ ε ∆ −− ψ∆−−≈ψ∆−ψ∆−−= (145) For transistors with thin oxide, e.g. nodes 90nm and below, the supply voltage is about 1V, V T ~0.2V, and ⅓ to ½ of overdrive voltage (V G -V T ) is “lost” in quantum effects and depletion of polysilicon gate. For overdrive in the range of 0.8V, the corresponding E ox =V ox /t ox <0.6V/3nm~2MV/cm, and the barrier lowering ∆Φ due to E ox is in the range 0.2-0.25eV, according to eq. (143). This is about 6% barrier lowering for electrons at SiO 2 interface, which would result in small increase of 3% for the tunneling attenuation distance λ, according to WentzelKramer-Brillouin (WKB) approximation by eq. (23). So, the 1/f noise S VG ∝ λ due to ∆n fluctuation would not be affected hardly by barrier lowering, when the gate dielectric is SiO 2 . Taking an average number from Table 2, the same calculation for HfO 2 suggests a barrier lowering of about 0.25eV/1.25eV~20%, or increase of 10% for S VG ∝ λ. On the other hand, the barrier lowering is found pronounced in thin dielectrics, suggesting that WKB approximation is not accurate for these dielectrics [164], and the noise in MOS transistors with high-k dielectrics is relatively high. We see here an open question on how to implement the barrier lowering into ∆n fluctuation model for 1/f noise in MOS transistors. The use of constant value for tunneling attenuation distance λ estimated from WKB in ∆n fluctuation model is not precise when the physical oxide thickness is less than 5nm. Also, for
54 of 286 the case of stacking several dielectrics in the gate insulator, the assumption for one value for λ is rough. A suggestion is given in [49] for how one can modify the ∆n fluctuation model when two materials are used in the gate insulator stack. The suggested modification uses weighting functions A and B, and it is given by ( ) 2t22ox1ox1t11ox eff t N)t,t,f(BN)t,f(AN λ+λ=λ , (146) where for each dielectric layer in the stack, λ 1 and λ 2 are the corresponding tunneling attenuation distances, and N t1 and N t2 are the corresponding oxide trap densities. The functions A and B depend both on the thicknesses t ox1 , t ox2 of the dielectric layers and the frequency f. More details will be given shortly in section IV.2 when discussing the gate dielectric profiling. The second consequence for ∆n fluctuation model, owing to ∆x q and ∆x d , is that the gate capacitance is bias dependent and it is less than the capacitance C ox of gate insulator stack. Thus, the derivation for S VG based on eqs. (139), c.f. eqs. (89), (98), (101), (104) and (109), and the assumption R=n’/(n’+n*) in eq. (123) become approximate, since the depletion capacitance of polysilicon gate and the “capacitance” arisen from the distance of the inversion layer centroid are not very large (ε Si /1nm~10µF/cm 2 ) when compared to the gate insulator capacitance larger than 1µF/cm 2 . This consequence also reflects in the next issue. Issues with condition vi. This condition for ∆n models of 1/f noise in MOS transistors states that the charge captured in the oxide is approximately at the semiconductor-dielectric interface. The distance x ti from interface, where the charge is trapped, is negligible as compared to dielectric thickness. This condition is clearly stated in [56], where also is given that t ox tiox N t xt nδ − =δ , (147) where t ox is the physical thickness of the gate insulator, x ti is the distance from semiconductor-insulator interface to the position of the trap in the oxide, δN t is variation in the number of trapped charges at x ti and δn is the change in the number of carriers in the channel of MOS transistor. Eq. (147) arises from Coulomb (“image charge”) balancing of the trapped charge between capacitances from the position of the trap to inversion layer and gate, where the charge is “mirrored”. Consider that the charge is trapped between the two dielectric layers in the gate stack in Figure 26, where the right-hand layer is SiO 2 interfacial layer with thickness x ti and permittivity ε SiO2 , and the left-hand layer is HfO 2 with thickness (t ox −x ti ) and permittivity ε HfO2 . The capacitance (per unit area) C xti from the trap to the inversion layer is 2 SiO ti Si q SiO ti xq,SiSiOxti cmF10 1 x x x C 1 C 1 C 1 222 µ + ε ≈ ε ∆ + ε =+= , (148) including the capacitance C Si,xq due to the displacement of the inversion layer centroid from the semiconductorinsulator interface. The corresponding capacitance to the gate conductor is 2 HfO tiox Si d HfO tiox xd,SiHfOgti cmF10 1 xtxxt C 1 C 1 C 1 222 µ + ε − ≈ ε ∆ + ε − =+= , (149) including the capacitance C Si,xd due to depletion of polysilicon gate. From the Coulomb balancing, we have
55 of 286 gti g xtigtixti t C nq C nq CC Nq δ = δ = + δ . (150) The fluctuation in the channel charge, therefore, is ttxtit gtixti xti NNRN CC C nδ≤δ=δ + =δ , (151) where R xti has the same meaning as R in eqs. (123) and (124) of coupling between oxide trap and channel charges, but R xti depends on the position of the oxide trap, and therefore, from the time constant τ of the trap, since τ=τ o exp(x ti /λ), according to eq. (99). For the example taken above, we have 2 2 2 2 2 2 110 1 1 110 HfO ox ti gti xti xti xti gti gti xti HfO HfO ox ti SiO t x C CF cm RC C C C t x F cm ε − + µ = = = + + ε ε − − + εµ , (152) which reduces to the expression in the brackets of eq. (147), neglecting the depletion in polysilicon gate and quantum effects in the channel and taking uniform dielectric (ε HfO2 = ε SiO2 ). The actual expression for R xti is more complicated, if one considers that the trap is not at the boundary between dielectrics in the gate insulator stack. Also, R xti is a function of bias and tunneling time constant. For the simple case of uniform dielectric and depletion in polysilicon gate and quantum effects neglected, one has ( ) ( ) ox o gtixti xti xti t ln 1 CC C Rττλ −= + =τ , (153) and R xti has to be included in the evaluation of the integral for superposition of Lorentzian spectra, since it multiplies R in the general expression, c.f. eq.(124) of the unified 1/f noise model for MOS transistor. Thus, ( ) = τ τ τπ+ τ µλ ≈ D S max o D V 0V 2 xti 2 2 D I dv f21 dR4 'n R *N L kTqI S , (154) but the integral in the square brackets is solved analytically only for the case R xti =1, to the best of our knowledge. Provided that the position of x ti ~2 nm of the slow traps inside the oxide is a significant portion of thin oxides t ox ~3-5nm, and also the depletion of polysilicon gate and quantum effects are not explicitly considered in the coupling parameters of the unified model, then an issue arises that these has to be included, in order to preserve the physical consistence of the model when applied to aggressively down-scaled MOS transistors. Conclusions to issues (i)-(vi). Owing to the above issues, the accuracy of the ∆n-∆µ model for 1/f noise in modern transistors with thin and stacked gate dielectrics is not very large. However, when carefully calibrated, the model is convenient for compact modeling, circuit simulation, and it is vital in providing information for comparisons in a qualitative manner for the properties and quality of gate stacks. One interesting class of techniques based on ∆n model for characterization of the oxide trap profile is now discussed.
56 of 286 IV.2. Charge trap profiling of gate dielectrics The ∆n model for 1/f noise in MOS transistors, see eqs. (104) and (105), uses approximation with uniform distributions, both in energy and space, for oxide traps and interface states. Such distributions result in g(τ)∝1/τ distribution for the time constants of the traps, e.g. eq. (100) for oxide traps, and produce 1/f noise from superposition of the Lorentzian spectra of the fluctuation of the individual traps – see again eqs. (72) to (75) and Figure 19. The deviation of noise power spectrum density from 1/f is used to evaluate the departure from uniform trap distribution, and in this way, provides profiling of the traps, by means of energy or distance, since either of them modifies the 1/τ distribution for the time constants of the traps [73]. The duality of capture energy and tunneling distance is addressed in [147] and earlier by eqs. (102) and (103), and the separation of spatial and energy profiles of the traps is made at assumptions for physical consistence at pre-determined characterization model [73]. IV.2.1. Spatial profiling of trap density Simple approach The simplest approach for spatial profiling of trap density N t (x ti ) in the oxide depth at distance x ti from semiconductor-insulator interface is used in [140] for a MOS transistor with gate stack of 2.1nm SiO 2 interfacial layer at semiconductor interface and 5nm HfO 2 on top of it. The assumption is that at given frequency f=f i , the traps with time constant τ i =1/(2πf i ) are dominant in the 1/f noise S(f), since the Lorentzian spectrum of these traps has maximum contribution in the quantity f i ×S(f i ). At this assumption, one writes from eqs. (99) and (101) that τπ λ= π =τ= λ τ oi ti i i ti o f2 1 lnx f2 1 x exp , with (2πf i τ o )<<1, (155) ( ) ( ) ( ) λ = λ = iVi 2 ox tittit 2 oxi V fSf kT WL q C xNxN WL kT C q f 1 S G G , (156) and taking λ≈constant≈0.1nm, both x ti and N t (x ti ) can be obtained, as shown in Figure 27 for the abovementioned MOS transistor. Note that the oxide trap profiling by using this simple approach is qualitative, as mentioned in [140], since details are not elaborated when N t is not uniform, when two materials with different tunneling attenuation distances λ SiO2 ≠λ HfO2 are used, and when the overlap in the spectra of traps with different time constants is neglected. The later caused that the increase of N t in Figure 27 is detected at x ti =1.8nm rather than at 2.1nm, where the interface SiO 2 -HfO 2 was. This is similar to the broadening of step changes in trap densities when measured by charge pumping techniques [147]. In the origin of this broadening is the integral form that describes the superposition of Lorentzian spectra of individual traps with similar time constants, c.f. eqs. (72) to (75). It is shown in [68] that the superposition integrals are equivalent to a convolution between the trap distributions in space and energy, since both distributions affect the distribution of trap time constants, which is used as integration variable. The effect of the convolution is that the slope of 1/f noise in respect to the frequency may vary, rather than a step in the level of 1/f noise magnitude at particular frequency to be observable [2, 56, 68]. Mathematically, the reason is that the function fτ/[1+(2πfτ)²] is not sharply peaking at 2πf=1/τ, in order to produce a step when integrating it. To demonstrate the problem, we adopt the approach from [49], as follows. Gate stacks
63 of 286 IV.3. RTS noise in MOS transistors As the size of MOS transistors is becoming smaller and smaller, then the charge capture and emission process by individual traps is increasingly distinguishable. This process results in a random bistable fluctuation in time domain, such as Random Telegraph Signal (RTS), and therefore the random bistable fluctuation is also called RTS noise. The RTS noise in the drain current of MOS transistors is usually measured and then referred to the gate terminal as a voltage by using the transconductance g m of the transistor [74]. It is cumulatively observed that the amplitude of the gate referred voltage from individual RTS noise matches well with addition and removal of one elementary electronic charge q to the gate oxide capacitance [74], that is, ox G m D WLC q C q V g I==∆= ∆ . (178) IV.3.1. Time constants of RTS noise The time constants of the two states of individual RTS match with the predictions of Shockley–Read–Hall theory for generation-recombination process [68, 74, 107, 112, 148, 166]. Consequently, the superposition of larger number of individual RTS, each of which having a Lorentzian spectrum, is found to coincide with 1/f noise in larger-area MOS transistors [68, 74, 112, 148]. The above findings are made over a period of about 50 years and has been reviewed in [68] at the time of entering the sub-micrometer technologies. In studies of MOS transistors with very thin oxides, RTS noise is observed also in the gate leakage current [148]. The above summary of findings implies a coherent picture for RTS noise in MOS transistors. However, looking at the details or individual measurements, one will observe significant deviations and will meet with difficulties to manage long time records, and to link them to spectra and models for noise. In fact, there is no standard procedure for analysis and compact model for RTS noise in MOS transistors. The RTS noise is also nonmonotonically biasand temperature dependent [58, 68, 137, 161, 167], it varies between different time records captured from one sample [68, 74], and between nominally identical samples [137, 166]. On the other hand, the amplitudes of RTS noise can be large in sub-micron area MOS transistors [74, 112, 137] and RTS noise becomes important issue that the designs have to overcome, e.g. in CMOS imagers with correlated double sampling [162]. Generally, three parameters describe the individual RTS noise component. For MOS transistors, these are amplitude ∆I D (or ∆V G ), and capture and emission time constants of the trapping GR center, respectively. The studies in [68, 74, 107, 112, 148, 166] confirmed that the time constants follow very well the Shockley–Read– Hall (SRH) statistics, which established eq. (63). In most cases, the emission time constant τ e is a weak function of the biasing (it increases when increasing V G -V T ), whereas the capture time constant τ c decreases when the biasing (V G -V T ) increases, owing to the increase of the carrier concentration n’ in the channel of the MOS transistors. There is a possible exception in sub-threshold regime of operation of the MOS transistor when charges from the bulk are trapped, and the bias dependence of the time constants can be inverted for this case [161]. There are experimental difficulties by obtaining the values for time constants since a processing of long records captured in time domain is necessary, especially when two or more RTS are present simultaneously [68, 74, 112], but overall there is no other significant issue with the RTS time constants. When the values are obtained, they match well the predictions of the theory and with the corner frequency of Lorentzian spectrum related to RTS, and given by
64 of 286 ( ) ( ) 2 2 D I f21 F1FI4 S D τπ+ τ−∆ = or ( ) ( ) 2 2 G V f21 F1FV4 S G τπ+ τ−∆ =, with ( ) ( ) 2 ce ce ce F-1F and 111 τ+τ τ τ = τ + τ = τ. (179) where F is the trap occupancy factor, given by Fermi-Dirac statistics. Therefore, we do not extend the discussion for the time constants of RTS further. IV.3.2. Amplitude of RTS noise There are some discrepancies between values for the RTS amplitudes ∆I D and ∆V G , measured in time domain, estimated from spectrum using eq. (179) and predicted by eq. (178). We focus on this issue, because the amplitudes of RTS noise will increase as the device size decreases with every next generation of technology nodes of minimum feature gate length L min , since RTS due to individual traps will dominate the low-frequency noise and, according to eq. (178), ∆V G and ∆I D will increase as 1/(WL)∝1/L min ², because C ox cannot be increased proportionally furthermore [3]. The characterization of MOS transistors usually uses eq. (178) rewritten in normalized form for the RTS current, given by oxD m D m G D m D D WLC q I g C q I g V I g I I==∆= ∆, (180) where ∆V G =q/(WLCox)=∆V FB is regarded as modulation flat band voltage V FB in MOS transistor owing to trapping of single electron at semiconductor-dielectric interface [58, 156]. This equation corresponds to the generic expression for coupling of noise, being a square-rooted version of eq. (1) with coupling coefficient √K=1. The direct application of this equation against measured data usually shows some discrepancies, as illustrated in Figure 29. One can see in the figure that the overall evolution of the RTS amplitude is captured by eq. (180), but the amplitude is underestimated at low bias, whereas it is overestimated at high bias. This implies that the coupling between trap and conduction layer involves also other factors, and the coupling coefficient √K≠1 is also bias dependent. The reasons of this discrepancy correspond to the issues with the ∆n model for 1/f noise in MOS transistors, since the origin of the models is the same. Correlation between mobility variation and trap occupancy One effect neglected in eq. (180) is the mobility variation correlated to the trap occupancy. This correlation was investigated in [58] for RTS amplitude, considering Coulomb and phonon scattering at lower and higher biasing of MOS transistor. The essence of this investigation is that, increasing the bias from weak to strong inversion, the channel carrier mobility µ increases at low bias, since µ is limited by Coulomb scattering [see eqs. (134) and (135)], whereas the channel carrier mobility µ decreases at high bias, since µ is limited by phonon and roughness scattering [see eqs. (131) and (133)]. Consequently, via the surface potential, at low bias the derivative ∂µ/∂∆V FB >0 and adds to the change ∂n’/∂∆V FB >0 in carrier concentration n’ in the MOS channel, causing the RTS amplitude to be higher than that given by eq. (180), whereas at high bias the derivative ∂µ/∂∆V FB <0 and subtracts from ∂n’/∂∆V FB >0, causing the RTS amplitude to be lower than that given by eq. (180). At a crossover bias (V G,cr , I D,cr , g m,cr ), when ∂µ/∂V G ≈0, eq. (180) matches the measurement. Noticeably, we can observe qualitatively exactly the same behavior in Figure 29, as well as in fig. 8 in [107], and to the first order of approximation we can suggest an empirical expression for coupling coefficient √K µ for correlated mobility contribution to RTS amplitude that is in the form of eq. (121) for the coupling coefficient for 1/f noise in MOS
65 of 286 transistors with correlated ∆n-∆µ fluctuation. −θ+≈ µcr,m cr,D m D g I g I 1K → µµ =∆= ∆K WLC q I g KV I g I I oxD m FB D m D D → µ ∆=∆ KVV FBG (181) For physical justification, the ratio I D,cr /g m,cr ∝C ox (V G,cr −V T )∝n’ cr corresponds to carrier concentration n’ cr in the channel at which crossover between Coulomb and phonon scattering occurs and ∂µ/∂V G ≈0. Position of the trap along the channel and depth of the trap in the oxide A second detail neglected in eq. (180) is the position of the trap in the oxide and in the channel. Position of the trap along the channel. Considering the channel conduction at different spots under the gate, it is deduced in [167] that the RTS amplitude is function not only on channel carrier concentration n’, but also on the lateral electric field at the position of the trap along the channel. Assuming that one carrier is trapped at a particular spot in the conductive channel of the MOS transistor, the relative RTS amplitude is then given by [167] n 1 E E G G I I 2 avgavg 2 D D D D µ µ = ∆ = ∆ , (182) where G D =I D /V D is the channel conductance between drain and source terminals, ∆I D and ∆G D are the RTS amplitudes, µ and E are local carrier mobility and lateral electric field in the channel at the spot of charge trapping, µ avg and E avg are average values for carrier mobility and lateral electric field carrier, and n is the total number of carriers in the channel, respectively. By inspection of MOS transistor equations, it can be shown that 1/n≈g m q/(I D WLC ox ). Therefore, eq. (182) can be rewritten in the form of eqs. (180) and (181), as oxD m 2 avgavg 2 oxD m E D D WLC q I g E E WLC q I g K I I µ µ == ∆, with 2 avgavg 2 E E E Kµ µ ≈ (183) being a coupling coefficient corresponding to variation of the lateral electric field E in MOS channel. Obviously, the lateral electric field at the source side of the channel is lower than E avg and traps located at this side will produce RTS with amplitude lower than that predicted by eq. (180). In contrary, the lateral electric field at the drain side of the channel is higher than E avg and traps located at the drain side will produce RTS with amplitude higher than that predicted by eq. (180), especially when the MOS transistor is in saturation mode of operation. Perhaps this is the reason why the RTS amplitudes from different traps scatter in one sample and between identical samples, as illustrated in Figure 30 from [68]. In this figure, 58 different RTS are shown from 12 nominally identical nMOS transistors biased at the same condition (V G −V T )=(3.8−1)V, V D =0.1V, V S =0, I D =6.7µA. The transistors are from 2.5µm technology node, but the mask was W=L=2µm being below the minimum feature size. From DC measurements, it was estimated that the effective size of the samples was W=0.5µm and L=0.75µm, but tolerances are not reported in [68]. The circles in Figure 30 present the relative RTS amplitudes ∆I D /I organized in [Fig. 16 in 68] in a scatter plot sorted by increased values of ∆I D /I D and formally numbered from 1 to 58, as shown in the bottom horizontal axis of Figure 30. We assume that the traps are uniformly distributed along the channel length L, so number 0 corresponds to source edge of the channel and number 59 corresponds to drain edge of the channel. In this way,
66 of 286 we scale the top horizontal axis of Figure 30, using z/L=(Number of RTS)/59, where (z) represents the position of the trap along the channel length L. Next assumption is that the carrier concentration in the channel is nearly constant, since (V G −V T )≈2.8V, while (V D -V S )=V D =0.1V and the transistors were operating deep in ohmic regime. Therefore, we expect linear increase of lateral electric field E from source to drain with an average value of E avg =V D /L≈1.3kV/cm. According to eqs. (182) and (183), the square root avgDD EEII ∝∆ , and by multiplying with the value above for E avg , we obtain the lateral electric field E for each data point, as shown with squares in Figure 30, which are fitted with the linear function E∝z/L. This linear-fit function was used in eq. (183) to calculate ∆I D /I D ∝(E/E avg )², as shown with thick line passing through the circles in Figure 30, estimating also C ox ≈95nF/cm² and EOT≈37nm, which are not stated in [68], but are reasonable for 2.5µm technology node. Apart from 4 data points of high RTS amplitudes on the right side of Figure 30, the agreement between measured (circles) and calculated (thick line) from eq. (183) values for the relative RTS amplitudes ∆I D /I D is good, despite the many assumptions stated above. The agreement leads to the conclusion that the variation of RTS amplitude in MOS transistors is vulnerable to the position of the trap along the channel, since the lateral electric field is different at different positions. The 4 data points that deviate from the rest of the data in Figure 30 are probably from one of the 12 devices. The sizes of this device probably deviate 20%-30% from the nominal values, which is reasonable for the case of transistors with effective sizes L and W being about 1/4-1/5 of the minimum feature size of the technology node. Unfortunately, the information in [68] is aggregated among all 12 samples, and we cannot inspect the details further. The analysis above for relation between trap position along the MOS transistor channel and amplitude of RTS noise is in general agreement with other works. It is argued in [133] that the flat band perturbation, which is in the origin of eq. (180) via ∆V G =q/(WLC ox )=∆V FB , neglects the local perturbation in the surface potential and the associated charge carrier density at a particular position in the channel, and thus underestimates both the RTS amplitude and 1/f noise as a superposition of RTS, when the transistor operates in strong inversion and saturation, i.e. inversion layer charge densities at source and drain sides are n’ S >2φ t C ox /q and n’ D <<n’ S . At this condition, as mentioned earlier, the 1/f noise magnitude is underestimated about 2 times by the flat band perturbation model. For RTS amplitude, the difference can be larger, and this was used to locate positions of traps in the channel in [166]. The methods in [166] require knowledge for details in the transistor and employ formulation in terms of surface potential and numerical iterations. This approach, although accurate, is too complicated for the purpose of experimental noise characterization, in which the measured data scatter, and in this way, do not allow for obtaining very precise modeling. The advantage of the methods in [166] is that the position of the trap along the channel can be found simultaneously with the depth of the trap in the oxide. Now we discuss the latter. Depth of the trap in the oxide. Based on earlier works [168], it is suggested in [112] that the relative RTS amplitude ∆I D /I D can be given as −η= ∆ ox ti oxD m D D t x 1 WLC q I g I I, (184) where t ox is the physical thickness of the gate insulator, x ti is the distance from semiconductor-insulator interface to the position of the trap in the oxide. The factor (1−x ti /t ox )=(t ox −x ti )/t ox is a coupling coefficient √K ti and it is the same as in eq. (147). The parameter η corresponds here to the product of coupling coefficients √K µ √K E in eqs.
67 of 286 (181) and (183) above, but it is left as a fitting parameter in [112]. Therefore, in order to obtain information for (t ox −x ti )/t ox , the ratio of the capture τ c and emission τ e time constants is used in [112], since τ c and τ e depend on x ti , see eq. (112), and on bias via Fermi level, see eq. (63). The explicit expressions are given in [166], but the overall bias dependence is [112] ( ) G s ox ti tox ti tG ec Vt x 1 1 t x 1 V ln ∂ ψ∂ − ϕ − ϕ −= ∂ ττ∂ , (185) where ϕ t =kT/q≈0.026V is the thermal voltage at room temperature T=300K, V G is the gate bias voltage and ψ s is the surface potential in the MOS channel. The surface potential ψ s is a complicated function of V G (and other biasing voltages applied to MOSFET), and the precise evaluation of the derivative ∂ψ s /∂V G is inconvenient and not necessary when the spread in the values for τ c and τ e is large, which is usually the case of noise measurements. Therefore, for practical cases of characterization, one can obtain an approximate value for the derivative ∂ψ s /∂V G , using several general relations in MOS transistors from [169], as follows, in which inversion charge Q inv , oxide capacitance C ox and depletion capacitance C d are given per unit area. Assume that the measurement was when the MOS transistor operated either in weak (sub-threshold) or in strong (above threshold) inversion regimes, that is |V G -V T |>0.1V. In weak inversion, when V G is below the threshold voltage V T , the inversion charge Q inv is negligible and the gate bias is spread across the series connection of oxide capacitance C ox and depletion capacitance C d . Therefore [pp.75 and 86 in 169] 1.5 ... 1constant C C 1 2 1 V ox d s s G ≈≈+=α= ψ γ += ψ∂ ∂, in weak inversion, (186) neglecting the interface states’ capacitance C it . The parameter α can be obtained from sub-threshold slope ∂V G /∂log 10 (I D )=2.3αϕ t [p.175 in 169] of the transfer characteristic I D (V G ). The parameter γ is the body bias coefficient and can be obtained from ∂V G /∂V B , where V B is a bias voltage applied to the bulk (body) under the MOS transistor channel. However, γ is not needed in this analysis. At the same time, the ratio g m /I D between transconductance g m and drain current I D in MOSFET in weak inversion also depends on α, and it is given by [p. 173 in 169] tD m 1 I g ϕα = (187) So, from eqs.(186) and (187), we get that 1 ... 6.0constant I g 1 V D m t G s ≈≈ϕ= α = ∂ ψ∂ , when the MOSFET is in weak inversion regime. (188) For strong inversion regime, according to [p.87 in 169], we write ( ) ( ) tinv G inv s inv inv G s G 2 1 Q V Q Qln Qln VV ϕ∂ ∂ = ψ∂ ∂ ∂ ∂ = ψ∂ ∂ (189) G inv inv t G s V Q Q 2 V∂ ∂ ϕ = ∂ ψ ∂ (190)
68 of 286 in linear (ohmic) mode, when (V D -V S )<<(V G -V T ), and where V DS =(V D -V S ) is drain-source bias voltage and V T is threshold voltage of MOS transistor. Also, the inversion layer charge Q inv is approximately constant along the channel, and it is ( ) W LI VVCQ D TGoxinv µ =−≈ . (191) where µ is the carrier mobility in the transistor channel of length L and width W. So, eq. (190) becomes G D D tD G D t G s V I I 2 W LI V W LI 2 V∂ ∂ϕ = µ∂ ∂ µ ϕ = ∂ ψ∂ , or (192) TG t D m t G s VV 2 I g 2 V− ϕ≈ϕ= ∂ ψ ∂ , in strong inversion and ohmic mode. (193) To obtain an expression for ∂ψ s /∂V G in at strong inversion regime saturation mode, we look closer at the charge sheet model for the drain (to source) current I D with neglected diffusion, given by [p.157 in 169] ( ) 2D,inv 2S,inv ox D QQ C2L W I− α µ = , (194) in which the inversion layer charge densities Q inv,S and Q inv,D are at the source and drain sides of the channel, respectively, and Q inv,S and Q inv,D are given with the bias voltages V S and V D applied to these sides, as [p.121 in 169] ( ) ( ) .0Q otherwise ,VV if ,VVCQ and ,0Q otherwise ,VV if ,VVCQ D,invPDDPoxDinv, S,invPSSPoxS,inv =<−α= = < − α = (195) The obvious approximation for the pinch-off voltage V P in the equation above is given by [p.158 in 169] α − ≈ TG P VV V . (196) The transistor is in saturation mode, if V D >V P , and in ohmic mode, if V D <V P . Note, I D , Q inv,S and Q inv,D depend on V G only via V P , and α is assumed constant, because the depletion capacitance C d in eq.(186) does not change significantly with V G , especially in strong inversion. In saturation mode, (V D >V P ) and Q inv,S >>Q inv,D ≈0. Therefore, eq. (194) is reduced, and using V S =0, which is the normal case when the body and source of MOSFET are tied together, one gets ox D S,invinv C2 W L I QQ α µ =≈ , (197) W L I Cox2 2 g V I W L I Cox2 2 1 Cox2 W L I VV Q D m G D D D GG inv µ α = ∂ ∂ µ α = α µ∂ ∂ = ∂ ∂ , (198) and inv D mD D m G inv Q I2 g Cox2 W L I I2 g V Q=α µ = ∂ ∂ . (199)
69 of 286 This means that one can substitute in eq. (190) the quantity D m G inv inv I2 g V Q Q 1= ∂ ∂ , (200) and to obtain TG t D m t D m t G s VV 2 I g I2 g 2 V− ϕ≈ϕ=ϕ= ∂ ψ ∂ , in strong inversion and saturation mode. (201) By comparing eqs.(188), (193) and (201), one can summarize that D m t G s I g Vϕ≈ ∂ ψ ∂ , in any mode of operation of MOSFET, (202) 1 ... 6.0 V G s ≈ ∂ ψ ∂ , in weak inversion regime (sub-threshold, V G is below V T ), (203) and ∂ψ s /∂V G gradually decreases at gate biasing above threshold, as TG t G s VV 2 V− ϕ = ∂ ψ ∂ , in strong inversion regime (above threshold, V G above V T ). (204) The accuracy of eqs. (202), (203) and (204) is within a factor of 2 (6dB), which is sufficient for analysis of RTS noise measurement data, which scatter usually more. So, eq. (185) can be rewritten to useful for experimental characterization forms of ( ) D m ox ti ox ti tG ec I g t x 1 t x 1 V ln −− ϕ −= ∂ ττ∂ , in any mode of operation of MOSFET, or (205) ( ) + ϕ −≈ α − ϕ − ϕ −= ∂ ττ∂ ox ti tox ti tox ti tG ec t x 21 11 t x 1 1 t x 1 V ln , in weak inversion, and (206) ( ) TGox ti ox ti tG ec VV 2 t x 1 t x 1 V ln − −− ϕ −= ∂ ττ∂ , in strong inversion. (207) The equations include parameters that can be easily obtained from DC measurements of the MOSFET transfer characteristic (V T , g m , α from sub-threshold slope). So, once the capture and emission times are characterized from RTS noise measurements at several bias points for V G , then the derivatives in the left-hand side of the equations can be found by a simple fitting of slopes in semi-log plot of ln(τ c /τ e ) vs. V G . Then, by substituting in one of eqs. (205), (206) or (207), the ratio x ti /t ox of the distance x ti from the trap to semiconductor-insulator interface to the physical thickness t ox of the gate insulator can be estimated, and the coupling coefficient √K ti for the trap position in the oxide can be found as −= ox ti ti t x 1K . (208) Note that the above equations are appropriate for approximate experimental characterization and volume tests when the scattering in the data is large. If a precision analysis of few samples is required, one should use rigorous models, such as this in [166], which can also locate the position of the trap along the channel of
70 of 286 MOSFET. The price is, of course, that the surface potential ψ s has to be obtained prior to analysis of the oxide trap, and this requires specific information, numerical simulations and optimizations, which might be not always available, or affordable, owing to time, expertise or other constraint in the experimental practice. Amplitude of RTS (concluding remarks). To summarize, the analysis of amplitude of RTS noise in MOS transistor requires large number of waveforms captured in time domain, from several samples and various biasing. The waveforms need to be processed so that the evolution of the amplitudes and time constants of RTS from individual traps with biasing and over the samples is observed. Then, the data have to be fitted to n 1 K WLC q I g K I I oxD m D D == ∆ , (209) where the coupling coefficient √K=√K µ √K E √K ti accommodates several dependences, and n is the total number of carriers in the channel of the MOS transistor, with n≈WLC ox (V G −V T )/q in the ohmic regime of operation of the MOS transistor at V D <<(V G −V T ), and twice smaller n≈½WLC ox (V G −V T )/q in the saturation regime of operation of the MOS transistor at V D ≥(V G −V T ). The coupling coefficient √K µ takes into account for the contribution of mobility variation, √K µ is given by eq. (181) and the parameters in this equation should be chosen so that the modeled and measured relative RTS amplitudes ∆I D /I D are proportional when the gate bias voltage V G is varied. As the initial values in eq. (181), one can use the first order mobility degradation coefficient θ~0.3…1 V − 1 and the ratio I D,cr /g m,cr ∝C ox (V G,cr −V T ) of the current to transconductance at gate bias corresponding to maximum mobility in strong inversion regime (V G,cr above threshold V T ). The coupling coefficient √K E takes into account for the variation of the lateral electric field and carrier velocity in the channel of MOS transistor, and √K E is given by eq. (183). If data from many samples measured in strong inversion and ohmic modes are available, then the value for √K E can be obtained from a scatter plot of relative RTS amplitudes ∆I D /I D as discussed above, and the position of the traps along the channel length can be estimated. Otherwise, knowledge for the lateral electric field and mobility has to be provided by other means, e.g. simulation of structure, in order to obtain surface potential along the channel [166]. Other approaches are also possible, e.g. swapping drain and source, to discriminate between traps located closer to source or closer to drain sides of the channel – see references in [166]. The coupling coefficient √K ti takes into account for the depth x ti of the trap in the gate insulator of thickness t ox . At assumption for tunneling mechanism for charge exchange between the trap and channel carriers, the ratio x ti /t ox , can be estimated using eqs. (205), (206) and (207) and DC parameters of MOS transistor from the gate bias dependence of capture τ c and emission τ e time constants of individual RTS, by the slope of ∂ln(τ c /τ e )/∂V G . Then, √K ti = (1−x ti /t ox ), as given by eq. (208). The advantage of the procedure above is that it is based only on DC measurements and waveform captures, and device simulation is not needed. However, the procedure requires large number of measurements and extensive processing of waveforms by establishing relation between data from different measurements and keeping track of the evolution of RTS parameters of individual traps. Nevertheless, tools and methods for semi-automated extraction of RTS parameters are reported (in [74, 107, 112, 166] using correspondence between ∆I in time domain and low-frequency plateau of Lorentzian spectrum in frequency domain, in [58, 74, 107] using histograms of drain current, in [28, 29] using discontinuity of waveform), although they should be further
71 of 286 organized to make the tests feasible for applications in the volume production in semiconductor industry. The above analysis of RTS noise is based on charge trapping. Alternative suggestion for RTS noise from mobility fluctuation is given in [170] for diodes, at condition when I<∆I<q/α H τ where τ is minority carrier life time and α H is the Hooge parameter – see eq. (6). Situation I<∆I of DC current smaller than RTS amplitude has been observed in carbon nanotube pMOS-like field-effect transistors [171]. IV.4. Figures of merit for MOS transistors The intensive research on MOS devices resulted in many figures of merit (FOM) that have been used to compare different technologies and explore the scaling of these devices. Each FOM was suggested in order to emphasize particular feature of the devices or circuit, which used these devices, or to facilitate a model or design. Consequently, the values from different FOM became difficult to compare each to other. In this section, we will discuss several FOM that are commonly used for the low-frequency noise in MOS transistors in attempt to relate the different FOM each to other. IV.4.1. Definition To set up the discussion, first it is helpful to state what FOM is and how it differs from device parameters or physical quantities. In principle, FOM is a customized expression that combines several device parameters and physical quantities according to particular model or targeting particular application. For example, the DC value I D of the drain current in MOS transistor and the power spectrum density S ID of drain current are physical quantities, but the normalized noise S ID /I D ² is a figure of merit for the ratio noise to DC in the transistor, and S ID /I D ² may (or may not) vary with frequency and bias, depending on what noise in a transistor is addressed. If the flicker noise is the concern, then one assumes a model with 1/f scaling rule for S ID and can evaluate the SPICE parameter K F =fS ID /I D ² according to eq. (2), but K F is also a FOM, since it is derived from another figure of merit, by using additional scaling rule for the frequency dependence of S ID and the 1/f dependence is canceled in K F just for convenience. Certainly, K F is a good figure of merit for flicker noise in BJT, since in most cases the 1/f noise is coupled from IFO via the transconductance g m , as discussed in section III.2. “Differences in BJT fabrication, IFO”, and g m /I C ≈1/φ t ≈constant in wide range of biasing conditions. However, K F is not the best choice for MOS transistors in strong inversion regime, since g m /I D ∝1/(V G -V T ) and the bias dependence of 1/f noise is not cancelled in K F . On the other hand, if the shot noise in BJT is addressed at higher frequencies, then the appropriate figure of merit for normalized noise is S IC /I C ~2q, which is frequency independent in principle, but not K F . Moving further to RF range, one usually uses the so-called “Noise Figure” or “Noise Factor”, which is a completely different FOM in its basis, and RF Noise Figure is a function of the ratio between device and thermal (Nyquist) noise at certain conditions for impedance matching. Thus, the different FOM depend on device, models and ranges, and one should explicitly state the expressions, the origin and the purpose of the normalization used for particular case of interest, since all normalizations are FOM, they are valid at certain conditions, and even the most popular FOM have counterparts. IV.4.2. Input and output referred noise – scalability of normalized noise MOS transistor In both cases, the word is for the 1/f noise S ID in the drain current I D , which is at the output terminal (drain) of the MOS transistor, and it should be clearly stated that noise S IG in the gate leakage I G of MOS transistors with
72 of 286 ultra thin oxides is not included in the input and output FOM, and it is separately analyzed, if present, with one exception in [110]. The practice is that output noise S ID is referred to the input (gate) terminal as a “gate voltage” S VG by using the most general expression for coupling from input to output via transconductance g m , which is GD V 2 mI SgS = , (210) and it follows from eq. (1). Consequently, S VG is a FOM, because it is a derived quantity, it cannot be directly measured as a voltage, but it is convenient, since it is weakly dependent on the bias of MOS transistor, it scales properly with the area WL of the transistor and it has been well explained in terms of oxide trapping, oxide capacitance, mobility degradation, etc., as discussed in previous sections. To compare different MOS transistors, the most popular FOM is the product WLS VG at low gate overdrive (V G −V T )~0.1V and frequency 1Hz, which is used in ITRS [3] – see eq. (106) for more details. At these conditions and at lower gate overdrive, the number fluctuation due to trapping at the oxide dominates, and S VG ≈S FB , where S FB is regarded as flat band voltage noise – see eq. (109). At higher gate overdrive, (V G −V T )>0.2V, the mobility fluctuation may significantly contribute to the noise. As explained in [156], if the mobility is correlated to the oxide trapping, then S VG becomes a quadratic function of (V G −V T )∝I D /g m , and an additional FOM=(S VG ) 0.5 vs. (V G −V T ) is used to obtain the carrier scattering coefficient [57, 89, 92, 139, 156] – see eq. (121) and the paragraph after it. If the mobility noise is not correlated to oxide trapping, then S VG ∝(V G −V T ) and the slope of this dependence is used to obtain the Hooge parameter α H , which follows from to eqs. (11) and (88), and it will be discussed further in IV.4.4. “Physical figures – trap density, Hooge parameter, scattering parameter” along with other FOM derived from S VG . The normalization of output noise is a FOM of form G D V 2 D m 2 D I S I g I S = . (211) It is widely used for MOS transistors as the intermediate step in analyses, such as for inspection of number fluctuation, or obtaining S VG and other quantities. However, this normalized noise is rarely used to obtain the SPICE parameter K F except for cases when comparing noise in MOS and BJT at similar biasing current targeting specific circuit application, e.g. low power RF oscillators [22, 23], BiCMOS circuits [83] and radiation resistance or sensitivity [102, 103, 104, 165]. Comparison BJT-MOS transistors As mentioned above, K F is an essential FOM for BJT, because it varies a little in wide range of biasing, but the problem is that K F in MOS transistors is a strong function of biasing above threshold voltage V T , and it is approximately a reciprocal function of gate overdrive (V G −V T ). Strictly speaking, comparisons based on K F in MOS transistors are valid only at fixed gate overdrive. The better way to compare MOS and BJT is to refer K F in BJT to the base terminal as a voltage noise S VB , and then to compare S VB and S VG as suggested in ITRS [3]. The conversion of K F into S VB is straightforward for BJT, because g m /I C ≈1/φ t ≈constant – see again eq. (3) and the discussion between eqs. (28) and (30), because
79 of 286 total noise S V = (S VG +S VSH +S Vth ) is the sum of all noise contributions to the input plane, as “seen” at the input of the amplifier (the gate terminal of the MOS transistor), and S V is shown with thick black lines in the left-hand plot of Figure 35 as the upper boundaries of the shaded areas. The shaded areas in this figure is the NF, and the values for NF [converted in dB – see again eq. (226)] as function of frequency f and R s are given in Figure 35 on the right-hand side by the contour plot, in which the dashed lines correspond to shaded areas in the left-hand plot of Figure 35. Increasing the signal source resistance R s , several observations can be made in the in the left-hand plot of Figure 35, in which the noise spectra are shown for three values of signal source resistance R s ={10kΩ,1MΩ,100MΩ}. - First, at low frequency, f<10Hz, the 1/f noise of the MOS transistor dominates in the total input noise S V , as along as R s <100MΩ, and S V is virtually unchanged, although NF decreases (shaded areas become smaller at higher R s ), because the reference thermal noise voltage increases. Thus, the reduction in NF at low frequencies at high R s does not imply reduction of noise, and since the noise level is constant, then the signal to noise ratio is also unchanged. At R s ~200MΩ, the thermal noise dominates in the total noise S V at 10Hz, and NF reaches local minimum, but S V begins increasing with R s . At higher R s , the shot noise takes over at low frequencies, and both NF and S V increase. For clarity in the figure, this situation is not shown. - Second, at high frequency, f>100kHz, the total noise S V changes a little with R s , since the thermal noise and shot noise are attenuated, owing to the low impedance of the parasitic capacitance C s at the input plane. Thus, S V is constant or even decreases with R s to the level of the white noise in S VG of the transistor, but NF increases, since the reference thermal noise is attenuated by approximately (2πfR s C s )². While a constant S V ~S VG set by the noise of MOS transistor is similar to the first case at low frequency, the signal to noise ratio at high frequency decreases with R s , since any signal from the signal source is attenuated in the same proportion (2πfR s C s )² as the thermal noise, which is different from the first case. - Third, at medium frequencies of few kHz, the 1/f noise of MOS transistor is lower, and the thermal noise S Vth from R s can reach levels above 1/f noise before being attenuated by the square of capacitance conductance (2πfC s ) 2 . In this case, the input noise S V ~S Vth follows closely the thermal noise, S V increases with R s , while NF reaches minimum values. The global minimum is known as Minimum Noise Figure NF min , and it is in the locus of the contour plot in the right-hand side of Figure 35. NF min is a conditional figure of merit, because it depends on biasing, frequency and impedance Z s . The condition for Z s suggests noise impedance matching, at which the noise from the transducer has the least relative contribution to the total noise, and NF min estimates this contribution in comparison to the thermal noise. From the contour plot in the right-hand side of Figure 35, we estimate NF min <1.02=0.1dB at R s ≈2MΩ and f=10 kHz. Any deviation from these conditions increases NF, as depicted in the contour plot in the right-hand side of Figure 35 with labeled arrows for the factors that cause increase in NF. The substitution in eq. (228) with the values for NF min and the corresponding R s implies that the noise resistance of the MOS transistor is R n <2MΩ×2%≈40kΩ, which is much lower than any electrical impedance of the gate terminal of the transistor at 10kHz, e.g. the impedance 1/(2πfC g )~20MΩ for the gate capacitance C g ~0.8pF in the example. Also, it should be noted again that the input referred noise level and SNR might be not at the optimum at the conditions for NF min , because the thermal noise also varies with the impedance Z s of circuit at its input plane, the low-frequency applications consider several frequency decades at low loading of signal source, thus do not care for power matching between source and load, and the SNR usually is not determined in respect to thermal noise,
80 of 286 since other types of noise dominate both in the signal source and in the amplifier. In RF applications, in contrast, the thermal noise is dominant in signal sources, the frequency range is “narrow”, e.g. few to 30 percent around the central frequency even in the so-called ultra wide band systems, the impedances are predetermined in a range between 30Ω and 600Ω, usually close to the characteristic impedances of 50Ω or 75Ω, and the impedance matching is very important in order to prevent from electromagnetic wave reflections, standing waves and to obtain power gain at high frequency. Therefore, NF is convenient and of high importance for RF applications. Definition of NF for RF applications. To meet the objectives mostly in applications such as low noise RF preamplifiers, another definition for NF is derived from the generic definition of eq. (226). First of all, the impedance matching between signal source impedance Z s and amplifier input impedance Z i are taken into accout and the reference thermal noise power is not the full power 4kT, but the available power from the signal source at matched condition Z i =Z s* , where Z s* is complex conjugated of Z s ; and the available thermal noise power from the signal source is [196] kT R4 kTR4 R4 S S s s s Vth avbl,th === , available thermal noise power from RF signal source, (230) where R s is the real component of Z s , and the noise voltage from the source is divided by 2, that is equally, between Z s and Z i at the matched condition Z i =Z s* . This reduction of reference power in RF noise figures to ¼ of thermal noise is clearly stated in [197], but rarely mentioned in recent publications. So, some RF noise figures are, as following. The input referred noise figure is T T 1 kT S 1 S S NF ee avbl,th avbl,in in +=+== , (231) where S in,avbl is the noise in the amplifier being referred as available power from signal source, S e =S in,avbl −S th,avbl is the excess noise added from the amplifier, also referred as available power from signal source, and T e =S e /k is the equivalent excess noise temperature, corresponding to S e . The standard reference temperature T o is 290K, and if the physical temperature is different, then a correction with ratio T/T o is made, as discussed in [197]. The excess noise ratio ENR, for example for noise sources, is given in respect to T o , and ENR=10dB×log10(S e /kT o )=10dB×log10(T e /290K). Since S in,avbl is input referred (actually, source referred), it does not exist as a physical signal that can be measured at the input of the amplifier. Therefore, one usually uses the output referred noise figure NF out , which is obtained from the input referred by multiplying the nominator an denominator of eq. (231) with the so-called Transducer Power Gain G T , and NF out is kTG S SG S SG SG NF T out avbl,thT out avbl,thT avbl,inT out === , (232) where S out is the measured noise power at the output of the amplifier, and G T S th,avbl is the reference value for the output noise power that corresponds only to the thermal noise from signal source. The transducer power gain G T is the ratio of the power delivered to load with impedance Z L in the output of the amplifier to the power available from the signal source with impedance Z s at the input of the amplifier, and G T can be obtained from S-parameter measurements of the source, amplifier and load, according to
81 of 286 ( )( ) 2 Ls2112L22s11 2 L 2 s 2 21 T SSS1S1 11S GΓΓ−Γ−Γ− Γ− Γ− = , (233) where S 11 , S 12 , S 21 , S 22 are the S-parameters of the amplifier, Γ s =(Z s -Z o* )/(Z s +Z o )=“S 22s ” is the reflection coefficient of the signal source (“S 22 ” of the source), Γ L =(Z L -Z o* )/(Z L +Z o )=“S 11L ” is the reflection coefficient of the load at the output of the amplifier, all these measured with a network analyzer with characteristic impedance Z o =Z o* =50Ω, for example. It is important to note that there are several noise figures for RF. For example, the publications usually report minimum noise figure [175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187], which is bias-frequency dependent and for optimum impedance of the RF network for the lowest noise figure. As mentioned above for the noise resistance, the earlier works [175, 176, 177, 178, 179, 180, 181, 182] address the RF noise from the perspective of the general theory of RF networks, replicating impedance mismatches in test setups with impedance tuners, while the later publications [183, 184, 185, 186, 187] relate the RF noise closely to device parameters and circuit applications, involving also simpler test setups to obtain the device noise parameters in cost-effective manner. The driving force for simplification of the RF noise analyses is normally that the RF noise is just one of many other performances that one needs to investigate, e.g., during reliability analyses by hot-carrier stress of RF amplifiers [198, 199]. We should also note that the optimum impedance for the minimum noise figure is neither the impedance Z o of the network nor the impedance for maximum power gain. Thus, the noise figure of an amplifier or mixer in a given RF application might be a lot larger (in the range 4-20 dB) compared to the minimum noise figure (in the range from fraction to few dB) of the used transistors. Signal to noise ratio. The most popular definition for noise figure is in terms of decrease of signal to noise ratio (SNR) in the output of the amplifier, SNR out , as compared to SNR in at amplifier input, or precisely, SNR s in the output of signal source. This form of NF SNR was historically first introduced [200] due to its clear meaning for the practice, when the noise in the signal source is thermal, e.g. antenna of receiver. NF SNR can be obtained from the generic definition for NF by assuming a signal with an arbitrary power level P s provided from the signal source and amplified to output power P L =G T P s on the load of the amplifier. Since the gain G T =P L /P s is the same also for the noise, then NF SNR is ( ) ( ) out in outL s avbl,inTsT avbl,ths avbl,th avbl,in SNR SNR SNR SP SNR SGPG SP S S NF ==== . (234) Evidently, NF SNR does not explicitly state what is the reference noise power at the input, and one may carry out a procedure for evaluation of NF SNR by measuring SNR first at the source plane and then at the output of amplifier. Since NF SNR is formally derived from the generic definition of NF, then one may decide that NF SNR =?=NF and to substitute in the generic definition for NF in order to evaluate the input referred noise S V , getting + + =+==== noise, RFfor , kTR S 1 or noise, LFfor , kTR4 S 1 S S 1NF?NF SNR SNR s RF,V s LF,V Vth V SNR out in , (235) from which the power spectrum density of the input referred noise voltage S V of the amplifier will be estimated
82 of 286 as ( ) ( ) SNR,V out in sSNRssRFV, LF,V V S SNR SNR 1kTRNF1kTR?NF1kTRSor 4 S S= −=−==−== . (236) We put the question in the equations, because the equality is valid only if SNR in is measured in respect to the thermal noise, which is not possible in the practice directly, because the thermal noise is the ultimate noise floor, and both the signal source and the measurement instrument for noise power are expected to have higher noise levels. Consequently, S V determined from NF is the correct, whereas S V,SNR calculated from NF SNR has a value different from S V . This is because, taking reference noise level S ref ≠kTR s , we have ( ) VV ref s VrefT sT ref s s out in sSNR,V SS S kTR SSG PG S P 1kTR SNR SNR 1kTRS ≠= + −= −= , (237) and, since the signal source used to supply with the signal P s has most likely noise S ref >S Vth =kTR s , then the use of NF SNR will underestimate the actual value for the input referred noise of the amplifier under test with the ratio S Vth /S ref . FOMs for non-thermal sources of noise In summary, the above discussion on noise temperature, noise resistance and NF implies that these are conditional figures of merit that relate the noise to the thermal noise, and thus, they are applicable when the origin of the noise in the signal source is thermal. There is a wide range of applications in the medium frequency range, e.g. material characterization, ultrasound imaging and electronic identification tags (RFID), where these figures of merit are essentially useful. The noise figure is particularly important for RF applications and currently intensive research is undergoing to resolve many issues with matching and characterization uncertainty. Furthermore, the thermal noise is dominant in the channel noise of MOS transistors, and therefore, the figures of merit that use the thermal noise as the reference are important. However, when the origin of the noise is not thermal, the above figures of merit are not very suitable. These are cases for BJT, reverse biased and avalanche diodes and non-resistive leakages (due to tunneling in insulators, especially in SOI transistors), for example, in which the shot noise is the major concern. As for the low-frequency range, the dominant noise is not originating from thermal noise, both in signal sources and in the devices, and the input referred noise is usually given with a pair S Veq and S Ieq of equivalent voltage and current noise sources, respectively. For MOS transistors, the input referred voltage noise is dominant, and it is given by S Veq =S VG in most cases, as discussed earlier in section IV.1. “Models and predictability”, while for BJT the input referred current noise is dominant, and it is given by S Ieq =S IB , as discussed in details in section III.3. “Crossover between different noise sources in BJT”. Most of the input referred noise voltages are physically non-existing quantities that represent equivalently noise sources inside the devices, while the most of input referred noise currents are actually fluctuations in biasing or leakage currents and are physically existing quantities in the input terminal of the devices and circuits. Interestingly, from the ratio S Veq /S Ieq we can derive other figures of merit, R eq and T eq , given by eqeqIeq eq Veq 2 eq Ieq Veq kT2RS R S R S S== = . (238)
83 of 286 The significance of R eq is that it determines whether the voltage or current input noise will dominate at particular resistance R s of signal source or circuit impedance. If R s <R eq , then the voltage noise is dominant, whereas if R s >R eq , then the current noise is dominant. At R s =R eq , both sources have equal contribution to the input referred noise, which can be referred to equivalent noise temperature T eq and equivalent power 4kT eq of the input referred noise. By expressing in terms of voltage noise and using eq. (228), for example, we can relate to noise figure, as ( ) min eqeq eq s s eqeq s eqeq 2 seqeq s Ieq 2 sVeq Vth V NF T Tmin 1 T T 1 R R R R 5.0 T T 1 kTR4 R2kTRR2kT 1 kTR4 SRS 1 S S 1NF =+≥+≥ ++= + += + +=+= (239) Taking into account that S Veq and S Ieq vary with the frequency, for example 1/f noise at low frequencies, the significance of last line of expressions is the following. The left-hand expression shows that NF has a local minimum when R s =R eq (f), and the local minimum [1+T eq (f)/T] is given with the second expression. For illustration, consider again the contour plot in the right-hand side of Figure 35 for a frequency about 200Hz. At low R s =1kΩ, the 1/f noise voltage S VG is dominant, but increasing R s to 1GΩ, the shot noise current 2qI ESD takes over. At R s =R eq ~30MΩ, the 1/f noise and the shot noise have equal contributions, NF reaches local minimum of about 0.5dB=1.12, and since the curves are for room temperature, then T~300K and T eq (f=200Hz)~36K. Increasing the frequency, the local minimum decreases, and for f~10kHz and R s ≈2MΩ, the global minimum for NF=NF min <0.1dB=1.02 is reached, which corresponds to min(T eq )<7K. This is given with the left-hand expressions in eq. (239). In this way, we demonstrate that the ratio between voltage and current noise is related to the noise resistance R eq and noise temperature T eq by eq. (238), and at condition R s =R eq , the contribution of voltage and current noise are equal, the noise figure has a minimum with value of (1+T eq /T), as given with eq. (239). Recalling that the original derivation for RF noise figure has extracted NF min from input referred voltage and current noise sources [201], then NF min should have similar meaning in the case of frequency dependent impedances, as the discussed above for low frequencies. We did not find recent work that provides deep insight on the physical significance of the four RF noise parameters NF min , r n and magnitude and angle of Γ opt , that participate in the popular equation for RF noise figure ( ) ( ) 2 opts s n optmin 2 opt 2 s 2 opts noptminRF YY G R )f(YNF 11 r4)f(NFNF −+= Γ+ Γ− Γ−Γ +Γ= , (240) where at given frequency f, the RF noise figure NF RF has a minimum in respect to signal source matching Γ opt , or admittance Y opt , and when the reflection coefficient Γ s or admittance Y s =G s +jB s of the signal source deviates from Γ opt or Y opt , then NF RF increases with a “rate” given by the parameter r n =R n /Z o , or R n /G s , respectively. The relation between power Y s =Y in * and noise matching Y s =Y opt is still not well elaborated for implementation in design procedures.
84 of 286 Another issue for RF noise figure is that the reference noise is solely attributed to the real component R s =1/G s of the impedance of the signal source, and all reactances are assumed noiseless. This is somewhat in contradiction with the general expression for thermal noise in eq. (218), which does not discriminate complex impedances and suggests looking at the input signal plane not discriminating between source and load at it. Thus, along with the many technical difficulties in measuring RF noise, some more general research is expected in near future, since the RF applications reached maturity in millimeter wavelengths, and the reactances and distributed loss dominate in these circuits, while the equation for noise uses lumped parameters. Again, the thermal noise is not dominant in the low-frequency noise, and all figures of merit related to thermal noise are lacking of physical significance from the perspective of non-linear device physics at low frequency. Therefore, other figures of merit are usually used for low-frequency noise, and these are discussed next for MOS transistors. IV.4.4. Physical figures – trap density, Hooge parameter, scattering parameter For MOS transistors, three physical parameters are usually used as figures of merit for the 1/f noise. These are oxide trap density N t for number fluctuation model, scattering parameter for models with correlated mobility fluctuation and Hooge parameter α H for uncorrelated mobility fluctuation and noise from contact resistance. The physical significance and the issues that have arisen with these parameters were discussed in section IV.1. “Models and predictability”. Here we mention that the most popular procedure to discriminate number and uncorrelated mobility fluctuation is that in [156], which analyzes the behavior of normalized noise S ID /I D2 against the behavior of (g m /I D ) 2 versus bias. Consider eq. (88). If S ID /I D2 ∝(g m /I D ) 2 , then the 1/f noise in MOS transistor can be referred as gate voltage noise S VG with origin number fluctuation according to eq. (104), whereas, if S ID /I D2 ∝(g m /I D ), then the uncorrelated mobility fluctuation is in the origin of 1/f noise, since in strong inversion (g m /I D )∝1/(V G −V T )∝1/n eff is inversely proportional to number of carriers n eff in MOS channel, while both in strong and in weak inversion regimes I D ∝n eff , and so, S ID /I D2 ∝α H /n eff ∝1/I D . Obviously, the noise is referred to the gate terminal of MOS transistors, and when including the correlated mobility fluctuation, the gate referred 1/f noise voltage to the first order of approximation is ( ) [ ] ( ) TG H ox 2 TG t 2 ox V VV WLC q f 1 VV1 WL NkT C q f 1 S G − α +−θ+ λ = , (241) as follows from eqs. (88), (90), (104), (120) and (121) for strong inversion regime of operation of MOS transistor above threshold voltage V T in linear mode, for example. The tunneling attenuation distance λ and gate capacitance per unit area C ox depend on the gate dielectric, and the mobility degradation coefficient θ depends on the electric field, but to the first order of approximation they can be taken constants for the purpose of comparison in terms of figure of merit. Taking the values for silicon MOS transistor with SiO 2 gate insulator of thickness t ox =EOT and with permittivity ε e =ε SiO2 =8.85×10 -14 F/cm tunneling attenuation distance λ e =0.1nm, one can rewrite the last equation as ( ) [ ] ( ) TGH e 2 TGt e 2 e V VV WL qEOT f 1 VV1N WL kT qEOT f 1 S G −α ε +−θ+ λ ε = , (242) where the quantities in the large brackets are taking care for device scaling, and N t , θ and α H can be used as figures of merit, which have physical meaning of equivalent trap density, scattering parameter and Hooge
85 of 286 parameter, respectively, for number fluctuation, correlated mobility fluctuation and uncorrelated mobility. Note that, if the scattering parameter θ is constant, then the correlated mobility results in quadratic dependence on gate overdrive (V G −V T ), while the uncorrelated mobility suggests a linear dependence, if α H is bias independent. This difference in the bias dependences is suggested in [156] as the criterion to discriminate correlated and uncorrelated mobility fluctuations. Many publications suggest that the Coulomb scattering causes bias variation of correlated mobility fluctuation, as discussed earlier by the help of eqs. (116) to (138). The figure of merit for the equivalent oxide density N t is shown in Figure 36. For nMOS transistors, the data are collected from [22, 47, 48, 49, 50, 51, 72, 82, 88, 89, 90, 91, 92, 93, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 149, 145, 202, 203, 204, 205, 206] for silicon transistors and from [109, 155, 207] for transistors that use germanium in the structure, e.g. to provide strain in the lattice. For pMOS transistors, the data are collected from [47, 52, 88, 94, 95, 96, 97, 100, 102, 104, 105, 109, 124, 125, 126, 128] for silicon transistors and from [52, 57, 79, 109, 208] for SiGe and SiGeC pMOS transistors. The data are stored in numerical form in [209]. While in the past, the uncorrelated mobility fluctuation was found to dominate in pMOS transistors, later publications imply the opposite, when EOT<6nm. The second observation in Figure 36 is that there is a crossover for N t at EOT<4−5nm. In “thick” oxide transistors with EOT>4nm, the equivalent oxide trap density is low with an average of N t ~10 17 cm − 3 eV − 1 , as shown by the right-hand histogram in Figure 36. In “thin” oxide transistors, however, N t (and thus, the 1/f noise) increases inversely with EOT with a steep slope of m=2.5 in magnitude. The scattering in the data is large and the slope was estimated by adjusting the symmetry in the left-hand histogram for the quantity N t EOT m . Since the slope m>2, then the trend at low EOT suggests that the 1/C ox ² scaling rule for 1/f noise is not followed anymore for transistors with high-k gate dielectrics of high permittivity. Many reasons for higher N t are suggested in the above references, e.g. larger tunneling attenuation distance λ in high-k dielectrics, non-uniformity of dielectric structure, and other. One of the other is the increased doping in the channel of MOS transistor, which is necessary in order to compensate for drain induced barrier lowering. The higher doping results in intensive Coulomb scattering at low bias, which effectively increases gate referred 1/f noise voltage at (V G −V T )~0.1V, as discussed in section IV.3. “RTS noise in MOS transistors”. On the other hand, this biasing condition is usually assumed for extraction of the value for N t , since the term θ(V G −V T )<<1 is assumed in eq. (242). However, as follows from eq. (138), the relative contribution of correlated mobility fluctuation due to Coulomb scattering increases when the bias is reduced, and the figure-of-merit form in eq. (242) may overestimate the oxide trap density. Interestingly, the slope of increase of N t at low EOT in Figure 36 is with value 0.5 higher than 2, which one can expect from the term for Coulomb scattering in eq. (138). To avoid the Coulomb scattering effects, we have used in Figure 37 only data for high overdrive voltage (V G −V T )>0.3V. At this condition, one expects that the effective parameter θ for correlated mobility fluctuation reflects phonon and surface scattering, according to eqs. (136) and (137); and θ is given by t ox 4 soxs 15 s ox s NEOT 1 C C/Vs10~ with ,C Vs10~ with , q C × ∝µ∝ αµα αµα =θ − (243) depending on whether the electron charge q is included in the definition for the scattering parameter α s , that is, whether the number fluctuation model uses concentration n’ of channel carriers or their charge concentration
86 of 286 q×n’. In eq. (243), the gate oxide capacitance per unit area C ox suggests that θ∝1/EOT, and according to eq. (136), the carrier mobility μ suggests θ∝1/N t . None of these dependences can be observed in Figure 37 from the published data in [47, 48, 49, 50, 51, 95, 100, 106, 149] for nMOS transistors, in [155] for strained on SiGe layer and control nMOS transistors, in [47, 52, 88, 94, 95, 100, 105, 124, 126, 128] for pMOS transistors, and in [52, 57, 79, 208] for SiGe and SiGeC pMOS transistors. The data are stored in numerical form in [209]. The above discussion clearly indicates that there are problems in the characterization of 1/f noise in terms of physical figures of merit. One partial solution to the problems could be, if the expression for correlated mobility fluctuation is split into two terms, one for Coulomb scattering at low bias and another for phonon and roughness scattering at high bias. The resulting equation for the gate referred 1/f noise voltage is ( ) [ ] ( ) TG H ox 2 TGTGC te 2 ox V VV WLC q f 1 VVVV1 WL NkT C q f 1 SG− α +−θ+−θ+ λ = , (244) but this equation has a problem when the parameter θ C associated with the Coulomb scattering dominates, since the equation reduces to ( ) ( ) TG H ox TG 2 C te 2 ox V VV WLC q f 1 VV WL NkT C q f 1 SG− α +−θ λ = , (245) and the correlated and uncorrelated mobility fluctuations cannot be discriminated experimentally, because they have the same bias dependence. Nevertheless, both fluctuations in this case are mobility fluctuations and perhaps it is not so critical to know their contributions separately. If this is acceptable, then one can define “apparent” Hooge parameter α HC from the parameter θ C for correlated mobility fluctuation caused by Coulomb scattering by the equation 2 0C te 2 C ox te HC NkT C NqkT µ µ λ=θ λ =α , with qEOTq C2 SiO 0C oox 0C C ε µ µ ≈ µ µ =θ (246) for strong inversion regime, where the mobility μ o ≈μ can be regarded as the maximum value for carrier mobility at low electric fields [57]. By taking typical values μ≈100cm²/Vs, μ C0 ≈10 8 cm/Vs and C ox ≈10 − 6 F/cm² for transistors with EOT~3.5nm, then θ C ≈2.5 V − 0.5 , and 1/f noise due to Coulomb scattering dominates at (V G −V T )>0.16V, since 1VV TGC >−θ . Therefore, eqs. (245) and (246) are applicable for (V G −V T )~0.2V, and since kT=0.026eV and λ e ≈0.1nm, then from eq. (246) we get α HC ≈2.6×10 − 4 for N t ~10 18 cmˉ³eVˉ¹, the latter taken from the trend in Figure 36 for EOT~3.5nm. Interestingly, while “apparent”, α HC is in the range of meaningful values for Hooge parameter, which shows one more time that there is a convergence between different models for 1/f noise. IV.4.5. Performance figures – RF to LFN Denormalization rules for FOM SVG The mostly accepted performance figure of merit for 1/f noise in MOS transistor is FOM SVG , it is related to the gate referred voltage noise S VG , as given by eq. (106), and it originates from number fluctuation model. FOM SVG is the power spectrum density of S VG at 1Hz in a transistor with gate area of 1μm² and at low gate overdrive
87 of 286 voltage (V GS -V T )≈0.1V. FOM SVG takes into account the general scaling rule for reciprocal dependence between noise level and device active area, given by eq. (14), and therefore allows for examination other factors and scaling rules that impact the 1/f noise and for comparisons between devices, technologies and so on, as has been illustrated in Figure 20 and Figure 22. The details related to FOM SVG are discussed in section IV.1. “Models and predictability”. Here, we briefly give the denormalization rules for FOM SVG by the help of the following equation 2 mVI 2 S V gSS and WL m1 f FOM SGD VG G= µ = , at specified bias condition for FOM SVG . (247) Note that the bias dependence of the noise is neglected in FOM SVG at an assumption that the bias dependence of the noise in relative units is similar in different MOS transistors, and also, it is assumed that the only one size dependence is the scaling rule with the reciprocal of the area of otherwise identical devices and the noise is exactly 1/f. Evidently from the discussions in preceding sections, the bias, frequency and size dependences vary several decades, and the models are devoted to capture these dependences. There are many issues when denormalizing FOM SVG with bias and when the noise is not 1/f. Therefore, FOM SVG is just a helpful figure of merit with significance for general comparisons, while, in practice, the actual realization of the low-frequency noise in particular MOS transistor and circuit made of it will be different from FOM SVG . Corner frequency f c between flicker (1/f) noise and white noise (FOM fc ) Other figure of merit is the corner frequency f c between flicker (1/f) noise and white noise, which we denote here as FOM fc ≡f c . FOM fc is used often in the practice in the fields of electronic oscillators, frequency converters and sampling systems, because of two reasons. First, the design equations are given in terms of f c , and second, the white noise in MOS transistors follows very closely the theoretical derivations made on assumption for channel thermal noise. The power spectrum density of the channel thermal noise is given by [169, 174] ( ) −µ== TGoxwh,I 2 mwh,V VVC L W kT4Sg S DG , MOS transistor white noise in linear regime, (248) ( ) −µ= TGox VVC L W 3 2 kT4 , MOS transistor white noise in saturation regime, (249) D qI2= , MOS transistor white noise in sub-threshold regime, (250) where S ID,wh is the white noise in drain current, S VG,wh is the gate referred voltage white noise, which is “coupled” back for convenience via the transconductance g m , according to eq. (1). These equations are for long-channel approximation, and one usually applies multiplicative correction factor γ for excess noise due to short channel effects, body leakage, gate induced or avalanche noise [174], but we omit it for clarity in analyses of lowfrequency noise, since the variations in flicker noise are much larger than γ, while γ plays essential role in RF range. We rewrite eqs. (248), (249) and (250) in terms of gate referred voltage white noise S VG,wh = S ID,wh /g m ², because we have the expressions for the gate referred flicker noise S VG in previous sections, e.g., eq. (241) above. The resulting equations for S VG,wh are
88 of 286 : in linear mode Doxm VC L W gµ= . So ( ) 2 D TG ox 2 m wh,I wh,V V VV W L C kT4 g S S D G − µ == , MOS transistor white noise in linear regime, (251) in saturation mode ( ) TGoxm VVC L W g−µ= . So ( ) 2 4 1 3 G V ,wh ox G T kT L S C W V V =µ − , MOS transistor white noise in saturation regime, (252) in sub-threshold mode t DD m I qkT I gϕ == . So − ≈= = t TG t ox t DD whV VV W L C kT I q q kT I q S G ϕ ϕ µ ϕ exp 14 4 122 2 2 , , white noise in sub-threshold regime, (253) when using ( ) αϕ − ϕαµ≈ t TG 2 toxD VV exp2C L W I for the MOS transistor current in sub-threshold regime [p.177 in 169] with α≈1 by neglecting depletion capacitance – see eq. (186). The corner frequency FOM fc ≡f c between flicker noise S VG and white noise S VG,wh is defined obviously as the frequency at which S VG = S VG,wh , because at f<f c the flicker noise dominates S VG in the power spectrum density of the noise, while S VG,wh dominates at f>f c . To obtain expressions for FOM fc ≡f c , we use the last three eqs. (251), (252) and (253) for the white noise S VG,wh in different regimes of operation of MOS transistor, and the components for the 1/f noise S VG for different noise mechanisms in eq. (241). - In sub-threshold regime, we use the number fluctuation model ∆n for the flicker noise S VG , and FOM fc is found using eq. (253) for the white noise S VG,wh , as wh,V t TG t ox t 2 oxc V GG S VV exp 1 W L C kT4 4 1 WL NkT C q f 1 S= ϕ − ϕ µ = λ = , ∆n in sub-threshold. (254) D t G 2 t TG t ox t 2 c TG f I V exp L VV exp C Nq f VV n FOM c∝ ϕ ∝ ϕ − ϕ µ λ == < ∆ . (255) - When using eq. (252) for the white noise S VG,wh in saturation regime above threshold, the number fluctuation ∆n for the flicker noise S VG results in FOM fc , given by
95 of 286 ternary SiGeC alloy. Data (b) and (d) are for silicon nMOS and pMOS transistors, in which an increase of noise and N t is observed when using nitridation of silicon oxide dielectric in order to increase the gate breakdown voltage. Data (c) are from the attempts to increase the oxide capacitance using high-k dielectrics in Si nMOS transistors, which are accompanied also with increase of noise and the corresponding values for N t . Data (b), (c) and (d) imply that a departure from SiO 2 in gate dielectrics causes increase of 1/f noise. Data (e) are from several publications, which reported variation of 1/f noise when changing the MOS channel alloy from Si to SiGe, then using high-k dielectrics in the gate stack in transistors with SiGe alloy in the channel, followed by semiconductor-on-insulator (SOI) structures based on Si and on Ge. The reported values for N t increased in this order of increasing complexity of the MOS structures. Data (f) are from a research, in which annealing after complete fabrication of SiGe MOS transistor was carried out. The post-annealing with inert Argon gas reduced N t , while the post-annealing with water vapor increased N t , owing to removal or adding of oxide traps, respectively. Overall, data (a) to (f) imply that there is an increase of the 1/f noise when using composite materials either for semiconductor or insulator, or adding non-uniformity in the transistor structure, which, perhaps, degrades the lattice quality in the regions related to the current transport in MOS transistor, e.g. conductive region of the channel and semiconductor-dielectric interface. These observations are cumulatively established, since the introduction of interfacial layer (~0.8−1.2nm) of SiO 2 between semiconductor and high-k dielectric [49, 50, 92, 99, 100], as well as the insertion of Si cap layers (~2−8nm) on the top of SiGe conduction channel [231, 232] usually result in lower noise, perhaps due to smoother transition between different materials and reduced density of point defects at the interfaces. This is further confirmed by using the strained Si cap layer as the conductive channel rather than the SiGe layer. Since the Si lattice on the top of a SiGe layer is strained, then the electron mobility in Si can be increased, but the lattice quality of strained Si and the uniformity of SiO 2 grown on the top of Si layer are well preserved. Consequently, as shown with data (h) in Figure 42, the noise in nMOS transistors with strained Si channel is only slightly higher than the noise in “regular, control” MOS transistors with unstrained silicon, given by data (g), respectively. It is interesting to mention that the way of introducing the strain in Si transistors is important for the 1/f noise. If the strain is applied from a region which does not participate in the current transport, then there is virtually no change in the 1/f noise. This has been demonstrated in [126], where a Si 3 N 4 layer on the top of the gate for compression of Si layer below the gate (and enhancement of the hole mobility in pMOS transistors) did not increase the noise, while the use of SiGe for drain and source, which also compress the silicon in the channel, resulted in increase of the noise, because the SiGe regions participate in the path of the current flow and the outdiffusion of Ge toward the channel is possible during device processing. For locally strained devices like the above pMOS transistors [126] or small area MOS transistors, the strained Si layer does not introduce additional noise, since the point defects are few in the small devices. In larger-sized MOS transistors, however, threading dislocations in the strained layer are statistically occurring, and the noise can increase. The areal dependence of this increase was modeled in [82] using Poisson distribution for the extra traps in the dislocations and geometric averaging. This one more time confirms that the increase of the noise in SiGe and strained-Si MOS transistors is due to material defects in the path of the current transport, but it is not a fundamental material property of SiGe
96 of 286 alloys and strained Si lattices, which is somehow in contrast to the explanations in [127] for the effect of the mechanical stress from shallow trench isolation in pMOS transistors. V.1.3. Modeling the 1/f noise in SiGe MOS transistors Although the above observations are well established experimentally, the modeling and the understanding of the physics behind the 1/f noise in SiGe MOS transistors are still not very certain. In most cases, the 1/f noise models for silicon transistors are used, c.f. portion or the whole eq. (241), with some enhancements for Hooge noise from drain and source contact regions [208] and modifications in the correlated mobility term, splitting it into two parts for screened Coulomb and phonon scattering [52, 57]. For the case of SiGe channel and Si cap layer on the top of it, it is suggested in [232] to use current partitioning between the potential well in SiGe at the Si interface and the well in Si at gate oxide. This partitioning can be lumped formally into two transistors, one in SiGe layer and another in the Si cap layer, which are driven by the same flat-band noise voltage S FB . The resulting equations are 2 SiGe,m SiGe,D SiGeoxsSiGe,m Si,m Si,D SioxsSi,mFBn,I g I CR1g g I C1gSS D µα++ µα+= µ∆−∆ (271) α + α = 2SiGe,D SiGe SiGe,H 2Si,D Si Si,H Hooge,I I n I nf 1 S D (272) Hooge,In,II DDD SSS += µ∆−∆ (273) where S FB is the same as for Si MOS transistors and given by eq. (109), the total noise S ID in the drain current is a sum of correlated number-mobility fluctuation S ID, Δ n −Δμ and Hooge noise S ID,Hooge , α s is the scattering parameter given by eqs. (127) and (136), C ox is the gate oxide capacitance per unit area, and the parameter R~0.1−0.2 is taking into account the remote interaction between SiGe layer and oxide traps, separated by the Si cap layer. The other parameters are respectively for Si cap layer and SiGe channel, and they are transconductances g m,Si and g m,SiGe , carrier mobilities μ Si and μ SiGe , Hooge parameters α H,Si and α H,SiGe , DC currents I D,Si and I D,SiGe , and total number of carriers n Si and n SiGe in the corresponding layer. It has been shown in [232] that this model can produce non-monotonic dependences for the noise levels as function of bias, thickness of cap layer and Ge content in the SiGe layer. However, the above equations assume that the Si and SiGe parts of the transistor conduct independently by neglecting the charge transfer between the two potential wells, which occurs when the transistor is operating in saturation regime well above the threshold in strong inversion, and the charge from Si layer at the source side is transferred into SiGe well along the channel length in the region close to pinch off point. Furthermore, the model requires a precise knowledge for the structure, e.g. layers’ thicknesses, concentrations, mobility, etc., which are not always available, and the correct partitioning can be done only after numerical simulation of the structure by using a semiconductor simulator. The issue with this noise model for SiGe MOS transistors, in fact, is that there is no mature characterization technique which can obtain the model parameters with unique values. V.2. Forward body bias in MOS transistors – not a panacea, but it helps In many circuit applications, especially in low-voltage circuits, the potential of the body and source of MOS transistors are different. Examples are dynamic threshold configurations [233, 234, 235], in which the body and
97 of 286 gate of the MOS transistor are tied together in order to obtain higher transconductance, differential amplifiers, when the transistor pair and biasing current source are in one well or in the substrate, or the body of the MOS transistor is used as input [236, 237, 238], cross-coupled pairs for oscillators [239, 240], in which the body is used for tuning of the oscillation. Different approaches for power management in digital circuits [241, 242, 243, 244] also use the body bias as a control to bring inactive portions of the circuit in sleep mode, applying reverse body−source voltage V BS , or applying forward V BS to accelerate the circuit in active mode and at low supply voltages. The low-frequency noise in MOS transistors, however, is sensitive to the body biasing, as illustrated in Figure 43 with data from [245] for a pMOS transistor from 0.18μm technology. In the left-hand plots, the data for several quantities, which represent the 1/f noise at 1Hz, are shown at different body−source bias voltages V BS , varying also gate bias voltage V G , and at unchanged drain bias voltage V D =0.6V. The different V BS are from -0.6V (reverse) to +0.6V (forward). Reverse or forward body bias is in respect to the direction for conduction of the body−source p−n junction. The quantities DC drain current I D , small signal transconductance g m and their ratio in the horizontal axes are obtained from I−V curves of the transfer characteristics. The insets in the figures show that the these quantities depend on the gate overdrive (V G −V T ), thus on carrier concentration n’∝(V G −V T ) in the transistor channel, since the values for I D and g m (and their ratio, consequently) overlap when V BS was varied, but (V G −V T ) is kept constant. In similar manner, the effect of body biasing on 1/f noise in one nMOS transistor was attempted to be explained in terms of carrier density n’ in [51], but measured data at only one constant gate bias were reported in this publication. However, the case is different for the low-frequency noise, as one can see from the data series for the noise in Figure 43. In particular, the 1/f noise is not a unique function of areal carrier density n’ in the channel, because at given values for the quantities in the horizontal axes, and therefore constant n’ as follows from above discussion, the noise decreases by a transition from high gate bias and reverse body bias toward low gate bias and forward body bias. This transition indicates a decrease of noise when the current flow is moved from the semiconductor-dielectric interface (surface channel at reverse body bias and high gate bias) toward the bulk of the semiconductor (buried channel conduction at forward body bias and low gate bias). Such transition in more pronounced form (from Δn noise at surface MOS channel to Hooge noise in the bulk JFET channel) is observed in [246] for a MOS−JFET SOI structure with gates around the conduction body. Further discussion on the crossover from surface to bulk noise is given section VI. “ Noise in advanced transistor structures ”. As follows from the above observations, there is an issue related with the body biasing of MOS transistors. The problem is that all the flicker noise models are derived at the assumption for surface conduction and charge sheet approximation for channel, and the noise models allow for bias variation only of the areal concentration n’ of charge carriers in the MOS transistor channel. Consequently, the noise parameters, either in number fluctuation or mobility fluctuation models, have to be varied with the body bias V BS . Therefore, the compact noise models used for computer simulations are unable to reproduce effects related to the depth of the channel and the simulations of some specific connections of MOS transistor can be inaccurate. For example, applying reverse bias to MOS transistors in voltage controlled oscillator (VCO) will underestimate the phase noise, and using the dynamic threshold configuration (with gate and body tied together) will require another set of values for the coefficients in the noise model. Qualitatively, the channel depth in pMOS transistor is depicted in the right-hand plots of Figure 43 for reverse,
98 of 286 zero and forward body bias (plots from top to bottom) at similar charge sheet concentrations (the dotted shapes for the charge carrier layers denote same area). Quantitative examples for the charge concentrations in Si and SiGe MOS transistors can be found in [247]. To the best of our knowledge, there is no 1/f noise model for MOS transistors that considers the volume distribution of the carriers in the depth of the channel, which can possibly explain the body bias dependence with variation of current density in terms of Hooge noise in its integral form of eq. (13), neither 1/f model, which considers variable distance from oxide traps to carriers in the channel, a distance which is usually taken zero for surface channels or constant in SiGe channels capped with Si layer, and which can possibly explain the body bias dependence with variation of Coulomb interaction between oxide traps and channel carriers in terms of number fluctuation models. Overall, the charge sheet approximation works well for DC and AC modeling of MOS transistor [233, 248], but it is not accurate for the noise. Nevertheless, the practical rule is that the forward body bias V BS reduces the noise in relative units, while the reverse V BS increases the noise, as illustrated in Figure 44 with data from [52, 245] for the gate referred 1/f noise S VG in Si and SiGe pMOS transistors in saturation and linear regimes of operation, respectively. A similar observation for the dependence of gate referred noise voltage on body bias can be found in [247] for pMOS Si and SiGe, in [249] for nMOS and pMOS transistors from a 130nm technology in weak inversion, in [250, 251] for SOI MOS transistors operating in dynamic threshold (body tied to gate) and normal (body tied to source) connections. The data in [249, 250, 251] have been explained in terms of number fluctuation model for the 1/f noise in MOS transistors, see eqs. (104), (109), (120) and (126), and sheet approximation for channel charge carriers coupled to trapped oxide charge via oxide C ox and depletion C d capacitances (per unit area). From the analyses in these publications, it can be shown that the gate referred noise voltage is a result of coupling to the total capacitance (C ox +C d ) seen at semiconductor-dielectric interface. That is, ( ) 2 m D effdoxs t 2 dox 2 m I V g I CC1 WL NkT CC q f 1 g S S D G µ+α+ λ + ≈= . (274) This equation reproduces the behavior of the 1/f noise as function of the body bias, because the depletion capacitance C d increases with forward V BS , resulting in decrease of S VG , and C d decreases with reverse V BS , resulting in increase of S VG . Also, the equation is in agreement with the general equations (4) and (5) for noise coupled from charge fluctuation. Eq. (274) was derived in [251] for dynamic threshold (DT) configuration (body tied to gate) in the form of FB 2 m D effoxs ox d 2 mI S g I C C C 1 1 gS D µα± + = , with WL NkT C q f 1 S t 2 ox FB λ = (275) where the oxide charge trapping is referred to the gate and attributed to noise in flat-band voltage V FB , see eq. (109), and it was experimentally validated by the similar values for oxide charge density N t obtained from DT connection after correction (1+C d /C ox )² and normal connection (body tied to source) without correction. The attenuation of the noise by (1+C d /C ox )² was also confirmed in [249] for weak inversion, where it was also observed that the body bias has no effect in strong inversion due to screening effect that the inversion layer charge has on C d . There are also other issues with the physical consistence of eq. (274), such as the question why
99 of 286 C d is omitted when body and source are tied together, and the explanations on these issues are mostly qualitative in the publications cited above. Overall, eq. (274) is useful for the practice, because (1+C d /C ox )²<2, which is within the experimental inaccuracy for characterization of N t from noise measurements at low gate overdrive voltage (V G −V T )<0.2V, while at higher overdrive, when the term for correlated mobility fluctuation dominates, the capacitances cancel in (274), and the values for the scattering parameter α s are unaffected by the choice C=C ox or C=(C ox +C d ). Thus, the addition of C d in the number fluctuation model does not compromise with the convergence of the model, and gives a straightforward way to introduce the body bias dependence in this model. VI. Noise in advanced transistor structures There are several reasons to focus on advanced silicon transistor structures and noise in them [3]. The semiconductor industry has been based for more than 40 years on scaling of transistor dimensions to achieve performance gains, utilizing tremendous investment in infrastructure to the highest possible degree. While the role of nanoscale devices in meeting future computing and communications applications is not clear at present, there are significant limitations that arise with nanoscale devices and will impact their usefulness. In particular, the near-term applications require nanoscale devices to be functionally and technologically compatible with silicon transistors, at least for using semiconductor manufacturing and design infrastructures, and for interfacing to scales accessible by human. On the other hand, there are many difficulties by downscaling silicon devices, even CMOS transistors, for which the advances are the most, because fundamental limitations for charge transport, electrostatic control and accuracy of device fabrication are reached in planar technologies, in particular bulk CMOS, and binary logic (state variables) based on electric charge are vulnerable to random errors due to low signal to noise ratio, and thermal and material instability. In order to improve the electrostatic control, the advanced silicon transistors utilize the vertical dimension of the structures, by developing approaches originally introduced for SOI. In this section, we review the results obtained for several 3-D transistor devices, Fin FET, Two-, Threeand All-Around-Gate FETs, along with complications for the low-frequency noise in SOI MOS transistors. At the other end, the enhanced 1-D current transport in carbon nanotubes and semiconductor nanowires attracted the attention recently, and we also review in this section results presently available for the low-frequency noise in devices based on 1-D charge transport. VI.1. From SOI toward gate all around In attempt to enhance the electrostatic control of the surface charge transport in MOS transistors, the body of SOI MOS transistors is made thinner and the semiconductor bulk is removed and replaced with insulator. The removal of the semiconductor bulk resolves problems related to drain induced barrier lowering (DIBL) caused by the high electric field in the bulk of MOS transistors owing to drain bias and uncontrolled by the gate bias [252]. The corresponding structures are known as partially depleted (PD) and fully depleted (FD) SOI MOS transistors, which are depicted in Figure 45. VI.1.1. Partially depleted SOI MOS transistors The transition from bulk MOS (Figure 45a) to partially depleted SOI MOS (Figure 45b) is accompanied with reduced control of body potential V B . In bulk MOS, the bias voltage of the body terminal sets reliably the level of V B , because any excess charge generated or entered in the body (e.g. due to impact ionization or valence band tunneling through thin gate oxides) is sunk-out through the low impedance of the conductive bulk, as illustrated
100 of 286 by arrows in the figure. Thus, the depletion region under the MOS channel (dashed line in Figure 45a) is fixed by the gate bias accordingly, and it does not vary randomly, once the body terminal (normally tied to source terminal) is connected to low impedance bias. Consequently, the low-frequency noise in the channel current is coupled only from the gate oxide trapping and charge transport in the inversion layer, as discussed in previous sections. However, the control of the body potential V B is weak in PD SOI MOS (Figure 45b), because the thin bulk has high impedance to the body terminal, or the body terminal is removed in order to reduce the size of the transistor. Therefore, the excess charge generated or entered in the body increases the body potential V B so that the junction body-source becomes forward biased and the excess charge flows toward the source terminal overcoming the impedance of the body-source junction. In this way, the body potential V B becomes dependent on the body current I B and the effects related to this dependence are depicted in Figure 46 with data from several publications [250, 252, 253, 254, 255]. As shown in Figure 46a, when the drain V D and gate V G bias voltages increase, then the electric fields at the drain side of the channel and in the gate of the MOS transistor increase, causing leakage, impact ionization and valence band tunneling currents I B , which flow into the body [250, 251, 252, 253, 254, 255, 256, 257, 258, 259]. The body current I B has to be readily sunk by the body-source junction, since all other interfaces surrounding body in SOI MOS transistor are either insulators or reverse biased junctions. Once I B becomes larger than the (reverse) saturation current I B0 of the body-source junction, then the body voltage V B increases, as shown in Figure 46b, and the variation in I B with bias causes changes in V B . The increase of V B reduces the threshold voltage V T of the transistor and undesirable kink in the transistor DC characteristics occurs, as illustrated in Figure 46c. Consequently, the noise S IB in I B causes voltage noise S VB =Z B2 S IB in the body potential, where Z B is the AC impedance of the body to ground, which is further coupled to the inversion layer as excess voltage noise S VG,ex by the body coefficient ∂V T /∂V B . Since I B is low, usually less than a nanoampere, then the 1/f noise in I B is negligible, and the shot noise in I B dominates, because the charge carriers related to I B cross the gate and drain potential barriers to enter the body, and body-source barrier to exit from the body. Assuming uncorrelated processes, then shot noises for currents entering from gate and drain into body (G&D→B) and exiting from the body toward source (B→S) are summed, and the noise S IB in I B is BSBBBD&GBI qI4qI2qI2S B =+≈ →→ , (276) where q=1.6×10 − 19 C is the electron charge, S IB is frequency independent (white noise) conservatively for the range of low frequencies, and S IB is proportional to the body current I B . The body impedance Z B =1/(G B +j2πfC B ) can be lumped into a simple RC equivalent circuit [97, 251, 258, 255] of parallel connection of body-source junction dynamic conductance G B ≈I B /φ t and body capacitance C B . The later has several components, e.g. depletion capacitance under the MOS channel, junction capacitances of the drain and source junctions, bottom oxide (back-gate) capacitance and other, depending of the SOI MOS transistor layout, but C B varies less with bias, as compared to G B , which is proportional to I B , and I B is close to exponential function of V G and V D . Thus, Z B acts as a low-pass filter for S IB , and the voltage noise S VB in the body potential has a Lorentzian spectrum, given by
101 of 286 ( ) ( ) 2 0 V 2 BB B I 2 BV ff1 0S fC2jG qI4 SZS B BB + = π+ ≈≈ , (277) where with φ t =kT/q being the thermal voltage, the low-frequency plateau S VB (0)=4qI B /(G B ) 2 ≈4qφ t ²/I B of the Lorentzian noise is inversely proportional to I B , while the corner frequency f 0 =G B /(2πC B )≈I B /(2πC B φ t ) is proportional to I B , and S VB is reproduced by the body coefficient ∂V T /∂V B as gate referred excess voltage noise S VG,ex given by ( ) ( ) 2 0 ex,V 2 B T Vex,V ff1 0S V V SS G BG + = ∂ ∂ ≈ , (278) where (∂V T /∂V B )≈constant is a weak function of the bias of the transistor. Thus, since I B increases with gate and drain biases of the PD SOI MOS transistor, then the excess Lorentzian noise evolves with the bias of the PB SOI MOS transistor. The evolution of the excess Lorentzian noise is as shown in Figure 46e as function of the drain bias at constant gate overdrive (V G −V T )=constant. Apparently, at given low frequency and varying the drain bias, one will observes noise “overshot” due to the evolution of the S VG,ex with V D . The overshot is correlated to a biasing condition at the onset of the kink in DC characteristics, as illustrated in Figure 46d, but note that the position of maximum of the overshot is dependent on the selected frequency and it is not unique function of the biasing condition [250]. In fact, one can mistakenly conclude that that the noise power has a maximum at particular bias. This issue is discussed in [88, 259], where is shown that the total “power” P VG,ex of the gate referred excess noise voltage in a wide frequency band is a weak function of the bias, when considering that ( ) 2 B T B 2 B T B t ex,V0 ff 0 ex,Vex,V V V C kT V V C q 0Sf 2 dfSP G 0 GG ∂ ∂ = ∂ ∂ ϕ = π ≈= >> , (279) which follows from above analysis, because the body capacitance C B and the body coefficient (∂V T /∂V B ) do not vary much with gate and drain biases. The experimental data in [88] imply that P VG,ex increases with gate and drain biases mostly due to faster increase of f 0 , as compared to the decrease of S VG,ex (0), which can be attributed to the change of the body capacitance with bias. Detailed analyses in [255, 259] demonstrated that the thermal noise in S IB can be also included, as well as the matching to experimental data is very well, when using the accurate models for the capacitances and resistance associated to the body of the MOS transistor. The above characterization approach was also confirmed for many SOI structures, in which the impact ionization [250, 251, 254, 255, 259] or the gate valence-band tunneling [88, 250, 253, 257] prevails, as well as for twin-gate MOS transistors [253, 257] and at different depletion levels of the body of PD SOI MOS transistor, achieved by varying the bias of the back gate of the SOI MOS transistor [88, 97, 256]. To summarize, for MOS transistors with floating partially depleted or undepleted bodies, an excess noise is coupled from the fluctuation S VB of the body potential into drain current noise S ID =(g mb )²S VB by the body-drain transconductance g mb . S VB can have several components, but the major contribution is from shot noise in body currents, owing to drain and gate leakages at high electric fields, and S VB has a Lorentzian spectrum, since the
102 of 286 shot noise is with white spectrum and it is filtered by the body impedance Z B =1/(G B +j2πfC B ), where G B is larger or equal to the conductance of body-source junction and C B depends on the layout of the structure, and it is larger than the depletion capacitance of the MOS transistor. Any additional connection to the body reduces Z B , and thus, reduces the excess noise. The excess noise is in close connection with the general equation (1) for coupling of noise, and it does not require “superficial” explanations or identification of extra generationrecombination noise sources, as one can find in old publications, e.g. [254]. Also, if the excess noise is due to body currents, then the corner frequency of the excess Lorentzian noise is the same as the corner frequency of the output AC conductance g d of the drain [255, 258], and both corner frequencies are associated with Z B and kink effect due to frequency dependent variation of body potential when the body is floating. The above discussion provides a coherent picture, when extrapolating back from partially depleted SOI MOS transistors to bulk MOS transistors. Consider again Figure 45. The body capacitance is large in bulk MOS, and the body impedance is reduced, when the body terminal is connected to the source terminal. Thus, the excess noise in bulk MOS is low, although this noise was observed [260]. VI.1.2. Fully depleted SOI MOS transistors The body is missing in fully depleted (FD) SOI MOS transistors, Figure 45c, and one cannot attribute noise to fluctuation of body potential. The experiments [255] and computer simulations [259] also showed that the kink effect and excess Lorentzian noise are suppressed when the body is fully depleted, but the FD SOI MOS transistors are not completely free from excess Lorentzian noise. This noise is just with smaller magnitude and high corner frequency (>100kHz), indicating high conductance G B of the “body”-source junction [255]. The reduction of G B is explained with reduction of “body”-source junction barrier, since the “body” potential V B at source junction is modified by the front and back gate biasing in FD SOI MOS transistors, and it is different from the Fermi potential V F set by doping of the body in partially depleted and bulk MOS transistors. Thus, the reduction of the barrier potential is ΔV rb =V B -V F and the corresponding (reverse) saturation current I B0,FD of the “body”-source junction in FD SOI MOS transistor is increased ϕ ∆ = t rb 0BFD,0B V expII , with 0BBFD,0B III >>>> (280) as compared to body current I B and junction saturation current I B0 in partially depleted MOS transistors. Consequently, at a given body current I B , the “body”-source junction conductance G B,FD in FD SOI MOS transistor increases significantly, according to t B t 0BB B t FD,0B t FD,0BB FD,B I II G III Gϕ ≈ ϕ + =>> ϕ ≈ ϕ + = , (281) and according to eq. (277), the excess Lorentzian noise has reduced magnitude and higher corner frequency, which are weakly dependent on the body current I B in FD SOI MOS transistors. An alternative explanation for the reduction of the Lorentzian noise FD SOI MOS transistors is given in [259], referring to the strong coupling between the front and back gates, which reduces the “body” transconductance g mb , and prevents the shot noise from source junction being transferred to the drain. The explanation is qualitative and does not provide insight for the shift of the Lorentzian corner frequency. Nevertheless, both explanations in [255] and [259] consider bipolar effects and modulation of potentials at body-source junction,
103 of 286 which is viewed by the dashed lines in Figure 45c. The issue is that such modulation is generally neglected in MOS transistor models, and the source-channel interface is not considered in low-frequency noise models, to the best of our knowledge. VI.1.3. Effects of oxide traps in the back gate insulator As the silicon film becomes thin in fully depleted SOI MOS transistors, the oxide traps in back gate insulator (usually called buried oxide, BOX) also contribute to the low-frequency noise, and the most common approach is to divide the conduction in MOS transistors in front and back channels. Then, one applies the noise models separately for the two channels, and the noise in the drain terminal is the sum of the contributions of the two channels plus the noise from the semiconductor bulk between them in depletion mode MOS transistors. This approach was taken long time ago, for example in [93], and results in a simple equation, such as BulkVolumeBackGateFrontGateI SSSS D ++= , (282) which is correct only if the three conduction paths are independent and isolated each from other, so that the noise source related to one of them does not affect the other. Obviously, this is not true, because the capacitances of the insulators and of the semiconductor couple the potentials from the front gate all along to the back gate, and the coupling is significant in FD SOI MOS structures, since the semiconductor film is thin. Partial solutions based on the above eq. (282) are possible and used in device characterization when the dominant conduction path is one, i.e. only front channel [88, 93, 246], only back channel [93], or only the bulk volume [93, 246], the later in the case of depletion mode MOS transistors. At crossover biasing regimes [93, 246], one always observes excess noise mostly with Lorentzian spectrum, and since the Lorentzian spectrum varies with the bias, then one attributes this variation to “superficial” traps [93, 246], while the excess noise is probably due to frequency limited coupling, similarly to the case for the excess Lorentzian noise in partially depleted SOI MOS transistors, as discussed just above. The partial solution of eq. (282) is proposed in [88] for the front channel of FD SOI MOS transistors, following the approach for noise coupling to the bottom of the channel [251]. The proposal assumes two capacitances, C F ≈C ox of the front gate oxide above the conduction channel, and below the conduction channel C S ≈(1/C d +1/C BOX ) − 1 ≈C BOX , where the back oxide capacitance C BOX is much smaller than either C ox or the capacitance C d of the depleted silicon film in FD SOI MOS transistor. Then, the coupling (attenuation) coefficient γ=∂V T /∂V B of the front gate threshold voltage V T to the back gate bias voltage V B is taken as from the capacitive divider C F −C S , given by 1 C C C C CC C ox BOX F S SF S <≈≈ + =γ , (283) where the highly conductive inversion charge layer is assumed disconnected from the source terminal, thus from ground, an assumption not stated, but seen from equations in [88]. Finally, the flat band voltage noise S FBB from back oxide is referred to the front gate as γ²S FBB and added to the flat band voltage noise S FB of the front oxide, resulting in gate referred voltage, given by FBB 2 FBVG SSS γ+= , (284) where S FB and S FBB are expressed also according to eq. (101) for the number fluctuation model, and the noise in the drain current becomes
104 of 286 FBB 2 mb FB 2 mFBB 22 mFB 2 mVG 2 mI SgSgSgSgSgS D +≡γ+== , (285) since the coupling coefficient γ=∂V T /∂V B =g mb /g m also is the ratio between the front gate transconductance g m and back gate (or body) transconductance g mb in the MOS transistor (see page 369 in [169], for example). The latter result can be derived easily from the general eq. (1) for coupling of two independent voltage noise sources into fluctuation of one current. It is worth mentioning that a guide line [88] can be provided by the substitution of the oxide trap number fluctuation model from eq. (101) in eq. (284), which gives ++λ λ +≈ γ+= 2 d BOX ox BOX t BOX,tBOX MOS Bulk VG FB FBB 2 FB SOI FD VG C C C C 1N N 1S S S 1SS . (286) Assuming that the same type of oxides are used for front and back gates, λ BOX =λ and N t,BOX =N t , the last equation suggests that the 1/f noise can increase twice in fully depleted SOI MOS transistors with thick back gate dielectrics and ultra-thin semiconductor films, when only the front gate is used as input. On the other hand, there should be a reduction of noise in double gate and Fin FETs, since the two gates are tied together, the gate area is doubled, thus S FB is a half, and the denominator in the last expression is always larger than four when C BOX =C ox . Evidently, the back-front coupling in advanced structures with increased electrostatic control should not increase the noise, provided that the quality of gate dielectrics and gate area are not reduced. Such reduction of the noise is demonstrated for Fin FET [94] and cylindrical gate-all-around MOS transistor [137], reporting low values for oxide trap density 1.5×10 17 eV − 1 cm − 3 and less than 0.5×10 17 eV − 1 cm − 3 , respectively. Also, the analysis of the 1/f noise from the back gate in FD SOI MOS transistors indicates one more time a convergence between noise models for different structures in a close relation with the general principle of noise coupling, expressed by eqs. (1), (2), (3), (4) and (5). We demonstrate below that the capacitive coupling in the semiconductor body can reduce the noise in multiple gate MOS transistors, by using eq. (4) with the DC current cancelled. VI.1.4. Capacitive coupling in the semiconductor body Bulk MOS transistor Consider first the obvious bulk MOS transistor with gate capacitance WLC ox and small depletion capacitance WLC d , as depicted on the top in Figure 47. As a statistical variance, the oxide charge fluctuation S Q =WLS Qo is proportional to the gate area WL, where S Qo is the charge fluctuation per unit area, and it is a constant at an assumption for uniform spatial distribution of oxide traps. At given biasing condition, the transconductance eq. (4) is g=g m =g o W/L, where g o is the transconductance of square-shaped gate. The capacitance in eq. (4) is the capacitance seen by the charges at semiconductor-dielectric interface and C=WL(C ox +C d )≈WLC ox . The substitution in eq. (4) gives the equation for the drain current noise S ID | N=1, one gate , given by ( ) [ ] 2 ox Qo 2 m 2 dox Qo 2 2 o 2 Q 2 mgate one ,1N D I WLC Sg CCWL WLS L W g C S gS ≈ + == = (287)
111 of 286 [ ] ( ) j )j(i 4 ci c 2 cH 2 i 2 ci 2 i 2 i j tot,F j 2 tot tot L 'n L L r r 1 Kf I S ρ+ρ αρ+αρ ρ+ρ π == (304) Similarly to eq. (300) for the resistances, the system of equations in eq. (304) has two unknowns, Hooge parameter α H in nanowires and characteristic unit-area noise α c [cm²] for the contact to the nanowires, and ideally two samples are needed to determine α H and α c . And also similarly, in practice one needs to carry out a guided optimization-discrimination procedure over several samples, in order to obtain reasonable values for α H and α c in each sample. Note that even after careful characterization in [271], the values of the parameters related to the noise in nanowires and contacts scatter over several decades in fig.6 in this publication, while the values for the resistance scatter within less than a half decade in worse case in fig.4 for the same samples. Evidently, the results in [271] symptomatically imply that the issues in nanowire and nanotube devices are related mostly to the access to these devices. If there is a contact problem for the DC performance of these devices, increasing the device resistance with 50%, then this problem actually dictates the low-frequency noise, pushing the noise 1-3 decades above the noise levels in the nano-conductor. The solutions of the contact problems were also recognized in ITRS predictions [3] as enabling factor to have access to advanced nano-devices. The problem with the contact is even more pronounced in single carbon nanotube devices. In [272], multiwall carbon nanotubes (MWCNT) were placed to cross the gap between gold contact pads, using atomic force microscope (AFM), and the resistance and low-frequency noise of these samples were measured at different low bias currents (I DC ) from room temperature to cryogenic temperatures 77K and 4.2 K, down to 1.5K. It is observed in [272] that the low-frequency noise is 1/f at room temperature and at 77K, but at 4.2K and 1.5K, the noise spectra showed large Lorentzian components, as illustrated in the left-hand plot of Figure 50. Interesting observations for the behavior of the samples are made in [272] at 4.2K and 1.5K, when changing the direction of the bias current, as illustrated in the right-hand plots of Figure 50. The top plot shows that the time constants of the Lorentzian spectra are bias dependent, they are different at different bias polarities, but the time constants do not depend on the temperature. The middle plot shows that the prefactor S 0 /τ=4(ΔI/I DC )²τ/(τ 1 +τ 2 ) also varies with bias, while from eq. (179) one expects (ΔI/I DC )=constant for a particular trap and τ/(τ 1 +τ 2 ) =F(1-F)=constant in conductive materials, in which the Fermi level, and thus the trap occupancy, are weak function of the bias. In the bottom plot, the contact resistance is larger than the resistance of the nanotube. The differential resistance is bias dependent, nearly ∝1/I DC for |I DC |<1μA, and also not fully symmetrical in respect to the direction of the bias current. In addition, the resistance is proportional to ~7mV/I DC , rather than to 0.13mV/I DC , which one would expect from the ratio φ t /I DC with thermal voltage φ t ≈0.13mV at temperature T=1.5K. This indicates tunneling junctions at the contacts, accompanied with Coulomb blockade, charge trapping and strong bias dependent Random Telegraph Signal (RTS) noise, which resulted in Lorentzian noise spectra at temperature 4.2K and below. Provided that 4.2K is very low temperature, then Shockley–Read–Hall (SRH) statistics for the process of trapping and de-trapping is not realistic, because the thermal velocity ν th is low in eq. (63) and the experiments in
112 of 286 [272] did not show temperature dependence between 1.5K and 4.2K. A variation of the prefactor S 0 /τ with the bias is not expected from SRH statistics, since the Fermi level is almost independent of bias in conductive materials and the trap occupancy, c.f. F(1-F) in eq. (179), should not vary with bias in these materials, unless there is a junction interface and potential bending in it. Therefore, RTS is explained in [272] with trapping and de-trapping in terms of tunneling to charge trap at the contact interface between nanotube and metal. So, instead of using of SRH relations from eq. (63), the capture and emission time constants are expressed in terms of tunneling, given by ( ) −τ= γ− λω =τ o DC o o I I expV1 d exp 1 . (305) Here, ω o is attempt frequency, d is effective tunneling distance, λ is tunneling attenuation distance, γ is parameter related to the shape of tunneling barrier, τ o =exp(d/λ)/ω o is tunneling time constant at equilibrium (no bias), and the characteristic tunneling current I o =λ/(dγR eff ) is a fitting parameter that accounts for the effective resistance R eff =V/I DC of the region, where the tunneling occurs. While the slopes for τ as function of the bias current in the right-hand plot on top of Figure 50 confirm the exponential dependence of eq. (305), an interesting observation in [272] is that the fitting parameters τ o and I o in eq. (305) are different for different time constants and they are not the same for capture and emission time constants τ 1 and τ 2 , since the prefactor S 0 /τ=4(ΔI/I DC )²τ/(τ 1 +τ 2 ) also varies with the bias, as shown in the right-hand plot of Figure 50 in the middle. This implies that tunneling distances, the attenuation distances and the barrier shapes vary, which is possible by having junctions at the weak contacts between nanotube and metal pads. Nanotube FETs The low-frequency noise in field-effect transistor configuration of nanotubes was recently also addressed. In these devices, one or several nanotubes, or thin film of random nanotube network bridges between metal contacts, and the gate is usually a conductive substrate, on the top of which the gate oxide and the metal pads are placed. These devices appear to be very noisy, affecting even DC measurements, as illustrated with several “I−V” curves in Figure 51. In Figure 51(a) and (b), transfer I−V characteristics of field-effect transistors (FETs) based on single CNT measured at room temperature are shown. The values for the currents scatter between 10% and 20% around a trend. The trend is similar to the I−V curve of pMOS transistors. Since the scattering is large in Figure 51(a), the current and the threshold voltage (cross point of two lines) are analyzed in terms of stochastic resonance in [273], demonstrating that CNT FET can be used as a detector for signals below threshold voltage. Figure 51(b) illustrates that the current and its scattering are larger in ambient atmosphere, and they are reduced in vacuum [274]. Figure 51(c) depicts the cracked “I−V” curves of single CNT FET at cryogenic temperatures [171], owing to giant and bias dependent RTS noise with amplitudes 30% to 60% of “DC” current. Note that the currents and the transitions between the segments in the plot are different at opposite directions of the current flow, which is similar to the observations for multiwall CNT devices, discussed just above. On the other hand, in contrast to single CNT devices, the I−V characteristics in Figure 51(d) are smooth when the FET is based on a random network of single-wall CNTs [275]. A hysteresis is evident when the device was operating in air, using the silicon wafer as solid-state gate. Interestingly, the hysteresis is reduced when a liquid solution is used to mediate between electrochemical gate comprised by a pair of reference (Ag/AgCl 4M KCl or a saturated calomel) and by
113 of 286 a working (Pt) electrodes. The improvement when using liquid gate was attributed in [275] to enhanced electrostatic control and suppression of charge trapping effects. It was observed that the threshold voltage of the liquid gate CNT FET is a function of pH and concentrations in the chemical solution. Referring to the discussion on the multiwall CNT given above, we note, however, that the problems in the single CNT FETs can be due to contacts, rather than due to traps around the CNT, since the contact of metal to a network of CNTs and addition of electrolyte at this contact, as it was in [275], would greatly improve the repeatability of the contact. Unfortunately, the noise from the contact was not addressed in [171, 273, 274, 275], perhaps, due to a lack of scalable model for noise from contacts when interfacing 1D to 3D current transports. Obviously following the style in the initial publication on noise for CNT films [267], the noise is first phenomenologically investigated, and then related to noise model for mesoscopic devices, e.g. for MOS transistor noise model, relying on the similarity that exists to some extend between CNT FET and MOS transistors. Note again in Figure 51 that all devices at all measurement conditions behave similarly to p-type MOS transistors – a fact, which is experimentally observed and reported many times in the literature, although not very well justified theoretically, since the band structure of semiconducting CNT is quite symmetrical above and below the band-gap – see fig. 3 in [270], for example. To illustrate the situation when investigating the noise in CNT FET and MOS transistors, we present a typical outcome from noise experiments in Figure 52. The data are from [274] for measurement of single CNT FET in vacuum (18mTorr) at room temperature. It is stated in the publication that this CNT FET has 4μm gap between gold electrodes and it was made as described in [171]. That is, CNT diameter is d=1−3nm and the gate is silicon wafer with t ox =500nm thermal SiO 2 . Considering the information from the publication, the CNT FET has length L=4μm equal the gap between electrodes, and taking an average diameter d=2nm, the channel width of the CNT FET is W=πd≈6.3nm. Assume that there is no gap between wafer surface and CNT, and consider d<<t ox , then the gate capacitance per unit area is C ox =7nF/cm² for the SiO 2 gate dielectric with thickness t ox =500nm. The experiments in [274] were carried out in linear mode of operation of CNT FET, at low drain bias voltage |V D |<(|V G -V T |-0.5V). Therefore, a rough estimation for the total number of carriers n can be made by assumption for uniform charge density in CNT, resulting in n=WL|V G −V T |C ox /q, which is approximately 11 electron charges per one volt of gate overdrive voltage |V G −V T |. Figure 52a presents results based on measurement, in which the gate bias was varied, while the low drain bias was constant −V D =0.1V. For the set of biasing points {V G , I D }, both the input referred (gate) noise voltage S VG and output (drain) noise current S ID are reported in [274]. From these, we obtain the transconductance g m =(S ID /S VG ) 0.5 ≈17nA/V (almost constant, as expected for linear mode of operation of FET transistors) and draw the evolution of the ratio I D /g m with bias current in the Figure 52a on top. For operation of FET in linear mode, we expect I D /g m =|V G −V T | and I D ∝|V G −V T |, and we observe linear dependences with slope 60MΩ for I D /g m and slope 40MΩ=1/[(W/L)μC ox V D ] for the relation between |V G −V T | and I D as function of the bias current I D . From the latter dependence, |V G −V T | vs. I D , shown just under the plot for I D /g m , we have estimated mobility μ≈24000 cm²/Vs. Comparing to crystal semiconductors, the value for mobility is impressive, but it is somehow in the middle of the range 4000-120000 cm²/Vs reported in [270], thus it is reasonable. Having the above information for the sample handy, we pursue analysis of noise in terms of mesoscopic noise models for MOS transistors.
114 of 286 For the number fluctuation model with correlated mobility fluctuation, as shown in the middle of Figure 52a, we plot the square root √S VG of the power spectrum density of the input (gate) noise voltage S VG at 1Hz, referring the reported data for S VG from 40Hz to 1Hz, by S VG (1Hz)=40Hz×S VG (40Hz). The constant in the linear fit to √S VG yields flat-band noise voltage, and from eqs.(101), (109) for noise from tunneling and trapping in gate oxide, we get Hz/VeVcm104.5]cmeV[N WLHz1 NkT C q Hz/V1036)Hz1(S 232231 t t 2 ox 26 FB −−− − ××= × λ = ×= , (306) when substituting the values for the parameters of sample and using kT=0.026eV for thermal energy at room temperature and tunneling attenuation distance λ=0.1nm. Therefore, we obtain N t =(36×10 − 6 V²/Hz)/(5.4×10 − 22 eVcm³V²/Hz)=6.7×10 16 eV − 1 cm − 3 , which is a reasonable value for the trap density is SiO 2 , see Figure 42, despite that S FB is high. To evaluate the parameters for the correlated mobility fluctuation, see the discussion after eq. (121), we use the bias dependent term in the linear fit of √S VG , which is I D 0.4MΩ. Since |V G −V T |= I D 40MΩ, as seen from the transfer |V G −V T | vs. I D curve in Figure 52a, then TGTG FB V VV1 Hz/mV6 M40Hz/M4.0 VV1 S S G −θ+= ΩΩ −+= , (307) or θ=0.01/0.006=1.67V − 1 , which is in the range observed for MOS transistors (see Figure 37). There are two definitions for scattering parameter α s , as shown in eq. (243). Using one of them, given by eqs.(127) and (137), then α s =qθ/(μC ox )=1.6×10 − 15 Vs. The other definition gives α s =θ/(μC ox )=10 4 Vs/C. Both values are within ranges deduced from Si MOS transistors – see eq.(243). Thus, both by N t and α s , the number fluctuation can be justified for CNT FET. The 1/f noise in CNT FET can be justified also in terms of intrinsic (Hooge) noise. We calculate the SPICE parameter K F =f×S ID /I D ² and plot it in the bottom of Figure 52a. Obviously, from eqs.(6) and (15), K F =α H /n decreases, when the total number of charges n=11×|V G −V T | increases with gate bias, as mentioned above. The calculated values for the Hooge parameter α H are shown above the plot for K F in Figure 52a. The values scatter, but the average α H ~2×10 − 3 is a reasonable value for conductors, and by this value, the 1/f noise in CNT FET can be also justified as mobility noise. The above discussion implies that 1/f noise in CNT FET is easily explained in terms of downscaled mesoscopic models for Δn and Δμ fluctuation, when data as function of gate bias are used. However, in the published analyses, which are similar to the above, there are details which are neglected. We show these details in Figure 52b. In this figure, from top to bottom, although V D <<V G , the relation between gate bias and drain current is not linear, the relation between drain bias and drain current has a step at low voltage, the normalized noise in terms of K F is function of drain bias at low drain bias levels, and K F obtained from experiment with variable gate bias is different from K F obtained from experiment with variation of drain bias, even for the bias point {–V G =2V,
115 of 286 −V D =0.1V}, which was common in the two experiments. For this bias point, even the DC currents were different, as shown with arrows in the figure. All these indicate that the access contact to CNT, trapping and barrier at it, might be significant noise sources, as deduced for CNT films, networks and multiwall CNT [268, 272]. To the best of our knowledge, there is no publication that explains these “small” details in the behavior and noise related to them in CNT FET. The explanations for the contact effects and functionalized surface of CNT by adsorption or by changing of chemical environment (air or other atmosphere [274], or pH of liquid [275, 276]) on noise in CNT FET are still qualitative. Analysis of the ratio K F /R≡A/R We address now the consequences from the empirical observation made in [267] that ratio K F /R≡A/R is approximately constant in CNT devices – see again eqs. (15), (292) and (293), where A≡K F is the SPICE parameter as defined in eq. (11). The CNT devices are typically arranged in thin-film structures, as shown in Figure 53a, b and c. The CNT networks have conductive branches, which are shown with arrows in these figures. Since the transport in CNT is 1-D, then the conduction branches are independent each from other. Having on average L carbon nanotubes in each conduction branch and W conduction branches in the device, then we can represent the CNT percolation network by and idealized resistor network, as shown in Figure 53d. For simplicity, we will assume that the resistance R O and the noise v o ² of each CNT have the same values, that the number L of serially connected CNTs in each conduction branch is the same in the network, that the number of identical parallel conduction branches is W, and that the noise v c ², which may originate to contact between nanotube and metal electrode, is the same for every single conduction branch. By these idealizations, when applying external bias voltage V (or current I), the DC and noise currents in each branch are I O =I/W=V/(LR O ) and i o ²=(v c ²+Lv o ²)/R O ²=i²/W, respectively, and one can easily find the total resistance R=V/I=LR O /W and voltage noise v²=R²i² of the circuit for the case of current biasing I. In this way, the normalized noise of the CNT network is +=+===== 2 2 o 2 2 c O 2 2 o 2 2 c 2 2 2 2 2 2 norm V v V Lv R R V v W L V v W 1 ... R r I i V v S , (308) and in terms of the empirically observation in [267] that K F /R≡A/R is approximately constant, we get +∝= 2 2 o 2 2 c O F norm V v V Lv R 1 R K R S f ≈constant, (309) where we assume 1/f noise, and the parameters resistance R O =constant of single CNT and number L=constant of nanotubes in a conductive branch are constants for given sample (and given gate bias, if the CNT device is a TFT transistor). Note that S norm /R and K F /R do not depend on the number W of parallel conductive branches. It is evident from eq. (309) that the empirical observation in [267] requires two conditions for the noise sources in CNT networks. The first condition is that the noise from individual CNT has to scale with the bias, but not with the number L of serially connected CNTs. The second condition is that the noise from contacts also has to scale with bias and it has to decrease with the number L of serially connected CNTs. The bias dependence is easily reproducible by mesoscopic models for noise. The dependence on the number of CNTs serially connected in conductive branch is, however, not. Let us take the first condition, for example, and try to analyze in terms of Hooge eq.(15) at given bias voltage, assuming number of carriers n o in a single
116 of 286 nanotube. For branch with one nanotube we have v o1 ²/V²=α H /n o . For a branch with L nanotubes, Lv oL ²/V²=α H /Ln o , according to the equivalent circuit of the network. In contrast to the expectation from eq. (309) for v oL ²/V²= v o1 ²/V², we get v oL ²/V²=α H /L²n o ≠α H /n o =v o1 ²/V². The condition will be satisfied, if one assumes that the total number Ln o of carriers in the branch decreases when increasing the number L of serially connected CNTs, that is n o ∝1/L². We did not find a way or publication to justify physically such dependence for the number of carriers, although many publications use the empirical observation in [267] for comparisons. The issue is that n o ∝1/L² dependence cannot be derived from any mesoscopic model for noise. Nevertheless, it seems that the noise in nanowire [271, 277, 278, 279, 280, 281] and CNT [268, 272, 274, 275, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292] devices scales according to the rules for mesoscopic devices. This is illustrated in Figure 54. The publications discussed above argue that the dominant sources of 1/f noise in CNT and nanowire devices are number fluctuation and from contacts. Therefore, we calculate the surface area and cross-section area of the CNT and nanowires, and plot versus these areas and versus data for MOS and bipolar transistors in Figure 54a and b, respectively. When the information for the size and number of CNTs in the devices was not stated, we assumed that CNT diameter is 3nm and the resistance due to one CNT is about 100kΩ. Interestingly, the data in Figure 54 show that the geometrical scaling rule for 1/f noise, the smaller is the area – the noisier the device is, also applies for nanowires and nanotube devices. The comparison of surface area dependence of the 1/f noise in Figure 54a to data for MOS transistors implies that the noise in nanowire (NW) and carbon nanotube (CNT) devices can be explained in a manner similar to the models for the 1/f noise in MOS transistors, because the noise in NW and CNT devices is less than and in the range of the noise in MOS transistors, although the scattering is large (σ dB =9.7dB). Thus, one can assume surface origin for the noise in NW and CNT devices. The comparison of cross-section area dependence of the 1/f noise in Figure 54b to data for bipolar transistors implies that the contact noise in NW and CNT is less than the noise in bipolar transistors, the scattering of the data is less (σ dB =5.6 dB ), but the noise rapidly increases in single-wall CNT FETs with single or small count of CNTs. Thus, one can deduce a crossover between dominant noise sources, since the normalized noise (K F ) increases steeper than 1/area in small-area CNT devices, as compared to the lower noise in NW samples with several parallel nanowires. Qualitatively, the crossover is from bulk noise in NW devices, to surface noise in multiple CNT, toward injection noise (either tunneling or thermionic) in single CNT devices. The published data scatter over several decades, and a reliable estimate for the crossover points is not possible. Therefore, one can find a variety of models and explanations for the noise in semiconductor nanowire (NW) and carbon nanotube (CNT) devices, which causes difficulties when comparing devices from different publications. Again looking at Figure 54, one can see that the CNT devices are noisy, but in relative units, not noisier than MOS and BJT, if the latter are scaled down to the sizes of carbon nanotubes. This demonstrates one more time that there is convergence of noise models and behaviors from very large down to very small devices. The issue is that the normalized 1/f noise (at 1Hz) in nano-devices is larger than K F >10 − 4 . Therefore, these devices will be difficult to use deterministically, since for an application that requires 4-5 frequency decades, the peak-to-peak noise becomes more than 20%, according to
117 of 286 ≈ − min max F DC pkpk f f lnK6 V V . (310) In three sentences, although the physical origin is not very well determined, the 1/f noise in semiconductor nanowire (NW) and carbon nanotube (CNT) devices scales according the rules for mesoscopic devices. Since the area (either surface, or cross-section of contacts) of NW and CNT is very small, then the 1/f noise in these devices is a limiting factor for the use in the practice. The single nanotube devices seem are not anymore deterministic, that is, they are behind the down-scaling barrier set by the 1/f noise. VI.3. Between 3D and 1D – the graphene and transition metal dichalcogenide 2D transistors The reduction of the mobility in ultra-thin silicon body SOI and FINFETs (transistors with 3D charge transport), and the difficulties in the mass-production of nanowire and nanotube transistors with 1D charge transport, brought the interest in exploring graphene and transition metal dichalcogenide transistors, which have semiconducting “body” of single to few atomic layers and 2D charge transport. The 2D transistors attempt to utilize the better electrostatic control in field-effect transistors with thinner body, the high intrinsic mobility of the graphene and the apparent advantage of atomic layer growth of metal dichalcogenides, along with the compatibility with the lithography for planar devices in the microelectronic manufacturing. However, the properties of the 2D semiconducting layers deviate from the properties of well-understood crystalline layers in the 3D silicon transistors, inheriting also from the quantum effects in the 1D nanowires and nanotubes. Typical issues in the 2D transistors are the poor contact with the metal electrodes of the device terminals and the non-covalent (van der Waals) bonding between the atomic layers. The latter, basically, implies that the 2D semiconductor is a stack of several atomic layers with increased spacing and energy barriers between the atomic layers, but not a homogeneous layer as in the crystalline 3D semiconductors. Below, we illustrate the consequences for the low-frequency noise in 2D transistors with an example from [293] for a MoS 2 (a metal dichalcogenide) transistor. Figure 55 (a) shows the barriers ϕ v in the energy diagram and the spacing d c between the MoS 2 monolayers in the spatial cross-section schematic diagram. The barriers and the spacing are due to van der Waals bonding between the MoS 2 monolayers, which is weaker than the covalent bonding in the MoS 2 monolayers and in the crystalline 3D semiconductors. According to these diagrams, the authors of [293] consider the following physics and relations. The noise is due to Δn fluctuation of the trapping in the gate oxide, combining several processes that affect the time constants of the trapping and the noise measurement. One process is the Shockley–Read–Hall (SRH) recombination at semiconductor-dielectric interfaces with time constant τ o ∝1/n inversely proportional to the carrier density n∝(V GS -V T ) in the semiconducting layers and the gate overdrive voltage (V GS -V T ), thereof. A second process is the charge tunneling to/from traps at different distances in the gate dielectric, which randomizes τ o in a range of larger values, resulting in band-limited 1/f noise. A third process is the additional increase of the time constant values for charges from semiconducting monolayers non-adjacent with the gate dielectric, owing to the energy barriers and spacing due to van der Waals bonding of the semiconductor monolayers. The fourth consideration is that 1/f noise is band limited and measurable only when the frequency band of the spectrometer (2Hz – 1000Hz) and the band limited 1/f noise overlap. This fourth consideration is essential for the explanation of the non-monotonic dependence of the normalized noise S o as function of the gate bias V GS shown in Figure 55 (b).
118 of 286 Figure 55 (b) shows the normalized noise S o referred to 1Hz, S o =average(f×S I (f)/I DS ²), averaged over logarithmically spaced frequencies f in the range from 2Hz to 1000Hz, vs. the gate bias voltage V GS . From left, S o reduces with the gate overdrive voltage (V GS -V T ), because the time constant τ o ∝1/n of the (first) Shockley– Read–Hall (SRH) process is inversely proportional to the carrier density n’∝(V GS -V T ) in the monolayer adjacent with the gate dielectric. The measured noise is 1/f, owing to the (second) process of the charge tunneling to/from traps at different distances in the gate dielectric, which randomizes τ o in a range, resulting in band-limited 1/f noise. One can deduce mathematically the reduction of S o and higher n’ by using fraction of τ∝1/n’ in the numerator of the integrand in eq. (73). The (third) process of additional increase of the time constant values for charges from semiconducting monolayers non-adjacent with the gate dielectric brings the upper boundary 1/(2πτ” min ) of frequency range of the band-limited 1/f noise from the non-adjacent semiconducting monolayer below the lower boundary of 2Hz of the spectrometer (the fourth consideration above), when the carrier concentration n”∝(V GS -V T ) in the non-adjacent monolayer is low at low gate overdrive voltage (V GS -V T ). Therefore, the noise associated with the non-adjacent semiconducting monolayer was not measured and missing in the left-hand side of the plot of S o in Figure 55 (b). However, increasing the overdrive voltage (V GS -V T ), the carrier concentration n”∝(V GS -V T ) in the non-adjacent monolayer increases, τ” min ∝1/n”, the upper boundary 1/(2πτ” min ) of the frequency range of the band-limited 1/f noise increases, reaching the spectrometer range 2Hz-1000Hz at V GS ≈5V, and the overlap of the 1/f noise and spectrometer frequency ranges increases, resulting in increase of S o to a peak value at V GS =20V. At higher V GS >20V, the overlap of the frequency ranges of the 1/f noise from the non-adjacent semiconducting monolayer and the spectrometer is full, but S o reduces at increasing V GS , because the time constants τ”∝1/n” of the noise, owing to the (first) Shockley–Read–Hall (SRH) process (as above for the noise from the adjacent monolayer in the left-hand side of the plot of S o in Figure 55 (b)) In summary, the low-frequency noise in 2D transistors follows the noise behavior in 3D transistors, but energy barriers and spatial spacing between monolayers bring the noise parts from different monolayers outside the ranges for noise spectrum measurement, which may cause apparently spurious non-monotonic data series for the noise levels, e.g., as function of bias, as shown in Figure 55 (b). The main difference from the noise in the 3D transistors is that each monolayer in the semiconductor film of the 2D transistor is likely contributing by bandlimited 1/f noise in different frequency ranges, and the measurements can miss band-limited 1/f noise at very low frequency, e.g., below 1Hz. Thus, an extrapolation of 1/f noise spectra measured at higher frequency toward lower frequency is uncertain for 2D transistors. The band-limited 1/f noise in 2D transistors is actually a predicator to the “peculiarities” observed in 1D nanowire and carbon nanotube transistors, discussed in the preceding Sec. VI.2. Nanotubes and nanowires – 1D seems too noisy. VII. Impact of LFN in circuits The impact of low-frequency noise (LFN) depends on the purpose of the electronic circuit. Since the variety of electronic circuits is large, then it is generally impossible to look at every single case of application. We have selected two of them: radio-frequency (RF) circuits and sensors. Current efforts in RF circuits are to reduce the supply and power of the electronic circuits and to increase their speed. In the first part of this section, we shall discuss the impact of the low-frequency noise on the performance of low-voltage and low-power circuits and the up-conversion of LFN in RF circuits. Finally, the implications of the noise in sensors are briefly addressed in
119 of 286 sub-section VII.3. Noise in sensors. VII.1. Trading power for noise The advances in CMOS technologies put them as the preferable choice for building low-voltage and low-power electronic circuits. This is because the threshold voltages (~0.2V-0.4V) of modern MOS transistors is a fraction of the turn-on voltage (~0.6V) of silicon bipolar transistors (BJT), the static current consumption of CMOS pairs is negligible especially in digital circuits and the input gate leakage current of MOS transistors is much lower than the base current of BJT at given input bias. Also, the diversity of functions integrated in CMOS circuits is larger than that in BJT circuits at higher density of integration. However, the low-frequency noise emerges as a problem in low-voltage and low-power electronic circuits. Some manufacturers of integrated circuits provide in datasheets, for example in [267], that the product I Q ×S VIN of the quiescent current I Q and input referred noise S VIN is a good figure of merit for the noise in amplifiers, but it is not possible to minimize the product below certain limit. Here, we shall study details that are related to the product I Q ×S VIN in BJT and MOS amplifiers by using characteristic values deduced in previous sections for the parameters of the transistors. The circuit in Figure 56 is a typical topology of low-frequency amplifier with voltage feedback. The transistors, which mostly determine the noise performance, are the amplification transistor T A and the loading transistor T L in the first differential stage. These transistors are surrounded by a dashed line and can be MOS or BJT in BiCMOS technologies, as depicted in the figure. For simplicity, assume that the common node in the differential pair T A −T’ A is grounded for AC signals and the voltage V i and the resistance R i are ½ of the actual voltage magnitude and impedance of signal source connected between input nodes IN−IN’. The differential amplifier DA in the second stage usually is with low impedance R L , and DA suppresses the noise from biasing current source I 1 . The noise from reference circuit can be filtered out by the capacitor C connected to node REF. When the loading transistors are identical, the noise from node REF also results in in-phase signal, which is suppressed by DA. The noise contribution of the amplification transistor T A at the input terminal IN of the circuit has voltage S VIN and current S IIN components. The input referred voltage noise S V,TA of the amplification transistor T A contributes directly to S VIN . To the first order of approximation, the current component is S IIN =S V,TA /(g IN )², where is g IN is the input conductance of the transistor T A . Since the loading transistor T L does not have a connection to the input, then T L does not contribute to input noise current, and the noise from the T L is referred to the input node IN of the circuit in Figure 56 as a voltage noise by S I , TL /(g m,TA )², where S I , TL is the output noise current of T L . Therefore, the total input referred voltage of the amplifier noise is += +=+= 2 T,m T,m T,V T,V T,V 2 T,m T,m T,VT,V 2T,m T,I T,VV A L A L A A L LA A L AIN g g S S 1S g g SS g S SS , (311) where g m,TA and g m,TL are transconductances and S V,TA and S V,TL are input referred noise voltages of transistors T A and T L , respectively. From this equation is clear that the relative contribution of T L to the input referred noise of the amplifier is proportional to the ratio of noise levels in T L to T A and it is a quadratic function of the ratio of the transistors’ transconductances. Note that the ratio g m,TA /g m,TL cannot be varied freely, because the same DC
120 of 286 current flows through T A and T L , since T A and T L are connected in series, as seen in Figure 56, and I TA =I TL =I DC =I 1 /2. Depending on whether T A and T L are MOS or BJT, S V,TA and S V,TL are given later by eqs. (314) and (315) for 1/f noise, and by eqs. (317), (319) and (320) for white noise. The current noise at the input of the circuit is π βϕ = ϕ ≈+= MOSfor ,fWLC2 BJTfor , I I g with ,SgSS ox t C t B IN I 2 INT,VI LEAKAIN , (312) We have discussed this conversion for BJT by eqs. (19) and (20), and will illustrate again with several examples below, when the conversion holds. The additional noise from gate leakage or protection diode leakage is LEAK 2 LEAK F I qI2I f K S LEAK LEAK += , (313) which is the sum of 1/f and shot noise components, since the noise is due to overcoming of junction or insulator barrier – see eqs. (215) and (220) for gate leakage. For simplicity, we will neglect the noise from leakage, although it can be significant in MOS transistors with very thin gate insulators. The characteristic relations and values for the parameters of the transistors, as deduced from the previous sections, are now summarized. Form the trends in Figure 15, discussed by eqs. (213) and (214), the input referred 1/f noise voltage of a transistor is WL Hz/Vm103.1 f Hz1 WL FOM f Hz1 S 229 S f/1,V VG G µ× == − , for a MOS transistor, (314) or E 2212 E S f/1,V A Hz/Vm108.3 f Hz1 A FOM f Hz1 S VB B µ× == − , for a BJT. (315) Here, WL is the gate area of MOS transistor and A E is the emitter area of BJT. The term 1Hz is added to match the dimensions. One expects that the input referred 1/f noise voltage of the amplifier will be higher, if replacing the BJTs of emitter area A E with MOS transistors of same gate area WL~A E , according to WL A 300~ WL A Hz/Vm108.3 Hz/Vm103.1 S S EE 2212 229 f/1 V V B G µ× µ× = − − . (316) When using the data from ITRS [3] shown in Figure 1a, the ratio in the last equation is between 100A E /(WL) and 300A E /(WL), which implies that one should use large MOS transistor in order to achieve the low-noise performance of smaller-area BJT in terms of input referred 1/f noise voltage. As shown by eq. (263), the white noise in the collector current is a sum of the collector current shot noise and the coupled shot noise from the base current. When referred to the input base terminal by the transconductance of BJT, the input referred white noise voltage of BJT is
127 of 286 nMOS from nodes with L min =90nm to 130nm (EOT~2.8nm, C ox ~1.2μF/cm², μ~200cm²/Vs). The triangles are for I DO =1μA and L=0.2μm, and correspond to advanced MOS nodes with L min <65nm (EOT~1.4nm, C ox ~2.5μF/cm², μ~300cm²/Vs), which is expected to be used also in analog applications in the near future. For convenience, and in order to show the gate overdrive (V G −V T ) in the horizontal axis at the top of Figure 57, the plots for MOS amplifiers are given versus the level of channel inversion (I Dsq /I DO ), which is a dimension-less quantity, rather than versus current density. The corner frequency f c between 1/f noise and white noise is then plotted at the top-right in Figure 57, for MOS amplifiers that use transistors with the abovementioned parameters, according to ( ) ( ) ϕ>−> ϕ − ≈ ϕ>ϕ−<− ϕ − ≈ × ϕ ≈ ++ϕ = tTGD t TG DO Dsq tDtTG t TG DO Dsq 2 t 2 DO S DO Dsq 2 t DO Dsq 2 DO S c 2VVVfor , 2 VV I I 3V and ,2VVfor , VV exp I I q2 L I FOM Hz1 I I 25.05.0q2 I I L I FOM Hz1f VG VG (337) as follows from eq. (336) for the condition S(1/f)=S(white). Note that the plots for MOS amplifier in Figure 57 are for low-impedance case, when the signal source resistance R i <<2πfWLC ox , and the plots represent the input referred voltage noise of MOS amplifier, whereas, the plots for BJT amplifier are for high-impedance case, when the signal source resistance R i >>z B =βφ t /I C , and the plots represent the input current noise of BJT amplifier via S VIN =S IIN z B ², as discussed earlier after eq. (327). Several useful comparisons between MOS and BJT amplifiers can be made in Figure 57 for the product I DC S VIN . For the white noise, the product is independent of transistor size, but increases with bias level (I Dsq /I DO ) in MOS amplifiers, whereas, the product is independent of bias level J C =I C /A E , but it depends on the current gain β of BJT used in the amplifier. For the 1/f noise, the product increases with the bias level in both MOS and BJT amplifiers, but the product is independent of the current gain β in BJT amplifiers, whereas, it increases as I DO /L²∝μC ox /L² in MOS amplifiers. The corner frequency f c decreases with the current gain β in BJT amplifiers, owing to higher white noise, whereas, f c increases as C ox /L² in MOS amplifiers, owing to higher 1/f noise. Consequently, f c is usually in the kHz range for BJT amplifiers, whereas, f c can reach GHz range in MOS amplifiers. Thus, one can observe both 1/f noise and white noise in low-frequency BJT amplifiers, whereas, the 1/f noise dominates in the entire low-frequency range below 1MHz in MOS amplifiers with short gates, e.g. L<1μm. Note again that the comparison is between voltage noise in MOS amplifiers and current noise in BJT amplifiers, the later multiplied by the square of input resistance of BJT. The relations for the product I DC S VIN in eq. (326) for BJT amplifiers and in eq. (336) for MOS amplifiers are generic, and we illustrate in Figure 58 how they are reflected in commercially available integrated amplifiers from different fabricators, Analog Devices, Linear Technology, and Texas Instruments and Burr-Brown. From the datasheets of the amplifiers, we have collected the values for 1/f and white noise, maximum quiescent
128 of 286 current I Q and bandwidth. As shown in the left-hand plots of Figure 58, on the top for BJT and below for MOS amplifiers, increasing I Q , the bandwidth increases, the white noise decreases, while the 1/f noise scatters generally. Obviously, the datasheets provide performance parameters of the amplifiers and the details for the designs are omitted. Since the current I DC and the sizes of the transistors in the input stage are not stated in datasheets, then we assume that I DC is a portion of the quiescent current I Q of the amplifier, and I DC scales approximately linearly with I Q , in order to preserve the bandwidth of the amplifier. Then, the current density in the input transistors of BJT amplifiers is estimated from the corner frequency f c between 1/f and white noise, according to eq. (327), and using typical values for β=150 and FOM SVB =3.8×10 − 12 μm²V²/Hz from eq. (315). Re-arranging the data against the current density, as illustrated at the top-right plot of Figure 58, the product 0.5I Q ×S VIN for BJT amplifiers fits with the predictions of eq. (326) given earlier for the design window in Figure 57. In particular, the 1/f noise is in agreement with the general trend (solid squares) for 2J C FOM SVB , having a distribution illustrated with bellshaped curve with standard deviation 5.5dB on right of the trend. The product 0.5I Q ×S VIN for white noise (triangles) also became with trend independent of the current density, having a distribution with a peak within the limits given by the design window for β=50 (solid diamonds) and β=500 (solid triangles), and standard deviation also 5.5dB. The similar values for the standard deviations for 1/f noise and white noise, as well as the increase of the bandwidth as function of current density, imply that the scattering around the trends is due to differences in the design of the amplifiers, but in general, the predictions from of eqs. (326) and (327) hold for BJT amplifiers, by taking 50% of I Q in the product I DC ×S VIN . The high value of 50% is empirical, and it is justified by the fact that the input referred voltage noise of BJT is normally lower in low impedance circuits, see again the discussion after eq. (327), and one has to take 10-20 times the biasing current I DC of the input stage of the amplifier in order to compare to eq. (326), which is derived for the dominant current noise in high-impedance circuits. Indeed, as given in the datasheet of the ultimately low noise amplifiers LT1028 and LT1128 from Linear Technology, the DC current is 1.8mA in the input stage, and it is a significant portion (~25%) of the quiescent current 7.5mA. As for MOS amplifiers, it is not possible to calculate gate overdrive and gate length from the information in datasheets. Therefore, the portion of I Q in the product I DC ×S VIN was varied until no data point left below the ultimate minimum of 2qφt² for the product, as follows from eq. (336) and assuming negligible noise contribution of the loading transistor. Using 0.02I Q ×S VIN , the results are shown in the bottom-right plot of Figure 58. In agreement with the prediction from eq. (336), the product with the white noise (triangles) is low at low I Q , and it increases slowly at higher I Q . There is bimodal distribution in the data, which corresponds to I DO (W/L)=2μA and 150nA, and for these values, the predictions from the term 2×2qφt²[0.5+(0.25+I Dsq /I DO ) 0.5 ] of eq. (336) is illustrated by solid lines through the triangles. Along with the increase of the bandwidth, this confirms that the current density (I Dsq /L²) increases with I Q , and according to the prediction from eq. (336) for the 1/f noise, the product 0.02I Q ×S VIN is nearly proportional to I Q for 1/f noise, as shown in the bottom-right plot of Figure 58 with circles and a trend line through them. The distribution in the data is wide, having standard deviation about 7dB, with limits shown by solid lines. The upper limit corresponds to 5V MOS with EOT=10nm, L=3µm, WL=450µm², low input capacitance of 1.6pF, and FOM SVG =500 μm²µV²/Hz. The low limit corresponds to 30V MOS with EOT=30nm, L=40µm WL=16000µm², high input capacitance of 18pF , and FOM SVG =100
129 of 286 μm²µV²/Hz. The capacitances are within the range given in the datasheets, the oxide thicknesses correspond to the maximum supply voltages in the datasheets, and the values for FOM SVG are within the range for analog MOS in ITRS [3]. Overall, the right-hand plots in Figure 58 demonstrate that the relations given in eq. (326) for BJT amplifiers and in eq. (336) for MOS amplifiers for the product I DC ×S VIN are valid, and the relations are observed in the commercially available integrated amplifiers as extension of the design window in Figure 57 toward low current densities, I C /A E for BJT or I DO /L² for MOS. Therefore, the IC manufactures use I Q ×S VIN as the figure of merit for their low-power low-noise amplifiers – see [295] for a datasheet from Linear Technology, for example. VII.2. Upand cross-conversion in phase noise VII.2.1. Definitions The noise in ideal linear circuits is additive, that is, the noise and the signal occur simultaneously and independently each from other. However, many circuits used for signal generation and processing are not linear, they are also time variant, and the low-frequency noise is “multiplied” in these circuits, even when the circuit operates at high frequency. This effect is known as up-conversion of low-frequency noise, in which the properties of the high frequency signal, e.g. amplitude, phase, delay, are affected by the low-frequency noise. The up-conversion of noise degrades the spectral purity of the signal, widening the spectrum of the signal, and in time domain it causes jitter in the transitions of the stationary cycling signals. There are many works devoted on up-conversion of low-frequency noise, and the up-conversion was reviewed in [296] with emphasis on BJT circuits. The problem of up-conversion can be introduced when looking at the modifications of ideal signal, when passing through a circuit. The modifications are assumed as modulations, and given generally by [ ] [ ] )t(tf2sin)t(V)t(V sas ϕ ∆+π∆+= , (338) where Δ a is fluctuation in amplitude and Δ φ is fluctuation in phase of the original signal V o =V s sin(2πf s t), the latter with amplitude V s and frequency f s . The frequency f s of RF signals is usually high as compared to the frequency range of the spectrum of the low-frequency fluctuations Δ a and Δ φ . Therefore, Δ a and Δ φ appear as amplitude and phase modulations of the original signal, and respectively, are regarded as amplitude and phase noise. (Other forms of signal modification can be also assumed, for example, as in perturbation phase noise theories, which we will discuss later.) The noise from devices contributes to both amplitude and phase noise in eq. (338). Both contributions result in broadening of the spectrum of the signal around frequency f s . For the amplitude noise, the product Δ a (t)sin(2πf s t) results in convolution Δ a (Δf)*δ(f s −Δf) in frequency domain between noise spectrum Δ a (Δf) at low frequency Δf<<f s and spectral line δ(f s −Δf) of the signal. The convolution creates side-lobes around signal frequency f s at offset frequencies Δf, which can be further modified by the shape of the amplitude-frequency response of the load, e.g. LC tank in RF circuits. In similar manner, the phase noise also creates side-lobes around signal frequency f s , since sin[2πf s t+Δ φ (t)]=sin(2πf s t)cos[Δ φ (t)]+cos(2πf s t)sin[Δ φ (t)] in time domain, and when converted in frequency domain, it also results in convolution δ(f s −Δf)*Δ φ (Δf)/Δf. (The division on Δf occurs, because one can take the spectrum of phase noise at Δf as constant at f s >>Δf, having spectrum inversely proportional to Δf after Fourier transformation. Precisely, the division is due to the fact that the frequency is time derivative of the
130 of 286 phase, as shown later.) The term Δ φ (Δf)/Δf results in “addition” of 1/Δf² slope to the frequency dependence of low-frequency noise, when addressing power spectrum densities (PSD) at high frequency. In other words, the PSD of phase noise decreases as 1/Δf² steeper, when compared to PSD of the noise source that causes the phase fluctuation. The above explanations are heuristic, but they capture the main properties of amplitude and phase noise. The typical spectrum of an oscillator is shown in Figure 59a for a CMOS oscillator [297]. When taking one side of the spectrum, either below or above oscillation frequency f s , one obtains the single side band (SSB) noise. The ratio of the power spectrum density of SSB noise to the magnitude of the carrier signal at the oscillation frequency defines the normalized spectrum of the phase noise, PhN, according to ( ) ( ) ( ) ],Hz/dBc[in,PhNlogdB10 [1/Hz]unit in , P fS fPhN 10 carrier SSB ×=⇔ ∆ =∆ (339) where P carrier is the power of the signal at the output of the oscillator and S SSB is the power spectrum density of SSB noise at Δf=f−f s frequency offset from oscillation frequency f s . One usually reports the data for PhN in units [dBc/Hz], which are according the second line of the equation. Figure 59b illustrates the PhN spectrum obtained from the spectrum of the oscillator in Figure 59a, showing also the two typical slopes, 1/Δf³ and 1/Δf², in PhN, which are result of up-conversion of 1/f and white noise, respectively, by “addition” of 1/Δf² to the slopes of the low-frequency noise, as mentioned above and discussed later. The amplitude noise is not a severe problem, since most practical circuits possess amplitude limiting mechanism [296], e.g. automatic gain control in oscillators and limiters for mixers, resulting in substantial suppression of the amplitude noise. However, the phase noise is serious problem, since a phase feedback is difficult, and the spectral broadening of the oscillation signal due to phase noise is perhaps the most critical issue for reference oscillators, in which the oscillator is free-running. Therefore, from here to the end of this sub-section, we focus on phase noise in oscillators. VII.2.2. Phase-noise in oscillators There are three approaches studying the phase noise in oscillators. These are time invariant, cyclostationary and perturbation approaches. Time invariant approach The time invariant approach was first suggested in [298], it is used widely, and a detailed review of the approach is given in [296]. It is assumed the obvious configuration of an electronic oscillator, in which an amplifier with noise figure NF, see eq. (228), has a frequency dependent LC−R s feedback so that at frequency (2πf s )²=1/(LC) the feedback is positive without phase shift, and the circuit generates sinusoidal signal at f s . An assumption is made in [298] that power spectrum density (PSD) of the phase fluctuation is equal to the ratio of amplifier noise to oscillation power. This assumption is based on eq. (338) in the following manner. [ ] { } [ ] )t(d)t(tf2cos)t(tf2sind V )t(dV ss sϕϕϕ ∆∆+π=∆+π= , (340) which rewritten for noise becomes [ ] 2 S S)t(tf2cos V S s 2 2 s Vϕ ϕϕ ≈∆+π= , (341)
131 of 286 since the time invariant approach takes the average value of cos²(x)=0.5. Then, by taking into account that V RMS =V s /√2 for sinusoidal oscillation, then the voltage noise of the amplifier is expressed with noise figure, as ∆ ∆ +==== f f 1NF P kT2 NF P kT2 NF RV kT4 V NFkTR4 V S c wh carriercarrier s 2 s 2 s s 2 s V , (342) where NF wh is effective noise figure of for white noise in respect to the resistance R s of the LC tank at resonance frequency f s , and Δf c is effective corner frequency between 1/f and white noise, as introduced in [298]. Thus, the PSD of the noise in the phase is ∆ ∆ +≈ ϕ f f 1NF P kT2 2 S c wh carrier , (343) where S φ /2 is SSB component corresponding to one side-lobe of noise. The fluctuation S φ in the phase is applied to the LC tank, which converts it in frequency noise S Δ f , as follows. The LC tank has (-3dB) bandwidth Q f 2 BW s =± , (344) where Q is the quality factor of the resonator used in the feedback of the oscillator. When the frequency offset is small, Δf<±BW/2, then the relation between phase and frequency in the LC tank is approximately ( ) s f Q2 f= ∆∂ ϕ ∂ ± , at f≈f s and f−f s =Δf<±BW/2. (345) Therefore, d(Δf)=dφ×f s /(2Q), which rewritten for noise S φ /2 in ±BW/2 is 2 S Q2 f S 2 s fϕ ∆ = , at f≈f s and f−f s =Δf<±BW/2. (346) For frequency deviation larger than ±BW/2, the phase of the LC tank does not change significantly with frequency. Therefore, the frequency and phase fluctuations are directly coupled each to other, without modification from LC tank. Owing to the general relation 2π(df)=∂(dφ)/∂t between phase φ and frequency f, which also holds for the output of the oscillator, then phase fluctuation can be equally rewritten as frequency fluctuation [298] ( ) ( ) ( ) 2 fS ffS 2 f ∆ ∆=∆ ϕ ∆ , at f~f s , but f−f s =Δf>±BW/2. (347) Rewritten for the oscillator output, the last relation is ( ) ( ) ( ) ( ) fPhN P S 2 fS f fS carrier SSB 2 f ∆== ∆ = ∆ ∆ θ ∆ , at f≈f s and any f−f s =Δf, (348) where S θ is noise in the phase and S SSB is the single-side band PSD of the oscillator output signal. Combining eqs. (346) and (347), as suggested in [298], and substituting in eq. (348), one gets
132 of 286 ( ) ( ) ( ) ( ) ∆ + ∆ = ∆ ∆ ==∆ ϕ ∆2 s 2 f carrier SSB fQ2 f 1 2 fS f fS P S fPhN , at f≈f s and any f−f s =Δf. (349) When substituting with eq. (343), the expression for time invariant approach for phase noise in oscillators becomes as ( ) ∆ + ∆ ∆ +==∆ 2 sc wh carriercarrier SSB fQ2 f 1 f f 1NF P kT2 P S fPhN , at f≈f s and any f−f s =Δf, (350) where the noise figure NF wh is for white noise in respect to the resistance R s of the LC tank at resonance frequency f s , as mentioned above. In RF oscillators generating in the frequency range of GHz, one usually uses LC tanks with Q<100, and f s /(2Q)>10MHz. Therefore, the last term in the square brackets dominates, and practically for Δf<1MHz ( ) ( ) ( ) ∆<∆∆ ∆>∆∆ ∝ ∆ ∆ ∆ +≈∆ conversion-up noise 1/f ,ff if ,f/1 conversion-up noise white,ff if ,f/1 fQ2 f f f 1NF P kT2 fPhN c 3 c 2 2 sc wh carrier (351) Consequently, one observes two slopes, 1/Δf³ and 1/Δf², in oscillator phase noise, as illustrated in Figure 59b. Eqs. (350) and (351) are known as Leeson formula for phase noise [298]. Certainly, the time invariant approach captures very essential features related with the phase noise. These are the abovementioned two frequency slopes, and the requirements for high quality factor Q of the resonator, low noise figure and low corner frequency between 1/f and white noise in the amplifier. However, there are issues, because Δf c is different, usually lower, than the corner frequency f c between 1/f and white noise in the amplifier, and also, there is no accurate expression that relates NF wh to the noise figure NF of the amplifier, the latter given for low-frequency noise by eq. (228), for example. Cyclostationary approaches The above issues were found to originate to circuit asymmetry and high-order harmonics in oscillators, after applying cyclostationary analyses. Earlier experiments, such as that in [299], showed that improving the circuit symmetry, and thus reducing harmonics related to non-linearity, decrease the SSB lobes and phase noise. The cyclostationary analyses are area of active research [300, 301, 302], they are lengthily to be presented here in full, and there is no review available at present, which compares different approaches, in computer simulators. The approach is introduced in [303], it is sketched in [304], the derivation of the equations is presented at [305], and summarized in [212], as ( ) ( ) ( ) 2 * n s n,m N 0m N 0n m s 2 f/1 2 0 s f2 1 i f S i f f S i f fPhN ∆ ∂ ∂ ∂ ∂ + ∆ ∂ ∂ ≈∆ = = , (352) where ∂f s /∂i m are sensitivities of the (oscillation) frequency f s to changes in harmonic with number m=0,1…N, [∂f s /∂i n ]* are complex conjugated values of the sensitivities of the (same) harmonics with number n=0,1…N, and the white noise power spectrum densities S m,n are from all noise sources weighted by magnitude of the harmonics (e.g. for shot noise in BJT, S m,n =2qI m − n , where I m − n is the magnitude of the (m−n) th harmonic in the BJT current). Note that the 1/f noise, S 1/f , and any other low-frequency noise, contributes only by the DC
133 of 286 component i 0 of the signal, as shown with the term in front of the sum, whereas the high frequency white noise around all harmonics also contributes to the phase noise around the fundamental harmonic with frequency f c . In other words, the up-conversion of low-frequency noise is only via “DC component”, the 0 th harmonic, of the signal derivatives in oscillators, while the white noise around all harmonics adds on the top of the up-converted low-frequency white noise, according to harmonic balance method for cyclostationary analyses of phase noise. Thus, in oscillators, in which the signals are large and harmonics are present, the phase noise is not solely due to up-conversion of low-frequency noise; and the corner frequency Δf c between 1/Δf² and 1/Δf³ phase noise components is lower in oscillators as compared to the corner frequency f c between low-frequency white and 1/f noise of the transistors biased at the same DC conditions as in the oscillators. This is also confirmed by perturbation methods for analysis of phase noise, which are summarized below. Owing to the importance of balance between DC, amplitude of oscillation and linearity for phase noise, we will discuss the tradeoff immediately after the paragraphs for the perturbation approach, as the third design consideration – see later and below eq. (398). Perturbation approach The perturbation approach for analysis of phase noise assumes that the noise perturbs the operation of the noiseless oscillator. As mentioned earlier, eq. (338) is not the only way to describe the modification of ideal signals. The perturbations can be assumed and introduced in various manners, e.g. in amplitude Δ a and phase Δ φ by eq. (338), or in other form, such as ( ) [ ] ( ) [ ] { } +∞ −∞= ∆+π+∆=∆++∆= k tskatoa ttkf2jexpV)t(ttv)t()t(V , (353) where v o (t)=ΣV k exp(j2πkf s t) is the un-perturbed (ideal) signal of the oscillator with fundamental frequency of oscillation f s , and also having harmonics with amplitudes V k , which are the Fourier coefficients of v o . The noise perturbation is usually assumed in Δ a as an amplitude perturbation, and the response of the oscillator converts it into jitter Δ t or, equivalently, in phase deviations 2πkΔ t . Thus, the phase deviation Δ φ =2πf s Δ t around fundamental harmonic (k=1) is the phase noise modulation in eq. (338), and eq. (353) also implies that there is phase noise modulation around all harmonics of the oscillator, which, indeed, are k-times larger than the phase noise modulation Δ φ of the fundamental harmonic, where k is the number of the harmonic. Among the several phase noise theories that use perturbation, we discuss two, since the mathematical derivations are lengthy, while the results from the different analyses appear to converge each to other. A list of publications that deal with phase noise theory is provided at the end of this sub-section. First approach. One approach to perturbation phase noise analysis is to look at the differential equations that describe the oscillator, by adding perturbation in the equations, and obtain insights for the behavior of the oscillator signals in time and frequency domains. Such approach was taken in [306], the comprehensive derivations are published in [307] and [308] for white noise and colored noise (1/f noise and band-limited Lorentzian noise), respectively, and the approach was summarized in [309], as follows. ( ) ( ) ( ) ( ) ( ) [ ] ( ) 2 2 ffw 2 s ffw 2 s 2 k SSB s f)f(Scckf )f(Scckf V S kf,fPhN ∆+∆+π ∆+ =≈∆ , (354) where k(=1, 2, 3…) is the number of the harmonic of interest with frequency (kf s ), c w is a parameter that reflects
134 of 286 the contribution of white noise sources in the oscillator, the sum is over noise sources in the oscillator, c f are parameters that weight the contribution from colored noise sources, the latter having power spectrum densities S f , and c f S f is normalized noise, that is, if S f is in unit A²/Hz, then c f is in unit 1/A². The expression in the square brackets is small, e.g. less than 10Hz, and it causes the phase noise spectrum to level-off at small frequency deviations. The parameter c w can be determined from the circuit analysis both in time and frequency domains [307], but the calculation involves finding solution of the equations based on the oscillator circuit equations, which are also differential and non-linear in principle. Nevertheless, at particular biasing and in particular circuit, c w has a single value, and it is shown in [307] that c w is also related to the variance of the jitter, or m 2w,t w t c ∆ σ = , (355) where t m is the duration of observation (measurement) of the jitter, and σ Δ t,w is the standard deviation of the jitter Δ t in eq. (353) during the observation and due to white noise. One can assume only the white noise causing the jitter, that is σ Δ t,w ≈σ Δ t , by neglecting other components in σ Δ t , resulting from 1/f or Lorentzian noise. Then, if the jitter measurement is at the n th period of the clock of a digital signal (after slope triggering of an oscilloscope by the clock signal, for example), then t m =nT s =n/f s and from histogram of the time of the signal transitions around nT s , one can calculate σ Δ t and estimate c w from eq. (355). Then, one can measure the phase noise spectrum and from the region with 1/(Δf)² slope in the spectrum, to obtain other estimate for c w =PhN×(Δf/fs)², according to eq. (354), and verify whether the contribution of white noise is the dominant in the jitter, which would be true, if the values for c w obtained by both measurements are close. One note should be made here, that Gaussian white noise was used in the deviations of eqs. (354) and (355), while in the practice, the jitter may have pattern dependent component, which might be not Gaussian. Thus, a simple pattern should be used in jitter measurement, but not pseudo random sequence. For the clock signals, one useful relation between cycle-to-cycle jitter and phase noise, which has been derived in a simple manner and verified experimentally in [310], is ( ) ( ) ( ) s 3 s 2 sm 2w,t f,fPhN f f f/1t ∆ ∆ ==σ ∆ , cycle-to-cycle jitter. (356) The expression for cycle-to-cycle jitter follows directly from eqs. (354) and (355) when setting observation (measurement) interval t m =1/f s reciprocal to of the oscillation frequency f s and consider the phase noise in the fundamental harmonic (k=1) at dominance of c w in the nominator and dominance of (Δf)² in the denominator of eq. (354). Apart from the complications to calculate the parameters c f from differential equations of the oscillator, another problem in eq. (354) is a singularity that occurs at very small offsets Δf→0, if the colored noise S f =K/f is 1/f noise. In such case, eq. (354) reduces to Δf/[(πkfs)²c f K]∝Δf→0, and the phase noise reduces, instead to increase or level off, when the frequency offset Δf decreases. Such reduction was never observed for phase noise in the practice of free running oscillators. (The reduction is evident in fractional PLL with sigma-delta modulator.) To remedy, a low-frequency corner f cf for the spectrum of 1/f noise was introduced in [308], below which frequency S f is constant. Thus, the 1/f noise at very low-frequency Δf→0Hz and DC was limited to S f (0Hz)=4/f cf in perturbation theory, by terming the flicker noise spectrum as
135 of 286 ( ) ( ) ( ) 0ff , f 4 ... f f2 2f2 4 f 1 f2 f artg f2 4 f 1 f2x dx 4fSS cf cfcf cf f2 2 BLf cf →>= + π + π π −≈ ππ −= π+ == ∞ , spectrum of 1/f noise limited to low frequency f cf . (357) The autocorrelation function R f =R 1/f and the variance σ f =σ 1/f (related to the jitter) of 1/f noise with lowfrequency corner f cf , respectively, are [308] ( ) ( ) ( ) tfEI2tRR cff/1f == , autocorrelation function of 1/f noise, (358) ( ) ( )( ) ( ) ( ) 2 cf cf 2 cfcfcfcf 2f/1 f tfEItftf1tfexp1tf2 2t +−−+− =σ , variance of 1/f noise, (359) with the exponential integral EI(x) being ( ) ∞ − = 1 dz z xzexp )x(EI . (360) For comparison, for RTS, burst and GR noise, which have low-frequency band-limited normalized Lorentzian spectrum of ( ) ( ) 2 cf BLf f f 21 1 fSS π+ == , Lorentzian spectrum (361) the autocorrelation function and the variance are [308] ( ) ( ) ( ) tfexp 2 f tRR cf cf BLf −== , autocorrelation function of noise with Lorentzian spectrum, (362) ( ) ( ) cf cfcf 2 BL f tfexp1tf t−+− =σ , variance of noise with Lorentzian spectrum, (363) and the white noise has jitter variance (σ Δ t,wh )²=c w t, as follows from eq. (355) given earlier. Note that variances are functions of the time, and the complexity of the functions increases when the type of the noise changes from white noise, through Lorentzian noise, to flicker (1/f) noise. In addition, the observation of phase noise or jitter is never for infinite time, but for the time window t m of the measurement, that is for the acquisition time by spectrum analyzers or oscilloscopes. The time window t m modifies the results for variance, and at high time and frequency resolution, so that both frequency and time being assumed continuous, it was shown in [308] that the variance σ Δ t for the jitter Δ t , see again eq. (353), during the observation time t m can be calculated by
136 of 286 ( ) ( ) ( ) τττ−+=σ ∆ m t 0 ffmwm 2t dRtc2tct , (364) using eqs. (358) and (362) for autocorrelation functions of 1/f noise and Lorentzian noise, respectively, by performing the calculation in time domain, or equivalently by calculation in frequency domain ( ) ( ) ( ) ( ) +∞ ∞ − ∆ π π− +=σ df f2 ft2jexp1 fSc2tct 2 m ffmwm 2t , (365) using eqs. (357) and (361) for power spectrum densities of 1/f noise and Lorentzian noise. The summation in the last two equations is along noise sources in the oscillator circuit. To meet the assumptions for Gaussian random variables, which is used in the derivations, t m has to be in large enough steps Δt m , so that the results for σ Δ t,i =σ Δ t (i×Δt m ) and σ Δ t,k =σ Δ t (k×Δt m ) by any values of i and k must satisfy the condition ( ) [ ] ( ) [ ] ki if ,0 kif2exp kif2exp 2k,t 2 2 s 2 2i,t 2 2 s 2 ≠≈ σ−π− σ−π− ∆ ∆ , when using eqs. (364) and (365). (366) Second approach. Noticeably, the use of the equations above meets with difficulties in the practice, they might be mapped in numerical methods, but then, they become vulnerable for quantization errors. Therefore, other approach for building of solvers for phase noise and jitter was used in [309]. The approach is based at macro model for the relation phase-frequency, in particular, the phase deviations are integrated frequency deviations. That is, the macro model is an ideal integrator in time domain. The circuit noise sources are weighted and summed into power spectrum density S ω of a macro noise source at the input of the integrator, and the output of the integrator is the phase deviation Δ φ or jitter Δ t =Δ φ /(2πkf s ) of the oscillator, where k is the harmonic number and f s is the frequency of oscillation. The weighting is with the parameters s w and c f are according to eq. (354), and the macro noise source at the input of the integrator is ( ) ( ) ∆+=∆ ω fSccfS ffw . (367) Note that S ω is a normalized power spectrum density in unit 1/Hz, as mentioned after eq. (354). In frequency domain, the transfer function of the integrator is H I =1/(j2πΔf). The observation (or measurement) is practically for finite time t m , and then, it is repeated, if desired. In other words, the integration captures the evolution of the increment of the phase and its variations (e.g. jitter or phase noise) only for time t m . Then, the integrator is “zeroed” before the next observation. This is exactly what happens by triggering the measurement in real instruments. Let us neglect the pause between single measurements. Every single measurement will give the increment of the phase and its variation between current and delayed with t m integrations. In such situation, the delay-difference operator in frequency domain is H tm =1−exp(−j2πΔf×t m ), and H tm is a transfer function that multiplies the transfer function H I of the integrator. In this way, the observation time window is introduced in [309], and the power spectrum density S φ of the phase at the output of the integrator for observation time t m becomes, as
143 of 286 thus, from eq. (375), ISF also peaks during transitions. Detailed investigations of phase noise in ring oscillators can be found in [315, 317, 318, 319, 320], from which one can have simple analytical expressions for phase noise as function of number of stages, supply, frequency, transistor sizes, etc., and comparisons to other type of oscillators, and regimes of operation, as well as relations for jitter, that match with the discussion on eqs. (371) and (372) earlier. In contrast to ring oscillators, in LC oscillators, the cyclostationary currents and voltages are not in phase, which results in large differences between ISF and effISF, especially in single transistor oscillators, such as the Colpitts oscillators, in which the cyclostationary transistor current is only for a fraction of the half period of oscillation, and max(|effISF|)=max(|ISF|×α)<max(|ISF|)×max(α). VII.2.3. Suggestions for the design of oscillators Several useful considerations in design of oscillators can be deduced from eq. (382). Minimize the up-conversion of low-frequency noise The first is that the DC component ISF dc =c o /2 has to be minimized in order to minimize the up-conversion of low-frequency noise, according to the middle line of eq. (382). This will also reduce the conversion of stationary RF white noise (e.g. thermal noise or shot noise due to quiescent current) by the term (c o )²∑S RF (kf s ) in the bottom line of the equation. As shown in [321], ISF dc is minimized when having particular symmetry in the waveforms of the oscillator, either even periodic waveforms f(t)=f(−t±nT), or half-wave symmetric waveforms f(t)=−f(t+T/2±nT). From practical point of view, both symmetries put requirements for identical transitions in oscillator waveform. This is depicted in Figure 61, in which the ideal waveforms are given with solid lines, and the distorted waveforms are drawn with black dash-lines, giving a rise of non-zero ISF dc , and thus, of upconversion of low-frequency noise into phase noise. The even symmetry requires identical transitions by mirroring in time, and the half-wave symmetry requires identical transitions by mirroring in amplitude, the latter also at delay of exactly half period. From spectral point of view, the even symmetry results in harmonic expansion of the oscillator waveform, ∑cos(2kπf s t+kφ o ) with φ o =constant, that is in-phase or equally delayed harmonics. The half-wave symmetry requires that the oscillator signal is free from even harmonics and the duty cycle of the signal is 50%. As seen from Figure 61, differences in rise and fall times of the circuit signals cause deviation from desired ideal waveforms. The circuits shown with gray color in the figure help to improve the signal symmetry. In particular, the complementary pMOS cross-coupled transistor pair works in anti-phase of the nMOS transistor pair. Therefore when nMOS switches on, pMOS switches off (and vice versa), which provides that one has always a transistor switching on and other transistor switching off during the time of signal transitions, and thus, smaller difference between rise and fall times and better even periodic symmetry, as illustrated by gray dash-lines in the bottom-left figure. Experiments that confirm the effect of reduction of phase noise when using complementary cross-coupled pairs can be found in [322]. The introduction of current limiting (“starving”) transistors in series with the main switching transistors in the ring oscillator minimizes the difference between the charging (I P , from pMOS) and discharging (I N , to nMOS) currents that flow through the node capacitance during transitions, and, thus, the rates dv/dt during signal rise and fall are equalized, resulting in signal with better half-wave symmetry, as illustrated by the gray dash-line in the bottom-right figure. Obviously, limiting the current, the circuit is slower, and the oscillation frequency in the ring oscillator with current “starving” is expected to decrease, as compared to the initial circuit without current “starving”. Experiments and analysis that confirm the effect of reduction of phase noise in ring oscillators when using current limiting (“starving”) transistors to adjust the currents of nMOS and pMOS transistors can be found in
[Document text truncated for crawler view.]