scieee AI-readable full text Open interactive document viewer

Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences – Is this Relevant for Speech in Noise?

Hauth, Christopher F.; Brand, Thomas

Full text

Submitted to Acta Acustica Template provided by EDP Sciences Research article 1 Adapting the EC Model to the Frequency Dependency of Binaural2 Masking Level Differences – Is this Relevant for Speech in Noise?3 Christopher F. Hauth1⋆and Thomas Brand1 4 1Carl von Ossietzky Universit¨at Oldenburg, Department f¨ur Medizinische Physik und Akustik and Exzellenzcluster5 “Hearing4All”, 26129 Oldenburg, Germany6 Abstract – Several binaural speech intelligibility models simulate human binaural processing using the7 Equalization-Cancellation (EC) mechanism, incorporating processing inaccuracies based on tone-in-noise8 detection at 500 Hz. Despite this, such inaccuracies are typically assumed to be frequency-independent in9 binaural speech intelligibility models. This study examines the validity of that assumption and explores10 how auditory filter bandwidth affects binaural masking release. Experiment I measured tone detection11 thresholds across frequencies from 250 to 2000 Hz in 12 normal-hearing and 5 hearing-impaired listeners,12 varying the interaural phase difference (IPD) of noise between 0 and 5π. Results were compared with13 EC model predictions. Binaural inaccuracies showed a low-pass effect, reducing binaural masking level14 differences (BMLDs) at higher frequencies. However, the EC model inaccurately predicted BMLDs at15 250 Hz. Experiment II involved a subset of listeners performing speech-in-noise intelligibility tasks with16 low-pass filtered speech. This allowed assessment of whether tone-in-noise detection could predict speech17 recognition thresholds (SRTs), particularly in hearing-impaired individuals whose low-frequency hearing18 remained near normal. Findings revealed that adapting the model’s binaural inaccuracies improved SRT19 predictions beyond audiometric thresholds alone.20 Keywords. Binaural, Tone-in-Noise, Binaural Processing Inaccuracies, Speech in Noise, Hearing Impairment21 1. Introduction22 Several binaural speech intelligibility models have successfully been used to predict the outcome of speech intel-23 ligibility experiments in different acoustic scenarios ranging from anechoic conditions to reverberant conditions [e.g.24 1–8]. These models can be separated into a front-end, which mimics human auditory processing, and a back-end,25 which quantifies the usable speech information in the acoustical signal. The front-end usually consists of a band-26 ⋆Corresponding author: [email protected] 1 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences pass filterbank [e.g. a gammatone filterbank; 9] to mimic the frequency selectivity of the auditory system and an27 equalization-cancellation (EC) mechanism [10] to model the effective binaural auditory processing of human listen-28 ers. The back-end usually consists of an intrusive model like the Speech Intelligibility Index [SII; 11], the Speech29 Transmission Index [STI; 12], the short-time Objective Intelligibility [STOI; 13] measure or a non-intrusive model [for30 an overview see e.g. 14].31 The EC model [10] uses interaural differences in the target signal and the interfering signal to improve the signal-to-32 noise ratio (SNR) especially at low frequencies. In the EC model, first, the left ear and right ear channels are equalized33 in level, such that the interaural level difference (ILD) is compensated. This processing approximates the compressive34 behavior of the cochlear. In a second step, the interaural time difference (ITD) or the interaural phase difference (IPD)35 between the left and right ear is equalized such that the SNR at the output of the EC stage, which is calculated by36 subtracting left and right ear channel from each other, is improved. The EC processing can be used to describe the37 outcome of narrow-band experiments like tone-in-noise detection tasks, but also for broad-band signals like speech,38 where independent EC processing across frequency bands is assumed. In this way, the EC stage can account for the39 binaural release from masking, which is observed when a target source and an interfering source differ in their ITDs or40 IPDs, or – more ecologically relevant – if they differ in their spatial location. The EC model, however, overestimates the41 binaural release from masking especially in anechoic conditions [2,10,15,16]. To compensate for this overestimation,42 binaural processing inaccuracies have been incorporated into the EC model in order to limit the resolution of the43 equalization process in level and time. Consequently, no perfect cancellation can be achieved if the binaural processing44 inaccuracies are used and the SNR improvement is limited. The binaural processing inaccuracies have first been derived45 by Durlach [10], where the ITD inaccuracy was assumed to be a normally-distributed random variable with zero mean46 and a standard deviation of 105 µs. This has been modified by vom H¨ovel [16], who fitted the binaural processing47 inaccuracies to binaural tone-in-noise detection thresholds obtained in an experiment by Langford and Jeffress [17]48 for jittering the ITD and to binaural tone-in-noise detection thresholds obtained in an experiment by Egan [18] for49 jittering the ILD. Both measurements were obtained for a single frequency of 500 Hz with normal-hearing listeners.50 In the experiment by Langford and Jeffress [17] tone detection thresholds were obtained in broadband noise, where51 the ITD of the noise was varied. In the experiment by Egan [18], the tone detection thresholds were obtained, while52 the ILD of the noise was varied. The processing inaccuracies derived by vom H¨ovel [16] are assumed to be normally53 distributed random variables with zero mean, and – in contrast to Durlach [10] - a standard deviation which is a54 function of ITD and ILD, respectively. These standard deviations of the processing inaccuracies are given by55 σδ=σδ0·[1 + |∆| ∆0 ],(1) and56 σϵ=σϵ0·[1 + ( α α0 )p],(2) 2 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences Figure 1. Standard deviations of the ITD error σδand ILD error σϵin the EC mechanism derived by vom H¨ovel [16]. The standard deviation of the ITD error is assumed to grow linearly with increasing ITD, while the ILD error in dB grows exponentially with increasing ILD (in dB). with σδ0= 65µs, ∆0= 1.6ms, σϵ0= 1.5dB , α0= 15dB, and p=1.6 [16].57 These values have for example been used in the binaural speech intelligibility models by Beutelmann et al. [2] and58 Andersen et al. [5], Hauth and Brand [7], and Hauth et al. [8]. Figure 1shows the standard deviations of the ITD59 and ILD inaccuracies as derived by vom H¨ovel [16]. The standard deviation of the ITD processing inaccuracy is60 defined in the time domain and assumed to grow linearly with ITD. The standard deviation for an ITD of 0 µs (and61 thus the intercept) was set to 65 µs. The standard deviation of the ILD processing inaccuracy is assumed to grow62 exponentially with increasing absolute value of the ILD. In the binaural speech intelligibility models mentioned above,63 these processing inaccuracies are applied in each frequency band, and therefore, are assumed to be independent of64 center frequency. However, it is unclear if this assumption is valid, because, to our knowledge, a frequency dependency65 of the binaural processing inaccuracies has not yet been investigated and compared to predictions of binaural speech66 intelligibility models using the EC model.67 Binaural unmasking, which relies on ITDs or IPDs, gets worse with increasing frequency[e.g. 19]. The same holds68 for the ability of the auditory system to preserve the temporal fine structure of a stimulus, which gets also worse69 with increasing frequency and is often called “phase locking”. The loss of phase locking has been regarded as the main70 reason for the reduction of BMLDs with increasing frequency [20,21]. The time domain implementation of the binaural71 processing inaccuracy in the EC mechanism generally accounts for this finding for the N0Sπor NπS0condition (where72 either the phase of the speech or the noise is interaurally inverted, while the phase of the respective other is not). The73 constant ITD inaccuracy leads to an increasing IPD inaccuracy with increasing frequency resulting in a decreasing74 BMLD with increasing frequency . However, this implementation implies that the predicted BMLD at 250 Hz is even75 higher than at 500 Hz, which has not been observed in binaural tone detection experiments so far. For example,76 van de Par and Kohlrausch [19] investigated the effect of masker bandwidth and center frequency on binaural release77 from masking. Their data indicate that the binaural release from masking gets smaller for frequencies below 500 Hz78 in the NπS0condition (see their Figure 1). Taken together, it is still not clear, whether predicting BMLDs in tone-79 in-noise detection experiments at different center frequencies might require frequency dependent binaural processing80 inaccuracies in the EC model.81 Moreover, this study analyzed if the binaural processing inaccuracy can be described as a function of ITD and how82 predictions with the EC model incorporating these binaural processing inaccuracies change for large ITDs. Physiological83 3 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences measurements in mammals suggest that the binaural neurons respond best to an ITD which is within a half-cycle (π)84 of the corresponding frequency band [22]. That’s why nearly all binaural models restrict the binaural processing stages85 to use temporal binaural cues within the so called “π-limit”. Both questions were addressed in Experiment I, where86 the frequency dependency of BMLD curves was investigated in 12 listeners with normal hearing and 5 listeners with87 high-frequency hearing loss. The results obtained in the listening experiments were compared to BMLD predictions of88 the EC model.89 In Experiment II, binaural speech intelligibility experiments were performed by a subset of the same listeners in order90 to test the hypothesis, if tone-detection performance is reflected in binaural speech reception thresholds (SRTs), i.e., if91 better tone detection is related to better SRTs. We expected, that this experiment might reveal differences in binaural92 unmasking between listeners with normal and with impaired hearing, which could potentially be used to improve SRT93 predictions for listeners with impaired hearing. In this study, the binaural speech intelligibility model by Beutelmann94 et al. [2] was used, which considers the individual pure-tone audiogram using a threshold simulating noise. However,95 it has been shown that the effect of hearing impairment was not fully accounted for by incorporating the audiogram96 alone as it does not reflect, for instance, supra-threshold deficits [23,24] and potential individual binaural processing97 capabilities. In a study conducted by Neher et al. [25] it was shown that listeners with hearing impairment differed98 in their binaural intelligibility difference (BILD), which is the binaural release from masking by spatially separating99 the target speech source from the interfering noise source in a speech intelligibility experiment, and BMLDs at 500 Hz100 even though they were matched in age, hearing loss (according to the pure-tone audiogram), and cognitive factors.101 This finding suggested that supra-threshold deficits are reflected in binaural tone detection thresholds and should,102 therefore, also be accounted for in binaural speech intelligibility models.103 The SRTs of Experiment II were predicted using BSIM with individual adapted binaural processing inaccuracies,104 which were derived based on the binaural tone detection thresholds from Experiment I. It was hypothesized that these105 individually adjusted binaural processing inaccuracies lead to better predictions of the individual SRTs. Both listeners106 with normal hearing and with hearing impairment performed SRT measurements for low-pass filtered speech with a107 cut-off frequency of 1500 Hz, in order to exclude effects of high frequency hearing loss and to focus on the frequency108 region, where binaural unmasking can be assumed to be largest.109 2. Methods110 Listeners111 12 listeners with normal hearing and 5 listeners with impaired hearing participated in Experiment I. 8 of the112 12 normally hearing listeners and the hearing-impaired listeners of Experiment I also participated in Experiment II.113 Normal hearing was guaranteed by standard pure-tone audiometry at the frequencies 125, 250, 500, 1000, 2000, 3000,114 4000, 6000, and 8000 Hz, where none of the thresholds exceeded 20 dB HL. Hearing impaired listeners with nearly115 normal hearing at frequencies up to 1000 Hz were chosen in order to reduce low-frequency audibility effects on the116 4 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences Figure 2. Audiograms of the listeners with hearing impairment. The left and right panels show the audiogram of the left and right ear, respectively. BMLD measurements. Figure 2shows the individual audiograms of the listeners with impaired hearing. Up to 1000 Hz,117 these listeners had audiometric thresholds, which were better than 30 dB HL.118 Experiment I: Binaural Tone-in-Noise Detection119 The first experiment was a binaural tone-in-noise detection task, similar to the BMLD measurements by Langford120 and Jeffress [17]. Binaural tone-in-noise detection thresholds were determined for carrier frequencies of 250, 500, 750,121 1000, 1500, and 2000 Hz for listeners with normal hearing and for 500, 750, 1000, and 1500 Hz for the listeners with122 impaired hearing. Tone detection was investigated in Gaussian white noise, which was bandpass filtered between 100-123 4000 Hz. The broad-band noise was preferred over narrow band noise because it allows for an easier discrimination124 between target tone and interfering noise. Moreover, the experiments by Langford and Jeffress [17] and Egan [18] were125 also performed using broadband noise. A 3-alternative-forced-choice (AFC) 1-up-2-down procedure converging to the126 70.7% correct point on the psychometric function was used to determine tone detection thresholds.127 Apparatus128 The stimuli were generated using MATLAB (MathWorks, Natick, MA, USA) using the AFC Toolbox (Version 1.4),129 developed by Stephan Ewert at Carl von Ossietzky Universit¨at, Oldenburg, Germany, and presented binaurally via an130 RME Fireface UC soundcard (Audio AG, Haimhausen, Germany) and HD 650 headphones (Sennheiser, Wedemark,131 Germany). The sound output was calibrated to dB SPL using a Br¨uel&Kjaer (B&K, Nærum, Denmark) 4153 artificial132 ear, a B&K 4134 half-inch microphone, a B&K 2669 preamplifier, and a B&K 2610 measuring amplifier. The noise133 level was set to 75 dB SPL and the level of the tone was varied to find the individual threshold. The experiments were134 conducted in a double-walled, sound-attenuated booth.135 5 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences Stimuli136 The target tone was always presented diotically (S0), i.e., the same signal wais presented to both ears. According137 to Langford and Jeffress [17], the tested ITDs of the noise were selected to mirror fixed IPDs for each tested frequency138 ranging from 0 to 5πin steps of π/2 leading to frequency dependent ITDs. For example, ITDs ranging from 0 to 5 ms139 in steps of 0.5 ms were used at 500 Hz, while the ITDs ranging from 0 ms to 10 ms in steps of 1 ms were used at140 250 Hz. This method was chosen, because binaural release from masking is largest using IPD of πin either the tone141 or the noise.142 Experiment II: Binaural Speech-in-Noise Experiments143 In Experiment II, a binaural speech-in-noise test was conducted with the same listeners as in Experiment I. Speech144 intelligibility experiments were conducted using the Oldenburg Sentence Test [OlSa, 26–28] in speech-shaped stationary145 noise with the same long term average spectrum of the speech. In the remainder of this article, this noise is referred146 to as “Olnoise”. OlSa sentences consist of five-word-sentences with a fixed grammatical structure noun-verb-numeral-147 adjective-object, where each word is randomly selected from a list of 10 words. To determine the SRT for 50% correctly148 understood words, an adaptive procedure was used for controlling the level of the speech [Equation 9, 29] and the149 SRT was estimated using a maximum likelihood fit. Each SRT was determined using a list of 20 sentences, which was150 randomly selected out of 45 lists. Because the main focus of this study lied on the binaural processing capabilities of151 the listeners, speech and noise stimuli were modified as follows: Both speech and noise signals were low-pass filtered152 with a cut off frequency of 1500 Hz in order to minimize the influence of the high frequency hearing loss for the listeners153 with hearing impairment while keeping their ability for binaurally processing the signals, which is most effective at low154 frequencies. To be able to compare the results obtained for normal-hearing listeners and hearing-impaired listeners,155 the listeners with normal hearing were provided with the same stimuli as the hearing-impaired listeners. In total, 4156 SRTs were obtained for each listener: 1) monaurals SRT for the left ear, 2) monaural SRTs for the right ear, 3) diotic157 SRTs, i.e., the same signals are presented to the left and right ear (N0S0), and 4) dichotic SRTs, where the phase158 of the noise was inverted between both ears (NπS0). This was done in order to be able to compare the outcome of159 the speech intelligibility tests with the binaural tone-in-noise detection experiment in a frequency region, where the160 listeners with hearing impairment showed nearly normal hearing thresholds according to the audiogram.161 Model predictions162 The binaural speech intelligibility model [BSIM, 2] consist of two stages: the front-end, that predicts the improve-163 ment of the SNR due to binaural unmasking and better ear listening, and the back-end that predicts the speech164 intelligibility. In the front-end, the EC model is applied in order to predict the improvement of the SNR due to binau-165 ral processing. To achieve this, the input signal is filtered into 30 ERB-spaced [30] frequency bands using a gammatone166 filter bank [9]. The listener’s individual hearing threshold is considered by adding a Threshold Simulating Noise (TSN)167 to each ear. This TSN is calculated by adding the frequency specific dB HL values to the threshold of normal hearing168 6 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences listener in dB SPL as defined in EN-ISO389-7 [31], where the threshold is determined binaurally in the free-field and169 is referred to as Minimum Audible Field (MAF). The TSN is uncorrelated between the ears, so that the EC model170 cannot cancel it. Afterwards, EC processing is applied in each frequency channel to improve the SNR. Finally, the171 resulting SNR is compared to the monaural SNRs at the left and right ear channel and the maximum of all three172 alternatives is considered as the result of the first stage.173 The model’s back-end uses the output of the front-end and estimates the speech intelligibility using the Speech174 Intelligibility Index [SII, 11]. The SII is calculated essentially as a weighted sum of the band-specific SNRs. Before the175 summation, the band-specific SNRs are limited to a range from -15 dB and 15 dB. The weighting is given by a band176 importance function, which mimics the importance of the individual frequency bands for human speech recognition,177 resulting in an SII value between 0 and 1. In this study, the NNS (various nonsense syllable tests where most of the178 English phonems occur equally often) weighting function [11] was used. The resulting SII values can then be mapped179 to intelligibility values. In this study, only SRT values are analyzed. Therefore, only the reference SII at the SRT is180 required, which we set to an SII value of 0.2. This values matches the SRT of -7.1 dB of listeners with normal-hearing181 for OlSa sentences in broadband speech shaped noise [28]. The front-end of the model without using the back-end was182 applied to the signals of Experiment I in order to evaluate how closely the tone-in-noise threshold of the model can be183 predicted by BSIMs front-end. This was done in two ways: Firstly, using the frequency independent processing errors184 derived for 500 Hz by vom H¨ovel [16] in the EC processing, in order to test the hypothesis of frequency independent185 binaural processing inaccuracies. And secondly, using optimized frequency dependent binaural processing accuracies,186 in order to test the hypothesis, that this can improve the prediction accuracy.187 The complete BSIM consisting of front-end and back-end is then applied to the results of Experiment II. Again,188 this was done using the original frequency independent binaural processing inaccuracies and the frequency dependent189 processing accuracies derived based on the results of Experiment I. The underling hypothesis of this approach is that190 the same enhancement mechanism is used by the human auditory system for narrow-band tone-in-noise tasks and for191 broad-band speech-in-noise tasks.192 3. Results193 Experiment I: Binaural Tone-in-Noise Measurements194 Figure 3shows the BMLDs obtained in the binaural tone-in-noise detection task for the 12 listeners with normal195 hearing and the BMLDs predicted by the EC model. The BMLD denotes the relative improvement of the tone detection196 threshold in the dichotic condition relative to the diotic condition. The BMLD is high if the ITD of the noise results197 in a phase shift corresponding to odd (1, 3, 5) multiples of πat the tested frequency, and low if the ITD corresponds198 to phase shifts of even (2, 4) multiples of π. This results in a periodic pattern of the BMLD, which can be observed in199 Figure 3and is referred to as “Jeffress curves”. The largest BMLD of approximately 11 dB can be observed at 500 Hz200 for an ITD of the noise of 1 ms, corresponding to a phase shift of π, which is in line with results from the literature [e.g.201 17]. With increasing frequency, the maximum achievable BMLD is reduced by approximately 4 dB at 1000 Hz and by202 7 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences Figure 3. Binaural Masking Level Differences (BMLD) measured in 12 listeners with normal hearing (gray diamonds). The tested frequencies of the tone were 250, 500, 750, 1000, 1500, and 2000 Hz. The IPD of the noise was varied from 0 to 5πin steps of 0.5π, leading to a frequency dependent ITD. For example, an IPD of πcorresponds to an ITD of the noise of 1ms at 500 Hz and to 0.5 ms at 1000 Hz. Additionally, the IPD of 3π 4corresponding to an ITD of 0.75 ms at 500 Hz was measured for all frequencies, which corresponds approximately to the maximum anatomically achievable ITD considering human head size. Two model versions were used: one with 1 ERB wide auditory filters and one with 2.3 ERB wide auditory filters. 6 dB at 2000 Hz. This finding is in line with literature [e.g. 19], where a correlation between the binaural release from203 masking and the capability of phase locking on the auditory nerve was observed. Interestingly, the BMLD observed204 for an IPD of 3π 4, which corresponds to an ITD of 750 µs at 500 Hz, is always as good as for an IPD of πfor all tested205 frequencies. For the tested frequency of 250 Hz (top-left panel of Figure 3), the BMLD is reduced compared to the206 500 Hz condition: For an IPD of the noise of π, the BMLD at 250 Hz is 8 dB, which is 3 dB lower than the BMLD207 obtained at 500 Hz. The reason for this is two-fold: Firstly, the diotic threshold is lower at 250 Hz than at 500 Hz,208 because the auditory filter can be assumed to be narrower at 250 Hz than at 500 Hz. According to the definition of209 the equivalent rectangular bandwidth (ERB = 24.7·(4.37 ·f[kHz] + 1) [30]), the reduction of masking by using the210 filter at 250 Hz and the same broad band noise can be up to 1.8 dB. This can also be found in the data, where the211 median diotic threshold at 250 Hz is reduced by 1.25 dB compared to the diotic threshold at 500 Hz, which is slightly212 lower than the theoretical value according to the ERB formula described above (see Figure 4).213 Secondly, the dichotic thresholds are statistically significantly increased by approximately 2 dB in the 250 Hz condition214 compared to the 500 Hz condition (t-test, p=0.021). In addition to the effect of center frequency on BMLD, also an215 effect of the tested IPD can be observed: With increasing IPD of the noise, the maximum achievable BMLD decreases216 from πto 3πand from 3πto 5πalmost independently of center frequency. Figure 5shows the differences between217 8 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences Figure 4. Diotic (N0S0) and dichotic (NπS0) tone detection thresholds obtained at 250 Hz (open circles) and 500 Hz (open diamonds). Figure 5. Difference between BMLDs (∆BMLD) for an IPD of πand an IPD of 3π(left panel) and for an IPD of 3π and an IPD of 5π(right panel). A negative ∆BMLD indicates that the tone detection threshold obtained at an IPD of πis lower (better) than at the tone detection threshold obtained at an IPD of 3π. BMLDs (∆BMLD) between πand 3π(left panel) and between 3πand 5π(right panel). Note, that the ∆BMLDs are218 identical to the differences between the dichotic tone detection thresholds, as the diotic thresholds are compensated219 due to the difference calculation. The largest difference in dichotic tone detection thresholds can be observed at 250 Hz,220 where the median difference is -4.8 dB. At 500 Hz, a difference of -4 dB is observed, which is reduced to -2.5 dB at221 750 Hz, to -2 dB at 1500 Hz, and to 0.8 dB at 2000 Hz. A Bonferroni corrected pair-wise comparison revealed a222 statistical difference in the BMLD differences between 250 Hz and 750 Hz as well as between 250 Hz and 2000 Hz.223 Moreover, the difference observed at 500 Hz was statistically significant different from the difference observed at224 2000 Hz. No statistical difference was observed in the tone detection threshold differences between an IPD of 3πand225 5πat the tested frequencies. The differences are in the range from -2.8 to -3 dB independent of frequency, except for226 250 Hz and 150 Hz, where the difference is below -2 dB.227 9 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences 2. Rainer Beutelmann, Thomas Brand, and Birger Kollmeier. Revision, extension, and evaluation of a binaural speech intel-391 ligibility model. The Journal of the Acoustical Society of America, 127(4):2479–2497, April 2010. ISSN 0001-4966. . URL392 http://asa.scitation.org/doi/10.1121/1.3295575.393 3. Rui Wan, Nathaniel I. Durlach, and H. Steven Colburn. Application of an extended equalization-cancellation model to394 speech intelligibility with spatially distributed maskers. The Journal of the Acoustical Society of America, 128(6):3678–395 3690, December 2010. ISSN 0001-4966. . URL http://asa.scitation.org/doi/10.1121/1.3502458.396 4. Sam Jelfs, Mathieu Lavandier, and John F. Culling. Revision and validation of a binaural model for speech intelligibility397 in noise. Hearing Research, page 9, 2011.398 5. Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, and Jesper Jensen. A method for predicting the in-399 telligibility of noisy and non-linearly enhanced binaural speech. In 2016 IEEE International Conference on Acoustics,400 Speech and Signal Processing (ICASSP), pages 4995–4999, Shanghai, March 2016. IEEE. ISBN 978-1-4799-9988-0. . URL401 http://ieeexplore.ieee.org/document/7472628/.402 6. Alexandre Chabot-Leclerc, Ewen N. MacDonald, and Torsten Dau. Predicting binaural speech intelligibility using the403 signal-to-noise ratio in the envelope power spectrum domain. The Journal of the Acoustical Society of America, 140(1):404 192–205, July 2016. ISSN 0001-4966. . URL http://asa.scitation.org/doi/10.1121/1.4954254.405 7. Christopher F. Hauth and Thomas Brand. Modeling Sluggishness in Binaural Unmasking of Speech for Maskers With406 Time-Varying Interaural Phase Differences. Trends in Hearing, 22:2331216517753547, 2018. . URL https://doi.org/10.407 1177/2331216517753547.408 8. Christopher F. Hauth, Simon C. Berning, Birger Kollmeier, and Thomas Brand. Modeling Binaural Unmasking of Speech409 Using a Blind Binaural Processing Stage. Trends in Hearing, 24:233121652097563, January 2020. ISSN 2331-2165, 2331-410 2165. . URL http://journals.sagepub.com/doi/10.1177/2331216520975630.411 9. Volker Hohmann. Frequency analysis and synthesis using a gammatone filterbank. Acta Acustica united with Acustica, 88412 (3):433–442, 2002.413 10. N. I. Durlach. Equalization and Cancellation Theory of Binaural Masking-Level Differences. The Journal of the Acoustical414 Society of America, 35(8):1206–1218, August 1963. ISSN 0001-4966. . URL http://asa.scitation.org/doi/10.1121/1.1918675.415 11. ANSI S3.5-1997. Methods for Calculation of the Speech Intelligibility Index, 1997.416 12. H. J. M. Steeneken and T. Houtgast. A physical method for measuring speech-transmission quality. The Journal of the417 Acoustical Society of America, 67(1):318–326, January 1980. ISSN 0001-4966. . URL http://asa.scitation.org/doi/10.1121/418 1.384464.419 13. Cees H. Taal, Richard C. Hendriks, Richard Heusdens, and Jesper Jensen. An Algorithm for Intelligibility Prediction of420 Time–Frequency Weighted Noisy Speech. IEEE Transactions on Audio, Speech, and Language Processing, 19(7):2125–2136,421 September 2011. ISSN 1558-7916, 1558-7924. . URL http://ieeexplore.ieee.org/document/5713237/.422 14. Yong Feng and Fei Chen. Nonintrusive objective measurement of speech intelligibility: A review of methodology. Biomedical423 Signal Processing and Control, 71:103204, January 2022. ISSN 17468094. . URL https://linkinghub.elsevier.com/retrieve/424 pii/S1746809421008016.425 15. Rainer Beutelmann and Thomas Brand. Prediction of speech intelligibility in spatial noise and reverberation for normal-426 hearing and hearing-impaired listeners. The Journal of the Acoustical Society of America, 120(1):331–342, July 2006. ISSN427 0001-4966. . URL http://asa.scitation.org/doi/10.1121/1.2202888.428 16. Harald vom H¨ovel. Zur Bedeutung der ¨ Ubertragungseigenschaften des Aussenohrs sowie des Binauralen H¨orsystems bei429 Gest¨orter Sprach¨ubertragung (On the importance of the transmission properties of the outer ear and the binaural auditory430 16 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences system in disturbed speech transmission). Ph.D. Dissertation, RWTH Aachen, 1984.431 17. Ted L Langford and Lloyd A Jeffress. Effect of Noise Crosscorrelation on Binaural Signal Detection. The Journal of the432 Acoustical Society of America, 36(8):1455–1458, August 1964. .433 18. James P. Egan. Masking-Level Differences as a Function of Interaural Disparities in Intensity of Signal and of Noise. The434 Journal of the Acoustical Society of America, 36(10):1992–1992, October 1964. ISSN 0001-4966. . URL http://asa.scitation.435 org/doi/10.1121/1.1939216.436 19. Steven van de Par and Armin Kohlrausch. Dependence of binaural masking level differences on center frequency, masker437 bandwidth, and interaural parameters. The Journal of the Acoustical Society of America, 106(4):1940–1947, October 1999.438 ISSN 0001-4966. . URL http://asa.scitation.org/doi/10.1121/1.427942.439 20. Leslie R. Bernstein and Constantine Trahiotis. The normalized correlation: Accounting for binaural detection across center440 frequency. The Journal of the Acoustical Society of America, 100(6):3774–3784, December 1996. ISSN 0001-4966. . URL441 http://asa.scitation.org/doi/10.1121/1.417237.442 21. P. M. Zurek and N. I. Durlach. Masker-bandwidth dependence in homophasic and antiphasic tone detection. The Journal443 of the Acoustical Society of America, 81(2):459–464, February 1987. ISSN 0001-4966. . URL http://asa.scitation.org/doi/444 10.1121/1.394911.445 22. David Mcalpine, Sarah Thompson, Katharina Von Kriegstein, Torsten Marquardt, Timothy Griffiths, and Adenike Deane-446 Pratt. A π-Limit for Coding ITDs: Neural Responses and the Binaural Display. In Birger Kollmeier, Georg Klump, Volker447 Hohmann, Ulrike Langemann, Manfred Mauermann, Stefan Uppenkamp, and Jesko Verhey, editors, Hearing – From Sensory448 Processing to Perception, pages 399–406, Berlin, Heidelberg, 2007. Springer Berlin Heidelberg. ISBN 978-3-540-73009-5.449 23. Laurel H. Carney. Supra-Threshold Hearing and Fluctuation Profiles: Implications for Sensorineural and Hidden Hearing450 Loss. Journal of the Association for Research in Otolaryngology, 19(4):331–352, August 2018. ISSN 1525-3961, 1438-7573.451 . URL http://link.springer.com/10.1007/s10162-018-0669-5.452 24. David H¨ulsmeier and Birger Kollmeier. How much individualization is required to predict the individual effect of suprathresh-453 old processing deficits? assessing plomp’s distortion component with psychoacoustic detection thresholds and fade. Hearing454 Research, 426:108609, 2022. ISSN 0378-5955. . URL https://www.sciencedirect.com/science/article/pii/S0378595522001770.455 25. Tobias Neher, Kirsten C. Wagener, and Matthias Latzel. Speech reception with different bilateral directional process-456 ing schemes: Influence of binaural hearing, audiometric asymmetry, and acoustic scenario. Hearing Research, 353:36–48,457 September 2017. ISSN 03785955. . URL http://linkinghub.elsevier.com/retrieve/pii/S0378595517301193.458 26. Kirsten Carola Wagener, Volker K¨uhnel, and Birger Kollmeier. Entwicklung und Evaluation eines Satztests f¨ur die deutsche459 Sprache I: Design des Oldenburger Satztests [development and evaluation of a sentence test for the german language i: Design460 of the oldenburg sentence test]. Zeitschrift f¨ur Audiologie, 38(1):4–15, 1999.461 27. Kirsten Carola Wagener, Thomas Brand, and Birger Kollmeier. Entwicklung und Evaluation eines Satztests f¨ur die deutsche462 Sprache Teil II: Optimierung des Oldenburger Satztests [development and evaluation of a sentence test for the german463 language ii: Optimization of the oldenburg sentence test]. Zeitschrift f¨ur Audiologie, 38(2):44–56, 1999.464 28. Kirsten Carola Wagener, Thomas Brand, and Birger Kollmeier. Entwicklung und Evaluation eines Satztests f¨ur die deutsche465 Sprache Teil III : Evaluation des Oldenburger Satztests [development and evaluation of a sentence test for the german466 language iii: Evaluation of the oldenburg sentence test]. Zeitschrift f¨ur Audiologie, 38(3):86–95, 1999.467 29. Thomas Brand and Birger Kollmeier. Efficient adaptive procedures for threshold and concurrent slope estimates for psy-468 chophysics and speech intelligibility tests. The Journal of the Acoustical Society of America, 111(6):2801–2810, June 2002.469 ISSN 0001-4966. . URL http://asa.scitation.org/doi/10.1121/1.1479152.470 17 C. Hauth et al.: Adapting the EC Model to the Frequency Dependency of Binaural Masking Level Differences 30. Brian C. J. Moore and Brian R. Glasberg. Suggested formulae for calculating auditory-filter bandwidths and excitation471 patterns. The Journal of the Acoustical Society of America, 74(3):750–753, September 1983. ISSN 0001-4966. . URL472 http://asa.scitation.org/doi/10.1121/1.389861.473 31. International Organization for Standardization. Acoustics — reference zero for the calibration of audiometric equipment474 — part 7: Reference threshold of hearing under free-field and diffuse-field listening conditions. ISO 389-7:2019, 2019. EN475 ISO 389-7:2019, Edition 3.476 32. Rainer Beutelmann, Thomas Brand, and Birger Kollmeier. Prediction of binaural speech intelligibility with frequency-477 dependent interaural phase differences. The Journal of the Acoustical Society of America, 126(3):1359–1368, September478 2009. ISSN 0001-4966. . URL http://asa.scitation.org/doi/10.1121/1.3177266.479 18