A Blind Binaural Real-Time Model for Listening Effort Evaluated Using Continuous Subjective Listening Effort Rating
Full text
Submitted to Acta Acustica Template provided by EDP Sciences Research article 1 A blind binaural real-time model for listening effort evaluated2 using continuous subjective listening effort rating3 Martin Berdau1⋆, Daniel-Jos´e Alcala Padilla1, Thomas Brand1, Christian Rollwage2, and Jan Rennies1,2 4 1CvO University Department f¨ur Medizinische Physik und Akustik und Exzellenzcluster “Hearing4All”, 26129 Oldenburg,5 Germany6 2Fraunhofer Institute for Digital Media Technology, Branch of Hearing Speech and Audio Technology, 26129 Oldenburg,7 Germany8 Abstract – A blind binaural real-time model for estimating listening effort (LE) was developed. The model9 consists of a binaural front-end, followed by a monaural back-end. As front-end, a novel blind real-time10 implementation of the binaural speech intelligibility model (BSIM) was developed, which models spatial11 release from masking by considering binaural unmasking and better-ear listening simultaneously. As the12 back-end, a neural network was used, which was trained on inputs and outputs of a LE prediction model13 based on phoneme classification called Listening Effort prediction from Acoustic Parameters (LEAP). A14 novel method for evaluating binaural real-time models of LE was developed, where simulated scenes with a15 target speaker and a noise interferer were used, which were either co-located or spatially separated. Dynamic16 changes were introduced to the scene by abruptly altering the signal-to-noise ratio and/or reverberation17 time. Participants continuously rated subjectively perceived LE using a slider interface with LE categories,18 while listening to the scenes via headphones. The model accurately predicted subjective LE, especially19 changes in signal-to-noise ratio and binaural benefits. It also predicted detrimental effects of reverberation20 as observed in the experiment, although the impact of reverberation was slightly overestimated. Human21 response times were estimated for further tweaking the models integration time.22 Keywords. Listening effort, Real-time, Spatial hearing23 1. Introduction24 Estimating speech intelligibility (SI) or listening effort (LE) in real-time can be valuable for several online speech25 processing applications. For example, a hearing aid device could utilize such estimates to choose the optimal algorithm26 for varying listening conditions. This study proposes a real-time model to predict LE, along with a novel experimental27 ⋆Corresponding author: [email protected] 1
M. Berdau et al.: Blind binaural real-time model for listening effort assessment method to capture subjectively perceived LE in dynamically changing listening conditions.28 Regarding the assessment of speech perception, SI has been a subject of research for several decades, while in recent29 years LE emerged as another relevant speech perception measure [1,2]. SI and LE are related in a sense that SI is low30 when LE is high and vice versa. However, while SI might be very high even when noise is present, LE still provides31 additional information due to the fact that still some effort is required to understand speech. As a result, changes in32 LE can be observed even if SI is at ceiling [3], making LE more relevant for desired comfortable listening conditions,33 e.g., in case of a positive signal-to-noise ratio (SNR).34 Methods of assessing LE can be divided into physiological, subjective and performance-based procedures: Common35 physiological procedures are electroencephalography (EEG) [4,5], pupillometry [6,7] and skin conductance [8].36 Common subjective procedures are questionnaire assessments [9,10], ratings on a scale and categorical ratings [11,12].37 Performance-based procedures typically include a second (not necessarily auditory) task while performing a listening38 task, and LE is then operationalized in terms of performance measures in the secondary task [13,14]. Several models of39 SI and LE consider that listening with two ears can improve speech recognition, and that those measures are affected40 by the spatial constellation of sound sources and the listener: Sound coming from a lateralized position arrives earlier41 at the ipsilateral ear. Likewise, the signal arriving at the contralateral ear is attenuated due to the head-shadowing42 effect. The difference in time of arrival (or in phase) and in level caused by lateralization is referred to as interaural43 time differences (ITD) or interaural phase differences (IPD) and interaural level differences (ILD), respectively. The44 extent of those differences depends on the degree of lateralization and, in general, SI is improved and LE is reduced45 when speech and noise sources have different ILDs and/or ITDs [15]. Two main factors contribute to this spatial46 release from masking (SRM): binaural unmasking (BU) and better-ear listening (BEL). BEL occurs when ILDs cause47 an improvement of the SNR. BU describes an additional unmasking effect that occurs when speech and interferer48 differ in their ITDs (or IPDs). Note, that BEL and BU can occur to different amounts in different frequency regions.49 To account for spatial effects in realistic listening scenarios, prediction models for SI and LE should therefore implement50 BU and/or BEL. One approach to model both effects at once has been proposed with the binaural speech intelligibility51 model (BSIM) [16,17], which implemented the Equalization Cancellation (EC) model [18]. The EC model assumes that52 BU can be modeled by removing ILDs and ITDs of the interferer from both ear signals (Equalization) and subtracting53 them (Cancellation), yielding an effective SNR improvement due to destructive interference. The BSIM first applies a54 gammatone filterbank to each ear signal, before applying the EC [18] processing in each filterband independently. Note55 that due the fact that the EC model is applied only to these bandpass-filtered signals, it has no practical relevance for56 the predicted masking effect whether the EC processing is implemented using time processing or phase processing. EC57 processing is done by determining the set of ITDs and ILDs that maximizes the SNR. Then, for each filterband out of58 both unprocessed ear signals and the EC-processed signal, the one yielding the highest SNR is chosen for determining59 SI using the speech intelligibility index (SII) [19]. Consequently, the model does not distinguish between BEL and BU,60 but models them in one common process.61 A shortcoming of these models is that they are not applicable in real-world scenarios, since they require the use of62 clean source signals, which are not available to, e.g., a hearing aid. To overcome this limitation, blind models have63 2
M. Berdau et al.: Blind binaural real-time model for listening effort been proposed that do not require auxiliary information and can derive their predictions based only on the signals64 actually arriving at the ears. One notable extension to the BSIM in that regard was introduced in [20], where the65 binaural processing was re-worked to work blindly. However, for estimating SI from the processed, signals the SII [19]66 was used as a back-end, which still required clean source signals. The model framework was further extended in [21],67 which proposed a model consisting of blind stages only: the blind BSIM [20] was combined with a blind back-end called68 Listening Effort prediction from Acoustic Parameters (LEAP) [22–24] instead of the SII. LEAP utilizes a phoneme69 classifier to make predictions of SI and LE. While the model framework of [21] could, in principle, predict SI and LE70 from the binaural ear signals alone, it still runs offline and cannot be used to assess SI or LE in real-time, which is the71 target application of the present study. Therefore, the model framework is further developed, partly simplified, and72 implemented more efficiently, resulting in a blind and real-time capable prediction model of LE as described in detail73 below.74 The real-time model is intended to realistically predict how human LE perception changes in time-varying acous-75 tic conditions. As far as we know, there is no established method to measure time-varying subjective LE in human76 listeners. One common way to measure subjective LE is to present a stimulus (e.g., a sentence) and then have the77 listeners rate the perceived effort on a categorical scale [12]. During the presentation of the stimulus, the acoustic78 scene (SNR, reverb, spatial source configuration, ...) is usually kept constant. In the present study, we were interested79 in how perceived LE changes when acoustic parameters change abruptly, and how the predictions of the developed80 model align with these changes. To this end, a continuous LE assessment method was developed as described below.81 Conceptually similar continuous psychoacoustic assessment methods were, e.g., used in the context of perceived loud-82 ness for stimuli altering in level over the course of time [25–27]. Participants were able to continuously rate loudness83 by either touching a switch next to seven presented loudness categories [25,26] or using a slider-like interface on a84 computer screen [27]. However, those experiments were dedicated to investigating the relationship between instanta-85 neous loudness and overall loudness and not to evaluate continuous model predictions. One example of a real-time86 assessment of speech stimuli for evaluating a model was proposed in [28], where listeners continuously rated speech87 quality while time-varying distortions were introduced. The course of speech distortions was pre-generated and called88 profiles. Each participant heard the same two profiles while rating speech quality continuously using a hardware slider89 with five quality categories. The subjective ratings were then compared to predictions of a speech quality model with90 time-varying output computed for time windows shifted along the stimuli.91 The present study employed a similar concept to investigate perceived LE: time-varying profiles of speech subjected92 to different amounts of noise and/or reverberation were created in conditions with separated and co-located sound93 sources. These profiles were continuously rated by listeners using a novel, touchscreen-based interface to investigate94 the perceptual dynamics of LE. Additionally, reaction times of participants related to sudden changes in noise and95 reverberation were estimated. The subjective ratings of LE and model predictions were quantitatively compared to96 assess the model’s prediction accuracy. The analyses also allow to determine necessary integration times for real-time97 LE evaluations, as a first step towards model-guided hearing devices in the future.98 3
M. Berdau et al.: Blind binaural real-time model for listening effort Figure 1. Overview of the front-end model. Horizontal dashed lines indicate a switch between time and spectral domain. The vertical solid line indicates the frequency dependent split of gammatone-filterbands into EC bands and BEL bands. Processing stages marked with * use the overlap-add procedure afterwards. 2. Methods99 2.1. Blind real-time model100 The model framework of this study was the same as in [21]. A binaural front-end processed the left and right ear101 signals using an EC mechanism combined with BEL. It produced a binaurally enhanced mono output signal, which102 was then further processed by a LEAP back-end. However, the model stages had to be modified and re-implemented103 in order to improve computational efficiency to make it applicable for real-time usage as described below.104 2.1.1. Model front-end105 The front-end takes key features of the BSIM version from [20] and implements those for real-time usage in a106 block processing scheme. An overview of the front-end model can be seen in Figure 1. The model takes the left and107 right ear signals as blocks of audio xL t[k] and xR t[k] as inputs. Here, tdenotes time frames and kdenotes samples. An108 overlap-add procedure with an overlap of 50% is used for the block processing. Therefore, blocks of audio are of length109 2K= 1024 samples and updated by frames of length K= 512, where Kis referred to as the hopsize. The model110 operates at a sampling frequency of 16 kHz leading to a hopsize of 32 ms and blocksize of 64 ms. The left and right111 ear spectra XL/R t[f] are computed using the fast Fourier transform (FFT):112 XL/R t[f] = FFT nw[k]·xL/R t[k]o.(1) 4
M. Berdau et al.: Blind binaural real-time model for listening effort Here w[k] denotes a square-root Hann window of length 2K, which is required, since the same window is used again113 in the reconstruction phase of the processed signal using the overlap-add procedure. The index fdenotes frequency114 bins. Only half the spectrum up to the Nyquist frequency of 8 kHz is used for efficiency since real-valued signals are115 processed, resulting in spectra of length F=4K 2+ 1 = 1025. Before applying the FFT, zero padding to 4Kis applied116 to the block of audio xL/R t[k] to avoid circular convolution artifacts, when doing spectral filtering later.117 A gammatone filterbank [29] is used since the auditory processing is assumed to work for auditory bandpass filters118 independently. Filters are applied in the spectral domain instead of time domain for efficiency to obtain the bandpass-119 filtered spectra YL/R t,b [f] according to:120 YL/R t,b [f] = Hb[f]·XL/R t[f].(2) Hb[f] denotes the transfer function of the gammatone filterbank from [29], which were calculated from the filter121 coefficients. The index bdenotes auditory filters with b= 1, ..., B, where B= 29 is the total number of bandpass122 filters.123 The front-end assumes simultaneous EC processing and BEL. Filters with a center frequency below fsplit = 1.5 kHz124 are processed using the EC procedure, which are 15 filterbands in total. Conversely, filters with a center frequency125 above fsplit are processed using BEL selection, which are 14 filterbands in total.126 For each EC filterband, first, the instantaneous ILD is estimated from the ear spectra YL/R t,b [f] prior to equalization,127 as128 ILDinst t,b = 20 ·log10 sPf|YL t,b[f]|2 Pf|YR t,b[f]|2!.(3) The actual ILD values ILDt,b are computed by taking the average over the frames of the past 300 ms to simulate129 the sluggishness of the human binaural processing [30]. The first equalization step is then performed resulting in130 ILD-equalized spectra YL,eq1 t,b [f] and YR,eq1 t,b [f] for the left and right ear:131 YL,eq1 t,b [f] = YL t,b[f]·10−0.5ILDt,b 20 , YR,eq1 t,b [f] = YR t,b[f]·10+0.5ILDt,b 20 . (4) Equalization is applied in a symmetric fashion, where the left ear is attenuated by half the ILD, while the right ear is132 amplified by half the ILD.133 Then, ITDs are calculated by first estimating the instantaneous IPD for each filterband. For this, the cross power134 spectral density (CPSD) between the right and left ear ΦLR t,b [f] is computed bandwise. The frequency bin-wise phase is135 retrieved by computing the angle of the CPSD and performing phase unwrapping over frequency bins. Instantaneous136 IPDs are then computed as:137 IP Dinst t,b =X f ∠ΦLR t,b [f]|ΦLR t,b [f]| Pf|ΦLR t,b [f]|,(5) 5
M. Berdau et al.: Blind binaural real-time model for listening effort where ∠denotes the angle operator to retrieve the bin-wise unwrapped phase.138 Instantaneous ITDs are then calculated applying the following equation:139 IT Dinst t,b =IP Dinst t,b 2·π·cfb .(6) Here cfbdenotes the center frequencies of each auditory filterband. Also, instantaneous ITDs are computed by taking140 the average over the frames of the past 300 ms resulting in IT Dt,b. The second equalization step is then performed141 resulting in fully equalized spectra YL,eq2 t,b [f] and YR,eq2 t,b [f] for the left and right ear according to142 YL,eq2 t,b [f] = YL,eq1 t,b [f]·e−j2πf·0.5·IT Dt,b , YR,eq2 t,b [f] = YR,eq1 t,b [f]·e+j2πf·0.5·IT Dt,b . (7) ITDs are also equalized in a symmetric fashion. Finally, the cancellation step is performed. Similarly, as done in [20],143 two approaches are used: The first one is supposed to increase the SNR by means of destructive interference minimizing144 signal energy of an interfering source by subtracting the equalized ear signals. This is the usual way the EC model [18]145 is applied. The second one aims to increase the SNR by means of constructive interference maximizing signal energy146 of a target source by adding the equalized ear signals. Therefore, the signals are computed according to147 Ymin t,b [f] = YL,eq2 t,b [f]−YR,eq2 t,b [f], Ymax t,b [f] = YL,eq2 t,b [f] + YR,eq2 t,b [f]. (8) Ymin t,b [f] denotes the signal obtained by using the level minimization strategy, while Ymax t,b [f] denotes the signal obtained148 by using the level maximization strategy. The effectiveness of the cancellation step in terms of improved SNR is very149 dependent on the accuracy of estimated ILDs and ITDs. Also, ITDs of target and interferer, respectively, have to be150 different for the binaural processing to produce a notable SNR improvement. In order for the minimization strategy151 to work, estimated ILDs and ITDs have to match the actual ones of the interferer signal. Likewise, the maximization152 can only lead to an increase in SNR, if estimated ILDs and ITDs are close to the ones of the target signal.153 The decision which one of the two alternative spectra Ymin t,b [f] and Ymax t,b [f] is chosen is done in the selection step: Since154 this is also supposed to work blindly, a measure is needed to determine which EC strategy leads to the greater SNR155 improvement. For this, a very simple speech modulation detection approach inspired by the speech-to-reverberation156 modulation energy ratio (SRMR) by [31] is used.157 Since the SRMR requires analytic signals for each gammatone filterband, the analytic time signal is calculated using158 the inverse FFT as:159 ˆyt[k] = iFFT {u[f]·Yt[f]}, with u[f] = 2 if f < F 1 if f=F 0 if f > F, (9) 6
M. Berdau et al.: Blind binaural real-time model for listening effort where u[f] denotes an adapted step function centered at the Nyquist frequency at frequency bin Fand ˆyn[k] denotes160 the analytic signal in the time domain of any spectrum Yn[f]. Zero-padding introduced before applying the FFT in161 Eq. 1is reversed by cropping the length of ˆyn[k] to 2Kagain.162 Signal reconstruction is achieved by using the overlap-add method:163 ˆyola t,b [k] = w[k]·ˆyt−1,b[k] + w[k+K]·ˆyt,b[k+K].(10) Here w[k] denotes the square-root Hann window, as used in Eq. 1, and ˆyt,b any bandpass-filtered analytic signal in164 time domain of length K.165 Following the steps of the SRMR [31], the envelope of any analytic signal ˆyola t,b [k] is computed as the absolute value166 of ˆyola t,b [k]. In [31] the envelopes were routed into a modulation filterbank of eight modulation filters, where the lowest167 four were assumed to represent speech and the highest four assumed to represent noise. We found this approach as168 too costly in terms of processing speed, which is crucial for an application in real-time. Instead, one modulation filter169 for detecting modulation typical for speech and one for detecting modulation typical for noise are used. Butterworth170 bandpass filters of first order are used with a center modulation frequency of 8 Hz for detecting speech modulation171 and 32 Hz for detecting noise modulation. Both filters have a bandwidth of two octaves each, resulting in a -3 dB172 crossover of filters at a modulation frequency of 16 Hz.173 The modulation energy for speech and noise ϵs/n t,b is calculated using174 ϵs/n t,b = m·K X k zs/n t,b [k]2,(11) where zs/n t,b [k] represents an internal buffer of 256 ms length, which is updated for every frame. The length was chosen as175 in [31], to detect speech modulations at modulation frequencies below 8 Hz. Note, that in contrast to [31], where whole176 signals were processed offline, here, no averaging of modulation energy over time frames was performed. Additionally,177 since an estimate of the modulation energy is required for each gammatone filterband, no averaging of energy is178 performed over gammatone filterbands.179 Finally, a ratio of speech-to-noise modulation energy is calculated:180 Ratiot,b =ϵs t,b ϵn t,b .(12) The selection whether to use the minimization or maximization strategy is done by checking whichever out of the181 signals ˆymin t,b [k] and ˆymax t,b [k] yields a higher ratio for each EC filterband. Similarly, BEL is implemented for each BEL182 filterband by checking, whichever ear signal out of ˆyL t,b[k] and ˆyR t,b[k] yields a higher ratio. Selected filterband signals183 are added again using overlap-add according to Eq. 10 for blending consecutive selected frames using a Hann window.184 Finally, the real parts of selected bands ˆysel t,b [k] are added up to one final output signal yout t,b [k], which is the output of185 the front-end model:186 yout t[k] = B X b Re{ˆysel t,b [k]}.(13) 7
M. Berdau et al.: Blind binaural real-time model for listening effort 2.1.2. Model back-end187 An end-to-end version of LEAP serves as the back-end. It was developed with the aim of using LEAP in real-time188 applications, for which its latency and computational efficiency had to be reduced significantly. To achieve this, a deep189 learning model was trained on the outputs of the original LEAP model in order to mimic its behavior of estimating LE.190 This way, the time-consuming operations of LEAP are avoided whilst retaining most of its estimation performance.191 The neural network architecture shown in Figure 2was used. It is based on QuartzNet [32] and primarily consists of192 one-dimensional convolutions (1D-Conv). At the input layer, mel-spectrograms of M= 35 time frames and N= 100193 frequency bins are fed into the model. They are then processed by a 1D-Conv layer with kernel size of 3 along the time194 dimension and a rectified linear unit (ReLU) for activation. Zero padding is used to keep the data shape constant.195 The main processing stage consists of time-channel separable convolutions (TCS-Conv), also known as depthwise196 separable convolution, which are based on 1D-Conv operations that utilize the channel dimension in order to process197 two-dimensional data. TCS-Conv comprise a depthwise convolution (D-Conv) and a pointwise convolution (P-Conv).198 First, a D-Conv is applied where for each of the Nchannels (frequency bins) separate kernels that span across all199 Mtime frames are used. For P-Conv, a single kernel of size 1 that spans across all Nchannels yields outputs for200 each individual time frame. All input and output channels share a weighted connection with a total of N2weight201 parameters, for a constant number of channels. A single TCS-Conv, as used in this work, requires MN +N2+ 2N202 model parameters.203 In the main processing stage, five blocks, each containing five TCS-Convs with ReLU activation, are applied. The204 data shape is retained at N×Mby using zero padding and keeping the number of channels constant for all Conv205 operations. An additive skip connection spans from the beginning to the end of each block and contains an extra206 P-Conv layer.207 After these TCS-Conv blocks, feature data are processed by means of a 1D-Conv layer with kernel size 3 without zero208 padding. Max pooling is applied such that the time dimension is reduced to a single frame, resulting in a vector of209 length N. Next, two fully connected (FC) layers with intermediate ReLU activation reduce the channel dimension to210 a single frequency bin, leaving a single scalar value as their output. The output is then scaled to a value between 0211 and 1 with a sigmoid function. Mapping this value owith212 ˆ le = o·12 + 1 (14) yields the final estimation of listening effort ˆ le. Overall, the neural network comprises 458,301 trainable model213 parameters.214 The model was trained on mixture signals of German speech and noise. Training and validation sets were generated215 in the same way, but with different, non-overlapping splits of the employed data corpora. Speech data were taken from216 the German version of CommonVoice 7.0 [33], a crowdsourcing-based corpus with a wide variety of speakers and signal217 quality. Since only 8 % of the 1035 hours of data were recorded by female speakers, most of the utterances recorded by218 males were discarded in order to achieve a balanced representation of both genders. The data with validated transcripts219 8
M. Berdau et al.: Blind binaural real-time model for listening effort Figure 2. Neural network structure of end-to-end LEAP. were used for the training set, those with unvalidated scripts for the validation set. Speechocean King-ASR-187 [34], a220 corpus containing about 100 hours of German speech data, was additionally used for the training set. The combined221 speech data of both corpora were used a total of three times to reach around 1000 hours of speech data. Noise data were222 taken from several different corpora: BBC Sound FX [35] contains a variety of noise types, ranging from human-made223 sounds to engine noises. MUSAN [36] comprises 929 noise recordings and more than 42 hours of music of different224 genres. The speech data of MUSAN were not used in this work. 997 noise recordings from the Corpus provided as part225 of the Deep Noise Supression (DNS) Challenge [37] were also used. Training data were built using all the noise from226 the BBC and DNS corpora, as well as 2/3 of the music of MUSAN. The remaining 1/3 of MUSAN’s music as well as227 all of its noise data were used for the validation set.228 Pairs of speech and noise signals were randomly selected and the noise adjusted to match the speech signal’s length.229 This was done by truncation if the noise signal was longer than the speech signal or by repetition if it was shorter.230 Speech and noise were added at a random SNR of -10, -6, -2, 2, 6, 10, 14, 18, or 22 dB. SNRs were applied using the231 active speech level according to ITU P.56 [38]. 28 % of the total data were clean speech without any noise. 95 % of232 the data were convolved with one of 60k simulated room impulse responses (RIR), originating from the corpus of [39].233 In order to make the model more robust against other disturbances besides noise, the training set was extended with234 9
M. Berdau et al.: Blind binaural real-time model for listening effort Figure 7. Median of predicted LE plotted against the median of subjective LE in ESCU for each section. Data points are grouped by the six different profiles. The black line represents the linear regression line, while the results of the linear regression analysis are displayed in the top left corner. and 0.81, respectively. The slope of the linear regression line was far below 1.0 except for the condition T60 only in364 S0, while the bias across profiles varied in a range of 1.95 to 4.60 ESCU. While the root mean sqare error (RMSE) for365 all profiles was at 1.0 ESCU, it was notably lower for individual profiles.366 367 3.3. Binaural benefit368 To analyze the binaural benefit, the same steady-state LE values calculated above were used. For each participant369 and section ending, subjective LE values for spatially separated sources were subtracted from subjective LE values for370 co-located sources. The same was done for the model predictions. The results grouped by SNR and T60 are displayed371 as violin plots in Figure 8. IQRs are shown as boxes, where the white line displays the median and whiskers are372 Table 2. Results of a linear regression analysis performed on subjective and predicted LE for each profile as well as all profiles combined. Columns show values for R2, slope as well as bias of the regression line and the RMSE. Profile R2Slope Bias /ESCU RMSE /ESCU SNR only (S0N0) 0.95 0.76 2.18 0.74 T60 only (S0) 0.92 1.06 2.05 0.69 SNR+T60 (S0N0) 0.84 0.59 4.60 0.84 SNR only (S−45N+45) 0.83 0.72 1.95 0.96 T60 only (S−45) 0.94 0.81 2.78 0.50 SNR+T60 (S−45N+45) 0.81 0.64 3.92 0.87 All 0.87 0.73 2.85 1.0 16
M. Berdau et al.: Blind binaural real-time model for listening effort Figure 8. Spatial release of LE for participants (blue) and the model (gray) due to lateralization of sound sources for the SNR only (upper panel) and T60 only (lower panel) profiles visualized as violin plots with box plots. Values are grouped by SNR and T60 values, respectively. Black boxes display IQRs, where the median is indicated as a white line inside. Whiskers are added as black lines. Note that for the T60 only profile and a T60 of 0 s no box and whiskers were drawn since most data points are at 0 ESCU. added as lines. Note, that only data from the profiles adapting single parameters are shown, since for SNR+T60 not373 all possible combinations of SNR and T60 values were used. In case of the SNR being adapted, the median binaural374 LE benefit for participants and the model decreased with increasing SNR. For the lowest presented SNR of -12 dB,375 median values for participants and the model amounted to 2.5 ESCU, respectively. The median LE benefit dropped376 17
M. Berdau et al.: Blind binaural real-time model for listening effort to 0 ESCU for participants as well as for the model at the highest SNR.377 In case of the speech source being lateralized to -45°for the T60 only profile, the median LE release was at around378 0 ESCU for participants regardless of presented T60. In contrast to that, the median predicted LE release by the model379 amounted to about -1 ESCU, indicating a higher LE being induced by the lateralization. However, with increasing380 T60, the predicted spatial release of LE increased to about 1 ESCU for a T60 of 5 s.381 3.4. Response times to abrupt changes382 The cross-correlation between each participant’s time-varying LE assessments and profiles was calculated for each383 of the six profiles to estimate the response times of participants similar to [28]. The point in time of each cross384 correlation peak was determined to obtain the delay of participants.385 The results can be seen as violin plots in Figure 9. Median response times were determined in a range from 1.54386 to 2.16 s, while overall response times appeared to be greatest for the SNR+T60 profile and lowest for the T60 only387 profiles by a slight margin compared to the response times for the SNR only profiles. A Shapiro-Wilk test revealed388 that not all distributions were normally distributed. Therefore, a Kruskal-Wallis test was applied, which returned that389 at least one distribution differed significantly from one other distribution (H=23.47, p<0.001). Post-hoc Dunn’s tests390 with Bonferroni correction revealed that in S0N0response times for the SNR+T60 profile differed significantly from the391 one for the T60 only profile at a significance level of α= 0.05. Response times for the SNR+T60 profile in S−45N+45 392 differed significantly at a level of α= 0.01 from the ones for the T60 only profile regardless of spatial configuration.393 4. Discussion394 The time courses in Figure 4to 6indicate that predicted LE is in line with subjective LE. This is supported by395 the high correlation values and fairly low RMSE values shown in Table 2. The main discrepancy was that the model396 Figure 9. Distribution of estimated response times in seconds for all six profiles. Significant differences are indicated with * and ** for significancy levels of α= 0.05 and α= 0.01, respectively. 18
M. Berdau et al.: Blind binaural real-time model for listening effort overestimates LE in the presence of reverberation.397 On the one hand, the model might overestimate the effect of reduced ITDs and ILDs. On the other hand, the back-end398 might not be able to handle the type of reverberation produced by TASCAR [41], although it was trained on a very399 large RIR dataset. The outputs of the LEAP model, which was used to train the proposed back-end, might also be400 biased with regard to reverberation. A short analysis was performed to find the origin of the mismatch, where the401 end-to-end LEAP model was replaced by the original version, and both LEAP versions were also run separately for402 each ear to obtain a better-ear prediction without binaural preprocessing. It was found that, in general, the overesti-403 mation decreased when using the original LEAP version, while it decreased further when running the LEAP version404 on the better ear signal only. Those results suggest that both the blind BSIM front-end and the end-to-end LEAP405 back-end contributed to the overestimation of LE in reverberant conditions. The proposed model could further be406 extended to incorporate hearing loss and adapted to being able to handle speech masking, which currently would lead407 to underestimated LE due to the back-end favoring the presence of speech in general.408 Apart from that, the low slope values combined with a positive bias suggest that the model tends to avoid to output409 extremely low and extremely high ESCU values. For high LE, this can be partially explained by the fact that par-410 ticipants were able to rate LE on a scale of up to 14 ESCU, while the model was trained on a scale of up to only411 13 ESCU. However, as can be seen in the time courses in Figure 4to 6, the median subjective LE does very rarely412 exceed values of 13 ESCU.413 While R2is very high for profiles adapting just one parameter, R2is notably lower for the SNR+T60 profile.414 As shown in Figure 8, the separation of speech and noise for the SNR only profile led to a binaural benefit in subjective415 LE as well as predicted LE, which increases with decreasing SNR. The data for participants for the T60 only profile,416 where only one speech source was presented, support this interpretation, since only diffuse reverberation was present.417 Surprisingly, here the model predicts increased LE by about 1 ESCU in the anechoic setting if the speech source is418 lateralized. Furthermore, a very subtle binaural benefit of about 1 ESCU is predicted for high reverberation.419 When comparing the results for subjective and predicted LE to the findings from [21], the overall binaural benefit420 in ESCU appears to be smaller. They observed subjective LE ratings for speech in noise using an American matrix421 sentence test [47]. Subjective LE ratings were tracked for static SNR values ranging from -25 to +10 dB for different422 spatial settings. Here, the speech source was located in front of the receiver and the noise source was moved to different423 positions, which differs from the procedure in this work. The S0N100 condition in [21] is the closest to the S−45N+45 424 in this study judging by the angular distance between the speech and noise sources. Comparing the difference in LE425 for the S0N100 and S0N0conditions in [21] yields a binaural benefit of about 6 ESCU at an SNR of -10 dB. This is a426 huge difference to the binaural benefit of 2.5 ESCU at an even smaller SNR of -12 dB in this study.427 The response times for judging LE appeared to be a little higher when SNR and T60 were adapted simultaneously,428 while the average response time was just below 2 s. Compared to the findings of continuous speech quality in [28],429 where response times around 1 s were observed, this is a large difference. A comparable response time was only achieved430 by very few participants (Figure 9). Nevertheless, there are some notable differences between [28] and the procedure431 described here. A simple explanation could be that speech quality degradation caused by signal distortion is faster to432 19
M. Berdau et al.: Blind binaural real-time model for listening effort rate than induced LE due to changes in SNR and T60. In some cases, parameter changes could have had a very little433 effect, so that participants reacted slower. This can be seen for example for SNR changes of just 3 dB in Figure 4or434 for changes in T60 of 3 s and of 5 s in Figure 5. Another reason for the difference in response times could have been435 caused by the different groups of participants. In [28] participants were recruited from the lab and thus probably had436 experience with speech quality measurements or comparable experiments. Therefore, it might have been easier for437 them to rate speech quality. Participants recruited in this study have mostly not taken part in any hearing experiment438 evolving around LE, which is why they might have had to come up with an idea, how to rate LE over the course of the439 experiment. Another factor might have been the use of abrupt parameter changes only, while in [28] gradual parameter440 changes were observed as well. While the occurrence of an abrupt change in this study was unpredictable, gradual441 changes could be easier to follow once the onset of such a parameter change is noticed. This could have lowered the442 average response times estimated by calculating the cross-correlation of subjective LE courses and the profiles. Lastly,443 the overall mental load throughout the experiment from this study could have been higher, since participants had to444 also pay attention to the story contents and stimuli of 5 min were considerably longer than the 40 s stimuli from [28].445 The response times estimated in this study could be used to determine smoothing constants of the model outputs.446 However, the actual human processing time to perceive LE is still unknown, since the response time also includes the447 delay to transmit the percept of LE on the slider interface.448 5. Conclusions449 This work introduced a blind binaural model for predicting LE, which runs in real-time.450 A method for assessing subjective continuous LE was introduced to evaluate the proposed model. The presented451 lightweight and easy setup could in theory be used for in-the-field experiments with focus on ecological validity.452 The results showed that the model can quite accurately predict subjective LE for changes in SNR. However, the effect453 of reverberation is slightly overestimated.454 Acknowledgments455 We thank Michel B¨urgel, who further adapted the speech material and the questionnaires from [43] in his Master’s456 thesis, which we used in this study. Also, we thank Lukas Wandelt for helping with implementing the slider interface.457 Funding458 This work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) Project ID459 352015383 - SFB 1330 A1.460 Conflicts of interest461 The authors declare no conflict of interest.462 20
M. Berdau et al.: Blind binaural real-time model for listening effort Data availability statement463 The experimental data will be made available on request. The code of the binaural BSIM model will be published464 and the link added here.465 Ethical approval466 All participants received hourly compensation and gave informed consent for their participation in the experiments.467 The methods were approved by the ethics committee of the University of Oldenburg (protocol Drs.-Nr. 04/2018).468 References469 1. Ronan McGarrigle, Kevin J Munro, Piers Dawes, Andrew J Stewart, David R Moore, Johanna G Barry, and Sygal Amitay.470 Listening effort and fatigue: What exactly are we measuring? a british society of audiology cognition in hearing special471 interest group ‘white paper’. International journal of audiology, 53(7):433–445, 2014.472 2. M Kathleen Pichora-Fuller, Sophia E Kramer, Mark A Eckert, Brent Edwards, Benjamin WY Hornsby, Larry E Humes,473 Ulrike Lemke, Thomas Lunner, Mohan Matthen, Carol L Mackersie, et al. Hearing impairment and cognitive energy: The474 framework for understanding effortful listening (fuel). Ear and hearing, 37:5S–27S, 2016.475 3. Jan Rennies, Henning Schepker, Inga Holube, and Birger Kollmeier. Listening effort and speech intelligibility in listening476 situations affected by noise and reverberation. The Journal of the Acoustical Society of America, 136(5):2642–2653, 2014.477 4. Jonas Obleser, Malte W¨ostmann, Nele Hellbernd, Anna Wilsch, and Burkhard Maess. Adverse listening conditions and478 memory load drive a common alpha oscillatory network. Journal of Neuroscience, 32(36):12376–12383, 2012.479 5. Corinna Bernarding, Daniel J Strauss, Ronny Hannemann, Harald Seidler, and Farah I Corona-Strauss. Neural correlates480 of listening effort related factors: Influence of age and hearing impairment. Brain research bulletin, 91:21–30, 2013.481 6. Sophia E Kramer, Theo S Kapteyn, Joost M Festen, and Dirk J Kuik. Assessing aspects of auditory handicap by means of482 pupil dilatation. Audiology, 36(3):155–164, 1997.483 7. Adriana A Zekveld, Sophia E Kramer, and Joost M Festen. Pupil response as an indication of effortful listening: The484 influence of sentence intelligibility. Ear and hearing, 31(4):480–490, 2010.485 8. Carol L Mackersie and Heather Cones. Subjective and psychophysiological indexes of listening effort in a competing-talker486 task. Journal of the American Academy of Audiology, 22(02):113–122, 2011.487 9. Sandra G Hart and Lowell E Staveland. Development of nasa-tlx (task load index): Results of empirical and theoretical488 research. In Advances in psychology, volume 52, pages 139–183. Elsevier, 1988.489 10. Stuart Gatehouse and William Noble. The speech, spatial and qualities of hearing scale (ssq). International journal of490 audiology, 43(2):85–99, 2004.491 11. Heleen Luts, Koen Eneman, Jan Wouters, Michael Schulte, Matthias Vormann, Michael Buechler, Norbert Dillier, Rolph492 Houben, Wouter A Dreschler, Matthias Froehlich, et al. Multicenter evaluation of signal enhancement algorithms for hearing493 aids. The Journal of the Acoustical Society of America, 127(3):1491–1505, 2010.494 12. Melanie Krueger, Michael Schulte, Thomas Brand, and Inga Holube. Development of an adaptive scaling method for495 subjective listening effort. The Journal of the Acoustical Society of America, 141(6):4680–4693, 2017.496 21
M. Berdau et al.: Blind binaural real-time model for listening effort 13. Penny Anderson Gosselin and Jean-Pierre Gagn´e. Use of a dual-task paradigm to measure listening effort utilisation d’un497 paradigme de double tˆache pour mesurer l’attention auditive. Inscription au R´epertoire, 34(1):43, 2010.498 14. Jani Johnson, Jingjing Xu, Robyn Cox, and Paul Pendergraft. A comparison of two methods for measuring listening effort499 as part of an audiologic test battery. American journal of audiology, 24(3):419–431, 2015.500 15. Jan Rennies and Gerald Kidd. Benefit of binaural listening as revealed by speech intelligibility and listening effort. The501 Journal of the Acoustical Society of America, 144(4):2147–2159, 2018.502 16. Rainer Beutelmann and Thomas Brand. Prediction of speech intelligibility in spatial noise and reverberation for normal-503 hearing and hearing-impaired listeners. The Journal of the Acoustical Society of America, 120(1):331–342, 2006.504 17. Rainer Beutelmann, Thomas Brand, and Birger Kollmeier. Revision, extension, and evaluation of a binaural speech intel-505 ligibility model. The Journal of the Acoustical Society of America, 127(4):2479–2497, 2010.506 18. Nathaniel I Durlach. Equalization and cancellation theory of binaural masking-level differences. The Journal of the507 Acoustical Society of America, 35(8):1206–1218, 1963.508 19. ANSI ANSI. S3. 5-1997, methods for the calculation of the speech intelligibility index. New York: American National509 Standards Institute, 19:90–119, 1997.510 20. Christopher F Hauth, Simon C Berning, Birger Kollmeier, and Thomas Brand. Modeling binaural unmasking of speech511 using a blind binaural processing stage. Trends in Hearing, 24:2331216520975630, 2020.512 21. Jan Rennies, Saskia R¨ottges, Rainer Huber, Christopher F Hauth, and Thomas Brand. A joint framework for blind513 prediction of binaural speech intelligibility and perceived listening effort. Hearing Research, 426:108598, 2022.514 22. Rainer Huber, Arne Pusch, Niko Moritz, Jan Rennies, Henning Schepker, and Bernd T Meyer. Objective assessment515 of a speech enhancement scheme with an automatic speech recognition-based system. In Speech Communication; 13th516 ITG-Symposium, pages 1–5. VDE, 2018.517 23. Rainer Huber, Melanie Kr¨uger, and Bernd T Meyer. Single-ended prediction of listening effort using deep neural networks.518 Hearing research, 359:40–49, 2018.519 24. Rainer Huber, Hannah Baumgartner, Stefan Goetze, and Jan Rennies. Asr-based, single-ended modeling of listening effort?520 a tool for tv sound engineers. In Forum Acusticum, pages 2441–2445, 2020.521 25. Sonoko Kuwano and Seichiro Namba. Continuous judgment of level-fluctuating sounds and the relationship between overall522 loudness and instantaneous loudness. Psychological research, 47(1):27–37, 1985.523 26. Seiichiro Namba, Sonoko Kuwano, and Hugo Fastl. Loudness of road traffic noise using the method of continuous judgment524 by category. In Noise as a Public Health Problem. Swedish Council for Building Research, 1988.525 27. Hugo Fastl, Sonoko Kuwano, and Seiichiro Namba. Assessing the railway bonus in laboratory studies. Journal of the526 Acoustical Society of Japan (E), 17(3):139–148, 1996.527 28. Martin Hansen and Birger Kollmeier. Continuous assessment of time-varying speech quality. The Journal of the Acoustical528 Society of America, 106(5):2888–2899, 1999.529 29. Volker Hohmann. Frequency analysis and synthesis using a gammatone filterbank. Acta Acustica united with Acustica, 88530 (3):433–442, 2002.531 30. Christopher F Hauth and Thomas Brand. Modeling sluggishness in binaural unmasking of speech for maskers with time-532 varying interaural phase differences. Trends in hearing, 22:2331216517753547, 2018.533 31. Tiago H Falk, Chenxi Zheng, and Wai-Yip Chan. A non-intrusive quality and intelligibility measure of reverberant and534 dereverberated speech. IEEE Transactions on Audio, Speech, and Language Processing, 18(7):1766–1774, 2010.535 22
M. Berdau et al.: Blind binaural real-time model for listening effort 32. Samuel Kriman, Stanislav Beliaev, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason536 Li, and Yang Zhang. Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions. In ICASSP537 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6124–6128. IEEE,538 2020. ISBN 978-1-5090-6631-5. .539 33. Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay540 Saunders, Francis M Tyers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. arXiv preprint541 arXiv:1912.06670, 2019.542 34. Beijing Haitian Ruisheng Science Technology Ltd. Speechocean: King-asr-187, n.d. URL https://en.speechocean.com/543 datacenter/details/1557.html. Accessed: 2022-01-13.544 35. BBC. BBC Sound Effects, 2020. URL http://bbcsfx.acropolis.org.uk/. Accessed: 2025-04-29.545 36. David Snyder, Guoguo Chen, and Daniel Povey. Musan: A music, speech, and noise corpus. arXiv preprint arXiv:1510.08484,546 2015.547 37. Chandan KA Reddy, Vishak Gopal, Ross Cutler, Ebrahim Beyrami, Roger Cheng, Harishchandra Dubey, Sergiy548 Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, et al. The interspeech 2020 deep noise suppression chal-549 lenge: Datasets, subjective testing framework, and challenge results. In INTERSPEECH, 2020.550 38. ITU. Objective measurement of active speech level, international telecommunications union (itu-t) recommendation p.56,551 2011.552 39. Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L. Seltzer, and Sanjeev Khudanpur. A study on data augmentation of553 reverberant speech for robust speech recognition. In 2017 IEEE International Conference on Acoustics, Speech, and Signal554 Processing, pages 5220–5224, Piscataway, NJ, 2017. IEEE. ISBN 978-1-5090-4117-6.555 40. Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.556 41. Giso Grimm, Joanna Luberadzka, and Volker Hohmann. A toolbox for rendering virtual acoustic environments in the557 context of audiology. Acta acustica united with acustica, 105(3):566–578, 2019.558 42. OHRKA. Ohrka.de - kostenlose H¨orabenteuer f¨ur Kinderohren, 2012. URL https://www.ohrka.de/hoeren/559 abenteuerlich-lustig/zwerg-nase. Accessed: 2025-04-29.560 43. Bojana Mirkovic, Martin G Bleichner, Maarten De Vos, and Stefan Debener. Target speaker detection with concealed eeg561 around the ear. Frontiers in neuroscience, 10:349, 2016.562 44. Wouter A Dreschler, Hans Verschuure, Carl Ludvigsen, and Søren Westermann. Icra noises: Artificial noise signals with563 speech-like spectral and temporal properties for hearing instrument assessment: Ruidos icra: Se˜nates de ruido artificial con564 espectro similar al habla y propiedades temporales para pruebas de instrumentos auditivos. Audiology, 40(3):148–157, 2001.565 45. Elisabeth Hering. Kostbarkeiten aus dem deutschen M¨archenschatz. BUCHFUNK Verlag, 2011.566 46. Alvy Ray Smith. Color gamut transform pairs. ACM Siggraph Computer Graphics, 12(3):12–19, 1978.567 47. Birger Kollmeier, Anna Warzybok, Sabine Hochmuth, Melanie A Zokoll, Verena Uslar, Thomas Brand, and Kirsten C568 Wagener. The multilingual matrix test: Principles, applications, and comparison across languages: A review. International569 journal of audiology, 54(sup2):3–16, 2015.570 23