Full text
An acoustic data collection for speech enabled technologies in moving vehicles DATASET DOCUMENTATION 1
Contents 1 Introduction 4 2 Data Acquisition Process 5 2.1 Microphone setups .......................... 5 2.2 Measurement of IRs ......................... 5 2.3 Recorded conditions ......................... 6 3 Synthesis Model 6 4 Additional parameters and functions 8 5 Naming Conventions 8 5.1 Speaker Locations .......................... 8 5.2 Windows Condition ......................... 9 5.3 Speed Condition ........................... 9 5.4 Different Versions ........................... 9 5.5 Ventilation Level ........................... 9 6 Car Recordings 10 6.1 Volkswagen Golf ........................... 11 6.1.1 Distributed Microphones ................... 11 6.1.1.1 Impulse Response Recordings ........... 11 6.1.1.2 Controlled Noise Recordings ........... 12 6.1.1.3 Uncontrolled Noise Recordings .......... 13 6.1.1.4 Ventilation Recordings ............... 14 6.1.1.5 Car Audio ..................... 14 6.1.2 Microphone Array ...................... 14 6.1.2.1 Impulse Response Recordings ........... 14 6.1.2.2 Controlled Noise Recordings ........... 15 6.1.2.3 Uncontrolled Noise Recordings .......... 17 6.1.2.4 Ventilation Recordings ............... 17 6.1.2.5 Car Audio ..................... 18 6.1.2.6 Steering Vector ................... 18 6.2 Smart forfour ............................. 19 6.2.1 Distributed Microphones ................... 19 6.2.1.1 Impulse Response Recordings ........... 19 6.2.1.2 Controlled Noise Recordings ........... 20 6.2.1.3 Ventilation Recordings ............... 21 6.2.1.4 Uncontrolled Noise Recordings .......... 21 6.2.1.5 Car Audio ..................... 22 6.2.2 Microphone Array ...................... 22 6.2.2.1 Impulse Response Recordings ........... 22 6.2.2.2 Controlled Noise Recordings ........... 23 2
6.2.2.3 Uncontrolled Noise Recordings .......... 24 6.2.2.4 Ventilation Recordings ............... 24 6.2.2.5 Car Audio ..................... 25 6.2.2.6 Steering Vector ................... 25 6.3 Alfa Romeo 146 TS .......................... 26 6.3.1 Distributed Microphones ................... 26 6.3.1.1 Impulse Response Recordings ........... 26 6.3.1.2 Controlled Noise Recordings ........... 28 6.3.1.3 Uncontrolled Noise Recordings .......... 29 6.3.1.4 Ventilation Recordings ............... 29 6.3.1.5 Car Audio ..................... 29 6.4 Honda CR-V ............................. 30 6.4.1 Distributed Microphones ................... 30 6.4.1.1 Impulse Response Recordings ........... 30 6.4.1.2 Controlled Noise Recordings ........... 31 6.4.1.3 Uncontrolled Noise Recordings .......... 32 6.4.1.4 Ventilation Recordings ............... 32 6.4.1.5 Car Audio ..................... 32 6.4.2 Microphone Array ...................... 33 6.4.2.1 Impulse Response Recordings ........... 33 6.4.2.2 Controlled Noise Recordings ........... 34 6.4.2.3 Uncontrolled Noise Recordings .......... 35 6.4.2.4 Ventilation Recordings ............... 35 6.4.2.5 Car Audio ..................... 36 6.4.2.6 Steering Vector ................... 36 3
CAVEMOVE 1 Introduction CAVEMOVE is a research project dedicated to the collection of audio data for the study of voice enabled technologies inside moving vehicles. The recording process involves (i) recordings of acoustic impulse responses, which are acquired at static conditions and provide the means for modeling the speech and car-audio components and (ii) recordings of acoustic noise at a wide range of both static and in-motion conditions. Data are recorded with two different microphone configurations and particularly (i) a compact circular microphone array or (ii) a distributed microphone setup. This document provides a description of the audio recordings and acoustic impulse responses that comprise the open access dataset of CAVEMOVE. The open access dataset is available at 16 kHz sampling rate and 24 bits bit depth and is approximately 8.1 GB in size. All recordings and impulse responses are available in the form of 8-channel .wav files. The latest version can be downloaded from zenodo. Note that the open access dataset is only a subset of the full CAVEMOVE dataset. It is recommended to use this dataset along with the corresponding python API that can be downloaded from CAVEMOVE github. It is important to note that the CAVEMOVE API relies on specific naming conventions for obtaining the audio mixtures corresponding to specific scenarios, and this document lists all the naming conventions required for correct use of the API. For any questions with respect to CAVEMOVE dataset or API, feel free to send an email to Andreas Symiakakis at [email protected] or Nikos Stefanakis at [email protected] CAVEMOVE project is funded by the Institute of Computer Science of the Foundation for Research and Technology-Hellas (FORTH). 4
2 Data Acquisition Process 2.1 Microphone setups All recordings are acquired using the M-audio M-Track Eight usb audio interface, at 48kHz sampling rate and at 24-bit sample size. Eight omnidirectional microphones (Shure SM93) placed at different locations are used for recording acoustic noise and IRs. These are lightweight lavalier microphones with a flat frequency response from 80 to 20000 Hz. Two different microphone configurations have been used until now, as described below. 1. Distributed setup: While there are slight variations in the setup between different cars, the microphones are distributed in a way so that six of the microphones are at the front part of the car, while the other two are placed on the headliner at the back of the car (see Figures 1,6and 10). This setup ensures that there is at least one microphone close to each passenger. All microphones were stabilized with blue tack and were covered with a windshield. 2. Microphone array: the eight microphones are mounted on a plastic circular case, forming a circular microphone array of radius equal to 5 cm. The circular array is placed on top of the dashboard, between the driver and the front passenger (see Figures 3and 7). Microphones were mounted on the array without a windshield. 2.2 Measurement of IRs Acoustic Impulse Responses (IRs) are measured for different passenger locations in each car. Basic locations covered include that of the driver, the front passenger, the rear-left, rear-middle and rear-right passenger, assuming that the passengers are at normal height and that they are facing in-front. In several cases, IRs are also measured for slight displacements, e.g. by assuming that the passenger’s head is shifted a few centimeters from the nominal location, or assuming that the head is rotated. All IRs are captured using Room Eq Wizard with a sweep tone as the excitation signal. The loudspeaker used for these measurements is the Talkbox from NTi which is especially designed for resembling the frequency response and radiation characteristics of the human voice. The loudspeaker is calibrated and it has built-in excitation signals which allow us to excite the car interior with acoustic powers that correspond to specific speech effort, particularly normal speech effort (referenced as 60 dBA at 1 m) and high speech effort (referenced as 70 dBA at 1 m). This process is very important as it will allow us to calculate the signal levels that correspond to specific speech efforts, consequently allowing us to scale the speech components so that when mixed with the noise components, a realistic balance is achieved. Additionally to the acoustic paths from the passengers to the microphones, IRs are also measured for the built-in audio system, whenever available. This process involves recording the response from all car loudspeakers simultaneously. 5
It was accomplished by directly feeding the excitation signal required for IR estimation into the auxiliary input of the audio system, or by reproducing it through the CD player. Attention was paid so that any adjustments related to equalization settings or panning were neutralized. 2.3 Recorded conditions As basic factors that affect the composition of noise in a car we considered the following: car speed, window aperture and ventilation/air condition level. With respect to these factors, it has been attempted to have them at fixed values for continuous temporal segments so as to achieve stationary noise conditions. A basic set of so-called ”controlled” conditions is recorded in each car, spanning a specific range of driving speeds and window apertures. Particularly regarding window aperture, we consider four states designated as 0, 1, 2 and 3. Regarding the driving speeds, in each car we cover at least the range from 50 km/h to 110 km/h with steps of 10 km/h. All combinations of window apertures and driving speeds are covered in each case, excluding those that it was not possible. Additionally to these in-motion recordings, static recordings were obtained for capturing the noise produced by the built-in ventilation or air-conditioning system. Two or three different levels of ventilation power were considered in each car and microphone setup. While all ”controlled” recordings are designated with respect to the window aperture and the car speed, additional recordings are available for each car and microphone setup, in which case the aforementioned conditions are not designated. These are called ”uncontrolled” noise recordings and they capture a variety of sound sources that are not usually found within the ”controlled” recordings. The total duration of uncontrolled noise recordings is approximately 20 min for each car and microphone setup. 3 Synthesis Model The ultimate goal of the data and python API delivered in the context of CAVEMOVE is to allow the engineers and researchers to easily synthesize the microphone signals that correspond to a particular scenario. Assuming that a target car and microphone setup has been chosen, the synthesis process can be compactly described as: Y=S(p, Ls, w, x) + A(La, w, z) + N(s, w) + V(l, w) (1) Briefly, each bold capital symbol is a N×Msignal matrix, where Mis the desired number of microphone channels and Nis the duration of the synthesized sound excerpt in samples. Srepresents the filtered speech components, Arepresents the interference components produced by the built-in audio system (when available), Nare the noise components corresponding to the particular driving 6
condition and Vare the noise components associated to the ventilation/aircondition functionality. Assuming that all these sound components are independent from one another, the final microphone signal can be synthesized as in Eq. 1, by means of simple superposition. As it can be seen, each component is produced as a function of user defined parameters. A brief explanation of these parameters is as follows; •w: is an integer taking values in the range [0, 1, 2, 3], with each integer value representing a different condition with respect to the windows’ apertures, as described in more detail below. •x: is a one-dimensional vector with the dry (ideally anechoic) speech recording that is used for the particular scenario. The user is responsible for providing an appropriate speech recording in PCM format. The audio recording can be of any length and sampling rate (it will be automatically converted to the target sampling rate). •p: it can be selected from a given set of string values, each value corresponding to a particular location of the passenger inside the car cabin. For each combination of p and w the corresponding impulse response is automatically loaded and used to produce the filtered speech components by means of convolution with x. •Ls: it is a user defined value, in dBA, corresponding to the speech effort. While this can be any non-negative value, we recommend values in the range between 60 and 70, with 60 corresponding to normal speech effort and 70 to a high speech effort. •La: is a user defined value, in dBA, corresponding to the mean A-weighted acoustic level of the audio program reproduced from the built-in audio system. The corresponding interference components are automatically scaled through this value so that the signal level at a reference microphone matches the desired sound level. Again, any non-negative value is applicable and it is up to the user to set it to a realistic value. •z: is a one-dimensional vector representing the audio program. The user is responsible for providing an appropriate audio file in PCM format. The audio file can be of any length and sampling rate. Audio files with more than one channel are accepted but will be automatically downmixed to a monophonic audio signal. •s: is an integer value within a given range, where each value corresponds to a different speed in km/h. The noise components corresponding to particular speed, window aperture, car and microphone setup are automatically loaded. •l: it is an integer with values in the range [1, 2, 3], each value corresponding to a different level of the ventilation/air-conditioning system, so that 7
higher levels produce more noise. The synthesis approach is designed in such way that the resulting length Nof the output signal Y, as well as of all matrices in Eq.1matches the length of the dry speech signal xwhen converted to the target sampling rate. To achieve this, we have incorporated a mechanism that recycles the noise sequences inherent to the construction of A,Nand Vas many times necessary in order to match the length of x. Note that in most cases this will not be necessary if the dry speech signal does not exceed 25 seconds of duration. 4 Additional parameters and functions An additional parameter that affects the returned microphone signals is the use of correction gains (or not), which will apply a scaling to the produced signals so as to compensate for slight deviations in the input channel sensitivities. The sensitivity values were measured with each microphone connected to a specific input channel in the sound card, subjecting it to a pink noise sound field produced inside a semi-anechoic chamber that we have in our facilities. Although the differences between minimum and maximum sensitivity do not exceed 3.5 dB, we recommend use of correction gains, which is also the default option in all relevant function calls. Our API provides functions that are specific to each one of the aforementioned components (i.e. speech, noise, ventilation and built-in audio system) but there is also a function with which the user can synthesize all four different components with a single function call. Additional auxiliary functions are provided for example: •to automatically crop or extend the duration of different components, so that they can be all added together to produce a mix, •to derive the IRs associated to a particular passenger location and microphone setup, •to derive the steering vector associated to a particular frequency and look direction with respect to the center of the circular array, which is handy for designing beamformers in the frequency domain. 5 Naming Conventions 5.1 Speaker Locations Acoustic Impulse Responses are provided for at least the following locations inside the car: •d: driver •pf: front passenger 8
•prl: rear-left passenger •prr: rear-right passenger •prm: rear-middle passenger 5.2 Windows Condition Regarding window aperture, we consider four states designated as 0, 1, 2 and 3: •w0: corresponds to completely closed windows. •w1: corresponds to windows of the driver and front passenger being slightly open (approx.10 cm). •w2: corresponds to completely open front windows and closed back windows. •w3: corresponds to all four windows being completely open. 5.3 Speed Condition Regarding the driving speeds, in each car we cover at least the range from 50 km/h to 110 km/h with steps of 10 km/h. Speed condition is symbolized by the letter s followed by the speed e.g. s70 or s30. 5.4 Different Versions Some noise recordings come with more than one versions. To discriminate between different versions we append the file name with “ ver1 ”, “ ver2 ” etc. Also, some versions are obtained on coarse road surface, in which case we append the name with “ coarse”. 5.5 Ventilation Level Two or three levels of ventilation power were considered in each car: •v1 •v2 •v3 The version number is indicative of ventilation power, therefore 1 corresponds to the lowest and 3 to the highest ventilation setting. 9
Figure 4: Speaker position pf60. 11. s40 w1 ver2 12. s40 w2 13. s40 w3 14. s50 w0 15. s50 w1 ver1 16. s50 w1 ver2 17. s50 w1 ver3 18. s50 w2 19. s50 w3 ver1 20. s50 w3 ver2 21. s60 w0 22. s60 w1 ver1 23. s60 w1 ver2 24. s60 w2 25. s60 w3 26. s70 w0 ver1 27. s70 w0 ver2 28. s70 w0 ver3 29. s70 w1 ver1 30. s70 w1 ver2 31. s70 w2 ver1 32. s70 w2 ver2 33. s70 w3 ver1 34. s70 w3 ver2 35. s80 w0 36. s80 w1 37. s80 w2 ver1 38. s80 w2 ver2 39. s80 w3 ver1 40. s80 w3 ver2 41. s80 w3 ver3 42. s90 w0 ver1 16
43. s90 w0 ver2 44. s90 w1 ver1 45. s90 w1 ver2 46. s90 w2 ver1 47. s90 w2 ver2 48. s90 w2 ver3 49. s90 w3 ver1 50. s90 w3 ver2 51. s100 w0 ver1 52. s100 w0 ver2 53. s100 w0 ver3 54. s100 w0 ver4 55. s100 w1 ver1 56. s100 w1 ver2 57. s100 w2 ver1 58. s100 w2 ver2 59. s100 w2 ver3 60. s100 w3 ver1 61. s100 w3 ver2 62. s110 w0 63. s110 w1 ver1 64. s110 w1 ver2 65. s110 w2 ver1 66. s110 w2 ver2 67. s110 w3 68. s120 w0 69. s120 w1 70. s120 w2 6.1.2.3 Uncontrolled Noise Recordings Approximately 20 minutes of audio content, recorded under conditions that are not designated with respect to the car speed or the window aperture, is available. This content is available in the form of 8-channel .wav files of various durations, with the shortest duration equal to 5 s. For this particular car and microphone arrangement there are 128 such audio recordings, labeled as unc xxx, where xxx is within the range from 001 to 128. The audio recordings span an average sound level from 61.4 dBA to 82.3 dBA and are sorted in ascending order with respect to the sound level, so that unc 001 corresponds to the quietest audio recording and unc 128 to the loudest one. 6.1.2.4 Ventilation Recordings 1. v1 w0 2. v1 w1 3. v1 w2 4. v1 w3 5. v2 w0 6. v2 w1 7. v2 w2 8. v2 w3 9. v3 w0 10. v3 w1 17
11. v3 w2 12. v3 w3 6.1.2.5 Car Audio Impulse responses from the car audio were obtained for all four windows apertures. 1. w0 2. w1 3. w2 4. w3 6.1.2.6 Steering Vector Angles for the different locations of passengers required to construct the steering vector are provided Figure 5. Figure 5: Angles for the different locations of passengers. 18
6.2 Smart forfour 6.2.1 Distributed Microphones The locations of the microphone array inside the car and numbering of the microphones can be seen in Figure 6. Figure 6: Locations and numbering of the distributed microphones inside the cabin. Six microphones are located in the front and two in the back. 6.2.1.1 Impulse Response Recordings 1. d48 w0 2. d48 w1 3. d48 w2 4. d58 w0 5. d58 w1 6. d58 w2 7. pf w0 8. pf w1 9. pf w2 10. prl10r w0 11. prl10r w1 12. prl10r w2 13. prl w0 14. prl w1 15. prl w2 16. prm10d w0 17. prm10d w1 19
18. prm10d w2 19. prm w0 20. prm w1 21. prm w2 22. prr10l w0 23. prr10l w1 24. prr10l w2 25. prr w0 26. prr w1 27. prr w2 6.2.1.2 Controlled Noise Recordings 1. s0 w09 2. s0 w1 3. s0 w2 4. s30 w0 coarse10 5. s30 w1 coarse 6. s30 w2 coarse 7. s40 w0 8. s40 w1 ver1 9. s40 w1 ver2 10. s40 w2 11. s50 w0 ver1 12. s50 w0 ver2 13. s50 w1 ver1 14. s50 w1 ver2 15. s50 w2 ver1 16. s50 w2 ver2 17. s60 w0 ver1 18. s60 w0 ver2 19. s60 w0 ver3 20. s60 w1 ver1 21. s60 w1 ver2 22. s60 w2 ver1 23. s60 w2 ver2 24. s70 w0 ver1 25. s70 w0 ver2 26. s70 w1 ver1 27. s70 w1 ver2 28. s70 w2 ver1 29. s70 w2 ver2 30. s80 w0 ver1 31. s80 w0 ver2 32. s80 w1 ver1 33. s80 w1 ver2 34. s80 w2 ver1 35. s80 w2 ver2 36. s90 w0 ver1 37. s90 w0 ver2 38. s90 w0 ver3 39. s90 w1 ver1 40. s90 w1 ver2 41. s90 w2 ver1 9s0: Idle conditions, stopped car. 10coarse: Coarse road surface. 20
42. s90 w2 ver2 43. s90 w2 ver3 44. s100 w0 ver1 45. s100 w0 ver2 46. s100 w1 ver1 47. s100 w1 ver2 48. s100 w2 ver1 49. s100 w2 ver2 50. s100 w2 ver3 51. s100 w2 ver4 52. s100 w2 ver5 53. s110 w0 ver1 54. s110 w0 ver2 55. s110 w1 ver1 56. s110 w1 ver2 57. s110 w1 ver3 58. s110 w2 59. s120 w0 60. s120 w1 61. s120 w2 6.2.1.3 Ventilation Recordings 1. v1 w0 ver1 2. v1 w0 ver2 3. v1 w1 4. v1 w2 ver1 5. v1 w2 ver2 6. v2 w0 7. v2 w1 ver1 8. v2 w1 ver2 9. v2 w2 ver1 10. v2 w2 ver2 11. v3 w0 12. v3 w1 13. v3 w2 ver1 14. v3 w2 ver2 6.2.1.4 Uncontrolled Noise Recordings Approximately 20 minutes of audio content, recorded under conditions that are not designated with respect to the car speed or the window aperture, is available. This content is available in the form of 8-channel .wav files of various durations, with the shortest duration equal to 5 s. For this particular car and microphone arrangement there are 118 such audio recordings, labeled as unc xxx, where xxx is within the range from 001 to 118. The audio recordings span an average sound level from 59.5 dBA to 81.3 dBA and are sorted in ascending order with respect to the sound level, so that unc 001 corresponds to the quietest audio recording and unc 118 to the loudest one. 21
6.2.1.5 Car Audio Impulse responses from the car audio equipment were obtained for all three window apertures, state 0, 1 and 2. 1. w0 2. w1 3. w2 6.2.2 Microphone Array The locations of the microphone array inside the car and numbering of the microphones can be seen in Figure 7. Figure 7: Location of the microphone array inside the car and numbering of the microphones. 6.2.2.1 Impulse Response Recordings 1. d w0 2. d w1 3. d w2 4. pf w0 5. pf w1 6. pf w2 7. pfr w0 8. pfr w1 9. pfr w2 10. prl w0 22
11. prl w1 12. prl w2 13. prm10l w011 14. prm10l w1 15. prm10l w2 16. prm10r w012 17. prm10r w1 18. prm10r w2 19. prm w0 20. prm w1 21. prm w2 22. prr w0 23. prr w1 24. prr w2 Figure 8: Locations prm10l (left) and prm10r (right). 6.2.2.2 Controlled Noise Recordings 1. s0 w013 2. s0 w1 3. s0 w2 4. s30 w0 coarse14 5. s30 w1 coarse 6. s30 w2 coarse 7. s40 w0 8. s40 w1 9. s40 w2 10. s50 w0 ver1 11. s50 w0 ver2 12. s50 w0 ver3 13. s50 w1 14. s50 w2 ver1 15. s50 w2 ver2 16. s60 w0 11prm10l: Rear middle passenger. Moved 10 cm towards the left (Figure 8). 12prm10r: Rear middle passenger. Moved 10 cm towards the right (Figure 8). 13s0: Idle conditions, stopped car. 14coarse: Coarse road surface. 23
17. s60 w1 ver1 18. s60 w1 ver2 19. s60 w2 20. s70 w0 ver1 21. s70 w0 ver2 22. s70 w1 ver1 23. s70 w1 ver2 24. s70 w2 25. s80 w0 ver1 26. s80 w0 ver2 27. s80 w1 28. s80 w2 ver1 29. s80 w2 ver2 30. s90 w0 ver1 31. s90 w0 ver2 32. s90 w1 ver1 33. s90 w1 ver2 34. s90 w2 ver1 35. s90 w2 ver2 36. s100 w0 37. s100 w1 ver1 38. s100 w1 ver2 39. s100 w2 40. s110 w0 41. s110 w1 42. s110 w2 43. s120 w1 6.2.2.3 Uncontrolled Noise Recordings Approximately 20 minutes of audio content, recorded under conditions that are not designated with respect to the car speed or the window aperture, is available. This content is available in the form of 8-channel .wav files of various durations, with the shortest duration equal to 5 s. For this particular car and microphone arrangement there are 111 such audio recordings, labelled as unc xxx, where xxx is within the range from 001 to 111. The audio recordings span an average sound level from 58.2 dBA to 78.8 dBA and are sorted in ascending order with respect to the sound level, so that unc 001 corresponds to the quietest audio recording and unc 111 to the loudest one. 6.2.2.4 Ventilation Recordings 1. v1 w0 2. v1 w1 3. v1 w2 4. v2 w0 5. v2 w1 6. v2 w2 7. v3 w0 8. v3 w1 9. v3 w2 24
6.2.2.5 Car Audio Impulse responses from the car audio equipment were obtained for all three window apertures, state 0, 1 and 2. 1. w0 2. w1 3. w2 6.2.2.6 Steering Vector Angles for the different passenger locations required for constructing the steering vector are provided in Figure 9. Figure 9: Angles for the different locations of passengers. 25
6.4.1.3 Uncontrolled Noise Recordings Approximately 20 minutes of audio content, recorded under conditions that are not designated with respect to the car speed or the window aperture, is available. This content is available in the form of 8-channel .wav files of various durations, with the shortest duration equal to 5 s. For this particular car and microphone arrangement there are 116 such audio recordings, labelled as unc xxx, where xxx is within the range from 001 to 116. The audio recordings span an average sound level from 56.6 dBA to 80.2 dBA and are sorted in ascending order with respect to the sound level, so that unc 001 corresponds to the quietest audio recording and unc 116 to the loudest one. 6.4.1.4 Ventilation Recordings 1. v1 w0 2. v1 w2 3. v2 w0 4. v2 w2 5. v3 w0 6. v3 w2 6.4.1.5 Car Audio Impulse responses from the car audio were obtained for all four window apertures in the distributed microphone setup. 1. w0 2. w1 3. w2 4. w3 32
6.4.2 Microphone Array The location of the microphone array inside the car and numbering of the microphones can be seen in Figure 15. Figure 15: Position of the microphone array inside the car cabin and numbering of the microphones. 6.4.2.1 Impulse Response Recordings 1. d55 w027 2. d55 w1 3. d55 w2 4. d55 w3 5. d63 w028 6. d63 w1 7. d63 w2 8. d63 w3 9. pf w0 10. pf w1 11. pf w2 12. pf w3 13. prl w0 14. prl w1 15. prl w2 16. prl w3 17. prm w0 18. prm w1 27d55: Driver position. 55cm distance from wheel. 28d63: Driver position. 63cm distance from wheel. 33
19. prm w2 20. prm w3 21. prr w0 22. prr w1 23. prr w2 24. prr w3 6.4.2.2 Controlled Noise Recordings 1. s0 w029 2. s0 w1 3. s0 w2 4. s0 w3 5. s30 w0 coarse30 6. s30 w1 coarse 7. s30 w2 coarse 8. s30 w3 coarse 9. s40 w0 ver1 10. s40 w0 ver2 11. s40 w1 ver1 12. s40 w1 ver2 13. s40 w2 ver1 14. s40 w2 ver2 15. s40 w3 ver1 16. s40 w3 ver2 17. s40 w3 ver3 18. s50 w0 19. s50 w1 20. s50 w2 ver1 21. s50 w2 ver2 22. s50 w2 ver3 23. s50 w3 24. s60 w0 ver1 25. s60 w0 ver2 26. s60 w0 ver3 27. s60 w0 ver4 28. s60 w1 ver1 29. s60 w1 ver2 30. s60 w2 31. s60 w3 32. s70 w0 33. s70 w1 34. s70 w2 ver1 35. s70 w2 ver2 36. s70 w3 ver1 37. s70 w3 ver2 38. s80 w0 39. s80 w1 40. s80 w2 ver1 41. s80 w2 ver2 42. s80 w3 43. s90 w0 44. s90 w1 ver1 45. s90 w1 ver2 29s0: Idle conditions, stopped car. 30coarse: Coarse road surface. 34
46. s90 w2 ver1 47. s90 w2 ver2 48. s90 w3 ver1 49. s90 w3 ver2 50. s90 w3 ver3 51. s100 w0 ver1 52. s100 w0 ver2 53. s100 w1 ver1 54. s100 w1 ver2 55. s100 w2 ver1 56. s100 w2 ver2 57. s100 w3 ver1 58. s100 w3 ver2 59. s110 w0 60. s110 w1 ver1 61. s110 w1 ver2 62. s110 w1 ver3 63. s110 w2 ver1 64. s110 w2 ver2 65. s110 w3 ver1 66. s110 w3 ver2 6.4.2.3 Uncontrolled Noise Recordings Approximately 20 minutes of audio content, recorded under conditions that are not designated with respect to the car speed or the window aperture, is available. This content is available in the form of 8-channel .wav files of various durations, with the shortest duration equal to 5 s. For this particular car and microphone arrangement there are 112 such audio recordings, labelled as unc xxx, where xxx is within the range from 001 to 112. The audio recordings span an average sound level from 61.4 dBA to 82.2 dBA and are sorted in ascending order with respect to the sound level, so that unc 001 corresponds to the quietest audio recording and unc 112 to the loudest one. 6.4.2.4 Ventilation Recordings 1. v1 w0 2. v1 w1 3. v1 w2 4. v1 w3 5. v2 w0 6. v2 w1 7. v2 w2 8. v2 w3 9. v3 w0 10. v3 w1 11. v3 w2 12. v3 w3 35
6.4.2.5 Car Audio Impulse responses from the car audio were obtained for all four windows apertures. 1. w0 2. w1 3. w2 4. w3 6.4.2.6 Steering Vector Angles for the different locations of passengers required to construct the steering vector can be seen in Figure 16. Figure 16: Angles for the different locations of passengers. 36