Title: Virtual audiovisual everyday-life environments for hearing aid research. Authors: Maartje M. E. Hendrikse, Gerard Llorach, Volker Hohmann and Giso Grimm. Version 3: 23/03/2020 In version 3 the documentation of the channel order for the loudspeaker recordings was corrected. Furthermore, the calibration levels were added that are needed for playback of the loudspeaker recordings. The virtual audiovisual environments in this database are published in conjunction with the paper Hendrikse et al. (2019). These virtual audiovisual environments were designed based on a number of situations with high importance and/or occurrence in everyday-life of young and elderly normal-hearing and hearing-impaired persons. They were designed to have a big range of different target sources (one or multiple live speakers or loudspeakers, close or further away) and distractors (concurrent speakers or other sounds, close or further away, moving or stationary). Their purpose was to enable the measurement of realistic head-, eyeand torso-movement and EEG during listening tasks in everyday-life situations. A database of the measured movement behavior and EEG of 21 young normalhearing and 19 elderly normal-hearing subjects is also available (see Related identifiers). The methods and an analysis of the subjective experience of these environments are published in Hendrikse et al. (2019). The environments include: a living room (watching the news on TV), a lecture hall (listening to a lecture), a cafeteria (listening to a 4-person conversation), a train station (listening to the announcements) and a street (listening to a 4person conversation at a bus stop or waiting for someone at a cross-section). The environments contain speech material in German. 3D videos of the environments can also be found on Youtube: https://www.youtube.com/playlist?list=PL-v3EIoK6ZarlOvYqCu7LpDext3Iafwtr General Information: This version contains for all environments: Audio (as WAV-files): o Target signal: binaural recording, Ambisonics recording, loudspeaker signals for our setup o Noise: binaural recording, Ambisonics recording, loudspeaker signals for our setup Video (as MP4-files): video of the cylindrical screen warped to panoramic view, including binaural Target+Noise sound for easy headphone listening. Detailed description of the environments, including timeline and licenses for all Blender objects and sound files that were used. The audio is provided as loudspeaker signals matching our setup for reproduction. The positions of the loudspeakers for the provided loudspeaker signals are in Table 1. The calibration levels that should be used for playback of the loudspeaker signals are provided in the file ‘audio_ldspkREC_caliblevels.mat’. The calibration level is the SPL level which a full-scale rectangular signal would achieve. This means that a full-scale sine wave of a sound file with a calibration level of 110 dB SPL would result in 117 dB SPL. In TASCAR, this is the ‘caliblevel’ value that should be provided with the ‘sound’ object, see the manual for more detailed info (http://tascar.org/manual.pdf). Moreover, binaural recordings (for playback over headphones) and 3rd order Ambisonics recordings are provided. The binaural signals were generated by convolution of the loudspeaker signals with HRIRs of a head-and-torso simulator Brüel & Kjær Type 4128C with artificial ears Brüel & Kjær Type 4158C (right) and Type 4159C (left) including preamplifiers Type 2669 (see Kayser, 2009, for more details). The HRIRs were measured in the reproduction room via the same loudspeakers. Files were stored in a standard RIFF-WAVE container in PCM format with two channels (left, right) and IEEE Float sample format. The 3rd order ambisonics files follow the Ambisonics Channel Number with Furse-Malham normalization (Chapman, 2009). They are packed into a standard RIFF-WAVE container in PCM format with 16 channels and IEEE Float sample format. The Ambisonics transmission format is independent of the loudspeaker layout, and can be converted to loudspeaker signals for any loudspeaker array using a custom decoder matched to that loudspeaker array. To create a decoder for a loudspeaker array, for example the Ambisonic Decoder Toolbox (https://bitbucket.org/ambidecodertoolbox/adt.git) could be used as described in Heller & Lee (2012). The audio is provided as Target and Noise separately, to allow the calculation of, e.g., signal-to-noise ratio. The videos show the projection onto a cylindrical screen and also include the binaural audio recordings for easy headphone listening.
Table 1: Loudspeakers positions for the provided loudspeaker signals. Channel Azimuth (deg) Elevation (deg) Distance to center (m) 1 11.25 0 1.76 2 33.75 0 1.76 3 56.25 0 1.76 4 78.75 0 1.76 5 101.25 0 1.76 6 123.75 0 1.76 7 146.25 0 1.76 8 168.75 0 1.76 9 191.25 0 1.76 10 213.75 0 1.76 11 236.25 0 1.76 12 258.75 0 1.76 13 281.25 0 1.76 14 303.75 0 1.76 15 326.25 0 1.76 16 348.75 0 1.76 17 (subwoofer) 55 -38 2.1 18 (subwoofer) 125 -38 2.1 19 (subwoofer) 235 -38 2.1 20 (subwoofer) 305 -38 2.1 21 150 -38 2.1 22 90 -38 2.1 23 30 -38 2.1 24 330 -38 2.1 25 270 -38 2.1 26 210 -38 2.1 27 240 34 1.2 28 300 34 1.2 29 0 34 1.2 30 60 34 1.2 31 120 34 1.2 32 180 34 1.2 The acoustic environment was made with the acoustic modeling toolbox TASCAR (Grimm et al., 2015). The visual environment was made with the 3D modeling software Blender (https://www.blender.org/). For the environments, sound files and Blender objects were used that were created by someone else and have a license. Licenses and attributions are listed below in the detailed description. The final rendered video and audio files are distributed with CC BY-NC-SA 3.0 license as constrained by the licensed objects. For ease of use and to prevent license violations, we provide the rendered audio and video files here, as they do not require detailed knowledge of the software TASCAR and Blender and do not allow access to the individual licensed objects. With this, reproduction of the experiment described in Hendrikse et al. (2019) is possible, except for the movement parallax simulation (changing the camera and acoustic receiver with the head position of the subject). For further information, please contact Prof. Dr. Volker Hohmann (
[email protected]).
Detailed description of environments: Below is described per environment what is happening and which sound sources and licensed objects are used. All diffuse sound sources are 1st order Ambisonics recordings unless described otherwise. For details of the rendering methods see Grimm et al. (2015). Living room In the living room, the listener is sitting on the sofa and is listening to the news on TV. There is one distractive female speaker making comments about the news and a person eating chips. Environmental noises include a fireplace and some noises from the kitchen. The levels and positions of the sound sources over time are shown in Figure 1. Figure 1: Positions in azimuth relative to listener position and levels (1 dB/degree) at listener position of all point noise sources in the living room environment. Target sound source: TV: recording of a female speaker reading Text 1 from “Texte für die neurologische Rehabilitation” by Bernhard Riedel, published by NAT-Verlag Hofheim, 2014 (https://www.natverlag.de/programm/textverarbeitung/texte/). Sound file published with permission of NAT-Verlag. Recordings of a female speaker reading the text were made by Theresa Nüsse. A left and right TV loudspeaker were modeled, including loudspeaker distortion. Noise sound sources: Female commenting on news: own recordings, Joanna Luberadzka, CC BY-SA 3.0 Fireplace: own recording, Giso Grimm, CC BY-NC-SA 4.0 Eating chips: agalaxy / freesound.org (194463__agaxly__eating-chips.wav), CC 0 Fridge: ancorapazzo / freesound.org (181616__ancorapazzo__fridge-hum.wav), CC BY 3.0 Dishwasher: cmusounddesign / freesound.org (71953__cmusounddesign__bm-dishwasher.wav), CC BY 3.0 Watercooker: own recording, Maartje Hendrikse, CC 0
Licensed Blender objects: Chair: captain_neodym / free3d.com (ncxk6dgdg83k-Sessel.rar), Personal Use License Sofa: benzin / free3d.com (w317ai1bvxts-sofa.zip), Personal Use License Stanford Bunny: Stanford University Computer Graphics Laboratory / http://graphics.stanford.edu/data/3Dscanrep/ Coffeetable: 3dhaupt / free3d.com (8h4bmee1bmo0-Wood_Table_with_glasplatte.zip), Personal Use License Coffee cups: cesarmrb2002 / free3d.com (hxpf9urgo0e8-TAZA.blend.rar), Personal Use License Table & chairs: pcpunch59 / free3d.com (nqh8isb919mo-TableChaisse.zip), Personal Use License TV: 3dhaupt / free3d.com (ku9p58i3hmo0-Samsung_Smart_TV_55_Zoll.zip), Personal Use License Lantern: banasiewicz / free3d.com (rpiw33woh8n4-lampion.rar), Personal Use License Plant: 3dhaupt / free3d.com (gz8we4f4o7b4-Indoor_plant_3.zip), Personal Use License Poster umbrella couple: Emanuel M. Ologeanu / etsy.com (https://i.etsystatic.com/6182845/r/il/a9d149/584623322/il_570xN.584623322_59w9.jpg) Poster Hamburg: unknown Lecture hall In the lecture hall, the listener is sitting in the audience and is listening to the lecture about a toolbox for acoustic scene creation and rendering (TASCAR). There are some noises from the audience (sniffing, coughing, sneezing, sighing, turning pages, writing) and at some point there is a paper airplane flying by, making a sound while landing. The levels and positions of the sound sources over time are shown in Figure 2. Figure 2: Positions in azimuth relative to listener position and levels (1 dB/degree) at listener position of all point noise sources in the lecture hall environment.
Target sound source: Lecturer: own recording, Giso Grimm, CC BY-NC-SA 3.0. Sound sources simulated as direct sound from position of lecturer and two loudspeakers (with some loudspeaker distortion) on the sides. Noise sound sources: Gertrude: Sigh: splicesound / freesound.org (218309__splicesound__girl-female-inhale-exhale-sighbreathing.wav), CC 0 Pencil writing: deleted-user-3720673 / freesound.org (209666__deleted-user-3720673__pencilwriting.wav), CC BY 3.0 Turning pages: sinjohnt / freesound.org (325603_sinjohnt_turning-pages.wav), CC BY-NC 3.0 Max: Coughing: deleted-user-7146007 / freesound.org (383783_deleted-user-7146007_man-coughingcough.wav), CC 0 Clearing throat: shorzie / freesound.org (407722_shorzie_male-throat-clearing-various.wav), CC BYNC 3.0 Turning pages: sinjohnt / freesound.org (325603_sinjohnt_turning-pages.wav), CC BY-NC 3.0 Julie: Coughing: loisgwenllian / freesound.org (142604__loisgwenllian__cough1.wav), CC 0 Clearing-throat: aryanotstark / freesound.org (407616_aryanotstark_clearing-throat.wav), CC BY-NC 3.0 George: Sniffing: otakua / freesound.org (219907_otakua_snif03.wav), CC BY 3.0 Sighing: stubb / freesound.org (389651_stubb_breathe-2.wav), CC 0 Charlie: Sneezing: aderumoro / freesound.org (238432_aderumoro_sneeze-2-young-lady.wav), CC BY 3.0 LuiYang: Coughing: joedeshon / freesound.org (266019_joedeshon_double-cough-01.wav), CC BY 3.0 John: Sneezing: inspectorj / freesound.org (352177_inspectorj_sneeze-single-a.wav), CC BY 3.0 Coughing: jacklmurr27 / freesound.org (393733_jacklmurr27_natural-male-cough.wav), CC 0 Sniffing: mihirfreesound / freesound.org (410979_mihirfreesound_male-sniff.wav), CC BY-NC 3.0 Paper airplane: own recording, Maartje Hendrikse, CC 0
Cafeteria In the cafeteria, the listener is sitting at a table and is listening to the conversation between 2 female and 2 male speakers at the table. The cafeteria was tested with two different conditions: listening only (no extra task) and dual task (listening to the conversation while doing a hand-eye coordination task to simulate eating). Therefore two versions with different target conversations are available. The noise consists of a background dialog and background conversation, laughter, music (diffuse) and diffuse background noise with babble and cutlery noise. The positions and levels over time of all point sources relative to the listener position are plotted in Figure 3 and 4. Figure 3: Positions in azimuth relative to listener position and levels (1 dB/degree) at listener position of all point noise sources in the cafeteria_listeningonly environment (Story 1). Figure 4: Positions in azimuth relative to listener position and levels (1 dB/degree) at listener position of all point noise sources in the cafeteria_dualtask environment (Story 5).
Target sound source: Conversations: Maartje Hendrikse; Giso Grimm; Gerard Llorach; Volker Hohmann / http://doi.org/10.5281/zenodo.1257333 (Story 1&5), CC BY-NC 3.0 Noise sound sources: Music (diffuse): "Lighthouse" by Dokapi from http://opsound.org/artist/dokapi/ (dokapi-continental_drift01-lighthouse_44.wav), CC BY-SA 3.0. In the simulation, the music was played over 6 loudspeakers located in different parts of the cafeteria and only the 1st order reflections (no direct sound) were allowed to make it diffuse. Background conversation: own recordings, Merle Gerken, CC BY-NC 4.0 Background dialog: Institut fuer Phonetik und digitale Sprachverarbeitung Christian-Albrechts-Universitaet zu Kiel / Kiel Corpus of Spoken German, commercial license Diffuse background noise mensa: own recording, Giso Grimm, CC BY-SA 3.0 Laughter: lmbubec / freesound.org (119450__lmbubec__girl-laugh.wav), CC 0 sandyrb / freesound.org (92408__sandyrb__close-female-laugh-049.wav), CC BY 3.0 Licensed Blender objects: Meat: kascad3 / free3d.com (jiq2nwj8gtmo-Raw_meat.zip), Personal Use License Plants and trees: unknown Train station In the train station, the listener is standing on the platform and is listening to the announcements. The noise consists of a background conversation, trains arriving and departing, people with trolleys walking past, a ticket validation machine and diffuse background noise recorded in a train station. The positions and levels over time of all point sources relative to the listener position are plotted in Figure 5. Figure 5: Positions in azimuth relative to listener position and levels (1 dB/degree) at listener position of all point noise sources in the train station environment.
Target sound source: Train announcements: own recordings, Matthias Vormann, CC BY-NC-SA 3.0 Noise sound sources: Train noises: own recordings, Giso Grimm, CC 0 Diffuse background noise: own recording, Giso Grimm, CC BY-NC-SA 4.0 Trolleys: own recordings, Maartje Hendrikse, CC 0 Background conversation: Maartje Hendrikse; Giso Grimm; Gerard Llorach; Volker Hohmann / http://doi.org/10.5281/zenodo.1257333 (Story 7), CC BY-NC 3.0 Ticket validation machine: own recording, Giso Grimm, CC 0 Street_active The listener is standing at a bus stop and is listening to a conversation between 2 female speakers and 2 male speakers. On the sidewalk there is a bicycle driving by and a mother with pram walking by singing lullabies. On the street there are cars, a bus, a truck and a rescue car driving by. Noises from a nearby school can be heard. A train is driving by in the distance. Diffuse background noise consists of birds singing and far-away traffic noise. The positions and levels over time of all point sources relative to the listener position are plotted in Figure 6. Figure 6: Positions in azimuth relative to listener position and levels (1 dB/degree) at listener position of all point noise sources in the street_active environment. Target sound source: Conversations: Maartje Hendrikse; Giso Grimm; Gerard Llorach; Volker Hohmann / http://doi.org/10.5281/zenodo.1257333 (Story 6), CC BY-NC 3.0 Noise sound sources: Background dialog: Institut fuer Phonetik und digitale Sprachverarbeitung Christian-Albrechts-Universitaet zu Kiel / Kiel Corpus of Spoken German, commercial license Car noises: