Full text
ii DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Licença concedida aos utilizadores deste trabalho Atribuição-NãoComercial-SemDerivações CC BY-NC-ND https://creativecommons.org/licenses/by-nc-nd/4.0/
iii Agradecimentos A escrita de uma Tese de Doutoramento é um processo curioso. Quase sempre penoso, frequentemente frustrante, mas que perto do fim se reveste de um prazer difícil de explicar. Celebrar a ultrapassagem dos muitos momentos difíceis deste processo, só faz sentido se puder ser partilhado com as pessoas da nossa vida. Assim, sem que a ordenação seja sinónimo de importância, quero agradecer aos meus colegas e amigos de laboratório pelo apoio diário durante os últimos anos. Foram eles que viveram de perto este processo, vindo muitas vezes deles o apoio e a crítica construtiva que se revela fundamental. Ao Bruno, à Liliana, e à Sandra agradecer tudo aquilo que me ensinaram e o espetacular acolhimento de um miúdo lhes foi parar às mãos à uns 10 nos atrás. Ao Emanuel, aos Joões (Pedro e Lamas), e à Joana companheiros de luta dos últimos anos, a paciência, a ajuda, e o incrível profissionalismo de um grupo de jovens investigadores que foram capazes de criar um espaço de trabalho ao qual me orgulho de pertencer. To Paolo, Glenn, Yi Zhang, and Paul Jones I want to thank you the opportunity you gave me to work in a wonderful place and the good moments we enjoyed in Maryland. Aos colegas do Centro de Computação Gráfica, instituição que me acolheu no início desta tese como bolseiro e que me permitiu terminá-la já como membro integrado da sua equipa de investigação. Ao Pierre, ao Tiago, ao Diogo, ao Pedro e a Nocas, a amizade verdadeira que foi sempre reafirmada a cada oportunidade que me deram para escapar do trabalho e entrar num mundo sempre muito cómico e de uma loucura saudável. Uma das melhores coisas de terminar uma tese, é poder deixar de ter desculpas para recusar os convites dos amigos. Aos meus orientadores, Professor José Creissac Campos e Professor Jorge Santos, por terem sempre acreditado em mim, ignorando sucessivos adiamentos para rematar isto e estando sempre presentes quando a necessidade de orientação surgia. Não podia ter pedido melhor orientação. Se um dia estiver no vosso papel, certamente imitarei a vossa postura esperando que compreendam que a imitação é a melhor forma de elogio. À minha família, que mesmo não sabendo muito bem o que faço acreditam que é a coisa mais importante do mundo. Não é essa a melhor demonstração de carinho que alguém pode receber? Em tudo o que faço, o vosso apoio é fundamental. Por último, à pessoa a quem dedico todo este trabalho esperando que a sua conclusão me permita mais facilmente retribuir todo o amor demonstrado ao longo destes últimos seis anos. À Mariana, o meu mais profundo obrigado e a promessa de recuperar para nós o tempo que mereces.
iv STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. University of Minho, 29/04/2019 Full Name: Carlos César Loureiro Silva Signature:
v Using Predictive and Descriptive Cognitive Models for Evaluation of Interactive Computing Systems The advent of Interactive Computing Systems (ICSs) has released the great potential that computers have for influencing everyday human activity. The establishment of Human Computer Interaction (HCI) as an academic discipline has helped develop a sense that serious consideration and understanding of how humans interact with computers is a pressing necessity for both the computer scientist and the developer. One of the traditional ways of using HCI research to foster the development of better technology is by resorting to the use of Interaction Models – or simplifications of the reality capable of being operationalized in such a way that provide valuable information in order to adapt a given system to its user. Two types of models, often respectively referred as Predictive Cognitive Models (PCMs) and Descriptive Cognitive Models (DCMs), have emerged from the HCI research in the last decades, and they have been used with the primary goal of formalizing empirical insights and foreseeing design and interaction problems. Throughout this thesis, we will review the prolific history of user-research in HCI while showing contributions of both interaction-modeling approaches for the development of innovative and safer user interface technology. As a way of proving the benefit these approaches can bring to the development of today’s ICSs, we will demonstrate their application to the evaluation and development of two types of ICSs that are at the forefront of current HCI challenges: Immersive Virtual Environments (IVEs), and Safety-critical Interactive Medical Devices . Although distinct in principle, these ICSs have some communalities that render them interesting use-cases in the scope of this thesis. Both systems are becoming increasingly widespread, and due to their particularities, both present new interaction challenges that are absent from more traditional HCI use-cases, such as the personal computer. The main outputs of this thesis are thus twofold: the outline of new PCMs and design guidelines that can help the development of increasingly immersive virtual environments; and the outline of a new DCM of use-error with safety-critical devices in addition to a set of design guidelines that can help mitigate these errors and foster the development of increasingly safer medical devices. Keywords: Descriptive Cognitive Models; Immersive Virtual Environments; Interactive Computing Systems; Predictive Cognitive Models; Safety-Critical Medical Devices.
vi Utilização de Modelos Cognitivos Preditivos e Descritivos para Avaliação de Sistemas Interativos de Computação O surgimento de Sistemas Interativos de Computação (SIC) potenciou o impacto e a influência que os computadores exercem sobre a atividade quotidiana dos seres humanos. A consolidação da Interação Humano-Computador (IHC) como uma disciplina académica, ajudou a estabelecer a ideia de que a séria consideração e o estudo aprofundado da forma como os utilizadores interagem com os computadores, é uma necessidade premente tanto para o cientista da computação, como para o técnico que desenha e desenvolve novos SIC. Uma das formas tradicionais de utilizar os métodos de IHC para fomentar o desenvolvimento de melhor tecnologia, é recorrendo ao uso de Modelos de Interação – simplificações da realidade capaz de serem operacionalizadas de modo a fornecer informações valiosas para adaptar um sistema ao seu utilizador. Da investigação em IHC emergiram dois tipos de modelos, frequentemente referidos como Modelos Cognitivos Preditivos (MCP) e Modelos Cognitivos Descritivos (MCD), usados com o objetivo principal de formalizar resultados empíricos e prever problemas de design da interação. Ao longo desta tese, iremos rever a prolifica história da investigação com utilizadores em IHC e demonstrar como ambas as abordagens de modelação da interação apoiaram o desenvolvimento de tecnologia, segura e inovadora, de interface com o utilizador. De forma a demonstrar que estas abordagens podem trazer o mesmo benefício ao desenvolvimento dos SIC atuais, descreveremos a sua aplicação à avaliação e desenvolvimento de dois tipos de SIC, na vanguarda dos atuais desafios da IHC: Ambientes Imersivos e Virtuais (AIVs), e Dispositivos Médicos Interativos e de Segurança-Crítica . Embora distintos na sua gênese, estes SIC têm algumas semelhanças que os tornam casos de uso interessantes no âmbito desta tese. Ambos os sistemas estão cada vez mais difundidos e, devido às suas particularidades, apresentam novos desafios de interação que estão ausentes nos casos de uso tradicionais da IHC, como o computador pessoal. Desta forma, os resultados desta tese dividem-se em duas principais áreas: o desenvolvimento de novos MCPs e diretrizes de design para apoiar o desenvolvimento de ambientes virtuais cada vez mais imersivos; e o desenvolvimento de um novo MCD de erro de uso em dispositivos de segurança-crítica, além de um conjunto de diretrizes de design que podem ajudar a mitigar esses erros e promover o desenvolvimento de dispositivos médicos cada vez mais seguros. Palavras-Chave: Ambientes Imersivos Virtuais; Dispositivos Médicos Interativos e de Segurança-Crítica; Modelos Cognitivos Descritivos; Modelos Cognitivos Preditivos; Sistemas Interativos de Computação.
vii Table of Contents PART I General Introduction 1. THE REVOLUTION OF INTERACTION ....................................................................................... 18 1.1. At the Dawn of Computation ........................................................................................ 18 1.2. The Contribution of Human-Computer Interaction ......................................................... 20 1.3. Approach .................................................................................................................... 24 1.4. Thesis Overview ........................................................................................................... 27 PART II PCMs and Immersive Virtual Environments 2. THEORETICAL BACKGROUND ON IVEs AND PCMs ................................................................. 29 2.1. What are Immersive Virtual Environments? ................................................................... 29 2.2. What are the problems with IVEs and how can we solve them? ..................................... 32 2.3. On Understanding and Modeling Perception and Performance for the Development of IVEs …………………………………………………………………………………………………………………...36 2.4. What is Psychophysics and how can it change the way we interact with computers ....... 40 2.5. Human Predictive Cognitive Models ............................................................................. 42 2.6. The current state of IVEs technology ............................................................................ 65 3. EVALUATION OF HUMAN PERCEPTION AND PERFORMANCE AND DEVELOPMENT OF PREDICTIVE COGNITIVE MODELS ........................................................................................... 77 3.1. Audiovisual Perception of Synchrony ............................................................................ 78 3.1.1.On the Role of Temporal Coincidence ................................................................. 78 3.1.2.Assessing Audiovisual Synchrony in an IVE ......................................................... 83 3.2. Audiovisual Unity Assumption ...................................................................................... 99 3.2.1.On the Role of Spatial Coincidence ..................................................................... 99 3.2.2.Assessing the Spatial Limits of Unity Assumption .............................................. 101 3.3. Perception of Visual Distances in Virtual Environments ............................................... 111 3.3.1.On the Need to Access Basic Perceptual Phenomenon to Validate Simulators for Studying Complex Perception-Action Behavior ........................................................................ 111 3.3.2.Measuring Distance Perception: A Comparison Between Real-world and Simulated Scenarios ............................................................................................................................. 114 3.4. Spatial Discriminability of Auditory Virtual Stimuli ....................................................... 121 3.4.1.On the Rising of Commercial Applications for 3D Sound .................................... 121 3.4.2.Measuring Auditory Location of Virtual Stimuli as a Function of the Audio Device 123 3.5. Conclusion ................................................................................................................ 128
xiv Figure 63. Use error types in our taxonomy. Figure 64. Classifying use errors with our Taxonomy. Figure 65. Use error classification with Zhang et al.’s taxonomy Figure 66. Two examples of Use Errors and their classification using our taxonomy (left) and Zhang’s taxonomy (right). Figure 67. Screenshots of the main tools provided by the PVSio-web environment. Top image – Prototype builder; Bottom image - model Editor and a model snippet generated from emuchart editors. Figure 68. Game Scenario, consisting in three rooms. Figure 69. Game flowchart. Figure 70. Close-up of the Massimo Radical 7® Pulse Co-Oximeter® a monitoring device that keeps track of patients’ oxygen saturation. Figure 71. Screenshots from the game environment. Top-image is the control room; Left-botton image is the fMRI room; and Right-bottom image is the Operating Room. Figure 72. Top image, interface shown during the game in order to prompt the player to open the door. Down image, physical visualization of the near sensor in Blender Game. Figure 73. A simplified Game Logic example for an opening door action. Figure 74. Medical devices currently modeled and available in our SBME. The top-left one is the Massimo Radical 7® Pulse Co-Oximeter® a monitoring device that keeps track of patients oxygen saturation, pulse rate, and perforation index; the bottom-left one is the BBraun® Infusionmat IV Pump, a programmable syringe based infusion pump; the right one is the Alaris® GP Volumetric Pump an infusion pump suited for drugs therapy, blood transfusions, and parenteral feeding Figure 75. Use Error Taxonomy for interaction with medical devices, presented in Section 5.1.2. This taxonomy was systematically derived from two important DCMs, the Rasmussen’s Decision-Ladder Framework [180] and Reason’s Generic Error-Modelling System [199]. List of Tables Table 1. Different task conditions in [52] and respective predicted (using Equation 1) and observed timeto-completion. Table 2. Different task conditions in [57] and respective predicted and observed time to completion. Table 3. Times for operators in the KLM (adapted from [69]) and updated with [11] values for pointing. Table 4. Comparison of KLM predictions for the same editing task using mouse-based operations and keyboard-based operations.
xv Table 5. SOAs values for the difference presentation distances. Negative values indicate that sound step was presented before the visual feet touched the ground, and positive values indicate that sound was presented after the visual step. Table 6. Individual values of the PSS and WTI. The values are presented in ms and for the several distances in each stimulus condition. In the last column the equations and the values of adjustment for each of the linear functions fitted to the individual data are presented. The asterisks signal the participants that had some background knowledge about the thematic of the study. Table 7. Performance by condition and experimental phase Table 8. Data grouped by training listening device. Table 9. Summary of the distinctions between skill-based, rule-based and knowledge-based errors (adapted from [199]). Table 10. Taxonomy alignment. Table 11. A core set of interaction design principles from two usability engineering standards (ANSI/AAMI/IEC 62366:2009 and HE75:2009) and regulatory guidance documents. Table 12. Design principles suitable for mitigating use errors in the cognitive stages of the Decision-Ladder model. Table 13. Assessment: - Not satisfied; ⁓ - Partially satisfied; - Improved; - Satisfied. Table 14. Summary of the IVEs design guidelines outputted from the experiments in Chapter 3. Abbreviations Part I UI User Interfaces SRI Standford Research Institute NLS oN-Line System ARC Augmentation Research Center ACM Association for Computing Machinery ICS Interactive Computing System HCI Human-Computer Interaction PCM Predictive Cognitive Model DCM Descriptive Cognitive Models GUI Graphical User Interfaces EICS Engineering Interactive Computing Systems ICS Interactive Computing Systems IVE Immersive Virtual Environment SBME Simulation-Based Medical Education Part II HMD Head Mounted Display VR Virtual Reality AR Augmented Reality WTI Window of Temporal Integration RV Reality-Virtuality EWK Extent of World Knowledge AV Augmented Virtuality ID Index of Difficulty CRT Cathode Ray Tube CRT Choice Reaction Time KLM Keystroke-Level Model MHP Model Human Processor EPIC Executive Process-Interactive Control CPS Central Production System VIVED Virtual Visual Environment Display CGI Computer-Generated Images
xvi OLED Organic LED display FoV Field of View CAVE Cave Automatic Virtual Environment TORE The Open Reality Experience SPL Sound Pressure Level ITD Interaural Time Difference ILD Interaural Level Difference HRIR Head Related Impulse Response HRTF Head Related Transfer Function ISM Image Source Method RIR Room Impulse Responses FDA Food and Drugs Administration MAUDE Manufacturer and User Facility Device Experience database AECL Atomic Energy of Canada Limited ETCC East Texas Cancer Center CWA Cognitive Work Analysis GEMS Generic Error-Modelling System GOMS Goals, Operators, Methods, Selection WTI Window of Temporal Integration TOJ Temporal Order Judgment SOA Stimulus Onset Asynchrony PSS Point of Subjective Simultaneity PLW Point Light Walker VE Ventriloquism Effect POI Point of Interest JND Just Noticeable Difference UA Unity Assumption FMDT Frontal Matching Distance Task RWS Real-World Scenario IVEPH Photorealistic Immersive Virtual Environment IVENPH Non-photorealistic Immersive Virtual Environment Part III CPS Cyber-Physical Systems ICE Integrated Clinical Environment HFCF Human Factors Classification Framework CPOE Computerized Physician Order Entry GEMS Generic Error Modeling System SBME Simulation-Based Medical Education
Part I General Introduction
18 1. THE REVOLUTION OF INTERACTION There will always be plenty of things to compute in the detailed affairs of millions of people doing complicated things. Vannevar Bush, As We may Think 1.1. At the Dawn of Computation The term Computeture can be traced to the writings of the 1th century roman author Pliny the Elder when, in his Naturalis Historia, he describes how the breadth of Asia could be properly calculated – “… latitudo sane computetur ” – from the Ethiopic Sea to Alexandria on the Nile [1]. By the first half of the 20th century, the word computer had already been in intensive use, primarily to mean a human activity (in the 18th and 19th century), later to name a profession (in the early 20th century), and in the 40’s of the past century it started to be applied as a term to describe a new technology capable of performing calculations of a complexity and speed unachievable by its human counterpart. In the 50’s, computers where still conceptualized as fundamentally powerful calculators and as such user interfaces (UIs) where merely a way of conveying instructions towards computing a given result. The first computers UI were essentially a programing interface and, frequently, a poorly usable one (see [2] for a concise history on electronic analog computing). Nevertheless, in 1945 some computer scientists were already grasping the potential computers could have in influencing everyday human activity. In that year, Vannevar Bush, first dean of engineering at the MIT and one of the developers of the early analog computers, published a detailed description of the Memex 1 - a system for storing and retrieving scientific information. His description of this conceptual system placed a great focus on new mediums of user interaction such as screens, photography cameras to register and catalog new information, and speech-recognition systems for digital stenography (Figure 1). 1 The description of the Memex system appeared first in the Atlantic Monthly on July 1945. In the previous years, Vannevar Bush had been the science advisor for President F.D. Roosevelt, playing that role during World War II. After the war, the great investment in military research would have to be redirect to new endeavors and, in Bush’s opinion, a project to develop the Memex systems could be the ideal replacement for the Manhattan project, again joining the best minds of a generation only now for a peaceful purpose during a time of peace [4].
19 A B Figure 1. Panel A , rendition of the camera used by the scientist - Memex user – to capture and catalog new information. Panel B , a diagram of the Memex desktop-like system as described by Bush in his article entitled “As We May Think” [3]. Images from LIFE. Although Bush’s bold vision was certainly difficult to materialize at that time, it contributed to the idea of the computer as an enabler of information access, storing, and sharing amongst its users. Following Bush’s vision 2 , during the 1950’s, Douglas Engelbart started a major research project at the Stanford Research Institute that included the development of new user interfaces for computer systems, of which the computer mouse was one of the components. The original concept was a unified system, which for the first time had in consideration all the interaction process ranging from displaying information to the user through a digital medium, to taking into account users’ input and act upon that information (Figure 2). The final goal of the system, which became known as the oN-Line System (NLS), was to enable a wide variety of user interaction with digital information that, in Engelbart’s words, would contribute to “augmenting human intellect” . Suitably, Englerbart’s laboratory was named the Augmentation Research Center (ARC) and its vision was that of increasing the effectiveness of individuals through the use of computer digital systems in an easy and effective way. With a good measure of correct forecast, Engelbart argued that these systems could fundamentally change the way people collaborate, work, and ultimately think. 2 Engelbart’s reports of the NLS quoted extensively Bush’s article, “As We May Think”.
20 A B Figure 2. Panel A, shows one of the several developed desktop configurations for the NLS. All of these configurations included a CathodeRay Tube (CRT) display, a computer mouse, its complimentary chord set, and a keyboard. It is important to note how this arrangement is similar to today’s desktop computers. Panel B , shows what had become the most popular mock-up design for the NLS (in this image operated by Engelbart), the Hermann Miller design. Images from Stanford.edu. The NLS was presented in 1968 at the Association for Computing Machinery (ACM) Fall Joint Computer Conference in San Francisco, and Engelbart’s demonstration of the overall system became known for posterity as “the mother of all demos” 3 . In addition to new user interfaces such as the computer mouse and its associated keyset, Engelbart’s also showcased some early demonstrations of today’s wellknown technology such as video-conferencing, collaborative text editing, outlining tools, and hyperlinks [4]. In a way, Engelbart was the first to materialize successfully what we now know as an Interactive Computing System (ICS). Once serious consideration was taken on how these systems would be helping humans in everyday tasks, capturing and understanding how humans interact with computers became a pressing necessity for both computer scientists and developers. 1.2. The Contribution of Human-Computer Interaction In its official curricula, the ACM defines Human-Computer Interaction (HCI) as: “… a discipline concerned with the design, evaluation and implementation of interactive computing systems for human use and with the study of major phenomena surrounding them. ” (p.5) [5]. As we will see some of the seminal works in HCI have emerged as early as in the mid 50’s (see Section 2.5 and the works of Paul Fitts), nonetheless HCI as a research field became formalized in the early 1980s, initially as a specialty area in computer sciences but rapidly expanded as an interdisciplinary 3 The full version of the Demo can be viewed here: https://web.stanford.edu/dept/SUL/library/extra4/sloan/mousesite/1968Demo.html
21 field that incorporated diverse concepts and approaches from a collection of semi-distinct fields of research and practice in human-centered informatics. It is a matter of ongoing discussion if we should define HCI as an autonomous scientific discipline since it appears that some of its characteristics fail at popperian aspects of a scientific discipline definition, as it is the case of its multidisciplinary character, the lack of a clear methodology, and the weak separation between practice and research (see [6], [7] or [8] for a debate on this issue). Nevertheless, it is of common agreement that HCI provides, in the words of Yvonne Rogers [9], “ prescriptive knowledge ”, giving advice on how to design an ICS. In addition, as pointed out by Alan Dix, there is an increasing sense of community within practitioners and researchers in computer and cognitive sciences that pursue two main common goals [10]: 1Improve interactions between computers and their users by making them more useable and receptive to the users’ needs ( primary goal ); 2Design systems that reduce the barrier between the users’ cognitive model of what they want to accomplish and the computer’s understanding of the users’ task ( long-term goal ). In this sense, we can classify HCI as an academic discipline, despite knowing that HCI implies sometimes not just scientific, but also engineering and even craft work. Interestingly, the aspects that render HCI a peculiar epistemological case are also the ones that give HCI its strength and its potential to continually find innovative computational solutions. If we take, for instance, the unclear separation between craft, engineering and science and the frequent iterative nature of the work in HCI, we surely can view these characteristics as making part of a rather empirical approach as opposed to the more scientific traditional ones. At the same time, however, it is not difficult to acknowledge the advantages of such approach when the goal is to come out with an exemplary product, in terms of its usability. Similarly, if we consider the multidisciplinary character of HCI, we can always argue that it is responsible for broadening too much its academic frontiers and give HCI that typical aspect of an embryonic scientific area – the difficulty in defining itself. Nevertheless, we should also acknowledge that the coexistence between such different areas within the HCI umbrella has been one of its main practical advantages, since it allows the interconnection of previously unrelated knowledge and contributes for the advance of each one of these disciplines (forming the HCI universe) individually. According to the ACM curricula on human-computer interaction [5], different disciplines, such as cognitive psychology, sociology, human factors, product design, and industrial engineering, function as supporting disciplines for computer
22 science in the scope of HCI, much as physics serves as a supporting discipline for civil engineering, or as mechanical engineering serve as a supporting discipline for robotics. The design of many modern computer applications inevitably requires the design of components of the system that allow for interaction between the system and its user, and these components typically represents more than half a system’s lines of code [5]. In the scope of computer science, HCI is responsible to guide us during the definition of the functionality a system will have, on how to bring it out to the user, and finally, on how to build and to test its design. One of the traditional ways of using HCI research to foster the development of better technology is by resorting to the use of Interaction Models . A model is a simplification of the reality capable of being operationalized in such a way that provides valuable information in order to adapt a given system to the subject being modeled. Their typology can be quite diverse and range from a formal language to a verbal-analytical description of behavior [11]. These two types of models are often respectively referred as predictive and descriptive models. HCI often uses both types of models with the goal of formalizing empirical insights and foreseeing design and interaction problems. Some HCI research methodologies were specifically developed to evaluate user’s interaction performance. The outputs of these research methodologies are often condensed in a Predictive Cognitive Model (PCM) based on the quantitative analysis of the user’s behavior. A PCM is a model that predicts how a given piece of information (input) will affect human behavior (output), and might be represented through an equation that predicts the outcome of a dependent variable based on the value of one or more independent variables – or predictors [11]. In the historical context of the HCI discipline, the independent variable has often been a given displayed piece of information or a specific design choice for an input device, while the dependent variable is often the time or speed when performing a task, its accuracy, and/or error rate. Nevertheless, any dimension of human behavior is capable of serving as the output of a PCM, as long as it is feasible to be measured precisely on a continuous or ratio scale [11]. The final goal of PCMs is the prediction of human performance, but in order to achieve their goal these models should give account of how information is processed and what is the role of human perception, memory, attention, and motor skill (of the user) in the outcome. Descriptive Cognitive Models (DCMs), on the other side, are applied more often during the conception or early development stages of a new interface, or during the study of use-error, as a “forensic tool”. As MacKenzie puts it, “they emerge from a process so natural it barely seems like modeling” (p. 233) [11]. The process Mackenzie is referring to is the partition, analysis, and description of the interaction process. This purely analytical process has since long been conducted by experts, analysts,
23 and developers in order to obtain insight on how the interaction flow, the design features, or the information content of an interface might lead to performance deficits, faulty interactions or use errors [12]. As we will see in Chapter 2 and Chapter 4, history is prolific in showing us contributions of both HCI interaction-modeling approaches for the development of innovative and safer user interface technology. User studies using both quantitative and qualitative methodologies have helped spark the development of highly diversified interactive computing systems. As pointed out by John Carroll, until the late 1970s, the interaction with computers was reserved to information technology professionals and a small group of dedicated hobbyists. In this way, the need for HCI was kept hidden under the existence of expert users only [7]. Nevertheless, this was dramatically changed around 1980, when the emergence of personal computers – and the advent of direct manipulation on graphical user interfaces (GUI) – turned HCI into an indispensable tool, once the use of computer systems as consumer goods became widespread (thus, fulfilling the realization of Engelbart’s vision) 4 . Today’s interactive computer systems encompass a great variety of platforms ranging from the single-user desktop computer to the use of ubiquitous systems in public spaces. This is evident in the ACM conference on Engineering Interactive Computing Systems (EICS), which covers topics of interest so diverse as: design and development of systems incorporating new interaction techniques and multimodal interaction; multi-device interaction ; mobile and pervasive systems ; safety-critical interactive devices; and immersive virtual environments (among others) . The last two ICSs of the list above are at the forefront of today’s HCI challenges 5 and will be the main use cases for the work developed in the scope of this thesis. Throughout the next chapters, we will be illustrating the contribution HCI can give to the evaluation and development of these two types of interactive computing systems: 1 Immersive Virtual Environments (IVEs), can be regarded as the epitome of interactive computer systems due to their ability to appeal to the different senses of the user during the interaction process. Nevertheless, knowledge about perception and performance – two fundamental steps 4 Around the time of Engelbart’s NLS, Ivan Sutherland was also providing a demonstration of direct manipulation interfaces in his 1963 PhD thesis on the Sketchpad, a system that supported the users’ manipulation of graphic objects using a light-pen that could interact with the image displayed on a CRT. These seminal works and the advent of the firsts CAD/CAM systems (the General Motors one, in 1963) led to an all new range of technology and interface ideas, and soon the first commercial systems to make extensive use of direct manipulation interfaces were on the market: the Xerox Star in 1981, the PERQ in 1981, the Apple Lisa in 1982, and the Apple Macintosh in 1984. 5 In addition to the EICS topics of interest, the top HCI conference, Computer Human Interaction (CHI), has a track dedicated to Interaction techniques, Devices and Modalities with mostly IVEs related topics such as: haptic and tangible interfaces, 3D interaction, augmented/mixed/virtual reality, wearable and on-body computing, sensors and sensing, displays and actuators, muscleand brain-computer interfaces, and auditory and speech interfaces. Another important conference in HCI, INTERACT, has a special track on safety/critical interactive systems with topics as automation, critical interactive systems, healthcare, human error, human work interaction design, safety, security, space, training, transportation.
30 Breaking down this definition, we can agree that there is a first part to it that is close to the one provided from the Merriam-Webster for “virtual reality”, and this is intended in order to establish the technological requirements of an IVE. This part states that the environment is computer-generated and it is responsive to users’ actions, thus it requires sensors that capture human behavior and actuators that render appropriate sensorial stimulation. On the other hand, the second part completes the MerriamWebster definition for “virtual reality” by adding the reference to these environments’ immersive property by stating that they should allow for naturalistic interactions with virtual elements. While, normally, users do not lose the ability to discern reality from virtuality when interacting with an IVE 6 (nevertheless, see [16]), quite often it is possible to capture realistic reactions to virtual elements [17], [18]. These realistic reactions are only possible if, indeed, the user becomes absorbed, immersed in the virtual world. Thus, we can say that IVEs have the ability to create, in its users, the illusion of being in a place other than where they actually are, or of having a coherent interaction with objects that do not exist in the real world – normally referred to as a feeling of presence [19]. Feeling of presence is an important subjective experience conveyed by the use of IVEs which, in the words of Sheridan ([20], p. 120), can be summarized as: “… a sense of being physically present with virtual objects ”. The feeling of presence renders the virtual environment credible and greatly contributes for an acceptable interaction with virtual objects. In order to convey this type of sensorial experience there are some features that an IVE should present: 1Unnoticeable hardware: An IVE should aim to convey the most naturalistic interaction possible. In this sense, obstructive hardware can be a problem as it functions as a constant reminder of the simulation limitations. A Head Mounted Display (HMD) might weight around 470 grams 7 , while reading glasses weight around 100 grams. The more the hardware of an immersive system goes unnoticed (consider not only the weight of the display and peripheral equipment, but also the existence of cables, markers, and similar components) the better immersion it will convey (see e.g. [21] and [22]); 6 Nevertheless, people are actively trying to create set-ups where they can lose discernment of reality while using VR apparatus. In March of 2019, Jak Wilmot, the co-founder of Atlanta-based VR content studioDisrupt VR, spent a total of 168 consecutive hours (1 week) in an VR headset, either seeing VR environments or pass-through computer generated images of its living room. The original report of this experience can be read here: https://futurism.com/guy-week-vrheadset/ 7 470 grams is the spec weight for both the Oculus Rift Consumer Version and the HTC Vive. The Sony HMZ-T3W specifications claim a weight of 320 grams, although battery, processor, and cable are not accounted for.
31 2Real-time update of the immersive environment: a quite important feature of IVEs is stimulation update based on user’s position and orientation tracking. The immersive capacity of a system will increase greatly if the image and/or the sound conveyed are dependent on the head movement of the user (and, obviously, congruent with the physical changes expected in the real world). This is, as we will see, still one of the mains sources of problems in today’s IVEs and one of the areas that can benefit the most from the capacity to adequately calibrate realtime responses against what is expected by users; 3Multimodal stimulation: The feeling of presence and consequent immersion appears to increase with the number of sensorial modalities that an immersive system is able to stimulate [23]. Theoretically, if the purpose is to convey full immersion , all the five senses should be stimulated using audiovisual Virtual Reality (VR) in combination with haptic stimulation and force feedback, smell, and taste replication. Nowadays there is no system that satisfactorily conveys a sensation of full immersion, nevertheless, multimodal immersive environments are a long existing reality and are regarded as more immersive than unimodal immersive systems; 4User-environment interactivity: Interactivity between user and the three-dimensional representations in the virtual world is paramount for an increasing sensation of immersion. Participants who were allowed to interact (e.g. move, rotate, catch) with stimuli of a virtual environment, report a higher feeling of immersion than passive observers [19]. Immersion and feeling of presence in a virtual environment are inseparable and positively correlated with the measured quality of the user experience in an IVE [24]. Thus, today’s developers aim to come out with ever increasingly immersive virtual reality experiences and any technological advancements in that sense puts them one-step ahead of the competition on a multi-billion dollar market 8 . Nevertheless, while technological developments on the four crucial properties for presence (unnoticeable hardware; real-time update; multimodal stimulation; interactivity) are constantly reaching new levels of performance, little progress has been done in order to understand the fundamental part of the immersion process: the user . In the next section, we explain why the current main challenge, when developing IVEs, 8 Augmented Reality and Virtual Reality Market are valued at USD 4.21 and 5.12 Billion dollars respectively, in 2017. Both are set to grow at a CARG above 30% up to 2023. The world’s biggest investment fund, the SoftBank’s Vision Fund managed by Masayoshi Son, has one of its highest investments on Improbable , a British company aimed at creating virtual environments indistinguishable from real world. See The Economist May 12th 2018 edition – “The Son Kingdom” report.
32 is not a technological one, but instead it is precisely a matter of understanding the interaction between these systems and the human element. 2.2. What are the problems with IVEs and how can we solve them? IVEs have the core goal of conveying the illusion of presence of (or in) something that, in fact, does not physically exists [24]. The systematic study and development of computerized IVEs can be traced to the early works of Ivan Sutherland and his student Bob Sproull that developed what has become known as the first Head Mounted Display (HMD), presented in 1968 (Figure 5). This was a system so heavy that had to be suspended from the ceiling thus inspiring its name, The Sword of Damocles 9 . This HMD was the first to use computer-generated visual stimuli capable of adapting to user head position. By today’s standards, it was a very primitive system of Augmented Reality (AR), both in terms of interface design as in terms of graphical representation of the virtualized stimuli. It was only capable of comprising simple wireframe cubic rooms in the binocular display, nevertheless the perspective of the room shown to the user was already dependent on the user’s position. The Sword of Damocles was a typical seetrough AR display, because images were conveyed on half silvered prisms, enabling a view of both the synthetic imagery and the real world surroundings (see definition of mixed reality , Section 2.3). Figure 5. The Sword of Damocles, a see-through Augmented Reality device. Ever since this computer science’s milestone, IVEs have gained crescent notoriety, alternating periods of high notoriety with periods of apparent vanishing from the scientific and academic debate [25]. Recent theoretical and technological developments have contributed for some consolidation of IVEs as a 9 The name is a reference to an anecdotic passage in Cicero’s Disputations , where the roman orator tells the story of Damocles a courtier of Dionysius II of Syracuse. Damocles once said to the king that he, as a great man of power, wisdom, and luxury was a truly fortunate man. Therefore, hearing this, the king offers to switch places with Damocles on the condition that he had to feel the same endangerment to his life as the king daily felt. To simulate this sense of endangerment, a huge sword was hanged, held at the ceiling by a single hair of a horse’s tail, right above the throne and directly facing the head of the new throne owner, Damocles. After a short period of time, Damocles ended begging Dionysius to switch places again.
33 consumer good technology and for the emergence of an increasingly number of VR and AR applications in the most varied fields, such as: science and data visualization, medical and surgical devices, telepresence systems, flight and driving simulators, therapy tools, digital arts, cinema, and videogames. Because of its inherent potential to directly interact with the human senses, immersive environments that make use of VR or AR have since long been regarded as in line to become the next predominant humancomputer interface. Carolina Cruz-Neira and her team suggested this possibility following the enthusiasm with the first Cave Automatic Virtual Environment (CAVE) system [26] (see Section 2.6), and more recently [27] and [28] referred immersive operative systems like the HoloLens, as in line for the next revolution regarding new concepts of user-interface (UI). However, in order to turn IVEs into a serious contender for the next predominant interface paradigm some current technologic and human factors limitations have to be overcome and, additionally, development approaches more focused on the analysis of human perception and action ought to be pursued. As soon as 1965, Sutherland had already outlined the ultimate challenge for IVEs developers in what has become known as the Sutherland’s challenge : “ Make that (virtual) world look real, sound real, feel real, and respond realistically to the viewer’s actions ” 10 [29]. More than fifty years later, and in a time where IVEs are becoming of widespread use, users still complain about their odd perceptual effects and performance deficits. These deficiencies affect all four factors involved on the triggering of the feeling of presence, mentioned in Section 2.1: 1 - Unnoticeable hardware: Today we already have wireless IVEs in two main forms: Smartphone VR Headsets (such as the Samsung Gear VR) and projection-based systems like the CAVE. It is true that in both systems users are not restricted by cable connections, however, while for Smartphone VR Headsets VR applications are still confined to passive-viewer and low quality video streaming, CAVE systems allow for user interactions with high-resolution virtual objects in large projection surfaces (room-like). Unfortunately, the CAVE system is not considered an electronic consumer good due to the high costs involved in assembling such 10 Sutherland’s conceptualization of the ultimate display also included some incursions in more extravagant predictions, such as: “The ultimate display would, of course, be a room within which the computer can control the existence of matter. A chair displayed in such room would be good enough to sit in. Handcuffs displayed in such a room would be confining, and a bullet displayed in such room would be fatal” [29].
34 virtual environment systems 11 . In terms of cost and quality, a middle-range solution is the PCbased premium VR headsets, commonly referred to as HMDs (such as Oculus Rift, HTC Vive, and the PlayStation VR). However, currently these systems still require cable connection 12 and the weight and size of these headsets keep raising complains from the users [24], [30]. In order to achieve wireless professional level HMD-based IVEs, developers will have to wait for 5G connection promises of wireless high throughput and low-latency connections to hold true [31]. 2 - Real-time update: Handling the delay between tracking the user’s head position and congruent change in the projection of a 3D image or of an auralized (3D) sound is one of the major challenges in the development of IVEs [32]. This type of delay is commonly referred to as end-to-end delay [33], and it is the consequence of the accumulated latency in several individual components, including tracking devices, signal processing and communication, and output devices. Consequences of end-to-end latency can range from output errors in terms of the visual perspective or 3D sound position, to more serious consequences as simulator motion sickness [34] [35]. End-to-end latency in IVEs results in temporal-spatial contradictory sensorial inputs to the user’s visual and vestibular systems. While the vestibular system might indicate that the user is tilting his head downwards 45 degrees, the visual render might still be displaying a 0 degrees perspective due to update lag resulting from end-to-end delay. These visualvestibular incongruences are the main etiological factor of motion sickness [36] and some studies indicate that the delays in linear visual oscillations (also known as heave motion) are particularly prone to cause motion sickness [37]. Some authors state that end-to-end delay should be below 15-20 ms in order to prevent simulator motion sickness [31], nevertheless, while some manufacturers already claim performance levels around these figures, controlled measurements report otherwise [38]. This is even more problematic in wireless collaborative IVEs, since current 4G communication involves a minimum of 25ms latency in controlled ideal operation conditions [31]. 3 - Multimodal stimulation: Truly multimodal IVEs, developed as such since the conception phase, are still scarce. Despite few exceptions from the past (such as the Morton Heilig’s 11 A Christie Mirage DLP projector can surpass the price tag of 50 000€, depending on the model. 12 Recently, speculation has surface about an Oculus Rift new generation headset codenamed Santa Cruz. Reports are stating that this model will be wireless, possibly taking advantage of 5G connection. The company has stressed that this is still a prototype and not a commercial product. See: https://www.theverge.com/2017/10/12/16463844/oculus-santa-cruz-standalone-headset-prototype-hands-on
35 Sensorama (see [39]) and the visuo-tactile VR system outputted from the European project PURE-FORM [40]), IVEs are predominantly unimodal. Moreover, multimodal IVEs are usually restricted to the actuation on two sensorial modalities – such as visual IVEs coupled with motion actuation (e.g. driving simulators), or audiovisual IVEs where stereoscopic imagery is coupled with spatialized audio. Current platforms of IVE development (such as Unity, Blender or Odeon) do not offer the possibility of developing with the same level of control for more than one modality. They are either fully focused on the development of the visual IVE (Unity and Blender) while offering comparatively poor add-ons for spatialized audio, or they are exclusive for acoustic modeling (Odeon). Moreover, users’ different sensorial modality channels operate with different timing constrains (see Section 3.1), which coupled with the end-to-end latency problem renders the synchronization of multimodal IVEs quite a challenging task. 4 - User environment interactivity: Seamless and convincing interaction with virtual objects is still one of the main challenges in IVEs development. This is a two-part problem where seamless and convincing are utterly difficult to combine in the same interaction actuator. On one hand, interaction is hampered due to the requirement to use peculiar peripherals (such as the HTC Vive wireless controller or the Oculus Touch VR controller), which require a learning process and prevent the use from performing natural movements (i.e., grabbing, catching, and pushing virtual objects) due to handling requirements of the controller. Nevertheless, seamless means of interaction are started to be employed mainly by using optic motion capture systems such as the ViconTM and KinectTM 13 . On the other hand, convincing interaction will require technologically complex (and not that seamless) tactile and haptic actuators, so the user can both interact and sense the physicality of the virtual object. Today’s haptic-actuators are mainly relying on vibration and force-feedback, which still require contact, or at least close proximity with fairly obstructive pieces of equipment [41] (see [42] for one of the less obstructive solutions). Contactless haptic actuators as ultrasound and air vortices are starting to be explored [43], however they are still very limited regarding their ability to simulate complex stimuli features such as size, shape, and texture of an object. 13 Although Kinect started as an prized success – was on the Guinness Book of World Records as the fastest-selling consumer electronics device with 8 million units in 60 days – Microsoft recently announced it will stop its development and production on the grounds of diminishing financial returns (see https://www.popularmechanics.com/technology/gadgets/news/a28783/microsoft-ends-production-of-kinect-attachment/).
36 As we can understand from the examples above, there are still several problematic issues with today’s IVEs. While some of these problems translate in almost purely technical challenges (reducing the equipment weight, for instance, or develop integrated virtualization platforms for multimodal IVEs), others will require user studies and fundamental research in Human Perception in order to support the development of technical solutions. Take, as an example, the above-mentioned difficulty in synchronizing the output of stimuli from different sensorial modalities. To resolve this problem, an in-depth knowledge about the human perceptual mechanisms that allow for synchrony perception is required. For instance, it is known that humans still perceive as synchronous an audio-visual stimulus where sound lags image for some milliseconds. This capability has been attributed to the existence of a Window of Temporal Integration (WTI) [44]. Nevertheless, we still need to know if this tolerance to lagging sound allows accommodating the typical audio-visual mismatches due to end-to-end delays found in the current IVEs. Moreover, we need to understand if this WTI is affected by some other characteristics of the IVEs, such as the distance at which the virtual stimulus is presented to the user. We will be looking in more detail into this problematic in Section 3.1. Because models of human perception and performance are the end-result of this empirical approach, as well as the benchmark of the evaluation/development iterative process described in Figure 4, in the following section, we will explore in more detail how these models can guide us in the development of increasingly immersive virtual environments. 2.3. On Understanding and Modeling Perception and Performance for the Development of IVEs Accurate predictions about how users perceive and interact with a computerized environment are of the foremost importance in the design of computer systems that emphasize usefulness and usability. Furthermore, any attempt at developing an interactive system should put the human, the user , in a central position that defines all the subsequent discussion and design [10] [13]. Knowing and understanding the perceptual mechanisms that allow users to perceive the natural world should be indispensable for the design of any satisfying interactive environment which makes use of audiovisual virtual or augmented reality. Computer generated immersive environments are broadly classified into two different categories: VR environments and AR environments. The distinction between these two is not a procedural or technical one; rather it is more of a performance-based distinction. The Reality-Virtuality (RV) continuum of Milgram and Colquhoun [45] has been widely used to clarify this distinction and provides a good theoretical
37 framework to derive a taxonomy on immersive environments. In the reality-virtually continuum the fundamental distinction is between Real Environments and Virtual Environments , that are located on opposite ends of the continuum (see Figure 6). Figure 6. Reality-Virtuality Continuum (adapted from [45]). The location of any computerized immersive environment along this continuum coincides with its location along a parallel Extent of World Knowledge (EWK) continuum , highlighting in this way, the importance of knowing about both the physical world and the human perceptual mechanisms that govern the perception of these physical signals. As depicted in Figure 7, on the right end of the R-V continuum are virtual environments (which need to be completely modeled in order to be rendered) implying, in this sense, full knowledge about the physical properties of the world and how they are perceived by humans. However, if we increase the granularity of the R-V continuum we can identify Augmented Reality (AR), as an environment where more than half of the scenario is real environment with no modeling whatsoever; and Augmented Virtuality (AV), as an environment where more than half of the scenario is virtual environment. These two types of intermediate stages are sometimes referred as Mixed Reality (see Figure 7). Figure 7. Reality-Virtuality Continuum, AR, and Augmented Virtuality (adapted from [45]).
38 A clear example of the close relation between the EWK continuum and the R-V continuum is the simulation of binocular depth cues in order to provide visual depth perception. The technique of Stereoscopy can be traced back to the first half of the nineteenth century [24], although initial demonstrations of the effect were developed when the principles of retinal or binocular disparity were discovered in the XVI century [46]. In stereoscopy, depth is simulated based on the fact that we naturally process two slightly different images of the same visual scene – one for each eye – and the bigger the offset between the two images the closer the object is perceived (see Figure 8). The differences between the images that each eye processes can be calculated based on the person’s eye separation and thus, binocular depth cues can be easily simulated if we control the image that each eye is seeing and if that image is adjusted for the position of the correspondent eye. A B Figure 8. Retinal image offset due to binocular disparity (A). Measurement of pupillary distance (B) performed after fixating on an object around 6 meters away and by marking each eye fixation with a marker. Use of the pupillary distance as a parameter to define the geometry of stereoscopic scene in some virtualization platforms such as in the Game mode of the Blender software [47] (C). The relevance of human’ studies for the development of IVEs is latent in the parallelization between the R-V and the EWK continuum. The most important principle in the design of immersive environments states that the environment should convey an accurately replication of the geometric and temporal characteristic of the real world. This does not mean that computerized environments must model precisely all physical attributes of a visual or an auditory real world’s scenario – in fact this is, and it is expectable that will remain being, technologically impracticable. What this principle really means is that the perception that a user has in a VR or AR environment, should be quantitatively indistinguishable from a C
39 correspondent scene perception in the real world. In the words of Machover and Tice in their seminal paper on Virtual Reality [48]: “ …the quality of the experience is crucial… the ‘reality’ must both react to the human participants in physically and perceptually appropriated ways, and to conform to their personal cognitive representations of the microworld in which they are engrossed. The experience does not necessarily have to be realistic – just consistent ” (p.15). Helping to measure performance in order to convey a consistent experience is where the discipline of HCI can play a key role. Applying HCI methodologies to the development and evaluation of immersive environments becomes an issue of the outmost importance, primarily when we think about the current lack of knowledge about: (1) human perceptual mechanisms – e.g., the mechanisms governing natural audio-visual asynchronies, the binding of perceptual cues to yield visual and auditory depth perception, and audiovisual recalibration phenomena; (2) human performance in IVEs – e.g., realistic interaction with virtual elements, auditory and visual depth perception and performance validity on high fidelity simulators; Knowledge about these two dimensions (human perception and human performance) is central to replicate the natural tridimensional world interaction conditions in IVEs. Thus, in the IVEs context, the use of HCI interaction-modeling methodologies and empirical research techniques can both give us some insight about the human perceptual mechanisms that are crucial for the design of a highly immersive VR or AR environment and, at the same time, allow us to accurately evaluate human performance of an IVE user and compare it with models of human performance in the real world. This comparative evaluation had widespread used in the development of PCMs that have boosted HCI influence in the design of digital interfaces. The same approach should be replicated for the development of IVEs, particularly if we expect them to become a credible and widespread UI option.
46 Figure 13 . Pointing devices tested by Card and colleagues. Image from [56]. Similarly to the original Fitts’ experiment, Card et al. [56] selection targets varied in distance to the cursor’s initial position and in size (depending on the number of characters of the text target). The authors were able to compute different task’s index of difficulty and manage to apply a linear regression to each one of the interaction conditions. The linear models accounted for more than 80% of the variation in time to reach the target as a function of the task index of difficulty, for all the interaction conditions. Thus, design choices based on PCMs were demonstrated feasible as the researchers could accurately predict how the user would perform with each one of the devices in different tasks of varying difficulty index. Figure 14. Linear Relation between the task ID and its time-to-completion, for a reaching task with different input devices – continuous input devices (right) discrete input devices (left). Data from [56], adapted using WebPlotDigitizer [55]. Prediction Models for Performance time of different interaction modes on Card et al. (1978)
47 By applying Fitts’ law models to their data, Card et al. were able to show that continuous input devices were a better design choice and that among those, the mouse allowed for the fastest and more accurate 17 level of performance (see Figure 14). More recently Mackenzie and Jusoh [57] following the design choice perspective of Card and colleagues, tested performance on a typical reciprocal tapping task using a remote pointer versus a mouse. The distance to hover between targets were 40, 80, or 160 pixels, while the width of the target could be 10, 20, or 40 pixels. Table 2 shows real performance and performance predicted by the Fitt’s law models. Table 2. Different task conditions in [57] and respective predicted and observed time to completion. A W We IDe Real Time (ms) Mouse Predicted Time (ms) Mouse Real Time (ms) Remote Pointer Predicted Time (ms) Remote Pointer 40 10 11,23 2,322 665 605 1587 1473 40 20 19,46 1,585 501 487 1293 1222 40 40 40,2 1,000 361 362 1001 954 80 10 10,28 3,170 762 798 1874 1884 80 20 18,72 2,322 604 648 1442 1564 80 40 35,67 1,585 481 505 1175 1259 160 10 10,71 4,087 979 973 2353 2259 160 20 21,04 3,170 823 792 1788 1872 160 40 41,96 2,322 615 621 1480 1507 Mean Time 643 643 1555 1555 Figure 15. Linear Relation between the task effective ID and its time-to-completion, for a pointing task with different devices – Mouse (black dots) Remote pointer (grey dots). Data from [57], adapted using WebPlotDigitizer [55]. 17 Card, English, and Burr [56] also explore the error-rate for each type of interaction. Results rank devices in the same way as performance time. y = 203,94x + 158,64 R² = 0,9702 y = 435,21x + 520,2 R² = 0,9557 0 500 1000 1500 2000 2500 012345 Time to reach target (ms) Effective Index of Difficulty (IDe) Prediction Models for Performance time of different interaction modes on Mackenzie et al. (2001) Mouse Remote Pointer
48 Once again, it is interesting to point out the accuracy of the models in Figure 15, that yielded predictions that were on average precise to the millisecond (see last row on Table 2). This added accuracy might be due to a small, although important, change to the original Fitts’ equation, proposed by Fitts in 1964 and extensively used by Mackenzie in different applications [58] [59] [60]. In 1964, Fitts suggested that when building a predictive model one should account for the variability of human motor performance during “hits” (or on-target) responses by defining the target width ( W) as: Equation 4 . 𝑊𝑒=4.133×𝑆𝐷𝑥 , where 𝑆𝐷𝑥 is the standard deviation of the distribution of “hits”. Using what was named effective width ( We ) [53], we are considering motor variability by taking into account a consequent 4% probability for boundary errors and accommodating this variability by adapting the considered width of the target accordingly. This added improvement of the Fitt’s models’ predictions illustrates the continuously adapting nature of PCMs. By incrementally accounting for ever-detailed descriptions of human performance (based on incremental knowledge about human cognition and behavior), PCMs have the tendency to become more complete and to yield increasingly accurate predictions. Hick-Hyman Law PCMs are also suitable to explain intrinsically cognitive tasks as shown by the Hick-Hyman law for choice reaction time [61] [62]. Similarly to Fitt’s Law, the Hick-Hyman law also takes the form of an information processing equation in order to make predictions about the reaction time in a task of target selection. In the original experiment Hick [61], presented the participants with a set of 10 pea lamps, arranged irregularly and connected to a set of 10 Morse keys (one dedicated to each finger). One random lamp would turn-on every 5 seconds and the participant had to press the associated key to turn it off. There was no logical arrangement between the lamp and associated key position and thus, participants had to first learn the association scheme. Crucial to the experiment, Hick presented the participants with several conditions that varied in the number of stimuli, ranging from sets of just 2 lamps to sets of 10 lamps (Figure 16).
49 Figure 16. Author’s rendition of the experimental apparatus based on Hick’s 1952 experiment description, from [63]. Hick observed that average choice reaction time would increase logarithmically with the number of different options to choose from [61]. This relation could be predicted by the following equation: Equation 5. 𝐶ℎ𝑜𝑖𝑐𝑒 𝑅𝑒𝑎𝑐𝑡𝑖𝑜𝑛 𝑇𝑖𝑚𝑒 (𝐶𝑅𝑇)=𝑎+𝑏 𝑙𝑜𝑔2(𝑛+1) , where CRT is measured in milliseconds, a and b are constants empirically determined and n is the number of options. When plotting the real data of choice reaction time as a function of the degree of choice, Hick found that the data dispersion could be well explained by the equation 𝐶𝑅𝑇=0.518 𝑙𝑜𝑔(𝑛+1). Thus, choice reaction time appears to be governed by an information theory principle, where additional alternatives function as added entropy and the human could be said to be functioning at almost full capacity in situations where the degree of choice was around 8 to 10 different options (Figure 17).
50 Figure 17. Data from subject A (Hick himself 18 ) in experiment 1. Adapted from [61] using WebPlotDigitizer [55]. Almost at the same time Hick was conducting his experiments, Hyman was arriving to similar conclusions in a somewhat more complex experimental design [62]. Hyman’s experiment had two parts: a first one where each lamp had the same probability to be turned on; a second part where probability of being turned on would differ from lamp to lamp allowing the participants to anticipate that certain lights had a higher probability of being turned on. This difference is quite important because it reflects the conditions of certain systems where specific operators, widgets, or telltales are more frequently involved in the interaction loop. For the information processing model describing choice reaction time, this distinction has some mathematical implications, namely on the part concerning system entropy: 𝑙𝑜𝑔2(𝑛+1). Thus, for a set of alternatives with different probabilities, the information entropy value (H) should be: Equation 6 . 𝐻= ∑𝑝𝑖 𝑙𝑜𝑔2(1 𝑝𝑖+1) 𝑛 𝑖=1 , where n is the number of alternatives and 𝑝𝑖 is the probability of the ith alternative. Consequently, if alternatives have equiprobable entropy, H is equal to: 𝑙𝑜𝑔2(𝑛+1), and it is said the system works at maximum information rate or maximum entropy. Conversely, when alternatives are not equiprobable, the entropy of the stimuli is reduced [Seow2005] and choice reaction time improves, mainly because the user learns to anticipate certain occurrences. Thus, although Hick demonstrated that 18 Although not acceptable by today’s standards, it was common practice in the first half of XX century for experimental psychologists to participate as subjects in their own experiments. The rational was that by having devised controlled ways to measure cognitive, perceptual, or motor human performance, results would not be affected by knowledge about the experiment’s goal. Iconic examples of self-experimentation were the pioneer of human memory study Hermann Ebbinghaus (1850-1909) and the social psychologist Stanley Milgram (1933-1984). For a reflection on self-experimentation see [236]. y = 0,518 log10 (n+1) R² = 0,9888 0 100 200 300 400 500 600 0 2 4 6 8 10 12 Choice reaction Time (ms) Degree of Choice (n) Choice reaction time as a function of the number of alternatives (Hick, 1952)
51 one could reduce entropy by reducing the number of alternatives, Hyman played with the probabilities of the stimuli to yield different amounts of entropy and study how choice reaction time could be described as function of H [63]. In its Experiment II, Hyman [62] comprised 8 different conditions, varying in stimuli set size and probability for each alternative, which resulted in the presentation of stimuli arrangements with entropy ranging from 0.47 to 2.75. Figure 18, shows how reaction time was affected by the different levels of entropy. Figure 18. Data from experiment II. Adapted from [62] using WebPlotDigitizer [55]. Hyman results demonstrate once again how linear models can be used to explain and predict human performance, this time in a purely cognitive task (i.e., Hyman task did not require a motor action, instead participant’s provided verbal answers). With the extension of Hyman’s findings, Hick’s law was renamed Hicks-Hyman law and is commonly represented in its linear form as: Equation 7 . 𝐶𝑅𝑇=𝑎+𝑏𝐻𝑡 , where a and b are empirically determined constants and Ht is the entropy as defined in Equation 6. Implications of the Hick-Hyman law are obvious for information systems design. Knowing the performance ceiling in choice tasks can help developers and designers to adapt the interface by reducing the information content or by creating hierarchical presentations of the information, based on probability of use. Moreover, in safety critical systems where choice performance plays an important role (as in control rooms or in the operationalization of medical devices), warnings, alarm systems and mitigation procedures can be implemented taking into account the typical human choice reaction time. Considering y = 191,31 + 192,57x R² = 0,9852 0 200 400 600 800 0 0,5 1 1,5 2 2,5 3 Choice Reaction Time (ms) Entropy (H) Choice reaction time as a function of entropy [Hyman1953]
52 the human cognitive capacity might help to prevent information overload in systems that require limited reaction time from the user. PCMs based on the Hick-Hyman law were also found suitable to describe selection of items in hierarchical menus [64] [65], control displays [66] and data visualization [67], and to describe the operation of mode-based applications in a tablet interface [68]. Nevertheless, and as pointed out by Seow, while “Fitts’ Law has enjoyed and continues to receive a great deal of attention in the field [HCI] the same cannot be said for the Hick-Hyman Law [that fail to] gain momentum in the field [63]”. According to Seow, there are two main reasons why the Hick-Hyman Law has fallen short of the impact other models have caused in HCI: 1. Difficulty in Application – To apply the Hick-Hyman Law in its traditional formulation, one must first codify equivalent events into equiprobable or non-equiprobable alternatives in order to define information entropy. In today’s complex interfaces, comprising a variety of multidimensional stimuli and different informational content, this is a difficult task to accomplish. This limitation often conveys the application of the law to simplified and highly controlled experimental set-ups which are unrepresentative of the current interfaces; 2. Levels and type of performance – While Fitts’ Law captures and predicts a type of human performance mainly related to dexterity and motor behavior – which fostered its application in the study of control systems –, Hick-Hyman Law addresses a cognitive task with dimensions which are harder to directly observe and measure. Despite their limitations, these influential PCMs made a lasting mark in HCI by first showing how controlled experimentations with user’s yields valuable predictive models for both motor performance and cognition, allowing for informed design choices and incrementing the fundamental knowledge on interaction between humans and computing systems. In fact, Fitts’ Law and Hick-Hyman Law became default principles in more ambitious human PCMs intended to describe the entire perception-action loop involved in human-computer interaction, such as the Keystroke-Level Model (KLM) [69].
53 The Keystroke-Level Model In 1983, Card, Newell, and Moran, published one on the most seminal books in HCI called “ The Psychology of Human-Computer Interaction ” [70]. In their book, besides presenting the case for the articulation of cognitive and computer scientists in favor of the development of better user interfaces, Card et al. also presented a general PCM – the Keystroke-Level Model (KLM). Following the tradition initiated by Paul Fitts research, the KLM also uses the knowledge about the human motor system – inferred from simple experimental controlled tasks – as a basis for detailed predictions about user performance in compounded tasks. This model splits the execution phase into five different physical motor operators, a mental operator and a system response operator: K – Keystroking, the time it takes to strike a key; B – Pressing a mouse button; P – Pointing, moving the mouse (or a similar device) into a target; H – Homing, or the time that it takes to switch the hand between the mouse and the keyboard; D – Drawing lines using the mouse; M – Mentally preparing for a physical action; R – System response, which may be ignored if the user does not have to wait for it. The execution of a task will involve interleaved occurrences of the various operators. The KLM predicts the total time for the execution of a task by adding the component times for each of the above activities, in such a way that the total time T is: Equation 8. 𝑇𝑒𝑥𝑒𝑐𝑢𝑡𝑒 = 𝑇𝐾+𝑇𝐵+𝑇𝑝+𝑇𝐻+𝑇𝐷+𝑇𝑀+𝑇𝑅 , where the total time of each operator is the sum of all its occurrences in the described interaction process In many examples, the system response time is defined as zero – when needed, this response time can be measured by observing the system and be taken into account in the calculations. The times for the other operators all depend on the skills of the user and they can be obtained from empirical data with preliminary tests. Card, Moran, and Newell [69] gave us a general indication of the time for each operator, based on a sample of 1280 user-system-task interactions, comprised of various combinations of 28 users, 10 systems, and 14 tasks (see
54 Table 3). Although individual predictions may be interesting, the power of KLM lies in comparison between different systems. Having a detailed description of the methods to perform key tasks, we can use KLM to tell different systems apart by predicting, for example, which one allows for the faster interaction. This is considerably cheaper than conducting lengthy experiments and, furthermore, the systems do not even have to exist in a physical form. Table 3. Times for operators in the KLM (adapted from [69]) and updated with [11] values for pointing. Operator Remarks Time (ms) K PRESS KEY Time varies with tipping skill: Expert Typist (135wpm) Good typist (90 wpm) Average typist (40wpm) Typing complex codes Poor typist (unfamiliar with keyboard) 80 120 280 750 1200 B MOUSE SCROLL Down/Up 100 P POINTING WITH A MOUSE (Includes terminating mouse-button click) Fitts’ law Indicative average time 𝑀𝑇=159+204[log2(𝐴 𝑊+1)] 1100 H HOMING KEYBORAD-OTHER DEVICES 400 D DRAW 𝑛𝐷STRAIGHT-LINE SEGMENTS OF TOTAL LENGTH 𝑙𝐷 900𝑛𝐷+160𝑙𝐷 M MENTALLY PREPARE depending on the task type, Hick-Hyman Law might be applied Indicative average time 𝐶𝑅𝑇=𝑎+𝑏[∑𝑝𝑖 𝑙𝑜𝑔2(1 𝑝𝑖+1) 𝑛 𝑖=1 ] 1350 R RESPONSE BY SYSTEM System and command dependent values, measured empirically t
55 In their original KLM experiment, Card and colleagues compared 14 tasks performed in different text and graphics editors. Figure 19, shows the predicted vs observed time of execution of each task. The position of the data points along the diagonal indicates that predicted and observed time are quite similar, staying under a 10% error rate in 12 out of 32 tasks. Figure 19. Observed versus KLM predicted time for text editing tasks. Data adapted from [69] using WebPlotDigitizer [55]. Using an example adapted from [11], we can understand the usefulness of KLMs for implementation comparisons and design choices. Consider the task of changing the second term on the equation depicted in Figure 20 to boldface and Arial font, using (1) only the mouse and the Windows Word GUI, and (2) using only the keyboard. Figure 20. Mouse operation for changing the second term of the equation to Bold and Arial font in Microsoft Word 2016®. ① to ⑤, represents the operation’s order. 1 2 4 8 16 32 64 1 2 4 8 16 32 64 Observed Execution Time (s) Predicted Execution Time (s) Observed VS Predicted in the KLM validation Expriment [Card1980] POET SOS DISPED MARKUP DRAW SIL ALL ① ② ③ ④ ⑤
62 was one condition of an experiment by Nilsen [79]. In Nilsen’s experiment, a digit was first showed to the subject who would then proceed by clicking on a target, causing a vertical menu of digits (from 1 to 9) to appear in a random order below the cursor. The subject’s experimental task was to point and click on the previously showed digit. Nilsen’s results showed that the time to select the correct digit was a function of its location on the randomly ordered menu, and that it was a fairly linear function with a slope of about 100ms per item. Modeling Nilsen’s task in EPIC required estimating a parameter for how long it takes to recognize the text label for digits, and where on the retina this recognition happens. Then, two models were constructed: (1) a serial search model that corresponded to a one-at-time hypothesis checking for visual search, where the eye is moved to the next object down the menu and if it matches the sought for target-item, a pointing movement to the that item is initiated; (2) an overlapping search model where the parallel processing possible in EPIC is fully exploited because two functions are actively working at the same time, one that is moving the eye as rapidly as possible from one item to another (SACCADE-ONEITEM) and the other that is monitoring the work memory (STOP-SCANNING, see Example 1) expecting the emergence of the target item to stop the search and start the movement of the cursor to the target. Example 1. Production rules for the serial search model (left) and for the overlapping search model (right), from [76]. Figure 23 shows the results from the predictions made by the two models using EPIC, and the results from Nilsen’s experiment with humans. As we can see, there is a considerable difference between
63 the serial-search and the overlapping-search model for the item selection time predicted as a function of the item position. Clearly, the model that best explains the real human performance is the overlappingsearch model , a model that exploits in a more satisfactory way the multitasking nature of some human performances. Once again, this result reveals the ever-improving nature of PCMs, which tend to improve their predictions with the increment of the accurate and completeness of their models of human cognition. Figure 23. Menu selection times of the observed data from [79] with the predicted times by two EPIC models [76], the Serial Search Model (SSM) and the Overlapping Search Model (OSM). Similar to EPIC, the ACT-R also consists of a set of modules dedicated to a specific kind of information but capable of contributing with outputs for the integrated cognition process. [77] illustrates the basic architecture of ACT-R as a visual module – capable of identifying objects in the visual field –, a manual module – to issue commands and control the actuation of the hands –, a declarative module – for retrieving information from the memory – and, finally a goal module – to keep track of the current goals and long-term intentions (Figure 24). 0 1000 2000 3000 4000 0 1 2 3 4 5 6 7 8 9 10 Item Selection Time (ms) Item Position Search and Selection Performance [Kieras1997] Observed OSM SSM
64 Figure 24. Information architecture of the ACT-R version 5.0 adapted from [77]. In parentheses are main brain regions where these functions have been identified and that the ACT-R attempts to modulate. The coordination of these modules is achieved by a set of rules executed in the Central Production System (CPS), intended to modulate the functioning of our central nervous system. The CPS is not sensitive to most of the activity occurring in the different modules, but rather it only responds to the information deposited in the buffers of those modules. The original architecture of the ACT-R was not committed with how many modules existed and consequently new modules or production mechanisms were implemented as part of the core system. For instance, in 2006 a new version of the ACT-R code was released (ACT-R 6.0) which included a new mechanism in the CPS called dynamic pattern matching [80], and more recently dedicated modules for speech and audition where also developed. Similarly to EPIC, the ACT-R is available in a Common LISP distribution 20 , and its framework allows developers to create new modules or partially modifying the behavior of the existing ones, by changing their standard parameters. The general accuracy of some PCMs is, in fact, notable and it is highlighted here because it demonstrates quite well the mutual contribution between cognitive and computer sciences. PCMs in HCI tried to integrate aspects of perception, attention, short-term memory operations, planning, and motor behavior in a single (executable) model, at a time when most cognitive science models addressed only 20 ACT-R software can be downloaded here: http://act-r.psy.cmu.edu/software/; an implementation in Python can be downloaded here: https://github.com/jakdot/pyactr
65 isolated laboratory phenomena [7] – and this attempt obtained some success. In an almost symbiotic way, HCI PCMs have also been benefiting immensely from periodic actualizations based on the latest advances in experimental psychology. As we saw, the tuning of these interaction models often relies on the definition of fixed and variable parameters that are only accessible through psychological and human factors experimentation on real users. This relationship of mutual benefit is becoming increasingly clear for professionals of both areas, and the future of both areas may lay on the fact that they depend on each other – technology designed for humans will always require human studies and the knowledge frontier on human cognition and performance may only be extended using new technology. Having in mind this symbiotic relation between cognitive and computer sciences within the HCI field, we can say that one of the most important questions for the work presented in this thesis is, what can be the practical use of merging experimental psychology and computer sciences on the development of immersive environments? As we could see from the PCMs presented before, the knowledge about human characteristics can add a lot to the modeling of human performance and subsequent design and implementation in computerized systems that require human interaction. By looking at the current state of IVEs, we will get a clear sense of the existing limitations and of how our approach can help to surpass them. The importance of adopting an iterative logic of user evaluation / re-design (described in more detail in Figure 4), is even more clearly understood when we analyze the limitations that current IVEs still present. By bringing back the empirical approach of the first HCI researchers of the like of Fitts, Card, and Mackenzie, and by resorting to controlled user testing, we will propose new PCM and design guidelines targeted at improving such IVEs. At a time when IVEs are being considered as the next revolution in terms of UI, it is only natural that the HCI principles of controlled user testing are transposed, adapted, and applied to IVEs development and research. 2.6. The current state of IVEs technology Visual IVEs As pointed out by Jerald [24], today’s visual IVEs are implemented in one of three ways: HMDs, world-fixed displays (screen or projection-based displays), and hand-held displays. In this section, we will review the first two, since hand-held displays (such as tablets) are out of the scope of highly immersive environments. HMDs are, indisputably, the currently most explored support in the development of IVEs. In 1984, the NASA Ames research center presented the Virtual Visual Environment Display (VIVED), a
66 monochromatic stereoscopic HMD, capable of displaying computer-generated images (CGI) with a frame rate that varied from 20 Hz up to 60 Hz depending on the complexity of the graphics [81]. The use of HMDs as the preferred support for IVEs continues today, as Oculus Rift® became the first enterprise to successfully commercialize this type of solution as a consumer good (Figure 25). Figure 25. PC-based premium HMDs currently commercially available. Top-Left is the Oculus Rift Consumer Version 1; top-right is the HTC Vive; down is the PlayStation VR from Sony (images distributed under a CC-BY 2.0 license). Independently of the type of displays, an HMD requires the presentation of different images to each eye, thus making extensive use of binocular cues (binocular disparity and convergence) to provide visual depth perception. HMD’s are comparably easy to transport and set-up and its price is also an important advantage when compared with other projection-based or world-fixed displays systems. Moreover, HMDs have been developed to address all the spectrum of the R-V Continuum with purely VR, mixed reality, and AR devices already on the market (Figure 25). Figure 25. The Samsung Odyssey and the Microsoft HoloLens, two of the state-of-the-art equipment for mixed reality and augmented reality.
67 However, the problems and limitations of HMD are many. We have already referred some of them in Section 2.2, nevertheless, we will provide next some additional detail and link the limitations to different types of HMDs technology. One general limitation is the end-to-end delay, with current performance failing to achieve the 15-20ms value that might prevent motion sickness [31]. Nevertheless, threshold values for latency perception vary greatly with the experimental condition (i.e. type of head movement, differences in the experimental protocol, and individual variability) [82]. Having built a laboratory HMD with an end-to-end latency of 7.4ms, Jerald [83] identified a participant capable of noticing a delay of just 3.2ms. In AR HMDs, where the real and synthetic images are merged in space and time, the latency problem becomes even more visually noticeable because any latency will result in displacement between CGI and real world scenarios. [24], suggest that the threshold should amount to no more than just 1ms for this type of displays, mainly due to the possibility of participants judging latency in relation to real-world fixed objects. A list of negative effects of latency is provided by [24] and, in addition to motion sickness, also includes: degraded visual acuity; degraded performance; breaks-in-presence; and negative training effects. A second problem with optical see-trough HMDs is light intensity. Light conditions of the real and the virtualized world have to be carefully controlled; otherwise, if the synthetic imagery is too bright relative to the ambient light, the real environment will not be visible – the reverse being also true. A final difficulty with optical see-trough displays is that of occlusions. If a virtual object should appear in front of a realworld object, it will generally appear to be semitransparent (holographic-like) and this transparency increases as a function of the occluded real-world object’s luminosity. Fuchs and Ackerman [84], described an interesting problem of visual displacement in mixed reality HMDs, an equipment that conveys to the user a view of the real world through one or more video cameras mounted on the front of the HMD. A mixed-reality HMD merges virtual imagery with camera recordings of the real world, thus giving the user a completely CGI environment. The problem is that without carefully optical considerations, it is quite difficult to align the camera’s view with the normal viewing axis of the user’s eye. The image sent to each of the user’s eyes is taken from a camera perspective other than that of the observer’s eyes, and this could distort the user’s sense of depth because the stereo pairs have an effective pupillary distance different from that to which the user is accustomed. The interesting perceptual result of this incongruence is that the user might get a false sense of height, particularly for near field objects.
68 Other performance limitations are general to all current HMDs, starting with the Field of View (FoV) problem. Humans have a horizontal FoV of about 160 degrees (≈114 degrees of binocular vision, ≈40 degrees of peripheral non-binocular vision) and a vertical FoV of about 120 degrees [50]. The FoV has been linked to a greater sense of immersion ([85], [86]), nevertheless current HMDs are still limited regarding this aspect. Information available about the FoV specifications of existing devices is contradictory, with specifications from manufacturers differing from measurements performed by independent users and technology reviewers. One of the few reviews that presents values for both horizontal and vertical FoV indicates a horizontal value of 100º FoV and a vertical value of 110º FoV for the HTC Vive, while the Oculus Rift CV1 presents values around 80º FoV horizontal and a vertical value of 90º FoV 21 . Published controlled measurements are scarce, nevertheless, [86] found values around 90º FoV for the Oculus Rift CV1. For the AR headset Hololens, the FoV limitation is even more problematic with a reported horizontal value of 30º FoV and a vertical value of 17.5º FoV. The Headset-fit and weight are also reported as important usability issues affecting HMDs [24]. Loose-fitting HMDs can cause unwanted small scene motions, while overly tight HMD will cause too much head pressure and discomfort. This problem is positively correlated with the weight of the equipment, which is around 470 grams for the Oculus Rift CV1 and the HTC Vive and can amount to around 579 grams for the Microsoft HoloLens. IVEs based in world-fixed displays, have the capacity to deal with the generality of the usability issues linked to HMDs, since they remove the physical affixation of the projection to the user’s head, nevertheless they are not immune to the perceptual problems that also affect HMDs. In 1993, CruzNeira, Sandin, and DeFanti presented their work on the design and implementation of the Cave Automatic Virtual Environment (CAVE), and according to these authors, the CAVE came (at the time) as the first system that fully met the standards that defined a VR system. In their words: “… a VR system is one which provides a real-time viewer-centered head-tracking perspective with a large angle of view, interactive control, and binocular display. ” [26] (p.135). Contrary to most of the prior technology, the first CAVE presented at SIGGRAPH ’93 was conceived as a tool for scientific visualization and, therefore, high performance and accuracy in the representation of the real world was needed in order to convince leading-edge computational scientists to use the CAVE 21 Review accessible on https://www.vrheads.com/field-view-faceoff-rift-vs-vive-vs-gear-vr-vs-psvr
69 in alternative to other more traditional immersive environments. In this sense, the goals that inspired the CAVE engineering effort were: 1. The desire for higher-resolution color images and good surround vision without geometric distortion; 2. Less sensitivity to head-rotation induced errors; 3. The ability to mix VR imagery with real devices (like the user hand, for instance – which leaves out the necessity of DataGloves -like equipment or other peripherals to interact with the virtual environment); 4. Minimize the existence of attachments and disturbing hardware around the user; 5. The desire to couple to networked supercomputers and data source for successive refinement. The original CAVE, developed at the University of Illinois was a 3x3x3 meters cubic structure, comprising three projection screens on vertical walls (the front one and the two lateral ones, with respect to the user) and one projection screen for the floor.To “transform” cube sides in projection planes, a CAVE system frequently uses a Window Projection Paradigm in which the projection plane and projection point relative to the plane are specified, thus creating an off-axis perspective projection. To get a differentiate image to each eye separately, the most common solution is to use frame sequential stereo with synchronized shutter glasses, a solution that immensely reduces the flicker effect. Images are rear projected in the walls, so that participants in the CAVE do not cast shadows on the projection screens – only in the ground, when the projection comes from the top of the user. The user’s tracking was originally performed with PolhemusTM tethered electromagnetic sensors, but currently it is made resorting to high precision motion tracking optical systems, such as the ViconTM or the QualisysTM systems. By combining high precision tracking with high performance computation 22 , CAVE systems are ideal to convey IVEs with high-resolution images and lower end-to-end latencies. CAVE systems might constitute a solution for the limited FoV problem and for the usability issues of obstructive hardware on HMDs. A CAVE system provides a surround projection and requires no peripherals other than lightweight stereoscopic glasses and retroreflective markers for user tracking. Nevertheless, a CAVE configuration is not exempted of usability issues. One of the most relevant issue is the space the user is confined to, often a 3m2 surface surround by 3m2 walls. Thus, natural navigation 22 In world-fixed displays IVEs computers can be housed further from the display, thus avoiding any space limitation. State of the art graphic boards and processors can be used to construct computer clusters that might optimize the performance of the IVE assuring high resolution images and low end-to-end latency.
70 (i.e., walking) in the real room is restricted and the user has to resort to commands in order to navigate the virtual space. Some CAVEs have dwelt with this problem by applying omnidirectional treadmills (Figure 26), thus allowing for continuous walking. However, this solution also carries some problems. As we can see in Figure 26, the user as to use a safety harness, adding obstructive equipment to the simulation, and the treadmill control to adapt to some natural walking patterns such as accelerations, decelerations and turns is quite difficult to obtain mainly due to the necessity of predicting user’s waking pattern to adapt timely [87]. Figure 26. A CAVE visualization system with an omnidirectional treadmill from the U.S. Army Research Lab. From [88] A more convenient solution for the space restriction problem is the adoption of a Power-Wall configuration. By simply opening the left and right sides of a CAVE, one can get a power-wall of about 9 meters wide by 3 meters height (Figure 27). This allows for plenty of additional room to lateral movements (though maintaining the same limitation in forward movements), while keeping FoV values considerably high (FoV values will increase as the user’s get closer to the projection surface). Figure 27. The author using a CAVE-like system in a power-wall configuration. Image from the Visualization System located at the Center for Computer Graphics (CCG), Guimarães-Portugal.
71 Recently a new configuration was developed for CAVE systems, which is an interesting combination of the traditional CAVE and the power-wall solution. In 2018 Antycip Simulation, a French company that develops visualization systems, presented the TORE (The Open Reality Experience) (Figure 28). This system was first installed in Lille University and is described as an edge-less CAVE capable of completely surrounding the user, while still providing considerable space for natural movement. Figure 28. The TORE system at Lille University, France. Images from Antycip Simulation. Although impressive and clearly beneficial in usability terms, CAVE-like systems have the main drawback of being highly expensive to acquire and to maintain. A CAVE or a power-wall that requires three projectors can cost around 150 000 € for the projection system alone, and the TORE system installed in Lille University required 20 high-resolution 3D projectors. Thus, contrary to what is starting to happen with HMDs, CAVE visualization systems are not on the spectrum of a consumer good, but are almost exclusively used for research applications, military training, driving simulators, and as part of permanent exhibitions on museums, planetariums, and scientific/educational centers. After revisiting some of the most important visual IVEs technology that were developed since Sutherland’s “ultimate display” we can safely state that the Sutherland challenge of a true VR environment has not been yet accomplished. Although, a CAVE-like environment is the best candidate for the visual component of a full VR environment since it deals with most of the usability problems of HMDs, this type
78 Visual distance perception – To understand how to accurately measure visual distance perception in large projection screens and how to use this as a performance measure that can informs us about the IVE degree of fidelity; Virtual sound location – To understand how auditory location might be affected by the use of specific sound output equipment. All the experiments described in this chapter were approved by the Direction Board of the Doctoral Program in Informatics, Department of Informatics, University of Minho. All participants gave their written informed consent. The experiments were conducted in accordance with the principles stated in the 1964 Declaration of Helsinki 24 . All these experiment’s implementations were made available (at https://github.com/CarlosCCG/Thesis2019) and were developed in Open-source software, thus consisting an ideal tool in order to foster the development of PCMs in the IVEs research and development community. 3.1. Audiovisual Perception of Synchrony 3.1.1. On the Role of Temporal Coincidence As we pointed out in Section 2.1, stimulating more than one human sense (also referred to as multimodality) is one of the essential characteristics of highly immersive virtual environments. Nevertheless, careless implementation of different modality actuation channels, without much consideration of how humans perceive multimodal scenes or without high control of the temporal relationship between the different output channels (e.g. image projection and sound output), can in fact have the opposite effect. To the more general phenomenon of perceiving bimodal stimulus as such, we call it Unity Assumption (UA) (mostly studied with visuo-tactile an audio-visual stimulation) [97]–[99]. The UA theory states that a stimulus is perceived as a multimodal audiovisual stimulus, for instance, “… if the two sensory modalities are providing information about a sensory situation that the observer has strong reasons to believe (not necessarily consciously) signifies a single (unitary) distal object or event.” [100]. Furthermore, the level of certainty with which we perceive an object or event as multimodal appears to 24 World Medical Association Declaration of Helsinki – Ethical Principles For Medical Research Involving Human Subjects. Can be downloaded here: https://www.wma.net/policies-post/wma-declaration-of-helsinki-ethical-principles-for-medical-research-involving-human-subjects/
79 be positively correlated with the number of amodal stimulus properties shared by the two sensory modalities. Amodal properties are physical properties that might be shared by different perceptual modalities, such as: spatial location ; temporal pattern ; size ; shape ; orientation ; intensity ; motion ; texture [100]. In particular, audio-visual stimuli appear to be best perceived as a unitary audiovisual stimulus when both modalities present [99]: Temporal coincidence; Spatial coincidence; Motion vector coincidence; Causal determination; In IVEs, these first two amodal properties – temporal coincidence and spatial coincidence – are of special relevance, particularly when we are dealing with systems with measurable end-to-end delays and targeted at applications requiring precise temporal and spatial fidelity. Thus, in this section we will explore the role of temporal coincidence, while in Section 3.2 we will explore the role of spatial coincidence. Humans are especially sensitive to audio-visual temporal relations, and thus in order to accurately implement an audiovisual IVE we must play particular attention to the aspect of audiovisual synchrony. Guidelines on how to measure, control, and set-up the appropriate temporal relation between auditory and visual stimuli of an IVE are dependent on knowledge about the fundamentals of human audiovisual synchrony perception – a phenomenon that is far from being clear and well understood [44]. Humans face intricate problems to perceive synchrony in audiovisual real-world events because of the relative timing of visual and auditory inputs. These problems are related to both physical and neural differences underlying sound and light propagation and human processing of these stimuli. When a natural audiovisual event occurs, the visual and auditory signals are synchronic at the origin, since they are caused by the same physical source and thus emitted at the same time. However, there are significant differences in the propagation time for light and sound, as sound takes about 3.4ms to travel 1 meter 25 and light travels approximately 299 792 meters in 1ms. Nevertheless, in our daily life and within a certain distance range, audiovisual stimuli are still perceived as synchronic at the source. Given the abovementioned sound delay (in relation to light) of about 3ms per meter traveled from the perceived 25 The standard value for the speed of sound (343 meters per second) is calculated as its value at 20°C and traveling in the elastic medium of air. Nevertheless, this value is dependent from both the temperature and the type of elastic medium through which the sound is been propagated.
80 object, it is not obvious how certain audiovisual stimuli are perceived by the user as an audiovisual unitary phenomenon. These constrains make the problem of audiovisual synchronization especially relevant for the development of audiovisual IVEs, mainly because common range distances from stimuli sources to observers are often large enough to create a considerable gap between the arrival times of visual and auditory signals in the real-world. Thus, when developing an audiovisual IVE, one should know if accurately modeling of this variable audio-image delay is relevant or not. Although physical realism would advise us to consider the slower propagation time of auditory stimuli, in terms of computational cost it would be unnecessary to do so if users are not sensitive to these delays. Moreover, implementing an audio delay based on the simulated distance of the audiovisual stimuli requires not only adaptation of the auditory and visual scene in the VR platform, but also complex measurements of each system (hardware) output channels. In order to know the audio-visual temporal relation of the output channels, we need to measure it externally, resorting to light and acoustic sensors, since the end-to-end delay from the video and the audio channels vary as a function of the graphic board, the sound card, the tracking system, or even the defined projection frame-rate [32]. Moreover, differences in propagation time are not the only aspect to consider in audiovisual synchronization. Humans also present differences at the level of transduction times of the human perceptual channels [101] [102], only this time with sound being transduced faster (~ 1 ms; see [103]) than light (~50 ms). However, as these neural temporal differences are constant, observers become adapted to this difference due to a long history of exposure to such “veridical” neural lags [104], due to a phenomenon called perceptual temporal recalibration [105]. Nevertheless, it remains to be explained how can we perceive audiovisual events as unitary when the audio and the visual streams will physically arrive to the observer at variable asynchronies as a function of the distance of presentation. Several studies have shown that audiovisual integration does not require temporal alignment between the visual and the auditory stimuli. We still perceive as synchronic visual and auditory stimuli that are not received or emitted at the same time (e.g. [106]–[112]). Notwithstanding, temporally mismatched stimuli can only be perceived as synchronic when keeping the onset difference between sound and image within certain temporal limits, termed Window of Temporal Integration (WTI) [44] [113]. In multisensory perception, this phenomenon can be defined as the range of temporal differences on the onset of two or more stimuli of different modalities where these are still best perceived as a unitary multisensory stimulus. As Vroomen and Keetels [44] pointed out, the main reason why signals from different sensory modalities are perceived as being synchronic despite these differences, is that the brain
81 judges as synchronic two stimulation streams that arrive within a certain amount of temporal disparity. Thus, another important piece of information for IVEs developers would be to know what are the limits of audiovisual WTI and how those relate with the distance of the audiovisual scene. By knowing this, they should aim at keeping systems end-to-end delays of each modality inside the values tolerated by the WTI. Research on this phenomenon has provided us with surprising findings. A large number of studies on audiovisual temporal alignment have found that we perceive stimuli from different modalities as being in maximal synchrony if the visual stimulus arrives at the observer shortly before the auditory stimulus (e. g., [106], [109], [110], [114]). This finding has been termed the vision-first bias [113] [44]. In a work that boosted the scientific discussion on the vision-first bias, [110] used a temporal-order judgment (TOJ) psychophysical task (see Figure 32) to assess the perceived temporal relation between the emission of a sound (a burst of white noise) and a brief light flash. The flashes of light were displayed from LEDs located at distances of 1, 5, 10, 20, 30, 40 and 50 m from the participant. The sound was always transmitted by headphones but was compared with the visual stimuli at different distances. In their paper published in Nature , Sugita and Suzuki reported that the Stimulus Onset Asynchrony (SOA) that provides the best perception of synchrony is always a positive one (i.e., sound lagging image) and most importantly, when the distance of visual stimuli increases, larger sound lags are observed at the Point of Subjective Simultaneity (PSS) (i.e., the temporal relation between sound and image which elicits the higher percentage of “synchronous” answers, see Figure 32). Their results are roughly consistent with the velocity of sound, at least up to 20 m of visual stimulus distance, and can be quite well predicted by a linear model based on this physical rule. Thus, and according to [110], it seems that the brain considers sound propagation velocity when judging synchrony, relying on distance information to compensate for the natural differences on the stimuli’s propagation velocity. Other studies have also pointed to a perceptual mechanism of compensation for differences in propagation velocity (e.g., [106], [109], [115]), further suggesting that we resynchronize the signals of an audiovisual event by shifting our PSS in the direction of the expected audio lag.
82 Figure 32. Graphically representation of the PSS and the WTI (corresponding to the Just Noticeable Difference – JND) on a Simultaneity Judgement (SJ) Task – where the question posed to the participant is of the type: “is the presented audiovisual stimulus synchronous regarding visual and auditory streams?” – and on a Temporal Order Judgment (TOJ) task – where the question posed to the participant is of the type: “what was the stimulus presented first, the visual or the auditory one?” Nonetheless, other researchers failed to observe such compensatory effect [116], [117]. Lewald and Guski [116] tried to replicate the findings of Sugita et. al [110] in a less artificial setting. Using the same kind of stimuli (sound bursts and LED flashes) but co-located, they found no distance compensation. In fact, the PSS shifted in the opposite direction. In this experiment, participants had the best perception of synchrony when auditory and visual signals were synchronic in their arrival at the observer’s sensorial receptors (i.e. sound leading image on onset). As Lewald and Guski pointed out, “this conclusion is in diametral opposition to the study of Sugita and Suzuki” (p. 121). According to these authors, there are some limitations in Sugita and Suzuki’s study that might have been affecting the ecology of the stimuli (i.e., ecology in this sense refers to the capacity for the stimuli to simulate the real world). Firstly, the sound stimuli were not co-located with the visual stimuli and consequently there was no auditory distance information. Secondly, the luminance of the visual stimuli was increased to compensate for the light intensity attenuation with distance and, by doing this, they kept the perceived stimuli’s luminance constant, thus providing a cue that is incongruent with the expected information on distance increment. Nevertheless, both the studies of [98] and [99] also present potential problems in the simulation of distance, namely: 1. By conducting the experiment in open-field, Lewald and Guski failed to provide the optimal conditions for auditory distance information. One of the most powerful auditory depth cues is
83 the ratio of energies of direct and reflected sounds [93], which is almost absent in open-field stimulation. Also, the most important depth cue in a situation of open-field – loudness – is frequently and erroneously perceived as the level of the sound itself in the absence of relative loudness cues, which can cause misjudgments of stimuli distance; 2. The use of artificial stimuli, such as flashes and beeps, eliminated two other relevant depth cues: familiar size of the visual stimuli and familiar loudness of the auditory stimuli. 3. Arnold and colleagues manipulated the angular size and velocity (i.e. retinal size and velocity) of their visual stimuli to ensure that the size and velocity of the stimuli appeared constant while distance increased. This poses the same problem of incongruent depth cues pointed out by Lewald [116] to the work of Sugita [110]. Additionally, there are limitations that are common to both groups of studies. Recent findings show that in the study of audiovisual synchrony, biological motion stimuli are preferable over rigid motion or stationary stimuli [107], [112]. For instance, the use of point-light displays with biological motion has been previously pointed out as an important factor in simultaneity judgment tasks. In a study where participants had established a baseline PSS to an audiovisual stimulus consisting of footage of a professional drummer playing a conga drum, Arrighi and collaborators found that the PSS was not different when the visual stimulus was a computerized abstraction of the drummer with the same movement (biological) presented in the footage. However, when the same artificial visual stimulus had an artificial motion pattern (constant velocity) the PSS was significantly different, even when the frequency was the same in both conditions [107]. Petrini, Holt, and Polick have found that a non-natural orientation of a point-light drummer can affect the simultaneity judgment of non-musical expert participants, thus showing that naturalistic representations are preferred in this kind of tasks [118]. 3.1.2. Assessing Audiovisual Synchrony in an IVE Considering the above controversy and taking into account the main critiques regarding previous studies on compensatory mechanisms, in this experiment we report findings that offer a clearer answer to the following questions:
84 1 How can we perceive an audiovisual event as such, if the audio and the visual signals physically arrive at different times to the observer? 2 Does distance play a role in a possible mechanism of perceptual compensation? 3 What are the levels of audiovisual asynchrony that users of IVEs can cope with without clearly perceiving an audiovisual mismatch? In order to answer these questions, we developed an audiovisual synchrony judgment task, capable of being performed in an IVE, and resorting to audiovisual biological motion stimuli. Biological stimuli such as PLW also allow for high levels of control and manipulation of critical parameters for visual and sound depth perception. Namely, angular and familiar size, angular velocity, elevation, intensity, contrast, and perspective for visual distance judgments; and sound pressure level, high frequency attenuation, and reflected sound for auditory distance judgments. Thus, in order to assess the existence of a mechanism that compensates for the differences in propagation velocity we presented the audiovisual biological stimulus at several distances in a controlled environment simulating the depth conditions of the real physical world. Additionally, by manipulating the number of depth cues presented in the stimuli we were able to assess if a perceptual mechanism that compensates for audiovisual asynchronies does exist and if it has a relation with distance of presentation of the audiovisual stimulus. In short, we expect to support the argument of perceptual compensation if we find a positive relation, close to the rule of physical propagation of sound, between the shift of the PSS (towards an audio lag) and the distance of the stimuli. We also expect this relation to be dependent on the number and quality of the depth cues. The results from this experiment will be discussed in light of their implications for the conceptualization of the mechanisms of simultaneity constancy and development of audiovisual IVEs. Method Participants: Four participants, aged 21-33 years old, underwent visual and auditory standard screening tests and had normal hearing and normal, or corrected to normal, vision. All were voluntary university students or researchers, and all gave written informed consent to participate in the study. Two had some background knowledge about the thematic of the study and the remaining two were naïve as to the purpose of the experiment.
85 Stimuli and Materials: The experimental tasks were performed in a darkened room at the Center for Computer Graphics in the University of Minho. For the stimuli presentation, we used a cluster of PCs with NVDIA® Quadro FX 4500. A 3chip DLP projector Christie Mirage S+4K with a resolution of 1400x1050 pixels and a frame rate of 60Hz was used for the projection of the visual stimuli. The area of projection was 2.80 m high and 2.10 m wide. The sound was presented using a computer with a Realtek Intel 8280 IBA sound card through a set of Etymotics ER-4B in-ear phones. Latencies between visual and auditory channels were measured and adjusted using a custom-built latency analyzer consisting of an Arm7 microprocessor coupled with light and sound sensors. We programed this experiment using the XML-based BioMoSe 26 language (see Example 2). BioMoSe ( Biological Motion Segmentation ) is a custom-made software, developed at the Laboratory for Visualization and Perception of the University of Minho, which uses the Open Graphics Library (OpenGL) to control and parametrize the displaying of the visual and auditory stimuli. As a virtualization platform, we used VR Juggler [119], an Open Source platform for virtual reality applications developed by the CruzNeira research group at Iowa University and thus specially tailored for controlled displaying of audiovisual stimuli in CAVE-like environments. VR Juggler is particularly useful for multiscreen applications, since it allows us to configure the correct partitioning between projection nodes and accurate edge blending in order to avoid bad continuity projection problems. Additional materials used in the implementation of this experiment can be found at: https://github.com/Carlos-CCG/Thesis2019/tree/master/ExpSynchrony 26 Extensible Markup Language
86 Example 2. Biomose declaration of the main settings of the experimental scenario (Top panel); XML declaration of the scene visual and auditory stimuli (Bottom panel). Three experimental conditions differing from each other in the number of depth cues were presented. The “audiovisual depth cues” condition presented visual and auditory depth cues coherent with the simulated presentation distance. In the “visual depth cues” condition only the visual depth cues varied coherently with the simulated presentation distance, but sound had no distance cues (anechoic Scene duration; Inter-stimulus interval (ISI); Scenes presentation sequence; Background color of the ISI; Responses codification; Video dumping; Camera initial position and settings; Virtual Room Acoustic and visual stimuli
87 sound without room-acoustic information). The “reduced depth cues” condition presented both visual and auditory depth cues impoverished. A) Auditory Stimuli. In the “audiovisual depth cues” condition the auditory stimulus consisted in a binaural sound recorded using a Brüel&Kjaer® head and torso simulator Type 412SC. An anechoic step sound was emitted trough a Brüel&Kjaer® omnidirectional Loudspeaker Type 4295 inside a sports pavilion 24 m wide, 50 m long and 11 m high. The step sound was recorded at distances of 10, 15, 20, 25, 30, and 35 m from the head and torso simulator, located 6 m from one end of the sports pavilion (see Figure 33). The anechoic step sound came from the database of controlled recordings from the College of Charlston [120] and corresponds to the sound of a male walking over a wooden floor and taking one step. In the two remaining conditions, the auditory stimulus was an auralized sound of the above-referred anechoic recording, with directional cues matching the visual stimulus, but in free field (thus without any room-acoustic information) and without air attenuation with distance. Figure 33. Set-up for acquisition of the step sounds binaural recordings. We used an omnidirectional speaker (Bruel&Kjaer Type 4295) to emit the anechoic step-sounds and an Head & Torso simulator (Bruel&Kjaer) to record the binaural stimuli. B) Visual Stimuli. Different visual stimuli according to the experimental condition were presented to participants. In both the “audiovisual depth cues” and the “visual depth cues” conditions the visual stimuli were PointLight Walkers (PLW) moving in a front-parallel plane to the observer and taking one step aligned with the center of projection. The PLW was walking inside a simulated room with the same dimensions of the sports pavilion referred before. The PLW was composed of 13 white dots (54 cd/m2) that moved against a black background (0.4 cd/m2) and was generated in the Laboratory of Visualization and Perception (University of Minho) from motion data captured using a Vicon® system with 6 MX F20 cameras and a
94 Figure 37. Proportion of “synchronized” answers as a function of the SOA for a data pool of distances of 10, 15, 20, 25, 30 and 35 m in the three experimental conditions. Each proportion was calculated using a pooled data of 160 answers from the 4 participants. Panel A shows the answer distributions for distances of 10, 20, and 30 meters; panel B shows the answer distributions for distances of 15, 25, and 35 meters. A fit of a Gaussian function was performed in order to get the PSS and WTI values for each distance of stimulation. An arrow indicates the PSS for each stimulus distance. In both the “audiovisual depth cues” and “visual depth cues” conditions, the peak of the Gaussian curve progressively moves towards a higher sound delay as the stimulus distance increases. This increment is generally lower in the “visual depth cues” condition, where there is a difference of only 20ms between the lowest (in the 15 m presentation) and the highest (in the 35 m presentation) PSS, while in the pooled data of the “audiovisual depth cues” condition this difference is of 60ms (with the lowest PSS at the 10 m presentation and the highest at 35 m). However, in the “reduced depth cues” condition the peak of the Gaussian curve hardly moves from one distance to another (especially in the odd group of distances) and when it moves, it does so in the direction of a decrement in sound delay. This was the only condition where several PSSs with a value close to zero (see distances 25 and 30 m) were found. A One-way ANOVA shows significant differences between conditions regarding the PSS (F (2, 15) = 11.9, p < .01), and the Scheffé Post-hoc tests revealed that these differences are significant when we compare the PSSs in the “reduced depth cues” condition with the PSSs in the “visual depth cues” condition (p < .01) and with the PSSs in the “audiovisual depth cues” condition (p < .05).
95 Figure 38, plots the increment in the PSS as a function of the increment of the stimulus distance regarding the first distance of presentation for each condition. Here, we can compare the way the PSS changes across conditions with a model for internal compensation of the slower propagation velocity of sound. A linear function was fitted to the PSSs obtained in the three conditions. A good adjustment (r2 = .94, F (1,4) = 84.8, p < .01) on the fitting of the function y = 2.6x to the results in the “audiovisual depth cues” condition was obtained. In the “visual depth” condition the best linear function was y = 0.9x , with a good fit (r2 = .93; F (1,4) = 63.84, p < .01). Similarly, a linear function was fitted to the PSSs obtained in the “reduced depth cues” condition ( y = -.75x ), but this time with only with a rough adjustment (r2 = .5, F (1,4) = 5.77, p < .1). Figure 38. Increment in the PSS as a function of increment in the stimulus distance, regarding the first distance of presentation. Black squares correspond to the theoretical values predicted by a mechanism that compensates for differences in propagation velocity. Red dots are the PSS found for each distance in the “AV depth cues” condition. Blue triangles are the PSS found for each distance in the “visual depth cues” condition and orange triangles are the PSS found for each distance in the “reduced depth cues” condition. A fit of a linear function was performed in each group of data. Panel A shows the graph for the pooled data, and Panel B show the same graph for the individual data. Figure 37 and Figure 38 clearly show that the conditions “audiovisual depth cues” and “visual depth cues” present different tendencies when compared with those from the condition “reduced depth cues”. While the PSS from the “audiovisual depth cues” and “visual depth cues” condition increases with distance, the PSS from the “reduced depth cues” condition seems to slightly decrease with distance. In fact, correlation tests show that the PSS in the “audiovisual depth cues” condition is positively and significantly correlated with distance (rsp = 1, p < .001) and the same is true for the PSS in the “visual depth cues” condition (rsp = .943, p < .01). On the other hand, PSS in the “reduced depth cues” condition
96 is just marginally correlated with distance (rsp = -.77, p < .1), and in the opposite sense: higher PSSs are associated with lower distances. Discussion: Our study aimed at verifying if a perceptual mechanism that compensates for audiovisual asynchronies as a function of the presentation distance exists, and if so, how can it be expressed in a PCM capable of guiding IVE developers in the implementation of accurate aural-visual synchronization. To do so, we presented audiovisual biological stimuli at several distances and with a wide range of stimulus onset asynchronies in an IVE. Crucially for this study, we manipulated the amount of distance cues in three experimental conditions: audiovisual depth cues, visual depth cues and reduced depth cues. Results from the first two experimental conditions revealed a systematic shift of the point of subjective simultaneity in the direction of greater audio lags with greater distance. Therefore, results support the existence of an internal compensation mechanism for varying stimuli delays with distance. Interestingly, compensatory evidences were not found in the “reduced depth cues” condition. Indeed, that condition was so impoverished regarding distance perception cues that only visual angular velocity – a relatively poor distance cue – was available. Therefore, the internal compensation mechanism appears to be dependent on the amount and quality of depth cues simulated in the IVE. This interpretation is further supported by the fact that in both conditions where evidence for such mechanism was found, the steepness of the function was not the same. In the “audiovisual depth cues” condition there was a steeper function, closer to that expected from the actual physical audio-visual delays. We conclude that the proposed compensation mechanism might not work in an all-or-nothing way, but there might exist intermediate levels of compensation. However, at this stage we are not able to provide an exact account on what is guiding a compensatory pattern of response in some conditions. It could be argued that such a mechanism emerged due to the enhanced realism of the stimuli, or due to the causal relation between visual and auditory stimuli. In any case, we can assume that both factors contributed to a greater perceived unity between the multisensory signals. Future studies should focus on the relative effects and weights of different cues, different stimuli, and different settings and, with this in mind, a critical review of previous studies in this field might reveal that data obtained so far was not necessarily contradictory, but mostly the result of different experimental setups. Despite all these limitations, some hypotheses on what is guiding compensation can be further discussed, considering our results. First of all, assuming that our stimuli were perceived as co-localized in the “audiovisual depth cues” condition, co-localization of auditory and visual stimuli seems to be an
97 important factor to exhibit compensation for the relatively slow speed of sound. However, co-localization does not seem mandatory for compensation of sound propagation velocity. Alais and Carlile have found evidence of compensation for sound propagation velocity by providing auditory depth cues while keeping the visual stimuli at a fixed distance [106]. In their study the ratio of direct-to-reverberant energy was used as an auditory depth cue, but the visual stimulus was fixed at 57 cm from the observer and primarily used as a reference point in time. Furthermore, their results show that the compensation effect relies on the robustness of the auditory depth cues. Thus, the authors concluded that reliable auditory depth cues together with a task-relevant situation are sufficient in order to activate compensation for the sound propagation velocity. These results shift the focus from co-localization to specific depth cues, when we try to uncover the reason for having compensation in some experimental conditions. Our results agree with the idea that a powerful auditory depth cue is necessary in order to get evidence for sound propagation velocity compensation: only when we add the binaural recordings of sound steps do we get clear evidence for such compensation. Our work was carefully designed taking into account several problems pointed out in previous studies. More specifically, we avoided the use of simple and stationary stimuli with poor distance cues. We also used audiovisual stimuli with a causal relation between them, as well as familiar stimuli. By using a PLW and real step sounds, we provided the participant with a type of stimulus that occurs in everyday life. Indeed, the finding of compensation for sound propagation velocity with such stimulus might be evidence that, apart from relying in some depth cues, this mechanism also uses knowledge from previous experience. In real-life situations, we are always exposed to a certain audio delay. Despite this delay being highly variable, a long time of exposure to it could lead to some kind of temporal recalibration. This approach is in accordance with the work of Heron and collaborators where, after a brief phase of exposure to the natural sound lag of a distant audiovisual stimulus, participants shifted lag expectations [121]. Moreover, the work of Virsu and collaborators has shown that simultaneity constancy can be learned in natural interactions with the environment and without explicit feedback [122]. Effects of simultaneity constancy appear to be long-lasting and modality specific. In addition, studies comparing the perception of synchrony in adults and infants have found that the thresholds for asynchrony detection are modified as we get older [123], further showing that synchrony perception is affected by our history of exposure to certain audio delays. For IVEs developers, these findings highlight the importance of including naturalistic sound delays in highly immersive audiovisual virtual scenes. As we can see from our data, IVEs users can judge audiovisual synchrony and are sensitive to small variations in the audio-visual temporal relation. Thus,
98 the rule of delaying the sound output about 3ms per meter, depending on the distance of the audiovisual stimulus, should be applied. Moreover, and by looking at the size of the WTIs from our pool of participants, we can understand how important it is to accurately measure end-to-end delays of IVEs and to struggle to keep these delays at a minimum level. On average WTIs for our participants where about 125ms (σ = 48ms), which means that this is the range of asynchronies around the PSS at which the participant will still respond with at least 75% of “synchronous” answer. It is important to note though, that this range is evenly distributed around the PSS and therefore is moveable, meaning that due to vision-first bias and effects of distance on the PSS this WTI shifts towards being increasingly tolerable to sound lag values and increasingly intolerant to vision lag values. In conclusion, from this work we can derive three important design guidelines for IVEs developers: 1. End-to-end delays of visual and auditory output should be measured and controlled in a way that a visual and an auditory stimulus modeled to be perceived synchronously, should not lag more than 62.5ms (half of the average WTIs value); 2. On an audiovisual synchronic scene, the vision-first bias should be taken into consideration and sound should never precede image; 3. The natural occurring sound lag should be taken into account and modeled when developing the IVEs. The rule for a buffer responsible for defining the sound delay should be: Equation 10. 𝑦=𝑦0−𝑘+ (3.4𝑥) , where 𝑦0 is the measured difference between the audio and the visual end-to-end delay of a specific IVEs, k is a constant delay inserted to compensate for this difference, and x is the distance-to-user of the simulated audiovisual event. All variables are in milliseconds. In this work, we have been looking into how temporal disparity affects perception of audiovisual stimuli in IVEs. In the next experiment, we will look into the role of spatial disparity and its effects on audiovisual perception in IVEs.
99 3.2. Audiovisual Unity Assumption 3.2.1. On the Role of Spatial Coincidence The design and equipment constraints of most IVEs require output of the different modality stimulus to be transmitted through non-co-localized means. For example, in a projection-based audiovisual IVE, the visual stimulus is projected on a screen while the auditory stimulus is conveyed through headphones or a set of speakers. This is particularly relevant if we consider that often the goal of an IVE is to present a unitary audiovisual stimulus at a given depth and position in a 3D space. As we saw in Section 3.1, this goal can be achieved by using convincing simulations of both visual and auditory depth cues. Binocular and pictorial depth cues allow users to perceive visual stimuli beyond the projection screen; while binaural cues and room acoustics allow users to perceive sound in space. Thus, in a perfect audiovisual virtual system, the projection screen fades away and gives room to perfectly defined objects in distance and the headphones wear off as the sound is perceived to be outputted from an external source in space and not from the apparatus in your head. The problem though is that, as we saw in Section 2.6, perfect virtualizing systems do not yet exist. While congruent audiovisual depth cues are simpler to manage in systems without user tracking, consider again the problem of end-to-end latency in interactive IVEs. Both visual and auditory end-to-end latency will result not just in temporal mismatch but also in physical displacement of the virtualized audiovisual object relative to the users’ head. Moreover, if the end-to-end delay is different for the visual and the audio output, users will also experience physical displacement between the relative position of the visual and the auditory stimulus. Thus, it should be quite relevant for IVEs developers to know and to be able to accurately measure the just noticeable amount of physical disparity between the visual and the auditory stimuli – i.e., the displacement value above which an audiovisual stimulus stops being perceived as such and starts to be perceived as two separated stimuli, one visual and one auditory. This value would give developers information about how manageable their end-to-end latency is (and if it is interfering with the understanding of the audiovisual scene); while a method to accurately measure this value, would give developers a technique to test mitigation measures. One of the main perceptual effects developers rely on in order to deal with audio-visual spatial incongruence is simply trusting on a well-known human perceptual phenomenon called the Ventriloquism Effect (VE) [124], [125]. The VE, also known as capture effect , happens when sound is perceived as coming from the location of a visual stimulus, despite being spatially displaced. The term is based on the trick used to give the impression of a dummy speaking, when in reality is its operator emitting the sound.
100 The VE is constantly present when we are watching a movie either at home or at the cinema, as sound is coming not from the actors’ location on the screen but from speakers located elsewhere. In real setting environments, the limits of VE were quantified to be around 20º azimuth, depending on the direction of the spatial disparity – the limit is smaller for displacements in the frontal area (around the 0º azimuth) and higher for displacements in medial areas (around the 45º azimuth) [126]. Although we can trust in VE to deal with most of the problem of audio-visual displacement in many audiovisual applications, we should not rely on VE if we want to design an IVE where precise auditory location is required. Kytö and collaborators [127], described an AR application where stereoscopic visual and binaural auditory stimuli where used in order to signalize points of interest (POI) in a navigation system. Because these POIs would have to be precise in the tens of angular units and because they needed to use acoustic POIs to direct user attention to portions of the virtual environment outside the visual FoV, they needed to know the limits of VE in order to effectively use these acoustic POIs (i.e., the POI sound should not be captured by anything being visual displayed in the scene as that could trigger an error in the users navigation). In order to accurately measure the limits of VE, Kytö and colleagues used an Oculus Rift AR system and Pure Data for audio rendering using HRTFs from the CIPIC database [128] for auralizing the auditory stimuli. They presented the participants with multiple possible combinations of auditory-visual stimulus displacement (see Figure 39) and asked participants to respond to the question: “Is the sound (audio POI) coming from the speech bubble (visual POI)?” . Figure 39. Set-up of Kytö et al. experiment. Top figure, shows all the possible combinations (in angular position in azimuth) of the audiovisual stimulus presented to the participant [127].
101 Using the Just Noticeable Difference (JND, see Section 3.1.1) as the limit at which 75% of the participant’s answers reported separation between visual and auditory stimulus, Kytö and colleagues found values between 32º and 45º, depending on the visual angle (mean JND was 38º). Similar to results from [126], participants needed further sound displacement to judge audio-visual separation for visual presentations around the 45º azimuth, than for visual presentation around the 0º azimuth location. Thus, the main guideline to extract from [127] is that, if we want to clearly spatially separate the visual from the auditory stimulus in an IVE, we should do so by keeping a spatial disparity greater than 32º-45º depending on visual direction. Of course, this interesting result can be interpreted from another perspective, which is: if the spatial disparity between visual and auditory stimulus is below 15º (average value of the 25% level of answers “stimulus are separated”), they are most likely to be perceived as an unitary audiovisual stimulus. As we saw with the work of Kytö and colleagues [127] AR navigation systems are one of the applications requiring precise spatial and temporal fidelity (see also [129] for an AR application on acoustic navigation). Other applications, such as interactive virtual concert halls or multi avatar VR environments, would greatly benefit of a clear knowledge of the spatial limits of audiovisual UA. 3.2.2. Assessing the Spatial Limits of Unity Assumption Thus, in this experiment we wanted to device a way of accurately measure the limits of UA in an IVE. Differently from [127], we were interested in measuring audiovisual unity assumption in a set where multiple visual stimulus compete for the correct attribution of the auditory source. This is closer to applications such as virtual concert halls or multi avatar environments, where the user should correctly bind the sound to each agent in the virtual scene (e.g., a music sound to an instrument in an acoustic setting, or step-sounds to avatar in a simulated room). Our main goal with this experiment is to accurately define absolute spatial thresholds for a correct unity judgment between an auditory stimulus and its correspondent visual stimulus in a conflicting task. Furthermore, we also wanted to explore if those limits are dependent or independent from stimuli distance.
102 Experiment 1 – Method Participants: Four participants, aged 19-30 years old, underwent visual and auditory standard screening tests and had normal hearing and normal vision. All were voluntary participants, and all gave written informed consent to participate in the study. Stimuli and Materials: The experimental tasks were performed in the set-up described in Section 3.1.2. Additional materials used in the implementation of this experiment can be found at: https://github.com/Carlos-CCG/Thesis2019/tree/master/ExpUnityAssumption A) Auditory Stimuli. The auditory stimuli were step sounds from the database of controlled recordings from the College of Charlston [Marcell2000]. They correspond to the sound of a male walking over a wooden floor and taking two steps at exactly the same velocity as the visual stimuli. These sounds were auralized as free field by a MATLAB routine with HRTFs from the CPIC database 28 [128]. With this custom MATLAB routine, a binaural sound would be generated through the application of the correct ILD and ITD cues (based on a geometric model describing sound source position in the 3D space) and the nearest HRTF would be choose from the CPIC database in order to perform the last step of head-related frequency modulation. A set of 20 different position sounds were generated and used in this experiment (from azimuth -45º to azimuth 45º, in steps of 5º plus two steps auralized at 3º and -3º azimuth). B) Visual Stimuli. Visual stimuli were two PLWs (similar to the ones described in Section 3.1.2) located at 20 meters in a front-parallel plane to the observer. The PLWs were back-to-back and equidistant from the center of the projection screen. Throughout the duration of a trial, each PLW would take two steps, but with the translational component removed. C) Auditory and Visual Stimuli Relation. The audio-visual scene consisted in the presentation of the two PLWs at a given separation degree and the output of sound co-located and synchronized (a sound delay of 68 ms was added, based on the 20 m of distance of the visual stimuli) with just one of the PLWs. Thus, the angular distance between the PLWs would randomly vary from trial to trial on a range of 45º horizontally, in steps of 5º plus a stimulus with just 3º of separation (Figure 40). 28 Available at https://www.ece.ucdavis.edu/cipic/spatial-sound/hrtf-data/
103 Figure 40. A graphical representation of the auditory and visual stimuli relation in this experiment. There were a total of 20 different audiovisual stimuli (10 angular distances between PLWs x 2 sound bindings, left PLW or Right PLW). Each stimulus was presented 40 times, with a total duration of 4.5s (3s of stimulus presentation + 1.5s of inter-stimulus interval). Thus, the total experiment took 1 hour to be complete for each participant (four sessions of 25 min. each). Procedure: Participants were first briefed on the purpose of the experiment and then proceeded to seat in a position aligned with the center of a 2.80 x 2.10m projection screen. The participant’s head rested on a chin-rest to prevent lateral movements. The task participants were asked to perform was: “Please pay attention to the following audiovisual scenario. You will see to PLWs at a certain lateral distance from each other and taking two steps. You will hear the sounds of two steps synchronized and co-localized with one of the PLWs. The sound will be co-localized with either the PLW from the left or the PLW from the right, please if you think the sound is from the left PLW press the left button of the mouse, if you think the sound is from the right PLW press the right button of the mouse.” Results: Figure 41, shows the response distribution for each participant. A Maximum Likelihood Estimation (MLE) function was fitted to the data of each participant, using RStudio and a package called ‘quickpsy’ [130]. The MLE is based on a Logistic function of the type: JND will correspond to the minimum distance necessary to perform the correct audiovisual binding at a 75% level. 3º, 5º, 10º, 15º, 20º, 25º, 30º, 35º, 40º, 45º
110 A B C Figure 44. Pooled data for the “sound only” condition and for the audiovisual conditions at 20m (A), 10m (B), and 10m same angular disparity (C). The graphs show percentages of responses “right” as a function of the PLWs’ degree of separation. We can observe on Figure 44 that there is not a clear relation between JNDs and distance of presentation. Values vary between 9.77º and 14.03º for presentations at 10 meters, and 12.78º for presentation at 20 meters (replication of experiment 1, which obtained a JND of 9.51º). Nevertheless, one interesting result is the lower JNDs for audiovisual presentations (however, see 20 meters condition), meaning that spatial coincidence between two signals of different modalities might help to the precision on this task. This is a result that might indicate a benefit of multisensory presentation for localization tasks, and wich might be further investigated in future research. Discussion: In both experiments, we found smaller JNDs, in the audio-visual tasks, than those found in both [126] and [127]. As we referred above, this can be an indication that having conflict stimuli of one modality (e.g. visual) competing for the binding of the second modality (e.g. auditory) helps performing a correct binding in a UA judgment task. Moreover, it also appears to enhance localization PSE = -1.84 JND = 17.98 PSE = -5.41 JND = 9.77 PSE = -4.31 JND = 14.03 PSE = 2.30 JND = 20.56 PSS = -3.54 JND = 12.78 PSS = 1.87 JND = 12.02
111 performance when compared with unimodal auditory tasks, at least at shorter distances from the observer (“sound only” conditions in experiment 2). The two experiments taken together allow us to produce additional guidelines for IVEs developers regarding spatial simulation of audiovisual stimuli. [127] advised developers to use a maximum angular disparity of 15º if they want to convey the perception of one audiovisual POI in an AR scene, or conversely to use angular disparities between 32º and 45º if the goal is to convey the perception of two separate stimuli, one visual and one auditory. From our experiments, we can conclude that in order to avoid wrongfully UA attributions, IVEs developers should: 1. Avoid audio-visual spatial mismatches that render the audio at a position closer than 9.5º - 12º azimuth from a conflicting visual stimulus; 2. When possible, co-localized visual stimuli should be used in order to improve spatial discriminability of auditory stimuli; In the next experiments, we will explore some phenomena of unimodal perception in IVEs and we will continue to demonstrate examples of psychophysics application as a method to study and improve perception in these environments. 3.3. Perception of Visual Distances in Virtual Environments 3.3.1. On the Need to Access Basic Perceptual Phenomenon to Validate Simulators for Studying Complex Perception-Action Behavior One of the most valuable applications of IVEs is as a platform for controlled experimentation or training. Using IVEs removes the risks inherently associated to ecological scenarios and brings the advantages of controllability to environments that are, perception wise, close enough to the real-world scenario that we may say they have ecological validity . If IVEs were capable of producing an exact replication of the real-world event with action-perception correspondence, we could refer to the IVEs as in a state of stimulus fidelity [133]. However, both technological limitations of the immersive system and the user’s prior knowledge of physically being in a place other than the one conveyed by the simulation, make impossible for contemporary immersive virtual systems to reach such a state. As a result, IVEs
112 users often notice differences between the simulation and the simulated environment, also called nonidentities . Accordingly, IVEs developers should always perform two types of control activity when designing a new virtual environment for any interactive application. The first control activity consists in comparing real-world scenarios and IVEs, to assess if the participants are noticing non-identities (assessment of the IVEs stimulus fidelity). The second control activity consists in assessing if the detected non-identities impair performances, i.e., whether the performances in an IVE are similar to the participant’s performances in real-world scenarios (assessment of the IVE action fidelity , also known as face validity [133], [134]). Psychophysical experimentations offer a quantifiable evaluation of human performances in perception-action related tasks, thus enabling comparisons with user’s performances in the real world and consequently making possible to access the user’s identification of IVEs non-identities and the IVEs action fidelity [135]–[137]. Furthermore, another way of looking into the action fidelity of an IVE is to try a replication of known real-world perceptual phenomena in an IVE. These control activities become even more important as today IVEs are being widely employed as training platforms as well as in a wide variety of psychological experimentation such as on visual [138] and auditory [139], [140] perception and action, spatial cognition [141], and even on social interaction [142], [143]. In the realm of controlled experimentation, some IVEs are being developed to study perception-action behaviors as complex as crossing the road (Figure 45). Before performing control experimentation to study complex perception-action behaviors, an IVE developer must have means to evaluate the participants’ perceptual mechanisms wich are crucial to the experimental goals. For instance, in the case of using pedestrian simulators like the ones in Figure 45 to study road crossing behavior – where the pedestrian has to accurately judge the distance of an approaching vehicle, calculate the time it would take him to get to the other side of the road, and judge the risk of a collision – the experimenter has to guarantee that the IVE is conveying an accurate sense of distance perception . In fact, accurate distance perception as since long being regarded as critical for interaction and navigation in IVEs [136], [144].
113 A B Figure 45. Two pedestrian simulators, that resort to projection-based visual IVEs, currently being used in transportation research. Panel A, is the University of Leeds Highly Immersive Kinematic Experimental Research (HIKER) simulator, Image from the University of Leeds. Panel B, is the pedestrian simulator located at the Center for Computer Graphics, Guimarães and developed in the context of the ANalysis of PEdestrians Behaviour (ANPEB) project. In the present study, we used a psychophysical experiment that we thought is ideal to evaluate stimulus fidelity and action fidelity of an IVE used in a distance perception task. We used a psychophysical task named Frontal Matching Distance Task (FMDT) [145] in order to evaluate participant’s distance perception. Comparable measurements of distance perception in an IVE and a real-world scenario are particularly appropriate to access both the stimulus and the action fidelity of an IVE. It is known that humans have a tendency to underestimate egocentric distances in a FMDT [145]–[147]. Experiments using a range of distances that goes from 5 meters up to 25 meters revealed a pattern of response showing that egocentric distance is perceived as compressed, but at a nearly constant ratio [145]. Gathering data from a FMDT in a real-world scenario and in an IVE with different conditions varying the level of simulation realism allows to check for perception of non-identities in the simulated environment and, furthermore, will provide new insight into the importance of photorealism in IVEs (see [148], for an
114 account of the debate). Moreover, looking at patterns of progressively underestimation with increasing distance will show if participants are sensitive to some perceptual effects similar to those that affect observers when in interaction with real-world environments. 3.3.2. Measuring Distance Perception: A Comparison Between Real-world and Simulated Scenarios Methodology The FMDT was first thought as an egocentric distance perception task by Li and collaborators [145] (although similar methodologies were already presented in [149] and [150]). Results on egocentric distance estimation tasks have been revealing that egocentric distance is normally perceived fairly linearly, however far from accurate. The typical pattern of response, on this type of task, is an underestimation of distance by a nearly constant ratio [145]. The FMDT is a particularly suitable task to assess stimuli fidelity of an IVE, because it only requires the presentation of two marks with a given distance between them on a frontal-parallel to the observer plane. Thus, transposition of the FMDT from the real world to an IVE is fairly easy and the data obtained in the latter environment is directly comparable with the data obtained in a real-world scenario. Participants: Five participants, aged 24-29 years old, underwent visual standard screening tests and had normal, or corrected to normal, vision. All were voluntary university students or researchers and all gave written informed consent to participate in the study. All the five participants were naïve as to the purpose of the experiment. Stimuli and Materials: 1. Real-World Scenario (RWS): The experimental task in the RWS condition was performed in a grassy open-field located in Campus de Azurém at the University of Minho (see Figure 46). Two collaborators were used as marks and the participant was standing facing one of them. Two sisal yarns were laid and stretch on the ground to form perpendicular lines. The participant was instructed to walk on top of the sisal yarn in order to ensure their perpendicularity to the marks’ plane. The sisal yarns had different lengths so that its end could not be used as a distance cue. The distances between the two collaborators used as marks and between the participant and the facing mark were measured using a Bosch™ DLR130k laser measure.
115 Figure 46. FMDT in the RWS condition. Participant performing the FMDT in an outdoor location at the University of Minho. 2. Immersive Virtual Environment (IVE): The experimental task in the IVE conditions was performed in a darkened room at the Centro de Computação Gráfica in the Campus of Azurém, University of Minho. For the stimuli presentation we used a cluster of PCs with NVDIA® Quadro K5000 graphic boards and the virtual environment was designed in the Blender software with projection conFigd and computed through the Blender VR software. Two 3chip DLP projectors Christie Mirage S+4K and one DS+6K-M SXGA+ DLP Christie projector with a resolution of 1400x1050 pixels and a frame rate of 120Hz were used for the stereoscopic projection (using active stereoscopy) of the visual stimuli. The area of projection was a power-wall like screen of 2.80 m high and 9 m wide. A Vicon® motion capture system composed of six cameras was used to track the participants head position and update the projection accordingly (see Figure 47). Two experimental conditions were presented, differing from each other in the level of photorealism of the virtual environment. Figure 47. FMDT in the IVE conditions. Participant performing the FMDT in the IVEPH condition. The structure with retroreflective markers on the participant’s head is used in order to track participant’s head position and rotation.
116 The “Photorealistic Immersive Virtual Environment (IVEPH)” condition presented a similar environment to the one in the RWS condition, with two avatars as marks located in a grass field with trees at the back (far enough to preclude their use as a relative size cues). Using a Photo Research® Photometer/Colorimeter Model PR 655 that outputs values of color space and luminance, the colors of the grass and the sky in the IEPH condition were adjusted to match the colors of this elements in the RWS condition. A model of atmospheric scattering was use in order to simulate aerial perspective (see Figure 48 – A). A B Figure 48. IVEs rendition of the FMDT. (A) Depiction of the virtual environment used in the IVEPH condition (B) Depiction of the virtual environment used in the IVENPH condition. The “Non-photorealistic Immersive Environment (IVENPH)” presented the same environment as in the IVEPH, however the colors of all the presented textures suffered a monochromatic negative transformation (see Figure 48 – B). Procedure: In each trial, the participant was presented with two marks (two persons in the RWS condition and two avatars in the IVEPH and IVENPH conditions) in a given fronto-parallel plane (L-shape arrangement, see Figure 49.The participant’s task was to position him(her)self at a distance from mark 1 equal to the distance separating marks 1 and 2.
117 Figure 49. Schematic representation of the FMDT. Participants had to move towards or recede from Mark1 in order to equalize their egocentric distance to Mark 1 with the distance between Marks. The participant’s initial distance to mark 1 and the distances between mark 1 and mark 2 varied across trials. Three initial participant-mark 1 distances (8, 14, and 20 meters) and nine mark 1-mark 2 distances were used (5 to 21 meters by step of 2 m). The participants performed 2 sessions of 27 trials (3 participant’s initial distances x 9 distances between marks) in each environment condition, all presented in a random order. At the beginning of each experimental session, the following instructions were given: “You are going to participate in a study on distance perception. In this task you will have to judge the distance between two persons, the ones located in front of you, and adapt your distance to the person directly in front of you so that it matches the distance between the two persons. You can move towards or recede from the person in front of you as many times as you wish, however you cannot move laterally.” In the RWS, participants moved forwards and backward by walking and indicated to the researcher when they reached the appropriate location and the researcher then measured the participant-mark 1 distance. In the IVEs conditions participants moved forwards and backward, by moving themselves in the space available in front of the projection screen and by using respectively the right and left button of a mouse. They pressed the wheel button when they were in the correct position and their location was saved by the computer before starting the next trial. The virtual FMDT was implemented in Blender (release 2.78c) [BlenderFoundation2018], an opensource 3D computer graphics software, written in C, C++, and Python, with an integrated game engine called Blender Game.
118 Commands and Logs Using Blender’s game logic we implemented the controls for starting the experiment, controlling the movement of the participant’s position, and confirm the participant’s final position in each trial (Figure 50). Figure 50. A simplified game logic for the virtual FMDT. UP and DOWN keys or the LEFT or RIGHT mouse buttons would, respectively, approach or recede participant’s camera position in relation to Mark 1. Thus, the participant could combine physically approaching to Mark 1 (by walking towards the projection screen) with the use of the game commands to walk the last meters in order to perform the FMDT. The camera’s height, rotation, and distance to origin is defined through motion capture data coming from the Vicon® system, but a camera offset locked in the Z axis (use as depth) is possible using the above mentioned commands for approach or recede (see red circles in Figure 50). The final distance-to-mark1, takes into account tracked position and camera offset in the Z axis. A Python script runs every time the participant confirms the final trial answer by pressing the middle button on the mouse. The script logs the participant final position and defines the next trial, as follows: INITIALIZE an array with all possible positions between marks [10 positions] INITIALIZE an array with all possible initial positions participant-to-mark [3 positions] IF Position Matrix is not created: INITIALIZE a Matrix with all possible position combinations INITIALIZE an array with a SHUFFLE of all possible combinations
119 ELSE Save trial number, distance between marks, and participant position POP one trial from the Positions Matrix ASSIGN a position from the position Matrix to each mark PLUS a new participant initial distance Avatars The avatars consisted in two identical female characters in a normal standing position. The avatars were generated using the Make Human software (version 1.1.1) [151] and later rigged in Blender using a 30 bones armature, allowing to change their final pose. This project can be downloaded at: https://github.com/Carlos-CCG/Thesis2019/tree/master/ExpDist Results: The individual analyses showed similar response patterns across participants in the FMDT (i.e., the condition that showed the highest and lowest mean error was the same for all participants). Therefore, individual data were pooled and mean values were analyzed using regression analysis. Figure 51 shows the fit of linear functions to the data pooled as a function of the experimental condition. Data distribution, grouped by condition, conformed well to the linear fittings (mean Pearson’s r=.99). Due to a general underestimation of egocentric distance, participants placed themselves too far from mark 1 when matching the mark1-mark2 distance, in all conditions. However, there was a significant effect of the condition on the FMDT mean error (Wilks’ Lambda = .38, F(2,134)= 110.8, p<.01), due to the fact that the mean error in the RWS condition (2.3m), was lower than the mean error in the IVEPH (5 m), and IVENPH (6 m) condition (both p<.01). The two latter conditions were also different (p<.01). Figure 51. Final participant’s distance as a function of the distance between marks. The black dashed line represents veridical performance (i.e. no error in the FMDT). Red dots are FMDT mean results for the RWS condition. Blue triangles are FMDT mean results for the IVEPH condition. Green squares are FMDT mean results for the IENPH condition. Mean Error on the FMDT: 2.3 m Δ = 1.10 Y-int = 0.86 Mean Error on the FMDT: 5 m Δ = 1.23 Y-int = 1.90 Mean Error on the FMDT: 6 m Δ = 1.22 Y-int = 3.04
126 Figure 54. Polar graphics with the absolute mean error as a function of the stimuli’s position. As we can see from Figure 54, the localization errors are higher in intermediate azimuths and lower on the ear plane and on frontal regions. This pattern of response is present with both equipment; however, there are globally lower errors in the headphones condition and that is even more clearly observed in the extreme presentations (ear plane and frontal regions). In a second analysis, we grouped the participants by audio output device used during training sessions. In doing this, we wanted to understand how congruency regarding devices used on training and experimental sessions might affect performance on auditory location. Table 8. Data grouped by training listening device. N = 8 Data Grouped by Training Listening Device - Headphones Azimuth Pre-Training Azimuth Post-Training Headphones Abs. Mean Error 17.24º (SD=4.97) 13.84º (SD=5.24) In-earphones Abs. Mean Error 17.64º (SD=5.61) 20.13º (SD=10.33) N = 8 Data Grouped by Training Listening Device – InEarphones Azimuth Pre-Training Azimuth Post-Training Headphones Abs. Mean Error 16.36º (SD=4.94) 13.93º (SD=4.25) In-earphones Abs. Mean Error 19.85º (SD=7.97) 15.81º (SD=4.85) Headphones In-Earphones
127 From Table 8 we can see that keeping congruency (grey cells) between listening devices used during training and experimental phases, gives rise to generally lower absolute mean errors of sound localization in the post-training phase. Incongruence between training and experimental session listening device disrupted completely the benefits of training in the case of participants that used in-earphones in experimental phases. A mean decrement in performance of about 2.5º is observed for these participants, from pre to post-training session (also the mean value presents more variability). Nevertheless, incongruence did not prevent learning and better performance in post-training sessions for participants that used headphones in experimental phases. Figure 55 presents the distribution of the mean error as a function of the stimuli position, for the congruent sessions (same audio output device in training and experimental sessions). In Figure 55, positive errors indicate misjudgments in sound location towards the ear plane, while negative errors indicate misjudgments of sound location towards the frontal plane (azimuth 0º). Figure 55. Mean error distribution and direction as a function of the stimuli position, for the congruent sessions. Positive errors indicate misjudgments in sound location towards the ear plane, negative errors indicate misjudgments of sound location towards the frontal plane. Interestingly, it is possible to observe that positive errors are predominant, meaning that when misjudging location participants are prone to locate the stimulus as closer to the ear plane. Finally, as headphones are more permeable to external noise when compared with in-earphones, we conducted a test to verify if the results obtained in silent conditions would hold in conditions with added environmental noise. Thus, we replicated this experimental protocol for eight new participants in a set-up in which the environmental noise reached the 56 dB(A) SPL. In these environmental conditions, participants had an absolute mean error of 18.52º azimuth for the pre-training session, and an absolute Azimuth (Degrees)
128 mean error of 14.86º azimuth for the post-training session. These results differ on an average of 1.16º, when compared with results of congruent sessions using headphones. Discussion: We presented a valuable method to assess listener’s spatial perception and evaluate performance between two audio devices. In the comparison between these particular models, headphones appeared to be the best solution for presentation of auralized sound and we should further investigate the benefits of using large housing with open back headphones. The fact that large housing headphones may allow individualized pinnae and ear-canal modulation over the non-individualized HRTFs, might be an important factor in the final performance outcome. Nevertheless, we can reduce the differences between devices if short training sessions are included and the same audio output device is used between training and test. In-earphones can benefit greatly of maintaining congruency between experimental and training phases. Future work should exhaustively compare between several types of audio output devices and should investigate how performance is affected by the introduction of binaural room acoustic cues. From this work, we can conclude additional guidelines for IVEs development: 1. Choosing audio output devices should be guided by the evaluation of user’s perception in controlled conditions; 2. Open-cascade headphones, which allow individualized pinnae and ear-canal modulation over the non-individualized HRTFs, are preferable over listening devices that have in-ear sound output. 3. When using non-individualized spatial cues, a short training session is advisable and can result in considerable improvements in location tasks. However, for the training improvements to be transferrable from training sessions to generic use situations, one should maintain the displaying or audio output device. 3.5. Conclusion In this chapter, we argued that the empirical approach, which was prominent at the dawn of HCI, could be particularly useful in this era of renewed and widespread use of IVEs. In fact, the same empirical
129 approach that resulted in the development, and ultimately adoption, of certain UIs that are still part of today’s personal computers will play the same role in order to define which IVE systems will be further developed and eventually be part of the long desired IVE as an UI. Thus, the main value of this section is in the experimental tools and protocols here described and their pertinence for the development of PCMs. Accordingly, its success should be measured by how IVE researchers and developers adopt these experimental methodologies. As we pointed out at the end of Section 1.3, there are two types of user’s cognitive models capable of providing insights into design and interaction problems. In Part II, we have been providing practical demonstrations on the role of PCMs, thus in Part III we will the same for DCMs.
Part III DCMs and Safety-Critical Interactive Computing Systems
131 4. THEORETHICAL BACKGROUND ON MEDICAL DEVICES ICSs AND DCMs In this Chapter we will introduce the reader to safety-critical ICSs, which will be our use case to develop new DCMs capable of detecting instances of potential use-error and guiding the definition of design guidelines capable of mitigating them. Due to the consequences of use errors when using such safety-critical ICSs, these devices constitute the perfect use cases to illustrate the usefulness of DCMs, an approach specially tailored to be applied in early design stages or during the investigative process towards understanding use-error in a particular UI. Medical devices are a category of safety-critical ICSs, especially prone to use-error with drastic consequences. Moreover, the area of medical devices is an emergent area in terms of new software and increasingly complex UIs. These devices stopped being manipulated by experts only in controlled (however sometimes harsh) environments such as the emergency room and began to be used also by patients with any kind of background and in a wide range of contexts of use. Next, we will better expose what are safety-critical interfaces and what are the HCI challenges in medical devices. 4.1. What are Safety-Critical Interfaces? Complex safety-critical systems are defined in [166] as: “… a system whose safety cannot be shown solely by test, whose logic is difficult to comprehend without the aid of analytical tools, and that might directly or indirectly contribute to put human lives at risk, damage the environment, or cause big economical losses. ” As Marco Bozzano points out on his book Design and safety Assessment on Critical Systems [167], this definition is peculiar because it chose to merge two different concepts that could be defined separately, complexity and criticality . Nevertheless, the motivation to presenting them together has to do with the fact that these concepts often go hand in hand with safety-critical systems, where there is a steady trend to increasing complexity as these systems become gradually more prevalent. There are many traditional applications of safety-critical systems in areas such as aviation and aircraft flight control, military and aerospace, and energy systems such as the nuclear industry. Moreover, the complexity of
132 current communication and information systems is contributing to the development of new areas of application of safety-critical systems in industries such as the automotive and healthcare. In the scope of this thesis, we are mainly concerned with the aspects related to the UI of safetycritical medical devices. Thus, complexity is relevant as long as it influences the quality of the interaction, as it often does. The increasing complexity of the software layers of a safety-critical system – that are “invisible” to the user – often result in medical devices that have a great number of functions, and where distinction between states is sometimes hard to comprehend and test [167]. The history of safety-critical medical devices can be traced back to more than one century ago, and similar to the history of IVEs, it begins with analogic devices and evolved to rather complex interactive computing systems. In 1895, Wilhelm Conrad Röntgen, accidently discovered the X-rays when manipulating a Crookes tube and within a year machines of the sort of the one in Figure 56 (panel A) where being deployed in medical facilities. In terms of interface, these first generation medical devices relied on knobs and knife switches, and had little or no displaying of information (with the exception of gauges on early XX century diagnostic devices, see Figure 56, panel B). A B Figure 56. Early experiments with X-ray machines using Crookes tube apparatus (Panel A), image under Public domain - published in USA before 1923. An example of an early commercial electrocardiogram machine manufactured by Cambridge Instruments (Panel B), image under Public Domain. With advances on the development of electronic components and a better understanding of the fundaments of bioelectricity, smaller programmable portable devices, such as the first wearable pacemaker (Figure 57, Panel A) became possible. In 1957, Medtronic started to commercialize wearable cardiac pacemakers that had controls to adjust the pulse to given amplitudes and durations, based on the particularities of each patient [168]. Today’s wearable devices, such as portable infusion pumps, have numerous functions that allow the patients to define and schedule the values to be infused,
133 monitor the progression of the treatment, and interpret complex logging data. These devices include wireless communication with hand-held peripherals and communication with software packages, allowing for storage of medical information and enhancing the process of therapy tracking (Figure 57, Panel B). A B Figure 57. A Medtronic wearable external peacemaker from 1958 (Panel A), image from [168]. The Medtronic MinimedTM 640G an insulin wearable pump released in 2015, image from Medtronic, Inc. Perhaps one of the best examples of the complexity of today’s safety-critical medical systems is the daVinci® Surgical System from Intuitive Surgical, Inc. The daVinci is a teleoperation system where the surgeon controls four robotic arms, which allow for minimal invasion surgery through video-laparoscopy (Figure 58). These surgical systems are particularly challenging, in terms of HCI, because they combine the challenges of IVEs – surgeons rely on stereoscopic images of the intervened region [169] and AR systems were developed to assist during the medical procedures [170]– and the challenges of operating using typical controls of a safety-critical medical system.
134 A B C Figure 58. The daVinci® Surgical System from Intuitive Surgical (Panel A), image from Intuitive Surgical. A stereoscopic endoscope (Panel B) and the visualization apparatus from the daVinci® master console (Panel C), images from [Nam2012]. Modern medical devices constitute a good example of the complexity safety-critical medical devices can attain, and are thus a particularly relevant case study for the contribution analytical HCI methods can bring to the safety enhancement of these interactive computer systems. From a software development standpoint, safety-critical systems requires the application of carefully designed processes (often, standardized processes) during the phases of specification, architecture, development, and verification. There are two distinctive approaches during specification and design of these systems [171]: 1. Specify and design error-free systems, proving that faults are not possible – however, this approach only works for small systems, which are sufficiently compact for formal mathematical methods to be used in the specification and design, and with no or little intervention by a human operator. [172] presents one of the few examples of this approach, where a productive dialogue between the developers of a dialysis machine, with no experience or knowledge of formal methods, and computer scientists using formal analysis tools proved the efficiency and safety of certain software components; 2. To aim for the first approach, but to accept that malfunction, faulty modes, or use-errors might happen and to contemplate, during specification and design, error detection and recovery
135 capabilities. Due to the complexity of current safety-critical systems, this is the prevalent approach. Thus, while there seems to be a high level of awareness among regulators and even developers for the need to apply standardized processes during development and evaluation of interactive safetycritical systems, as we will see, practice tells a different story. As Lyu famously exposed in his Handbook of software reliability engineering : “ The demand for complex hardware/software systems has increased more rapidly than the ability to design, implement, test, and maintain them. ” [173]. 4.2. What is the Problem with Medical Devices Interfaces? In the United States, when a medical device certified by the Food and Drugs Administration (FDA) is involved in the root cause of a medical accident, a description of the event, including causes that led to the medical incident, is included in the MAUDE 30 public database, and an investigation might be open to look into a possible recall of such device. One of the most severe incidents registered, and one that sparked a lot of research in the causes of medical device use error, was the Therac-25 medical incident. The Therac-25 was a computercontrolled radiation therapy machine, commercialized in 1982 with the main purpose of delivering a highly concentrated dose of radiation to a part of the human body with minimum impact to nearby tissues. Taking into account the goals and context of use of such a device, one could guess that this is a technology to be used by skilled operators and that depends on highly safe software. Nevertheless, and contrary to what might be the patient perceptions, these safety-critical devices are often developed by small companies, at the time, with few established practices of safety-critical software development. The Therac-25 was developed by a small team of engineers at a company called Atomic Energy of Canada Limited (AECL). While previous versions of the radiotherapy machine sold by AECL relied heavily on hardware control, this version of the Therac was the first generation to make extensive use of software control – developed by a single person in the company using PLP 11 assembly language and leaving the development process mostly undocumented. In a letter to FDA, the AECL admitted that the Therac-25 software evolved from the Therac-6 (an older machine from the same company) and that the current “program structure and software was using certain subroutines that were carried over to the Therac-25 around 1976” [174]. This migration of older software routines to a new machine and the removal of 30 The Manufacturer and User Facility Device Experience (MAUDE) database from the FDA can be visited, here: https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfmaude/search.cfm
142 In [188], the Decision-Ladder framework was further refined by identifying the cognitive shortcuts that are most important when analyzing human-machine interaction. These shortcuts are those that allow a rapid transition between three fundamentally different user's type of behavior and mental state: (i) Skill-Based Performance – highly practiced, mainly physical actions in which there is no conscious monitoring. Typically used to complete familiar and routine tasks, skill-based responses are generally initiated by a specific event (e.g. alarm, visual cue, the need to start a routine task). Tasks performed with this type of behavior are so familiar that little or no feedback information is needed to accomplish the task. (ii) Rule-Based Performance – in which the user applies a set of rules (learned through formal training, interaction with other experienced users) to a particular interaction. During rulebased performance the rules are being applied with minimal feedback from the situation, except perhaps for the detection of waypoints to indicate the correct progression of the rule-based procedure or sequence of actions [188]. The level of conscious control is intermediate between the skill-based performance and the knowledge-based performance. (iii) Knowledge-Based Performance – is adopted when a completely new task or an abnormal situation is presented for which the user has no set of rules, and a novel plan of action needs to be formulated. In these situations, trained users tend to try to find an analogy between the unfamiliar situation and some known patterns of events for which rules are available in the rulebased level. Once a plan or strategy is developed using knowledge-based information processing, the user might revert to the previous levels to accomplish the task. Since it may involve a process of trial and error in order to optimize the solution, knowledge-based behavior involves a significant amount of feedback from the situation. Because of its ability to account for different types of users in different contexts, the DecisionLadder framework became the most widely used information-processing model for use error classification [189],with applications ranging from Lean manufacturing [190], to aviation and air traffic control [191], [192], to submarine navigation [193] and military land combat [194]. Today's practitioner's guides on process plants design and safe operation are still almost entirely based on Rasmussen's work (see [195]). Furthermore, one of the main frameworks for the study of complex socio-technical work systems, the
143 Cognitive Work Analysis (CWA) [196] uses decision ladders as one of its outputs. Recently two publications further analyzed the impact of Rasmussen's ideas on the disciplines of human factors, HMI, and HCI (see [197], [198]). Both Norman’s and Rasmussen’s work have drawn attention to the fact that we should consider several cognitive steps that might be invisible to the external observer, as well as different types of performances when analyzing and describing the interaction process. James Reason, in 1990, built on this knowledge and focused his research on interaction failure, by developing a taxonomy capable of categorizing all the different instances of an interaction failure, thus creating one of the first use error taxonomies [199]. To design safer HMIs that prevent use errors and facilitate the recovery from use errors when they occur, developers need to have a clear understanding of the relation between HMI design aspects and use errors. A standard way to build this understanding is by using a use error taxonomy, such as the one proposed by Reason, which classifies use errors in accordance with systematic criteria [200]. Use error taxonomies are quite valuable for developers because they provide a reference to check what types of use errors typically occur during user-system interaction, as well as when these errors are likely to occur. In addition, a use error taxonomy can also help developers distinguish intricate differences between use errors stemming from different causes, and in turn devise effective measures to prevent and mitigate future occurrences (see Section 5). Reason’s taxonomy, the Generic Error-Modelling System (GEMS), tries to explain how different types of performance (distinguished based on Jens Rasmussen’s skill-rule-knowledge classification of human performance) relate with different types of use-error. Thus, GEMS guides the identification of three basic error types: skill-based slips and lapses - usually errors in highly repetitive, practiced tasks, due to motor variability (slips), or inattention or memory failure (lapses); rule-based mistakes – usually errors due to erroneous action plans as a result of picking an inappropriate rule or using a deficient rule, or caused by faulty rule-retrieving mechanisms coming from misconstrued view of the device state, an over-zealous pattern matching, or frequency gambling effects; knowledge-based mistakes - usually errors due to erroneous action plans as a result of incomplete/inaccurate understanding of system, confirmation bias, overconfidence, or cognitive strain.
144 According to Reason, Slips, Lapses and Mistakes may be distinguished based on certain dimensions, such as type of activity prone to their occurrence, attentional focus at the moment of use error, predictability of the use error occurrence, control mode of the operator at the moment of use error, ratio of error to opportunity of error, situational influences, ease of detection and relationship to change (see Table 9). Table 9. Summary of the distinctions between skill-based, rule-based and knowledge-based errors (adapted from [199]). DIMENSION SKILL-BASED Slips and Lapses RULE-BASED Mistakes KNOWLEDGE-BASED Mistakes Type of Activity Routine actions Problem-solving activities Focus of Attention Out of the task Directed at problem-related issues Control Mode Mainly by automated processors (schemata or stored rules) Limited, conscious processes Predictability of Error Types Largely predictable “strong-but-wrong” errors (actions or rules) Variable Ratio of Error to Opportunity of Error Absolute numbers high but opportunity ratio low Absolute numbers small but opportunity ratio high Influence of Situational Factors Low to moderate: Intrinsic factors (frequency of prior use) likely to exert the dominant influence Extrinsic factors likely to dominate Ease of Detection Detection is usually fairly rapid and effective Difficult, and often only achieved through external intervention Relationship to Change Knowledge of change not accessed at the proper time When and how anticipated change will occur unknown Changes not prepared for or anticipated Reason’s use-error taxonomy became almost a standard for consideration of different types of useerror in generic tasks. Although it became widely applied in industries traditionally concerned with the effects of use-error (such as the aviation, medical, and automotive industry), the GEMS model stays true to its designation and often offers an overly generic set of classifications that fits all industries but stays short of capturing the complexity involved in each one. The notion of the different types of behavior (skill, rule, and knowledge-based) and the consideration of different types of use-error (Slips, Lapses, and Mistakes), have become the cornerstone concepts of DCMs dedicated to describe the occurrence of use-error. Along the way some generic models which try
145 to describe all the interaction cycle in all situations (faulty and normal interactions), have also been proposed. The GOMS model of Card, Moran, and Newell [70], is probably the most important examples of a general DCM. GOMS is an acronym standing for: GOALS – descriptions of what the user wants to achieve. Goals can be divided in sets of different sub-goals; OPERATORS – basic actions that the user must perform in order to interact with the system and ultimately achieve the goals; METHODS – different strategies that the user can adopt in order to achieve the goal; SELECTION – description about which methods will the user apply. This might depend on particularities of the user, on system state, or on details about the goals. In a typical GOMS analysis, a high-level goal is decomposed into a sequence of unit tasks, all of which can be further decomposed down to the level of basic operators (see Example 3). Example 3. One possible GOMS description of the goal hierarchy for the task of photocopying an article from a journal (adapted from [10]). This goal decomposition rational involves detailed understanding of the user’s problem-solving strategies and of the application domain [10]. With their GOMS model, Card and colleagues, created a notation useful for describing the visible user’s procedural knowledge when interacting with a system and this allows the analyst to easily grasp possible contexts for faulty interactions, dead-ends, or redundancy in the interaction flow. Looking at Example 3 above, and as [10] pointed out, there is a post-completion error easily identifiable in the GOMS description. When the copy of the article is removed from the tray (outer goal satisfied!)
146 [REMOVE-COPY], the outer (main) goal is already satisfied, which might lead to users moving away from the machine without fulfilling the last goal [RETRIEVE-JOURNAL]. In fact, this is a quite common use error – returning from a photocopier with the copy and forgetting the original in the scanner. As suggested by [10], one way of avoiding this could be a design that forced the goal [RETRIEVE-JOURNAL] to be satisfied before [COLLECT-COPY] becomes available, thus observing a good design principle – the obligation to satisfy all the mandatory sub-task before signalizing the completion of the main task. Today, software exists for automatic generation of GOMS-like models upon description of an interface and some rules of interaction (see CogTool [71] and Cogulator [72]). DCMs are certainly capable of providing valuable information to developers, and ideally, they should be applied in earlier phases of design [12], as they are especially helpful as a preventive tool. Nevertheless, they can also be quite valuable as “forensic” tools to study what happened during a useerror event. The in-depth descriptions of the interaction processes DCMs provide are an ideal starting point to consider mitigation measures that can help to reduce or eliminate use-errors. In Chapter 4, we presented the theoretical background that justifies the use of DCMs as efficient user-research tools during the development of new UIs. We have made the case that analytical approaches that take into account all the intricacies of the interaction process can become powerful tools in order to guide the development and analysis of UIs in safety-critical medical devices. In the next chapter, we will describe the development of a new DCM (a new taxonomy for use-error in medical devices) and show how it can contribute for the development of safer and more usable safetycritical interfaces.
147 5. ANALYSIS OF HUMAN PERFORMANCE AND DEVELOPMENT OF DESCRIPTIVE COGNITIVE MODELS “The human factor cannot be safely neglected in planning machinery” Alfred Holt, 1878 In this Chapter, we will describe the development of a new DCM (a new taxonomy for use-error in medical devices) and show how it can contribute for the development of safer and more usable safetycritical interfaces. By doing this we will also demonstrate the role of analytical evaluation methodologies for the in-depth study of use-error in safety-critical UIs. We choose as a use case the safety-critical field of medical devices, where human error can have severe consequences and where the capacity to understand and describe the user interaction behavior is of crucial relevance. We will end this section describing the development of an IVEs for training with these safety-critical interfaces. 5.1. A Use Error Taxonomy for Improving Human-Machine Interface Design in Medical Devices 5.1.1. On the Importance of a Use-Error Taxonomy for Safety-critical Interfaces on Medical Systems Medical Cyber-Physical Systems (CPS) typically incorporate a Human-Machine Interface (HMI) that serves as a centralized or distributed portal for users to monitor and control the system. Consider for example the Integrated Clinical Environment (ICE) [201], an interoperable infrastructure that coordinates multiple medical devices, medical apps, and other equipment to accomplish a shared clinical mission. ICE systems often provide a centralized HMI to allow the users to monitor and control the devices connected to the system. The HMI might also include safety interlocks and other safety systems (e.g., a centralized smart alarm system) to facilitate effective and safe user-system interaction. It is thus critical to the safety of medical CPS to design safe HMIs that ensure expected user-system interactions and prevent potential use errors . Use error is an act of omission or commission performed by the user that causes a device to respond unexpectedly [202]. Preventing use errors has becoming a top priority in medical device design [177], [179]. From a system-engineering standpoint, use errors are often induced by flaws in the HMI design. In fact, investigations of incidents with medical devices usually reveal that HMI design flaws, rather than the lack of user training or inadvertent user behavior, constitute the main source of use errors [177], [178].
148 To design safer HMIs that prevent use errors and facilitate the recovery from use errors when they occur, developers need to have a clear understanding of the relation between HMI design aspects and use errors. A standard way to build this understanding is by using a use error taxonomy that classifies use errors in accordance with systematic criteria. Developers can use the taxonomy as a reference, to check what types of use errors typically occur during user-system interaction, as well as when the errors are likely to occur. In addition, a use error taxonomy can also help developers distinguish intricate differences between use errors stemming from different causes, and in turn devise effective measures to prevent and mitigate future use errors. Numerous use error taxonomies have been proposed for medical systems with different degrees of specificity, scope, and coverage (see [200] for a survey). Recent years have seen a surge in frameworks to classify use errors with medical devices and better understand their causes in system design. In [203], Leveson’s STAMP framework ( System Theoretic Accidents Models and Process ) is used as a basis to classify medical errors. STAMP is designed to support the analysis of causal factors not only at the level of unsafe actions committed by individual users, but also at management levels. While this broader view is certainly useful for healthcare providers to investigate systemand organizational-level causes of use errors, the error model used in the framework only coarsely classifies use errors into three categories: feedback, control action, and knowledge errors. In [204], a Human Factors Classification Framework (HFCF) [205] from the avionics domain is adapted to the analysis of medical device-related incidents. The framework is built on Reason’s GEMS [199] and similarly to the approach based on STAMP, HFCF considers use errors from a system-level perspective, and includes only five use error categories: decision errors, skill-based errors, perceptual errors, routine violations, and exceptional violations. In [206] and [207], participatory design methods were used to create use error taxonomies for Computerized Physician Order Entry (CPOE) and tele-medicine systems. Participatory design builds on focus group discussions involving relevant stakeholders, including human factors specialists, cognitive specialists, social scientists and clinicians. These taxonomies can be categorized in two main types: model-based taxonomies , which build on human cognitive models (an example taxonomy of this type is that proposed by Zhang et al. [183]); and data-driven taxonomies , which build on statistical data on use errors (an example taxonomy of this type is that presented in [208] for number entry errors). However, existing taxonomies have a variety of limitations that might result in incorrect, incomplete, or inaccurate classification of use errors. On the one hand, data-driven taxonomies often include ad-hoc error categories derived from statistics on medical
149 incidents. They are usually limited to the current understanding and knowledge about use errors with a specific system. On the other hand, while model-based taxonomies in general promote a more systematic classification of use errors and their applicability typically extends across multiple device types [200], usually they build on Norman’s action theory model [182], which oversimplifies the human decisionmaking as a sequential process with seven conceptual cognitive stages. Whilst Norman’s action theory provides mental scaffolding for reasoning about the causal relation between HMI design aspects and use errors, the model disregards an aspect of human cognition that is important in the medical domain: skilled behavior due to well-practiced activities (e.g., learnt through training) or related to actions (predominantly motor actions) that can be performed with little conscious attention. This causes the taxonomies built on Noman’s model to fall short when dealing with use errors committed by trained personnel, which is often the case with clinical operators of medical devices. The Decision-Ladder framework describes human problem-solving as a seven-stage cognitive process, similar to Norman’s action theory. As illustrated in Figure 62, these stages include: (i) goal formation , where one decides what needs to be done; (ii) intention formation , where one decides a strategy to achieve the selected goal; (iii) actions specification , where one identifies a concrete sequence of actions to implement the selected strategy; (iv) actions execution , where one performs the identified sequence of actions; (v) perception , where one monitors the effects of the actions; (vi) interpretation , where one develops an understanding of the perceived system state; and (vii) evaluation , where one decides whether the goal has been achieved. Failure to complete any cognitive stage is interpreted as a precursor to use error. Figure 62. The Decision-Ladder Framework.
150 The advantage of the Decision-Ladder framework however lies in that it captures the case where one can traverse the cognitive stages in a non-sequential order, by taking certain cognitive shortcuts to skip cognitive stages. The cognitive shortcuts are useful for representing “rule of thumb” solutions adopted by experienced users when solving common problems [185], [188]. It should be noted that, when used appropriately, cognitive shortcuts can greatly improve interaction performance and also reduce use errors. The Decision-Ladder model includes three types of shortcuts: Skill-based shortcuts : heuristics adopted by skilled users when performing highly practiced actions during tasks (mainly motor tasks) – skilled users can complete these tasks with little or no feedback from the device. Rule-based shortcuts : heuristics adopted by trained users when performing procedural tasks they have learned through training or from previous experience – trained users typically rely on waypoints to monitor progress and status of procedural tasks. Knowledge-based shortcuts : heuristics adopted by experienced users when facing unfamiliar situations – experienced user tends to formulate an action plan by finding an analogy between the unfamiliar situation and some known patterns of events, and then execute the action plan (likely using skill-or rule-based shortcuts). The use of cognitive shortcuts may vary across different users, depending on the heuristics they have learned and their past experience with a particular device. 5.1.2. A New Use-Error Taxonomy for Safety-critical Interfaces on Medical Devices In this section, we present a use error taxonomy for medical devices that aims to address the limitations of existing taxonomies and explore the benefits of using a cognitive process model that is more sophisticated than Norman’s action theory model. In particular, we consider Rasmussen’s decision-ladder framework [180]. Thus, this work main contributions are: (i) the development of a use error taxonomy for medical devices to help developers reason about use error types and their relation with HMI design; (ii) an initial evaluation of the benefits of the developed taxonomy with respect to existing taxonomies. We have followed the guidelines in [183] to develop our taxonomy:
151 Step 1: Identify generic use error types by applying systematically a human error model to the cognitive stages of the selected cognitive model. Step 2: Elaborate an interpretation of the generic use error types with examples of the error types in the medical domain and typical HMI design flaws contributing to them. The elaboration is expected to better explain how the taxonomy is applied to medical devices and medical CPS. Generic Use Error Types To identify generic use error types, we build on Reason’s Generic Error Modeling System (GEMS) [199], which is a de-facto standard approach. GEMS defines three generic error types: Slips - actions not carried out as intended; Lapses - missed actions due to temporary failure of concentration, memory, or judgement; Mistakes - errors due to erroneous action plans. Applying GEMS to the Decision-Ladder framework is conducted by exploring the possibility of instantiating each error type at each stage and shortcut of the framework. This results in 16 use errors types (see Figure 63): 13 types of information processing errors due to failure to complete one or more cognitive stages; and 3 types of cognitive shortcut errors due to misuse of cognitive shortcuts during inappropriate situations. Note that not all GEMS error types can be applied to all stages. For example, slips or lapses are not applicable to the interpretation stage, as this stage relates to knowledge-based behaviors of the user.