Design explorations of interactive virtual environments using volumetric capture and visualisation techniques
Abstract
Dissertação de Mestrado em Design e Multimédia apresentada à Faculdade de Ciências e Tecnologia
Full text
1 1 Maximilian Oliver Rubin Design explorations of interactive virtual environments using volumetric capture anD visualisation techniques A 250 yeAr timestAmp of the botAnicAl gArden of the university of coimbrA Thesis in the context of the Master’s in Design and Multimedia, advised by Professors Jorge Carlos dos Santos Cardoso and Pedro Filipe Martins Carvalho and presented to the Department of Informatics Engineering of the Faculty of Sciences and Technology of the University of Coimbra. September 2022
Kairos A 250 year timestamp of the Botanical Garden of the University of Coimbra Maximilian Oliver Rubin Advised by Professors Jorge Carlos dos Santos Cardoso and Pedro Filipe Martins Carvalho
Abstract Capturing environments and moments in time has always been something humans have been fascinated by – from paintings to analogue photography, to digital cameras. With the recent developments to AR and VR technologies, direct volumetric capture of reality has been gaining traction, with the potential of creating immersive, state-of-the-art visuals, blurring the lines between the real and virtual worlds. Capturing and transmitting space itself is one of the leading challenges in the field of immersive media. Capturing areas such as sites of cultural heritage is part of this equation and has generally been achieved either by 360-degree video or photography. However, these methods limit the user’s full sense of depth of the image. A solution to overcome this limitation consists of capturing real-world data through reality capture methods, resulting in a set of data points in space called a point cloud. This point cloud then passes through a reconstruction process, transforming it into a 3D polygonal mesh. However, this process limits its ability to provide a deeper sense of immersion and can introduce undesirable interpolations of reality. Nonetheless, by using point clouds as the basis of an immersive environment, we can explore new visual languages and alternate modes of perception. The objective of this thesis is the creation of a point cloudbased digital timestamp of the Botanical Garden of the University of Coimbra, by capturing reality and thus a moment in time that can be run and visualised in 3D space (through Virtual Reality), in real-time. This digital timestamp consists of three distinct locations within the Garden, containing century-old tree specimens. These areas were digitised through photogrammetry to create point clouds, which were then used as the basis of an immersive audio-visual experience whose main goal was to convey the ambiance of the Garden, along with showing people its natural processes usually hidden to the naked eye with the aim to rekindle our relationship with nature. This thesis resulted in multimedia content consisting of PC and VR based experiences, and 360 degree videos. Both PC and VR variants were evaluated by 22 volunteers, resulting in a positive response distribution for the user’s experience in relation to their relationship with the natural world. Keywords: Virtual Reality, Data Capture, Point Clouds, Photogrammetry, Immersive Media.
Resumo Capturar ambientes e momentos foi algo que sempre fascinou o homem – desde pinturas, a fotografia analógica, a câmaras digitais. Com os recentes avanços em AR e VR, a captura volumétrica da realidade tem ganho tracção, tendo o potencial para criar ambientes visuais imersivos, de útima geração, esbatendo as linhas entre o mundo real e o virtual. Capturar e transmitir o próprio espaço é um dos principais desafios na área dos meios imersivos. Neste sentido, a captura de áreas do património cultural é um objetivo recorrente e tem sido geralmente realizada quer por vídeo de 360 graus, quer por fotografia panorâmica. No entanto, estes métodos naturalmente limitam a sensação total de profundidade por parte do utilizador. Uma solução para ultrapassar esta limitação consiste em capturar dados do mundo real através de métodos de captura volumétrica, resultando num conjunto de pontos de dados no espaço também conhecido por “nuvem de pontos”. Esta nuvem de pontos passa por um processo de reconstrução, transformando-a numa malha poligonal 3D. Contudo, este processo limita a sua capacidade de proporcionar um sentido mais profundo de imersão. No entanto, ao utilizar nuvens de pontos como base de um ambiente imersivo, podemos explorar novas linguagens visuais e modos alternativos de percepção. O objectivo desta tese é a criação de um timestamp digital do Jardim Botânico da Universidade de Coimbra, através da captação da realidade e, portanto, de um momento no tempo que pode ser executado e visualizado em espaço 3D (através da Realidade Virtual), em tempo real. Este timestamp digital consiste em três locais distintos do Jardim contendo espécimes de árvores centenárias. Estas áreas foram digitalizadas através de técnicas de captura volumétrica para criar nuvens de pontos, que foram utilizadas como base de uma experiência audio-visual imersiva cujo principal objectivo é captar o ambiente do Jardim juntamente com mostrar às pessoas os seus processos naturais normalmente escondidos a olho nu com a finalidade de reavivar a nossa relação com a natureza. Esta tese resultou em conteúdo multimédia constituído por uma experiência baseada em PC e VR, e vídeos de 360 graus. As variantes PC e VR foram avaliadas por 22 voluntáros, resultando numa distribuição de resposta positiva para as experiências do utilizador em relação à sua relação com o mundo natural. Palavras-chave: Realidade Virtual, Captura de Dados, Nuvens de Pontos, Fotogrametria, Media imersiva.
1. Introduction
2 The essence of observing, describing and remembering information and personal experiences has been with us for as long as we can remember, and has been the very basis of our survival as a species. A primary instance of this can be seen in the prehistoric paintings and engravings on walls of caves and rocks, which portray the importance of our human need to communicate [1]. This form of communication left permanent markings for the future, leaving important messages for the generations to come. To this day, the basis of effective communication in the form of visual prompts hasn’t changed. We have only got better at it and are continuously refining and optimising the ways we share information. Now when we think about how far we have come in our ability to express our thoughts and ideas with the help of VR (Virtual Reality) technology, we have almost reached the tipping point of limitless personal expression and of making the imagination explicit. The desire to show to another human being or to an audience the real state of one’s imagination, perspective and ideas is quite an alluring concept. Turning ideas, written and spoken words into something we can see, can be considered the pinnacle of personal expression and transmission of information. The same can be said when someone tells us that we “speak clearly”, it’s a visual metaphor: they “see” what you mean. In other words, it would be like a new kind of communication, a form of sharing information that is similar to telepathy where thoughts and ideas can be shared visually. Technologies and projects based on VR are heading in this direction and the recent hype surrounding the Metaverse [2] is just one of the many examples of projects which will radically change the way we communicate and perceive reality as we know it. Because of this, there has been an influx of new methods and improved accessibility when it comes to capturing reality to facilitate the emergence of this upcoming digital realm [3][4][5]. Cheaper capturing devices, improved data processing methods and advancements in hardware performance (especially when it comes to GPUs (Graphics Processing Units)) are allowing this field to thrive, making it attractive for digital artists and content creators. Point clouds are the foundation of the process of capturing reality. At their core, point clouds are essentially files that contain a multitude of data points in 3D space that are created by performing a scan of an object or a location.
3 The most popular forms of attaining these scans is by using either a LiDAR (Light Detection and Ranging) device, or photogrammetry (the process of extracting 3D data from digital photographs). However, challenges still remain when it comes to transforming point cloud data, whether it be from a laser scanner or from photogrammetric capture, into realistic digital models. The conversion of the measured data into a high quality 3D polygonal surface model is still a resource-heavy task. To make matters worse, point clouds often consist of incomplete, noisy and/or sparse data, especially when capturing large outdoor scenes, which in turn further negatively influence the quality of the final model [6]. A small number of artists in the visual effects and virtual immersive fields have recently approached this challenge in a different way. They started exploring point clouds in their crude format, taking advantage of their noisy nature to discover new aesthetics and new visual languages. With this approach, they are able to take us on emotive journeys through digital fragments of reality. Most of these projects consist of pre-rendered videos while others consist of realtime immersive VR experiences. However, the latter are quite demanding, hardware-wise, usually take place in short lived exhibitions and typically follow a predetermined route within their virtual worlds, thus limiting the users’ ability for full control within the virtual setting. This project’s main goal was the creation of an immersive visual representation of a memory of the Botanical Garden of the University of Coimbra [7], by capturing reality and thus a moment in time that can be run and visualised in 3D space (through VR), in real-time. In addition to serving as a digital backup of the Garden and commemorating its 250th anniversary, this project allows anyone to access and experience this location remotely, even in face of restrictions such as ones imposed by pandemics or physical limitations, effectively allowing unlimited visits to this site of cultural heritage [8]. This project focused on three distinct areas within the Botanical Garden, each of which highlights memorable centenary-old tree specimens: the Tilia x europaea [9], the Erythrina Crista-Galli [10], and the Ficus Macrophylla [11]. Here we also simulated some of the area’s natural processes usually hidden to the naked eye, with the aim of evoking a sense of inquisitiveness within our users.
4 These natural processes were simulated over our point clouds and consist of the areas flora and fauna, including the tree’s vascular systems and underground root networks, following the premise that if we can observe the processes of nature in VR, we may become more conscious of its inner workings in real life, and thus gain new levels of appreciation for the natural world. These three areas are virtually interconnected, simulating snippets of a memory stroll through the Garden, allowing users to freely travel from one memory to the next. 1.1 Why point clouds? One of the main concerns behind VR is the challenge to develop interactive and immersive worlds that feel as close to reality as possible. Point clouds are typically converted into 3D polygonal meshes in an effort to create a close replica of what was captured, but this sometimes results in a “soft edged” appearance which does not appeal to the eye, especially at close proximity [6]. The fact of the matter is, point clouds are the basis of a captured surface and thus can be considered the fabric of (virtual) reality itself, metaphorically speaking. By visualising point clouds in their crude and diffused state, our focus gets taken away from the visual element’s mass, whilst maintaining a sufficient level of surface detail of what was captured for recognition, allowing our mind to focus more on the general atmosphere and tone of the locale. This opens new potentials to explore alternate modes of perception and new visual languages. So rather than using computer algorithms to fill in missing surface information with polygons, which are merely an interpolated interpretation of reality, we can take advantage of our brain’s innate ability to actively interpret our perceived reality by “filling in the gaps”, fabricating the illusion of reality, and thus allowing for a deeper, more authentic connection between the user and the virtual world [12][13]. In other words, even if some point cloud data is missing, its diffused and fragmented character still has the capacity to provide a greater sense of presence and authenticity compared to 3D polygon meshes. Moreover, this form of visualising data opens new gates to explore the way our memories and dreams work, emulating the way our subconscious mind stores and replays data. Similar to point clouds, dreams tend to be quite incomplete, fuzzy and sometimes unclear, yet we often seem to remember the general feelings they leave. That’s the real goal here, to be able to recreate the ambiance of a moment lost in time.
5 1.2 Motivation Several factors influenced the development of this project. First, in addition to serving as a digital backup of the Garden and commemorating its 250th anniversary, this project allows anyone to access and experience this location remotely, allowing unlimited visits to this site of cultural heritage. Second, as a technology-driven species, we have been slowly losing touch with nature and have become more distant to its life-giving benefits [14]. This project aims to trigger our innate “biophilia” [15] (a term used to describe our affiliation with mother nature, its positive impact on our mental health and our overall well-being), by rekindling our connection to the natural world, not as a tool to substitute nature and reality itself, but rather as a tool to enhance it. Third, by utilising and exploring the potential of the latest reality capture and point cloud visualisation techniques, this project seeks to effortlessly transmit the feelings of a moment captured in time through contemporary technology. Similar to a photograph, but in this case a type of photograph we can “jump into”, so to speak. Allowing us to explore the intricacies of an already geometrically complex scenery. Lastly, the conservation of the Botanical Garden is something that the University of Coimbra is working on and this has become a continuous endeavour. Belonging to one of UNESCO’s World Heritage Sites [16], it is paramount that the Botanical Garden preserves its integrity and essence. Natural disasters such as storms and the finite life-span of plant species are inevitable and have unfortunately taken out some historical trees and caused structural damage over the years. For instance, the iconic Washingtonia robusta palm tree, which at the time was the tallest plant in the Garden, got obliterated by a lightning strike in May of 1986 [17], not to mention the recent thunderstorm “Leslie” in 2018, which caused a lot of structural damage to the Garden [18] (figures 1, 3 and 4). Furthermore, there are also signs of deliberate acts of degeneracy by the public, such as carvings in the neoclassical Central Square (figure 2). Although these examples don’t all belong to the three chosen areas of capture, these instances are just to demonstrate how exposed the Garden is to external forces. That being the case, this project will also function as a digital timestamp, a sort of backup, by digitally preserving the areas within the Garden, with emphasis on one of its oldest trees: the Erythrina Crista-Galli, which is currently in a structurally vulnerable state (figure 5).
6 Figure 1. Figure 3. Figure 2. Figure 4. Figure 5. Figure 1: A missing Tilia tree in Alameda das Tílias, 05/11/21; Figure 2: Carvings in the Central Square marble, 05/11/21; Figure 3: Polypore mushrooms growing on the base of a cherry tree in Central Square, indicating trunk rot, 05/11/21; Figure 4: Damage to the fence’s neoclassical ornament; Figure 5: Current state of the Erythrina Crista-Galli, 05/06/22.
7 1.3 Goals and Results This project’s main goal consisted of creating a digital representation of the Botanical Garden of the University of Coimbra using point cloud-based capture techniques, by capturing reality and thus a moment in time that could then be run and visualised in 3D space, in real-time. Furthermore, by working in a digital environment, all constraints present in the real world were unleashed, meaning that the sky was the limit when it came to how we could manipulate and present these data. This lead to another goal which consisted of enhancing the sensation of being present in the Garden, along with evoking feelings of childlike wonder by sensory stimulation via the visual and auditory systems. Moreover, scientific information on the characteristics of each area were gathered for the end goal of simulating some of their natural processes. To further increase its immersive factor, a VR compatible version of this project was also produced. Last but not least, this work resulted in a publication accepted by the International Conference on Entertainment Computing, in the category of Works-in-Progress papers, that will be presented in November 2022 in Bremen, Germany. This submission will also be considered for the conference’s Student Competition. This immersive experience of the Botanical Garden was the end product of this thesis, which was made publicly available for PC, a VR headset, along with 360-degree videos to be viewed within a VR headset or through online video streaming platforms. Albeit not having the same immersive potential as the VR headset variant, the 360-degree videos were created to improve audience reach. Regarding the VR variant, a rather optimistic yet feasible goal was explored, which consisted of running the immersive experience within the Botanical Garden, that is, the creation of an insitu installation, where our audience could further indulge themselves in the Garden. First and foremost, to be able to achieve these goals, a few steps were taken in order to bring the most out of this experience and to enhance the sensation of being present in the Garden. These steps consisted of capturing as much data as possible in order to stimulate as many senses as possible within the digital environment. In this case, these encompassed two main areas: the visual and auditory. To be able to acquire these data, photogrammetric capture was used to record the visuals whilst a set of microphones were used for audio capture for the goal of creating a 3D ambient sound effect.
8 These acquired data were then cleaned, processed and refined for the final digital environment. These data were then composed within Unity’s real-time game engine, where further visual exploration was done to see how these data could be presented and enhanced to create a captivating and immersive dream-like experience. Next, in order to guarantee a quality experience, performance and stability tests were run to ensure that no issues occurred such as low frame rates or visual anomalies, as this would have negatively impacted our desired outcome. As a final point, a set of user evaluation tests were run to determine if our immersive experience fulfilled our goals. 1.4 Scope Naturally, some boundaries were set in order to achieve the best possible outcome with the limited resources available. This section presents some considerations on the project’s scope. A part of this project was dedicated to emulate the way we experience dream-like states and memories through the expressive visual potential of point clouds. However, we only scratched the surface of this matter and did not delve into the finer details of Oneirology (the scientific study of dreams), Neuroscience and/or Cognitive Psychology. Although our initial plan was to capture multiple locations within the Botanical Garden, the sheer size of the Garden had to be taken into account due to time constraints and technical limitations. From its thirteen hectares of ground, three distinct areas of interest were selected containing the following century-old trees: the Tilia x europaea, the Erythrina Crista-Galli, and the Ficus Macrophylla. Lastly, due to the wide-ranging field of data capture and sensing technologies, and for the sake of this thesis, we mostly discussed photogrammetry, as it is the basis of the practical component of this thesis, and did not address the deeper scientific and mathematical components of remote sensing.
9 1.5 Document Structure This document is structured as follows: Introduction, Introduction to Capturing Reality, Related Work, Methodology and Work Plan, Project Development, Evaluation and lastly, Conclusion. In the first chapter, an introduction to the field of study and motivations behind the thesis are presented, as well as its goals, objectives and scope. The second chapter, Introduction to Capturing Reality, serves to provide a deeper understanding of the field of capturing real-world data by laying the foundations of this matter. This includes its historical background, its methodologies, current challenges and limitations, followed by the latest contemporary methods and devices used in this field. The third chapter presents related work and projects to provide a better understanding of what has already been achieved in the field of using point cloud data in the context of immersive experiences. In this chapter, each case study is analysed for its workflow, motivations and inspiration, highlighting pertinent information that aided the development of the practical work featured in this thesis. The fourth chapter contains the Methodology and Work Plan, which defined the structure for the development of the practical component of this thesis, followed by a project time-line. This chapter served to guarantee a stable systematic workflow throughout the development of this thesis, including the methods and procedures that were used to best achieve our desired goals and objectives. The fifth, sixth and seventh chapters encompass all details regarding the development of the practical components of this thesis. This includes a small introduction to the site of capture: the Botanical Garden of the University of Coimbra, followed by a section of preliminary experiments which were conducted to get accustomed to the photogrammetric processing pipeline. Next, we thoroughly describe the development of Kairos, including all of its variants, together with all the steps that were taken to achieve the final immersive experience. Chapter eight, Evaluation, pertains to the testing phase of our project where we gathered and analysed user feedback of our immersive experience.
16 great concern. Apple recently introduced LiDAR scanning capabilities in its series of mobile phones and some tablets, though not as accurate as professional LiDAR devices, these still open up many interesting applications for hand-held scanning and are greatly improving audience reach and recognition [31]. Figure 10: Top Left: An illustration of a terrestrial laser scanning system (TLS) in the context of a forest. Available at: «https://ars els-cdn com/content/image/1-s2 0-S0924271616000204gr1_lrg jpg». Top Right: Photograph of a well known terrestrial laser scanner – the Leica P20 (manufactured in 2014). Available at: «https://www.researchgate.net/profile/Jim-Chandler-2/ publication/337849886/figure/fig1/AS:836700154654724@1576496270365/Leica-P20-terrestriallaser-scanner-with-1-HDS-target-used-for-registration-and png». Bottom: An example of a point cloud resulting from a TLS scan of a public park. Available at: «https://i pinimg com/originals/ de/18/bf/de18bf1c9b78e4b52981165b47796f3c jpg»
17 2.3 Photogrammetry Photogrammetry (Photos = light; Gramma = writing; Metron = measure) is defined by the American Society for Photogrammetry and Remote Sensing as, “The art, science and technology of obtaining reliable information about physical objects and the environment through processes of recording, measuring and interpreting photographic images and patterns of electromagnetic radiant energy and other phenomena” [32]. To recapitulate and to fully quantify the progression of photogrammetry throughout history, this field of study can be separated into four general stages based on improvements in technology. The first generation of photogrammetry started with the invention of photography, which initiated the investigation and experimentation phase of remote sensing. The second generation, considered the era of analogue photogrammetry, brought aerial surveying and stereo plotting techniques that were established and used during the two World Wars, turning photogrammetry into a reputable and popular surveying and mapping method. The mathematical theory was well known at this stage, but the lack of computational power hindered numerical solutions. The third generation, also known as the analytical photogrammetry generation, began in the 1950’s during the uprising of the computer, which allowed for significant advances to photogrammetric development. Hellmut H. Schmid was one of the first photogrammetrists who took advantage of the computer, developing the basis of analytical photogrammetry using matrix algebra. Block adjustment programs (a set of tools designed to improve the georeferencing accuracy) and aerial triangulation (the process for determining the correct position and orientation of each image in a series of aerial images) were subsequently introduced in the late sixties that greatly improved aerial triangulation accuracy. We are now in the fourth generation of photogrammetry. This began with the transition from analogue to digital, where digital images are used instead of aerial photographs. Photogrammetric processes are now managed by computer software, providing huge cost and time savings in producing typical photogrammetric outputs [33][34].
18 2.3.1 The Photogrammetric Processing Pipeline The wide availability of digital cameras and the recent progress in computer vision has led to the rise of accessible 3D modelling solutions. Photogrammetric capture naturally starts with the capture of real world data through photographs or video (as an image sequence). This captured data then gets fed into the photogrammetric processing pipeline (figure 11), consisting of a few stages that need to be followed in order to produce a successful digital asset. This pipeline generally consists of the following stages: camera alignment, generating dense cloud, build mesh, and build texture. During the initial stage, camera alignment, the photogrammetry software takes as input a set of images or video frames and their associated camera calibration information (if available). It then scans through every frame and searches for features. Feature detection determines salient points (features) in the images or video frames. These typically consist of distinctive attributes such as patterns, strong lines and edge points. These in turn help estimate the following positioning of each camera (frame) within a common coordinate system which is related to the world coordinate system. It is usually desirable to reliably detect these points at different viewpoints and radial distances to ensure that the same salient features are detected in different frames. After the feature points have been recognized in several images, they are internally validated and screened for false correspondences in a step called geometric validation. The validated features are then correlated within every frame to get putative matches which are then used to estimate the position of every camera frame within the scene [21]. As a result, a sparse point cloud and a set of camera positions are formed [35]. Note that this sparse point cloud is merely a reference of the resulting camera alignments and only a fraction of the total data. In other words, a “validation step” before starting any heavy duty computational work. During the second phase, generating dense cloud, 3D coordinates of the surface points are estimated based on the output of the previous phase. The line of sight from every camera to the object is reconstructed into what is called a ray cloud. The intersection of many rays determines the final 3D coordinates of the object [36]. Once this has been completed, a dense point cloud is generated where each matchable pixel gets its own x, y, z coordinates in 3D space [37] (figure 12). The third phase, build mesh, consists of generating the surface of the scene or object based on the previously produced dense
19 point cloud. This step connects each set of three adjacent points into a triangular face, resulting in a continuous mesh over the surface of the model. In the final step, build texture, original images are combined into a texture map and projected onto the mesh, resulting in a photo-realistic digital representation of the originally captured object. Figure 11: Left: The basis of the photogrammetric pipeline. Right: The principal procedures of photogrammetry. The dotted line indicates the connection between image features Pn,m, external image orientations Rn, tn, and the tie point P1. Available at: «https://www researchgate.net/figure/ The-principal-procedures-ofclose-range-photogrammetryThe-dotted-line-indicatesthe_fig2_347828301» Figure 12: Photogrammetric capture of the Three Graces by Antonio Canova, illustrating the resulting dense point cloud and camera alignments. Available at: «https://www factum-arte com/resources/images/fa/technology/ photgrammetry/capture3 jpg»
20 2.4 Challenges and Current Limitations Photogrammetry While many improvements have been made in the field of photogrammetry over the years, some hurdles still persist and it is important to be aware of its challenges and limitations to be able to assess each and every situation accordingly. Many of the algorithms used for photogrammetric processing fall short in the reconstruction phase (build mesh), resulting in a failed or problematic 3D mesh. Most photogrammetric processing software provide thorough documentation to help circumvent failed or problematic meshes, however, these following limitations should be considered in all contexts. 2.4.1 Occlusions and Number of Photographs Necessary While not so common in indoor capture studios, occlusions are a common issue when capturing outdoor objects or scenes. An occlusion is when an unwanted object gets in between the camera and the target object, resulting in missing data. However, this can be easily overcome by taking more images from different angles, providing sufficient data for the processing algorithm to compensate for the occluded frame. For heavy occlusions, capturing images every 10º is recommended. For objects with minimal occlusion, images taken every 20º is sufficient [20]. As a rule of thumb, when in doubt, take more photographs. 2.4.2 Feature Detection During the camera alignment process, the algorithm searches for feature points within multiple images. For example, taking images of blank walls or having unfocused shots will most likely cause the algorithm to fail. The more tracking features the scene or object has, such as patterns and/or strong distinctive attributes, the better the results. In some cases tracking markers can be added to the object to yield better results. 2.4.3 Reflective, Transparent or Glossy Surfaces Reflective, transparent and glossy surfaces are photogrammetry’s biggest concern. Objects of this kind have
21 little to no surface detail for the software to pick up on. There is currently no way around this other than to coat these reflective surfaces with something that will scan, this however is not a great solution and quite impractical, especially when it comes to large scenes containing buildings with a lot of glass, for example. 2.4.4 Environmental Conditions Photogrammetric capture is heavily dependent on weather conditions. Because of this and in order to improve the successful rate of a capture, we should be fully aware of the following environmental factors: surface movement due to wind or other forces, as this may decrease the ability for the processing algorithm to match scene points, resulting in a very noisy or incomplete point cloud. Excessive variation in ambient lighting during image collection, caused by changing cloud patterns, for example, is also something we should look out for [38]. Though not as severe as constant surface movement, changing in lighting conditions will impact the recognition of features. LiDAR 2.4.5 Reflectivity Limitations A notable drawback of LiDAR is that it does not work well in scenes with a lot of light reflections. LiDAR is also unreliable when used to capture non reflective surfaces due to the fact that this technology relies on the principle of light reflection [39]. 2.4.6 Environmental Conditions Since LiDAR uses visible lasers to measure distance, this technology is of limited use in harsh environments, especially in certain weather conditions such as fog, rain, snow and/or hail, making it essentially “blind” in these conditions. 2.4.7 Data Processing Requirement Since LiDAR can cover large areas at a fast rate, it tends to generate a lot of data. Meaning adequate equipment is crucial to be able to process and interpret the large amounts of data it is able to capture [39].
22 2.5 Contemporary Data Capture Today, data capture has become one of the main fields of interest surrounding the creation of realistic digital representations of the real world, regardless of the technology chosen to do so. When we ask ourselves what concerns we may have when it comes to capturing and representing reality, some questions may arise such as: “how should I capture this scene or object?” Or, “what methods should I use to achieve the best results possible?” First and foremost, we have to be clear on what we want to capture, in order to determine how we should proceed to capture these data as an asset or as a scene to be used in VR. Capturing single objects has generally been accomplished with ease within studios using multi-camera systems. Multiple cameras are placed around the object, fully encompassing it. Generally, the higher the camera count, the better, as more data will be available to feed into photogrammetric algorithms that in turn will produce its digital representation. Furthermore, being in a controlled environment (e.g. studio), generally allows for a higher quality capture, such as having full control over global lighting – a key factor for a successful capture. The most prominent state-of-the-art example of photogrammetric capture can be seen in Wolf Digital World’s spherical-shaped chamber (figure 13). With a diameter and height of 6 and 7 metres, respectively, and containing 200 Sony A7RIVA cameras, it can currently be considered the world’s most advanced indoor reality scanner. On the other hand, Leica Geosystems have recently announced their state-of-the-art BLK series autonomous LiDAR scanners: the BLK2FLY, BLK ARC and the BLK2GO (figure 14), providing easy data capture solutions for a wide range of settings. The Leica BLK2FLY is a drone-style laser scanner with advanced obstacle avoidance. It creates point clouds during flight and was designed to capture building exteriors, structures and the environment. This device can be set to autonomously capture difficult-to-access building locations, such as rooftops and facades. The Leica BLK ARC was designed to integrate with robotic carriers to provide autonomous mobile laser scanning with minimal or no human intervention. Users are able to plan a scanning path and then launch the BLK ARC to scan the desired area, autonomously.
23 Last but not least, the BLK2GO is a portable hand-held imaging laser scanner that recreates 3D spaces on the go. It is able to capture images and accurate point clouds in real time and uses SLAM (simultaneous localisation and mapping) technology to record its path through space. The BLK2GO automatically constructs point clouds from the moment you start the scanning session to when you turn off the device [40]. All images and 3D data get merged and aligned automatically, making this a big improvement to the notorious post manual point cloud registration and alignment process – a common step which was necessary in previous generation laser scanners. It is clear that reality capture offers unique opportunities across a range of industries and sectors, and technology is advancing rapidly. The continued implementation of automated procedures for analysing and extracting information from 3D point clouds will continue to provide versatility, ease of use and lower cost. This will make LiDAR and photogrammetry available to a wider range of use cases, including non-traditional users. Advances in computer vision, AI and machine learning are already adding new dimensions to capturing reality. Despite the recent improvements in these fields, we are still in the early stages of computational and software-based methods that will form the basis for future reality capture. Figure 13: Wolf Digital World’s sphericalshaped photogrammetric capture chamber. Available at: «https:// metahero io/»
24 Figure 14: Leicas BLK ARC, BLK2GO, BLK2FLY, from left to right, respectively. Available at: «https://blk2021.com/wp-content/uploads/2021/09/ARC_2GO_2FLY.jpg»
3. Related Work: Immersive Experiences using Point Cloud Data
32 Figure 18: Still from In the Eyes of the Animal, speculating the way mosquitos see the environment. Available in «https://www.flickr.com/photos/marshmallowlaserfeast/» Figure 19: Photograph of the in-situ installation of In the Eyes of the Animal showing the custom VR helmets along with the tactile subpack elements. Available in «https://www.flickr.com/photos/ marshmallowlaserfeast/»
33 3.4 Promenade Created by Italian visual artist Davide Quayola, Promenade [46] takes a more synthetic approach to capturing nature. This project emerged from a commission by the Swiss manufacturer of luxury mechanical watches Audemars Piguet. Moreover, the relationship between Quayola’s work and Audemars Piguet is related to their interplay between tradition and experimentation. In other words, it is a balance between the master revisiting the tradition while at the same time revisiting it and reinterpreting it in a new way. Promenade is a short film that presents a high quality bird’seye view of the forests of the Vallée de Joux in Switzerland, near the brand’s original watchmaking facility. Quayola used high precision LiDAR scanners to collect these data, then translated the forest into a wave of interconnected footage, transforming the documented landscape into digital art. A lot of Quayola’s work is centred around exploring different ways of seeing and discovering new visual languages with the help of technology, equivalent to studying things with “the eyes of machines”. Quayola himself stated that he is not particularly interested in the object he is trying to capture, he is more interested in the tension between the old language behind it and the new language he helps discover. Although Quayola uses high-precision LiDAR scanners, he still enjoys the small imperfections it brings during the translation process (the transformation of the forest into the digital) (figure 20). “They (LiDAR scanners) cannot see perfectly. That fascinates me. If you imagine an impressionist painting, you are reducing the resolution. You are not aiming for perfection. Limitation brings out the expression.” – Quayola [47]. Furthermore, Promenade plays with the merging of the real world with technology. During the course of the video, different visual effects alternate between the real and the digital, giving us a peculiar hybrid machine-like view of the forest. To further enhance these ephemeral graphics, a carefully crafted soundtrack was added and synchronised with these visual effects. Quayola’s bold exploration of these new visual languages push boundaries in the way we can extract and visualise real world data in the context of nature, by processing visual information with a different logic. Even though Promenade has stronger “techy” visuals and a “fossil style” colour palette
34 (figure 21), compared to the more natural organic nature of the desired outcome of the practical component of this thesis, it’s accurately designed soundtrack demonstrates just how important audio is to get a more tangible experience from within the digital environment. Figure 20: Render from Promenade, note how well LiDAR was able to capture the intricacies of the forest. Available in «https://whitewall art/lifestyle/quayola-captures-vallee-de-joux-audemars-piguet» Figure 21: Still from Promenade, illustrating its “techy” visual language. Available in «https://www audemarspiguet com/com/zh-hant/news/art/quayola-promenade html»
35 3.5 Clams Casinos’ Moon Trip Radio The series of videos used to accompany Moon Trip Radio’s album [48] is another great example of point cloud data used in a creative way. Developed by Prague based computational artist Jakub Valtar, his custom made visualiser takes a more psychedelic approach to volumetric data representation. Based on drone-captured footage and photogrammetry, the otherworldly visual language chosen for his pre-rendered videos accompany Casinos’ soundtracks and are reminiscent of the bioluminescent corals seen in the tropics. This is evident in the first song of the album: Clams Casino – Rune (figure 22), which takes you on a journey through what seems like a mountainous terrain composed of high contrasting fluorescent point clouds. As opposed to the previous related projects, this one proceeded to follow-up and conclude the reconstruction phase of the photogrammetric processing pipeline, that is, reconstructing the faces (mesh) of the captured scene. Though not so obvious at first glance, once we take a closer look we can see that the fluorescent point clouds are in fact resting on a dark mesh. This must have been decided due to the chosen high contrasting visual language, where a floating point cloud would perhaps render in a very noisy and unclear manner due to point clouds inherent “see through” character. Likewise to Quayola’s accurately designed soundtrack, the synergy of Jakub’s videos with Casinos’ music ascend our sensory experience to another level. For example, if we play Casinos’ intricate sonics without any visual stimuli, or play Jakub’s videos without any music, we lose half of their expressive potential in both cases. Jakub’s videos, Promenade and especially In the Eyes of the Animal teach us that in order to take full advantage of any given immersive experience, we should place a strong focus on stimulating as many of our senses as feasible, because higher immersive quality levels elicit higher levels of presence, which in turn increases the effectiveness of the overall experience.
36 Figure 22: Still from Clams Casino – Rune, showing the dark mountainous terrain composed of fluorescent point clouds. Available in «https://www.youtube.com/watch?v=sWeG68lNcVA»
37 3.6 Andy Shauf – ‘Clove Cigarette’ The picturesque track ‘Clove Cigarette’ [49], by Andy Shauf, is a song written about half-forgotten memories that took a new approach to storytelling and cinematography. Following that theme, the music video tours through Shauf’s mind as he revisits Toronto’s skyline restaurant, his backyard, and late summer nights. Directed by Colin Medley and Jared Raab – a production team based in Toronto, the video clip was produced using a variety of technologies including LiDAR, photogrammetry and a renowned video game engine – Unreal Engine. Sauf’s world was captured then digitally reconstructed into still, relit scenes which were then rendered out as one continuous video. Because the song was about half-forgotten moments, Jared and Raab wanted to create something dream-like and imperfect. After looking for solutions, they stumbled upon Ruben Fro’s works, that is when they knew they had found the right toolset and aesthetic to articulate the feeling of the song. However, instead of running everything within Unity, like Ruben, they decided to try out Unreal Engine, which was something they were already familiar with. The team initially wanted to create point cloud based animations, but having run a few tests of just a still scene with the addition of atmospheric lighting and smooth camera passages, they figured out that was all they needed. By following this approach, they were able to create a dreamlike journey through Andy’s world that suited his warm, hazy summer vibes. By working in a fully digital environment, they had a lot of versatility when it came to perfecting the camera tracks, adding post-effects, changing the lighting and even bringing foreign elements into the scene (figures 23 and 24). This project motivated our incentive to test Unreal Engine as an option to render our point clouds. This led to our discovery of its superior volumetric lighting capabilities and easier point cloud importing methods, compared to its rival – Unity.
38 Figure 23: Still from Andy Shauf – ‘Clove Cigarette’, showing its dream-like and imperfect nature, the subtle depth of field camera and lighting effects along with the burning cigarette post effect. Available in «https://www.youtube.com/watch?v=A8QPFtS2_fE» Figure 24: Still from Andy Shauf – ‘Clove Cigarette’, showing its post effects (bloom/fog), lighting (neon lights above restaurant entrance and taxi headlights) and foreign elements (taxi and blue bicycle). Available in «https://www.youtube.com/watch?v=A8QPFtS2_fE»
39 3.7 Realities – Cologne Cathedral Part of the practical component of this thesis addresses the preservation of the Botanical Garden as a digital backup. This is something that is common in the Cultural Heritage and Archaeological Documentation fields, which focus primarily on creating the finest digital replicas of reality possible. Realities.IO is a company that specialises in reality capture, particularly in the field of Cultural Heritage, whose main goal is to allow you to visit sites of interest all over the world. Realities.IO software is available for free on Steam and is currently compatible with the HTC Vive, Oculus Rift, and Windows Mixed Reality. Realities – Cologne Cathedral [50] allows you to tour through a digital replica of the Gothic architecture of the Cologne Cathedral – another one of UNESCO World Heritage Sites. This VR tour also includes a portion of the Cathedral where visitors are not normally allowed, adding an extra sense of excitement and wonder to the VR experience when you are able to enter its secret rooms and enclosed areas. Similar to Clams Casino – Rune music video, Realities.IO used photogrammetry and proceeded to reconstruct the resulting point clouds into fully textured 3D polygonal mesh (figure 25). This was evidently done to achieve the highest level of realism, however, its “soft edges” and “low-poly” appearance are still quite noticeable at close proximity (figure 26). In this context it can be seen as a fair decision, nonetheless, it still demonstrates the imperfect nature of photogrammetry. For example, by comparing this project to the fragmented form of La Cathédrale (figure 27), a project by Benjamin Bardou, we can clearly see how it differs from Realities.IO’s visual language. Although La Cathédrale is not as visually refined, its point cloud representation allows us to dig a bit deeper into its space and get a better feeling of its atmosphere. Realities – Cologne Cathedral is a clear example of our decision to use point clouds as a visual representation medium for this thesis rather than a reconstructed mesh, to avoid the “soft edge” appearance resulting from the photogrammetric mesh face reconstruction process.
40 Figure 25: Still from Realities – Cologne Cathedral, illustrating its reconstructed mesh framework. Available in «https://store.steampowered.com/app/708620/Realities__Cologne_Cathedral/»
41 Figure 26: Still from Realities – Cologne Cathedral, showing an example of the “soft edge” appearance. Available in «https://store.steampowered.com/app/708620/Realities__Cologne_ Cathedral/» Figure 27: Still from La Cathédrale, by Benjamin Bardou, demonstrating an alternative point cloud based visual language used in the context of a cathedral. Available in «https://youtu be/ Nt9W1TQrTGA»
48 Sep JanOct FebNov Mar MayDec Apr Jun Jul 1.Data Acquisition 2.Data Processing 3.Project Development 4.Evaluation and Optimisation 1.a 2.a 2.b 2.c 3.a 3.b 2.d 1.b 3.c 3.d 3.e 3.f 3.g 4.a 4.b 1.c Figure 30: Gantt chart of our initial work plan, demonstrating how these four phases would be carried out during the execution time of this thesis.
49 Changes to the Work Plan Due to several unexpected technical drawbacks described throughout chapters 5, 6 and 7, along with the development of additional showcase material, together with the composition of the work-in-progress paper for ICEC 2022, which was initiated during mid-project development, the four phases of our initial work-plan took a noticeable temporal shift. Figure 31 shows an updated Gantt chart of how the development of our project really went, along with an updated work plan. Here the main differences are: the separation of the processing of our photographic and audio data due to their complexity, followed by the addition of activity #5: Developing showcase material, where we developed a website and 360-degree videos. 1. Data Acquisition a. Explore and develop a workflow for capturing data; b. Collect photogrammetric data; c. Record audio. 2. Data Processing a. Develop a workflow for data refinement; b. Prepare captured data for point cloud processing; c. Create point clouds; d. Refine point clouds; e. Refine Audio. 3. Project Development a. Explore the expressive potential of point clouds; b. Develop the structure of the virtual universe; c. Explore ways of enhancing our experience within this digital universe by adding interactive elements; d. Compose the immersive experience (PC); e. Run preliminary evaluation tests; f. Develop the immersive experience compatible with a VR headset. 4. Evaluation and Optimisation a. Fine-tune the immersive experience to achieve the best quality-to-performance ratio for both PC and VR variants; b. Run a final set of user evaluation tests. 5. Developing showcase material a. Website development; b. Audio-visual material development.
50 Sep JanOct FebNov Mar MayDec Apr Jun Jul Aug 1.Data Acquisition 2.Data Processing 3.Project Development 1.a 2.a 2.b 3.a 2.d 1.b 3.b 3.c 3.d 3.e 3.f 4.a 1.c Sep 2.e 4.b 4.Evaluation and Optimisation Figure 31: Gantt chart of how our work plan really went during the time-span of this thesis. 2.c 5.b 5. Developing Showcase Material 5.b
5. Project Development
52 Before starting project development, a brief introduction and historical background of the Botanical Garden of Coimbra will be made in order to help acclimatise our notion of this location of cultural heritage. After this, we present a series of four preliminary experiments, focused on exploring the best solutions for point cloud data acquisition through photogrammetry. Finally, two initial prototypes are described. 5.1 Object of Study – The Botanical Garden The University of Coimbra’s statutes of 1772 stated that: “in a place in the vicinity of the University that is to be considered the most appropriate and competent, a said Garden will soon be established; so that every kind of plant may be cultivated there; and particularly those of which can fruit usefulness in the field of medicine, are to be expected” [55]. Located in the heart of the city, the Botanical Garden is an iconic space in Coimbra and a prestigious location in Portugal for its scientific contributions to botany, consisting of a plethora of transcontinental flora that take you on a journey to regions throughout the world. Belonging to UNESCO’s list of properties registered as World Heritage since 2013, this Garden has been expanding its visibility and consequently the presence of visitors, who, in addition to being able to participate in the area of scientific culture, can equally enjoy this public space for leisure, where they can meet for casual encounters and explore nature by strolling by rare and exotic plants from all parts of the world with other natural landscape enthusiasts. Furthermore, this Garden has also expanded its role in environmental education, as a recreational space for various activities surrounding mental health and well-being, tours, concerts and exhibitions. Looking back in time, the 18th century brought about a revolution of minds and significant scientific progress. The Botanical Gardens, therefore, were created in order to complement and encourage the fields of Natural History and Medicinal Studies of the University of Coimbra [56]. This context led to the creation of the Botanical Garden in 1772, by the administration of the Marquis of Pombal, on land donated by Benedictine friars. Initially named Horto Botânico, the Garden took up only the area that is still known as the Quadrado Central (Figure 32: Central Square, point 5).
53 The statutes of the University of Coimbra determined the cultivation of all sorts of medicinal flora, including plants of overseas origin. The cultivation of plants began in the Central Square in 1774 and was only completed in 1790, with the addition of a central fountain – an installation which still remains to present day. During this time, the Faculty of Medicine studied the healing properties of plants which led to an upgrade consisting of a series of rectangular beds for further cultivation of medicinal flora. Today, the Garden still keeps to this scientific component through a seedbank program. Published for the first time in 1868, this seedbank includes many Portuguese and exotic species, including several endangered varieties, which have played a pioneering role in the conservation of nature. At the present time, the Botanical Garden covers around thirteen hectares of land and can be divided into different levels, stairs and avenues, each containing their own unique space, flora and atmosphere. Below is an illustrative map of the Garden with enumerated areas depicting its most prominent locations (figure 32), followed by a closer look and description of the listed locations (figures 33, 34, 35, 36, 37 and 38). One of the main objectives of the Botanical Garden is to establish a closer link between the Garden and the public, by awakening their interest in science, nature and botany. This is one of the main reasons this location was chosen as the basis for this thesis, by merging similar ideas and objectives thus emphasising the Garden’s essence through the help of contemporary technology. Main Gate Baixo Relevo de Luís Carrisso Alameda das Tílias Tropical Greenhouse Central Square Cold Greenhouse Figure 32: Bird’s-eye view illustration of the Botanical Garden depicting its most prominent locations. Adapted from «https://www facebook com/media/ set/?set=a 424667424214088&type=3» 1 1 2 2 3 3 4 4 6 6 5 5
54 Figure 33. Figure 34. Figure 36. Figure 38. Figure 33: Main Gate entrance, containing the marble statue of botanist Félix de Avelar Brotero. Dates from 1887 [58] – photograph taken in the Botanical Gardens of Coimbra, 05/11/21. Figure 34: Photograph from the Central Square, looking up at the Main Entrance, containing the statue of the director of the Gardens in 1918 – Professor Luís Carrisso [57]. Available at «https:// www.facebook.com/JardimBotanicoUC/photos/?ref=page_internal». Figure 35: Alameda das Tílias, one of the most emblematic places in the Gardens, which reminisces the old 19th century public footpaths of European cities [57]. Photograph taken in the Botanical Gardens of Coimbra, 05/11/21. Figure 36: Tropical Greenhouse – one of the oldest iron architecture buildings in Portugal. Its perfect combination of iron and glass give this space its surprising beauty. Reopened in 2018 after a few years of restoration, it is home to mainly tropical and subtropical plants [57]. Available at «https://www.facebook.com/JardimBotanicoUC/photos/?ref=page_internal». Figure 37: Neoclassical fountain in the middle of the Central Square, surrounded by geometrically designed walls, hedges and raised beds containing a wide variety of flowering fauna, transmitting the entire atmosphere of Romanticism to this area [57]. Available at «https://www facebook com/ JardimBotanicoUC/photos/?ref=page_internal». Figure 38: Cold Greenhouse, built in the 1950’s, under the direction of Botanist Abílio Fernandes. This greenhouse contains flora of humid and shaded environments, surrounded by a cascading waterfall and a small stream that runs through its entirety. A female nude statue by sculptor Martins Correia also inhabits this area, named “Botânica’’, who symbolises the Science of Plants [57]. Available at «https://www facebook com/ JardimBotanicoUC/photos/?ref=page_internal». 4 6 12 Figure 35. 3 Figure 37. 5
55 5.2 Preliminary Experiments Before starting any definitive development, initial tests were run in an effort to get accustomed to the photogrammetric processing pipeline. Moreover, these tests were carried out to determine an efficient workflow to best satisfy the practical goals and objectives of this thesis. 5.2.1 Experiment nº1 – 360º Camera In an attempt to capture as much real-world data in the shortest amount of time, we tested a 360º 4k video recording device, more specifically, the Samsung Gear 360. This idea stemmed from the urban-based projects by Ruben Frosali (3.1 Memories of Tsukiji, page 26) and Benjamin Bardou (3.2 Shinjuku, page 29), as these were also recorded using 360º cameras. In that respect, we recorded a short 25 second video, where we followed a predetermined path that revolved around a well known Tilia tree of the Botanical Garden (figure 39), thus capturing a complete 360º view of the tree. In order to be able to transform this video clip into a point cloud, we first had to render the 360º footage into a compatible format required by the photogrammetric processing software Agisoft Metashape. This was accomplished with Samsung’s Gear 360 ActionDirector, a media processing software used to stitch the 360º video into a standard video frame (figure 40). Basic video colour correction was also done within ActionDirector, to bring out the shadows and to correct any overblown highlights in an effort to aid MetaShape’s feature detection process. This footage was then imported into MetaShape’s photogrammetric processing pipeline using its default processing parameters. This resulted in a point cloud of 27 million points (figure 41). A few notes can be taken from this initial test, namely the large amount of noise in the resulting point cloud with strange floating sky and cloud particles, as well as some evident alignment issues resulting in duplicate objects (figure 42). Nonetheless, this point cloud was then imported into Unreal Engine, using its built-in point cloud importer, to be further analysed in the context of a real-time virtual environment. At first, Unreal Engine was having some difficulties rendering the point cloud (figure 43), its structure was incomplete and parts of the point cloud would appear and disappear intermittently during viewport movement. However, this issue was bypassed by increasing the point cloud plugin’s
56 default 1 million point budget to a value greater than 4 million (figure 44). It is to note that the larger the point budget, the more detailed the image becomes, however, this comes with the burden of exponentially higher processing requirements, thus lower frame-rates. In this experiment and with the hardware used, we were able to render 20 million points on-screen, in real-time, at a stable frame-rate of 60 frames per second. Figure 39: Still from the raw 360º footage of the Tilia tree. Figure 40: Still from the now stitched 360º footage with improved shadows and reduced highlights.
57 Figure 41: Screen capture of the resulting point cloud within CloudCompare, note its noisy nature, strange misaligned tree branches and blue sky particles. Figure 42: Screen capture showing evident alignment issues resulting in duplicate objects.
64 Figure 53: Screen captures illustrating CloudCompare’s colour filtering addon following the removal of the blue sky points (Left: Original, Right: Cleaned). 5.2.5 First Prototype Having already explored different visual representations of our point clouds, we decided to opt for a middle ground between accuracy and noise/fragmentation. Displaying enough data to allow us to understand our surroundings, but not realistic enough so as to not remove ourselves from the dream-like state of reliving a memory. It was understood that having an initial high quality dataset to work with was the most sensible course of action. This is because we can always reduce the final resolution of our point cloud, but can hardly increase it. For this reason, capturing reality by taking high resolution photographs in RAW format was the chosen and final method of data capture. Our final preliminary experiment began with the capture of the entire Main Entrance of the Botanical Garden, where 450 photographs were taken. To not miss any data, several rounds of photographs were shot, following a method similar to emulating a one-eyed cyborg trying to make sense of this space, by revolving around each and every object of the scene multiple times at different angles (figure 54).
65 Similar to the previous experiment, these photographs were then colour corrected and re-touched within Adobe Lightroom before getting imported into Metashape. This time, in an attempt to further enhance accuracy and feature detection, all photographs were slightly sharpened and de-noised. Furthermore, we decided to increase some of Metashape’s parameters, namely the number of desired feature points for the algorithm to look for, as this was thought to increase the number of points in the outcome of the following build dense-cloud process. This parameter alone increased the number of points of our point cloud to 174 million, resulting in a rather large and heavy file (figure 55). Naturally, a lot of floating points and false correspondences were included in this sum and needed to be assessed within CloudCompare (figure 56). Floating blue sky points were removed with the colour filtering addon, then the general noise was reduced with a noise filter (by removing any orphaned floaters greater than a certain radius of each and every point), followed by a subsampler to lower the global point count to 50 million points (figure 57). Although not yet defined, this subsample value was purposely used to improve storage capabilities along with improving responsiveness and performance within Unreal Engine. This point cloud was then imported into Unreal Engine, where the finer details of our selected visual language were explored to get better accustomed to the capabilities of Unreal Engine. Within this virtual setup, several visual postprocessing features were explored to enhance the point cloud environment. Adhering to a more realistic approach, here we explored dynamic lighting, volumetric fog and shadows, to emulate the light rays we had observed radiating through the tree’s leaves and branches (figures 58, 59 and 60). This followed by the exploration of point shape and size, in an effort to find a balance between realism and the atmosphere of a dream-like state (figure 61).
66 Figure 54: Screen capture of the resulting sparse cloud within CloudCompare, showing its various angles of capture. Figure 55: Screen capture of the resulting point cloud within CloudCompare, resulting in a point cloud of 174 million points.
67 Figure 56: Screen capture within CloudCompare, illustrating its floating points and false correspondences. Figure 57: Screen capture within CloudCompare, showing the reduction of floating points and false correspondences after an imperfect clean-up process.
68 Figure 59. Screen capture within Unreal Engine, illustrating the volumetric fog and dynamic lighting effects implemented for added realism. Figure 58. Screen capture within Unreal Engine, illustrating the volumetric fog and light ray post-processing effects radiating through the tree’s leaves and branches.
69 Figure 60. Screen capture within Unreal Engine, showing the exploration of point size and shape. Figure 61: Screen capture within Unreal Engine, illustrating the resulting balance between realism and the atmosphere of a dream-like state.
70 5.4 Second Prototype: High fidelity Tilia x europaea After concluding the previous experiments and having established most of the methodology, a more complete, high fidelity prototype was developed to include the majority of the features expected in the final project. During the development of our initial prototype, we noticed that Unreal Engine’s built-in point cloud importer was showing some issues whilst rendering the points on screen. For reasons still to be determined, some point clusters would start to flicker depending on the player’s angle of view within the virtual environment, this issue became quite noticeable and negatively affected our user experience. Coincidentally, this issue is also apparent in Andy Shauf’s – “Clove Cigarette” music video (mentioned in section 3.6: page 37), which was also developed using Unreal Engine, where fragments of the scene would disappear and reappear in an abrupt fashion. After finding no solution to this problem, it was decided that we shift over to Unity’s game engine, as this was also software we were more familiar with. This high fidelity prototype consists of a single area of the Botanical Garden that was captured during November, 2021 containing the Tilia x europaea [59]. The development of this prototype was divided into six stages: 1. Point Cloud Representation, 2. Representation of Natural Processes, 3. Interactions, 4. Sound, 5. Locomotion and Area Traversal, and lastly 6. Preliminary Evaluation. Here further visual exploration was performed on our point cloud, together with different approaches to represent and highlight some of the scenes natural processes associated with the tree’s biology. Moreover, sound and user interaction which allow users to directly influence the behaviour of the experience was also explored, including locomotion and area traversal. Lastly, a set of preliminary user tests were performed to understand how people would react to the immersive environment. Made up of 275 photographs, this capture resulted in a dense cloud of 100 million points (figures 62 and 63). However, after running into some initial performance issues (low frame rate), it was discovered that we could only render approximately 20 million points at any given time, meaning that our initial idea of merging multiple areas of the Botanical Garden in a sandbox style open-world was out of the question. During the development of this prototype we
71 Figure 63: A closer view of the tree where we can see how the capture process was done. Figure 62. Screen capture within Metashape showing the alignment of the 275 photographs, resulting in a point cloud of approximately 100 million points. further discovered that the graphics API set within Unity’s preferences heavily influenced the performance of our scene, where Vulkan API would outperform DirectX11 by around 40 frames per second, both using the exact same scene parameter settings (figures 64 and 65).
72 Figure 64. Screen capture of the scene being rendered using the Vulkan API. Figure 65: Screen capture of the scene being rendered using DirectX11, note the change in frame rate compared to the previous figure (see top right stat. bars).
73 5.4.1 Point Cloud Representation Three point cloud rendering methods were explored to find a balance between visual fidelity and computational performance. Our point cloud was initially rendered as a static object with a fixed point size, however, its points would become large at close proximity and thus diminish the visual dream-like aesthetic we were looking for (figure 66). As an alternative, a “pixel absolute” rendering method was explored so that each point would occupy the size of a pixel at all times (figure 67). Nonetheless, with this method, the points farthest away in the scene would inevitably scale much larger than the ones closer to the camera (user), resulting in an unappealing contrast of point sizes. Moreover, during these tests we noticed a significant drop in performance depending on the set size of our points, where a smaller global point size would render much faster than a larger point size. As a third attempt, Unity’s VFX (visual effects graph) was used to render and animate our point cloud (via texture files containing its position and colour data), here its continuous point burst feature with the addition of force fields was explored. This method resulted in interesting visuals dominated by roiling textures (figure 68), however, it quickly became apparent how computationally heavier this method was compared to the first two. From these results it was decided that the initial fixed point size rendering method be used, this time with a smaller global point size, as the final rendering method (figure 69). This method not only created a more uniform look throughout the scene, but also produced a more photo realistic and consistent appearance at different resolutions (1080p, 1440p, and 4k). Figures 66 and 67: Comparison between “fixed point size” (Left) and “pixel absolute” (Right) point rendering methods.
80 Figure 79: Area traversal zones at either extremity of our area. ~3.64m ~4.82m ~5.01m Figure 80: Representation of the proximity system that controls the fade-out effect. The red points ( ) representing the proximity sensors, along with their respective distances to the user (camera). Zone 1 Zone 2 User (camera) Proximity Sensors Trigger Box
81 5.4.5 Preliminary Evaluation This prototype was only a fraction of our desired goal at the time, nevertheless, we wanted to perform preliminary tests with participants external to our team to understand how people would react to the immersive environment, and to also gather extra data to submit our WiP (Work-inProgress) submission for the International Conference on Entertainment Computing [62]. During a high-school visit to our University on the 27th of April 2022, we ran in-person playtests on sixteen students (figure 81). Each participant explored both developed variants of our virtual environment: one variant used a keyboard and mouse (with a fixed altitude) to move and explore the environment; the second variant, used a gamepad controller (with the ability to fly). During these tests we recorded user analytics, including play time, and user position and orientation tracking within the area. We subsequently asked our participants to answer a short feedback questionnaire where they were asked to describe their experience in five words, what variant they preferred (keyboard or gamepad controller), and lastly, a line for comments and/or suggestions. The average play time for the keyboard and mouse was 53 seconds whilst the gamepad controller was 62 seconds. The most popular words were “Realistic” (11.8%), “Fun” (11.8%), “Interesting” (9.4%) and “Incredible” (7.1%). Two thirds of the participants preferred the gamepad controller over the keyboard and mouse, however, this factor was suspected to be linked to the young demographic so further tests were planned to be carried out in the future, on a wider audience. In terms of user position tracking and despite the area having a central focal point (being the Tilia tree), we noticed that most users proceeded to explore the entire space, as can be seen in the user position trackers (figures 82 and 83). Furthermore, we also noticed that most of our participants had the urge to travel to the next area, wanting to explore more areas. Last but not least, some users managed to bypass the set boundaries (wall colliders) of our area, an issue that eventually got resolved by increasing the thickness of the wall colliders. In general, these preliminary results showed a positive evaluation and provided enough feedback to guide us to continue to expand and improve our virtual world. Figure 81: “Girls in ICT day” at the Department of Informatics Engineering of the University of Coimbra, where we ran in-person playtests on sixeteen highschool students.
82 Figure 83: User position tracking of the gamepad controller variant with the ability to fly. Figure 82: User position tracking of the keyboard and mouse variant with a fixed altitude. Each colour represents a participant and the arrows show in what direction the player was looking.
6. Development of Kairos
84 The development of the main practical component of this thesis was started by establishing the number of areas we were going to capture, what natural processes we were going to include within each area, how we would interconnect these locations along with how the user would interact with said areas. Having already visited the Botanical Garden multiple times, we opted for three locations highlighting some of the oldest iconic tree specimens of the Garden: the Tilia x europaea [9], situated near the Alameda das Tílias, the Erythrina Crista-Galli [10], located within the Central Square, and the Ficus Macrophylla [11], located beside the cold greenhouse (figure 84). Area 1 Tilia x europaea Area 3 Ficus Macrophylla Area 2 Erythrina Crista-Galli Figure 84: Map of the Botanical Garden depicting the three chosen locations of capture.
85 A strategy and course of action was devised on how we would capture these three locations. Initially, the capture of these areas was planned within the closest possible time frame, to conceive the idea of replaying a memory of a walk in the Garden, but because each location requires a large set of photographs, taking around 30 to 50 minutes per area, we decided to split the capture session into three separate parts: morning (for the Tilia), noon (for the Crista-Galli) and midafternoon (for the Ficus), simulating a day’s passage. Time was spent within each area before commencing our captures, by gathering and analysing information on the visible natural processes occurring in each area. Scientific information was then collected of the trees root system structures and their symbiotic relationships with their surrounding environment [63][64] (figure 85). This was inspired by the hidden underground world of microbes, with emphasis on mycorrhizal networks [65], that is, the symbiotic relationships that form between fungi and plants (figure 86). 6.1 Data Capture and Processing Data capture commenced in a chronological order, starting with the capture of the Tilia x europaea, followed by the Erythrina Crista-Galli and ending with the Ficus Macrophylla. A lot more care was taken during the capture process, as these data would be used for our final immersive environment. Because of this, a higher number of photographs were taken to improve the diversity of our dataset, and thus, the quality of our point clouds. Our initial plan to capture all locations on the same day failed due to the constant change of weather conditions, along with exceeding our camera’s limited battery and storage capacities. Consequently, each area got captured on different dates. Moreover, the first two areas became a priority due to their tree’s limited flowering periods, giving a span of about two weeks to capture and process these two. Fortunately the Ficus Macrophylla keeps its leaves and fruit (figs) all year round, making it less of a priority at the time. The successful captures of our areas took place on the 8th of June from 15:56pm to 16:30pm (Ficus Macrophylla), on the 9th of June from 11:52am to 12:38pm (Erythrina Crista-Galli) and on the 12th of June from 9:20am to 9:44am (Tilia x europaea) Our photographic dataset was imported and analysed within Adobe Lightroom. Here the process of colour correction ensued in an attempt to bring out as much detail from the Figure 85: “World Wide Web” – Illustration depicting trees symbiotic relationship with their surrounding environment. Available in: https://www arbioperu org/ en/blog-posts/mother-trees/ Figure 86: A small pine tree grown in a glass tank shows the amount of white, finely branched symbiotic mychorrhizal threads that help feed the tree. Available in https://www csuchico edu/ regenerativeagriculture/blog/how-fungus-might-saveus shtml
86 Figure 87: Colour correction, within Adobe Lightroom, of the area containing the Erythrina Crista-Galli, in an effort to bring out as much detail from the scene as possible. scenes as possible (figure 87). Lightroom’s “Match Total Exposure” feature was also used to automatically equalise the exposure value of all photographs in hopes to better Metashape’s feature tracking process. Although not recommended by Agisoft Metashape’s user manual, we had no choice but to compress our photographic dataset to the lossy compression format JPG, to avoid running out of memory (RAM) during Metashape’s camera alignment process. Albeit having already gained experience from our preliminary tests, new issues arose concerning failed camera alignment, running out of system memory, and failed object alignment within the scenes, consisting of multiple misaligned tree trunks and branches (figures 88 and 89). Paradoxically, these issues were only solved by reducing the size of our photographic dataset through manual selection, where photographs deemed redundant and problematic got removed, such as those with a lot of occlusions and ones with similar angles of capture. From an initial photographic dataset ranging from 700 to 1100 photographs per area, our final dataset consisted of a range of only 400 to 600 photographs. During the removal process, we reiterated the photograph alignment and “build dense-cloud” processes several times until we ended up with a successful point cloud.
87 The area containing the Erythrina Crista-Galli was the most problematic due to its high concentration of vegetation, low prominent feature-tracking points and the constant change in weather and lighting conditions during our capture sessions, which negatively influenced the accuracy and outcome of our point cloud. We ended up re-capturing the Tilia x europaea and Erythrina Crista-Galli several times due to the similar issues mentioned above. Figure 89: Representation of failed object alignment within the scene containing the Crista-Galli, note the duplicate branches and wooden tree support poles. Figure 88: Representation of a failed camera alignment within the Tilia scene, note the strange camera positioning at the bottom right corner.
88 When it came to the Ficus Macrophylla, both capture and processing of this scene were concluded in a single run with ease, perhaps this was due to the higher prominent feature count of the scene, together with more stable weather conditions during the capture session (figure 90). Figure 90: Representation of a successful camera alignment within the scene containing the Ficus (sparse cloud), where the blue rectangles represent every photograph taken. Figure 92: Localised merging, scaling and alignment of the Crista-Galli’s hollow trunk, using data from two separate point clouds (Left); Before (Centre); After (Right). Figure 91: Alignment of two separate point clouds of the same area (highlighted in red and yellow), in an effort to enrich our final point cloud. Despite having already carried out multiple capture sessions on the Erythrina Crista-Galli, each session resulted in an insufficient point cloud dataset. For instance, the first capture failed to reproduce the tree’s leaves, flowers and branches, but successfully managed to capture the tree’s hollow trunk with great detail. The second and third captures resulted in point clouds with a lot more leaf, branch and flower detail but failed at reproducing the tree’s trunk. This issue got resolved by merging, scaling and aligning sections of point cloud data from all three captures into one (figures 91 and 92).
89 6.2 Data Refinement Proceeding with the refinement of our point clouds, an initial manual removal of unnecessary and out-of-bounds points within Metashape was executed using its free-form selection tool (figure 93). Subsequently, all undesirable point cloud data such as blue sky and white cloud points got removed through Metashape’s colour selection tool, together with CloudCompare’s RGB filter tool, with a custom defined tolerance percentage for each and every colour removal iteration (figure 94). Following the major clean-up work, we were still left with a large volume of points that needed subsampling (i.e. decrease the number of points) to preserve performance for Unity. This process got initiated using CloudCompare’s “Remove duplicate points” feature, where points with a radius of less than 0.0099 units (CloudCompare’s units of distance) from one another got removed. This efficacious tool alone was able to remove more than half the total point count, freeing up space while maintaining visual coherence. These results conclude that our photogrammetry derived point clouds contained a large number of nearly overlapping points. Figure 93: Raw dense cloud of the area containing the Ficus tree, consisting of 344 million points (Top). Dense cloud after manual clean-up, now with 261 million points (Bottom).
96 Figure 106: The small Lasioglossum Dialictus (Left) and the larger Apis mellifera (Right), seen on the Tilia at the Botanical Garden. Several previously observed natural processes (that were seen during out visits to the Botanical Garden) were implemented. Out of these, two distinguishable bee species were observed and registered within the area containing the Tilia x europaea, which were at the time gathering nectar from the Tilia’s fragrant, pale yellow flowers: the small Lasioglossum Dialictus and the larger Apis mellifera (also known as the European honey bee), (figure 106). These were both simulated within VFX graphs and also include user collision detection. Moreover, the Tilia’s leaf position data was used as input for the spawn positions of the smaller bees (encompassing the entire tree, as was observed in real life). Two swarms of the larger honey bees were manually placed within the area, positioned in two separate locations. Here proximity sensors were implemented that trigger when the user gets too close, this incites the bees which then start revolving around the user in a more agitated fashion (figure 107). Figure 107: Representation of the Apis mellifera honey bees within Unity.
97 Several Pieris rapae butterflies were observed fluttering within the area containing the Erythrina Crista-Galli, in the rose patches adjacent to the tree (figure 108). These got simulated by animating two 2D textures (emulating wings) by scaling them following a sign wave (see figure 109). Likewise to the bees, these butterflies also react to the user’s presence and disperse when the user gets too close (figure 110). Figure 108: Pieris rapae butterflies. Available at: https://species wikimedia org/wiki/Pieris_rapae#/media/File:Pieris_rapae_ MHNT jpg Figure 110: Representation of the Pieris rapae butterflies within Unity. Figure 109: Left: screen capture of the butterflies VFX graph, showing the procedure of animating their wings by scaling two 2D textures following a sine wave. Right: Illustrative representation of the wing animation.
98 To accentuate the Crista-Galli’s red fragrant flowers, it was decided to simulate the smell of nectar and pollen being released from these (figure 111). Particle systems where positioned in the centre of each flower cluster, which release small particles (representing the nectar and pollen). Furthermore, each cluster contains its own proximity sensor that bursts particles if the user decides to collide or come into contact with the flowers. Having an initial burst quantity of 600 particles, for every consecutive collision, its burst quantity gets divided by three (figure 112). Figure 111: Flower cluster of the Crista-Galli, rich in nectar. Figure 112: Screen capture of the Crista-Galli’s flower particle systems – speculating the release of pollen and nectar.
99 During the capture session of the Ficus Macrophylla, it was noticed that a lot of its figs had fallen to the ground (figure 113), this is also evident in the final audio recording of the area, where we can hear figs dropping. Hence, it was decided to simulate these figs using two VFX graphs. The first randomly spawns figs on the ground on scene load (figure 115), the second emits figs from the tree (using the tree’s leaf-position data as input), where a random fig gets dropped (emitted) within a random time-space between four and eight seconds. A colour sample range was then added from an actual fig to diversify the colours of our particle system (figure 114). Additionally, a subtle force is applied to their negative y-axes, simulating gravity. For added realism, once these fall off the tree, they lightly bounce off the ground (via a floor collider) before coming to a halt. Small clusters of flies were also spotted within the area (figure 116) and implemented within the scene, following similar parameters used for the bees surrounding the Tilia, with the addition of trails (figure 117). As a last note, several birds were observed eating figs off the ground, however, these did not end up getting simulated due to time constraints. Figure 113: Photograph of fallen figs under the Ficus. Figure 114: Photograph of one of the Ficus’s figs, together with its colour gradient sample. Figure 115: Representation of fallen figs within Unity, highlighted in orange.
100 Figure 116: Small clusters of flies spotted near the Ficus. Figure 117: Small clusters of flies simulation within Unity.
101 6.4 Locomotion and Area Traversal Taking the feedback data from the preliminary evaluation tests into consideration, it was decided to keep both modes of user interface: the keyboard and mouse variant, and the game-pad controller variant, both configured using the “fly camera” mode (granting the ability to fly). Subsequently a script got implemented that continuously detects if a controller is connected or not and switches the input device accordingly. Below is the splash screen created for Kairos, containing a diagram of the mouse and keyboard controls as well as the gamepad controller.
102 Our final area traversal system is based on the implementation from the high fidelity prototype. Figure 118 shows a top view of the Botanical Garden containing all three captured areas, including an illustration of the area traversal system where existing pathways within each area were used as exit and entry points. Having noticed that the area containing the Erythrina CristaGalli included four pathways, it became apparent to take advantage of this by installing two exit and two entry points to and from the other areas – a middle ground between the three (figure 119). Adapted from the area traversal script used for our high fidelity prototype, this system controls the entry and exit points of our three areas, allowing the user to travel from one area to another. To achieve this, trigger boxes were implemented and controlled by a script that constantly calculates the user’s distance from the exits – allowing for a gradual scene fade-out. These sensors detect the user’s proximity and start to fade the scene once the user gets closer than 10 units of distance from either one of the three sensors. Subsequently, the user gets transported to the next area the moment they touch the trigger box. Figure 119: Top view of Area #2 Erythrina Crista-Galli, depicting its entry and exit points. Bottom Right: Snip of our area traversal script, of a single traversal rectangle, showing its three sensors (Objects 1, 2 and 3), their respective distances to the User (camera), and the Alpha value of our fade transition. Area #2 – Erythrina Crista-Galli Area #2 – Erythrina Crista-Galli Area #1 – Tilia x europaeaArea #1 – Tilia x europaea Area #3 – Ficus Macrophylla Area #3 – Ficus Macrophylla Figure 118: Top view of the Botanical Garden illustrating all three areas, together with their entry and exit points. Adapted from Google Maps.
103 Figure 120: Four of the five smartphones used for our first recording attempt. 6.5 Sound Design An experimental approach was taken to achieve our goal of creating a 3D ambient sound reproduction for our immersive experience, consisting of installing multiple microphones within each of the captured areas. Before commencing any audio recordings at the Botanical Garden, a preliminary audio test was performed in an effort to determine if our audio capture plan was in fact feasible. Two smartphones were employed for this test: the Sony XZ1 and XZ Premium, which were placed on a lawn (in a garden), separated by roughly five meters. Each phone was set to record at its highest available quality (WAV format with a 48kHz sample rate). After starting both recordings, a loud clap signified the start, to help synchronise and merge these into a single audio file. The resulting audio quality was decent (relatively speaking), but most importantly we had successfully created a binaural audio recording, by simulating sound as if it is being heard live (download link available: “Lawn 3D”, page 179). Having concluded this preliminary test, we returned to the Botanical Garden where we installed five smartphones (figure 120) which got evenly distributed within each scene. Unfortunately, this equipment got stolen just after concluding the last set of recordings, leaving us without our audio capture devices and audio files. Adhering to our initial plan, a second round of audio capture was done with a set of new microphones, including ones of a higher calibre: the H1 Zoom, the H2n Zoom, our Sony a6300 camera and lastly, a Sony XZ1 smartphone (figure 121). These devices were evenly distributed around each scene in a fashion resembling a square (figure 122). Figure 121: The H2n Zoom, Sony a6300 camera, Sony XZ1 smartphone and H1 Zoom – used for our final audio recordings. Figure 122: The positioning of our equipment within the scene containing the Tilia, resembling a square.
104 Due to unpredictable weather conditions, windscreens were made for these microphones in a hope to minimise unwanted wind noise. Each area was recorded (using a WAV format with a 48kHz sample rate) for a period of around 10 to 15 minutes. However, before initialising our audio recordings, microphone gain was fine tuned for the H1 Zoom, H2n Zoom and Sony a6300 to determine their optimal signal-to-noise ratio (figure 123). Having successfully recorded of all three of our areas, all audio files got imported into Premiere Pro, where the synchronisation, clean-up and enhancement processes were initiated. The synchronisation and clean-up of our clips was done by picking out the peaks in the audio tracks (representing the claps), followed by cutting out all unwanted sounds comprised of human chatter, wind noise and traffic sounds. The removal of these sounds was decided due to their overly distracting nature, where these were presumed too interfering for the delicate character of the areas, especially in regard to the areas containing the Tilia x europaea and Erythrina Crista-Galli A cross-fade audio effect was then used as a buffer, to smooth out the audio cuts between the separated (deleted) clips. Next, noise reduction, treble, bass and amplification adjustments were applied in order to unify our audio as a whole (figure 124). Ultimately, all tracks got imported into Unity. Here speakers were placed in the exact same positions as the microphones had been placed in the real world (figures 125 and 126). Each speaker was then transformed into a 3D sound emitter where its “spatial blend” parameter along with its min. and max. emission distances were adjusted to create our desired 3D audio effect. Figure 123: Microphone gain adjustments of the H1 Zoom and H2n Zoom, together with their respective windscreens. Figure 124: Audio editing within Premiere Pro, showing the audio adjustment parameters used in an attempt to unify all audio tracks as a whole.
105 Figure 126: Screen capture depicting the placement of the 3D sound emitting speakers within the virtual scene, in the exact same positions as the microphones had been placed in the real world. Figure 125: Microphone placement within the area surrounding the Crista-Galli – taking advantage of the garden’s shrubs as natural wind-shields.
112 7.1 Virtual Reality Setup and Development Our quest to implement a Virtual Reality variant initiated at the same time as our high fidelity prototype was being developed, where the possibility of running our immersive experience within a VR headset was explored. This began with the HTC Vive (2016) via Steam’s VR platform (SteamVR). Nevertheless, after importing our point cloud into Unity via Keijiro’s point cloud importer, our project failed to run – throwing an unexpected error. This got resolved by changing Unity’s graphics API from Vulkan to DirectX 11, however, this change also meant that we were getting a considerable performance reduction of about 20 frames per second. Besides this, stuttering issues became apparent due to high latency as a consequence of Vive’s constant minimum refresh rate of 90Hz (90 FPS) with a minimum latency of 11.1 ms at a resolution of 3024 x 1680, or 1512 × 1680 per eye. After delving into Steam VR’s settings we were able to slightly improve runtime performance by lowering the value of the “Throttling and Prediction” setting, by forcing the hardware to run at a fixed slower rate in order to provide an overall smoother experience. In spite of our efforts, the high latency issue persisted, resulting in an overall poor VR experience. As a last resort, our render method was switched over to Unity’s VFX graph, where our output points were chosen to be rendered as primitives (quads with a fixed size). Nevertheless, only when our point count was reduced to 2 million (originally 20 million) were we able to achieve an acceptable latency rate, this, however, drastically reduced the visual fidelity of our scene (figure 132). For this reason we decided to put the HTC Vive aside to look for another solution. Figure 132: Screen capture showing poor visual fidelity, resulting from the reduction of the point count to 2 million.
113 As a second attempt we tested the Oculus Quest 2 (2020). This equipment seemed promising due to its higher screen resolution and wireless tracking capabilities. To take advantage of this headset’s internal, Android-based operating system, we initially tried building our project as an app for Android, as this would permit the launch of our immersive experience wirelessly (as a standalone). However, with this method we were relying on the Quest’s Qualcomm Snapdragon XR2 chip to render all of our points. Unfortunately the XR2 was only able to render a meagre 200 thousand points before running into performance issues. Nonetheless, by connecting the Quest 2 directly to a PCVR ready computer, via a Link cable, were we able to take advantage of our higher performance PC hardware to render graphics, instead of the standalone mobile hardware in the headset. Moreover, this VR headset did not have any issues running off the previously problematic Vulkan graphics API, meaning we had an additional performance boost. Several benchmarks were performed using the Link cable for the purpose of determining the maximum quantity of points we could render at a sustainable frame rate. Although not perceptible at first glance (within the headset), we later discovered that we were getting low and variable frame rates, likely due to the fact that we were attempting to render our scene at a resolution of 3616 x 1840 (or 1808 x 920 per eye), at a constant refresh rate of 72 FPS (the Quest’s default refresh rate). For this reason our scene was changed to the Erythrina Crista-Galli – a computationally lighter scene due to its lower global point count. An initial performance optimisation was made by manually lowering the point count within the VFX graph as a means of reaching a stable frame rate of 72 FPS. Subsequently, all animations got imported into the scene (root systems, butterflies, trunk, flower and leaf animations). This, however, negatively impacted performance once again, this time making our scene fluctuate from 72 FPS to 36 FPS. A more structured approach was taken to try and stabilise the frame rate, as lowering the point count further would only lead to a greater loss in visual fidelity. This approach consisted of splitting the point cloud into several sections and subsections (ground, scene extremities, tree, tree’s leaves, flowers and scene bushes), as was done during the development of the PC variant of Kairos, allowing the reduction of our scene’s point count to be made in a more controlled fashion, by maintaining more detail (more points) at the focal point of the scene and reducing the point count at the extremities of the scene (figure 133). Figure 133: Process of separating the scene into sections and subsections in an effort to improve performance whilst maintaining as much detail as possible.
114 Two options arose when it came to user interaction: to work with the Quest’s controllers or experiment with its recently released hand tracking capability, the latter being able to detect our hands and fingers in real-time, through four IR (infrared) cameras located on the Quest’s headset, then displaying a model of them within the scene. This hand tracking system was explored first, as its new hand and finger detection system was released shortly after starting our VR development (released on April 19, 2022). In order to allow this system to work within Unity, the download and import of Unity’s OVRPlayerController was necessary, along with manually attaching its hand prefabs to both hands. Our goal consisted of allowing the user to roam around and interact with the scene solely using this hand tracking Figure 134: Screen capture of our final VR scene containing 2.8 million points. This, however, was not sufficient to stabilise the frame rate and a further reduction in the global point count had to be done. From the scenes original point count of 20 million points, our final VR area consisted of only 2.8 million points. As a last resort, the size of our point particles was increased to add more volume to the scene in an attempt to counteract its low point count (figure 134).
115 feature, through gestures and by attaching sphere colliders and attractors to either hand. We however were only able to successfully implement the latter and had to find another solution for user locomotion. Having failed to move around the scene solely with hand tracking enabled, it was decided to activate the Quest’s controllers, along with modifying the official OVRPlayerController script to allow the user to “fly” within the scene by pointing in the desired direction of flight via the headset’s rotation/tilt, and accelerating via the Left trigger button. Our next task consisted of merging the controller’s functionality with the hand tracking system, unfortunately the Quest 2 does not support the use of both input methods simultaneously. For this reason it was decided to keep both variants, as separate modes, where one can fly around the scene using the controllers, then, by placing down the controllers, Unity automatically switches to hand tracking where the user can then interact with the scene. Below is an illustrative diagram of the controls and input methods implemented for our Quest 2 variant.
116 The Right controller and Right hand contain a sphere collider and attractor to interact with the scene (figures 135, 136, 137, 138 and 139), however, an issue arose when trying to represent both controller meshes and hand meshes during these input transitions, where both meshes would overlap. Although imperfect, it was decided to prioritise the hand tracking system over the controllers, by hiding them. Moreover, the reason behind adding only one sphere collider to the Right controller and hand was decided due to Unity’s incapability of performing multiple collider and attractor forces (for both Right and Left outputs) on a single VFX particle system. The only way around this issue consisted of splitting each of our VFX-based point cloud particle systems in half (50% for the left hand and 50% for the right hand), this, however, resulted in an uneven attraction effect because only half the particles would be attracted to either hand, leaving some particles unaffected by their collider and attractor forces. Figure 135: Representation of the Quest’s hand tracking, together with the implemented sphere collider – highlighted in red. Figure 136: Demonstration of the hand collision with the tree’s trunk. Figure 137: Demonstration of a butterfly dispersing as the user tries to catch it.
117 Figure 138: Demonstration of the flower’s “burst” effect upon hand collision. Figure 139: Demonstration of the attracting force between the user and the tree’s energy. In comparison to the interactive elements of the Desktop (PC) variant, an addition of vegetation interaction was implemented in the VR variant, allowing the user to interact with the scene’s shrubs, along with the leaves of neighbouring trees. As a final point, all three locations were originally planned on being implemented for our VR variant, however, due to performance limitations, it was decided to solely include the scene containing the Erythrina Crista-Galli.
118 7.2 Passthrough Exploration In an effort to produce more content for the Quest 2, it was decided to explore its experimental passthrough API. Passthrough is a mixed reality-focused feature that allows the user to see a real-time 3D visualisation of the physical world in the Oculus headset via its IR sensors. This feature brought up the idea of creating an in-situ variant consisting of a mixed reality experience where the user could theoretically visualise the hidden processes of the Erythrina Crista-Galli by overlapping our point clouds with the physical world. A preliminary point cloud capture test of our residency was made before running any in-situ tests. This consisted of rendering our point cloud, via Unity’s VFX graph, over the Quest’s passthrough stream. After correcting the scale, rotation and position of our point cloud (relative to the physical world), we were able to create an interesting parallelism between the real and virtual worlds, allowing us to experience two universes simultaneously, so to speak (figure 140). Figure 140: POV of the Quest’s passthrough test, showing an overlapped representation of two realities.
119 Figure 141: In-situ equipment setup. An attempt to explore the possibilities of this idea was conducted in-situ (figure 141), however, due to technical issues related with the low portability and power requirements of our setup, we were not able to carry out our plan of action within the time-span of this thesis. Nevertheless, figure 142 shows a simulation of what our insitu variant could have looked like, which is something that may be of interest for future exploration.
120 Figure 142: A simulation of the in-situ variant using the Quest 2 (POV), where one could view the scene’s hidden processes, along with interacting with said processes through physical contact.
8. Public Feedback Evaluation
128 Not at all Participants Figure 154: How well Kairos improved our volunteers appreciation of the natural world. Greatly These were followed by a set of three multi-choice linear scale questions concerning Kairos’s impression on our volunteers. Here we asked how much the experience had improved their appreciation for the natural world, if it made them feel any closer to nature, and if the experience had increased their desire to explore nature in the real world. On a scale of 1 out of 10, the majority of our volunteers rated Kairos as having improved their appreciation for the natural world (average of 7.83), made them feel closer to nature (average of 8.55), and an increase in desire to explore nature in the real world (average of 8.27) (figures 154, 155 and 156). Not at all Participants Figure 155: How well Kairos made our volunteers feel closer to nature. Much closer
129 Last but not least, a final set of open questions were given, asking if any other issues were experienced during run-time, what they would add to the experience, followed by a section for comments and/or suggestions. Some volunteers expressed their desire for more interactive elements within the scenes, along with a guide (or a minimap) to show the user where to go, as some would become lost within the areas. Moreover, smaller suggestions were also left, such as the addition of a blue sky and stars. Not at all Participants Figure 156: If Kairos increased our volunteers desire to explore nature in real life. Very much so
9. Making Kairos Publicly Available
132 Showcase material was prepared with the intention of captivating the public’s interest and to be shared with the Botanical Garden’s social media platforms and web pages. This task got initiated through the development of a website (https://kairos dei uc pt/), where we explain the ideas and motivation behind our project as a brief introduction, figures depicting how it got developed, two download links for both Desktop and VR variants (figure 157), along with our paper that got accepted at the International Conference on Entertainment Computing. Figure 157: Screen captures of Kairos website. Top: front page. Bottom: Download section containing links for both PC and VR variants. Available at: https://kairos dei uc pt/
133 In addition to this website, and due to the hardware requirements necessary to run Kairos, we elected to create videos consisting of stereoscopic (5K resolution) and monoscopic (8K resolution) 360 degree videos of our captured areas. The stereoscopic variant being aimed for the Oculus Quest 2 and YouTube VR, and the monoscopic variant to be downloaded and viewed locally on a PC (available at https://kairos dei uc pt/). These videos consist of two minute long passages through our captured areas, which were created by setting the scene cameras to follow a pre-determined bézier curve within each area (figure 158). To maximise the output quality of our videos, it was decided to render our scenes as PNG image sequences. This encompassed the use of Unity’s in-built Recorder, where our camera was set to render a 360 degree view of its surroundings, with an output resolution of 8192 x 8192 pixels (stereoscopic over/under format), with a Cube Map Size of 4096 pixels and a frame rate of 60 FPS (figure 159). Figure 158: Screen capture within Unity, showing the camera’s pre-determined path of 360 degree capture.
134 Figure 159: The resulting stereoscopic (over/ under) 360-degree image render of the Tilia, with a resolution of 8192 x 8192 pixels and a stereo separation of 0.065. These image sequences were then imported into Premiere Pro and transformed into videos. In an attempt to guarantee smooth video playback of the stereoscopic variant, for the platforms mentioned previously, the H.265 video codec was used at a bit rate of 120 Mbps. Higher bit rates were tested in the effort to improve the visual quality of our videos, however, these would cause video playback issues within the Quest 2 and YouTube. Monoscopic 360 degree videos were subsequently rendered, using the same image sequences, but this time at a higher bit rate of 400 Mbps. This was done to maximise video quality, however, these are aimed solely for local (Desktop) video playback.
135 On a related note, a cooperation was developed between Kairos and the Department of Informatics Engineering’s “PlugNPlay” display system. This recently implemented system consists of an array of nine vertically stacked and interconnected monitors, providing the capacity of displaying content at a maximum resolution of 9720 x 1920 pixels (figure 160). The endeavour of executing Kairos on this system stemmed from a mutual curiosity to determine if Kairos could successfully be run on this high resolution display at a stable frame rate of 60 FPS (figure 161). Surprisingly this was the case, which lead to a new task: to create a Kairos variant as a demo. for the inauguration of PlugNPlay. Figure 160: PlugNPlay’s powerhouse, used to feed its nine displays, with the help of its RTX 3080 Ti graphics card. Figure 161: Initial testing of Kairos on the PlugNPlay system. A few adjustments needed to be made in order to fully harness the capabilities of PlugNPlay, this included adjusting the scene’s cameras FOV (Field Of View), to encompass the display’s wide aspect ratio (approximately 5:1) (figure 162), along with positioning multiple stationary cameras, at different angles, within each scene (totalling two per scene). A script was then added to automatically (or manually) transition between these cameras and scenes, similar to live wallpaper background (figure 163).
136 Lastly, a teaser trailer was created, consisting of an interlaced interplay between video footage from the Botanical Garden with its digital point cloud representation (available at: https://kairos dei uc pt/ ). A custom audio track, created by Coimbra-based music artist Samuel Cadima, was also added, which couples with a voice-over, as a prelude to the natural processes occurring at the Botanical Garden of the University of Coimbra. Figure 162: Pre-testing Kairos PlugNPlay variant within Unity, at a resolution of 9720 x 1920. Figure 163: Kairos running on the PlugNPlay system.
10. Conclusion