Full text
Tesis Doctoral Artificial Intelligence in Architecture A Journey from Engineering Applications to Generative Design and Conceptual Representations of Space Realizada por Jaime de Miguel Rodr´ıguez Para la obtenci´on del t´ıtulo de Doctor Dirigida por Fernando Sancho Caparrini En el departamento de Ciencias de la Computaci´ on e Inteligencia Artificial Programa de Doctorado de Ingenier´ıa Inform´atica 2024
Dedicado a mis Padres por su enorme paciencia, a Fernando Sancho por su inestimable apoyo y a mi T´ıa Triny Rodr´ıguez-Burgos por su dedicaci´on incansable al conocimiento.
Agradecimientos Quiero expresar mi m´ as profundo agradecimiento a Fernando Sancho, por much´ ısimas cosas, pero en especial, por haber valorado la innovaci´ on de mi trabajo en un estado muy incipiente, y en un momento muy dif´ ıcil de mi trayectoria. Sin la confianza que me brind´ o el contar con su respaldo, ni esta tesis, ni las aportaciones que m´ as estimo de la misma, hubieran sido posible. Tambi´ en a Joaqu´ ın Borrego, por haberme dado la oportunidad de comenzar mis estudios de doctorado en el Departamento de Ciencias de la Computaci´ on e Inteligencia Artificial de la Universidad de Sevilla, y por haber dirigido con enorme paciencia mi primera etapa en el programa. A Pedro Almagro, por su explicaci´ on en profundidad del proceso de entrenamiento de las redes neuronales en el Seminario de Aprendizaje Autom´ atico del departamento (generosamente organizado por Fernando Sancho), la cual me abri´ o la puerta a comprender muchos otros modelos que se mencionan en esta tesis. A Mar´ ıa Eugenia Villafa˜ ne, por haberme motivado y animado a acometer el estudio de m´ etodos generativos para la fusi´ on de tipos arquitect´ onicos, con vistas a la conferencia AAG 2018. A Antonio Morales, Mar´ ıa Victoria Requena y Emilio Romero, por introducirme en el campo de la ingenier´ ıa s´ ısmica y por todo el apoyo recibido durante los proyectos PERSISTAH y SIMRIS. A Matthew Peavy, por su generosidad en revisar mis trabajos y sus comentarios siempre acertados. A Pille Bunnell, por su serie de v´ ıdeos Dancing With Ambiguity, que fueron una gran fuente de inspiraci´ on para el desarollo del m´ etodo propuesto en esta tesis sobre la generaci´ on emergente de estructuras conceptuales. En relaci´ on al m´ etodo para la generaci´ on emergente de estructuras conceptuales, tambi´ en quisiera mostrar mi agradecimiento a Francisco M´ arquez, por sus excelentes clases sobre el mundo de la identidad y el mundo de la diferencia. Dichas clases despertaron mi inter´ es por los aspectos cognitivos del dise˜ no arquit´ ectonico en mis primeros a˜ nos de carrera, y me llevaron a encontrar los trabajos de Gregory Bateson sobre los que este m´ etodo se basa. A Yusuke Obuchi, por demistificar y darme las claves conceptuales para i
introducirme en el mundo la programaci´ on inform´ atica. A Fabio Gramazio, por descubrir en m´ ı la urgencia por la innovaci´ on, un reconocimiento que me dio la seguridad necesaria para acometer proyectos en disciplinas que me eran ajenas, como la Inteligencia Artificial y la Ciencia Cognitiva, protagonistas de esta tesis. A Ena Lloret, por el camino de investigaci´ on que recorrimos juntos en la Universidad ETH de Z´ urich, que no hubiera sido posible sin ella. Finalmente, un agradecimiento muy especial a Mª Carmen Cardona, por toda su ayuda en la gesti´ on de tr´ amites administrativos y de tantos apuros de ´ ultima hora a lo largo de estos a˜ nos. ii
Resumen Esta tesis explora el papel de la Inteligencia Artificial (IA) en el modelado conceptual del espacio, centr´ andose en la est´ etica arquitect´ onica, o venustas, y abordando tambi´ en utilitas (funcionalidad) y firmitas (estructura). El aspecto est´ etico, siendo el m´ as desafiante, es central en la investigaci´ on, que examina la interacci´ on entre la IA, la est´ etica del dise˜ no y las representaciones espaciales. Se examinan los desaf´ ıos y oportunidades que la IA presenta para la creatividad, el razonamiento conceptual y la representaci´ on espacial: tres pilares fundamentales en el discurso arquitect´ onico. Se analizan varias t´ ecnicas de IA, incluidas las redes neuronales para aplicaciones estructurales en ingenier´ ıa civil y s´ ısmica, y los modelos generativos para el an´ alisis urbano y la b´ usqueda generativa de formas no triviales. El An´ alisis Formal de Conceptos se introduce como una herramienta ´ util para el estudio de datos urbanos. La tesis contin´ ua analizando cr´ ıticamente los modelos generativos de IA, evaluando su potencial creativo, su capacidad para soportar representaciones complejas y ambiguas, y su idoneidad para operar a un nivel conceptual. Si bien la IA generativa muestra promesas en la generaci´ on de soluciones de dise˜ no creativas, est´ a limitada en la realizaci´ on de operaciones impulsadas por conceptos, particularmente en comparaci´ on con los modelos simb´ olicos, que destacan en el razonamiento, la composibilidad y la explicabilidad. La investigaci´ on finalmente profundiza en el problema de la representaci´ on espacial en el dise˜ no arquitect´ onico. Se revisan varias aproximaciones simb´ olicas y neuro-simb´ olicas, aunque ninguna aborda completamente los complejos aspectos cualitativos del espacio, fundamentales para la cognici´ on humana y la pr´ actica arquitect´ onica. En respuesta, esta tesis propone un nuevo modelo piloto que revisita los enfoques puramente simb´ olicos. Este modelo proporciona una base para la emergencia de estructuras conceptuales a partir de datos sensoriales, abordando el prolongado problema del paradigma simb´ olico en IA y ofreciendo una posible soluci´ on para la representaci´ on de objetos espaciales en el dise˜ no arquitect´ onico. En ´ ultima instancia, esta tesis contribuye al creciente di´ alogo entre la IA y el dise˜ no arquitect´ onico al proponer un nuevo marco para integrar la IA en los procesos creativos y conceptuales que definen la disciplina. Tambi´ en ofrece una evaluaci´ on cr´ ıtica de los m´ etodos actuales de IA y sugiere un enfoque novedoso que podr´ ıa allanar el camino para futuros avances en IA y Dise˜ no. iii
Abstract This thesis explores the role of Artificial Intelligence (AI) in the conceptual modelling of space, focusing on architectural aesthetics, or venustas, while also addressing utilitas (functionality) and firmitas (structure). The aesthetic aspect, which is the most challenging, is central to the research, which examines the interplay between AI, design aesthetics, and spatial representations. It examines the challenges and opportunities that AI presents for creativity, conceptual reasoning, and the representation of space —three pillars that are central to architectural discourse. Various AI techniques are analysed, including neural networks for structural applications in civil and seismic engineering, and generative models for urban analysis and form-finding. Formal Concept Analysis is introduced as a key tool for conceptual modelling in AI, applied to urban case studies. The thesis continues by critically analysing generative AI models, evaluating their creative potential, their capacity to handle complex and ambiguous representations, and their suitability to perform at a conceptual level. Although generative AI shows promise in generating creative design solutions, it is limited in performing concept-driven operations, particularly compared to symbolic models that excel in reasoning, composability, and explainability. The research finally delves into the problem of spatial representation in architectural design. A variety of symbolic and neural-symbolic approaches are reviewed, yet none fully address the complex, qualitative aspects of space central to human cognition and architectural practice. In response, this thesis proposes a novel proof-of-concept model that revisits purely symbolic approaches. This model provides a foundation for emerging conceptual structures from raw sensory data, addressing the long-standing symbol grounding problem in AI and offering a potential solution for representing spatial objects in architectural design. Ultimately, this thesis contributes to the growing dialogue between AI and architectural design by proposing a new framework for integrating AI into the creative and conceptual processes that define the discipline. It offers a critical assessment of current AI methods and suggests a novel approach that could pave the way for future advancements in AI and Design. iv
Contents 1 Introduction .................................. 1 1.1 Brief overview of Concept Theories . . . . . . . . . . . . . . . . . . . . 4 1.2 Concepts in Artificial Intelligence . . . . . . . . . . . . . . . . . . . . . 6 1.3 Discussion.................................. 10 1.4 Structure of the document . . . . . . . . . . . . . . . . . . . . . . . . . 13 1.5 Contributions ................................ 14 1.5.1 Use-Case Specific Contributions . . . . . . . . . . . . . . . . . 15 1.5.2 Contributions to Bridging AI and Architectural Design . . . . 16 1.5.3 Publications associated with the thesis . . . . . . . . . . . . . . 17 2 Engineering Applications of Artificial Intelligence in Architecture .................................. 20 2.1 Background ................................. 20 2.2 Earlyworks ................................. 24 2.2.1 Inductive Learning . . . . . . . . . . . . . . . . . . . . . . . . . 24 2.2.2 Conceptual Clustering or Learning by Observation . . . . . . 26 2.3 Current trends: Pattern Recognition and Deep Learning in Civil Engineering ................................. 28 2.3.1 Relatedworks............................ 36 2.4 Experimental application . . . . . . . . . . . . . . . . . . . . . . . . . . 38 2.4.1 Methodology ............................ 41 2.4.2 Results................................ 58 2.5 Discussion.................................. 64 3 Urban Intermezzo I ............................. 67 3.1 Introduction and literature review . . . . . . . . . . . . . . . . . . . . 67 3.2 Methodology ................................ 69 3.2.1 Pre-processing ........................... 70 3.2.2 Generation and augmentation of training and validation sets 72 3.2.3 Variational Autoencoder model . . . . . . . . . . . . . . . . . . 72 3.3 Results .................................... 76 3.4 Discussion.................................. 76 3.5 Conclusions ................................. 81 v
4 Generative Artificial Intelligence in Design ............ 83 4.1 Background ................................. 83 4.1.1 The connectionist leap forward . . . . . . . . . . . . . . . . . . 86 4.2 Methods ................................... 90 4.2.1 ML models as creative design engines . . . . . . . . . . . . . . 91 4.2.2 Foundation of generative ML methods . . . . . . . . . . . . . 94 4.3 Literature review: (architecture-related) . . . . . . . . . . . . . . . . . 97 4.4 Experimental application . . . . . . . . . . . . . . . . . . . . . . . . . . 102 4.4.1 Introduction............................. 102 4.4.2 Methodology ............................ 104 4.5 Experimentation and discussion of results . . . . . . . . . . . . . . . . 114 4.5.1 Conclusions............................. 134 5 Urban Intermezzo II ............................. 137 5.1 Introduction................................. 137 5.1.1 Urban Informational Ecosystem . . . . . . . . . . . . . . . . . 138 5.1.2 Housingmarkets.......................... 140 5.1.3 The Self-City platform . . . . . . . . . . . . . . . . . . . . . . . 141 5.2 Methodology (Formal Concept Analysis) . . . . . . . . . . . . . . . . 143 5.2.1 Contexts from urban digital footprints . . . . . . . . . . . . . . 144 5.3 Preliminary exploration . . . . . . . . . . . . . . . . . . . . . . . . . . 145 5.4 Conclusions and future work . . . . . . . . . . . . . . . . . . . . . . . 145 6 Conceptual representation of space .................. 149 6.1 Space ..................................... 150 6.2 The Symbol Grounding Problem . . . . . . . . . . . . . . . . . . . . . 155 6.3 Relatedworks................................ 157 6.4 Background method: Formal Concept Analysis . . . . . . . . . . . . . 163 6.5 Experimental proof-of-concept proposal . . . . . . . . . . . . . . . . . 164 6.5.1 Method................................ 169 6.5.2 Experimentation and results . . . . . . . . . . . . . . . . . . . 176 6.5.3 Comparison of results . . . . . . . . . . . . . . . . . . . . . . . 179 6.5.4 Discussion.............................. 181 6.5.5 Conclusions............................. 186 6.5.6 Futurework............................. 188 7 Discussion and conclusions ....................... 195 7.1 Discussion .................................. 195 7.2 Conclusions ................................. 208 vi
Bibliography .................................... 210 vii
6.4 Time series representation of the sample entity . . . . . . . . . . . . . 173 6.5 Formal context table resulting from step 1 . . . . . . . . . . . . . . . . 174 6.6 Resulting concept set after step 1 . . . . . . . . . . . . . . . . . . . . . 175 6.7 Concepts (intent) present across all curves in the sample-set . . . . . 177 6.8 Number of differing concepts with parameters: angle ......... 178 6.9 Number of differing concepts with parameters: angle,width . . . . . 178 6.10 Number of differing concepts with parameters: angle,x,y...... 178 6.11 Number of differing concepts with parameters: angle,x........ 178 6.12 Number of differing concepts with parameters: angle,width,x,y. . 179 6.13 Formal context table resulting from step 2 . . . . . . . . . . . . . . . . 191 6.14 Formal context after step 3 (removed redundant intervals) . . . . . . 191 6.15 Final formal context obtained . . . . . . . . . . . . . . . . . . . . . . . 192 6.16 Concepts C10 –C28 (all remaining concepts) . . . . . . . . . . . . . . . 193 xiv
1. Introduction This thesis is set out as an exploration of Artificial Intelligence (AI) and Machine Learning (ML) techniques in Architecture, with some touches also upon the field Urban Studies. Of course, that is an overwhelmingly broad topic, of which dozens of volumes can be written. Attempting to deliver a comprehensive account of all the different areas within the field of Architecture in which AI and ML are playing an important role is a very broad task. Instead, this thesis is articulated in a series of studies on Architecture and AI, which will be used as a vehicle to weave a research journey: from the domain of patterns to the complex and multifaceted realm of concepts. The way will be marked by case studies and discussions on both ends of the AI spectrum: symbolic and statistical methods (largely represented by connectionism). And also, by combinations of both, namely, neural-symbolic methods, which have been a hot research topic in recent years and remain so today. Architecture as a field has a very long history. Some early accounts of this discipline probably go back to the Egyptian physicist and architect Imhotep [4]. However, the earliest surviving written work on the subject of architecture is “De architectura” by the Roman architect Vitruvius in the early 1st century AD [5]. He established the three foundational principles of Architecture: firmitas,utilitas, and venustas (strength, utility, and beauty, respectively) [6]. Although much has evolved in the field since then, these principles still stand their ground today in one way or another; as well as quite a few Roman constructions (and perhaps this is not a total coincidence). In our current time, firmitas corresponds to the engineering side of Architecture. This is a field that has advanced enormously. And maybe because its objectifiable and more clear scientific nature, it is the one where ML has initially had a greater impact. In particular, the power of modern neural network models in tasks of complex multivariate regressions and pattern recognition among other tasks has yielded a plethora of unprecedented applications in this field. For this reason, firmitas, together with the study of patterns and complex regressions in engineering applications, will be the starting point of this journey. However, the other two attributes: utilitas and venustas, are more slippery in terms of scientific development, especially the latter. The role of function in Architecture had its apex during the Modern Movement (a.k.a. International Style, 1920s – 1970s), spearheaded by architects like Le Corbusier, Ludwig Mies van der Rohe, Walter Gropius, and Konstantin Melnikov, among other prominent pioneers 1
[7]. Another figure, Louis Sullivan, popularised the famous axiom ”Form follows function” [8] which called for prioritising utility in architectural design, in detriment of beauty. The pursuit of architectural beauty by itself was frowned upon during this period. As an example, the book “Ornament and Crime” by Adolf Loos [9], was one of the most influential titles among architectural practices and schools at the time. Despite the relatively recent emphasis that the field has stowed upon function, the application of AI or ML to assist the utilitas of buildings has not been enormously fruitful, generally speaking. The truth of the matter is that for average-sized buildings, fitting a set of programmatic requirements can be equally a relatively simple or extremely complex task. But in neither case does it have to involve a lot of data (which is a common marker for the use of ML). When the task is simple, it is obviously not necessary to involve AI, unless there is a need for a sort of en masse production of designs. Conversely, when the task is difficult, the complexity is typically driven by the need to find compromises between multidimensional aspects that operate mostly on a conceptual level. For example, an architect may have to find a spatial solution that produces a feeling of ’intimacy’, while at the same time balancing a natural-light geometry that is cost-effective and easy to build, possibly in addition to a few other requirements. In this sense, designing may be compared to solving a tricky puzzle. Not one with many pieces, but one with, so to speak, a couple dimensions or more added on top of the conventional three-dimensional space. Solving this puzzle requires not only a good framework for spatial reasoning, but also one that is able to interact with other design constraints. Among them, one may find structural parameters, economic variables and other objectifiable aspects; but also, others less tangible, like cultural habits or psychological factors that may (or may not) render a space suitable for those activities for which the space is designed. In other words, the utility aspects that concern the vast majority of buildings are not as objective or mathematical as the term function may suggest. Additionally, function is a rather flexible term, and its interpretation may vary widely as architects and designers push the boundaries of the discipline. However, these nuances may be overlooked when one considers large building configurations that require complex programmatic functionalities, such as international airports, large hospital compounds and facilities, industrial manufacturing or production plants, etc. In these cases, it is safer to speak of function in a more classical and objective sense. The more these building complexes behave like a system with clear objective constraints and well-defined goals, the greater their potential for AI applications. 2
In this sense, one area on which utilitas has definitely thrived is the field of Urban Analysis. Great examples can be found in the early complex systems modelling of urban city growth during the 1960s and 1970s [10,11], the study of cities as movement economies of Bill Hillier and Space Syntax in [12], and the presence of well-established journals that are riddled with research on cities using numerous AI and ML techniques (such as Computers, Environment and Urban Systems or Environment and Planning B: Urban Analytics and City Science). Following this trend, two urban studies are presented in this thesis: Urban Intermezzo I & II. The first is a en masse building typology analysis for the city of Seville using Variational Autoencoders, and the second is a real estate analysis of the same city based on Formal Concept Analysis (FCA) [13,14]. But, as already mentioned, dealing with more nuanced approaches to functionality and especially, when addressing aesthetics in Architecture and Design, is quite a challenging enterprise. Aesthetics is a highly subjective domain. Any computational approach or even any attempt to articulate some degree of objectivity around it will face great difficulties. To illustrate this point, one may observe some fundamental differences in the way engineering versus design practices are taught in today’s classrooms (at least in the western education system). While most engineering tasks are based on the principle of optimisation and finding the best design or solution to a given problem, aesthetic-based design stands quite on the opposite end of the spectrum: there is never, as a principle, one right answer to any given task. When optimisation, as a criterion, is thrown out the window, assessing and evaluating a design may seem like a hard task (and it is indeed). For this reason, among others, Design as a discipline has evolved to foster the idea of conceptual discourse or narrative as an integral part of the creative process. A conceptual narrative [15,16] can be understood as a tool through which designers can explain and make sense of their creations coherently. In other words, it is a coherent discourse based on a conceptual understanding of their creation process and their produced artefacts. It has both a component of self-guidance or self-understanding as well as of communication (to others). Without going into too much detail, the evaluation process of a student’s work in this domain often lies in assessing the coherence of the discourse presented within a certain logic framework. Typically, in a tutor-student scenario, the set of logic axioms against which the design is assessed is inferred by the tutor from the conceptual discourse itself. It is clear, therefore, that any computational approximation to Aesthetics and Design must establish a strong theory of concepts and conceptual representation, with an emphasis on semantics, because that is precisely the basic toolkit that these 3
subjects are dealt with in contemporary design practice. However, while patterns may be a pretty simple idea to understand, the research around concepts is very heterogeneous and multidisciplinary. Concept Representation constitutes an important research area for, at least, the following fields: Philosophy, Psychology, Neuroscience, Cognitive Science, Linguistics, and Artificial Intelligence. And to make matters even harder, there is quite a wide range of perspectives even in the very understanding of what concepts are [17]. However, it is important to acknowledge that some current ML models, such as Generative Neural Networks, have achieved a great deal in terms of producing creative outputs (in a loose sense of the term creative). These models can compose stunning images from text prompts and even generate new ones from training experience. While they have not been explicitly built to support concepts in the classical notion of the term, their important achievements suggest that they should be taken seriously into account in concept studies. As such, this volume dedicates a full chapter to exploring their potential for concept research. 1.1. Brief overview of Concept Theories Concepts are one of those notions that are easy to understand intuitively but very hard to define formally. Plato and Aristotle already theorised about concepts in one way or another, which speaks volumes of the importance and historic legacy of the idea of concepts. Plato believed that for every category of objects in the world, there existed a pure and uncorrupted instance of it in the world of Ideas or world of Forms [18,19]. Instead, for Aristotle, the world was made of substances (which he later replaced with the term essence). Substances were of two kinds: primary substances comprised independent objects composed of matter and form; secondary substances corresponded to larger groups or categories to which the former objects belonged [20]. Since then, the idea of a concept has been a constant presence in both philosophical and scientific arenas. The eminent philosophers Lock and Hume, for example, were early adopters of the Representational Theory of Mind (RTM) [21]. A theory that to date is still the most prevalent in the field [22]. RTM posits that thinking occurs in an internal system of representation. Beliefs and desires and other propositional attitudes enter into mental processes as internal symbols [21]. However, this view has important detractors that have proposed alternative perspectives. Some of them maintain that concepts should be rather understood as mental abilities [23,24,25] (for instance, the concept water, may tantamount to simply distinguishing water from other entities and drawing some inferences from it), while other detractors view concepts as abstract objects that mediate between 4
language-thought and referents (objects in the world, real or fictional). The main argument here is that a concept may exist outside of any human mind; there exist concepts that humans have never yet entertained only because of our intellectual limitations [26]. Their supporters also point out that a representational theory of mind poses the philosophical challenge that the same concepts may have different representations in every individual. Many supporters of the RTM are also in favour of what has been called the Theory Theory of Concepts [27,28]. According to this postulate, categorisation is a process that strongly resembles scientific theorising. The theory is able to explain some of the complexity around the categorisation aspect of concepts. For example, children have been reported to dismiss the importance of visual similarity when faced with a situation where a dog is intentionally altered to resemble a raccoon. They argue that even at a young age, children have a basic understanding of biology. According to the theory, being a dog goes beyond mere visual resemblance: it hinges on possessing the essence of dogs, whatever that may be [29]. Another advantage their supporters claim is that it helps explain conceptual development over time (concepts acquired in childhood that evolve into adulthood). The implications of this view have important consequences for AI, because it means that concepts are of an extremely subjective nature – extremely hard to formalise. Another important area of discussion is the nativist versus the empiricist view, which also has deep historical roots. Nativists, on the one hand, maintain that there are preexisting innate concepts (not learnt) and that the mind is wired differently for different domains (e.g. perception systems or different levels of abstraction) [21]. Empiricists, on the other hand, believe that there are few innate concepts (or none at all) and that most cognitive abilities are developed from simple cognitive mechanisms [21]. Neither Hume nor Kant, for example, supported the nativist approach [30,31]. However, the debate became much alive in the second half of the twentieth century after Fodor’s contributions to the conversation [32]. Fodor embraced radical nativism claiming that lexical concepts lack semantic structure, and consequently, virtually all lexical concepts must be innate. In this, he only allowed that complex concepts can be learnt, where learning is carried out by assembly or a combinatorial process. Not many (if any at all) have taken these radical perspectives, but not few have reevaluated their views on concept learning after Fodor’s arguments [33,34,35,22]. A further development within the empiricist view is the notion of embodied cognition, or the embodied mind thesis [36]. According to this approach, the concept of ‘water’ and the word ’water’, for example, ”acquire meaning by internal stimulation or reenactment of the perceptual motor and emotional experiences” 5
that are connected to seeing water, touching it or hearing it splash [37,38]. The embodied perspective considers that human concepts are influenced by the kind of body that an organism possess; because it is in their perception, action, and emotion systems where these concepts are grounded [39,40,41]. Some important authors like Harnad (postulator of the Symbol Grounding Problem that will be discussed in a later chapter), have gone so far as to suggest the replacement of the idea of concepts in Cognitive Science, for “inborn and acquired sensorimotor category-detectors and category-names combined into propositions that define and describe further categories” [42]. Embodied cognition has gained a very wide support among the scientific community, especially around Cognitive Science. This theory provides a strong understanding of concrete concepts (those with a clear tangible referent, e.g.: apple). However, according to some authors, it faces important challenges when attempting to explain abstract ones (those without a clear tangible referent, e.g.: love) [17]. Their reason for this problem is that abstract concepts are not only anchored in the sensorial domain. In fact, they argue, linguistic and social aspects play a key role in the acquisition of abstract concepts. Subsequently, they propose to extend the embodied approach to include linguistic and social experiences. Of course, part of this realisation is due to the rapid success of the distributional semantic hypothesis [43]. This view posits that meaning is computed statistically. Meaning is given by the cooccurrence of words in large masses (corpora) and it derives from the relationship between associated words rather than between words and their referents. It goes without saying that the recent development of Large Language Models (LLM) [44,45,46] is deeply impacting the perspectives on abstract concept learning. 1.2. Concepts in Artificial Intelligence In Artificial Intelligence and Machine Learning, several concept theories have been presented over the years. The early cyberneticians such as Wiener referred to the notion of universals to capture “what makes a square a square” [47]. The connectionist avenue gravitated more towards understanding concepts or universals as patterns. Pattern recognition gained enormous traction, exhibiting very powerful results, but was limited to the distribution of training data and facing important challenges in terms of explainability and compositionality. However, most approaches to concept learning have been proposed from the point of view of symbolic AI, typically focussing on semantics [48]. A good review of these approaches is discussed in the work of Goguen [49], which includes the Geometrical Conceptual Spaces of G¨ ardenfors, the Mental Spaces of Fauconnier, 6
the Information Flow of Barwise and Seligman, the Formal Concept Analysis (FCA) of Wille, the Lattice of Theories of Sowa, and the Conceptual Integration of Fauconnier and Turner. Next we focus on FCA, Mental Spaces and Conceptual Spaces, as they are the most inline with the subject at stake. Formal Concept Analysis views concepts as knowledge units and acts of cognition that are potentially independent of language [50]. In formal terms, concepts are binary relations between objects and their properties. The concept apple, for example, is defined by (i) the symbolic properties of apples (e.g.: is a fruit, grows in trees, has light colour inside, has thin skin, etc.) and (ii) all the objects that verify those properties (all apples). From this definition, and borrowing from Lattice Theory and Ordered Sets Theory [51,52], Wille is able to generate formal structures that express hierarchical relationships among concepts. Additionally, FCA provides a powerful engine capable of reasoning about objects and their properties upon the definition of a formal context (a set of objects and their properties). The simplicity and mathematical elegance of his theory have given it a relatively large success, and research on FCA is still strong in the current literature. Important extensions to Rough Sets [53], Fuzzy Logic [54] or Granular Computing [55] have been proposed since. In addition, a wide variety of applications have been developed throughout the years [56,57,58,59]. Wille’s motivations went far beyond the realm of Mathematics. The main aspiration of his body of work was to support rational communication in humans. As such, a great effort was dedicated to contextualising FCA in broader philosophical discussions and advances in psychology at that time. In particular, Wille attributed to Piaget’s foundational work on cognition and Seiler’s ideas contained in “Conceiving and Understanding” [60]. Furthermore, in [50], he gives a detailed account of FCA compatibility with Piaget’s school of thought. In particular, Wille supports Piaget’s view that concepts are naive and subjective theories that contain implicit and explicit assumptions about the world. In addition, the author endorses the notion that concepts are of an abstract and idealising nature, or that “concepts consider things and events out of a specific perspective and reconstruct only those aspects and relations which follow from the specific view”. Despite efforts, the impact of FCA on cognitive science has been moderate. Perhaps, one of the reasons is that, while it may well be compatible with some important cognitive theories, it certainly does not provide any explanation for many of the pressing questions and complexities around concepts discussed in the previous section. Fauconnier’s model of Mental Spaces [23] and of cognition in general, is extremely interesting. His motivation was to overcome the limitation of 7
Propositional Logics in the study of language and cognition: “Regardless of whether propositions play a role in semantic theory or natural language logic, sentences are not carriers of propositions” [23]. He argued that this model was flawed and that a more flexible approach had to be pushed forward, thus his proposal of the mental spaces model. In his view, it seems that the idea of concept is not of particular interest, at least not in a monolithic or static way. He argues that, contrary to what had been the norm in linguistic research so far, meaning is not contained in semantic objects (e.g. words, sentences). Instead, meaning is created in the mind ad hoc from a complex process in which many factors intervene, such as context, previous experience, etc. In this process, language is a trigger or a vehicle that can guide and influence the construction of meaning [61] but, in no case, meaning is contained in language itself. This perspective implies that concepts are highly dynamic and contextual. Therefore, the pursuit of defining concepts in a sort of essence-driven manner (what makes a square a square), is seen as misguided in the light of his work. Mental Spaces as a model, however, has not had perhaps all the impact it deserved. On the one hand, Fauconnier does not provide a mathematical formalisation like FCA, which makes it harder to adopt by the AI community. And, on the other hand, the model was rather complex in terms of the number of elements and special theoretical instruments that conformed it. This aspect leads to a certain level of arbitrariness when it is not supported by empirical evidence from psychology or neuroscience. In fact, most of his approach stemmed from the field of linguistics. Subsequently, many of the ideas around Mental Spaces had a limited legacy in Cognitive Science. Conversely, his later work on Conceptual Integration (or Conceptual Blending) with Turner [62] did prompt quite some volumes of literature. The Conceptual Blending Theory posits that the mind has a creative ability to combine disparate elements from different domains, leading to the emergence of unique, integrated mental spaces known as blends. As the authors claim, blends allow for the generation of novel meanings and insights. It must be credited that this line of research is still very much alive today, especially in the fields of Linguistics and Creativity [63,64,65,66]. G¨ ardenfors’ Conceptual Spaces [67] do not really fall under the category of symbolic methods. In his work, he points out the weaknesses of both symbolic and connectionist models. Symbolic models fall prey to the Frame Problem [68,69,70]. The upshot of this problem is that propositional representations are not well suited for representing causal connections or dynamic interactions. Additionally, symbolic models face the Symbol Grounding Problem (SGP) formulated by Harnad [71]. The SGP mandates that symbols should be given content (referents) in a sort of self-emergent process, and not by some deus ex machina procedure (e.g., 8
experts assign meaning or values to the symbols). G¨ ardenfors refers to connectionist models as a particular case of associationism [72] (greatly promoted by Locke and Hume), based on neural network architectures. The main issues he points out are (i) the need for large training sets, (ii) lack of explainability, (iii) poor cross-domain performance (a network cannot easily generalise what it has learnt from one domain to another, for example from audio input to image input) and (iv) similarities among patterns learnt by the networks cannot be established intrinsically or in a natural way. Conceptual Spaces, thus, are formulated as a third way. This third way is a geometrical form of representation (neither symbolic nor connectionist), where concepts are regions in a multidimensional space. Each dimension of this space corresponds to some domain qualities of the world. For example, when focussing on an object colour, one may consider three domain qualities: hue,saturation and brightness. Therefore, concepts in this space are defined by volumetric regions within a three-dimensional space defined by those domain qualities. With this setup, it is clear that the method is well positioned to ground concepts in perceptual data. Also, the accounts for similarities among concepts are natural to the system. Additionally, because concepts are regions in space, it is possible to establish spatial relations among concepts: hierarchical (a region inside another region), intersections, or more generally, any relation given by a definition of a region connection calculus [73]. In order to discretise the domain quality space into concrete concept-regions, G¨ ardenfors borrows the idea of prototypes from Prototype Theory in the line of Quine [74]. According to at least some aspects of this theory, concepts (usually) show graded membership. This means that some referents are more representative of the concept than others, hence the notion of prototype or prototypical referent (or prototypical point in the domain quality space). From these points, Concept Spaces suggest generalised Voronoi tessellation to discretise the space into finite concept-regions. From this strategy, it seems apparent that the task of deciding which are the prototypical qualities of a concept is not an easy one. Unless an explicit self-emergent mechanism is provided to account for the selection of prototypical referents, it seems that the method would still face important challenges with regards to the SGP. The angle of Concept Spaces is perhaps not one that has attracted too much attention from Cognitive Science, although the author does throw in some hints at biological and psychological research that share a similar geometric approach [75,76]. Indeed, the method, as he himself explains, is focused exclusively on constructive aspects of cognition rather than explanatory ones. However, as a 9
hidden patterns from the data, enabling actionable insights for urban planning agencies. Architectural Design: • Proposes a parametric data augmentation scheme for three-dimensional geometric samples. And also, a specific representation scheme for wireframe building structures with application in neural models. • Demonstrates a form-finding methodology to create novel building structures by combining features of other building typologies. 1.5.2 Contributions to Bridging AI and Architectural Design This research makes significant strides in the bridge between AI and architectural design, focusing on conceptual models. It builds the case for the importance of concepts in both design discourse and creativity. And it discusses the state of the art around concepts from the angle of Cognitive Science and also form the standpoint of formal AI implementations and generative ML methods. In generative AI, it critiques current methods under the lens of creativity and their ability to perform at a conceptual level, exploring their suitability for design disciplines. Most profoundly, the thesis proposes a new conceptual model for spatial representation that integrates perception and semantic structures, addresses the Symbol Grounding Problem, and paves the way for the development of an AI-based language for conceptual modelling. These contributions represent a new approach to using artificial intelligence in conceptual, creative processes central to architectural design. The main contributions in this area are summarised below, distinguishing those coming from the method proposed at the end of Chapter 6(BIGA) from the more general ones: General contributions: • Provides a critical discussion on the current generative AI models in relation to the subject of creativity. • Offers a survey of generative AI methods, especially in connection to design disciplines. • Delivers an analysis on generative models through the lens of their suitability to perform under a conceptual paradigm. Specific contributions from the BIGA model: 16
• Integrates perception and conceptual structures. The method provides a seamless integration between sensory perception and semantic conceptual structures without the need to combine separate AI paradigms, distinguishing it from neural-symbolic models. • Eliminates data labelling and extensive training. It removes the requirement for extensive data labelling and training, making it more accessible for practical applications and reducing reliance on resource-heavy AI training methods. • Enables explainable AI. By offering full explainability of how concepts are formed from raw data through atomic comparisons, the method enhances transparency in AI decision-making, contributing to the broader effort to make AI systems more interpretable. • Addresses the Symbol Grounding Problem. The model’s bottom-up approach to construct semantic structures from basic tokens contributes toward solving the long-standing Symbol Grounding Problem in AI. • Lays the foundation for a new AI language. The method hints at the potential for building a basic AI language from primitive tokens, offering a novel direction for future exploration in AI-driven semantic representation. 1.5.3 Publications associated with the thesis In a more broader sense, the thesis also examines the following aspects —cutting across both the connectionist and the symbolic paradigms in AI, and addressing the background contexts of creativity, space and design: • Critical Analysis of Generative AI: The thesis provides a detailed examination of the role of generative AI in creativity and design, highlighting its potential and limitations. • Evaluation of Symbolic Methods: FCA’s application in urban analysis and its conceptual capabilities are critically assessed, revealing both strengths and areas for improvement. • Integration of NSMs: The exploration of NSMs offers insights into how combining symbolic and connectionist approaches can address some of the current challenges in AI. • Novel Representation Scheme: The proposed method for conceptual representation of space introduces a new approach to integrating sensory data with semantic structures, paving the way for future research and 17
application in architectural design. In general, this thesis underscores the importance of integrating diverse AI approaches to address the complexities of architectural design. It highlights the need for continued exploration of hybrid models and novel methods to enhance the interplay between AI, creativity, and human cognition in shaping the architectural design discourse. Future research should focus on expanding these methods, exploring their practical applications, and further bridging the gap between AI and the qualitative aspects of design. Finally, these contributions have been made accessible to the research community through the following publications (in chronological order): •de Miguel-Rodr´ıguez J,Gal´an-P´aez J,Aranda-Corral GA,Borrego-D´ıaz J. Urban knowledge extraction, representation and reasoning as a bridge from Data City towards Smart City. In: Proceedings of the 2016 International IEEE Conferences on Ubiquitous Intelligence & Computing, Advanced and Trusted Computing, Scalable Computing and Communications, Cloud and Big Data Computing, Internet of People, and Smart World Congress (UIC/ATC/ScalCom/CBDCom/IoP/SmartWorld). Toulouse, France. 2016:968-974. Prior to the design of sociotechnical artifacts in cities, it seems important to extract the qualitative, quantitative opinions, sentiment, and feedbacks present in these data. This paper presents three solutions for mining these contents through Knowledge Extraction methods, as a previous step to the prospection of new smart services. •de Miguel-Rodr´ıguez J,Villafa˜ne M-E,Piˇskorec L,Sancho-Caparrini F. Generation of geometric interpolations of building types with deep variational autoencoders. Design Science. 2020;6:e34. doi:10.1017/dsj.2020.31 This work presents a methodology for the generation of novel 3D objects that resemble wireframes of building types. These results from the reconstruction of interpolated locations within the learnt distribution of variational autoencoders (VAEs), a deep generative machine learning model based on neural networks. •de-Miguel-Rodr´ıguez J,Morales-Esteban A,Requena-Garc´ıa-Cruz M-V, Zapico-Blanco B,Segovia-Verjel M-L,Romero-S´anchez E, Carvalho-Estˆev˜ao JM. Fast Seismic Assessment of Built Urban Areas with the Accuracy of Mechanical Methods Using a Feedforward Neural Network. Sustainability. 2022; 14(9):5274. https://doi.org/10.3390/su14095274 18
This paper presents a novel en-masse method to assess the seismic vulnerability of urban areas swiftly and with the accuracy of mechanical methods. The core of this methodology is the calculation of the capacity curves of low-rise reinforced concrete buildings using neural networks, where no building modelling is required. •de Miguel-Rodr´ıguez J,Sancho-Caparrini F. A Recursive Bateson-Inspired Model for the Generation of Semantic Formal Concepts from Spatial Sensory Data. arXiv. Published online 2023. doi:10.48550/arXiv.2307.08087. Available from: https://arxiv.org/abs/2307.08087. This paper presents a new symbolic-only method for the generation of hierarchical concept structures from complex spatial sensory data. The approach is based on Bateson’s notion of difference as the key to the genesis of an idea or a concept. •de Miguel-Rodr´ıguez J,Sancho-Caparrini F. A Bateson-Inspired Model for the Generation of Semantic Concepts from Sensory Data. AI Communications. 2024; ARTICLE IN PRESS (currently undergoing its second review after having submitted a major revision). This paper revisits symbolic-only approaches by introducing a new algorithm to create hierarchical concept structures from spatial sensory data. The method is based on Bateson’s idea of difference as the fundamental element of concept formation. •de Miguel-Rodr´ıguez J,Morales-Esteban A,Requena-Garc´ıa MV, Romero-S´anchez E. Automated Extraction of Building Typologies from a Digital Land Cadaster Using a Variational Autoencoder. World Conference on Earthquake Engineering, Milan, Italy. 2024; ARTICLE IN PRESS (pending only publication). This paper introduces a novel method to automate the identification of building typologies from the digital land cadastre of the city of Seville, eliminating the need for human intervention. The method is based on a variational autoencoder that learns to extract these typologies without explicit supervision. 19
2. Engineering Applications of Artificial Intelligence in Architecture 2.1. Background Before the wide spread of Deep Learning in Machine Learning, the initial approaches to its application in civil engineering were mostly oriented towards Expert Systems. These systems were regarded as some of the first applications of AI that actually achieved some important practical results [92,93]. In the words of Dana S. Nau, expert systems are “problem-solving computer programs that can reach a level of performance comparable to that of a human expert in some specialised problem domain”. Given this definition, the difference from a standard computer program may not become apparent. However, there is an important difference, in that expert systems employ a knowledge base that is articulated as a separate and independent entity. This knowledge base, in turn, is accessed and operated by a control system that is also a separate and independent module. Formerly, the process of knowledge acquisition required to form the knowledge base involved consulting experts on domain matters. This process had a number of problems, such as communication with experts, quality of expertise, conflicting expertise among different experts, or misunderstanding instructions, among others [94]. For many, the knowledge acquisition process had become the single most important challenge facing expert systems, in what became known as the knowledge acquisition bottleneck. In this context, ML methods appeared as a powerful alternative that had the potential to eliminate the aforementioned problems. These techniques posed a huge advantage in that they made substantially fewer assumptions about the data. The view was that ML can be used as a means to automate the generation of knowledge for its operation within expert systems [1,95]. This automated knowledge elicitation process was initially also referred to often as Machine Induction [96], because knowledge was built by capturing high-level relationships from a set of observable samples. This knowledge would then be applied to unseen situations. However, soon more automated techniques began to be used and the term fell out of vogue in favour of the current and more general one, Machine Learning. Some early classifications of the ML approaches available for automating the 20
knowledge acquisition process include: • Learning from examples, also known as Inductive Learning or Concept Learning. • Learning by observation, also known as Conceptual Clustering or Concept Formation. • Theory-driven learning. • Learning by discovery [95]. Although classifications vary according to the author and may also seem quite obsolete today, they offer historical insight into the early stages of AI applications in the field of engineering. According to the authors of this classification, inductive learning involves the acquisition of ’concepts’ at the high level in the form of decision rules and trees. Learning by observation, in turn, refers to the ability to produce ‘theories’ that account for a set of given facts or observations. The difference with inductive learning is that here the aim is to arrive at a reasoning engine that can be applied to the observable world, not beyond. In inductive learning, the aim is to extrapolate the findings to unobserved situations. Of course, this distinction is rather artificial and in practice, concept formation techniques such as FCA can also be applied to new unobserved situations in an inductive fashion. In contrast, decision rules and trees can also be used to engage in some level of reasoning or ‘theorising’ about the observable world. Moving forward, theory-driven learning, as its name suggests, involves the learning of conceptual knowledge directly from a formal domain theory. Because it is not common for such formal theories to be available in many of the areas where ML is applied in engineering and elsewhere, this group may not have accounted for a broad research adoption. Finally, learning-by-discovery is described as an approach in which knowledge is acquired by understanding or mining certain patterns or regularities in the data or among past problem solving experiences. Different authors have added more categories, such as causal learning,learning by analogy, or learning by experimentation. Other categories also include similarity-based learning and explanation-based learning [97]. In general, the distinctions between these groups are often quite loose and present important overlaps. Even more contemporary classifications of ML methods, like supervised versus unsupervised methods or symbolic versus probabilistic, can be found to be excessively rigid in certain cases (e.g., low-supervision models, connectionist reasoning models, etc.). Therefore, it might be more coherent to focus on the individual methods themselves rather than 21
attempting to follow such classifications too closely. In the next section, some of these methods will be explored further. During the emergence of ML for expert systems, some efforts have been made to provide unified frameworks to guide engineers in their implementation. In [1], Reich points out two important challenges present in the applications of ML to the field of civil engineering. Firstly, practical problems are often too complex to be handled by a single method. This circumstance led to the development of multi-strategy learning [98,99]. For example, empirical learning generally relies on numerous input examples while requiring minimal background knowledge. In contrast, explanation-based learning requires only a single example, but requires comprehensive background knowledge. Learning by analogy hinges on having background knowledge analogous to the input. Real-world applications rarely meet the criteria of single-strategy learning approaches. Hence, there has been a growing interest in constructing systems that fuse various learning strategies. Within these strategies, macro and micro approaches are distinguished. The macro approach involves the interaction of different learning programs under the non-trivial responsibility of the user to manage those interactions. The micro is targeted towards specific fine-grained tasks. Secondly, the other difficulty that had become apparent was that ML implementation in civil engineering and related practices was not as straightforward as taking an ML program and applying it to the data. Indeed, the process involves aligning the scope of applicability of ML tools with the specific nature of the task at hand. This alignment requires a deep grasp of ML techniques and the ability to be creative around them. To assist in this challenging endeavour, the author proposes a seven-step schema starting with (1) problem analysis, (2) data collection and knowledge, (3) problem representation, (4) method selection, (5) parameter settings, (6) evaluation and interpretation, and (7) solution deployment (see Fig. 2.1). These steps are merely informative of course, and such guidelines tend to proliferate over time with different variations and scopes depending on the inclination to a certain line of ML approaches, or on the application domain, etc. In a much more recent publication [100], aimed at the healthcare industry, for example, the authors propose a simpler four-step schema: (1) inception, (2) preparation, (3) development, and (4) integration. Clearly, research-oriented approaches do not account for aspects such as integration with a larger software architecture or interactions with users after deployment as much as industry-driven approaches, which is expected. Another recent paper [101], more oriented towards supervised learning, proposes a set of five steps: (1) justification 22
Figure 2.1: Diagram of Machine learning application steps. Image extracted from [1]. for using ML approaches, (2) data collection, (3) data preprocessing, (4) ML modelling and training, and (5) system testing and performance validation. It can be noted that, in this schema, data-related aspects play a major role, as is common in current ML models that operate on massive volumes of data. As a working guideline, combining the spirit of these and other recommendations from the literature, a generic framework may be proposed as follows: 1. Preliminary analysis 2. Data engineering 3. Method selection (model selection) 4. Training, parameter and hyperparameters 5. Testing, interpretation and evaluation 6. Deployment and interaction 23
As an example, the case study presented in this chapter follows this framework to a great extent. It starts with a preliminary analysis, identifying the need for rapid urban seismic assessment and the cost/accuracy trade-off associated with traditional macroseismic methods. The next step, data engineering, involves preparing a dataset based on structural building properties and seismic responses. Additionally, a data exploration exercise is also carried out in this step. A UMAP algorithm (Fig. 2.10) is applied to the data allowing to better understand its nature and providing insights for model selection. The model selected is a feed-forward neural network due to its capacity to generalise complex patterns and perform high-dimensional regressions. During the training phase, the model is tuned using data obtained from the calculation of 10k structures in a structural analysis software (SAP2000). During this process, different model parameters and hyperparameters are tested for improved accuracy. The testing and evaluation phase shows the effectiveness of the model, with its predictions aligning closely with mechanical methods while offering significantly faster processing times. Finally, a potential for deployment is suggested, positioning the model as a rapid seismic assessment tool for urban planning and disaster mitigation, emphasising its practical applications in emergency scenarios. 2.2. Early works 2.2.1 Inductive Learning Suppose Nreal-world observed examples are given, {e1, . . . , eN}which are defined from a set of attributes (properties), ei= (pi1, . . . , pim), and for each of them there is an observed classification, ci. The task of Inductive Learning is to induce from the above data a mechanism that allows inferring the classifications of each of the examples from the properties alone. If this is possible, this mechanism could be used to deduce the classification of new examples having only observed their properties. An important development in Inductive Learning was the ID3 algorithm (Iterative Dichotomizer 3) developed by Quinlan in 1986 [96]. ID3 is a type of decision tree algorithm based on Hunt’s Concept Learning System [102] but with an Information Theory approach. A decision tree consists of a set of decision nodes (interior) and answer nodes (leaves): • A decision node is associated with one of the attributes and has 2 or more branches coming out of it, each representing the possible values that the associated attribute can take. A decision node can be interpreted as a question 24
asked to the analysed example about one of its attributes, and depending on the answer it provides, the flow will take one of the outgoing branches. • An answer node is associated with the classification to be provided and returns the tree’s decision with respect to the input example. It is obvious that obtaining a decision tree that can predict the examples with 100% reliability will not always be possible, but the better the battery of examples available (e.g., no contradictions between classifications), the better the performance of the tree that can be built from them. Obviously, the construction of the decision tree is not unique. By applying different strategies when deciding in which order to ask the questions about the attributes, very different trees may be obtained. Also, their construction varies in complexity. Among all the possible trees, the objective is finding those that fulfil the best characteristics as prediction machines. Consequently, the challenge is to give an automatic mechanism for the construction of an optimal (or near-optimal) tree from the examples. The ID3 algorithm builds a decision tree from top to bottom, in a straightforward manner, without backtracking, and based only on the initial examples provided. It uses the concept of Information Gain (based on Shannon Entropy [103], which measures the degree of uncertainty of a sample of examples) to select the most useful attribute at each step, and follows a voracious method to decide which question offers the highest gain at each step, i.e., the one that best separates the current examples from the final classification: the lower the uncertainty associated with a given attribute, the closer the node associated with that attribute is to the root in the tree. ID3 is the predecessor of C4.5 (also known as Statistical Classifier) [104], one of the most popular decision tree algorithms used in classification today. In this extension, the algorithm can handle both categorical and continuous attributes and includes mechanisms to handle missing values and pruning trees to avoid overfitting, both of which are important limitations of ID3. Throughout the 1990s, these two were the most widely used inductive learning techniques in civil engineering [105]. Of course, other decision tree methods have also been developed, such as CART (Classification and Regression Trees) [106], CHAID (Chi-squared Automatic Interaction Detector) [107], and MARS (Multivariate Adaptive Regression Splines) [108], among other algorithms and variations. Random Forests (ensemble learning) were about to emerge as a top player in the year 2001 [109], and neural networks had begun to gain considerable traction by then as well. 25
setup of weights and biases (often initialised using random values), the output values can be calculated deterministically using the formula above [139]. Although only one output value is shown in the example, it is easy to project multiple output values in the same scheme. For example, for classification tasks, a network could be designed to have N possible output values through Nneurons in the output layer. The general idea is to achieve a distribution of weights and biases (trainable parameters) that when the inputs corresponding to an example of i-th category flow through the network, the two output values form a vector close to, say, ei(the unit vector in the i-th direction). When all functions in neurons are differentiable, this can be achieved by adjusting the training parameters through an iterative process known as backpropagation. When multiple examples of the categories are processed by the network with the ensuing adjustment of weights and biases, the network ’learns’ a distribution of trainable parameters that minimises the classification error. This process can be seen as a search for local minima in the error function or, as commonly referred to, a gradient descent along the loss function. Of course, there are many different algorithms to handle gradient descent, as it is in fact one of the most important factors in the training process, especially in models that feature a high number of layers and parameters. When dealing with architectures with many layers, they are known as deep learning models. The term deep learning is used quite loosely in the literature, without a strong consensus on what or how many layers exactly constitute deep learning. An in-depth discussion of this matter can be found in [140]. Going forward, this thesis will not attempt to make any hard distinctions between deep and nondeep learning. Some of the existing methods to compute the gradient descent will be discussed in the case study presented after this section. Figure 2.4: Perceptron scheme 32
Figure 2.5: general neural network diagram Within this general scheme, there are many different neural models and architectures possible. The most simple, as just described, is a Feed-Forward model. Other important models include Convolutional Neural Networks, Recurrent Neural Networks,Autoencoders,Variational Autoencoders,Transformers, and Generative Adversarial Networks. There are many other types and variations; however, these are arguably the main ones. In convolutional models [141], neurons are first interconnected within a local range, before connecting to the next layer. This is done to boost learning of features that depend on spatial proximity, as is, for example, the case in images. Recurrent models are characterised by feeding the output of a neuron back into the same neuron as an additional input. In this way, recurrent models are able to model state-dependant problems in a more efficient way, and are thus typically chosen to learn from dynamic systems (time-series). Autoencoders [142] are a special type of neural network in which, rather than having a set of categories-labelled examples, the model attempts to reconstruct the very same input being fed. Therefore, in autoencoders, the error function is simply defined by some measure of the difference between the input and output values, and is commonly referred to as reconstruction error. This trait makes autoencoders a type of unsupervised or self-supervised learning model. They are usually aimed at compression or denoising, and normally feature a smaller-sized layer in the middle that can be considered to contain a 33
condensed encoding of the inputs. Variational autoencoders are a particular case of autoencoders that together with generative adversarial networks constitute a cornerstone of generative AI. These models will be discussed in the next chapter. Finally, it is important to note that these are relatively fluid schemas, and many models are a mix of the types presented here. For example, it is common to introduce convolutional architectures in autoencoders or variational autoencoders as will be shown later. The following is a brief compilation of works. Tai Ng et al. [143,144] incorporated an ANN with a Bayesian method for the health assessment of a steel frame structure. Radhika et al. [145] proposed a wavelet-based change detection method using ANN and Support Vector Machine (SVM) for damage classification [145]. Then Alavi et al. [146] proposed a damage assessment approach based on probabilistic neural networks and Bayesian decision theory for SHM. Lee et al. [147] designed a theoretical model using ANN to predict the shear strength of the slender fibre reinforced polymer in reinforced concrete beams. Figueiredo et al. [148] authored an environmental variability study and damage detection using ANN, Mahalanobis distance, and singular value decomposition. Yan et al. [149] implemented a neural network and SVM model to assess damage in beams on ocean platforms. Parsad et al. [150] applied a neural network model to the prediction of compressive strength in self-compacting and high-performance concrete. And Chatterjee et al. [151] developed a multi-objective genetic algorithm for the calibration of a neural network model in the classification of reinforced concrete buildings. Dai et al. [152] designed a wavelet SVM-based neural network metamodel for reliability analysis in various structures. Finally, Butcher et al. [153] used ANN and extreme learning machine methods for SHM in concrete structures with mesh reinforcement. Sarkar et al. [154] used CNNs to characterise crack damage in composite materials. Abdeljaber et al. [155,156] employed one-dimensional CNNs for vibration-based structural damage detection, learning directly from acceleration data. Abdeljaber et al. [157] introduced a nonparametric damage identification method using CNN, effective with only two measurement sessions. Cha et al. [158] implemented a deep learning network to detect concrete cracks in tunnels without computing defect features, robust compared to traditional methods. Lee et al. [159] explored deep learning models and CNNs for structural analysis of a ten-bar planar truss, showing efficiency compared to conventional neural networks. Finally, in [160], CNN was used to identify structural damage, discovering unknown relationships between measurements and damage patterns. 34
4. Other methods: Zhou et al. [161] introduced a damage detection technique using cosine similarity measure. Zhang et al. [162] contributed a structural identification method employing pattern recognition and support vector regression (SVR). Laory et al. [163] developed a methodology to predict natural frequency responses of a suspension bridge using multiple linear regression, ANN, SVR, regression tree and random forest. Nagarajaiah et al. [164] studied damage detection based on sparse and low-rank data structures for structural dynamics, and Yang et al. [165,166] analysed recovery of structural vibration responses and damage localisation with low-rank matrix decomposition. Yepes et al. [167] proposed a multi-objective optimisation of high-strength reinforced concrete beams using Minkowsky metrics. Garcia-Segura et al. [168] implemented a reliability-based optimisation of post-tensioned concrete box-girder bridges under corrosion attack. Saridemir [169] employed genetic algorithms to determine the tensile strength split from the compressive strength of concrete. Yeh et al. [170] implemented a genetic operation tree to predict the compressive strength of high-performance concrete, while Cheng et al. [171] developed a genetic weighted pyramid operation tree for the prediction of compressive strength in high-performance concrete. Kiremidjian et al. [172,173] utilised autoregressive models for structural health monitoring (SHM) of bridge structures, while Gul et al. [174] and Yao et al. [175] used autoregressive models with a Mahalanobis distance-based outlier detection algorithm for damage detection in civil structures. As mentioned above, seismic engineering is a field with remarkable complexities involved. The inherent unpredictability of seismic events, coupled with the intricate nature of structural responses, poses significant challenges. Because of this, there is an important window of opportunity in using ML methods to address these challenges. Using the power of ML, engineers can analyse vast amounts of data, uncover patterns, and make predictions that were previously unattainable through traditional methods. Many authors and the engineering community have recognised this potential and have started to work intensely on ML applications for seismic engineering. Following this trend, this chapter introduces a neural network application designed to predict the seismic vulnerability of a large number of buildings based on simple geometric parameters. This approach aims to streamline the assessment process, providing rapid and accurate predictions of the stress-deformation curves of building structures under seismic actions. Developing fast and reliable methods to obtain these curves (capacity curves in seismic engineering lingo) can 35
significantly improve preparedness and mitigation strategies to alleviate the impact of earthquakes in urban settings. Before delving into the specifics of this application, a series of works related to ML applications in seismic engineering is presented. These studies underscore the progress and potential of integrating ML into the field, highlighting various innovative approaches and their impacts on the assessment and management of seismic risks. By situating this neural network application within its broader context, the aim is to illustrate its relevance and contribution to ongoing advances in seismic engineering. 2.3.1 Related works Neural networks and especially Deep Learning have been intensely applied in the field of seismic engineering. This area requires some of the most complex structural analysis models because of the nonlinearity induced by seismic actions. In particular, these nonlinearities can be of geometric type, e.g.: a structure is deformed to a point where one can no longer assure the verticality of the columns, thus requiring that second-order effects be taken into account. And also, nonlinearities may present themselves in material behaviour, e.g.: reinforcement steel working past the elastic or linear regime into the plastic domain. Because of the important challenge that these computations pose to ML, many researchers have resorted to Deep Learning in the quest to predict the nonlinear behaviour of structures under seismic action. However, neural-based methods are not always the best choice for the given task. Seismic engineering is a broad field, and some problems may require different approaches. For example, Gong et al. [176] implemented an earthquake-induced damage identification model in buildings using SVM, random forest, and KNN. In addition, Elwood et al. [177] proposed an approach based on fuzzy pattern recognition for the detection of seismic damage in concrete structures. However, connectionist models have been extremely powerful and have seen a very rapid widespread adoption. In this line, Soleimani et al. [178] incorporate ANN into state-of-the-art probabilistic seismic demand prediction methods for bridge components. Perez Ramirez et al. [179] a deep recurrent neural network model based on a nonlinear autoregressive exogenous model (NARX) is presented for the accurate prediction of the seismic response of large structures. Ruggieri et al. [180] developed an ML framework for the vulnerability analysis of existing buildings based mainly on CNN models over labelled photographs of these structures. Won and Shin [181] developed another neural network-based model that was able to rapidly predict seismic responses with soil–structure interaction effects and determine the corresponding seismic performance levels. Kwag et al. [182] 36
proposed a model to predict the seismic performance of the slope with relatively high accuracy and efficiency using ML methods. They compare ANN and SVM methods with a result much more favourable to the former. Finally, He et al. [183] present a DL approach based on a Gated Recurrent Unit (GRU) network to assess the seismic fragility of structures. The GRU network is used to create a surrogate model that captures the nonlinear relationship between seismic responses and mainshock-aftershock earthquakes. Moving on to works that are more specifically related to the objectives of the case study presented in the next section, a number of articles have been identified. Although most of them are based on neural architectures, other relevant methods have been included which employ other ML methods. In [184] ANN were used to predict the structural response of different floors of a building to earthquakes of various intensities. However, their study is limited to a modal analysis, whereas in the present work a full nonlinear static analysis is developed. Similarly, a very in-depth application of neural networks was carried out in [185] to predict seismicinduced stress in specific elements of a two-span and two-story structure. In contrast with the present research, their model is specific to a case study structure and trained networks cannot be used to predict stresses in similar but different structures. This is also the case for [179], where a recurrent neural network model with Bayesian training and mutual information is used for the prediction of response for large buildings. Recurrent nets are a common choice for time series and, more broadly, data that feature temporal correlations. Although not specific to response prediction, other works featuring recurrent models in the area of seismic analysis can also be found in the literature [186,187]. In light of the results obtained in the case study developed for this chapter, the feed-forward model architecture chosen here was sufficient to address the problem at hand. However, a recurrent approach would perhaps be better suited for the challenges described in the Future work Section (for instance, extending the proposed methodology to high-rise buildings). Finally, in [188] attention was paid exclusively to estimating seismic-induced demands on column splices, and in [189] a more generic analysis was performed, building an ANN that predicts damage in RC shear walls based on the inter-story drift. However, in this case, the drift per level must be calculated before making use of the neural network. ANN were also used as a classifier to establish in which damage-level category a certain structure would fall into, for a given seismic demand [190,191]. Their approach was therefore to use neural networks for pattern recognition based on structural parameters. Here, a multidimensional regression approach to the 37
application of neural networks is used to predict capacity curves, which provide a more detailed behaviour of the structural damage. In the case of [192], a very similar approach to the case study in this chapter was adopted to predict fragility curves with neural networks. However, since their study was based on dynamic analysis, capacity curves were not considered, in contrast to the static approach employed in this paper. Furthermore, fragility curves were not computed in a single run of the network as proposed by the present method. Instead, each run estimates an individual pair of coordinates, and then curves are re-built based on these estimations. ANN and other techniques were also used to predict the performance point of school buildings under seismic action [193]. Similarly, in [194] the same objective was aimed instead using a genetic algorithm, achieving similar levels of error. The advantage of this last approach is that their model provides a transparent mathematical formula for the prediction of the performance point. However, their study focusses on predicting the performance point directly without considering the capacity curves of the buildings. Other works used neural networks to make use of simplified approaches to determine the seismic performance of buildings. In [195], an experimental database was used to train a neural network with very few input parameters, seven in total, to predict stress and deformation values at specific locations within masonry-infilled RC frames under seismic action. The study aims at simplifying the complex modelling of mixed-element structures but is still limited in the scope of application due to a much-reduced number of input parameters. Finally, in [196], neural networks were used to predict a bilinear simplification of capacity curves with accurate results. In contrast, in the methodology that will be presented here, the original capacity curves are predicted without simplification, broadening their scope of application. 2.4. A method for en-masse seismic engineering analysis with neural networks The objective here is to develop an accurate technique for performing a seismic assessment of low-rise buildings by predicting their stress deformation curves with ML, employing only a simple set of geometrical parameters of the buildings, rather than engaging in a tedious modelling process. Today, it is common to use macro-seismic approaches to conduct seismic hazard assessment [197,198]. In these methods, the seismic action is first determined and then combined with the seismic vulnerability curves of the stock 38
building to calculate the seismic hazard. For example, in Europe, it is common to use the building classes of the RISK-UE project. In contrast, when studying a building in detail, a specific mechanical model is defined and calculated for it. This is much more accurate but requires many hours of work: the blueprints of the building must be obtained, the parameters of the materials must be determined, a model must be built by a specialist and, finally, results are obtained. This is obviously too time-consuming when calculating a large number of buildings. By shortcutting the modelling process with ML, this approach begets the advantages of both worlds (mechanical and macro-seismic) while having none of their drawbacks. In short, the main innovation and impact is that it sets a methodology for a fast en-masse analysis that preserves the accuracy of mechanical methods with negligible loss. In light of these advantages, this method is set to replace macro-seismic models for seismic hazard assessment of urban areas and real-time seismic evaluation tools. Previous research exists that uses ML methods to predict stress/deformation curves (of which capacity curves can be considered a specific type or subset) of various nature and different contexts, ranging from entire structures to structural elements and material stress tests. Capacity curves are a specific type of stress / deformation curve widely used to perform seismic vulnerability and damage assessments of structures. When calculating the capacity curves of a building, it has been shown that the nonlinear static pushover method yields accurate results, especially when dealing with low-rise buildings that respond mainly in the first vibration mode [199]. The methodology developed here allows for the prediction of these curves by: (i) using simple input parameters derived only from the geometric and material properties of the buildings, (ii) performing in a plastic regime, (iii) in high resolution of up to 100 points per curve, (iv) in a single and coherent process (rather than a point-by-point fashion), (v) for entire buildings with a great range of variability in size (limited to low-rise), and (vi) with immediate applicability to real world emergency relief use cases. These features are explained in the following paragraphs. When assessing an individual building, this method is substantially simpler than a full-fledged nonlinear time history analysis [200]. However, pushover analysis is still time-consuming, computationally expensive and requires advanced modelling expertise. In general, 3D-modelling of structures within engineering software packages is not an option when one needs to assess a large number of buildings. For this reason, there is an increasing volume of research advocating the use of ML techniques that allow bypassing these limitations. Specifically, ANN can 39
perform nonlinear modelling without prior knowledge of the relationships between input and output variables, and consequently, it is no surprise that many engineering disciplines are witnessing an intense engagement with neural network models to solve a wide range of challenging problems [201,202,203,204,205]. Seismic analysis requires accounting for structural behaviour in a plastic regime, and, for this reason, the problem at stake involves highly nonlinear calculations. In most works, stress deformation curves have been predicted using ML methods only in rather constrained settings and specific samples of structural elements. The present approach considers not only a full range of real-life building structures but also the modelling of plastic hinges in structural joints to account for the nonlinear performance of the structural materials (in this case reinforced concrete). Furthermore, the model faces a strong regression challenge by predicting the full capacity curve as a set of 100 points, which constitutes a remarkable resolution. In contrast to previous research where each stress/deformation pair is predicted separately one by one, the methodology proposed here tackles the full curve in one go, which means that all the curve points are delivered in a single run of the network. This allows for the prediction of the full curve in a single step, but also for where it stops. The end point of the curve is important because it can be indicative, for example, of the imminent descent of the structure into a mechanism leading to its total collapse. These two aspects are missing from previous stress-deformation prediction research, whereas in the work presented here it is done in a single and coherent process. The choice of ANN responds, on the one hand, to the fact that these models are naturally well suited for the challenges outlined above (strong nonlinearity and a large regression output of up to 100 points). And, on the other hand, it may be noted that Finite Element Methods (FEM), as used in the SAP2000 calculations of this study, aim at solving differential functions that arise from structural analysis. Neural networks, in turn, are naturally equipped to deal with continuous and differentiable data that allow for an error optimisation process through the gradient descent algorithm. In other words, the differentiable character of the data generated by FEM calculations is another aspect that makes the problem under study a good fit for a neural network model. However, with ANN, there is a well-known drawback in terms of poor or no explainability and the need for massive datasets. Explainability is very important in many cases, especially, but not only, when dealing with human-related data (medical or life insurance areas, for example). In the problem dealt with in this paper, a black-box approach might be forgiven in exchange for the efficiency and speed of the model which allows for fast relief of large urban areas in case of an earthquake. With regard to the need for large training sets, this can be very problematic when data is difficult to obtain. 40
However, this last issue might be offset by the fact that the methodology proposed here allows for the creation of complete datasets at will. As a first implementation, a typology of low-rise prismatic reinforced concrete (RC) buildings has been defined and a training set of more than 7,000 structures has been parametrically generated. The capacity curves of these models have been obtained by means of pushover analysis using SAP2000 software. After defining and training an appropriate neural network model, full capacity curves are predicted in a single run of the network, with a resolution of up to 100 points. This benchmark is substantially higher than previous research that employed less efficient and comprehensive approaches. The problem with this common approach is not only that it is slower, more tedious, and complicated, but also that the information of where the curve ends is completely lost. The methodology proposed allows (i) bypassing heavy and highly specialised computer calculations and complex nonlinear modelling as in [206,207,208], (ii) allowing fast and easy evaluations in emergency scenarios as in [209,191,210], and (iii) predicting missing data as in [192,211,212]. This research is part of the PERSISTAH project1, which aims to jointly assess the seismic vulnerability of primary schools in the Algarve-Huelva region and implement appropriate retrofit solutions. This region has been affected by some of the most famous earthquakes in Europe (20,21)[213,214]. Of 442 school buildings, more than 400 structures in the area are relatively simple low-rise RC buildings; therefore, they constitute an adequate object for the application of the method described in this paper. This research is also a first step in the SIMRIS project, where ML techniques are proposed to obtain the capacity curves of a large dataset of real buildings. 2.4.1 Methodology Parametric generation of structures In order to automate the generation of a training set for the neural network, the first step is to generate a large number of virtual structures, all different from each other but falling under a typology that is representative of the schools in the AlgarveHuelva region. To carry out this task, a set of parameters and value ranges have been chosen based on the list of school buildings dealt with in the PERSISTAH project. These parameters and their value ranges are defined in Table 2.1. For every structure, all slabs share the same height, and the same applies to the 1https://keep.eu/projects/21865/Projects-of-earthquaque-res-EN/ 41
t-SNE [224] and the more recent UMAP [225], which is based on manifold learning techniques and ideas from topological data analysis. With the UMAP algorithm, each of these vectors can be mapped in a bidimensional space, so that all the samples can be visualised simultaneously as (x,y)points in a chart. To track the relationships between the distribution of these points and some output metric, each point has been coloured according to the maximum shear of its corresponding capacity curve. With this colour code, a simplified sketch of how the input data distribution aligns with its output can be obtained as shown in Fig. 2.10. In particular, a certain level of pattern formation is observed, which is desirable for any ML task. However, it is also apparent that these patterns compose a complex spatial distribution which is not trivial to predict. Figure 2.10: UMAP projection of the input dataset. Each sample is colour-coded according to the maximum shear value of its corresponding capacity curve. Curve data pre-processing and post-processing Some studies like the ones described in the previous section have not attempted to predict complete curves within a single neural network architecture; instead, they have approached the problem by predicting individual (stress,displacement) 48
points, thus disregarding the prediction of where the curve ends. The methodology presented here explores the possibility of also predicting these end points by introducing the complete curve as the expected output of the network. This involves the ’pre’ and ’post-processing’ of the data points that build up the curves. In order to arrange the data for each curve in a way that can be processed by the neural network, all curves in both the training and validation sets must be normalised (Fig. 2.11) and defined by equal intervals on the displacement axis and the same number of points. Therefore, (i) a resampling of the curves has been conducted, and (ii) shorter curves have been completed with a stretch of constant shear value, as shown in Fig. 2.12. Figure 2.11: Normalised capacity curves of structures #11 (left) and #34 (right) as returned from SAP2000 Figure 2.12: Re-sampling of capacity curves of structures #11 (left) and #34 (right) with a resolution of 100 points and an interval of 0.01. Curve points to the left of the dashed line correspond to the original capacity curve after normalisation. Curve points to the right are included in order to complete short curves with a stretch of constant shear value. Once a curve is predicted, it is necessary to undo this horizontal stretch, so that the end point of the predicted curve can be obtained. Due to the stochastic nature of the neural network used in this study, the prediction of this horizontal part of the curve does not yield a perfectly straight line. Thus, it is necessary to design an 49
algorithm capable of capturing it, despite its irregularities. In this work, a simple algorithm has been designed and implemented for this purpose with satisfactory results, as follows: (1) A tolerance value (adjusted empirically to serve the curves in the present work) is established as 1% of the maximum shear value in the curve. (2) Count the number of times the curve changes from a negative to a positive slope angle or vice versa (this has been called a dribble). (3) Increase the tolerance value for every dribble occurrence beyond the second one by 0.2% with a maximum possible value of 5%. (4) If the next shear value stays within the range of the previous one plus/minus the current tolerance value, then the point is flagged as a potential curve-stop point and the flagged shear value is stored. (5) If the flag is up, then check if the current shear value stays within the range of the flagged shear value plus/minus the current tolerance. (6) If the current shear value stays within the aforementioned range, then count the number of consecutive occurrences of this event. (7) If the end of the curve is reached while the former conditions are met, then return the flagged point if it is at least the second consecutive occurrence calculated in the previous step. (8) If the current shear value leaves the range indicated in Step 5, then reset the dribble counter, the flagged point, and the flagged shear value. (9) If no condition is met at the end of the curve, then return the coordinates of the last point. Artificial Neural Network Loss function and error measurements The ultimate goal of the neural network is to predict the capacity curves of those structures stored in the validation set, given the set of simple input parameters associated with them, as described earlier. To measure how well this prediction occurs, the network requires a loss function which will feed error values into the optimisation process. These values are used to adjust the weights and biases of the network, effectively minimising the error during training. Therefore, the loss function establishes the metric against which the neural network learns. In [226] it was indicated that the Root Mean Squared Error (RMSE) and the Mean Absolute Error (MAE) are commonly used loss functions for the regression of continuous variables, i.e. the problem under study in the present paper. These loss functions are expressed as follows: RMSE =1/nsn ∑ i=1 (yi−y′ i)2 50
MAE =1/n n ∑ i=1 |yi−y′ i| Where nis the number of neurons in the output layer, yiare each of the expected output values, and y′ iare the corresponding predicted values. This value will be calculated for each training sample and backpropagated into the network to serve as a criterion for weight optimisation. After intense training, the MAE value will be as low as possible (and thus the prediction error will be minimum). Both errors are similar, but RMSE penalises large errors more severely, whereas MAE provides a linear penalisation. In addition, RMSE is less intuitive to interpret and is sensitive to the number of samples used for training. In [227] it is thoroughly analysed and suggested that error metrics based on absolutes rather than squares can perform better in regression problems. For all these reasons, MAE is chosen over RMSE. Because metrics like MAE might not be very intuitive in visual terms, the results presented here include three other expressions of the prediction error, just for the sake of clarity. First, the full area error % is defined as the percentage ratio between (i) the excess area enclosed between the true (or test) capacity curve and the predicted curve, and (ii) the total area delimited by the test curve, as expressed in Fig. 2.13. Second, the fitted area error % shares the same definition as above, but limiting the areas to the lowest of the last displacement points of both curves (test and predicted) as in Fig. 2.14. Furthermore, in the figures that follow, MAE has been expressed as a percentage of the area calculated as the sum of all individual MAE for each curve point (100 points), times the interval between points (0.01). Figure 2.13: Full area error scheme (sample #1062) Third and last, the last displacement error % refers to the prediction of the end point of the curve and is defined as the percentage ratio between (i) the absolute 51
Ld predicted Ld test Excess area Test curve area Figure 2.14: Fitted area error and last displacement (Ld) error schemes (sample #1062) difference between the predicted and true end points of the curve and (ii) the true end point, as shown in Fig. 2.14. These measurements, especially the fitted area error % and the last displacement error%, allow for the separate evaluation of (i) how well the curves fit each other in terms of shape and (ii) how well the network has predicted the end of the curve. Network architecture The first attempt to define a network architecture is to implement a standard feed-forward model. Because there is no theoretical methodology to establish the optimal configuration of the network [228], a stepped process, that is, a grid search, has been designed to determine the most appropriate architecture for the neural network model. More thorough methods such as optimising the architecture with genetic algorithms [229,230] can be explored in future work. In a first phase, a single layer network is setup with varying sizes of the hidden layer: (a) a hidden layer equal to the size of the input, (b) equal to the output layer, (c) a value in between, and (d) larger than the output layer, as seen in Fig. 2.15. This helps in assessing what range of sizes in the hidden layer suits the problem best. With this information at hand, another set of architectures of varying depth (i.e. a varying number of hidden layers) is evaluated. These layers use the previous size value as an initial or tentative layer size. It must be mentioned though, that this value is used only temporarily since more adjustments on the sizes of the hidden layers will be carried out later in the process. Deep neural architectures can be extremely powerful. When applied to regressions of continuous variables and time series (similar to the examples in this study), these models show a great performance [231,232,233]. However, there is a 52
Figure 2.15: Varying hidden layer sizes, (top left) h1=30, (top right) h1=65, (bottom left) h1=100, (bottom right) h1=135 trade-off between the complexity that a neural network can bring forward and the overfitting of the model to the data. Due to overfitting, deep architectures may perform very well on the training samples but poorly on the validation data. To test the interplay of these factors, three initial architectures are presented primarily according to the number of hidden layers (depth). In the following, Fig. 2.16 shows a diagram of these schemes. Figure 2.16: Varying depths in network architecture, 2 (left) and 3 (right) hidden layers. In both Figs. 2.15,2.16,xrepresents the input neurons of the network. As seen in previous sections, the total number of inputs for the network is fixed to 30 as this matches the number of parameters that have been used to generate the buildings. Therefore, hrepresents the hidden layers in the model. An initial size is temporarily set at 65 as it is the best performing size in a preliminary and intuitive evaluation, but more refined values will be explored later. Finally, yis the output 53
layer, which corresponds to the values of the capacity curves that the network is aiming to predict. The size of this output layer has a strong impact on the performance of the learning process. Small sizes in this layer reduce the level of difficulty of the predictions; however, a valid resolution for the curves must be ensured. For this reason, the number of neurons in the output layer has been initially set at 100 because it provides enough resolution to graph the capacity curve as discussed and it is slightly higher than the maximum number of points present in the original capacity curves as retrieved from SAP2000. However, variations in this output size will also be discussed later. From this outset, the main algorithms of the network will be explored, namely the activation functions that are implemented in each layer (which may be considered an algorithm when viewed as a whole) and the error optimisation algorithm. Activation function Neural networks are very sensitive to this election, and there are many possible activation functions used in neural networks, so it must be determined which one fits the problem best. In [234] it was observed that a Hyperbolic Tangent (Tanh) is a common activation function applied to the hidden layers of a neural network for similar problems as the one discussed in this study, while sigmoid activations are commonly used for the output layer (as they map from 0 to 1). For this reason, both Tanh and sigmoid activation functions will be tested in hidden and output layers, respectively. Alternatively, in a fine-tuning phase, Rectified Linear Unit (ReLU) activation has also been tested due to its common use in deep neural architectures [235] but with no positive results. Optimiser A preliminary choice for the optimisation algorithm is Stochastic Gradient Descent (SGD), which is the basic algorithm for neural network training [236]. Nevertheless, in a later stage, a second algorithm has been tested, namely the Adadelta optimiser [237], which is a variant of the SGD that presents a novel learning rate method per dimension for gradient descent by dynamically adapting over time. The application of this optimisation algorithm did not improve the error rate obtained using SGD. Preliminary adjustment of SGD parameters As explained in the methodology section, the first objective is to fix the architecture of the network and the parameters of the optimiser algorithm. In order to test various architectures reliably, a satisfactory set of parameters is required to 54
guide the backpropagation of the error. For this purpose, a simple initial architecture is defined. The initial conditions for the preliminary adjustment of SGD are shown in Table 2.3. SGD is adjusted through the following parameters: Learning rate (Lr), Decay, and Momentum (M). The results are measured against the validation set using MAE (Table 2.4 and Fig. 2.17). Network Architecture Layers X h1 Y Layer size 30 65 100 Network Parameters Layers X h1 Y Activation - Tanh (Hyperbolic Tangent) Sigmoid Weight initialisation - Random seed with range (0, 0.1) Random seed with range (0, 0.1) Bias initialisation - Random seed with range (0, 0.1) Random seed with range (0, 0.1) Training Parameters Epochs(ep) 200 Batch size 12 Shuffle samples at each epoch Yes Table 2.3: Initial conditions for preliminary Stochastic Gradient Descent (SGD) adjustment. Lr (Learning rate) Decay M (Momentum) MAE loss MAE loss MAE loss 0.15 0.0159 Lr/800 0.0158 0.80 0.0160 0.20 0.0157 Lr/1000 0.0156 0.85 0.0156 0.25 0.0156 Lr/1200 0.0155 0.90 0.0150 0.30 0.0156 Lr/1400 0.0154 0.95 0.0181 0.35 0.0154 Lr/1600 0.0150 Nesterov 0.0149 0.40 0.0157 Lr/1800 0.0152 0.45 0.0159 Lr/2000 0.0154 Fixed parameters Decay = Lr/1000 Lr = 0.35 Lr = 0.35 M = 0.9 M = 0.9 Decay = Lr/1600 Table 2.4: Experimentation and results of preliminary SGD adjustment. Network architecture configuration Once the optimisation parameters have been set, it is then possible to test different network architectures more effectively. In these tests, Tanh activation in the hidden layers and sigmoid activation in the output layer are fixed. 55
Figure 2.17: Loss/Epochs graph with selected SGD values of preliminary adjustment. In Table 2.5, the initial conditions for the variations in network architecture are shown. Then Table 2.6, Figs. 2.18–2.20 present the tests and results of the network architecture showing different variations over hidden layers. In Table 2.7, the results of the curve resolution (size of the output layer) are listed. Figure 2.18: Validation loss/Epochs graph for 30-205-100 layer scheme (1 hidden layer) Network parameter fine-tuning After obtaining the best error rates for all the configurations tested, a final architecture 30-65-65-100 is selected. From this setup, a more refined set of variations is executed. These variations include a second round of SGD optimiser parameters, as well as an attempt on the Adadelta optimizer, ReLu activation for hidden layers, and different batch sizes. The number of training epochs is always prolonged until the error does not improve significantly. 56
Network Parameters Layers X h1 Y Activation - Tanh (Hyperbolic Tangent) Sigmoid Weight initialisation - Random seed with range (0, 0.1) Random seed with range (0, 0.1) Bias initialisation - Random seed with range (0, 1) Random seed with range (0, 1) Training Parameters Lr (Learning rate) 0.35 Decay Lr/(8.ep) M (Momentum) Nesterov Epochs (ep) 800-1200 Batch size 12 Shuffle samples Yes Table 2.5: Initial conditions for network architecture variations. Layer scheme & size MAE loss X h1 Y Validation error Training error 30 30 100 0.0153 30 65 100 0.0135 30 100 100 0.0132 30 135 100 0.0131 0.0115 30 170 100 0.0132 30 205 100 0.0133 Layer scheme & size MAE loss X h1 h2 Y Validation error Training error 30 30 30 100 0.0133 30 30 65 100 0.0137 30 30 100 100 0.0137 30 55 80 100 0.0128 30 65 65 100 0.0126 0.0106 30 65 100 100 0.0131 30 100 100 100 0.0134 Layer scheme & size MAE loss X h1 h2 h3 Y Validation error Training error 30 30 30 30 100 0.0137 30 30 30 65 100 0.0133 0.0107 30 30 65 65 100 0.0138 30 65 65 65 100 0.1407 Table 2.6: Results of hidden layer/s architecture. 57
Figure 2.28: Fitted area error ˜ 5.0%. Validation sample #147. Figure 2.29: Ld <5.0%. Validation sample #639. Figure 2.30: Ld ˜ 21.32%. Validation sample #1971. 2.5. Discussion Following the methodology described in the previous Section, initial tests were carried out to find an optimal value for the SGD parameters of the network. The 64
Figure 2.31: Ld ˜ 60.0%. Validation sample #1926. results have revealed that with a simple initial architecture (30-65-100) the network is capable of high prediction accuracy. The tests on different network architectures yielded the best results for configurations that featured two hidden layers; in particular, the scheme that delivered the best results was 30-65-65-100. Architectures with only one hidden layer returned better results with larger sizes, but still not as competitive as the latter. The fact that the network clearly performs better when increasing its complexity accounts for the level of difficulty of the predictions. However, after a certain point, increasing the complexity of the model does not improve the results due to overfitting. Perhaps, with an even larger training set, these deeper architectures may improve the results, thus leaving room for future work. Regarding the size of the output layer, and rather counter-intuitively, smaller sizes of curve resolution have not improved the metrics. An initial output layer size of 100 was set because it was considered to have enough resolution for the problem at hand while not being excessively large for training. Interestingly, lowering this value proved detrimental while eventually increasing its size to 135 yielded equally accurate results. This can be explained by the fact that lower resolutions are less loyal to the calculation algorithms within the engineering software which produced the validation set of capacity curves (SAP2000), and since neural networks perform best with clear patterns, lower resolutions introduce harmful noise in the training process. Finally, in the fine-tuning stage, batch sizes played an important role in maximising the curve prediction accuracy. Although there may be some disagreement about the regularising effect of batch size [239,240], in these experiments it was found that larger batch sizes can prevent overfitting by regularising the network to some extent, because loss values are averaged for all 65
the elements in the batch and then backpropagated to adjust the weights and biases of the model. Table 9 shows how the lowest batch sizes had very good training results, but lagged when tested against the validation set, thus showing a more acute overfitting. The best performing batch size tested was 24 samples. In the methodology presented, no modelling of the specific building is required. Curves can be estimated with an average curve-area fitting above 97.6%, requiring only the basic geometric parameters of the building to be specified. The main conclusions of this study can be summarised as follows: • ANN provide an accurate approximation method for the nonlinear static pushover calculation of low-rise structures within a wide range of sizes and geometric configurations. • The accuracy of the method successfully addresses the shortcomings of current macroseismic approaches while remaining fast and efficient. • Stress-deformation curves in a plastic regime can be predicted with ANN in one go for entire buildings using only basic geometric parameters. For low-rise structures, this work achieves a curve area error below 2.7% and a resolution of up to 100 points. • The relative simplicity of the ANN architecture required to predict the capacity curves of low-rise buildings makes a strong case for future research of highrise structures using deeper networks and larger datasets. In future work, it may be interesting to include the use of genetic algorithms to evolve an even more optimal network architecture. Furthermore, a comparison with other regression methods or ML approaches would be desirable to contextualise the results obtained. In the long term, it might be interesting to explore a similar approach with dynamic analysis, high-rise buildings with higher vibration modes, and a wider variation of sectional and material properties. It should also be noted that this work has followed the Eurocode-08 capacity spectrum method, which has been proven to work well with low-rise structures. However, a research effort to predict capacity curves under a set of strong earthquakes that can produce greater nonlinearities would extend the applicability of the method presented here. All these cases (dynamic analysis, high-rise buildings, and stronger earthquakes) pose a greater challenge in terms of ML and may require the use of more specialised neural network architectures. In this regard, because capacity curves can be regarded as time-dependent series, network models better suited to handle dynamic input, such as recurrent neural networks [241,242], should be explored. 66
3. Application of generative Artificial Intelligence to seismic resilience: Automated extraction of building typologies from digital urban cadasters This chapter is meant to be a transition from engineering applications of neural networks to the world of neural models of generative AI. It is structured as a case study that builds on the work presented in the previous chapter. In particular, in the following sections, a method is proposed to extract building typologies from a digital land cadastre. In doing so, the intention is to expand the study on seismic demand prediction to the most relevant typologies of any given city. Thus, although this chapter can be seen as a continuation of the previous one on seismic engineering, it shows a more functional angle (utilitas) on an urban scale. At the same time, it introduces an example of the VAE model that will be explored in greater depth in the next chapter. The VAE model that will be used here targets image-based inputs of building roof-prints, which serves as an introductory case study to the more challenging three-dimensional problem presented later. 3.1. Introduction and literature review Recent advances in Machine Learning techniques are having a great impact on the field of macroseismic analysis. Until not so long ago, there used to be a trade-off between the level of detail of the analyses and their feasibility, particularly in terms of cost, time consumption and man power. Therefore, macroseismic studies devoted many efforts to finding effective strategies for the delivery of acceptable approximations to the assessment of large urban areas, employing a reasonable amount of time and resources. However, this status quo is being transformed at a very rapid pace with the development of ML applications for these macroseismic challenges. The previous chapter [243] presented an artificial neural network that was trained to predict the capacity curves of prismatic reinforced concrete (RC) buildings of up to four stories and spanning a very wide range of floorplan dimensions. The presented neural network model can output capacity curves from only basic geometric parameters, such as the number of spans in the X,Yand Z axes and the features of the beams and columns in the structure. As a consequence, emergency relief agencies can assess the seismic vulnerability of large numbers of buildings very quickly without engaging in tedious structural modelling. The work presented now is part of a larger effort in expanding these 67
developments to account not only for prismatic buildings, but for the major building types present in a given urban area. Recent work addressing the same objectives can be found, for example, in [244]. However, the approach here attempts to achieve en-masse seismic analysis at the level of individual building modelling. A first step in this direction entails identifying these types in an efficient way, which is the main objective of this case study. For the vast majority of buildings, the study of their typology can be addressed solely from the shape of their floorplan boundary or perimeter, also commonly referred to as ‘roof-print’. This is an important advantage, mainly due to the fact that roof-prints are typically available as part of a city’s geographic information system (GIS). For this case study, more than 100k of these shapes have been collected from the digital land cadastre of Seville (Spain) into a database both in vector and raster (pixel) format. The challenge, as mentioned earlier, is to extract or cluster these roof-print shapes into coherent types that are representative of parametrically similar structural definitions. This will allow the creation of large databases of building structures to train ANN models, as described in [244]. Shape Clustering is a well-developed field of research with important contributions dating back, at least, to the 1970’s (with even earlier studies by Sir D’Arcy Wentworth Thompson in 1917). Kendall defined shape as the geometric information that remains after filtering out location, scale and rotation [245]. Then in [246], Procrustean analysis (coined by Hurley & Cattell in 1962 [247]) was introduced to the analysis of shapes, and consolidated in “Procrustes Methods for the Statistical Analysis of Shape” [248]. The Procrustes method involves finding the optimal translation, scaling and rotation of a shape with respect to a reference shape. The objective is to minimise some error, or distance, established among the two sets of landmark points of each shape. Landmark points in turn are defined as representative points within the shapes that match across the population of shapes under study. More works following this initial phase were carried out in the early 2000s [249, 250]. Then, a more general and powerful method was proposed by [251], which was able to remove the need for landmark-based shapes and compute shapes in the form of full continuous curves. These achievements were further consolidated by other works such as [252,253]. Another important development was brought about by the work of Amaral et al. [254]. In their contribution “k-Means Algorithm in Statistical Shape Analysis”, the authors opened the door to the study of shape clustering using standard clustering algorithms. However, as in all previous studies, the result of these computations relied heavily on the error or shape distance selected. In [255], several metrics and appropriate clustering methods are evaluated for a set of twodimensional shapes that include Euclidean and non-Euclidean metrics such as the 68
Riemannian distance. Finally, some works have targeted the application of these shape clustering methods to the fields of Architecture and Design [256,257,258, 259]. Shape clustering as per the current state-of-the-art relies on prior knowledge of the shape space. For example, it needs the number of clusters to be specified beforehand, and, more importantly, there has to be a process to either derive the reference shape for each category or a reference shape has to be directly provided altogether for each category so that the distances or error metrics can be established upon. Therefore, unsupervised exercises are perhaps better suited for different clustering techniques and approaches. For this reason, a Variational Autoencoder (VAE) [260] model is tested for the purpose of extracting typologies from a land cadastre. This approach is also useful for filtering out typologies that are not representative or do not appear in sufficient numbers within an urban context. This is a natural effect that derives from the training aspects of neural networks, where less frequent samples in a training dataset are ’ignored’ by the model. Clustering with VAEs has witnessed a strong increase in contributions from the literature in recent years [261,262], including the use of convolutional architectures for image clustering [263], as is the scope of this case study. While the aforementioned studies focus on benchmarking against well-known datasets (achieving impressive results), here the objective is to test a straightforward model on a real-world problem, and it is expected that the application-specific potential and validity of the method may be assessed in the context of building type analysis within large urban areas. 3.2. Methodology The basic premise of the method is the use of a VAE to perform unsupervised clustering of images. A large set of pre-processed roof-print image samples is fed to the model for training. Once the VAE is trained, the latent space is approximated by a Gaussian distribution. Sampling from this space, a set of reconstructed images provide a summary of the building typologies in the original dataset (see Fig. 3.1). All of these steps are explained in detail throughout this section. The shapes studied are taken from the official land cadastre of the city of Seville in Spain, which is available as a vector-based ESRI shape file. The file contains the roof-prints of all the buildings and other built entities in the municipality of Seville, totalling more than 100k records. These outlines are in the form of closed polylines and are provided with a set of data fields indicating the number of floors, the date of entry in the database, the surface area and other information. On a preliminary visual exploration, the following observations are made: 69
Figure 3.1: Graphic summary of the methodology. • Polylines present an arbitrary number of vertices (for example, a more or less perfect rectangle may have a lot more than four points). • The outlines correspond not only to buildings but also to other urban structures such as retention walls, underground and surface parking etc. • The drafting quality of the outlines is rather deficient in many cases, whereby polylines often form sharp spikes that do not seem to correspond to the actual built objects. • Many buildings that are known to be orthogonal present some angle distortions that although not exaggerated, can still be identified as non-orthogonal in a quick visual check. To counter these issues, a number of preprocessing procedures are employed. 3.2.1 Pre-processing Because the number of points will be used further in the method to create different datasets, the first simplification procedure will attempt to eliminate irrelevant polyline points. The criterion is to eliminate a vertex that joins two segments with an angle smaller than a threshold value (see Fig.3.2). This threshold value (<π/40) has been adjusted so that curved building outlines are not affected. The second pre-processing step aims to identify and remove sharp spikes that result from ‘sloppy’ drafting. In this case, the strategy adopted has been two-fold: first, sharp angles (below a chosen threshold: <π/8) are flagged, and then those flagged angles belonging to short polyline segments (a second threshold is set here: <1m) are removed (see Fig. 3.2). Third, since the shapes under study should correspond to building typologies, 70
entities that are clearly neither buildings nor representative buildings are filtered out. To this category belong (i) very large (>900m2) or small (<100m2) entities, (ii) those that do not span a single floor in height, and finally, (iii) linear-like entities, defined as a minimum threshold ratio between their surface and perimeter length. Figure 3.2: Pre-processing steps #1 and #2: removal of sharp spikes and redundant vertices. Additionally, further pre-processing is carried out to address potential issues during the learning process of the VAE model. The dominant issue is that a VAE model can learn the distributions of position, size, and rotations across the samples, as discussed in the previous chapter [264]. Thus, similarly to the Procrustes idea described earlier, all samples are translated to the same position in space (0, 0)and scaled (isomorphic) to the same bounding-box size. Rotations, however, pose a greater challenge, because it is not possible to determine a reference rotation for any of the shapes without a prior heavy-duty study. Hinging on the design rationality of the vast majority of the building stock of a city, a fourth pre-processing strategy that identifies the principal vector directions of each building outline is proposed. In this approach, all segments of the outline polyline are clustered using a reference angle (<π/20), which means that in each cluster, any given pair of segments present an angle lower than the chosen threshold. The average angles or directions of the clusters that have the longest aggregate segment length (sum of the lengths of each segment in the cluster) become the principal directions. It is acknowledged that PCA (Principal Component Analysis) might be a more suitable approach to solve this problem. An example of principal directions is illustrated in Fig. 3.3. Finally, each shape is rotated to align the top-most principal direction to the Xaxis. Of course, this method does not guarantee that same-to-be shapes may not present a 90º, 180º or 270º with respect to each other, but it does reduce the rotation differences that might be learnt by the VAE to only those three options. 71
Figure 3.3: Pre-processing step #4: rotation along top-most principal direction and visualisation of principal directions. 3.2.2 Generation and augmentation of training and validation sets For the generation of the datasets, all the preprocessed vector outlines are converted into solid image-based shapes (black-coloured with greyscale antialias borders). For buildings with patios, the patio is coloured in the same colour as the background (white). The main reason being that VAE models and in general neural networks are very efficient at working with pixel-based data such as images. Because there are potentially quite a number of different building types within the dataset, the challenge for the model is rather huge. For this reason, and despite the already large number of samples available, a data augmentation scheme has been implemented: each sample is randomly rotated within a small angle range (±π/50) and also displaced randomly within a marginal range (±2.5% of the width and height of the image for Xand Yrespectively). With this approach, depicted in Fig. 3.4, it is possible to augment the dataset more than 100-fold. Two complete datasets (before splitting into training and validation) were generated for experimentation: a ’flat’ dataset (set A) where each sample was randomly rotated and translated 100 times, increasing the available data 100-fold, and also a ’progressive’ dataset (set B), where the number of instances randomly generated from each shape was proportional to its number of landmark vertices (number of instances =100 ∗(n+ (n−4)), where nis the number of vertices). The intent behind this idea is to allow more complex building types (assuming simpler shapes are over-represented) to increase their relative representation in the final dataset. In essence, the strategy devised for the progressive dataset aims to alleviate the well-known class imbalance problem [265]. Both datasets feature a number of samples above 1.5M and are split 70 / 30 into training and validation sets, respectively. 3.2.3 Variational Autoencoder model Finding the optimal neural model for a given dataset and loss function is a non-trivial task. In fact, there is no formal method to determine an optimal 72
Figure 3.4: Dataset augmentation through random rotations and displacements. architecture for deep learning models [228]. More so, in the case of VAEs, as their loss function includes not only the reconstruction error but also the Kullback-Leibler (KL) divergence [266] between the target and the obtained latent space distributions. However, there are methods that help to find a suitable model configuration, such as Genetic Algorithms [229,230]. These methods would involve very heavy computation power for the dataset at hand (beyond a regular parallel computing platform such as the CUDA engine – which has been utilised in the experimentation). They also require cutting-edge expertise in deep learning. The aim of this work is to prove that the method is suitable for the problem at hand, at the level of a proof-of-concept, and not to find an optimal performance of the model. In light of this context, many of the network configuration settings are drawn from previous research [264]. The experimentation is conducted in two phases: a preliminary exploration and a full-scale study. The objective of the preliminary exploration is to acquire a ‘feel’ for the model. In this phase, a reduced dataset is used of about 150k samples, which is 10% of the final dataset used. The reduction is achieved simply by limiting the generation of random samples from 100×to 10×, so that the diversity of building types remains constant. The range of architecture definitions of the VAE models designed for this preliminary phase is presented in Fig. 3.5, from the first and simplest iteration (a) to the best performing (b). This series of models features two 2D convolution layers of varying depth and convolution matrix size. The first is applied over the input layer that is 60×60 in size across all experiments. The size of the second one is exactly half (30×30) or a third (20×20), and the transition from one to the other is processed by a down-sampling layer. The method selected for the reduction in size is averaging (the value of each neuron in the target layer is computed as the average of the corresponding values of the source layer). At the 73
shape may not add much more valuable information to the models. Thus, from a methodological perspective, it seems as though the data augmentation factor used for phase 1 is within a better magnitude range for the problem at stake than the size implemented in phase 2. For this reason, no tests have been conducted for the full progressive dataset (set B) and the final shape extraction results have been taken exclusively from phase 1. With regard to the extraction of building typologies, a total of 24 to 26 different types have been extracted depending on criteria (Fig. 3.9). Most of them are sourced from the progressive dataset, as it performed better in terms of loss minimisation and also delivered the broadest variety of shapes (Fig. 3.10). Of course, there is a substantial overlap between both sets; thus some of the types in Fig. 3.9(up) can also be extracted from the flat dataset (set A). Therefore, Fig. 3.9(down) shows only those types that are not extracted or are not very clear from the set B. For example, D3 corresponds to a very common typology commonly referred to as ‘H-block’ but is not clearly depicted. However, E6 portrays a much more accurate representation of this typology. A similar case occurs between A6 and E2, although in this case the difference is barely visible (the latter is a bit clearer than the former). This type also corresponds to a very common building that is found in the old quarter of the city centre of Seville. It features a small patio on the back, while the fac¸ade articulates a built-in body that hosts a set of glazed balconies along each storey. Then, E3 (circular or round buildings), E4 (L-block) and E5 (chamfered L-block) have been extracted only from the set A(not present after sampling from PH1 B2). Finally, E1 (open C-block) could arguably be classified in the same category as C4. In such cases, it is difficult to establish clear criteria for when to draw the line between two categories. In fact, it may be argued further that C6 should belong to this category as well, and that the distinction between a C-block and an open C-block is rather weak and unsubstantiated. A similar situation also occurs between B1 and D4. In reality, these types of building correspond to very different architectural configurations. Thus, it may be suggested that the capacity of the model to provide this level of differentiation represents an important (and unforeseen) strength of the method. Overall, the models have captured the most well-known building types, and also many others that would have required long manual supervision to identify (such as A1, B6, C1, and B2). Furthermore, the method has filtered out many rare or uncommon typologies that are not representative of the urban fabric (for example, the buildings in Fig. 3.3). It is also very possible that the method has not been able to identify some relevant building types. Finally, it is very important to note that VAE are essentially generative. This means that they are able to create new shapes by interpolating two or more existing ones. In relation to the matter of this case study, it implies that some of the building types extracted and presented in Fig. 3.9 might in 80
fact be entirely fictional. In general, it is expected that these interpolations present a less defined image profile than their non-fictional counterparts. However, from a local expert’s point of view, the suspected fictions within the results presented are B2, B5 and D1. In future iterations, instead of sampling all across the latent space, it would be more appropriate to sample only from the points of the space that correspond to real samples. In this way, fictional reconstructions could be avoided. 3.5. Conclusions This case study has presented a method based on the VAE model to extract the most common building typologies from a digital land and real estate cadastre. The work has focused on the Seville cadastre as a case study to establish a proof-of-concept. The approach differs widely from conventional shape clustering analysis and offers a series of advantages that are specific to the problem at hand. First, no prior knowledge of the types within the dataset is required. For example, the number of clusters does not need to be specified beforehand. In fact, the output is not framed in terms of a rigid set of clusters, but rather displays a gradual transition between learnt categories. Second, less common building types that have little presence in the city are filtered out of the extraction by default; since neural models learn by iterative processes of loss optimisation through repetition, non-representative samples are naturally punished out of the learning space. And third, the method allows for the computation of subtle differences within categories, rendering multiple variations of the same cluster. This is due mainly to the ability of VAEs to pick up shape variations (such as, but not limited to, scale and rotation) during the learning process, as discussed earlier in the Methodology Section. Overall, the following conclusions can be derived from this study: • The method proposed is very sensitive to the data augmentation factor. Larger datasets that result from stronger augmentation factors do not necessarily improve the learning curves. This sensitivity is a well-known phenomenon in deep learning, and comprehensive reviews have been published addressing the problem [265]. • The models have been trained to a degree of accuracy just enough to prove the validity of the method. From an optimisation perspective, there is great room for improvement of the models implemented in this work. • The most common typologies have been successfully extracted (a total of 10). A substantial number of less obvious typologies have also been extracted (a total of 14 to 16 types depending on criteria). And up to three of the types are suspected of being generated artificially by the model (fictional). A plausible 81
solution to identify these fictional typologies could be implemented by tracking the latent space distance from each typology to its closest training sample. This distance would provide a metric of the ‘fictionality’ of the building types extracted. • The results show that the method presented is a suitable alternative to conventional shape clustering in the context of urban typologies, and as such, future work should expand and consolidate this research. As a last remark, this work is hoped to have a positive impact on the field of seismic emergency and relief. By providing tools that alleviate the burden of identifying and modelling en-masse the structural behaviour of buildings in urban areas, great benefits can be provided to hundreds of cities world-wide, not only with regard to economic savings, but also, and most importantly, in terms of human lives. 82
4. Generative Artificial Intelligence in Design This chapter focuses on Generative AI and its connection to the field of design. As mentioned at the outset of this thesis, those aspects in Architecture related to aesthetics (venustas) constitute the most challenging application scenarios of AI. In the last decade, the field of ML has witnessed a particular revolution on this front. The development of Variational Autoencoders (VAE) and Generative Adversarial Networks (GAN) in 2014, have opened up an ocean of possibilities for the interaction between AI, Creativity and Design. These two models shocked the world with their ability to ‘create’ fictional, yet extremely realistic, images of various kinds. Among them, these models were capable of generating faces of nonexistent celebrities with a remarkable feel of veracity [267]. Their success, exemplified by the latter and other equally impressive experiments (if not more), have consolidated in the birth of what is currently known as Generative AI. Of course, such developments spurred intense debates (and continue to do so) in Art, Philosophy, and Design circles, but also in the AI community itself. Important subject matters like the idea of authorship or creativity saw significant upheaval with each release of the ever-increasingly stunning results coming from these generative AI models. Given that the relevance of the topic is clear, the next sections aim to delve further into the connection between generative AI and the field of Architecture and Design. It should be noted, before moving ahead, that another important flagship has recently joined the fleet of generative AI; namely, Large Language Models (LLM). However, this chapter will consider only generative models oriented to work with non-symbolic data (e.g., pixels) and not those that focus mainly on semantic input, like LLMs. The reason is that some of the main challenges facing the collaboration between AI and Design is based precisely on the bridging of sensory perception and conceptual reasoning; the sensory part, of course, requiring the ability of analogue sensory data, i.e., non-symbolic data. It should be noted that the most recent research in this field combines LLMs architectures with non-symbolic models like VAEs or GANs, and therefore one must not abide by this distinction too strictly. 4.1. Background The possibility of building creative machines or machines with the ability to autonomously generate their own artefacts has captured the imagination of researchers and scientists since the early days of AI. Much of the initial 83
investigations revolved around the quest for artificial or autonomous ‘creativity’ rather than specifically honing on the notion of ‘generative’ machines or generative AI. In fact, the term generative was associated with slightly different ideas in AI until not so long ago. For example, in 2010, a Ph.D. thesis under the title ‘Generative AI’ [268] held merely a faint connection to what is actually known as Generative AI nowadays. Instead, the work hinged on neo-Cybernetics and the need for an AI paradigm that, in itself, was able to ‘generate’ intelligent systems. Furthermore, the 1980s saw the emergence of the research area ‘AI-based Generative Process Planning’ [269,270]. The works in this area were aimed at capturing in a computer program the logic used by a process planner to convert design information from engineering drawings to process plans. Another area of generative studies, especially in engineering and design, was spearheaded by models featuring autonomous agent ensembles. For instance, a paper from 2002 [271] describes situated cognition as the basis for group creative design, which is implemented through a multiagent model. Another, from 2007 [272], explores generative designs through an interactive and evolutionary system with ML. Although there is no question about the ‘generative’ capacities of agent-based approaches, especially in the realm of Architectural Design, this area deviates from the core methods of modern Generative AI. Finally, some early attempts were also made, even prior to the works mentioned above, in the field of language. A good example can be found in the chatbot ELIZA [273], which precedes the generative grammar studies that are discussed below. In contrast, research that focusses on creativity appears to be more in line with contemporary studies in generative methods. Also, this avenue involved deeper and more fundamental questions pertaining to AI as a whole and, as such, was backed by a broader community of renown scientists (Hofstadter, Rowe, Martindale, etc.). The pursuit of creativity was seen by these authors as an important part of the larger question of intelligence. In the words of Rowe and Partridge: ”Intelligence involves creative behaviour. Few would challenge this statement, but agreement on what is meant by ‘creative behaviour’ would be much harder to find...” [274]. Due to how central creativity was perceived in relation to the fundamental question of intelligence, important efforts were put into this topic. A particularly interesting legacy from those days is an account of five conditions for creative behaviour proposed by the aforementioned authors. A brief outline of each one is copied below for reference, as it provides a useful framework upon which to discuss some of the more recent works later in this thesis: • Firstly, it is necessary that knowledge is organised in such a way that the number of possible associations (the creative potential) is maximised. 84
• Secondly, it is necessary to tolerate ambiguity in representations. • Thirdly, there is a need for multiple representations. • Fourthly, the usefulness of new combinations should be assessable. It is of no use to create many combinations if they are all useless. • Lastly, any new combinations need to be elaboratable to find out their consequences. [...] The consequences discovered should be applied to the problem situation; the usefulness of a discovery is decided by its applicability. Following is a selection of initial efforts in artificial creativity that align with the previously mentioned directions. Some of the first attempts on creativity took place in the field of language and language-oriented domains. For example, in [275,276], generative language programs were developed with explicit support of syntax and grammar rules. In [277,278], generative experiments were carried out in the field of music composition. All of these works had incorporated an element of randomness in their generative engines, which made them very interesting at first but rather shallow at the end of the day. Here is a text created by Ractor [276] in 1984: ”Bill sings to Sarah. Sarah sings to Bill. Perhaps they will do other dangerous things together. They may eat lamb or stroke each other. They may chant of their difficulties and their happiness. They have love but they also have typewriters. That is interesting.” Or the more popular one: ”Reflections are images of tarnished aspirations.” As Rowe points out, creativity requires a sense of novelty, which sometimes can be achieved by introducing randomness in the same way that Ractor was built. However, in his own words, it seems that ”novelty is not enough”. This is a very interesting statement in light of the current developments in Generative AI. Other authors explored other areas of creativity. Lenat and Davis introduced the program AM [279] in 1982, which was built with the intention of discovering mathematical concepts autonomously. Other studies that contributed to this or similar models can be found in [280,281]. Discovery was also attempted in the domain of undirected graphs by Epstein [282]. Other areas of creativity research included meta-rule-based models [283] and analogy-based models [284]. Lastly, other approaches were more lenient towards decentralised and connectionist approaches, with important contributions from Hofstadter [285] and Minsky [286]. For instance, Rowe proposed the GENESIS model [287], which was rooted in Minsky’s influential ‘Society of Mind’ [288]. As a final remark, it is interesting to 85
review some of the thoughts shared back then on connectionist systems and their ability to become generative. As Rowe writes: ”These networks are generally used to model very low-level behaviour such as vision. Whilst they are good at generalising from a given data set, the learning procedures are very artificial and much research is still needed before connectionism could be applied to the larger problems of creativity.” All in all, the majority of these efforts were approached from the paradigm of symbolic AI (with the few connectionist exceptions mentioned). In such models, creativity is seen under a combinatorial perspective. Although all of these models are different in the way symbols are combined to form new structures, they are ultimately similar in their combinatorial nature. Some methods, usually based on tree-like knowledge representation schemes, such as the COBWEB algorithm [2], did present a hybrid strategy of combination and interpolation to achieve generative abilities. Other ones, like FCA, can also present generative capabilities with ease. However, all of these methods operate under the symbolic paradigm and therefore struggled greatly when dealing with sensory data originating in analogue sources (e.g., pixel-based images). 4.1.1 The connectionist leap forward With the increase in computation power and training data availability during the 2000s, neural network models began to make important strides in the arena of ML. By the mid-2010s, facial recognition had become an industry standard after the introduction of AlexNet in 2012 [289]. It was clear then that the connectionist paradigm excelled in the domain of raw sensory data. Soon after, as mentioned earlier, two powerful generative models were proposed: VAEs were published in 2013 [260], and GANs in 2014 [290]. VAEs were briefly introduced in the previous chapter and will be explained in more detail in the following section. GAN model consists of two networks that are trained together; one is sampling new instances from points of a latent space, while the other is an encoder learning to discriminate the generated instances (fakes) from the examples in the training set. When the generator is capable of producing instances that the discriminant cannot distinguish from true examples, then the GAN has achieved successful training. A useful analogy to understand GANs is that of the police and the thief. The thief (generator) is producing counterfeits of bank notes and the police (discriminant) has to detect counterfeits. Since the development of these two models, the field of Generative AI has triggered a landslide of impressive developments and is still operating in full swing today. 86
The first developments aimed to enhance the quality of the images generated by these models. The lowest hanging fruit in this regard involved the use of deep convolutional architectures within GANs [291] and VAEs [292]. However, in 2017, the size at which these models were able to recreate quality images was only about 256×256 pixels [267]. In 2019 the BigGAN model [293] pushed the boundaries of high-quality image generation to 512×512 pixels with results that were, in many cases, indistinguishable from real images. Another interesting area of exploration was Style Transfer, an implementation of transfer learning that could also be seen as a form of what is known as conditional generation. This approach used generative models to learn representative patterns across a set of uniform-style images, and then applied these networks to images of a different style. In doing so, the network was able to produce the same image recreated in the style that it had previously learnt. Among the first works to achieve significant results are Pix2Pix [294], CycleGAN [295] and StyleGAN [296]. An added advantage in these works was the possibility to increase the output resolution using the new input image as a sort of scaffolding. Further iterations of these models, like StyleGAN2, also consolidated the application of generative models to image upsampling tasks. In a short time, they reached resolutions above 1024×1024 pixels. An important area of research in the field was that of disentangled representations. Disentangled learning is a machine learning approach aimed at separating distinct, interpretable factors of variation within the data into individual, independent components or dimensions. According to Bengio et al. [297], in a disentangled representation individual latent units (e.g., the different dimensions of the latent space) respond to changes in single generative factors while remaining relatively unaffected by variations in other factors. For example, a model trained on a 3D object dataset might develop latent units that independently respond to factors such as object identity, position, scale, lighting, or colour. In such a representation, understanding one factor can generalise to new combinations of other factors. These developments [298,299,300] have been critical in providing more control when exploiting the generative abilities of these models, especially in conjunction with semantic prompts and queries. An early example can be found in the (latent space) vector arithmetic for visual concepts implemented in the work of Radford et al. in 2016 [301] (Fig. 4.1). In their work, the authors experimented with vector operations in latent representations of words proposed in [302], achieving effective visual operations. Disentangled representations have made important strides, since these works, both in quality and resolution, are an important part of the current tools available in the generative AI industry. A more in-depth review of recent developments can be accessed in [303]. Another relevant line of exploration came about through the pursuit of neural 87
Figure 4.1: Results of vector arithmetic for visual concepts in the latent space. discrete representation learning (also known as quantised latent space learning). Quantised models seek to learn a discretised representation of the latent space. This can be useful, for example, when the continuous space of latent variables needs to be mapped onto a symbolic or semantic structure for use in another model. The relevance of these possibilities, especially in combination with disentanglement, is that they enable a fair degree of compositionality, with some interesting text-to-image examples capable of generating, say, a horse with elephant legs. A landmark contribution in this space is VQ-VAE [304], where the proposed model is based on previous work on Vector Quantisation [305]. It achieves high-quality images, videos, and speech and is relatively easy to train. In their method, the prior (e.g., the imposed distribution on the latent space) is not static but is learnt dynamically during training. With this strategy, the model makes another important contribution: VQ-VAE solves the problem of posterior collapse in VAEs. Posterior collapse is a phenomenon observed in powerful VAE models where the network is capable of learning to reconstruct the inputs with negligible error, causing the latent space to collapse into a single point or a very narrow area. Some other methods to avoid this issue have been published in the literature, as it still constitutes an active area of research [306,307,308]. VQ-VAE has opened a line of research with substantial impact in the ML community. Further contributions can be found in [309,310] (including a discrete hierarchical approach), with some important commercial applications built on top of its foundational ideas [311]. Moreover, other important studies were advancing the field in different directions. Auto-regressive models were proposed in 2016 based on previous work [312,313,314] and were implemented in the projects PixelRNN and PixelCNN [315,316]. In the auto-regressive approach, each element (e.g., pixel) is generated 88
one at a time based on previously generated ones. More specifically, these models generate data sequentially by modelling the conditional distribution of each element given the previous elements. The flow-based model was developed through the work of Rezende and Mohamed [317] in 2015. As opposed to the VAE and the GAN models, their proposal hinges on building adaptable distributions in the latent space that can better fit the target data (more on this topic in the Methods Section). Later, in 2020, two other methods were released in the Generative AI space. Taking inspiration from statistical physics, energy-based models have been around since the early days [318], but they began to attract more attention in recent years [319]. These models provide a uniform framework for both probabilistic and non-probabilistic methods, and as such can be viewed as a generalisation of most probabilistic methods. In these methods, there is no distribution normalisation imposed on the latent space, and therefore those probabilistic models that do impose a normalised distribution can be seen as a particular case. The release of this constraint offers many advantages, but also some disadvantages. Specifically, the process of sampling new instances becomes less straightforward (it involves an iterative process based on Langevin dynamics [320]). Among the advantages, it may be highlighted that (i) energy-based models feature built-in compositionality with other models and also, (ii) while in both VAEs and Flow-based models the generator must learn a map from a continuous space to a possibly disconnected space containing different data modes (which requires large capacity and may not be possible to learn), energy-based models can easily learn to homogenise disjoint regions. Due to the strong benefits of the model, intense research is currently ongoing in an effort to bridge the results gap with other methods in the field. The other model also drawing inspiration from Physics is the Denoising Diffusion Probabilistic Model [321], also known as Diffusion Model for short. In essence, this model operates by progressively adding Gaussian noise to the training data, effectively corrupting it. The model is then trained to reconstruct the original data by reversing the noising process. Once training is complete, it can generate new data by sampling random noise and applying the learnt denoising process to it. The method is seen by the authors as a generalisation of auto-regressive models. Some advantages of the diffusion models are training stability, high-quality, diverse outputs, ease of training, and a lower tendency to overfitting than in the mainstream generative models. Due to their high efficiency, this method has seen an epic explosion of research and industry-led applications across the board at a very rapid development pace. For example, popular platforms for AI-generated text-to-image graphics, such as Stable Diffusion or Midjourney, employ it. More works on this approach include (i) the Latent 89
Kullback-Leibler divergence (KL) between the distribution (normally a Gaussian distribution) of choice and the obtained distribution: Loss =Loss(X,X′) + KL where, KL = n ∑ i=1 σ2 i+µ2 i−log(σi)−1 and, Z=Gaussian(σ,µ) This leads to the generalisation posed above, where the loss function is expressed as a function of both the latent and the output variables. This straightforward approach makes VAEs good candidates for early-stage experimentation. Other models like the very popular GAN or Diffusion Denoising Probabilistic Model present slightly more complex architectures and training processes. Especially, the former is known to be particularly hard to train, and the latter typically involve lengthy and impractical output generation times. Additionally, both of these methods involve a gradual composition process of images from random Gaussian noise. This technique may work well for some types of data, like pixel-based images, but may not be as suitable to other domains that are of less continuous nature. For example, in the architectural domain, building structures can be defined as graph structures; a series of interconnected columns and beams. This graph-like representation will be chosen for the case study presented later in this chapter, and therefore a VAE model that does not introduce this kind of Gaussian noise decomposition may be better suited for the experimentation. Other methods like the Flow-Based models or Energy-Based models discussed earlier have important advantages; however, their complexity and less mature phase of adoption may render them more suitable for a future phase of experimentation. Another important topic to discuss with regard to the idea of latent variables and latent spaces within the context of generative models is the subject of disentanglement. As mentioned above, latent variables can be useful in capturing features in the data that are not explicitly provided during training. For example, one may hope that in a deep learning architecture, the sequence of hidden layers will effectively learn a certain level or granularity of features; the deeper the layer, the higher the level of abstraction of the features captured. Very deep layers could detect, for instance, a certain direction or some relative positioning, while more shallow ones could detect a part of the mouth or the eye. This hypothesis has been 96
intensely explored throughout the deep learning paradigm, delivering mixed results [334,335,336]. Yes, this phenomenon does take place to a reasonable degree, but at the same time, it is extremely hard to predict what layers will capture what and to what extent. In fact, the nature and abstraction level of the features collected at each layer are very sensitive to the design of the model and the training data (including the order in which the samples are fed into the network). Consequently, they vary widely from one network configuration to the other, leaving the modeler with rather limited control over the process. Some works have explored the features learnt inside the hidden layers of deep learning models in very creative ways. An example can be found in the project known as ‘Deep Dream’ or ‘Inceptionism’ [337], where these features were used to reconstruct remarkably surreal images. Similarly, one would hope that by setting up a set of latent variables (dimensions of the latent space), the model will fit those features implicit in the data that are most prominent to each of the dimensions of the latent space. While this is partially true (and more so after the methods BetaVAE [298], FactorVAE [300] and InfoGAN [299] among others), it is also subject to the same caveats as in the case of deep learning. Indeed, these works do present an important degree of disentanglement that delivers very significant results. However, there is still little direct control over what features make it to each of the dimensions in the latent space. Additionally, the configuration of the network, that is, the size of the latent space, is not a flexible parameter capable of adapting to the input samples. Once a latent space size is set, the model will be trained (with massive amounts of data) with the given size from beginning to end. As a consequence, the amount of implicit disentangled features that can be potentially captured by the model is limited by the size of the latent space. 4.3. Literature review: (architecture-related) As mentioned in the previous sections, creating new instances based on the information of known references is a classical necessity within design disciplines. From an initial known design, it could be desirable to preserve some of its features by ‘importing’ these into a family of new designs, conceived as variations of the original. Generating this family of new designs is a task that can be addressed with either manual or computational methods. However, one could argue that both types of methods could be ‘algorithmically’ addressed: by identifying the design features that one may wish to preserve (so that they are found within the instances of the new family of designs), a set of rules can be specified so that the selected features are applied with different gradients, making each instance unique, yet 97
identifiable as part of a family of shapes design [338]. However utilitarian, tools of this nature may support the creative production within design disciplines by offering a dynamic way of obtaining variations of the known reference for further evaluation. In particular, and in contrast with a wide stream of ML research that focusses on 2D problems, the application of ML to the architectural arena demands a 3D approach. In this regard, machine learning techniques for 3D shapes are implemented in various different tasks of geometry manipulation for evaluation purposes. These tasks can be classified into the following categories: Single object classification [339,340], 3D pose estimation [341], multiple objects detection [342], scene-object semantic segmentation [343], 3D geometry synthesis-reconstruction [344] and many other categories that are serviceable to the technical aspects of working with 3D geometry. Initially, some machine learning research projects tackled the production of 3D objects through techniques associated with deep-generative models operating with 2D data sets. In this approach, a geometric 3D entity can be reconstructed from the information acquired from the 2D output of the model [345,346,347]. For example, one implementation of this technique in the field of architectural design, to generate new instances of 2D graphs representing architectural layouts of living units, finds inspiration in ‘composing high performing parts of separate design entries into a new whole’ [348]. This strategy takes advantage of the relative consolidation of well-tested models applied to 2D pixel data. While in many occasions they remain at the level of images (interior decoration styles or style transfer on building facades, etc.), in some others they leverage them in combination with other techniques so as to become functional in the 3D domain. However, there are also a good number of studies that address 3D problems directly, without intermediate 2D models. The following is a review of some of the most relevant works along this line in the context of the case study developed in the next section. In 3D studies of architectural designs and forms, representation is a key element. Pixel-based or voxel-based representations can have significant limitations in engineering and architectural design. Especially in architectural design, the large scale and complexity of artefacts can make voxel-based approaches prohibitively computationally expensive, even with advanced computing hardware. Therefore, research efforts have been dedicated to the search for suitable alternatives. In terms of data representation, some works have attacked this problem purely from a ML perspective, aside from any generative aspects. A sample of these works includes (i) MeshCNN [349], (ii) VoxNet: A 3D Convolutional Neural Network for Real-Time Object Recognition [339] and (iii) 98
Generative Deep Learning in Architectural Design (GDLAD) [350], all of them are interesting approaches that look at the problem from different angles. MeshCNN proposes a workflow in which 3D objects are represented as tetrahedron meshes, and the connectivity of the edges of each mesh informs the architecture of the model, accounting for the size of the convolutional layer. In this model, meshes are nonuniform representations of the shapes. Since these are essentially numbered point clouds, the objects result in irregular structures. In contrast, in the work of both Maturana (VoxNet) and Newton (GDLAD), the 3D objects present in the data set are denoted in a consistently sized 3D envelope, using voxels within this 3D space. The case study that will be presented later relies heavily on vectors denoting the connections between points of a wireframe structure in 3D space. Thus, it shares important strategic similarities with MeshCNN as opposed to the latter works (VoxNet and GDLAD). With regard to contributions purely in the field of Generative AI, the following works have been selected on the basis of their similarity in the objectives or strategy to the later case study: In ”3D Shape Synthesis for Conceptual Design and Optimization Using Variational Autoencoders” [351], a data-driven 3D shape design method is presented. It can learn a generative model from a corpus of existing designs and use this model to produce a wide range of new designs. The approach is based on an unsupervised VAE architecture to learn an encoding of the samples in the training corpus, without the need for an explicit parametric representation of the original designs. To facilitate the generation of smooth final surfaces, the method uses a 3D shape representation based on a distance transformation of the original 3D data, rather than using the commonly utilised binary voxel representation (as in VoxNet). The generator maps the latent space representations to the high-dimensional distance transformation fields, which are then automatically surfaced to produce 3D representations. The method is applied to the computational design of gliders that are subsequently optimised to achieve a certain physics-based performance. Another contribution ”Automated modular housing design using a module configuration algorithm and a coupled generative adversarial network (CoGAN)” [352], introduces an innovative approach to the design of modular housing. It uses an automated system that uses a module configuration algorithm and a coupled generative adversarial network (CoGAN). The module configuration algorithm automates the arrangement of different modules based on specific design criteria and constraints, ensuring optimal spatial configuration and functionality. Then, the CoGAN architecture consists of two GANs working together; one GAN generates 99
architectural designs, while the other evaluates and refines them to ensure quality and coherence. The proposed methodology aims to streamline and improve the modular home design process, making it more efficient and adaptable to various design requirements, showing that the system can produce diverse and innovative modular home designs that meet the specified requirements. The generated designs were evaluated for feasibility and aesthetic appeal, showing promising potential for real-world application. The work ”Diffusion Probabilistic Model Assisted 3D Form Finding and Design Latent Space Exploration” [353], employs diffusion models for the generation of novel designs. The method is applied to the spatial transformation of Taihu stones, a traditional Chinese garden element known for its intricate and natural forms. The methodology leverages the strengths of diffusion models to generate and explore complex 3D geometries, enabling innovative design possibilities. The authors use distributions in the latent space as a tool to explore the space of design possibilities, which is actually a common trait to all the works using models based on latent spaces. This use of the latent space allows designers to uncover a wide range of potential forms and transformations, fostering creativity, and pushing the boundaries and limitations of the designer’s imagination. The study shares many similarities with the case study in this chapter. In particular, its diffusion model aims to interpolate organic stone forms with orthogonal building-like shapes. In doing so, the authors expect the model to arrive at interesting blends of organic forms (resembling Taihu stones) that still possess some degree of orthogonality and are still feasible in terms of physical construction. An illustration of their results is shown in Fig. 4.6. A very interesting contribution ”VQ-CAD: Computer-Aided Design model generation with vector quantized diffusion” [354], focusses on learning implicit constraints in the domain of Computer Aided Design (CAD). Generative models are naturally equipped to perform continuous transitions across data. However, in some data domains, such as industrial or architectural design, not every shape or form can be considered a ‘valid’ design, since some designs might be impossible to fabricate or build. For example, a pin-hole that does not adhere to the standard size for market-available pins, or a column that is too thin to bear the load assigned to it. Thus, generative studies on CAD artefacts have often led to the explicit imposition of external constraints. An important aspect of their work is that it does not enforce any explicit constraints on the model, but instead the model learns them through examples and is hence incorporated into the model implicitly. The method is based on a vector quantisation latent diffusion architecture, with an implementation of hierarchical 100
Figure 4.6: Results of the Taihu stones project. code-book in the latent space. That is, the method uses a hierarchical discrete approach to the distribution of the latent space that allows operating and learning within a constrained domain. Additionally, the approach paves the way for text-based embeddings and prompt-based interactions with the generative engine, which they exploit in the second part of their paper through the CLIP model. In terms of data representation, the method engages with CAD primitives such as lines, arcs, extrusions, profiles, etc., that are hard-coded in the model with hierarchical dependencies (to form the code-book). In essence, VQ-CAD does not deal directly with raw data, but instead, as in the approach that will be presented here, there is a preliminary modelling of the spatial data that conforms the objects. Nevertheless, the method does present relevant advances with regard to the case study in this thesis, as it puts forth an interesting solution to the problem of design constraints. Finally, ”3D Diffusion or 3D Disfiguration?” [355] constitutes what can be considered a successful and industry-complete example of art generation within the Generative AI paradigm. The work uses diffusion models to create design blends across a set of object categories (traditional wooden chairs, urban benches, timber row-boats, etc.). The approach is based on a voxelised representation of the data and there are no design constraints imposed. After a visual exploration of the latent space, an intriguing set of chair-like and bench-like artefacts were selected by the artist. Subsequently, these artefacts were digitally fabricated and exhibited at the CVPR 2024 AI Art conference (Fig. 4.7). The different works discussed above highlight the relevance of generative 101
Figure 4.7: Chair artefacts exhibited at the CVPR 2024 AI Art conference. models in the field of design. They also show how the specific requirements of the field call for different data representation strategies. As pointed out before, while many of the most important milestones and achievements in ML and Generative AI have been showcased on pixel images, this representation medium may present important limitations in the domain of engineering and architectural design. Particularly for architectural design, the inherent large scale and complexity of the artefacts may render voxel-based approaches computationally expensive, even with the immense capacity of current computing hardware. Thus, other methods of representation need to be explored. In the following section, a case study is presented for the interpolation of building types. In this study, the buildings are expressed as mathematical graphs that carry the information on how the beams and columns are connected to each other. This approach differs from the strategies discussed above and may provide a valid alternative to the study of generative methods for architectural building structures. 4.4. Case study: VAE for the generation of hybrid building typologies 4.4.1 Introduction In design disciplines, the need for alternatives to a known design is a classic situation. It manifests itself as a search for variations that can resemble certain characteristics of the original known design, whilst being original in their own right. Deep generative models can help tackle such challenges by producing output samples that resemble features of the input sample. Being probabilistic 102
models, they allow for interpretable representations and measurable output predictions, in addition to the adaptability and learning scalability of deep neural networks. This research area is one of the most exciting and rapidly evolving fields of statistical machine learning [356]. The present case study focusses on the use of VAEs models [357], which are a special kind of autoencoder that enforce a continuous distribution in their latent space. As explained earlier, by sampling from a continuous latent space, these models generate new objects that inherit features from the samples present in the training set but are at the same time essentially unique. This study follows on previous work where the notion of a ‘connectivity vector’ was developed [358], which represented the geometric 3D object in a network fashion. This encoding resulted in a data set of high-dimensional vectors. Using the way data are encoded, geometric 3D objects can be expressed as tensor-shaped input data sets for training. Thus, the challenge becomes the balance of the large number of parameters in the model, especially when compared to the amount of samples in the training set. The work presented in this section handles input composed of high-dimensional vectors containing the data of geometrical objects. These objects will be called ’building types’, as they are simplified representations of the geometry of more sophisticated architectural building objects. These are inspired by the structural wireframes of these building geometries. By taking a simplified set of centre lines of the interconnected structural elements that compose an architectural object in 3D space, a wireframe is created. Thus, this wireframe is representative of the core geometry of the building type. Described by connectivity vectors, this geometrical object is used as input data for a VAE model, opening up a methodology within ML applied to 3D objects that is applicable to the field of architectural design. As discussed earlier, VAE models are shown to be able to generate many types of complex data. Although initially trained on sets of 2D images, the proliferation of these models in wider disciplines has driven the need to work with vector data and 3D geometry [359,360]. Various works within the design community that map the potential of this approach have already been conducted [361]. Powerful techniques inherent to deep-generative models such as sampling [362] and feature vector arithmetic [301] show great promise in the context of design. Deep-generative models typically require large data sets for training purposes. Working with representations of structural building types allows the implementation of parametric tools for data augmentation. This is because building types allow individual samples to have various discrete, yet observable, characteristics that enable them to be identified as part of a given family or type. In 103
the Experimental section, an account of the experiments conducted is presented. In these tests, a VAE model trains on an augmented dataset and learns to extract features that are characteristic of the input building types. In the last stage of the process, the generative capability of the network is used by sampling new points from the continuous latent distribution that the model has learnt. The decoder then outputs their corresponding connectivity maps that result in newly generated building types. This work is an attempt to improve the results obtained throughout the experiments conducted during the development of this method [358], leading to overfitting of the learning process of the model, such as the following: (i) due to the high dimensionality inherent to the technique used for encoding 3D objects, a large number of samples were required to train the model; (ii) limited geometrical variation of the types in training and validation sets, and (iii) models with densely connected layers resulted in a very high number of parameters (150 M+ trainable parameters). This case study explores the following solutions to the aforementioned problems: first, the development of a parametric data augmentation scheme that enhances geometrical variation. Second, the implementation of convolutional layers within the architecture of the model, helping to reduce the number of parameters of the model while maintaining a strong capacity to learn complex patterns. Lastly, a constrained variant of the parametric augmentation method that allows for the limitation of the feature spread of the geometries that compose the data set. Furthermore, this study seeks to serve as a starting point for exploring future avenues of generative design, especially when in search of alternatives to a known building configuration. The main contribution here is to provide new insights into how parametric augmentation techniques might improve VAE learning in the context of 3D building wireframes. Despite the fact that the current output is still very much a work in progress, the resulting 3D wireframes can be conceived as the first step towards the interpolation of geometries from a set of known input types. Thus, this methodology can serve for experimentation and can be further explored by disciplines of architectural design. This is only a snippet of the possibilities that this methodology can unlock. 4.4.2 Methodology This section is arranged into several parts: data representation, VAE, and network architecture. In the first part, a detailed description of the methods used for data representation is presented, followed by a description of the encoding method based on the concept of a 3D-canvas with voxelised wireframes. Then, a description of the 104
revised neural network architecture used is shown. The model learns a continuous latent distribution of the input data from which it is possible to sample and generate new geometry instances, essentially hybrids of the initial input geometries. Data representation: Connectivity map In order to train a VAE, it is crucial to prepare the 3D geometry in a way that can be parsed through the network. This means choosing the most compact way to represent the geometry, avoiding redundancies whilst retaining full information. The scheme presented here is based on a 3D-canvas, which consists of a rectangular 3D volume discretised in cube-shaped cells within which the input geometry is contained. Each cell of the 3D-canvas contains labelled connectivity vectors that can be activated or deactivated depending on the input geometry. These connectivity vectors represent wireframe segments in different orientations that are used to approximate the input geometry in 3D space. To keep the data representation compact and without overlaps, parallel connectivity vectors are discarded and only 13 vectors for each cell are considered, as shown in Figs. 4.8 and 4.9. To illustrate the scheme, simple geometry and connectivity vectors are depicted in Figs. 4.10,4.11, showing the same principle working in 3D space. input vector dim = height x width x 4 A, B, C, D 4 possible connections between nodes 2D canvas input vector dim = height x width x 4 A, B, C, D 4 possible connections between nodes 2D canvas B CA D Canopy with Canopy vector size for operations = cloudPt_h * cloudPt_w *4 2d Representation Figure 4.8: Diagram of the connectivity vector scheme for a 2D geometry. At the beginning of the routine, a 3D-canvas is generated ‘around’ the input geometry, so that the full extent of the input can be encoded within it. Wireframe line segments of the input geometry are then snapped to the grid defined by the cubeshaped cells of the 3D-canvas and corresponding connectivity vectors are mapped 105
Figure 4.16: Parametric generation of samples (CCTV). autoencoders, to slightly different variants of the same output. This generates a much smoother and interpolated latent space capable of producing new outputs that share common features from diverse inputs. As long as the VAE model does not provide a single encoding but a set of encodings that, with greater or lesser probability, could be the result of the encoder, the usual loss functions (representation losses) are not adequate to measure the error that the network presents during training. In order to solve this problem, from a theoretical point of view, a new factor is introduced in the loss function, called KL-divergence (Kullback–Leibler Divergence) [266], which, instead of measuring the distance between points, measures the difference between two probability distributions. 112
z = μ + εσ ε Sample from N(0,1) DKL[ N(μ,σ) | N(0,1) ] μ σ + = Loss Reconstruction Loss Input layer Training set 1 Training set 2 Hidden layer Hidden layer Hidden layer Latent layer Output layer Hidden layer Hidden layer Hidden layer Encoder Decoder Reconstructed set 1 Reconstructed set 2 ReLU activation sigmoid activation Figure 4.17: Generic diagram of the network architecture of a standard VAE. As usual in machine learning, some conditions can be imposed on the network so that it becomes able to learn a convenient distribution (usually a Gaussian distribution, as it offers simplicity in the implementation). Neural network architecture Neural network models present an extremely high variety of possible configurations. To date, there is no straightforward methodology to determine the optimal configuration of a model for a given data set. Neural network models are very sensitive to input data; what may work for a certain problem is likely to perform poorly for a different dataset. For this reason, the optimal setup of the model has to be found through heuristics, with the added difficulty that the search space is overwhelming. To complicate things further, learning problems may require the use of huge data sets, which make training processes a rather slow endeavour and, thus, hinder the possibility of massive testing through the search space. Some of the most relevant approaches to this problem have been the implementation of heuristics through genetic algorithms. Genetic algorithms allow the space of hyperparameters to be searched in a systematic way and have been shown to be very effective in determining efficient configurations for neural network models [365], especially when compared to more traditional grid-search approaches [366]. However, in the present work there is a large overhead in terms of computation due to the magnitude of the problem that is being dealt with. Working with 3D data sets implies that the networks have to learn features from very large samples, and therefore the models must be configured in a way that allows for accommodating a high level of complexity (the network must be ready 113
to approximate very complex functions). The combination of a powerful network architecture and a large data set (+50 k samples, 1 MB per sample) translates into highly time-consuming training for small tests. This circumstance entails a strong limitation in the number of experiments that may be carried out in a reasonable amount of time, despite employing one of the best GPU available in the market. Given the context, this early study does not focus on exhaustively searching for an optimal network configuration, but rather it simply attempts to find a network architecture that is capable enough of clearly separating geometric types in the latent space of the VAE. Finally, before moving on to the next section, Fig. 4.18 presents a complete diagram of the proposed methodology. z = μ + εσ ε Sample from N(0,1) DKL[ N(μ,σ) | N(0,1) ] μ σ + = Loss Reconstruction Loss Input layer Training set 1 Training set 2 Hidden layer Hidden layer Hidden layer Latent layer Output layer Hidden layer Hidden layer Hidden layer Encoder Decoder Reconstructed set 1 Reconstructed set 2 ReLU activation sigmoid activation Geometry Free Limited exact size of dataset Training Dense scheme Training Conv3D scheme Sampling from Latent Space New Geometries Figure 4.18: Workflow diagram of the proposed methodology. 4.5. Experimentation and discussion of results Throughout all experiments conducted in this study, the source data set is initially composed of 30K variations of a CCTV-inspired building and another 30K variations of a Hejduk-inspired structure. In the last experiments, this number was increased to 75K variations of each type. In the training examples generated, the 3D-canvas on which these samples are inscribed is set consistently as a grid of 21×21×21 units. This resolution is sufficient to distinguish structural types like 114
arches, wall and surface elements, volumes with cavities, openings, etc., allowing for the design and implementation of instances of neural networks which are trained for recognition and handling of data pertaining to the field of architectural geometry. In all cases, the connectivity vector for each point in the grid takes 13 values. Therefore, the number of input neurons for the VAE is 21 ×21 ×21 ×13 =120, 393. As the number of training parameters in a neural network grows with the number of input neurons, this size of the 3D-canvas is close to the limit of what can be reasonably dealt with in terms of available computational power. The latent space of the VAE is always fixed to 2 dimensions (2 neurons in the bottleneck of the autoencoder). In VAEs, this low value is justified by the fact that high dimensionality in the latent space has been shown to deliver poor results [357] and has been criticised for the effects of the ’soap bubble’ in the output [367]. Encoders and decoders are defined as symmetric as possible. Regularisation techniques such as dropout layers have not been used because they may compromise the definition of the reconstructed geometries. The reconstruction error metric used is Binary Cross Entropy, as is commonly recommended in VAE models [368]. However, the performance of the models will also be evaluated qualitatively from a design perspective, since the objective of this research is to provide a generative tool for design. In particular, the assessments will consider the uniqueness and hybridisation of features present in the generated geometries. The first pair of tests aims to compare the performance of (i) the data set generated by parametric augmentation, as explained above, and (ii) the previous augmentation strategy implemented in [358], which is based on a combination of geometry displacements and random noise in the values of the connectivity vector. For this purpose, both data sets are trained during 25 epochs under the same feed-forward network with only one hidden layer, for the sake of simplicity. The encoder configuration has an input layer of 21×21×21×13 neurons and 1 hidden layer of 512 neurons as seen in Table 4.1. The decoder is exactly symmetrical, and the latent space is two-dimensional. Input layer Hidden layer (H1) Latent space Hidden layer (H1’) Output layer Type - Dense Dense Dense - Size 120,393 (21x21x21x13) 512 2 512 120,393 Activation relu relu relu relu sigmoid Table 4.1: Network architecture of preliminary experiments. In the case of the displacement and noise dataset, training yields much lower 115
error rates than parametric augmentation (Tables 4.2,4.3 and Fig. 4.19). This is understandable, as the spectrum of geometries enabled by the new method is much more diverse and, thus, requires stronger learning capabilities from the VAE model. However, the latent space resulting from the latter shows that the network is already able to differentiate the two types of building to some extent (Fig. 4.20), while the latent space corresponding to the former is far from capable of rendering this distinction (Fig. 4.22). Furthermore, distinct structural features can be observed in the latent space corresponding to the parametric augmentation data set (Fig. 4.21), which are not present in any form in the other latent space. Finally, in Fig. 4.20, VAE model was able to spread the samples taking a larger portion of the latent space than in Fig. 4.22, where most samples are concentrated between the values (−2.5, 2.5)on the vertical axis. Although these are purely visual observations, it is advisable to approach this analysis from a more rigorous methodology in future work, especially in cases where the distinction is less obvious. Batch size Optimizer Learning rate Validation loss Validation loss (A) parametric augmentation (B) random noise + displacements 128 RMSProp 0.0005 1,261.57 401.35 ‘’ ‘’ 0.0010 1,246.69 400.72 ‘’ ‘’ 0.0015 1,279.31 399.61 ‘’ ‘’ 0.0020 1,311.84 405.06 Table 4.2: Hyper-parameters and results of preliminary experiments. Training loss Validation loss Parametric augmentation 1242.39 1247.02 Noise + displacements 467.56 471.63 Table 4.3: Training results for preliminary experiments. A plausible interpretation of this apparent contradiction would point to the possibility that the variety of the information contained in the data set generated through random discrete displacements (which are very limited in comparison with the volume of the data set) and random noise in the values of the connectivity vector is not rich enough to enable the learning of relevant features. If the network learns features that are not representative of the input, it may not be able to separate them in a latent space even if it succeeds in reconstructing them. Accordingly, it is concluded that the dataset produced by parametric augmentation offers better prospects for further training. However, it must be noted that there is a distinctive feature that clearly differentiates the two types: CCTV does not have diagonal elements. This may give the VAE a head start in the training process in terms of separating both classes in the latent space. In order to test for generality, 116
0 5 10 15 20 epochs 0 500 1000 1500 2000 2500 loss v_loss (parametric augmentation) t_loss v_loss (noise + displacements) t_loss Figure 4.19: Best training results for both the parametric augmentation data set and the previous random displacements + noise data set. Figure 4.20: Latent space. Encoded samples from the parametric augmentation data set (yellow and purple dots represent each of the two training categories). 117
Figure 4.21: Highlight of observed structural features in the latent space from the parametric augmentation data set. more types should be scrutinised. However, this early work attempts only to carry out an initial exploration on the potential of the methodology presented here. Upon proof of the potential benefits of the aforementioned parametric augmentation strategy, the second set of experiments attempts to establish whether a network architecture based on 3D convolutional hidden layers would outperform a model containing only linear layers that are densely connected. Convolutional layers can help keep the number of trainable parameters in check, as will be explained in the discussion of results. This is important because if the ratio to the number of training samples is overweighed by the parameters, then there is a very high risk of overfitting. Deep architectures can be extremely powerful; however, there is a trade-off between the magnitude of the problems that a neural network can solve and the overfitting of the network to the data, as shown in Fig. 4.23. For this set of experiments, two groups of network architectures have been selected after further preliminary testing. The first group (A-Conv) includes three convolutional schemes, and the second (B-Dense), three standard feed-forward configurations with varying numbers and sizes of hidden layers. The first model (C21-C7-D512) of the A-Conv group features two 3D convolutional layers in the encoder, leading to a linear dense layer of 512 neurons and a symmetrical decoder. 118
Figure 4.22: Latent space. Encoded samples from the random displacements + noise data set. Figure 4.23: Model complexity versus training and validation errors (overfitting problem). The latent space is again 2D (and will remain the same in all the experiments). The first of the layers is a 3D convolutional layer spatially arranged as 21×21×21 neurons with a depth of 13 channels. Convolutions are applied in 3×3×3 units with maximum overlap. The size of the second one is 7×7×7, and the rest of the configuration remains identical. The second model (2xC21-2xC7-3xD512) duplicates both convolutional layers and adds two more dense layers of 512 neurons at the end and beginning of the encoder and decoder, respectively. Repetition of convolutional layers has delivered good results when applied to deep neural 119
networks [369]. Finally, the third model (3×C21-3×C7-4×D512) adds another convolutional layer between each of the two pairs of convolutional layers of the previous model. Both of these new layers feature an increased depth of 39 channels. Also, at the end of the encoder, an additional dense layer is allocated maintaining the same configuration of the three preceding layers. The decoder, as always, is the exact mirror. The three architectures described above contain 982K, 5.78M, and 6.725M trainable parameters, respectively. For each, different learning rates have been tested, and the final configuration is shown in Tables 4.4–4.6. The training results for the A-Conv set are shown in Fig. 4.24. C21-C7-D512 Input H1 H2 H3 Latent Type - Conv 3D Conv 3D Dense Dense Size 120,393 21x21x21 7x7x7 512 2 Convolution filter - (3x3x3)x1 (3x3x3)x1 - - Optimizer RMSProp, Lr (learning rate) = 0.0017, Stride = 1 All activations are relu except output layer (sigmoid) Total parameters: 982k Table 4.4: C21-C7-D512 network architecture (showing only encoder for simplicity). 2xC21-2xC7-3xD512 Input H1-H2 H3-H4 H5-H7 Latent Type - Conv 3D Conv 3D Dense Dense Size 120,393 21x21x21 7x7x7 512 2 Convolution filter - (3x3x3)x1 (3x3x3)x1 - - Optimizer RMSProp, Lr = 0.0011, Stride = 1 All activations are relu except output layer (sigmoid) Total parameters: 5.78M Table 4.5: 2xC21-2xC7-3xD512 network architecture (showing only encoder for simplicity). 3xC21-3xC7-4xD512 Input H1-H2 H3 H4-H5 H6 H7-H10 Latent Type - Conv 3D Conv 3D Conv 3D Conv 3D Dense Dense Size 120,393 21x21x21 21x21x21 7x7x7 7x7x7 512 2 Convolution filter - (3x3x3)x1 (3x3x3)x3 (3x3x3)x1 (3x3x3)x3 - - Optimizer RMSProp, Lr = 0.0009, Stride = 1 All activations are relu except output layer (sigmoid) Total parameters: 6.72M Table 4.6: 3xC21-3xC7-4xD512 network architecture (showing only encoder for simplicity). In the B-Dense group (Tables 4.7–4.9), the first model (D2048-D512) bears two hidden layers that are densely connected for both the encoder and the decoder. The first hidden layer holds 2048 neurons and the second one 512, which is 495.35M 120
0 10 20 30 40 50 60 70 80 epochs 0 500 1000 1500 2000 2500 3000 loss v_loss (C21-C7-D512) t_loss v_loss (2xC21-2xC7-3xD512) t_loss v_loss (3xC21-3xC7-4xD512) t_loss Figure 4.24: Best training results of A-Conv scheme. trainable parameters. In the second model (4xD512), these two hidden layers are replaced by four identical hidden layers of 512 neurons, thus reducing the parameter count to 125.245M. However, this value may still be quite high for the dataset at hand. Finally, a third model (6xD112) is set up that contains up to six hidden layers of 112 neurons each, both on the encoder and decoder. However, this last model cuts the total parameter count down to 27.215M, which is still a remarkable figure. The training results for the B-Dense scheme are shown in Fig. 4.25. D2048-D512 Input H1 H2 Latent Type - Dense Dense Dense Size 120,393 2048 512 2 Optimizer RMSProp, Lr = 0.0012 All activations are relu except output layer (sigmoid) Total parameters: 495.35M Table 4.7: D2048-D512 network architecture (showing only encoder for simplicity). The best performing architectures of each scheme are 2xC21-2xC7-3xD512 and 6xD112. The first one features a total of four convolutional hidden layers and three dense hidden layers for both the encoder and the decoder. The second model is composed of six dense layers in between the input and the latent space and the same layers again in between the latent space and the output layer of the autoencoder. The performance of both models in terms of validation loss is relatively similar, as can be seen in Table 4.10 and Figs. 4.26,4.27. However, the 121
Figure 4.31: Geometry reconstructed from latent space (blow-up). Model 6xD112. particular, the parameters that have been most limited are width and length. Previously, these two dimensions were free to take up the complete size of the canvas, whereas in the restricted dataset, they do not exceed half the length or width of the canvas. These two strategies have been tested through two experiments. In the first one, the data set is increased and nothing else is altered, and in the second one, the data set is increased, and a limit to the variation of the parametric augmentation method has been applied as explained above. The results of these experiments are shown in Table 4.11 and Figs. 4.32–4.36. The data show a large reduction in validation loss (50%) and a latent space with a very clear pattern segregation. It can also be observed that the interpolation of geometry takes place along a very narrow passage in between the two clusters. A detailed visualisation of the output geometry for the second case (modified and increased dataset) is shown in Figs. 4.37,4.38. 128
2xC21-2xC7-3x512 Training loss Validation loss Reference dataset 1029.86 1036.01 Dataset increased 1335.11 1357.23 Dataset increased and modified 464.34 506.75 Table 4.11: Training results for 2 × C21-2 × C7-3 × 512 upon variations of the data set. Figure 4.32: Latent space of 2xC21-2xC7-3xD512 model with the ‘increased data set’. These last two experiments were an attempt to reduce the reconstruction error by providing a larger data set and limiting the range of variations of the samples that are generated through parametric augmentation. In Fig. 4.28, it can be observed that a large part of the transition gradient deals mainly with accommodating variations in size, rather than picking up morphological or topological features, which are central to this study. Thus, the motivation behind reducing parameter variation ranges, such as those applied to the width and length of the samples, is to improve the accuracy of the network by eliminating unnecessary information that leads to no relevant learning. In the first of these two experiments, the parameters for the generation of samples through parametric augmentation remained unchanged. However, the size of the dataset was more than doubled. In Fig. 4.32, the latent space corresponding to this test reveals that the model was totally incapable of 129
Figure 4.33: Latent space of 2xC21-2xC7-3xD512 model with the ‘modified and increased data set’. 0 10 20 30 40 50 epochs 0 500 1000 1500 2000 2500 3000 loss v_loss (best result from previous experiment) t_loss v_loss (dataset increased) t_loss v_loss (dataset increased and modified) t_loss Figure 4.34: Comparison of training results of 2xC21-2xC7-3xD512 model from the previous experiment with both the ‘increased data set’ and the ‘increased and modified data set’. separating the two types. The experiment thus suggests that increasing the number of training samples does not provide the network with better learning prospects. 130
Figure 4.35: Geometry reconstructed from latent space. Model 2xC-2xC-3xD512 using the increased and modified data set. This finding is rather interesting because the initial objective of using parametric augmentation was precisely to afford the possibility of generating large datasets as required by the complexity of the problem at hand. However, this result shows exactly the opposite: When the variations achieved through parametric augmentation are too broad, the shear scope of these may easily exceed the hypothetical benefit of providing more samples. In the second experiment, the parameters that guide these variations were heavily restricted in order to produce less heterogeneity while still generating the desired number of unique samples. The results (Figs. 4.35,4.36) prove that the strategy is effective in reducing the validation loss approximately around 50%, which is very successful. In addition, categories are clearly separated after encoding. However, the resultant latent space (Fig. 4.33) shows a hard fault line between the two training categories. This 131
Figure 4.36: Geometry reconstructed from latent space (blow-up). Model 2xC-2xC-3xD512 using the increased and modified data set. Figure 4.37: Rendered geometry reconstructed from latent space. Model 2xC-2xC-3xD512 using the increased and modified data set. 132
Figure 4.38: 3D printed geometry reconstructed from latent space. Model 2xC-2xC-3xD512 using the increased and modified data set. condition entails that interpolation of geometry may only happen in a very narrow area and that transitions will not be able to display smooth gradients. In Fig. 4.35, this effect can be clearly observed in the distribution of the generated geometries, and the region blow-up effectively reveals abrupt transitions along the fault line. It is unclear at this stage the reason behind this fallback into unsatisfactory interpolation spaces, especially despite the positive results in terms of both validation loss and spatial segregation in the latent space. Possible explanations may be connected to the particular geometries selected for the experiments, in the sense that they may not share common distinctive features that allow clean interpolations. Or perhaps, the network architectures implemented in this work lack the ability to capture global features from one or both input types. Finally, there is another possible factor that should be mentioned, one that touches upon the core of the methodology presented here and that should be taken into consideration for future work. ML algorithms are essentially optimisation models. Most of them function on the basis of the gradient descent, whereby a relevant local minimum of the loss function may be found. However, loss functions must be continuous and tractable and should not present frequent or large areas of null derivative (plateaus). If these areas are prevalent in the function, then the algorithm may not know in which direction to move when pursuing lower loss values. It may assume that it has already touched bottom, or it may trigger a 133
random decision regarding which direction to take, resulting in high volatility during the training process (as can be seen in Figs. 4.25,4.34) and low convergence. In fact, models that did not easily converge have been prominently present throughout much of the side work carried out during the present study. The alternative representation of geometry presented here is based on a connectivity vector of discrete values (0, 1). Furthermore, in this last experiment, variations among the training samples were minimised to avoid an excessive spread of features. This means that many samples shared large sets of identical values, facilitating the emergence of flat areas in the loss function. And where those values were different, a transition pattern was hard to find due to the discrete nature of the representation method chosen to approach the problem. It may thus be valuable when engaging in further work to look into some valuable research that is currently taking place to tackle deep learning problems in discrete spaces. Works along the lines of [370,371] attempt to find alternative representations of these spaces that smooth out the issues mentioned above. The adoption of the methods proposed in these studies may provide the answers that are required to improve the results presented in this work. 4.5.1 Conclusions Due to the growing international academic interest in generative machine learning methods and its wide commercial applications, the workflow presented in this case study builds on the emergence of a blooming field. Although developed almost ten years ago, the use and usefulness of VAEs in the context of architectural design remains mostly unexplored. The work presented here builds on the notion of a connectivity vector that is used to represent 3D mesh-like geometries, with the objective of facilitating their processing by neural networks in general and VAEs in particular. This representation was explored in [358], where a data set comprising two building types was generated through noiseand displacement-based enhancement. The results suggested on the one hand that noise did not help the network in identifying patterns. And, on the other hand, that the displacement approach was always prone to overfitting, because the larger the displacement space, the more the number of trainable parameters grew. In this context, the main objective of the present work has been to explore the suitability of an alternative augmentation method to assist in the generation of novel geometries with a VAE. This alternative method, parametric augmentation, has allowed very large data sets to be created without the need to increase the size of the 3D-canvas (and consequently the size of the input layer and total number of trainable parameters), 134
as was the case in the previous approach. Additionally, parametric augmentation is particularly efficient for increasing data sets of 3D geometries within the field of architectural geometry – especially when resembling building types, due to the discrete yet observable characteristics of each sample, as constituent of each type. The results presented show that the method was indeed successful in preventing the original overfitting problem. Consequently, the autoencoder was successful in reconstructing the geometries of the data set. However, the augmentation method has created another set of problems that have challenged the performance of the VAE. Firstly, results show that the feature spread produced through parametric augmentation can overwhelm the network and can hinder its ability to extract those features. This was made apparent when an increase in the size of the dataset worsened the reconstruction error rather than improving it. This limitation blocked the possibility of expanding the dataset to improve the definition of the geometries that are generated by sampling from the latent space. Thus, although there were smooth transitions across the two types present in the resulting geometries, these remained quite blurry. Secondly, when attempting to counter the latter issue by limiting the range of the parameters involved in building up the training set, it was found that despite being effective in further lowering the reconstruction error, transitions across types had been drastically reduced. An important takeaway from the experimentation is that there seems to be a more fundamental problem underlying the difficulties faced by the VAE. As discussed in the previous section, the representation based on a connectivity vector as implemented in this study creates a tough landscape to navigate with optimisation algorithms. This is mainly due to the discrete and sparse nature of the data set. Potential avenues of research that may shed light onto this problem touch upon transformation techniques from discrete spaces into continuous ones. Some of these have been earmarked for future work, as they are currently being pursued by several authors. Despite the difficulties, the final results show partial success in generating new geometries from the latent space that share a mix of features from the two types present in the training set, which was the initial objective of the research. The quality of these new samples achieves a certain degree of interpolation as can be observed in Figs. 4.37,4.38, although the definition of the resulting geometries may still be improved. Aside from the definition of the results, the main contribution of this work has been to explore an alternative avenue for the parametric generation of large datasets of 3D geometries, showcasing its problems and limitations in the context of neural networks and VAEs and pointing out potential solutions for future work. 135
In general, the proposed workflow challenges designers to acquire a critical perspective of the impact and potential of AI in our society and design practices. Generative neural network models have a large potential to redefine how architects and designers work with architectural precedents, that is, to use them directly as data for design generation. This case study aims to show the potential of such techniques and open a discussion about the future of machine learning in the context of geometry generation for the architectural design industry. 136
5. Concept representation and urban space: the case of real estate markets The previous chapter has explored the latest generative methods within the connectionist paradigm. The application of generative models to the field of Design has highlighted the importance of a suitable representation strategy for three-dimensional geometric data. Specifically, it has been shown that architectural artifacts can scale up the size of the neural input layers very quickly; to the point where even today’s most powerful hardware is rendered incapable of dealing with such amounts of training parameters. Therefore, alternative data representation strategies are necessary and are currently being explored. Additionally, neural network models have certain limitations when dealing with heterogeneous data. The input layers of these models are not adaptable during training, meaning that they must handle inputs of the exact same size and structure throughout. This is inefficient, as seen in the previous case study, because the network is forced to account for an input layer size that fits all the examples in the training data. Thus, it is relevant to explore other methods that may offer interesting alternatives to handling large-scale 3D data. This is especially pertinent in the case of even larger-scale studies, i.e. urban analysis. As in the previous Urban Intermezzo, this chapter also introduces and exemplifies through an urban analysis problem (utilitas), a technique that will be presented in more detail later: Formal Concept Analysis (FCA). The theory and tools provided by FCA allow for a different approach to the representation of data that operates on a higher symbolic level. Thus, a good part of the heavy lifting required by neural models that deal with raw data is alleviated. Of course, this comes with its own limitations, which will be discussed in the next chapter. Finally, this case study using FCA allows opening the perspective and contrasting the conversation on conceptual knowledge with the recent incursion in the connectionist domain. 5.1. Introduction Today, urban ecosystems produce data that enable one to manage, compare, and share huge amounts of information about the city. According to its nature, urban data can be grouped into different sets of information layers, which, as other information repositories and systems, are prone to problematic issues such as noise, inconsistencies, and ambiguity. 137
[282] Susan L Epstein. Learning and discovery: one system’s search for mathematical knowledge. Computational Intelligence, 4(1):42–53, 1988. [283] Masoud Yazdani. A computational model of creativity. In Machine learning: principles and techniques, pages 171–183. 1988. [284] Mark Keane. Analogical mechanisms. Artificial Intelligence Review, 2(4):229– 251, 1988. [285] Douglas R Hofstadter. On the seeming paradox of mechanizing creativity. Metamagical themas, pages 526–546, 1985. [286] Marvin Minsky. K-lines: A theory of memory. Cognitive science, 4(2):117–133, 1980. [287] Jonathan Edward Rowe. Emergent creativity: a computational study. 1991. [288] Marvin Minsky. Society of mind. Simon and Schuster, 1988. [289] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012. [290] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David WardeFarley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. In Proceedings of the International Conference on Neural Information Processing Systems (NIPS), pages 2672–2680, 2014. [291] Alec Radford, Luke Metz, and Soumith Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks, 2016. [292] Yunchen Pu, Zhe Gan, Ricardo Henao, Xin Yuan, Chunyuan Li, Andrew Stevens, and Lawrence Carin. Variational autoencoder for deep learning of images, labels and captions, 2016. [293] Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale gan training for high fidelity natural image synthesis, 2019. [294] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017. [295] Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks, 2017. 240
[296] Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks, 2019. [297] Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8):1798–1828, 2013. doi: 10.1109/TPAMI.2013.50. [298] Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. beta-VAE: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations, 2017. URL https://open review.net/forum?id=Sy2fzU9gl. [299] Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel. Infogan: Interpretable representation learning by information maximizing generative adversarial nets. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016. URL https://proceeding s.neurips.cc/paper files/paper/2016/file/7c9d0b1f96aebd7b5eca8c3edaa 19ebb-Paper.pdf. [300] Hyunjik Kim and Andriy Mnih. Disentangling by factorising, 2019. [301] A. Radford, L. Metz, and S. Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint. arXiv:1511.06434, 2016. [302] Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In C.J. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013. URL https://proceedings.neurips.cc/paper files/p aper/2013/file/9aa42b31882ec039965f3c4923ce901b-Paper.pdf. [303] Xin Wang, Hong Chen, Si’ao Tang, Zihao Wu, and Wenwu Zhu. Disentangled representation learning, 2024. [304] Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu. Neural discrete representation learning. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper files/paper/2017/file/7a98af17e 63a0ac09ce2e96d03992fbc-Paper.pdf. 241
[305] R. Gray. Vector quantization. IEEE ASSP Magazine, 1(2):4–29, 1984. doi: 10.1 109/MASSP.1984.1162229. [306] Yuri Kinoshita, Kenta Oono, Kenji Fukumizu, Yuichi Yoshida, and Shin ichi Maeda. Controlling posterior collapse by an inverse lipschitz constraint on the decoder network, 2024. [307] Yewen Li, Chaojie Wang, Zhibin Duan, Dongsheng Wang, Bo Chen, Bo An, and Mingyuan Zhou. Alleviating” posterior collapse”in deep topic models via policy gradient. Advances in Neural Information Processing Systems, 35:22562– 22575, 2022. [308] Zihao Wang and Liu Ziyin. Posterior collapse of a linear latent variable model. Advances in Neural Information Processing Systems, 35:37537–37548, 2022. [309] Ali Razavi, Aaron van den Oord, and Oriol Vinyals. Generating diverse highfidelity images with vq-vae-2, 2019. [310] Jialun Peng, Dong Liu, Songcen Xu, and Houqiang Li. Generating diverse structure for image inpainting with hierarchical vq-vae. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10775– 10784, 2021. [311] Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8821–8831. PMLR, 18–24 Jul 2021. URL https://proceeding s.mlr.press/v139/ramesh21a.html. [312] Geoffrey E. Hinton, Simon Osindero, and Yee-Whye Teh. A fast learning algorithm for deep belief nets. Neural Computation, 18(7):1527–1554, 2006. doi: 10.1162/neco.2006.18.7.1527. [313] Yoshua Bengio. Learning deep architectures for ai. Found. Trends Mach. Learn., 2(1):1–127, jan 2009. ISSN 1935-8237. doi: 10.1561/2200000006. URL https: //doi.org/10.1561/2200000006. [314] Hugo Larochelle and Iain Murray. The neural autoregressive distribution estimator. In Geoffrey Gordon, David Dunson, and Miroslav Dud´ ık, editors, Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, volume 15 of Proceedings of Machine Learning Research, pages 29–37, Fort Lauderdale, FL, USA, 11–13 Apr 2011. PMLR. URL https://proceeding s.mlr.press/v15/larochelle11a.html. 242
[315] A¨ aron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. Pixel recurrent neural networks. In Maria Florina Balcan and Kilian Q. Weinberger, editors, Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 1747–1756, New York, New York, USA, 20–22 Jun 2016. PMLR. URL https://proceedings.ml r.press/v48/oord16.html. [316] Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu. Conditional image generation with pixelcnn decoders, 2016. [317] Danilo Rezende and Shakir Mohamed. Variational inference with normalizing flows. In Francis Bach and David Blei, editors, Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1530–1538, Lille, France, 07–09 Jul 2015. PMLR. URL https://proceedings.mlr.press/v37/rezende15.html. [318] Yann Lecun, Sumit Chopra, Raia Hadsell, Marc Aurelio Ranzato, and Fu Jie Huang. A tutorial on energy-based learning. MIT Press, 2006. [319] Yilun Du and Igor Mordatch. Implicit generation and generalization in energy-based models, 2020. [320] Max Welling and Yee W Teh. Bayesian learning via stochastic gradient langevin dynamics. In Proceedings of the 28th international conference on machine learning (ICML-11), pages 681–688. Citeseer, 2011. [321] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. [322] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨ orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, June 2022. [323] Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. [324] Arash Vahdat, Karsten Kreis, and Jan Kautz. Score-based generative modeling in latent space. Advances in neural information processing systems, 34:11287– 11302, 2021. 243
[325] Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. [326] Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International conference on machine learning, pages 8162–8171. PMLR, 2021. [327] Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics (ToG), 36(4):1–13, 2017. [328] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: representing scenes as neural radiance fields for view synthesis. Commun. ACM, 65(1):99–106, dec 2021. ISSN 00010782. doi: 10.1145/3503250. URL https://doi.org/10.1145/3503250. [329] Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 7462–7473. Curran Associates, Inc., 2020. URL https://proceedings.neurip s.cc/paper files/paper/2020/file/53c04118df112c13a8c34b38343b9c10-P aper.pdf. [330] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/radford21a.html. [331] Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In Hal Daum´ e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 1691–1703. PMLR, 13–18 Jul 2020. URL https://proceedings.mlr.pr ess/v119/chen20s.html. [332] Stafford Beer. Designing Freedom. Wiley, 1974. [333] David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. Learning representations by back-propagating errors. nature, 323(6088):533–536, 1986. 244
[334] K Simonyan, A Vedaldi, and A Zisserman. Deep inside convolutional networks: visualising image classification models and saliency maps. In Proceedings of the International Conference on Learning Representations (ICLR). ICLR, 2014. [335] Anh Nguyen, Jason Yosinski, and Jeff Clune. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 427–436, 2015. [336] Xingchao Peng, Baochen Sun, Karim Ali, and Kate Saenko. What do deep cnns learn about objects? arXiv preprint arXiv:1504.02485, 2015. [337] Alexander Mordvintsev, Christopher Olah, and Mike Tyka. Inceptionism: Going deeper into neural networks, 2015. URL https://research.googl eblog.com/2015/06/inceptionism-going-deeper-into-neural.html. [338] J. Gips and G. Stiny. Shape grammars and the generative specification of painting and sculpture. In Proceedings of IFIP Congress 1971. North Holland Publishing Co., 1972. [339] D. Maturana and S. Scherer. Voxnet: A 3d convolutional neural network for real-time object recognition. In 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2015. Electronic ISBN: 978-1-4799-99941. [340] J. Wu, C. Zhang, T. Xue, W. T. Freeman, and J. B. Tenenbaum. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. arXiv preprint. arXiv:1610.07584, 2016. [341] J. Wu, T. Xue, J. J. Lim, Y. Tian, J. B. Tenenbaum, A. Torralba, and W. T. Freeman. Single image 3d interpreter network. arXiv preprint. arXiv:1604.08685v2, 2016. [342] S. Song and J. Xiao. Deep sliding shapes for amodal 3d object detection in rgbd images. In Proceedings of 29th IEEE Conference on Computer Vision and Pattern Recognition, 2016. [343] E. Kalogerakis, M. Averkiou, S. Maji, and S. Chaudhuri. 3d shape segmentation with projective convolutional networks. In Proceedings of the IEEE Computer Vision and Pattern Recognition (CVPR), 2017. oral presentation. [344] G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. Osman, D. Tzionas, and M. J. Black. Expressive body capture: 3d hands, face, and body from a single image. arXiv:1904.05866, 2019. 245
[345] T. Kelly, P. Guerrero, A. Steed, P. Wonka, and N. J. Mitra. Frankengan: Guided detail synthesis for building mass-models using style-synchonized gans. arXiv preprint. arXiv:1806.07179, 2018. [346] T. Wang, D. Ceylan, J. Popovic, and N. J. Mitra. Learning a shared shape space for multimodal garment design. In SIGRAPH ASIA 2018, volume 1806.11335, 2018. [347] S. Hoyer, J. Sohl-Dickstein, and S. Greydanus. Neural reparameterization improves structural optimization. arXiv preprint. arXiv:1909.04240, 2019. [348] I. As, S. Pal, and P. Basu. Artificial intelligence in architecture: Generating conceptual design via deep learning. International Journal of Architectural Computing, 16(4):306–327, 2018. [349] R. Hanocka, A. Hertz, N. Fish, R. Giryes, S. Feishman, and D. Cohen-Or. Meshcnn: A network with an edge. arXiv preprint. arXiv:1809.05910, 2019. [350] D. Newton. Generative deep learning in architectural design. Technology— Architecture+Design, 3(2):176–189, 2019. [351] Wentai Zhang, Zhangsihao Yang, Haoliang Jiang, Suyash Nigam, Soji Yamakawa, Tomotake Furuhata, Kenji Shimada, and Levent Burak Kara. 3d shape synthesis for conceptual design and optimization using variational autoencoders. In International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, volume 59186, page V02AT03A017. American Society of Mechanical Engineers, 2019. [352] Pedram Ghannad and Yong-Cheol Lee. Automated modular housing design using a module configuration algorithm and a coupled generative adversarial network (cogan). Automation in Construction, 139:104234, 2022. ISSN 09265805. doi: https://doi.org/10.1016/j.autcon.2022.104234. URL https: //www.sciencedirect.com/science/article/pii/S0926580522001078. [353] Yubo Liu, Han Li, Qiaoming Deng, and Kai Hu. Diffusion probabilistic model assisted 3d form finding and design latent space exploration: A case study for taihu stone spacial transformation. Computational Design and Robotic Fabrication, Part F2072:11 – 23, 2024. doi: 10.1007/978-981-99-8405-3 2. Cited by: 0; All Open Access, Hybrid Gold Open Access. [354] Hanxiao Wang, Mingyang Zhao, Yiqun Wang, Weize Quan, and Dong-Ming Yan. Vq-cad: Computer-aided design model generation with vector quantized diffusion. Computer Aided Geometric Design, 111, 2024. doi: 10.1016/j.cagd.202 4.102327. Cited by: 0. 246
[355] Immanuel Koh. 3d diffusion or 3d disfiguration?, 2024. Cited by: 0. [356] J. P. Cunningham. Columbia university, deep generative models. http://st at.columbia.edu/∼cunningham/teaching/GR8201, 2019. [357] D. P. Kingma and M. Welling. Auto-encoding variational bayes. In ICLR 2014, 2014. arXiv preprint arXiv:1312.6114. [358] Jaime De Miguel Rodr´ ıguez, Maria Eugenia Villafa˜ ne, Luka Piˇ skorec, and Fernando Sancho-Caparrini. Deep form finding: Using variational autoencoders for deep form finding of structural typologies. Proceedings of the International Conference on Education and Research in Computer Aided Architectural Design in Europe, 1:71–80, 2019. ISSN 26841843. doi: 10.5151/ PROCEEDINGS-CAADESIGRADI2019 514. [359] K. Gregor, I. Danihelka, A. Graves, D. Rezende, and D. Wierstra. Draw: A recurrent neural network for image generation. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015. [360] D. Ha and D. Eck. A neural representation of sketch drawings. arXiv preprint. arXiv:1704.03477, 2017. [361] J. Cudzik and K. Radziszewski. Artificial intelligence aided architectural design. In Proceedings of the 36th eCAADe Conference - Volume 1, 2018. [362] T. White. Sampling generative networks. arXiv preprint. arXiv:1609.04468, 2016. [363] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. MIT Press, 2016. [364] C. Doersch. Tutorial on variational autoencoders. arXiv preprint arXiv:1606.05908, 2016. [365] K. O. Stanley and R. Miikkulainen. Evolving neural networks through augmenting topologies. Evolutionary Computation, 10(2):99–127, 2002. doi: 10.1162/106365602320169811. [366] F.J. Pontes, G.F. Amorim, P.P. Balestrassi, A.P. Paiva, and J.R. Ferreira. Design of experiments and focused grid search for neural network parameter optimization. Neurocomputing, 186:22–34, 2016. ISSN 0925-2312. doi: https: //doi.org/10.1016/j.neucom.2015.12.061. URL https://www.sciencedirect. com/science/article/pii/S0925231215020184. [367] T. R. Davidson, L. Falorsi, N. De Cao, T. Kipf, and J. M. Tomczak. Hyperspherical variational auto-encoders. In 34th Conference on Uncertainty in Artificial Intelligence (UAI-18), 2018. 247
[368] A. Creswell, K. Arulkumaran, and A. A. Bharath. On denoising autoencoders trained to minimise binary cross-entropy. arXiv:1708.08487, 2017. [369] K. Simonyan and A. Zisserman. Very deep convolutional networks for largescale image recognition. ICLR 2015. arXiv:1409.1556, 2015. [370] D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. arXiv preprint arXiv:1811.04551, 2018. [371] H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree. Regularisation of neural networks by enforcing lipschitz continuity. CoRR, abs/1804.04368, 2018. [372] Gonzalo Aranda-Corral and Joaqu´ ın Borrego-D´ ıaz. Ontological dimensions of semantic mobile web 2.0: First principles. In Handbook of Research on Mobility and Computing: Evolving Technologies and Ubiquitous Impacts, pages 667–688. IGI Global, 2011. [373] Marcus Foth, Bhishna Bajracharya, Ross Brown, and Greg Hearn. The second life of urban planning? using neogeography tools for community engagement. Journal of location based services, 3(2):97–117, 2009. [374] Sabyasachi Basu and Thomas G. Thibodeau. Analysis of spatial autocorrelation in house prices. Journal of Real Estate Finance and Economics, 17:61–85, 1998. ISSN 08955638. doi: 10.1023/A:1007703229507. [375] Gianmarco I.P. Ottaviano and Giovanni Peri. The economic value of cultural diversity: Evidence from us cities. 11 2004. doi: 10.3386/w10904. URL http://www.nber.org/papers/w10904.pdf. [376] Neil Smith. New globalism, new urbanism: gentrification as global urban strategy. Antipode, 34(3):427–450, 2002. [377] Edward L. Glaeser, Joseph Gyourko, and Albert Saiz. Housing supply and housing bubbles. Journal of Urban Economics, 64:198–217, 9 2008. ISSN 00941190. doi: 10.1016/J.JUE.2008.07.007. URL https://www.sciencedirect.com/ science/article/pii/S0094119008000648. [378] Byeonghwa Park and Jae Kwon Bae. Using machine learning algorithms for housing price prediction: The case of fairfax county, virginia housing data. Expert Systems with Applications, 42:2928–2934, 4 2015. ISSN 09574174. doi: 10.1016/j.eswa.2014.11.040. URL https://linkinghub.elsevier.com/retrie ve/pii/S0957417414007325. 248
[379] Sherwin Rosen and Sherwin Rosen. Hedonic prices and implicit markets: Product differentiation in pure competition. JOURNAL OF POLITICAL ECONOMY, 82:34–55, 1974. URL https://citeseerx.ist.psu.edu/view doc/summary?doi=10.1.1.517.5639. [380] Jinyao Lin and Xia Li. Knowledge transfer for large-scale urban growth modeling based on formal concept analysis. Transactions in GIS, 20(5):684– 700, 2016. doi: https://doi.org/10.1111/tgis.12172. URL https: //onlinelibrary.wiley.com/doi/abs/10.1111/tgis.12172. [381] Benjamin Allbach, Sascha Henninger, and Eugen Deitche. An urban sensing system as backbone of smart cities. In REAL CORP 2014–PLAN IT SMART! Clever Solutions for Smart Cities. Proceedings of 19th International Conference on Urban Planning, Regional Development and Information Society, pages 55–64. CORP–Competence Center of Urban and Regional Planning, 2014. [382] Karima Kourtit, Peter Nijkamp, and John Steenbruggen. The significance of digital data systems for smart city policy. Socio-Economic Planning Sciences, 58: 13–21, 6 2017. ISSN 00380121. doi: 10.1016/j.seps.2016.10.001. [383] Wencheng Yu, Qizhi Mao, Song Yang, Songmao Zhang, and Yilong Rong. Social sensing: The necessary component of planning support system for smart city in the era of big data. Planning Support Science for Smarter Urban Futures 15, pages 231–244, 2017. [384] Geoff Boeing and Paul Waddell. New insights into rental housing markets across the united states: Web scraping and analyzing craigslist rental listings. Journal of Planning Education and Research, 37:457–476, 12 2017. ISSN 0739456X. doi: 10.1177/0739456X16664789. URL http://journals.sagepub.com /doi/10.1177/0739456X16664789. [385] M L Benedikt. To take hold of space: Isovists and isovist fields. Environment and Planning B: Planning and Design, 6(1):47–65, 1979. doi: 10.1068/b060047. URL https://doi.org/10.1068/b060047. [386] John R. Searle. Minds, brains, and programs. Behavioral and Brain Sciences, 3 (3):417 – 424, 1980. doi: 10.1017/S0140525X00005756. Cited by: 3218; All Open Access, Green Open Access. [387] P. Steadman. Architectural Morphology: An Introduction to the Geometry of Building Plans. A Pion publication. Pion, 1983. ISBN 9780850860863. [388] Jerry Jen-Hung Tsai and John S Gero. A qualitative energy-based unified representation for buildings. Automation in construction, 19(1):20–42, 2010. 249
[450] Amira Mouakher, Axel Ragobert, S´ ebastien Gerin, and Andrea Ko. Conceptual coverage driven by essential concepts: A formal concept analysis approach. Mathematics, 9, 11 2021. ISSN 22277390. doi: 10.3390/MATH9212 694. [451] Roberto G. Arag´ on, Jes´ us Medina, and Elo´ ısa Ram´ ırez-Poussa. Reducing concept lattices by means of a weaker notion of congruence. Fuzzy Sets and Systems, 418:153–169, 8 2021. ISSN 01650114. doi: 10.1016/J.FSS.2020.09.013. [452] Fei Hao, Erhe Yang, Lantian Guo, Aziz Nasridinov, and Doo Soon Park. On invariance of concept stability for attribute reduction in concept lattice. Lecture Notes in Electrical Engineering, 715:101–106, 2021. ISSN 18761119. doi: 10.100 7/978-981-15-9343-7\14. [453] Salah Eddine Boukhetta, J´ er´ emy Richard, Christophe Demko, and Karell Bertet. Interval-based sequence mining using fca and the next priority concept algorithm. CEUR Workshop Proceedings, 2729:91–102, 2020. ISSN 16130073. [454] Allen Newell. Physical symbol systems. Cognitive Science, 4:135–183, 1980. ISSN 03640213. doi: 10.1016/S0364-0213(80)80015-2. [455] Herbert A. Simon. Artificial intelligence: an empirical science. Artificial Intelligence, 77:95–127, 1995. ISSN 00043702. doi: 10.1016/0004-3702(95)0 0039-H. [456] Walter Pitts and Warren S. McCulloch. How we know universals; the perception of auditory and visual forms. The bulletin of mathematical biophysics 1947 9:3, 9:127–147, 9 1947. ISSN 1522-9602. doi: 10.1007/BF02478291. URL https://link.springer.com/article/10.1007/BF02478291. [457] Yann Lecun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. Nature, 521: 436–444, 5 2015. ISSN 14764687. doi: 10.1038/NATURE14539. [458] Artur D.Avila Garcez, Tarek R. Besold, Luc De Raedt, Peter Foldiak, Pascal Hitzler, Thomas Icard, Kai Uwe Kiihnberger, Luis C. Lamb, Risto Miikkulainen, and Daniel L. Silver. Neural-symbolic learning and reasoning: Contributions and challenges. AAAI Spring Symposium - Technical Report, SS15-03:18–21, 2015. [459] Pascal Hitzler, Md Kamruzzaman Sarker, Tarek R. Besold, Artur D’Avila Garcez, Sebastian Bader, Howard Bowman, Pedro Domingos, Pascal Hitzler, Kai Uwe K¨ uhnberger, Luis C. Lamb, Priscila Mac Hado Vieira Lima, Leo De Penning, Gadi Pinkas, Hoifung Poon, and Gerson Zaverucha. Neuralsymbolic learning and reasoning: A survey and interpretation. Frontiers in 256
Artificial Intelligence and Applications, 342:1–51, 2022. ISSN 09226389. doi: 10.3233/FAIA210348. [460] Md Kamruzzaman Sarker, Lu Zhou, Aaron Eberhart, and Pascal Hitzler. Neuro-symbolic artificial intelligence. AI Communications, 34:197–209, 1 2021. ISSN 0921-7126. doi: 10.3233/AIC-210084. [461] F. Fdez-Riverola and J. M. Corchado. Forecasting red tides using an hybrid neuro-symbolic system. AI Communications, 16, 2003. ISSN 09217126. [462] Zhi Hua Zhou, Yuan Jiang, and Shi Fu Chen. Extracting symbolic rules from trained neural network ensembles. AI Communications, 16, 2003. ISSN 09217126. [463] Yoshua Bengio. The consciousness prior. arXiv, 9 2017. doi: 10.48550/arxiv.1 709.08568. URL https://arxiv.org/abs/1709.08568v2. [464] Nello Cristianini. On the current paradigm in artificial intelligence. AI Communications, 27, 2014. ISSN 09217126. doi: 10.3233/AIC-130582. [465] Kevin Xia, Kai Zhan Lee, Yoshua Bengio, and Elias Bareinboim. The causalneural connection: Expressiveness, learnability, and inference. Advances in Neural Information Processing Systems, 13:10823–10836, 2021. ISSN 10495258. [466] Gregory Bateson. Steps to an Ecology of Mind : collected essays in anthropology, psychiatry, evolution, and epistemology. University of Chicago Press, 2 1999. doi: 10.7208/CHICAGO/9780226924601.001.0001. [467] Jaime F C´ ardenas-Garc´ ıa. The central dogma of information. 2022. doi: 10.3 390/info13080365. URL https://doi.org/10.3390/info13080365. [468] Jaime F. C´ ardenas-Garc´ ıa and Timothy Ireland. Bateson information revisited: A new paradigm. Proceedings 2020, Vol. 47, Page 5, 47:5, 5 2020. ISSN 2504-3900. doi: 10.3390/PROCEEDINGS2020047005. URL https://www.mdpi.com/250 4-3900/47/1/5. [469] Robert G. Frank Jr. Instruments, nerve action, and the all-or-none principle. Instruments, 9:208–235, 1994. URL https://www.jstor.org/stable/302006. [470] Jaime de Miguel Rodr´ ıguez. Concept emergence from complex sensory data: A connectionist model. November 2023. [471] Paul Vogt. Anchoring symbols to sensorimotor control. 2003. URL https: //web-archive.southampton.ac.uk/cogprints.org/3060/. 257
[472] Mariarosaria Taddeo and Luciano Floridi. A praxical solution of the symbol grounding problem. Minds and Machines, 17:369–389, 12 2007. ISSN 09246495. doi: 10.1007/S11023-007-9081-3. [473] Li Deng. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Processing Magazine, 29(6):141–142, 2012. doi: 10.1109/MSP.2012.2211477. [474] Franc¸ois Chollet et al. Keras. https://keras.io, 2015. [475] Farhad Soleimanian Gharehchopogh. An improved harris hawks optimization algorithm with multi-strategy for community detection in social network. Journal of Bionic Engineering, 20(3):1175–1197, May 2023. ISSN 25432141. doi: 10.1007/s42235-022-00303-z. URL https://doi.org/10.1007/s422 35-022-00303-z. [476] Shima Imani and Eamonn Keogh. Matrix profile xix: Time series semantic motifs: A new primitive for finding higher-level structure in time series. Industrial Conference on Data Mining, 2019-November:329–338, 11 2019. ISSN 15504786. doi: 10.1109/ICDM.2019.00043. [477] Mehdi Ayar, Ayaz Isazadeh, Farhad Soleimanian Gharehchopogh, and MirHojjat Seyedi. Chaotic-based divide-and-conquer feature selection method and its application in cardiac arrhythmia classification. The Journal of Supercomputing, 78(4):5856–5882, Mar 2022. ISSN 1573-0484. doi: 10.1007/ s11227-021-04108-5. URL https://doi.org/10.1007/s11227-021-04108-5. [478] Natalia D´ ıaz-Rodr´ ıguez, Alberto Lamas, Jules Sanchez, Gianni Franchi, Ivan Donadello, Siham Tabik, David Filliat, Policarpo Cruz, Rosana Montes, and Francisco Herrera. Explainable neural-symbolic learning (x-nesyl) methodology to fuse deep learning representations with expert knowledge graphs: The monumai cultural heritage use case. Information Fusion, 79:58–83, 2022. ISSN 1566-2535. doi: https://doi.org/10.1016/j.inffus.2021.09.022. URL https://www.sciencedirect.com/science/article/pii/S1566253521001986. [479] Zhiquan He, Wenming Cao, Jianhe Yuan, Zhihai He, and Zhi Zhang. Fast deep neural networks with knowledge guided training and predicted regions of interests for real-time video object detection. IEEE Access, 6:8990 – 8999, 2018. doi: 10.1109/ACCESS.2018.2795798. Cited by: 49; All Open Access, Gold Open Access. [480] Ilesh Dattani and Max Bramer. Utilizing symbol hierarchies and qualitative models for knowledge guided induction. Number 198, page 2/1–2/4, 1996. Cited by: 0. 258
[481] Khedidja Boulanouar, Allel Hadjali, and Mohand Lagha. A hybrid approach for linguistic summarization of time series. 2020 International Conference on Data Analytics for Business and Industry: Way Towards a Sustainable Economy, ICDABI 2020, 10 2020. doi: 10.1109/ICDABI51230.2020.9325701. [482] Katarzyna Kaczmarek-Majer and Olgierd Hryniewicz. Application of linguistic summarization methods in time series forecasting. Information Sciences, 478:580–594, 4 2019. ISSN 00200255. doi: 10.1016/J.INS.2018.11.036. [483] Elizabeth Black, Martim Brand˜ ao, Oana Cocarascu, Bart De Keijzer, Yali Du, Derek Long, Michael Luck, Peter McBurney, Albert Mero˜ no-Pe˜ nuela, Simon Miles, Sanjay Modgil, Luc Moreau, Maria Polukarov, Odinaldo Rodrigues, and Carmine Ventre. Reasoning and interaction for social artificial intelligence. AI Communications, 35:309–325, 9 2022. ISSN 09217126. doi: 10.3233/AIC-220133. [484] Oliver Lemon. Conversational ai for multi-agent communication in natural language. AI Communications, 35:295–308, 9 2022. ISSN 09217126. doi: 10.323 3/AIC-220147. [485] Jack Corrigan. The evolution of artificial intelligence (ai) spending by the u.s. government, 2023. URL https://www.brookings.edu/articles/the-evolu tion-of-artificial-intelligence-ai-spending-by-the-u-s-government/. Accessed: 2024-09-18. 259