scieee AI-readable full text Open interactive document viewer

Artificial Intelligence for knowledge discovery and generation

Domingo Roig, Oriol

Abstract

En els darrers anys, la indústria de la intel·ligència artificial ha aprofitat el poder de la computació conjuntament amb els models d'aprenentatge profund per construir aplicacions d'avantguarda. Algunes d'aquestes aplicacions, com ara assistents personals o bots de xat, depenen en gran mesura de Bases de Dades de Coneixement, un dipòsit de dades sobre dominis específics. Tot i això, aquestes bases de dades no només han d'ingerir fets nous constantment per actualitzar-se amb la informació més recent sobre el seu domini, sinó que també han de recuperar el coneixement ingerit de manera comprensible per a la gent, en la majoria dels casos. Centrat en aquest darrer punt, fer que el coneixement sigui fàcilment accessible pels humans, aquest treball es centra en donar accés automàticament a aquest coneixement mitjançant el llenguatge natural. Ho fem construint un model únic capaç d'extreure coneixements donat uns enunciats en llenguatge natural, així com de generar-los donat un cert coneixement. La solució proposada, una arquitectura basada en un Transformer eficient, s'entrena en un entorn semi-supervisat de múltiples tasques, seguint un règim d'entrenament cíclic. Els nostres resultats superen l'estat de l'art en l'extracció de coneixement per models sense supervisió, i també s'assoleixen resultats satisfactoris per la tasca de generació de text. El model resultant es pot entrenar fàcilment en qualsevol nou domini, amb dades no paral·leles, simplement afegint text i coneixement al respecte, gràcies al nostre marc d'entrenament cíclic. A més a més, aquest entorn semi-supervisat és útil per aconseguir un aprenentatge permanent.

Full text

Artificial Intelligence for Knowledge Generation and Knowledge Discovery Oriol Domingo Roig Universitat Politècnica de Catalunya Facultat d’Informàtica de Barcelona Escola Tècnica Superior d’Enginyeria de Telecomunicacions de Barcelona Facultat de Matemàtiques i Estadístiques Submitted in satisfaction of the requirements for the Degree of Bachelor in Data Science and Engineering Supervisor Dr. Marta Ruiz Costa-Jussà Co-Supervisor PhD Candidate Carlos Escolano Peinado June, 2021 I would like to thank my family for their unconditional support on pursuing this bachelor degree. Also to my supervisor, Marta R. Costa-Jussà for her guidance and patience throughout the whole scientific journey and writing process of this thesis. i Abstract In recent years, Artificial Intelligence industry has leveraged the power of computation along deep learning models to build cutting-edge applications. Some of these applications, such as personal assistants or chat-bots, heavily rely on Knowledge Bases, a data repository about specific domains. However, not only do these data bases need to constantly ingest new facts, in order to be updated with the latest information about its domain, but they also need to retrieve the ingested knowledge in a human friendly manner, in most of the cases. Focusing on the latter, making knowledge easily accessible by humans, this work focuses on automatically giving access to this knowledge through natural language. We do this by building a single model capable of extracting knowledge given natural language utterances, as well as, generating them given some knowledge. The proposed solution, an efficient Transformer architecture, is trained in a multi-task semi-supervised environment, following a cycle training regime. We surpass stateof-the-art results in knowledge extraction for unsupervised models, and reach satisfactory results for the text generation task. The resulting model can be easily trained in any new domain with non-parallel data, by simply adding text and knowledge about it, in our cycle framework. More relevantly, this semi-supervised environment is useful for lifelong learning. ii Resum En els darrers anys, la indústria de la intel·ligència artificial ha aprofitat el poder de la computació conjuntament amb els models d’aprenentatge profund per construir aplicacions d’avantguarda. Algunes d’aquestes aplicacions, com ara assistents personals o bots de xat, depenen en gran mesura de Bases de Dades de Coneixement, un dipòsit de dades sobre dominis específics. Tot i això, aquestes bases de dades no només han d’ingerir fets nous constantment per actualitzar-se amb la informació més recent sobre el seu domini, sinó que també han de recuperar el coneixement ingerit de manera comprensible per a la gent, en la majoria dels casos. Centrat en aquest darrer punt, fer que el coneixement sigui fàcilment accessible pels humans, aquest treball es centra en donar accés automàticament a aquest coneixement mitjançant el llenguatge natural. Ho fem construint un model únic capaç d’extreure coneixements donat uns enunciats en llenguatge natural, així com de generar-los donat un cert coneixement. La solució proposada, una arquitectura basada en un Transformer eficient, s’entrena en un entorn semi-supervisat de múltiples tasques, seguint un règim d’entrenament cíclic. Els nostres resultats superen l’estat de l’art en l’extracció de coneixement per models sense supervisió, i també s’assoleixen resultats satisfactoris per la tasca de generació de text. El model resultant es pot entrenar fàcilment en qualsevol nou domini, amb dades no paral·leles, simplement afegint text i coneixement al respecte, gràcies al nostre marc d’entrenament cíclic. A més a més, aquest entorn semi-supervisat és útil per aconseguir un aprenentatge permanent. iii Resumen En los últimos años, la industria de la inteligencia artificial ha aprovechado el poder de la computación conjuntamente con los modelos de aprendizaje profundo para construir aplicaciones de vanguardia. Algunas de estas aplicaciones, tales como asistentes personales o chat-bots, dependen en gran medida de Bases de Datos de Conocimiento, un depósito de datos sobre dominios específicos. Sin embargo, estas bases de datos no sólo deben ingerir hechos nuevos constantemente para actualizarse con la información más reciente sobre su dominio, sino que también tienen que recuperar el conocimiento ingerido de manera comprensible para la gente, en la mayoría de los casos. Centrado en este último punto, hacer que el conocimiento sea fácilmente accesible por los humanos, este trabajo se caracteriza por dar acceso automáticamente a este conocimiento mediante el lenguaje natural. Lo hacemos construyendo un modelo único capaz de extraer conocimientos dado unos enunciados en lenguaje natural, así como de generar texto dado un cierto conocimiento. La solución propuesta, una arquitectura basada en un Transformer eficiente, se entrena en un entorno semi-supervisado de múltiples tareas, siguiendo un régimen de entrenamiento cíclico. Nuestros resultados superan el estado del arte en la extracción de conocimiento para modelos sin supervisión, y también se alcanzan resultados satisfactorios para la tarea de generación de texto. El modelo resultante se puede entrenar fácilmente en cualquier nuevo dominio, con datos no paralelos, simplemente añadiendo texto y conocimientos al respecto, gracias a nuestro marco de entrenamiento cíclico. Además, este entorno semi-supervisado es útil para conseguir un aprendizaje permanente. iv List of Tables 3.1 Comparative table of the 5 released T5 models with respect to the number of parameters and the storage size. . . . . . . . . . . . . . . . 14 4.1 Corpus statistics of the WebNLG 3.0 English version. Train and Dev data sets are the same for both tasks (RE and SR), but they have specific test sets. Properties row reveals the number of unique DBpedia properties on each split. . . . . . . . . . . . . . . . . . . . . . . . . . 20 4.2 Number of the test instances for surface realisation and relationship extraction with respect to the different data types. . . . . . . . . . . . 20 4.3 Comparison between models’ performance at different phases, with different data amounts, regarding test set in Surface Realisation task. Bestinbold................................. 23 4.4 Detailed view of models’ performance (Surface Realisation) with regards to the three data types: seen categories, unseen entities and unseen categories. For each data amount, the best model is presented. 24 4.5 Comparative table of models’ prediction regarding seen categories (seen data) and unseen categories (unseen predicates) for the Surface Realisation task. For each data amount, the best model is presented. Golden standard is a human reference of the desired prediction. . . . 25 4.6 Comparison between models’ performance at different phases, with different data amounts, regarding test set on strict match in Relationship Extraction. Best in bold. . . . . . . . . . . . . . . . . . . . . . . 26 4.7 Detailed view of models’ performance (Relationship Extraction) with regards to the three data types: seen categories, unseen entities and unseen categories, in strict measurement. For each data amount, the best model is presented. . . . . . . . . . . . . . . . . . . . . . . . . . 26 4.8 Comparative table of models’ prediction regarding seen categories (seen data) and unseen categories (unseen predicates) for the Relationship Extraction task. For each data amount, the best model is presented. Golden standard is a human reference of the desired prediction.................................... 27 v 4.9 Summary of SOTA results on both, supervised & unsupervised learning, in the Surface Realisation task. Our best semi-supervised model is added, which does not surpass unsupervised SOTA, and the best baseline is also appended to the table. . . . . . . . . . . . . . . . . . 28 4.10 Summary of SOTA results on both, supervised & unsupervised learning, in the Relationship Extraction task. Our best semi-supervised model is added, which outperforms unsupervised SOTA regardless thematch.................................. 29 LIST OF TABLES vi List of Figures 1.1 Data source is text corpus, in which several entities are mentioned (highlighted). These entities with their corresponding relations (triples) are depicted in the Knowledge Base overview. The data base exhibits a graph structure, in which extracted facts about the data source are represented. Notice that doted lines represent triples that are not present in this text corpus, but they might have appeared on other samples. One example of a triple is: 〈High Renaissance, location, Italy〉, which stands for subject: High Renaissance, predicate: location, and object: Italy. . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.2 Overview of the problem. A Deep Learning model is used for both, discovering new facts given a source (Knowledge Discovery), as well as, providing an interface to the output of a given query (Knowledge Generation). For example, if the system is queried about all the monuments located in Barcelona, the model must translate the triples (query response) inside the red-dotted rectangle into a natural language representation. . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.3 Gantt Diagram of the Degree Thesis . . . . . . . . . . . . . . . . . . 5 2.1 Main task represents the ultimate task our model needs to accomplish (Source Language →Target Language). However, we use a model to predict (synthetic text) in the reverse direction (Target Language → Source Language), which is the Back Translation step. Thus, this synthetic text, aligned with its corresponding monolingual text, can be added to the data base used in our main task. . . . . . . . . . . . 8 vii 3.1 An overview of the text-to-text framework. Every task, including pretrained ones and our tasks, is cast as feeding the model some text, in which the task to be solved is specified (underlined text). In our tasks (bottom diagram), we refer to triples as "rel" and text as "lex", then, we can specify which task has to be solved with these tokens. For instance, "translate rel to lex" stands for "translate triples to text" (Surface Realisation task). . . . . . . . . . . . . . . . . . . . . . . . . 13 3.2 (1) The cycle training approach is based on two mapping functions f:K → T and g:T → K, however, our approach uses a single function (model). This framework introduces two cycle losses that capture the intuition that if we translate from one side to the other, and back again we should obtain the original sample: (2) forward cycle loss: x→f(x)→g(f(x)) ≃x, and (3) backward cycle loss: y→g(y)→f(g(y)) ≃y.......................... 15 3.3 This block diagram is an overview of our contributions. We focus on extracting data from only Wikipedia (text) and Wikidata (triples). Apart from that, we adapted the cycle training framework to use it with a single model (zα) in a multi-task set-up ( z:T → K and z:K → T ). ................................ 18 LIST OF FIGURES viii Figure 1.3: Gantt Diagram of the Degree Thesis •M3-Training: Model implementation and fine-tuning on both data sets, as well as, cycle training implementation and training. •M4-Test and Results: Experiments’ evaluation and comparison. •M5-Final Document: Thesis document preparation. 1.6 Incidents & Modifications The first incident I had to face was that collecting triples required running queries against a knowledge base, in SPARQL1(Resource Description Framework query language), which I had no prior experience, and neither the focus of the work was to learn it. However, cleaning and processing the collected unsupervised data, triples and text, took fewer days than the estimated. Hence, the deadline of the data preparation milestone was reached according to the plan. Another problem was that the selected T5 model2was showing some computational instability due to its storage size, which was hard to fit on the available GPUs. This 1SPARQL Query Language. 2This will be detailed on Chapter 3. Introduction 5 not only complicated the fine-tuning procedure during the training phase, but it also conditioned first experiments on the cycle training. Finally, the training procedure was tracked using the Weight & Biases3API to see the evolution and performance of each model. Unexpectedly, the Python API started to fail, producing an infinite waiting loop without putting the model into training mode, which was hard to debug. After several weeks, I received an e-mail reporting that the issue was due to the wandb client version. This issue produced a delay of one week during the training milestone. 1.7 Organisation The rest of this work is organised as follows. In Chapter 2 previous approaches reported in the literature are introduced, and then, the current solution is proposed in Chapter 3. Thereafter, the experiment’s results are explained in Chapter 4, before Chapter 5 presents the conclusions of this project. 3Developer tools for machine learning: Weight & Biases Introduction 6 Chapter 2 Related Work Relationship Extraction (RE) and Surface Realisation (SR) are tasks that have been tackled during the last decades [6] [27]. At the beginning of this century, models heavily relied on statistical methods [10], however, during the last years end-to-end deep learning (DL) approaches [1] [28] have surpassed those early statistical models reaching state-of-the-art (SOTA) results, following the tendency in other domains. 2.1 Author’s previous work Personally, I conducted some research on which Machine Translation (MT) techniques are worth applying to the SR task [13]. One of the key aspects of this work was that Back Translation (BT) can potentially lead to significant improvements in several automatic metrics, specifically in data samples that were not present in the original data set. Back Translation operates in a semi-supervised setup where both parallel corpora1 and monolingual data in the target language are available [33] [Figure 2.1]. The training regime is the following: 1. Train an intermediate model on the parallel data, which is used to translate from the target monolingual data to the source language, i.e. text-to-triples. It results in a parallel corpus where the source, triples, is synthetic MT output while the target is human written. 2. The generated synthetic parallel corpus is added to the original data set in 1A parallel corpora is a text placed alongside its translation/s 7 Figure 2.1: Main task represents the ultimate task our model needs to accomplish (Source Language →Target Language). However, we use a model to predict (synthetic text) in the reverse direction (Target Language →Source Language), which is the Back Translation step. Thus, this synthetic text, aligned with its corresponding monolingual text, can be added to the data base used in our main task. order to train a final model that will translate from the source (triples) to the target language (text). 2.2 Unsupervised learning It is widely known that DL models significantly improve at the expenses of data. However, there is always an implicit trade-off between generalisation and memorisation of the training data [15]. Most of the unsupervised techniques that are applied to DL models are designed to overcome such an issue, but also the data scarcity one [2]. Mainly, semi-supervised models learn from a small number of labeled samples and (usually a large number of) unlabeled samples [40]. In this work, a semi-supervised approach to overcome both issues, generalisation and data scarcity, is explored. Recently, Martin Schmitt et al. 2020 [32] presented the first approach to unsupervised text generation from KB, and simultaneously showed how it can be used for unsupervised RE. They proposed two different methods: a rule-based system; they considered as a baseline; and a neural sequence-to-sequence system; which considerably improved baseline results. Related Work 8 On one hand, the rule-based system relies on several steps such as preprocessing, removing stop words, part of speech tagging (similar to semantic parsers), heuristics and template linearisation. On the other hand, the neural model is a BiLSTM [21] architecture with dropout [35] on both, encoder and decoder. Moreover, the neural model is trained in a multi-task environment (more details on 2.3) and fine-tuned with noisy source samples [38]. Regarding their training regime, firstly, they obtained a language model for both graphs and text by training one epoch with the denoising auto-encoder objective [39], and later on iterative BT is applied. Qipeng Guo et al. 2020 [20] developed an unsupervised training method that can bootstrap from fully non-parallel triples and text data, and iteratively back translate between the two forms. This cycle training framework achieves SOTA results for unsupervised models in this domain. Nevertheless, their method relies on two different models, a T5 pre-trained model [31] for SR and a BiLSTM model for RE. Consequently, not only does this approach seek to have a common data representation between these two models, but this also constraints the optimization procedure as both models cannot be updated together, rather they are updated one-by-one, making their loss function non-differentiable. 2.3 Multi-task learning Multi-task learning is a learning environment in which several tasks are simultaneously learned by a single model. Such an approach offers advantages like improved data efficiency, reduced overfitting through shared representations, and fast learning by leveraging auxiliary information [11]. For these reasons, our final model is also trained in a multi-task set-up, i.e. the resulting model is capable of providing answers to both tasks, RE and SR. As previously mentioned, Martin Schmitt et al. trained their sequence-to-sequence model in a multi-task environment. In particular, the model is trained for both conversion tasks, sharing encoder and decoder. However, they tell the decoder which Related Work 9 type of output should be produced (text or triples) by means of cell state initialisation in the decoder side, with an embedding corresponding to the desired output type. Alternatively, Oshin Agarwal et al. 2020 [1] overcome the need of task specification at decoder level thanks to the input format of the chosen model. The model is a pre-trained T5 from Google, in which the task to be solved is simply identified by adding prefixes to the input2. Moreover, not only does their work solve RE and SR in a multi-task set-up, but they also build a bilingual model for English and Russian languages. Stated by the authors, these capabilities remarkably improves on unseen relations and Russian. Apart from that, they also perform data augmentation but not in an unsupervised manner, they rather pre-trained the original model on a parallel corpus of News data: WMT-News corpus (WMT) [36]. 2This will be later detailed on Chapter 3. Related Work 10 Chapter 3 Methodology In this chapter, we explain and formalise the problem, as well as, the approaches taken to solve it. We derive some mathematical concepts to provide a more detailed definition of our tasks: Relationship Extraction (RE) and Surface Realisation (SR), and semi-supervised framework. Apart from that, we present the system architecture for the baseline and cycle learning process. For both, we detail their background context, and how we have tailored them to our scenario. 3.1 Problem formulation We can formalise a Knowledge Base (KB) as a labeled directed graph [Eq 3.1] (K) over the triples in the data base. The nodes represent entities (subjects or objects) and edges represent the relationship between these entities. Notice that the direction of the edge embeds semantic information as it determines who is the subject and the predicate of the corresponding triple. Monolingual corpus, particularly text, can be easily formalized as a set of sentences, which are sequences of words [Eq 3.2] (T). K:= {(V, E) : ∀(s, p, o)i∈KB →(si, oi)∈V, (pi)∈E, arc(si, oi) = pi}(3.1) T:= {si:∀si= [w1, ..., wn]∈Corpus}(3.2) The supervised examples S: (x, y)∈ K × T are a pair of triples and pieces of text 11 that are aligned, i.e. there is a correspondence between the information embedded in the triples and the corresponding text. Given a data set of supervised examples, we can train a model to do RE and SR from S. In [Eq 3.3] we aim to generate text (triples) from triples (text) using some function fθ(gφ). Ideally, these functions are optimised over Sby means of a maximum loglikelihood estimation [Eq 3.4].        fθ(x)=[w1, ..., wm] = ˆy gφ(y) = {(s1, p1, o1), ..., (sr, or, pr)}= ˆx (3.3) J(θ, φ) = E(x,y)∼S [−log p(y|x;θ)−log p(x|y;φ) ] (3.4) Nevertheless, this formulation assumes that two different functions are used to solve both tasks, i.e. it does not consider the multi-task learning environment for a single model. In such a case, both functions (fθand gφ) are merged into a single one (zα). Thus, the above equation can be reformulated [Eq 3.5] according to such a learning process, which is the one followed in this work. J(α) = E(x,y)∼S [−log p(y|x;α)−log p(x|y;α) ] (3.5) 3.2 Baseline It has been shown [25] [24] that pre-training hard followed by an easy fine-tuning is a good strategy for Natural Language Processing tasks. Thanks to their transfer learning ability, fine-tuning T5 [24] achieves SOTA results on many benchmarks, and also mBART [25] reaches very competitive results. The advantage of pre-training hard is making the model capable of generalising better on unseen tokens. In this work, we are going to use the T5 model [31] as the main system, which will serve as a baseline for later comparison against the cycle training approach. The T5 model is a Transformer architecture, that it is roughly equivalent to the original Methodology 12 Figure 3.1: An overview of the text-to-text framework. Every task, including pretrained ones and our tasks, is cast as feeding the model some text, in which the task to be solved is specified (underlined text). In our tasks (bottom diagram), we refer to triples as "rel" and text as "lex", then, we can specify which task has to be solved with these tokens. For instance, "translate rel to lex" stands for "translate triples to text" (Surface Realisation task). one [37], however, it introduces an unified framework that converts all text-based language problems into text-to-text format as shown in [Figure 3.1][31]. The few modifications introduced by T5’s authors are that they removed the layer norm bias, placing the layer normalization outside the residual path, and used a relative position embedding rather than sinusoidal position signal or learned position embeddings. This model has been pre-trained on the "Colossal Clean Crawled Corpus" [31]: a collection of scraped and cleaned natural English text that is about 750 GB. Then, a single model is trained on a set of tasks, such as machine translation, summarisation, text classification, question answering, and so forth, following the unified framework, in which the model is fed some text for context and is then asked to produce some output text. This framework provides a consistent training objective for both, pretraining and fine-tuning, being the latter of special interest in our work as it eases the optimisation of our single model (zα) on both downstream tasks at the same time. After this pre-training process, they released 5 different models [Table 3.1]. In particular, we are going to use the T5-Base model due to its good performance and our limited computational power. Methodology 13 T5-Small T5-Base T5-Large T5-3B T5-11B Parameters 60 M 220 M 770 M 3 B 11 B Size 230 MB 850 MB 2.7 GB 10.6 GB 42.1 GB Table 3.1: Comparative table of the 5 released T5 models with respect to the number of parameters and the storage size. 3.3 Cycle training Cycle training was originally suggested as an image-to-image translation, rather than text-to-text (our current approach), a problem where the goal is to learn a mapping between an input image and an output image [41]. It was presented as a solution to the absence of paired examples. This solution was based on a cycle consistency loss relying on Generative Adversarial Networks [19]. The main constraint for using cycle training is that there must exist two complementary tasks that guarantees that the input of one task is the output of the other task, and vice-versa [Figure 3.2 (1)]. For instance, a set of triples can be fed into our model to generate some text. The resulting text can also be fed into this model to generate a set of triples, which should resemble the original ones [Figure 3.2 (2)]. The same procedure is applied in the reverse direction, i.e. starting from text and generating synthetic triples [Figure 3.2 (3)]. At this point, we can cycle-train our model since a reference of our hypothesis exists on both sides. These steps constitute the iterative loop in which cycle training is based on. If the previous constraint holds (existence of complementary tasks), which is our case, then, it is possible to build a bijective mapping function that given a variable xsatisfies x=f−1(f(x)), where f−1is the inverse function of f1. In our case, both functions represent the same model, hence, it must hold that f−1=f, which is an involutory function. This constraint needs to be slightly relaxed as our variable x, representing text or triples, is concatenated with an extra token at the beginning of the sentence in order to specify the model which output should be generated, and the output does not 1In [Figure 3.2] grepresents f−1. Methodology 14 176,000 unique entities with 532,288 triples and 240,024 paragraphs. Finally, we built two data sets, one for the triples with 271,095 instances (each instance has between 1 and 6 triples), and another one for the natural text, with 240,024 instances (one instance per paragraph) of 459.67characters on average length. 4.3 Evaluation metrics On one hand, the quality of the model for SR is evaluated using some of the most popular automatic text generation metrics: •BLEU: Scores are calculated, based on a modified form of precision, for individual translated segments by comparing them with a set of good reference translations. These scores are averaged through the whole corpus to reach a quality estimation ranging between 0-1, being the higher the better [29]. •TER: This score measures the amount of editing that has to be performed in order to attain an exact match to the reference translation. The final value is also between 0-1, however, the higher the worse [34]. •chrF++: Scores are computed based on a n-gram F-score with respect to its aligned reference, which also takes into account some morpho-syntactic phenomena2. The final estimation value is between 0-1, being the higher the better again [30]. The outputs were directly compared to its gold-standards without post-processing, with a maximum of 5 references. On the other hand, the quality of the parsed triples is harder to evaluate than plain text due to two main issues: triples format forgetting and triples ordering. Hence, post-processing filters are implemented to ensure triples format is consistent and it does not condition the evaluation results. Moreover, the evaluation script already deals with the triples ordering issue, which is provided by the WebNLG team3. 2The parameters were set to: character 6-grams and β= 2 3The evaluation script can be found here. Experiments 21 The metrics used are the well-known Precision, Recall and F1 score. However, four different ways to measure these metrics are investigated: •Strict: Exact match of the hypothetical triple with the reference triple is required. •Exact: Exact match of the hypothetical triple with the reference triple is required, but the entity type (subject, predicate, object) is irrelevant •Partial: The hypothetical triple should match at least partially with the reference triple, and the entity type is irrelevant. •Type: The hypothetical triple element should match at least partially with the reference triple, and the entity type should match with the reference. 4.4 Experiment results Two steps are executed on each experiment: fine-tuning and cycle training. Similar hyper-parameters have been used on both phases, and for the sake of simplicity, each iteration of the cycle training is configured with the same hyper-parameters. Specifically, we trained on mini-batches (8 samples) with accumulation steps (4 steps), at a learning rate of 2.0e−4(fine-tuning) and 1.0e−5(cycle training). The maximum number of epochs are set to 50 (fine-tuning) and 30 (cycle training) with 5 epochs of patience. The maximum source and target length is limited to 64 tokens. Moreover, the predictions during the cycle training are generated with (4) beam search using 64 tokens as the maximum source and target length as well. Repetition and length penalty are applied, 2.5 and 1.0 respectively, along early stopping. The number of cycle steps are restricted to 5. For the fine-tuning phase, we experimented with four different set-ups, decreasing the amount of data used: Large (100%), Medium (15%), Small (3%) and Extra-Small (1%). Moreover, we present the model’s results after the fine-tune step (baseline), after cycle training with unsupervised data (cycle) and after cycle training with unsupervised and supervised data (cycle + parallel). The Large model has been Experiments 22 BLEU (↑)TER (↓)chrF++ (↑) L-Baseline 44.68 0.51 0.54 L-Cycle 42.41 0.53 0.51 L-Cycle + Parallel 39.72 0.56 0.56 M-Baseline 42.58 0.51 0.60 M-Cycle 38.81 0.54 0.54 M-Cycle + Parallel 39.98 0.56 0.60 S-Baseline 31.08 0.64 0.53 S-Cycle 33.89 0.59 0.5 S-Cycle + Parallel 39.75 0.58 0.54 XS-Baseline 13.92 0.85 0.45 XS-Cycle 23.76 0.68 0.47 XS-Cycle + Parallel 29.30 0.68 0.54 Table 4.3: Comparison between models’ performance at different phases, with different data amounts, regarding test set in Surface Realisation task. Best in bold. cycle-trained with 40,000 synthetic samples, which can be different on each cycle iteration, whilst the other models have been cycle-trained with 30,000 samples. Surface Realisation First of all, we present the results obtained in the text generation task with regard to the test set [Table 4.3], later on this report, the best models are detailed in each category. In [Table 4.3], it is clearly shown that cycle training only improves baseline results when using an amount of data equal or lower than 3% (S and XS models). In fact, the S-Cycle + Parallel is capable of reaching M-Cycle + Parallel results, and surpassing M-Cycle ones. The greatest improvement using the semi-supervised approach is achieved by the XS model. In general, models do not suffer too much in terms of chrF++ score since it is the most constant score among the others. This indicates that words appearing in the generated text also appear in the golden standard references in the same ratio across the different models, however, these predictions have every-likelihood to differ in phrase construction and/or word ordering, which results in greater differences in the other metrics. Apart from that, we also observe that cycle training achieves better results when Experiments 23 Model Seen Categories Unseen Entities Unseen Categories BLEU TER chrF++ BLEU TER chrF++ BLEU TER chrF++ L-Baseline 51.70 0.49 0.57 45.44 0.49 0.55 38.00 0.52 0.51 M-Baseline 46.55 0.52 0.61 44.48 0.48 0.63 38.21 0.51 0.57 S-Cycle + Parallel 41.36 0.6 0.59 40.96 0.52 0.59 37.74 0.51 0.56 XS-Cycle + Parallel 30.45 0.74 0.52 30.73 0.66 0.56 27.70 0.65 0.54 Table 4.4: Detailed view of models’ performance (Surface Realisation) with regards to the three data types: seen categories, unseen entities and unseen categories. For each data amount, the best model is presented. unsupervised data (*-Cycle)4is used along the supervised one (*-Cycle + Parallel), during the training process. This behaviour might suggest that cycle training, with only unsupervised data, tends to understand more unseen words than its baseline, however, it might suffer from generating appropriate text structure due to synthetic data. Taking these results into consideration, we will only investigate, in more detail, the performance of the best model regarding each data amount. This analysis [Table 4.4] provides a performance overview on each of the three data types, explained earlier in this Chapter. S/XS-Cycle + Parallel models have a similar behaviour: constant performance on seen categories and unseen entities based on BLEU score; and small performance drop on unseen categories (3∇BLEU, 0.01∇TER, 0.03/0.02∇chrF++)5. Hence, the performance for the models using a small amount of fine-tuned data is quite robust to unseen tokens after cycle training is applied. However, L/M-Baseline drastically suffer in unseen categories, reaching a similar or worse (depending on the metric) performance than the S-Cycle + Parallel model. Notice that this model has not been trained with our framework, hence, this suggests that our semi-supervised approach really helps to better leverage performance across seen and unseen data as shown by small models. Finally, we can inspect the models’ output to better understand how they generate text [Table 4.5], based on our human criteria6. Despite we cannot display several examples, these patterns are broadly observed through different text generations. In the seen data, the S-Cycle + Parallel model generates text in an incorrect verb tense (past), and its text structure is not natural either fluent at all (this was 4"*" means any model: Large (L), Medium (M), Small (S) and Extra Small (XS) 5Remember that the lower the TER score, then, the better the model. 6Human evaluation for the whole test set has not been possible due to time constraints. Experiments 24 Model Seen Categories Unseen Categories <s> Aleksandr Prudnikov <p> height <o> 185.0 cm <s> FC Spartak Moscow <p> ground <o> Otkrytiye Arena <s> Aleksandr Prudnikov <p> club <o> FC Spartak Moscow <s> Brexpiprazole <p> instance of <o> Medication <s> Brexpiprazole <p> physically interacts with <o> Dopamine receptor D2 L-Baseline Aleksandr Prudnikov is 185 cm tall and plays for FC Spartak Moscow at the Otkrytiye Arena. Xpiprazole, which interacts with the dopamine receptor D2, is a type of medication. M-Baseline Aleksandr Prudnikov is 185.0 cm tall and plays for FC Spartak Moscow who play their home games at the Otkrytiye Arena. Brexpiprazole is a type of medication that physically interacts with the Dopamine receptor D2. S-Cycle + Parallel Aleksandr Prudnikov, who played for FC Spartak Moscow at the Otkrytiye Arena, has a height of 185.0 cm. Relative to brexpiprazole, it physically interacts with Dopamine receptor D2. XS-Cycle + Parallel Aleksandr Prudnikov, 185,0 cm tall, plays for FC Spartak Moscow and his club is Otkrytiye Arena. p> instance of o> medication brexpiprazole p> physically interacts with o> Dopamine receptor D2. Golden standard Aleksandr Prudnikov, 185 centimetre tall, plays for the Otkrytiye Arena based FC Spartak, Moscow. Brexpiprazole is a medicament that physically interacts with the Dopamine receptor D2. Table 4.5: Comparative table of models’ prediction regarding seen categories (seen data) and unseen categories (unseen predicates) for the Surface Realisation task. For each data amount, the best model is presented. Golden standard is a human reference of the desired prediction. earlier reported based on chrF++ score). Moreover, the XS-Cycle + Parallel model misunderstands some triples (or triples direction) since it generates: "(Aleksandr Prudnikov) his club is Otkrytiye Arena.", which should be "Otkrytiye Arena is the ground of (Aleksandr Prudnikov) his club" or similar. However, the L/M-Baseline outputs cover all the information contained in the original triples, as well as, it fluently translates them to text, but the medium model makes a small mistake as the pronoun "who" is not precededed by a ",". Surprisingly, in the unseen predicates, both L/M-Baseline models perform pretty well, but these triples are not so challenging as they are almost in natural language. However, S-Cycle + Parallel suffers from data coverage, since it does not mention that "Brexpiprazole" is a medication, but the worst translation is the XS-Cycle + Parallel by far. It seems as if the model did not understand the task, because it rather returned the input. Relationship Extraction Now, we show the results (strict measure) obtained in the parsing task as for the test set [Table 4.6], later on this section, the best models are analysed in each category. In [Table 4.6], there is an opposite trend to the one observed in the previous task, Experiments 25 F1 Precision Recall L-Baseline 0.394 0.389 0.401 L-Cycle 0.336 0.329 0.349 L-Cycle + Parallel 0.346 0.331 0.359 M-Baseline 0.273 0.275 0.274 M-Cycle 0.330 0.324 0.340 M-Cycle + Parallel 0.430 0.426 0.440 S-Baseline 0.379 0.385 0.376 S-Cycle 0.307 0.301 0.319 S-Cycle + Parallel 0.395 0.388 0.405 XS-Baseline 0.362 0.358 0.370 XS-Cycle 0.265 0.275 0.259 XS-Cycle + Parallel 0.372 0.367 0.381 Table 4.6: Comparison between models’ performance at different phases, with different data amounts, regarding test set on strict match in Relationship Extraction. Best in bold. since M-Cycle models already improve its baseline results, but the S-Cycle and XSCycle do not. Not until parallel data is added to the unsupervised one (*-Cycle + Parallel models), do we observe a boost in performance, regardless the model. Interestingly, M-Baseline significantly suffers in comparison to the other baseline models, however, cycle training really enhances its performance, which makes the M-Cycle + Parallel model the best one. We also conducted a deeper study [Table 4.7] on the performance of the best models, which are selected along the different data amounts. Remarkably, the greatest difference between models’ performance is attained in the seen categories, where F1 score greatly varies between: 0.472-0.343. However, this difference (0.129) is significantly reduced in the unseen categories (0.049), revealing that this behaviour might be a consequence of both, the cycle training environment and data reduction, since the greatest difference is in the seen domain. This seen domain, refers to samples in the Model Seen Categories Unseen Entities Unseen Categories F1 Precision Recall F1 Precision Recall F1 Precision Recall L-Baseline 0.429 0.426 0.433 0.459 0.456 0.464 0.345 0.335 0.356 M-Cycle + Parallel 0.472 0.469 0.478 0.440 0.436 0.447 0.394 0.386 0.408 S-Cycle + Parallel 0.390 0.386 0.397 0.421 0.416 0.430 0.386 0.377 0.400 XS-Cycle + Parallel 0.343 0.339 0.350 0.419 0.416 0.426 0.371 0.364 0.383 Table 4.7: Detailed view of models’ performance (Relationship Extraction) with regards to the three data types: seen categories, unseen entities and unseen categories, in strict measurement. For each data amount, the best model is presented. Experiments 26 Model Seen Categories Unseen Categories 185 centimetre tall Aleksandr Prudnikov played for the Otkrytiye Arena based FC Spartak, Moscow. Leonardo da Vinci was an Italian polymath of the High Renaissance. L-Baseline <s> Aleksandr Prudnikov <p> height <o> 185.0 (centimetres) <s> FC Spartak Moscow <p> ground <o> Otkrytiye Arena <s> Leonardo da Vinci <p> nationality <o> Italy <s> Italy <p> Renaissance M-Cycle + Parallel <s> Aleksandr Prudnikov <p> height <o> 185.0 (centimetres) <s> FC Spartak, Moscow <p> ground <o> Otkrytiye Arena <s> Leonardo da Vinci <p> occupation <o> polymath <s> Leonardo da Vinci <p> time period <o> High Renaissance S-Cycle + Parallel <s> Aleksandr Prudnikov <p> club <o> Otkrytiye Arena <s> FC Spartak, Moscow <p> height <o> 185 cm <s> Leonardo da Vinci <p> occupation <o> polymath <s> Leonardo da Vinci <p> time period <o> High Renaissance XS-Cycle + Parallel 185 centimetre tall Aleksandr Prudnikov played for the Otkrytiye Arena of FC Spartak, Moscow. <s> Leonardo da Vinci <p> occupation <o> polymath. Golden standard <s> Aleksandr Prudnikov <p> height<o> 185.0 cm <s> FC Spartak Moscow <p> ground <o> Otkrytiye Arena <s> Aleksandr Prudnikov <p> club <o> FCSpartak Moscow <s> Leonardo da Vinci <p> occupation <o> polymath <s> Leonardo da Vinci <p> time period <o> High Renaissance <s> Leonardo da Vinci <p> born <o> Italy Table 4.8: Comparative table of models’ prediction regarding seen categories (seen data) and unseen categories (unseen predicates) for the Relationship Extraction task. For each data amount, the best model is presented. Golden standard is a human reference of the desired prediction. train set, which is reduced on each model (100%-15%-3%-1%), but the test set is the same for all, hence, it is probably a seen domain for L/M-Cycle + Parallel, but not for S/XS-Cycle + Parallel. This is also supported with the fact that in the unseen domain all models reach similar scores unless the supervised one (L-Baseline). Another interesting aspect is that Recall is always a bit higher than Precision, among these models, so they likely return a high number of triples where some are relevant (most of the time triples are noisy so they are not strictly relevant). As in previous section, we can have a closer look at the models’ output to better understand how they parse text [Table 4.8], based on our human criteria7. Despite we cannot display several examples, these patterns are broadly observed through different parsing samples. In the seen categories, the M-Cycle + Parallel and LBaseline models almost extract all the relationships within the text, showcasing its high performance in the seen categories, which was shown in previous tables. The S-Cycle + Parallel model is not that robust, in fact, it builds two triples that mismatch the object (1st triple) and subject (2nd triple), likewise, it forgets some relevant triples. Interestingly, we observe that XS-Cycle + Parallel output seeks 7Human evaluation for the whole test set has not been possible due to time constraints. Experiments 27 Model BLEU (↑)TER (↓)chrF++ (↑) Oshin Agarwal et al. 2020 [1] 51.70 0.435 0.679 - supervised L-Baseline 44.68 0.510 0.540 - supervised Qipeng Guo et al. 2020 [20] 44.60 0.479 0.637 - unsupervised M-Cycle + Parallel 39.98 0.560 0.600 - semi-supervised Table 4.9: Summary of SOTA results on both, supervised & unsupervised learning, in the Surface Realisation task. Our best semi-supervised model is added, which does not surpass unsupervised SOTA, and the best baseline is also appended to the table. to have parsed triples since it returns the input text. This behaviour was already observed in [Table 4.5], but in the reverse task, where this model simply replicated the input triples in the unseen predicates. In [Table 4.8] unseen categories , M/S-Cycle + Parallel showcased what was earlier observed: they have a quite similar performance, but in this example, they fail to retrieve one triple (<s> Leonardo da Vinci <p> born <o> Italy). The XS-Cycle + Parallel model suffers even more from data coverage since it only retrieves 33% of the triples. In general, we observe that relevant triples are retrieved, however, they are likely to be noisy or corrupted (subject or object mismatch). If so, they are not counted as relevant by this strict measurement, hence, the Recall is a bit higher than the Precision, which was actually observed in previous tables. Apart from that, the L-Baseline model drops its triple extraction capacity in unseen categories showing that it might probably fail to understand the output structure, which can be seen in the second triple: the object is missing. However, models with cycle training might have seen more examples and learned a better representation, as they do not suffer that much. 4.5 State-of-the-art comparison To conclude this chapter, an overview of our current solution against state-of-the-art (SOTA) models is presented for both tasks. Experiments 28 On one hand, SR comparison [Table 4.9] shows that our supervised approach reaches similar scores, or even worse, than SOTA unsupervised model. Similarly, our best semi-supervised model performs quite worse than the SOTA unsupervised, but improves L-Baseline in chrF++ score. Notice that all these models rely on the same pre-trained T5 model, however, they differ on the fine-tuning strategy or the training regime. On the other hand, RE comparison [Table 4.10] reveals that our semi-supervised model outperforms previous SOTA unsupervised results by 0.1 point on average. Despite our model seems to be very competitive, it does not reach SOTA supervised performance, which also relies on a T5 architecture. We observe that our approach and the supervised one attain their best results in the entity type metric, whereas the unsupervised model (BiLSTM) reaches its highest performance in the partial match. Besides, the toughest metric is the strict one, which was used on all previous tables’ analysis, as it is demonstrated with the greatest performance drop by all the models. Model Match F1 Precision Recall Oshin Agarwal et al. 2020 [1] Exact 0.682 0.670 0.701 - supervised Type 0.737 0.721 0.762 Partial 0.713 0.700 0.736 Strict 0.675 0.663 0.695 M-Cycle + Parallel Exact 0.437 0.432 0.447 - semi-supervised Type 0.485 0.475 0.500 Partial 0.465 0.457 0.478 Strict 0.430 0.426 0.440 Qipeng Guo et al. 2020 [20] Exact 0.342 0.338 0.349 - unsupervised Type 0.343 0.335 0.356 Partial 0.360 0.355 0.372 Strict 0.309 0.306 0.315 Table 4.10: Summary of SOTA results on both, supervised & unsupervised learning, in the Relationship Extraction task. Our best semi-supervised model is added, which outperforms unsupervised SOTA regardless the match. Experiments 29 Chapter 5 Conclusions This project’s main goal has been to train an end-to-end multitask semi-supervised model for a Knowledge Generation and Knowledge Discovery system, so it can be used in industrial applications using Knowledge Bases, or in any other future work. As shown in this work, we have successfully accomplished this main goal, and achieve very competitive results, specifically, in the Relationship Extraction task. Through this journey, we have gained a deeper understanding of the T5’s unified text framework, as well as, the cycle training approach. The code for the system implementation will be released after the publication of our scientific paper. It took more time than expected to have the whole data collection pipeline, and afterwards, dealing with a custom implementation of the cycle training framework was also quite complicated. Thus at the end, it let us a little time to make more experiments on the cycle training hyper-parameters or try other data sets. However, we finally have trained several models and tried it with different configurations in order to obtain the highest performance on both tasks. To conclude, we have contributed with a system that is based on the T5 architecture [31]. This is a pre-trained model that is fine-tuned with different amounts of data, and then, they are trained through our cycle framework. This framework consists in translating text to triples, and vice-versa, using the model itself, so it obtains parallel data from fully non-parallel samples. These samples are added (optionally) to the original parallel data before training the final models, this process can be iteratively applied resulting in a lifelong learning loop. 30