Full text
ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships Anonymous COLING 2025 submission Abstract001 Natural Language Inference (NLI), also known 002 as Recognizing Textual Entailment (RTE), 003 serves as a crucial area within the domain of 004 Natural Language Processing (NLP). This area 005 fundamentally empowers machines to discern 006 semantic relationships between assorted sec007 tions of text. Even though considerable work 008 has been executed for the English language, it 009 has been observed that efforts for the Spanish 010 language are relatively sparse. Keeping this in 011 view, this paper focuses on generating a multi012 genre Spanish dataset for NLI,ESNLIR, par013 ticularly accounting for causal Relationships. 014 A preliminary baseline has been conceptualized 015 and subjected to an evaluation, leveraging mod016 els drawn from the BERT family. The findings 017 signify that the enrichment of genres essentially 018 contributes to the enrichment of the model’s ca019 pability to generalize.020 The code, notebooks and whole datasets for this 021 experiments is available at: https://zenodo.022 org/records/15002575 . If you are interested 023 only in the dataset you can find it here: https:024 //zenodo.org/records/15002371.025 1 Introduction026 In the field of Natural Language Processing, several 027 applications including Information Retrieval (IR), 028 Question Answering (QA), and Information Extrac029 tion (IE) necessitate machine comprehension of se030 mantic meaning in text corpora (Dagan et al.,2006). 031 However, understanding isolated sentences is insuf032 ficient in these contexts. Being able to grasp the 033 relationships between sentences is equally imper034 ative. Consequently, Natural Language Inference 035 (NLI) or Recognizing Textual Entailment (RTE) 036 has been developed to ascertain whether one sen037 tence, the hypothesis, can be inferred from a dif038 ferent sentence, the premise (Dagan et al.,2006). 039 Nonetheless, a binary label proved inadequate for 040 encompassing all possible semantic relationships 041 between sentences, such as contradiction (Giampic042 colo et al.,2008). Hence, a three-way labeling 043 scheme was devised, which comprises of Entail044 ment,Neutral and Contradiction classes. 045 This labeling technique has been employed in 046 the development of various contemporary bench047 marks such as RTE4 (Giampiccolo et al.,2008), 048 SICK (Marelli et al.,2014), which are small 049 datasets with less than 10,000 examples, SNLI 050 (Bowman et al.,2015), the first large scale 051 dataset with more than 500,000 examples, MNLI 052 (Williams et al.,2018), the multi-genre version of 053 SNLI, and XNLI (Conneau et al.,2018), a multi054 lingual subset of XNLI. While commendable clas055 sification performance has been achieved on these 056 benchmarks, the sciNLI (Sadat and Caragea,2022) 057 benchmark underlines the need for a novel label, 058 Reasoning, to encapsulate causal relationships in 059 genres such as scientific text, where ideas often 060 manifest as cause-effect sequences. The same 061 benchmark introduces a method for automated 062 premise-hypothesis pair extraction, diverging from 063 the standard practice of requiring human annotation 064 for sentence generation, thus facilitating the extrac065 tion of large volumes of examples without human 066 supervision. This method has also been success067 fully employed in the msciNLI benchmark (Sadat 068 and Caragea,2024) for multi-genre configurations. 069 In particular, these works (Sadat and Caragea,2022, 070 2024) state that linking phrases between sentences 071 are a consequence of their type of semantic relation072 ship, and that some linking phrases such as "thus" 073 and "therefore" show cause-effect relationships be074 tween sentences that most previous works did not 075 take into account. 076 However, limited research has been executed on 077 NLI in Spanish, with the exception of the Spanish 078 corpus included in XNLI. The first benchmark in 079 this area was created by generating sentence tem080 plates from a Spanish Question Answering corpus 081 (Peñas et al.,2006) but is relatively small having 082 1
2,962 examples and adopts a binary label for Entail083 ment. Despite being subjected to traditional mod084 els such as Support Vector Machines, the SPARTE 085 benchmark achieved an accuracy of 60.83 (Castillo, 086 2010). Another benchmark, INFERES (Kovatchev 087 and Taulé,2022), produced with human annota088 tion and exhibiting a three-way labeling technique, 089 scans across six domains extracted from Wikipedia. 090 It is rather more extensive with 8,055 examples and 091 includes a Contradiction category. Other studies 092 (Kauchak et al.,2019) have indicated that viable 093 results can be attained in medical domain text by 094 employing linking phrases to extract sentence pairs 095 and ensuing classification of sentence relationships. 096 Additionally, certain model-based approaches like 097 SILT (Huertas-Tato et al.,2022) and Facter-Check 098 (Martín et al.,2022) have demonstrated that BERT099 based models can achieve good classification per100 formance. However, these models may uninten101 tionally learn heuristics from annotation artifacts 102 (McCoy et al.,2019) and may even predict labels 103 by merely examining the hypothesis (Gururangan 104 et al.,2018).105 It is apparent from the existing Spanish NLI re106 search that there is significant room for improve107 ment. One avenue could be the creation of a larger 108 benchmark to facilitate the training of expansive 109 deep learning architectures. Ideally, such a bench110 mark would span multiple genres, thereby encom111 passing various writing styles and multitudinous 112 topics to enhance model generalization, akin to 113 MNLI and MSciNLI. Further improvements could 114 be made by including a label to account for causal 115 relationships, as exhibited in sciNLI and MSciNLI. 116 Not least, model evaluations should seek to mea117 sure vulnerability to learning artificial patterns 118 from the annotation artifacts in the benchmark in 119 addition to conventional classification metrics (Gu120 rurangan et al.,2018) (Naik et al.,2018). In light 121 of these considerations, we propose a novel Span122 ish NLI multi-genre dataset, ESNLIR, Reasoning123 Spanish-Natural-Language-Inference dataset, the 124 extraction methodology for which is based on 125 sciNLI. We apply BERT-based models to corpora 126 of multiple genres to establish a baseline of not 127 only classification metrics but also a method for 128 detecting potential annotation and model artifacts.129 With this in mind, this paper first defines the 130 dataset extraction, followed by the training and 131 evaluation results for the baseline and BERT-based 132 models, including insights into the model annota133 tion artefacts and stress tests.134 2 Dataset 135 The data collection process for ESNLIR is akin 136 to that outlined in sciNLI, focusing on searching 137 for linking phrases within extensive corpora. This 138 approach allows for the automatic extraction of 139 premise-hypothesis pairs, accompanied by a brief 140 label validation procedure to ensure a basic stan141 dard of data quality. In this section, we provide a 142 detailed description of this process, covering the 143 sources of corpora, genres, and cardinality. 144 2.1 Multi-Genre Corpora 145 Following the hypothesis from the MNLI dataset, 146 which suggests that incorporating multiple genres 147 enhances models’ generalization capabilities, we 148 selected 34 Spanish corpora to represent eight dis149 tinct writing genres: 150 • articles: Web articles and encyclopedia arti151 cles. 152 • books: Literature in general, consisting of 153 open access or public domain books. 154 • comments: Web comments on social net155 works, accounting for non-formal texts. 156 • legal: Legal documents from Colombia. 157 • clinical: Clinical cases extracted from open 158 source datasets. 159 • news: News articles extracted from multiple 160 sources. 161 • talks: Speeches extracted from TED, which 162 extend the non-formal texts group. 163 • theses: Theses extracted from the website 164 of the Universidad de Los Andes, Colombia, 165 with access for academic purposes. 166 The domain and genre of each selected corpora can 167 be found in Table 1.168 2.2 Premise-hypothesis extraction 169 For each corpus, an identical process is used to 170 extract premise-hypothesis pairs. Using a method171 ology similar to sciNLI, certain linking phrases are 172 identified as indicators of various semantic rela173 tionships between sentences, as demonstrated in 174 Table 2. Based on these criteria, for each document 175 in the corpus, sentences starting with any of the 176 linking phrases listed in Table 2are identified and 177 2
used as hypotheses in sentence pairs. To construct 178 these sentence pairs, the sentences that end just be179 fore the linking phrases and are located in the same 180 paragraph are taken as premises. Ultimately, the la181 bel for each premise-hypothesis pair is determined 182 based on the label allocated to the linking phrase in 183 Table 2. In executing this process, both premises 184 and hypotheses for examples from the contrasting, 185 entailment, and reasoning categories are extracted 186 from the same paragraph. Conversely, neutral ex187 amples are created by pairing sentences that do not 188 fall into the aforementioned categories and origi189 nate from different paragraphs. Specifically, there 190 are multiple methods to create neutral examples:191 • Both random: Pairing of sentences not previ192 ously used as premises or hypotheses in the 193 other classes.194 • First random: The first sentence is a previ195 ously used premise and the hypothesis is a 196 random, previously unused sentence.197 • Second random: The first sentence is a previ198 ously unused random sentence and the second 199 is a previously used hypothesis.200 • Both entailed: Both premise and hypothesis 201 have been used previously in other pairs, but 202 belong to different paragraphs in the same 203 document.204 Since the neutral class has a completely different 205 extraction method, which mainly differs in taking 206 text fragments from different paragraphs and not 207 using linking phrases, it is expected that the classi208 fiers will have some bias towards this class, causing 209 it to have better classification metrics than the oth210 ers.211 To illustrate the process, in the following para212 graph taken from Wikipedia, Sin embargo is a 213 linking phrase belonging to the contrasting class:214 Text in Spanish:215 "Formado en el diario El País de Madrid, fundó 216 y dirigió el periódico Siglo 21 de Guadalajara 217 desde su fundación en 1991 hasta el 30 de abril de 218 1997, día en el que renunció, para crear el diario 219 Público. Un día antes de que Zepeda anunció el 220 primer número de Público, hubo una operación que 221 buscó cerrar a Siglo 21, mediante la salida masiva 222 de sus empleados.Sin embargo,Siglo 21 sobre223 vivió a ese intento de desaparecerlo y, finalmente, 224 Zepeda se vio obligado a vender 66.66 por ciento 225 de Público a los propietarios de Grupo Multime226 dios, (que posteriormente renombraron a Público 227 como Milenio Jalisco, denominación que mantiene 228 hasta la fecha). En 1999 deja Público para asumir 229 la subdirección de El Universal en la Ciudad de 230 México, cargo que desempeñaría hasta 2001." 231 Text in English "Trained at El País newspaper 232 in Madrid, he founded and directed the newspaper 233 Siglo 21 in Guadalajara from its founding in 1991 234 until April 30, 1997, the day he resigned to cre235 ate the newspaper Público. A day before Zepeda 236 announced the first issue of Público, there was an 237 operation that sought to close Siglo 21, through 238 the massive departure of its employees.However, 239 Siglo 21 survived that attempt to disappear it and, 240 finally, Zepeda was forced to sell 66.66 percent of 241 Público to the owners of Grupo Multimedios, (who 242 subsequently renamed Público as Milenio Jalisco, 243 a denomination it maintains to date). In 1999, he 244 left Público to take over as assistant editor of El 245 Universal in Mexico City, a position he held until 246 2001." 247 In this context, the sentence highlighted in blue 248 can be considered as the premise, and the sentence 249 highlighted in red as the hypothesis. The premise 250 discusses an operation aimed at closing the mag251 azine Siglo 21, whereas the hypothesis contrasts 252 this by stating that the magazine ultimately sur253 vived. 254 After extraction, sentences are refined by re255 taining only those that are syntactically complete. 256 To achieve this, Part-of-Speech (POS) tagging 257 is performed on each sentence using the Spacy 258 es_core_news_lg model. Sentences are preserved 259 only if they contain both a subject (with tags such 260 as "NOUN", "PRON", "PROPN") and a predicate 261 (with tags such as "AUX", "VERB"). 262 2.3 Train-Val-Test split 263 In order to facilitate the evaluation of the dataset, 264 for each corpus all splits are extracted in a balanced 265 way. In addition, to avoid sharing pairs between 266 splits, a double stratification strategy is used to 267 generate the splits: 268 1. Separate articles into groups, one for each 269 split, to avoid shared articles between splits. 270 First a training and validation vs testing split 271 is made, then a training vs validation split 272 completes the grouping. 273 2. For each split find the class with the minimum 274 number of examples and downsample the rest 275 3
of the classes to have the same number of 276 examples.277 As shown in table 1, a maximum of 15000 exam278 ples have been selected for testing and validation 279 splits, but the balance of classes is maintained for 280 all splits. However, there is some imbalance in the 281 genre representation, which is caused by the dif282 ference in the number of articles found in each of 283 the corpora. Apart from that, some corpora were 284 selected to be used in the test only, in order to test 285 if the dataset offers some form of domain general286 isation. In particular, the following corpora were 287 selected for testing only for the following reasons:288 • esbooks__eswikibooks and es289 books__libreriaunal: It is expected that 290 the semantic relations learnt in other book 291 corpora, i.e. from traditional literature, 292 can help to learn fenomena in web books 293 and college books such as those found in 294 wikibooks and libreria unal respectively.295 • escomments__reddit and escom296 ments__suicide: It is expected that relations 297 learned from Twitter comments will help to 298 learn behaviours in these corpora.299 • eslegal__entrevistas_comision_verdad: Since 300 these interviews are informal, it is interesting 301 to check whether the models are able to learn 302 this type of relations from formal sources on 303 the same topics.304 • esmedical__BVS and esmedical__SPACC 305 and esmedical__TEI_ES: As the training split 306 contains information about medical theses, it 307 is expected that the models can generalise to 308 these corpora.309 • esnews__eswikinews: Similar to es310 books__eswikibooks, it is expected that 311 learning from traditional news sources can 312 help to predict web news.313 • estalks__TED: It is expected that a combina314 tion of non-formal sources such as twitter and 315 all the domains presented in the rest of the 316 training corpora can help to predict relations 317 in sentences belonging to talks.318 Finally, for each split all the corpora examples are 319 merged into a single split, resulting in a single 320 dataset with 7’325.356 training examples, 127.404 321 validation examples, and 128.412 test examples.322 3 Experimental setup 323 This section describes the experimental setup used 324 to train and evaluate the performance of several 325 models, and the dataset in general. 326 3.1 Models 327 The following models have been selected to provide 328 a performance baseline for further research: 329 • XGBoost (Chen and Guestrin,2016): This 330 is an ensemble of three models used as a 331 baseline for lexical representations, in fact a 332 Bag Of Words representation, to determine 333 whether lexical representations are sufficient 334 to classify the examples. For this model, 2000 335 estimators were used. 336 • BERTIN (la Rosa y Eduardo G. Pon337 ferrada y Manu Romero y Paulo Ville338 gas y Pablo González de Prado Salas y 339 María Grandury,2022): A BERT-based ar340 chitecture that removes the next sentence pre341 diction task. Specifically, it uses the Span342 ish pre-trained model bertin-project/bertin343 roberta-base-spanish from huggingfaces. 344 • XLMRoBERTa (Conneau et al.,2019): A 345 multilingual version of RoBERTa to test if 346 our dataset performs well on multilingual con347 figurations. This pre-trained model is also 348 extracted from huggingfaces, with the path 349 FacebookAI/xlm-roberta-base. 350 3.2 Metrics 351 As the test dataset is balanced, the accuracy and 352 f1_score macro are used to evaluate the perfor353 mance of the models. 354 3.3 Training and Fine-Tuning 355 The XGBoost model was trained from scratch 356 using a Bag Of Words representation, while the 357 BERT-based models were fine-tuned from the pre358 trained models up to a maximum of 6 epochs, with 359 early stopping via the f1_score metric, a batch size 360 of 64 and a learning rate of 2e-5. 361 3.4 Dataset artifact detection 362 To detect dataset artifacts the strategy applied in 363 (Gururangan et al.,2018) is used, this means that 364 for each one of the family of models mentioned 365 previously a model will be trained and evaluated 366 using only the premise of the sentence pairs. 367 4
dataset train|unique_articles train|total_examples val|unique_articles val|total_examples test|unique_articles test|total_examples esarticles__eswiki 75004 254756 1907 5976 2051 6568 esbooks__crawling 75122 2653328 209 6848 197 7448 esbooks__elchico 582 9600 69 696 91 1828 esbooks__eltec 37 620 10 124 8 100 esbooks__eswikibooks 0 0 0 0 135 672 esbooks__gutenberg 455 7404 44 776 55 904 esbooks__libreriaunal 0 0 0 0 402 6588 escomments__reddit 0 0 0 0 2009 2860 escomments__suicide 0 0 0 0 25 36 escomments__tweets 52573 66708 6463 8176 6504 8244 eslegal__entrevistas_comision_verdad 0 0 0 0 8 68 eslegal__informes_analisis_comision_verdad 3 52 1 4 1 4 eslegal__raw_sentencias_corte_colombia 18020 249424 385 7020 407 5556 esmedical__BVS 0 0 0 0 12 16 esmedical__TEI_ES 0 0 0 0 89 180 esnews__colombian_news 8043 18272 1134 2268 1140 2264 esnews__eswikinews 0 0 0 0 7 12 esnews__spanish_pd_news 207471 858268 1971 7460 1991 7972 estalks__TED 0 0 0 0 156 548 estheses__Centro Interdisciplinario de Estudios sobre Desarrollo 320 5460 26 380 26 520 estheses__Escuela de Gobierno Alberto Lleras Camargo 232 3308 21 276 22 440 estheses__Facultad de Administración 622 11664 56 944 53 696 estheses__Facultad de Arquitectura y Diseño 1204 12700 139 1444 131 1444 estheses__Facultad de Arte y Humanidades 623 11120 78 1200 63 912 estheses__Facultad de Ciencias 985 16036 123 1960 105 2224 estheses__Facultad de Ciencias Sociales 1676 50872 188 5544 207 6452 estheses__Facultad de Derecho 1600 32412 199 3908 181 3968 estheses__Facultad de Economía 1201 17828 149 2112 150 1980 estheses__Facultad de Educación 671 16264 73 2152 75 1944 estheses__Facultad de Ingeniería 5660 83284 503 7236 493 6904 estheses__Facultad de Medicina 19 132 4 16 4 28 estheses__Unknown 507 8456 65 880 66 836 Table 1: Corpora split example count. Some of the corpora is out-of-sample, meaning that it does not have examples in training and validation splits 5
Class Description Linking phrases contrasting The hypothesis contradicts or mentions a comparison, criticism, juxtaposition, or a limitation of something said in the premise. "sin embargo," "no obstante," "por otra parte," "por otro lado," "en cambio," "por el contrario," "al contrario," "en contraste," entailment The hypothesis generalizes, specifies or has an equivalent meaning with the premise. "en concreto," "concretamente," "especificamente," "precisamente," "en particular," "particularmente," "en especial," "es decir," "en otras palabras," "dicho de otra manera," "dicho de otro modo," "en otros terminos," "de hecho," "esto es," "o sea," "mejor dicho," "sobre todo," "justamente," "en resumidas cuentas," "en resumen," "en breve," "por ejemplo," "en sintesis," "en efecto," "en pocas palabras," "en una palabra," "recapitulando," "brevemente," "recogiendo lo mas importante," "como se ha dicho," "para ilustrar," neutral Premise and hypothesis are semantically independent of each other. reasoning The premise presents the reason, cause, or condition for the result or conclusion made in the hypothesis. "por lo tanto," "por tanto," "en consecuencia," "por consiguiente," "por ende," "por esa razon," "por eso," "de ahi que," "como resultado," "como consecuencia," Table 2: Classes and linking phrases used to group different relation types between premise and hypothesis pairs. 3.5 Stress tests368 To evaluate the robustness of the selected models 369 when training with the dataset, four of the strategies 370 for stress test generation from (Naik et al.,2018) 371 are used:372 • Length mismatch: Make the premise a lot 373 longer than the hypothesis, by adding the 374 expression: y verdadero es verdadero y ver375 dadero es verdadero y verdadero es verdadero 376 y verdadero es verdadero y verdadero es ver377 dadero, which does not alter the premise 378 meaning.379 • Negation: Add negation to the hypothesis 380 without altering the meaning with the expres-381 sion: y falso no es verdadero.382 • Overlap: Add the expression: y verdadero es 383 verdadero to the hypothesis to generate a word 384 mismatch between premise and hypothesis.385 • Spelling: Misspell a word in the premise.386 For each strategy a parallel test split is generated 387 from the original test split.388 3.6 Human validation389 To ensure that ESNLIR is useful to other re390 searchers, and given the current budget constraints, 391 a small balanced proportion of 2000 randomly se392 lected pairs from the dataset are annotated by a 393 group of 27 students who are asked to assign a sin394 gle label to each pair. Only those pairs where the 395 original label matches the majority of human labels 396 are retained. 397 Once the gold standard pairs have been extracted, 398 prediction is performed using XLMRoBERTa and 399 classification metrics (f1_score and accuracy) are 400 applied across the dataset and per genre to deter401 mine if the dataset would benefit from human an402 notation. 403 4 Results 404 4.1 Test performance 405 The Table 3presents the performance of the mod406 els on the test set. When compared to a majority 407 class baseline, the XGBoost model shows a per408 formance increase of 10 points. This suggests that 409 the presence of certain words can indicate the class 410 of a pair. As illustrated in Image 1, words such 411 as caso and si are frequently associated with the 412 contrasting class. 413 Evaluating BERT-based models reveals that their 414 semantic representations for sentence pairs are su415 perior, nearly doubling the performance of the XG416 Boost baseline. Notably, XLMRoBERTa, the mul417 tilingual model achieves the best overall perfor418 mance, which could be attributed to the larger size 419 of the dataset used for its pre-training. 420 model accuracy f1_score Majority class 0.25 0.25 XGBoost 0.35 0.348 BERTIN 0.663 0.664 XLMRoBERTa 0.676 0.676 Table 3: Performance of the baseline models in test 6
Figure 1: Wordcloud for contrasting class In-depth analysis of the classes reveals that for 421 BERT-based models, the ’neutral’ class is the sim-422 plest to identify. This indicates that distinguish423 ing the lack of semantic relationships between sen424 tences is more straightforward than identifying spe425 cific types of semantic relationships. On the other 426 hand, the selection method for this class is different 427 from the others, by not using linking phrases for 428 pair detection, and instead selecting sentences from 429 different paragraphs, which generates a bias that is 430 favourable for the classificaiton of this class. For 431 instance, the following pair is easily identifiable 432 as ’neutral’ because there are no shared entities, 433 and the first sentence discusses a business associa434 tion, whereas the second addresses processes for 435 opportunities:436 • Premise: ’Dentro de esta definición aparece 437 el concepto de asociaciones de negocios, en 438 las que una organización puede ser miembro 439 de muchas organizaciones virtuales’ (English: 440 Within this definition appears the concept of 441 business partnerships, in which an organiza442 tion can be a member of many virtual organi443 zations.)444 • Hypothesis: ’que no había un proceso for445 mal de localización de las oportunidades, sim446 plemente se cantaban las oportunidades y se 447 “peleaban” las mismas entre los segmentos’ 448 (English: that there was no formal process 449 for locating opportunities, opportunities were 450 simply sung about and “fought over” between 451 segments)452 Semantic relationships can be classified into var453 ious categories, among which contrasting is the 454 second easiest to identify. This may be the case 455 because contrasting sentences often involve similar 456 entities but contain predicates that are opposites. 457 For instance, consider a pair of sentences where 458 both discuss ongoing research. The first sentence 459 may state that the research reaches the state-of460 the-art level, while the second might indicate that 461 additional research is necessary: 462 • Premise: ’cabe resaltar que estos resultados se 463 encuentran de acuerdo con lo que se reporta 464 en la literatura’ (English: It should be noted 465 that these results are in agreement with what 466 is reported in the literature.) 467 • Hypothesis: ’una investigación más detallada 468 es requerida’ (English: a more detailed inves469 tigation is required) 470 Lastly, the most challenging categories are rea471 soning and entailment, which models often confuse 472 with each other more than with other categories, 473 as demonstrated in Figure 2. These confusions 474 may occur because some entailment pairs exhibit 475 specifications or generalizations that might be in476 terpreted as causal relationships, and vice versa. 477 For instance, consider the following two sentences 478 discussing the policies of the Colombian political 479 party UP. The first sentence addresses how these 480 policies failed to consider real-life changes in the 481 country’s situation, while the second sentence indi482 cates the party’s failure to acknowledge that their 483 policies were not sustainable in the long term: 484 • Premise: ’en particular, el programa de la UP 485 y sus políticas no prestaban atención al tipo de 486 cambio real como determinante de la posición 487 competitiva internacional del país’ (English: 488 In particular, the UP program and its policies 489 did not pay attention to the real exchange rate 490 as a determinant of the country’s international 491 competitive position.) 492 • Hypothesis: ’la UP no reconoció que sus 493 políticas no serían sostenibles en el mediano 494 plazo y que las limitaciones de capacidad 495 se convertirían en un obstáculo insuperable 496 para un rápido crecimiento’ (English: the UP 497 failed to recognize that its policies would not 498 be sustainable in the medium term and that 499 capacity constraints would become an insur500 mountable obstacle to rapid growth.) 501 The relationship can be unclear without proper con502 text. For some individuals, the pair relates to rea503 soning because the failure of the UP to consider 504 real-life changes makes it difficult for them to per505 ceive their policies as unsustainable. For others, 506 the pair is associated with entailment because not 507 7
model contrasting entailment neutral reasoning XGBoost 0.345 0.436 0.278 0.341 BERTIN 0.683 0.626 0.669 0.675 XLMRoBERTa 0.664 0.676 0.695 0.667 Table 4: Accuracy per class for baseline models in test model contrasting entailment neutral reasoning XGBoost (both sentences) 0.345 0.434 0.287 0.335 BERTIN 0.419 0.411 0.261 0.452 XLMRoBERTa 0.441 0.409 0.222 0.485 Table 5: Accuracy training and evaluating with only the premise stress test contrasting entailment neutral reasoning test 0.664 0.676 0.695 0.667 test_length_mismatch 0.513 0.524 0.755 0.654 test_negation 0.580 0.575 0.744 0.656 test_overlap 0.624 0.637 0.719 0.665 test_spelling 0.687 0.646 0.677 0.659 Table 6: Stress test accuracy per class for XLMRoBERTa acknowledging the unsustainability of the policies 508 can be seen as a result of disregarding the actual 509 changes in the country’s situation.510 Looking at the details of the correctly classified 511 reasoning pairs, it seems that both sentences have 512 nouns that are related through a process of trans513 formation, which may imply a causal relationship. 514 For example, in the following pair, the first sen515 tence talks about the cremation of a group of peo516 ple, while the second sentence talks about people 517 collecting their ashes because they believed they 518 were saints. In this case, ’quemados’ (burned) is 519 transformed into ’cenizas’ (ashes):520 • Premise: ’esa misma tarde fueron quema521 dos’(English: that same afternoon they were 522 burned)523 • Hypothesis: ’la gente quiso coger sus cenizas 524 creyendo que eran las reliquias de unos santos, 525 y por eso, para acabar con el mito, Felipe 526 hizo que tiraran las cenizas al Sena’ (English: 527 people wanted to take their ashes believing 528 that they were the relics of saints, and so, to 529 put an end to the myth, Philip had the ashes 530 thrown into the Seine.)531 4.2 Dataset artifact detection532 In a repeated training and evaluation setup, it was 533 observed that the dataset largely lacks annotation 534 artifacts. As demonstrated in Table 5, the accu535 racy of models trained and evaluated using only the 536 Figure 2: Confusion matrix for XLMRoBERTa premise is nearly comparable to the XGBoost base537 line, yet approximately 25 points lower than the 538 original models. This indicates that both sentences 539 are necessary for the models to effectively identify 540 the type of relationship in a significant portion of 541 the test set. Some performance in this setup may 542 be influenced by unique words within each class, 543 including examples such as misspellings, foreign 544 language words, and entity names: 545 • contrasting: ’tatuajes’, ’finiquito’, ’imnedi546 ata’, ’generation’, ’exmarido’, ’Songkhram’ 547 • entailment: ’impresionada’, ’Leguízamo’, 548 ’Computacional’, ’gallinero’, ’liviano’ 549 • neutral: ’problematizarse’, ’palurdos’, ’inde550 pendizados’, ’Barbusse’, ’Riopaila’ 551 8
• reasoning: ’aeronáuticos’, ’primitivamente’, 552 ’mitosis’, ’larson’, ’Marseille’, ’Marimbera’553 4.3 Stress tests554 The transformations applied to each stress test in 555 Table 6affect the classes in different ways. While 556 spelling has minimal impact on performance, the 557 contrasting and entailment classes are signifi558 cantly influenced by the length mismatch and 559 negation tests. This indicates that sentence length 560 and the presence of negations in the hypothesis sub561 stantially affect the discrimination criteria for these 562 classes.563 4.4 Genre performance564 The evaluation of performance by genre, as de565 picted in Table 7, indicates that accuracy does not 566 correspond directly to the number of examples in 567 the datasets. Interestingly, the second-best perform568 ing genre lacked any examples in the training split, 569 suggesting that the dataset enables models to gen570 eralize to out-of-domain contexts. Nonetheless, the 571 datasets with the poorest performance are associ572 ated with non-formal genres, such as comments 573 and talks. This could be because non-formal writ574 ing lacks the rigid semantic rules that other genres 575 possess, which limits their potential for generaliza576 tion.577 genre train|total_examples accuracy f1_score clinical 0 0.735 0.730 legal 253592 0.727 0.727 books 3021904 0.684 0.684 articles 254756 0.679 0.678 theses 269536 0.674 0.674 news 877096 0.671 0.672 comments 100956 0.647 0.646 talks 0 0.597 0.595 Table 7: Performance per genre for XLMRoBERTa 4.5 Out of domain performance578 In Table 8, it is evident that the highest out-of579 sample performance is observed in formal writing 580 genres. In contrast, the lowest performance is noted 581 in non-formal writing genres. This supports the 582 notion that the highly variable semantic rules in 583 these genres negatively impact their performance.584 4.6 Human validation results585 27 students were asked to annotate the 2000 val586 idated examples. To facilitate the annotation pro587 cess, a LabelStudio instance containing the dataset 588 corpus accuracy f1_score esmedical__TEI_ES 0.744 0.740 esbooks__libreriaunal 0.677 0.677 esbooks__eswikibooks 0.652 0.653 escomments__suicide 0.639 0.626 esmedical__BVS 0.625 0.598 estalks__TED 0.597 0.595 esnews__eswikinews 0.583 0.592 escomments__reddit 0.579 0.577 eslegal__entrevistas_comision_verdad 0.515 0.513 Table 8: Out-of-domain performance for XLMRoBERTa was made available in the cloud, where each anno589 tator was instructed to select a label for a set. This 590 annotation process resulted in each example having 591 more than one human label, and then only those 592 examples where the original label matched the ma593 jority of human labels were retained. This resulted 594 in an unbalanced dataset of 974 examples, with the 595 number of labels shown in tables 9and 10, where 596 all 8 original genres are included. The reduction to 597 less than half the original sample size may mean 598 that for a large part of the dataset, the label defined 599 by the linking phrase does not match the label that 600 a human would use. There are three possible rea601 sons for this, which require further analysis: lack 602 of domain knowledge, lack of clear instructions, or 603 incorrect use of linking phrases at the time the pair 604 source was written. 605 class number of examples contrasting 184 entailment 219 neutral 362 reasoning 207 Table 9: Validated dataset class count genre number of examples books 217 legal 71 clinical 51 theses 427 articles 34 comments 67 news 71 talks 34 Table 10: Genre distribution in the validated dataset Focusing on the metrics after pruning the dataset, 606 it can be seen in Table 11 that the human annotation 607 9