Full text
id182832 EFFICIENT FINETUNING STRATEGIES FOR MULTILINGUAL NEURAL MACHINE TRANSLATION GERARD SÁNCHEZ I MALTAS Thesis supervisor: CARLOSESCOLANOPEINADO(DepartmentofComputerScience) Degree:Master'sDegreeinDataScience Master's thesis Facultat d'Informàtica de Barcelona (FIB) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech 25/01/2024
i Acknowledgements I would like to express my sincere gratitude to my thesis supervisor, Carlos, for his guidance, support, and encouragement throughout my research. His expertise and insights have been invaluable to me, and I am truly grateful for the opportunity to have worked with him. I would also like to extend my deepest gratitude to Gerard, for his helpful feedback and suggestions on my thesis. His expertise and input greatly improved the quality of my work. Finally, special thanks to my family and friends who played an instrumental role in the development and completion of this thesis. Your support, encouragement, and assistance throughout the writing have been priceless to me.
ii Abstract Driven by the ambition to eliminate language barriers worldwide, Machine Translation has become a central area of interest in today’s artificial intelligence research. Despite significant advancements, the concentration of research and resources has been predominantly on a highresource languages. This discrepancy in language coverage points out a critical gap in the field. Recent breakthroughs in Machine Translation have seen the emergence of Multilingual Large Pre-Trained models, which have set new benchmarks across the field by enabling low-resource languages benefit from zero-shot translation. However, these models obtain high performances at the cost of requiring huge amounts of data and hardware resources. The focus of this thesis is to explore and formulate a fine-tuning strategy for a multilingual machine translation model, such as M2M100. Specifically, the project aims to extend their linguistic capabilities by incorporating new low-resource languages by fine-tuning specific language adapters using Low-Rank Adaptation methods. To evaluate the performance of the strategy, state-of-the-art techniques and evaluation metrics are employed, considering factors like scalability, catastrophic forgetting and zero-shot translation. The implemented approach has successfully developed a M2M100 language translator in a low-resource context resulting in a SacreBLEU score of 5.6 when just training a 13 % of the parameters while the full fine-tuning methodology reach a score of 7.43. Furthermore, this framework demonstrates its capability to produce more advanced and efficient machine translation models, which can deliver high-quality translations with reduced computational demands. Keywords – Machine Translation, Multilingual Models, M2M100, Fine-Tuning, LoRA
Contents iii Contents 1 Introduction 1 1.1 Background and Context ................................ 1 1.2 Problem Statement ................................... 2 1.3 Objective and Scope .................................. 2 1.4 Significance and Contribution of the Study ...................... 3 1.5 Overview of the Thesis ................................. 3 2 Literature Review 5 2.1 Sequence-to-Sequence Models ............................. 5 2.2 Tranformer Architecture ................................ 7 2.2.1 Self-Attention .................................. 9 2.2.2 Multi-Head Self-Attention ........................... 11 2.2.3 Masked Self-Attention ............................. 12 2.2.4 Positional Encoding .............................. 12 2.2.5 Point-Wise Feed-Forward Network ...................... 13 2.2.6 Residual Connections & Normalization .................... 13 2.2.7 Final Linear & Softmax Layer ......................... 14 2.3 Neural Machine Translation .............................. 14 2.4 Automatic Evaluation of Machine Translation .................... 15 2.4.1 SacreBLEU ................................... 15 2.4.2 COMET ..................................... 17 2.5 Multilingual Machine Translation ........................... 18 2.6 Zero-shot Translation .................................. 19 2.7 M2M100 ......................................... 19 2.7.1 SentencePiece Tokenizer ............................ 20 2.7.2 Architecture ................................... 20 2.8 Fine-Tuning ....................................... 21 2.8.1 Adapters ..................................... 22 2.8.2 LoRA ...................................... 24 2.8.3 LoHa ...................................... 26 2.8.4 LoKr ....................................... 27 3 Experimental Framework 28 4 Results 31 4.1 Supervised Translation ................................. 32 4.2 Data Scalability ..................................... 40 4.3 Catastrophic Forgetting ................................ 40 4.4 Zero-Shot Translation ................................. 41 5 Conclusions and Future Work 43 References 47
List of Figures iv List of Figures 2.1 Sequence-to-Sequence General Schema. ........................ 5 2.2 Transformer Architecture. ............................... 8 2.3 Self-Attention Module. ................................. 10 2.4 Multi-Head Self-Attention Module. .......................... 11 2.5 Feature-Based Fine-Tuning. .............................. 21 2.6 Adapters Fine-Tuning. ................................. 22 2.7 Full Fine-Tuning. .................................... 22 2.8 ∆WLoRA Reparametrization. ............................ 25 2.9 ∆WLoHA Reparametrization. ............................ 26 2.10 ∆WLoKr Reparametrization. ............................. 27 4.1 Percentage of Trainables Parameters vs. Percentage of Full Fine-Tuning in Experiments with α= 2r................................ 34 4.2 SacreBLEU Comparison: Training Attention Weights vs. Attention, Fully Connected Layers, and Head with Frozen Embeddings and Average Initialization. 35 4.3 SacreBLEU Comparison: Training Attention Weights vs. Attention, Fully Connected Layers, and Head with Trained Embeddings and Average Initialization. 35 4.4 SacreBLEU Comparison: Training Attention Weights vs. Attention, Fully Connected Layers, and Head with Frozen Embeddings and Random Initialization. 36 4.5 SacreBLEU Comparison: Training Attention Weights vs. Attention, Fully Connected Layers, and Head with Trained Embeddings and Random Initialization. 36 4.6 Percentage of Trainable Parameters vs. Percentage of Full Fine-Tuning in Each Experiment ....................................... 37
List of Tables v List of Tables 3.1 Training Configuration Parameters .......................... 29 3.2 Language Selection Choices .............................. 30 4.1 Evaluation of NLLB and Google Translation Using the SacreBLEU Metric on the FLORES Dataset. ................................... 31 4.2 Comprehensive Analysis of M2M100 Full Fine-Tuning Across Diverse Dataset Sizes. 32 4.3 SacreBLEU Results Using the FLORES Dataset for Attention Weight Training and Frozen Embeddings. ................................ 32 4.4 SacreBLEU Results Using the FLORES Dataset for Attention Weight and Embedding Training. .................................. 33 4.5 Results from Training Specific Transformer Components of M2M100. ....... 38 4.6 Results for Different LoHa Ranks ........................... 39 4.7 Results for Different LoKr Ranks. ........................... 39 4.8 Selected Fine-Tuning Strategy Parameters. ...................... 39 4.9 SacreBLEU Results for LoRA Training with Three Different Datasets. ...... 40 4.10 SacreBLEU Results for M2M100 Languages Translated to Basque. ........ 41 4.11 SacreBLEU Results for Directions M2M100 Languages to Basque. ......... 41
1 1 Introduction 1.1 Background and Context In the field of artificial intelligence, neural networks have become fundamental, demonstrating innovative capabilities in a variety of tasks, including Natural Language Processing (NLP). These models outperform at learning complex patterns and representations from large datasets. Among the diverse applications, Neural Machine Translation (NMT) stands out as a transformative technology, that improves communication across language barriers. Conceived as computational systems that translate text from one language to another, Machine Translation (MT) has been around since the 1940s, but its recent migration from statistical modeling to neural systems has pushed the technology to new frontiers enabling people all over the world to communicate, work, travel and access information easily and automatically. Nevertheless, there is a significant variation in the efficacy of this approach across different languages. Even though languages with large quantities of digital data (high-resource languages) have seen major improvements, languages with less digital data available (low-resource languages) remain challenging to get effective translation systems. The expansion of machine translation into low-resource languages faces significant technical challenges due to the high cost and logistical difficulties in obtaining training data. Without sufficient training data, standard techniques such as large-scale pre-training and model scaling, may not be suitable for upcoming needs and modern techniques. To address these challenges in low-resource translation, considerable emphasis has been placed on the use of multilingual models. These models, aim to build a single model to translate between any pair of languages. Multilingual translation models benefit from cross lingual transfer sharing information between similar languages, which benefits low-resource directions and enables translation between language pairs that have never been explicitly seen during training. For example, if a model has been trained on English-to-French and English-to-Spanish translations, it may also be able to translate between French and Spanish, even though it was never explicitly trained on a French-to-Spanish corpus. This capability arises from the model learning a deep, abstract representation of language that is not tied to any specific language pair. This is called zero-shot translation and is particularly valuable for low-resource languages as it allows more flexible and extensive use of these models extending their usefulness to a wider and rare range of language pairs.
1.2 Problem Statement 2 Beyond the development of multilingual models, a crucial strategy in advancing machine translation is the application of transfer learning, particularly through the fine-tuning of large pre-trained models. Fine-tuning involves adjusting a pre-trained model on a specific, often smaller, dataset to improve its performance for particular tasks or languages, proving especially beneficial for low-resource languages with limited data. This approach enables the model to improve translation accuracy in some fields or languages by reducing the resources needed for training from scratch. However, while fine-tuning offers significant benefits, it also presents a challenge in terms of resources. With numerous downstream tasks, the process becomes resource-intensive, increasing the demand for computational power and complicating model management. This inefficiency calls for innovative solutions, such as employing techniques that share parameters across tasks or developing more adaptive models that can handle multiple tasks with minimal adjustments. 1.2 Problem Statement Recently, large pre-trained models have become state-of-the-art as they obtain high performances. Despite the progress made in Neural Machine Translation, there is still an important requirement for more efficient and language-specific fine-tuning strategies: the cost of adjusting numerous parameters requires huge amounts of data and hardware resources. This situation highlights the importance of developing strategies to efficiently fine-tune multilingual neural machine translation models. A particular focus is needed on integrating new languages, that large pre-trained models have not been seen before, especially in the case of low-resource languages. The key challenge is to formulate a strategy that minimizes the number of learnable parameters while still achieving an outcome as effective as training with the full set of parameters. 1.3 Objective and Scope The primary objective of this master thesis is to evaluate the viability and effectiveness of LoRA as a fine-tuning method for encoder-decoder architectures. We will focus on M2M100 model in the specific context of machine translation from English to Basque. The study aims to assess the ability of LoRA to enhance translation quality while minimizing the number of trainable parameters, thus optimizing computational efficiency. Additionally, exploring extensions of LoRA, such as LoHa and LoKr presents an opportunity to further improve the adaptability and performance of the fine-tuned models.
2.2 Tranformer Architecture 9 scores for multiple linear projections of the input, providing a more comprehensive understanding of the relationships between words. 2. Feed-Forward Neural Network (FFNN): fully connected layer which is applied to each position of the output Z from the Multi-Head self-attention sub-layer. This sub-module enables the extraction of position-dependent and non-linear features of the previous representation. Each of these two sub-layers are preceded by a residual connection and followed by a normalization layer. • Decoder Component: Stack of decoders, like in the encoder component, formed of 6 decoder layers. Each layer comprises the following components: 1. Masked Multi-Head self-attention layers: The same mechanism used in the encoder, however, the information about future tokens is hidden from the model by setting attention scores to −∞ . This self-attention layer is only allowed to attend to earlier positions in the output sequence. 2. Encoder-Decoder attention: The procedure is similar to a multi-head self-attention layer, nevertheless, it conducts attention over the input of the masked multi-head self-attention and the output of the encoder to assist the decoder in focusing on relevant parts of the input sentence. 3. Feed-Forward Neural Network (FFNN). Identical to the encoder, preceding residual connections and subsequent normalization layers are placed in the components introduced in the decoder. • Linear & Softmax Layers: Responsible for transforming the ultimate vector generated by the decoder stack into words. The Linear layer is a simple fully connected neural network that projects the last vector into a logits vector with vocabulary size. Each position of the vector corresponds to the score of each word learned from the training dataset. Finally, the softmax layer converts scores into probabilities. Then, any technique explained Equations in 2.3,2.5, or 2.6 can be used to infer the next token. 2.2.1 Self-Attention Self-attention is a mechanism that enables neural networks to weigh the significance of different words in a sequence against each other. It is a key component of transformer architectures, which have been highly successful in NLP tasks. The self-attention mechanism allows the model to
2.2 Tranformer Architecture 10 consider the relationships between all words in a sequence simultaneously, capturing dependencies regardless of their distance apart. In the context of transformers, self-attention is applied to both the encoder and decoder layers. Figure 2.3: Self-Attention Module. Let’s focus on the self-attention mechanism within one layer of a transformer. The self-attention mechanism can be described as follows: 1. Query (Q), Key (K), Value (V) matrices: The input representation X∈Rn×dmodel is multiplied by the trained weight matrices WQ∈Rdmodel×dk , WK∈Rdmodel×dk , WV∈ Rdmodel×dv to get the Query matrix Q∈Rn×dk , the Key matrix K∈Rn×dk and the Value V∈Rn×dv matrix. These matrices are different representations of the input matrix X and they are employed to evaluate its own similarity using the Scaled Dot-Product. 2. Attention Score or "Scaled Dot-Product Attention" (reference): An attention function can be described as mapping a query and a set of key-value pairs to an output. Each token of the input sentence is compared to the other tokens to determine how much focus to place on other parts of the input sentence as we encode a word at a certain position. The matrix of outputs Z∈Rn×dvis computed as Z=Attention(Q, K, V ) = softmax(Q·KT √dk )·V(2.7) where after the computation of the dot product of the queries with all keys ( Q·KT ), it is divided by √dk to avoid large dot products pushing the softmax into regions where it has
2.2 Tranformer Architecture 11 extremely small gradients. Afterwards, the result is normalized using the softmax function which determines how much each word will be expressed at this position. The last step is to sum up the weighted value vectors to produce the output of the self-attention layer. 2.2.2 Multi-Head Self-Attention Multi-head self-attention is a powerful extension of the traditional self-attention mechanism, designed to enhance the model’s ability to capture diverse patterns within an input sequence. In this approach, multiple attention heads operate in parallel, independently introducing diversity and enriching the model’s comprehension. Figure 2.4: Multi-Head Self-Attention Module. Unlike a single attention head, multi-head attention enables the model to jointly attend to information from different representation subspaces at various positions. This contrasts with a single attention head, where averaging tends to inhibit such joint attention. The mechanism achieves this through the concatenation and projection of h heads using the matrix W0∈ Rh·dv×dmodel H=MultiHead(Q, K, V ) = Concat(Z1, . . . , Zh)·W0(2.8) where Zi = Attention ( Qi, Ki, Vi )with Qi = X·Wi Q , Ki = X·Wi K and Vi = X·Wi V having i∈(1, h) Each set of attention heads is randomly initialized during training. Post-training, these sets play a crucial role in projecting input embeddings from lower encoders/decoders into different representation subspaces. This comprehensive approach enhances the model’s understanding of
2.2 Tranformer Architecture 12 the input sequence, promoting richer and more nuanced representations. In [1], the dimensionality of the matrices are dmodel = 512,h= 8 and dk=dv=dmodel h= 64 2.2.3 Masked Self-Attention It might be required to eliminate attention links between certain word pairs. For instance, the decoder at token position t should not be able to attend to token position t + 1, as it would see the answer when training. This can be achieved by introducing a mask matrix in the mutli-head self-attention module in the decoder, denoted as M, which assigns a value of −∞ to entries where attention links should be deleted, and 0 otherwise MaskedAttention(Q, K, V ) = softmax(Q·KT √dk +M)·V(2.9) where M= [mij]and mij = 0if i≤j −∞ otherwise . 2.2.4 Positional Encoding The positional encoding module addresses the model’s lack of sequential information. Since the original architecture relies only on self-attention mechanisms without considering the order of tokens in a sequence, information about the relative or absolute position of the tokens is injected. To this end, the transformer adds a vector to each input embedding at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension dmodel as the embeddings so that the two can be summed. The positional encoding used in Vaswani are sine and cosine functions of different frequencies PE(k,2i)=sin(k 10000 2i dmodel )(2.10) PE(k,2i+1) =cos(k 10000 2i dmodel )(2.11) where k∈ [0 , n )is the position in the tokenized sentence and i∈ [0 ,dmodel 2 ]position in the dimension of the model. Having n tokens embedded as X = ( x1,··· , xn )with xi∈Rdmodel , the positional encoding in
2.2 Tranformer Architecture 13 the transformer would be PE(k)=(sin(k 10000 0 dmodel ), cos(k 10000 0 dmodel ),··· , sin(k 10000 dmodel dmodel ), cos(k 10000 dmodel dmodel )) (2.12) The intuition here is that adding these values to the embeddings provides meaningful distances between the embedding vectors once they are projected into Q/K/V vectors and during dotproduct attention. 2.2.5 Point-Wise Feed-Forward Network A unidirectional neural network is linked to each encoder and decoder module within the Transformer, involving two linear transformations separated by a point-wise application of a ReLU activation FFN(x) = max(0, xW1+b1)·W2+b2(2.13) where W1∈Rdmodel×dff , b1∈Rdff , W1∈Rdff ×dmodel and b2∈Rdmodel . While the linear transformations are the same across different positions, they use different parameters from layer to layer. The dimensionality of input and output is dmodel = 512, and the inner-layer has dimensionality dff = 2048. 2.2.6 Residual Connections & Normalization Residual connections are a technique that helps models train better. Instead of only propagating information forward through a series of layers, residual connections also allow information to flow directly through the layers via a shortcut connection. Layer normalization is a technique that helps models train faster by cutting down on uninformative variation in hidden vector values. This is achieved by normalizing the values to have a zero mean and unit standard deviation within each layer LayerNorm(Y=X+Z) = A⊙Y−µ σ+ϵ+B(2.14) where A = [ aij ]and B = [ bij ]are learnable scaling and shifting parameters, µ = [ µj ]and σ = [ σj ] are the mean and the standard deviation of each column of the input, and ϵ is a small value added for numerical stability.
2.3 Neural Machine Translation 14 2.2.7 Final Linear & Softmax Layer Transformers are typically used to parameterize a probabilistic model P ( Y|X ). Thus, a linear layer and a softmax layer are used to get the probability distribution of the next token from the decoder stack output P(Y|X) = softmax(Z·Wz) = Z·Wz Pn k=1 exp(Z·Wz)k (2.15) where Wz∈Rn×Vwith V the vocabulary of the trained model. 2.3 Neural Machine Translation Machine Translation is a subfield of computational linguistics that is focused on the task of automatically converting source text in one language to text in another language. Examples of the principal approaches are Rule-Based MT, Statistical MT and Neural MT. Neural Machine Translation is a method of machine translation that uses artificial neural networks to predict the likelihood of a sequence of words, typically modeling entire sentences in a single integrated model. This approach marks a significant shift from traditional rule-based or statistical machine translation methods. The paradigm that is addressed in Machine Translation is the following: Given a source language and a target language, the goal is to translate from the source language to the target language. Mathematically, the best sentence y is found with language lt , conditioned on the source sentence x in language ls p(y|x, ls, lt) = Ty Y t=1 p(yt|y<t, x, ls, lt)(2.16) Here is an overview of its key aspects: End-to-End Model. NMT models are trained end-to-end to maximize the translation performance, learning to map the input text in the source language to the corresponding text in the target language. Seq2Seq Model. NMT is often implemented as a Sequence-to-Sequence learning task with an encoder-decoder architecture. Key Technologies: •Recurrent Neural Networks (RNN). Initially, RNNs, especially Long Short-Term Memory (LSTM) networks, were popular in NMT due to their ability to handle sequences of variable
2.4 Automatic Evaluation of Machine Translation 15 lengths. • Attention Mechanism. Allowing the model to focus on different parts of the input sentence. • Transformers. The introduction of the Transformer model marked a significant leap in NMT performance. Transformers use self-attention mechanisms and do not rely on sequential data processing, allowing for more parallelization and faster training. 2.4 Automatic Evaluation of Machine Translation Automatic evaluation of machine translation is a critical area in the field of NLP and computational linguistics. It involves the use of algorithms and metrics to assess the quality of translations produced by machine translation systems without human intervention. Here are some key aspects of why automatic evaluation methods are important: •Faster and require fewer resources than manual evaluation. •Consistent and reproducible results, to compare different MT models. •Handling of large datasets. Understanding Automated Metrics Functionality. The main idea is to compare an MT system’s output with a reference translation created by humans, assessing the similarity between the generated translated sentence and the reference translation. The major assumption here is that the closer the machine translation is to the reference, the higher its quality. These comparisons are done focusing on words or n-grams to calculate precision scores. A n-gram refers to a sequence of n items extracted from a given text. These items could be phonemes, syllables, letters, words, or other similar elements. Subsequently, this thesis introduces SacreBLEU and COMET as the primary metrics for the automatic evaluation of machine translation systems. 2.4.1 SacreBLEU BLEU (Bilingual Evaluation Understudy) is a widely recognized metric for evaluating the quality of machine-translated text. Developed in 2002 [ 2 ], BLEU quantifies the similarity between a machine translation and human translations, focusing primarily on the precision of word choice and the correct sequencing of words. It employs a technique called n-gram matching, which compares fixed-length sequences of words (n-grams) from the generated output to the reference text. BLEU score, which ranges from 0 to 1, is known for its ability to provide quick, consistent
2.4 Automatic Evaluation of Machine Translation 16 evaluations, making it a key metric in the development and refinement of machine translation algorithms. The higher the value, the better the translations. To address the problems found in the original BLEU metric, such as difficulties in reproducibility on various systems and situations, sacreBLEU was introduced [ 3 ]. SacreBLEU is an open-source standardization of the BLEU metric, designed to provide consistent and comparable BLEU scores. It addresses the reproducibility issue by fixing the test sets and providing pre-defined tokenization methods, ensuring that BLEU scores are consistent and reliable among different research and development approaches. To gain a fundamental understanding of the basis for implementing BLEU and SacreBLEU independent of the tokenization methods, the BLEU score is mathematically defined as a combination of a precision measure and a brevity penalty: Modified n-gram Precision measures the precision of n-grams in the candidate translation by counting, for each n-gram, the maximum frequency of its occurrence in any single reference translation, and this count is then clipped to the maximum to prevent over-counting. The precision Pnfor each n-gram is calculated as Pn=Number of clipped n-gram matches Total number of n-grams in the candidate translation (2.17) Geometric Mean of Precision Scores. The overall precision P is the geometric mean of the precision scores Pnfor each n-gram size P= ( N Y n=1 Pn)1 N(2.18) Here, N is typically 4, meaning that BLEU considers up to 4-grams. Brevity Penalty. To penalize short translations, a penalty is included BP = 1if c > r e(1−r c)if c≤r (2.19) where cis the length of the candidate translation and ris the reference corpus length Finally, the BLEU score is calculated as BLEU =BP ×P(2.20) This score is often reported as a percentage from 0 to 100. Despite its mathematical rigor, it is
2.4 Automatic Evaluation of Machine Translation 17 important to remember that BLEU is just one of many metrics for evaluating machine translation and has its limitations, particularly in capturing semantic accuracy and grammatical correctness. 2.4.2 COMET The Crosslingual Optimized Metric for Evaluation of Translation (COMET) is an open-source framework designed to train Machine Translation (MT) metrics that correlate with various types of human judgments. Introduced in [ 4 ], it provides predictive estimates of human judgments in several forms, including Direct Assessments (DA), Human-mediated Translation Edit Rate (HTER), and metrics aligned with the Multidimensional Quality Metric (MQM) framework. COMET takes 3 inputs: sources (a list of source sentences), predictions (a list of candidate translations) and references (a list of reference translations). Architectural Overview. COMET supports two distinct model architectures, each with a different training objective and incorporating specialized components: • Estimator Model: This model directly regresses on a quality score. It is comprised of a cross-lingual encoder, which independently encodes the source, hypothesis, and reference into word embeddings. These embeddings are then passed through a pooling layer to create sentence embeddings for each segment. The sentence embeddings are combined and fed into a feed-forward regressor. The model’s training involves minimizing the Mean Squared Error (MSE). • Translation Ranking Model: Designed to minimize the distance between a "better" hypothesis and the reference as well as the original source, this model architecture includes a cross-lingual encoder followed by a pooling layer. It processes four segments: the source, the reference, a "better" hypothesis, and a "worse" one. The model is optimized using the triplet margin loss, aiming to refine the embedding space for more accurate MT evaluation. Both architectures use XLM-RoBERTa base [ 5 ] as the encoder model. The encoder produces token embeddings for each layer of the input sequence, mapping them into a shared feature space. The pooling layer then consolidates information from the most crucial encoder layers into a single embedding for each token, achieved via a layer-wise attention mechanism, followed by average pooling to derive sentence embeddings. Trained Models. COMET’s framework has been applied to train several models, each targeting specific types of human judgments: • COMET-HTER: This version of the estimator model regresses on HTER and was trained
2.5 Multilingual Machine Translation 18 using the QT21 corpus. • COMET-MQM: Another estimator model variant, it regresses the score trained on an internal MQM corpus. • COMET-RANK:Translation Ranking Model Trained with the WMT DARR corpus from 2017 and 2018, this model emphasizes minimizing the distance between more accurate translations and their corresponding references and sources. Through these models, COMET provides an improved and comprehensive approach to machine translation evaluation, aligning closely with various human judgment criteria and advancing the field of automated language translation assessment. Since the first publication of COMET several models have been released. In this thesis, Unbabel/wmt22-comet-da [ 6 ] is the model selected to assess translation accuracy. This training approach scales the scores between 0 and 1. This makes it easier to interpret the scores: a score close to 1 indicates a high-quality translation, while a score close to 0 indicates a translation that is no better than random chance. 2.5 Multilingual Machine Translation Multilingual Machine Translation (MMT) aims to build a single model to translate between any pair of languages. MMT models optimize computation when translating to many languages with a single model and share information between similar languages. This has the potential to improve the translation quality of low-resource language pairs by leveraging the ability to transfer knowledge from similar but higher-resource language pairs and data in similar domains but in different languages. Historically, multilingual systems haven’t matched the effectiveness of bilingual models in translations of the same language pairs, as the model’s resources need to be divided among several languages. This limitation was somewhat mitigated by expanding the model’s capacity, though this solution requires larger, more complex multilingual datasets that are challenging and time-consuming to compile. Most previous research has primarily concentrated on datasets and models centered around English, facilitating translations to and from English but not among non-English languages. This focus on English leads to a bias in the data and the models derived from it, which does not accurately represent real-world translation usage and often results in reduced effectiveness for translations that don’t involve English.
2.8 Fine-Tuning 25 Figure 2.8: ∆WLoRA Reparametrization. This approach has several advantages: • LoRA makes fine-tuning more efficient by drastically reducing the number of trainable parameters. • The original pre-trained weights are kept frozen, which means you can have multiple lightweight and portable LoRA models for various downstream tasks built on top of them. • LoRA is orthogonal to many other parameter-efficient methods and can be combined with many of them. • Performance of models fine-tuned using LoRA is comparable to the performance of fully fine-tuned models. • LoRA does not add any inference latency because adapter weights can be merged with the base model.
2.8 Fine-Tuning 26 2.8.3 LoHa Low-Rank Hadamard Product (LoHa), is similar to LoRA except it approximates the large weight matrix with more low-rank matrices and combines them with the Hadamard product. This method is even more parameter-efficient than LoRA and achieves comparable performance. To achieve better fine-tuning performance, a relatively large rank might be necessary, particularly when working with larger fine-tuning datasets or when the data distribution of downstream tasks greatly deviates from the pretraining data. However, this cloud leads to increased memory usage and more storage demands. The LoHA approximation was introduced in FedPara [ 19 ] developed for federated learning. One of the advantages of FedPara is that the maximum rank of the resulting matrix is larger than those derived from conventional low-rank decomposition (such as LoRA). More precisely, the new reparametrization Figure 2.9: ∆WLoHA Reparametrization. ∆W= (B1A1)⊙(B2A2)(2.26) where ⊙ denotes the Hadamard product (element-wise product), B1, B2∈Rp×r , A1, A2∈Rr×q , and r≤min(p, q), the rank of ∆Wcan be as large as r2.
2.8 Fine-Tuning 27 2.8.4 LoKr Low-Rank Kronecker Product (LoKr) [ 18 ], is a LoRA-variant method that approximates the large weight matrix with two low-rank matrices and combines them with the Kronecker product. This method is an extension of the KronA technique [ 22 ] for fine-tuning language models, and employs Kronecker products for matrix decomposition. Additionally, this technique can be applied to convolutional layers. A unique advantage of using Kronecker products lies in the multiplicative nature of their ranks, allowing us to move beyond the limitations of low-rank assumptions. Figure 2.10: ∆WLoKr Reparametrization. In summary, writing ⊗for the Kronecker product, the adapter layer is modified to ∆W= [C⊗(BA)] (2.27) The size of these matrices are determined by two user-specified hyperparameters: the factor f and the dimension r. With these, we have C∈Rup×uq,B∈Rvp×r, and A∈Rr×vq, where up= max(u≤min(f, √p)|pmod u= 0), vp=up. The two scalars uq and vq are defined in the same way. Additionally, LoKr includes a flexible option for low-rank decomposition, allowing users to decide whether to apply it specifically to the right block that emerges from the Kronecker decomposition. This added feature enables user control over the model’s fine-tuning process.
28 3 Experimental Framework In this section, we performed extensive experiments to compare different LoRA algorithms with various configurations. The goal was to evaluate the effectiveness in fine-tuning a multilingual machine translation model while minimizing the number of trainable parameters. The downstream task is the translation from English to Basque chosen to leverage cross-lingual transfer from the multiple languages already learned from the model. Additionally, extensions of Lora used in Stable Fusion models were explored, such as LoHa and LoKr. Finally, all the experiments are compared with the full fine-tuning of the model. Fine-Tuned Model. The experiments are carried out with M2M100 with 418M parameters, the distilled version of the original model, chosen for its multi-language capabilities. M2M100 as explained in Section 2.7, translates from 100 languages to 100 languages, thus, the main idea is to benefit from all the languages to learn a new one, the Basque. Baselines Models. No Language Left Behind [ 23 ] from Facebook AI and Google translator will be chosen to check how the multilingual machine translation state-of-art models perform. Tokenizer. The vocabulary dictionary dimension of the model will be extended to add new Basque tokens to help in the learning process. Thus, a new tokenizer has to be trained to generate these new sub-sequences. The original vocabulary has a size of 128k unique tokens. We trained with SentencePiece a Basque tokenizer with vocabulary 6000 subwords using the Unigram Model. The new dictionary has 131430 tokens as M2M100 tokenizer has some identic tokens as the learned Basque tokenizer. Software Framework. PyTorch [ 24 ] is the foundational framework used for building and training LLM models. Alongside PyTorch, it is used Transformers [ 25 ], allowing us to access state-of-the-art pre-trained models. Additionally, we employed Parameter-Efficient Fine-Tuning PEFT [ 26 ], which is a library for efficiently finetuning large pretrained models without updating all model’s parameters. Finetuning Strategies. M2M100 is fully fine-tuned to be compared with LoRA finetuning. Moreover, extensions of LoRA are tried, LoKr and LoHA. Metrics. The main metric chosen to assess the performance of the new fine-tuned model is SacreBLEU. Additionally. COMET-ML is also computed for further analysis. Datasets. The source of the datasets is OPUS [ 27 ] which is a growing collection of translated texts from the web. It is chosen 3 types of datasets to assess the scalability of the method.
29 Small [ 28 ], medium and large datasets [ 27 ]. The main dataset used in the experiments is the medium-size WIKIMEDIA dataset which are Wikipedia translations published by the wikimedia foundation and their article translation system. The small dataset TED2020 contains a crawl of nearly 4000 TED and TED-X transcripts from July 2020. The last dataset used is the considered large named EHUHAC which is a compilation of translations of 202 books of fables. For evaluation, the benchmarking datasets FLORES [23][29][30] are used. Training Process. In the following table is summed up the configuration used in all experiments for the training process Parameter Value/Description Trained Epochs 8 FP16 (16-bit Floating Point) True Batch Size 4 Gradient Accumulation Steps 2 Learning Rate 2·10−5 Weight Decay 0.01 Table 3.1: Training Configuration Parameters Experiment Configurations. Various parametrizations of LoRa are explored to evaluate its effectiveness in comparison to full fine-tuning. These are the parameters configured: • LoRA rank r : the rank of the update matrices, lower rank results in smaller update matrices with fewer trainable parameters. Different powers of 2 are tried to see how SacreBLEU changes: r∈ (1 , 2 , 4 , 8 , 16 , 32 , 64 , 128 , 256). The experiments are focused on large ranks as the use case is mentioned in [ 17 ]. It says that if the downstream tasks were in a different language than the one used for pre-training, retraining the entire model (similar to LoRA with r = dmodel ) could certainly outperform LoRA with a small r and for the extension we should use larger ranks. •LoRa α: scaling factor. Alpha is parametrized in terms of r,α∈(0.5r, r, 2r). • Embeddings: they are subject to training or remain fixed, and whether they are initialized with random values or through an averaging of the already trained embeddings of M2M100 to initialize the newly added tokens. Particularly when expanding the vocabulary of pretrained models, a notable challenge emerges in the initialization of embeddings for new words. The default method employed by Transformers, which relies on initializing these new word embeddings with small-norm random noise, can lead to a model bias. This bias results in the model showing an unfair preference for the new words, often assigning them a very high likelihood of being the correct choice. This phenomenon is primarily due to the
30 training dynamics where the logits (dot products in the softmax layer) are usually large and negative. Consequently, the new word logits, being close to zero, disproportionately influence the softmax function, since e0>> e−β for large values of β . Further justification can be found in [31]. • Combination of which layers of the architecture are trainable: Attention weights, fully connected layers, head, first layers of encoder-decoder, last layers of encoder-decoder, decoder layers and encoder layers. Performance Analysis: Once all experiments are done, different analyses are carried: • Comparison between LoRA, LoHA, LoKr, and the full fine-tuning in terms of SacreBLEU and trainable parameters. •Scalability assessment across different dataset sizes in terms of SacreBLEU. • Zero-Shot Translation: Evaluating the model’s capability to translate pair of languages not seen during training. Language Pair Reason for Selection Spanish - Basque Geographical Proximity French - Basque Geographical Proximity Japanese - Basque Similar Alignment German - Basque Different Origin and Linguistic Characteristics Basque - English Opposite Task Learned Table 3.2: Language Selection Choices • Catastrophic Forgetting: Assessing model retention of previously learned languages, involving five languages. Pairs of languages: EN-ES,EN-FR,EN-JA,JA-ES,DE-JA.
31 4 Results In the experimental framework section of this thesis, we conducted a comprehensive evaluation of various finetuning strategies implemented on a new language pair, English-Basque (EN-EU) using the medium-sized dataset. This study covered key areas such as supervised translation, data scalability, catastrophic forgetting, and zero-shot translation. Our approach for supervised translation was straightforward yet effective: we experimented with different fine-tuning techniques based on playing with the different parameters that LoRA offers and with different embedding strategies. The aim was to find a strategy that reaches an optimal compromise between trainable parameters and boost in SacreBLEU scores in contrast to the results from fully fine-tuning M2M100. Once the optimal strategy regarding LoRA and the embeddings is concluded, the next step is to explore various combinations of trainable layers. This investigation aims to determine which components of the transformer architecture exhibit enhanced learning capabilities or even might surpass the performance of the optimal configuration. Following the completion of this methodology, we will also employ additional techniques such as LoHa and LoKr to further assess whether these approaches yield improvements in our metrics. After this last experiment, we continue with the best configuration achieved so far. As a final point, data scalability, catastrophic forgetting, and zero-shot translation of the fine-tuned model will be evaluated to obtain more conclusions. Baselines. Before starting the analysis, an evaluation of the current state-of-the-art models in multilingual machine translation is performed. This step provides us with a clear reference point from which to advance our research. These are the results after inference with [ 23 ] and Google translator Experiment SacreBLEU # parameters Google Translate 14.06 - NLLB 14.90 1.3 B Table 4.1: Evaluation of NLLB and Google Translation Using the SacreBLEU Metric on the FLORES Dataset.
4.1 Supervised Translation 32 4.1 Supervised Translation In this section, the different experiments explained in Section 3are analyzed: Full Fine-Tuning (FT). During fine-tuning, the model is initialized to the pre-trained weights and biases, and all model parameters undergo gradient updates. In the following table, the test results after fine-tuning with the 3 different dataset sizes are displayed Dataset # Data SacreBLEU COMET WIKIMEDIA 61K 7.43 0.72 TED2020 10k 3.86 0.56 Ehuhac 0.6M 5.13 0.65 Table 4.2: Comprehensive Analysis of M2M100 Full Fine-Tuning Across Diverse Dataset Sizes. Table 4.2 illustrates a comparison of results obtained after fine-tuning all model parameters using three datasets of varying sizes. Contrary to the usual expectation that larger datasets yield better metrics, the highest scores for SacreBLEU and COMET are observed with the medium-sized dataset. As anticipated, fine-tuning with the smaller dataset results in the least favorable outcomes. This discrepancy could be attributed to the quality of the data or variances in the domain compared to the data used in pre-training. In our analysis, we will consistently compare the results of full-fine-tuning with those of a fine-tuning strategy that has been trained on the identical dataset. The selected dataset for further experiments will be the WIKIMEDIA dataset of 61K as it offers the best metrics. LoRA finetuning. To make fine-tuning more efficient, LoRA’s approach is to represent the weight updates with two smaller matrices through low-rank decomposition. Training attention weights. In Tables 4.3 and 4.4 are shown the SacreBLEU results when varying LoRA rank, LoRA alpha and different approaches for the embeddings when only updating the attention layers of all encoders and decoders, matrices WK, WQand WV Experiment LoRa Rank (r) Embeddings Alpha 1 2 4 8 16 32 64 128 256 Random Initialization α= 0.5r1.37 1.26 1.26 1.34 1.31 1.55 2.22 2.39 3.08 α=r1.30 1.29 1.47 1.46 1.63 1.88 2.65 2.92 3.62 α= 2r1.30 1.32 1.25 1.49 2.06 2.70 2.93 3.46 4.29 Average Initialization α= 0.5r1.38 1.47 1.73 2.15 2.45 2.89 3.33 4.01 4.25 α=r1.46 1.34 2.24 2.50 2.84 3.44 3.92 4.67 5.26 α= 2r1.80 1.97 2.46 3.07 3.36 4.05 4.77 5.06 5.60 Table 4.3: SacreBLEU Results Using the FLORES Dataset for Attention Weight Training and Frozen Embeddings.
4.1 Supervised Translation 33 Experiment LoRa Rank (r) Embeddings Alpha 1 2 4 8 16 32 64 128 256 Random Initialization α= 0.5r3.76 4.08 4.00 4.34 4.58 4.87 5.19 5.47 5.62 α=r3.91 4.24 4.51 4.56 4.70 5.0 5.34 5.76 5.98 α= 2r4.24 4.40 4.57 4.64 5.03 5.10 5.57 6.14 6.41 Average Initialization α= 0.5r3.89 3.94 4.15 4.37 4.72 4.75 5.13 5.45 5.73 α=r4.15 4.18 4.26 4.67 4.96 5.15 5.35 5.72 6.04 α= 2r4.27 4.27 4.48 4.84 5.25 5.36 5.68 6.01 6.42 Table 4.4: SacreBLEU Results Using the FLORES Dataset for Attention Weight and Embedding Training. After closely inspecting the data presented in Tables 4.3 and 4.4, it becomes evident that an increase in the rank of the matrices correlates with a rise in the SacreBLEU scores. This trend aligns with expectations, as a higher rank means the training of a greater number of model parameters. Another clear proof observed from the tables is that when the scaling factor ( α ) is doubled relative to the rank ( r ), there is a noticeable increase in the SacreBLEU scores. This observation implies that the knowledge acquired in the new matrices plays an important role in significantly contributing to the performance of the previously pre-trained parameters of the model. Furthermore, we conclude that freezing the embeddings, particularly when initializing the newly added Basque word embeddings with the average of the pre-trained model, results in a substantial improvement in the metrics observed. However, when these embeddings are not frozen, regardless of their initial configuration, they tend to converge to similar results. The results indicate that training the embeddings leads to higher performance compared to not training them. However, this approach increases the number of parameters by up to 40-50% of the model. In contrast, when the embeddings are not trained, LoRA updates only between 5-20% of the model’s parameters. The optimal results are attained with a rank (r) of 256, using average initialization, and an alpha value set to 2r. This configuration yields a SacreBLEU score of 5.6 when the embeddings are not trained and 6.42 when they are trained. We observed that for each rank of the matrices, this strategy consistently delivers the highest SacreBLEU scores. In Figure 4.1, it is calculated for the scaling factor ( α = 2 r ) and each embedding strategy, the number of parameters trained for each rank and SacreBLEULoRA SacreBLEUF ull to establish a comparison between LoRA and FT results. SacreBLEUF ull = 7 . 43 as is the result achieved with the medium-sized dataset training.
4.1 Supervised Translation 34 Figure 4.1: Percentage of Trainables Parameters vs. Percentage of Full Fine-Tuning in Experiments with α= 2r. The analysis reveals that the most favorable outcomes are obtained when training the embeddings for ranks 128 and 256 regardless of the embedding strategy. Achieving 80.89% and 86.35% of the full fine-tuning performance with 46.53% and 49.42% of the parameters trained, respectively. Furthermore, notable results are still achieved without training the embeddings. In this case, we reach 75.32% of the full fine-tuning performance while only training 13.41% of the parameters. In the plot, two clear effects can be seen in the curves. The effect of initializing the new embeddings as average when the embeddings are frozen and the effect of training the embeddings against when they are not trained. Training Attention Weights, Fully Connected Layers and Head. Continuing the exploration of LoRA’s capabilities, the subsequent tables present similar experiments to those previously conducted. This time, however, the focus extends to training not only the attention weights but also includes the fully connected layers situated within the various attention modules and the final head of the architecture. Notably, these experiments are conducted exclusively with larger ranks, ranging from 16 to 256, based on the observation that lower ranks (from 1 to 4) did not yield results as promising as the higher ranks.
4.4 Zero-Shot Translation 41 Language Pair Model EN-ES EN-FR EN-JA JA-ES DE-JPN M2M100 22.17 38.99 2.97 13.17 2.64 Full Fine-Tuning 1.18 1.17 0.09 0.59 0.07 LoRA Fine-Tuning 1.20 1.41 0.13 0.84 0.11 Table 4.10: SacreBLEU Results for M2M100 Languages Translated to Basque. After examining the table, it becomes apparent that catastrophic forgetting is indeed occurring. There is an obvious decline in the SacreBLEU scores from the original model, indicating a substantial loss of translation capability. When comparing the two fine-tuning methods, it is observed that the LoRA strategy shows a lesser degree of forgetting than full fine-tuning. This is most likely due to the reduced number of parameters trained in the LoRA approach. However, despite this relative advantage, the results still confirm the occurrence of catastrophic forgetting. An interesting advantage of the LoRA fine-tuning method, however, is the ability to go back to the model’s original functionality. This is achieved by deactivating the adapters updated during training, effectively unmerging the learned weights from the original matrices. Consequently, by simply unloading the LoRA modules, one can access either the improved capabilities for the downstream task provided by the adapters or the original, broader translation capabilities of the pre-trained model. 4.4 Zero-Shot Translation The next exploration is about analyzing how the translation from languages already learned in the pre-trained M2M100 to Basque work after finetuning the low-rank matrices. Additionally, it is also assessed the reverse downstream task, from Basque to English. The main goal here is to enable the machine translation model to translate between language pairs it has never explicitly been trained on. In the following table, the results are presented Language Pair Model EU-EN ES-EU JA-EU FR-EU DE-EU Full Fine-Tuning 1.15 4.30 3.13 4.29 4.62 LoRA Fine-Tuning 0.58 4.05 2.95 4.36 4.88 Table 4.11: SacreBLEU Results for Directions M2M100 Languages to Basque. As illustrated in the table, the model demonstrates its capability to learn new translation directions, specifically from languages included in the pre-trained model to Basque. The results from Full Fine-tuning (FT) and LoRA fine-tuning show a comparable level of performance. For
4.4 Zero-Shot Translation 42 translations from Spanish (ES) to Basque (EU) and Japanese (JA) to Basque (EU), the FT strategy yields marginally better results. Conversely, in the case of LoRA FT, translations from French (FR) to Basque (EU) and German (DA) to Basque (EU) are more effective. Particularly noteworthy are the results for FR-EU and DA-EU, where SacreBLEU scores of 4.36 and 4.88 respectively are achieved. These scores are remarkably close to the performance observed for English (EN) to Basque (EU) translations, which stand at 5.6. This highlights the model’s efficiency in adapting to new language pairs, especially in scenarios where the target language is underrepresented in the training data.
43 5 Conclusions and Future Work This study has proposed different fine-tuning strategies using Low-Rank Adaptation (LoRA) for the English-Basque (EN-EU) translation task focusing on an encoder-decoder architecture such as M2M100. The main focus of the research was the exploration of supervised translation, emphasizing the balance between trainable parameters and the performance of translation accuracy assessed by SacreBLEU and COMET. The primary task of this study was to identify an optimal fine-tuning strategy that employs the LoRA method with specific configurations. Our approach, characterized by a LoRA rank of 256 and an α of 512 applied in the attention weights of the architecture, and freezed embeddings with average initialization of the new Basque tokens, emerged as highly efficient. It remarkably achieved a significant proportion of the full fine-tuning performance with a 75.32% result while training 13.41% of the model’s parameters. When the embeddings were also trained alongside the same LoRA configuration, the performance nearly matched that of the full fine-tuning method, achieving a SacreBLEU score of 6.42. However, this improvement came at a significant cost: it required training about 50% of the model’s parameters, which is a substantial increase compared to the more efficient approach of freezing the embeddings. Contrasting full fine-tuning with the LoRA-based approach, it was noted that while the former yields the highest SacreBLEU scores, particularly with a medium-sized dataset, the latter presents an effective alternative. LoRA’s efficacy is highlighted when it focuses on fine-tuning specific components like attention weights, achieving near-comparable performance to full fine-tuning but with much fewer parameters involved. This finding reinforces the potential of the Low-Rank method for fine-tuning encoder-decoder architectures in Machine Translation. Moreover, we aimed to train additional linear layers included in the transformers of M2M100 model, such as the fully connected layers and the head, using strategies similar to those previously applied. Our findings offered a detailed insight of these strategies. We observed that a strategy involving average initialization of new embeddings did not consistently lead to the best results. This was especially evident when the scaling factor α was set to half the rank value (0.5r), which resulted in suboptimal performance, even with increased rank. However, a marked improvement was noted when the rank was set to 256 and α to twice the rank (2r), achieving SacreBLEU scores over 5, signifying a considerable enhancement in performance. Nevertheless, this benefit was accompanied by the drawback of increasing the number of parameters required. Our analyses indicate that training a wider range of the model’s components may yield comparable results to
44 those strategies only focus on attention weights. However, this is not always the case; certain strategies can lead to inconsistent and, at times, inferior results. The decision to train more components, again, presents the trade-off between increasing the parameters and potential performance gains. Another significant aspect of the study was the exploration of training specific parts of the M2M100 model’s architecture. Here, the research concluded that training only the decoder component achieved one of the most interesting results. This approach not only reached a high SacreBLEU score of 4.94 but also marked efficiency in terms of the minimal number of parameters trained, a 4.62 %. Conversely, training the encoder alone was less effective. This outcome can be largely attributed to the crucial role of the decoder in the process of generating the new target language. The exploration of variations of LoRA, specifically LoHa and LoKr, provided further insights. Beginning with LoHa, it was observed that despite these variations in increasing the number of trainable parameters, they did not exceed the performance of the optimal LoRA strategy. This finding implies that just increasing the number of trainable parameters does not always lead to better performance. On the other hand, LoKr presented a different scenario. The significant reduction in trainable parameters offered by this method resulted in underfitting in test results. In summary, within this specific framework, LoRA demonstrates superior performance compared to both LoHa and LoKr. The investigation was extended into data scalability, catastrophic forgetting, and zero-shot translation analysis. First, the research challenged the conventional belief that larger datasets always yield better results. Simply increasing the size of the dataset did not lead to significant performance gains. However, when compared to the full fine-tuning approach, it was observed that the larger dataset achieved the highest proportion of performance improvement. This highlighted the importance of data quality and relevance. In terms of catastrophic forgetting, the study found that the LoRA strategy exhibited less forgetting than full fine-tuning, likely due to fewer parameters being trained. Moreover, the flexibility of the LoRA method to switch back to the original model’s capabilities was highlighted as a notable advantage. In zero-shot translation analysis, the model demonstrated favorable results in translating new language pairs to Basque, showing adaptability to unseen languages in the training data. This research created a framework for more sophisticated and efficient machine translation models,
45 capable of delivering high-quality translations with reduced computational demands. Future directions for improving the fine-tuning strategy may include using the M2M100 model with 1.2 billion parameters and applying Quantized LoRA to update the model’s parameters. This approach aims to enhance performance while maintaining a similar number of trained parameters as achieved with the default version of LoRA. Extending the study to other lowresource languages by training more LoRA adapters, could also provide valuable insights into the applicability of these methods across diverse linguistic contexts as well as take advantage of the activation and deactivation of these adapters for further capability. Investigate the reasons behind the suboptimal results observed in supervised translation, particularly when average initialization is used and all linear layers (attention layers, fully connected layers, and the head) of M2M100 are trained. Last, explore other multilingual models like T5, assessing their potential and effectiveness in comparison to the M2M100 model.
46 To conclude this thesis, we provide a fine-tuning recipe to adapt a multilingual model to a new language using LoRA. Create your own M2M100 language translator in a low-resource setting This thesis describes several steps that are taken to learn a new language to M2M100 using LoRA. The recommended shortest path to replicate this for another language is to follow these steps: Vocabulary. Extend M2M100 vocabulary with new tokens for your target language using SentencePiece. The optimal size for your vocabulary depends on your language. In this research, it was used 6000 tokens. Embedding strategy. Initialize the new word embeddings with the average of the pre-trained embeddings and freeze the embeddings during training. LoRA rank. Set a high rank for the updated matrices in order to facilitate the learning of the new language. In this study, it was proposed r= 256. LoRA α . Set a scaling factor that gives more importance to the updated matrix. In this thesis, it was proposed α= 2r= 512 Finetuned layers Freeze all the layers except the attention layers of M2M100 to obtain a comparable result to the full fine-tuning method. If it is desired to get a closer solution, unfreeze the embeddings to train them. However, the time and the parameters will increase. .
References 47 References [1] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017. [2] Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA, July 2002. Association for Computational Linguistics. [3] Matt Post. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Belgium, Brussels, October 2018. Association for Computational Linguistics. [4] Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. COMET: A neural framework for MT evaluation. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online, November 2020. Association for Computational Linguistics. [5] Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. CoRR, abs/1911.02116, 2019. [6] Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. COMET-22: Unbabel-IST 2022 submission for the metrics shared task. In Philipp Koehn, Loïc Barrault, Ondřej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Alexander Fraser, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Paco Guzman, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Tom Kocmi, André Martins, Makoto Morishita, Christof Monz, Masaaki Nagata, Toshiaki Nakazawa, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Marco Turchi, and Marcos Zampieri, editors, Proceedings of the Seventh Conference on Machine Translation (WMT), pages 578–585, Abu Dhabi, United Arab Emirates (Hybrid), December 2022. Association for Computational Linguistics. [7] Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin. Beyond english-centric multilingual machine translation. J. Mach. Learn. Res., 22(1), jan 2021. [8] Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Katrin Erk and Noah A. Smith, editors, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany, August 2016. Association for Computational Linguistics.
References 48 [9] Taku Kudo. Subword regularization: Improving neural network translation models with multiple subword candidates, 2018. [10] Ankur Bapna and Orhan Firat. Simple, scalable adaptation for neural machine translation. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 1538–1548. Association for Computational Linguistics, 2019. [11] Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. Adapterfusion: Non-destructive task composition for transfer learning. CoRR, abs/2005.00247, 2020. [12] Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 2790–2799. PMLR, 2019. [13] Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. Adapterhub: A framework for adapting transformers. In Qun Liu and David Schlangen, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2020 - Demos, Online, November 16-20, 2020, pages 46–54. Association for Computational Linguistics, 2020. [14] Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. MAD-X: an adapter-based framework for multi-task cross-lingual transfer. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 7654–7673. Association for Computational Linguistics, 2020. [15] Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 4582–4597. Association for Computational Linguistics, 2021. [16] Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. Compacter: Efficient low-rank hypercomplex adapter layers. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, editors, Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 1022–1035, 2021. [17] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. [18] Shin-Ying Yeh, Yu-Guan Hsieh, Zhidong Gao, Bernard B. W. Yang, Giyeong Oh, and Yanmin Gong. Navigating text-to-image customization: From lycoris fine-tuning to model evaluation. CoRR, abs/2309.14859, 2023.
References 49 [19] Nam Hyeon-Woo, Moon Ye-Bin, and Tae-Hyun Oh. Fedpara: Low-rank hadamard product for communication-efficient federated learning. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. [20] Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. Few-shot parameter-efficient fine-tuning is better and cheaper than incontext learning. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022. [21] Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 3045–3059. Association for Computational Linguistics, 2021. [22] Ali Edalati, Marzieh S. Tahaei, Ivan Kobyzev, Vahid Partovi Nia, James J. Clark, and Mehdi Rezagholizadeh. Krona: Parameter efficient tuning with kronecker adapter. CoRR, abs/2212.10650, 2022. [23] James Cross Onur Çelebi Maha Elbayad Kenneth Heafield Kevin Heffernan Elahe Kalbassi Janice Lam Daniel Licht Jean Maillard Anna Sun Skyler Wang Guillaume Wenzek Al Youngblood Bapi Akula Loic Barrault Gabriel Mejia Gonzalez Prangthip Hansanti John Hoffman Semarley Jarrett Kaushik Ram Sadagopan Dirk Rowe Shannon Spruit Chau Tran Pierre Andrews Necip Fazil Ayan Shruti Bhosale Sergey Edunov Angela Fan Cynthia Gao Vedanuj Goswami Francisco Guzmán Philipp Koehn Alexandre Mourachko Christophe Ropers Safiyyah Saleem Holger Schwenk Jeff Wang NLLB Team, Marta R. Costa-jussà. No language left behind: Scaling human-centered machine translation. 2022. [24] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 8024–8035. Curran Associates, Inc., 2019. [25] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online, October 2020. Association for Computational Linguistics. [26] Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. Peft: State-of-the-art parameter-efficient fine-tuning methods. https: //github.com/huggingface/peft, 2022. [27] Jörg Tiedemann. Parallel data, tools and interfaces in opus. In Nicoletta Calzolari (Conference Chair), Khalid Choukri, Thierry Declerck, Mehmet Ugur Dogan, Bente Maegaard, Joseph Mariani, Jan Odijk, and Stelios Piperidis, editors, Proceedings of the Eight International
References 50 Conference on Language Resources and Evaluation (LREC’12), Istanbul, Turkey, may 2012. European Language Resources Association (ELRA). [28] Nils Reimers and Iryna Gurevych. Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2020. [29] Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. The flores-101 evaluation benchmark for low-resource and multilingual machine translation. 2021. [30] Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. Two new evaluation datasets for low-resource machine translation: Nepali-english and sinhala-english. 2019. [31] John Hewitt. Initializing new word embeddings for pretrained language models, 2021.