scieee AI-readable full text Open interactive document viewer

Model optimization for chess pieces classification

Vilà Sánchez, Diwa

Abstract

Aquest treball de fi de grau investiga l'ús d'arquitectures d'aprenentatge profund per a la tasca de digitalització automàtica de partides d'escacs. El treball previ és escàs i, tot i que ha mostrat resultats prometedors, les tècniques encara necessiten millores addicionals per permetre'n un ús pràctic en producció. La nostra primera contribució es va fer comparant el rendiment de ResNet50, NASNet Mobile, Mobilenet i VGG16 en la classificació de les peces d'escacs a partir d'una imatge de la peça. L'estudi va trobar que totes les arquitectures van ser capaces d'assolir una precisió satisfactòria, arribant al 98\% d'exactitud amb ResNet50. Aquestes arquitectures es van entrenar per desenvolupar un sistema capaç de recuperar les posicions en un tauler d'escacs a partir d'una imatge. En les millors condicions, el sistema pot digitalitzar una imatge d'un tauler d'escacs en 4 segons amb una exactitud del 75\%. Els resultats de les proves de producció destaquen la importància de l'ús de coneixement del domini que tinguem a la nostra disposició i una extracció cuidadosa de les dades, ja que creiem que la majoria de les imprecisions que el sistema cometeix es podrien resoldre amb la introducció d'informació del domini i una forma més òptima d'obtenir les caselles del tauler. L'estudi també proporciona informació sobre el consum d'energia durant l'entrenament de les Xarxes Neuronals Convolucionals.

Full text

id178224   MODEL OPTIMIZATION FOR CHESS PIECES CLASSIFICATION DIWA VILÀ SÁNCHEZ Thesis supervisor: SILVERIOJUANMARTÍNEZFERNÁNDEZ(DepartmentofServiceandInformation SystemEngineering) Thesis co-supervisor: SANTIAGODELREYJUAREZ(DepartmentofServiceandInformationSystem Engineering) Degree:Bachelor'sDegreeinDataScienceandEngineering Bachelor's thesis Facultat d'Informàtica de Barcelona (FIB) Escola Tècnica Superior d'Enginyeria de Telecomunicació de Barcelona (ETSETB) Facultat de Matemàtiques i Estadística (FME) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech Acknowledgments Fisrt of all, I would like to show my gratitude to my supervisor Dr. Silverio Mart´ınez and co-supervisior Santiago Del Rey for their support and advice. Special thanks to my family and friends for their unconditional support during the most difficult parts of the project. 1 Abstract This bachelor’s thesis investigates the use of deep learning architectures for the task of automatic digitization of chess games. Previous work is scarce and, although it has shown promising results, the techniques still need further enhancements to allow a practical use in production. Our first contribution was done by comparing the performance of ResNet50, NASNet Mobile, Mobilenet, and VGG16 when classifying chess pieces given a piece image. The study found that all the architectures were able to retrieve a high satisfactory accuracy, reaching a 98% accuracy with ResNet50. Those architectures were trained in order to develop a system able to retrieve the positions in a chess board given just an image. In the best conditions, the system is able to digitise a chess board image in 4 seconds with a 75% accuracy. The results from the production testing highlight the importance of the use of domain knowledge and a careful extraction of the data, as we believe that the majority of imprecisions the system makes could be solved with the introduction of domain information and a more optimal way to get the squares from the board. The study also gives some insight about the energy consumption of the Convolutional Neural Networks training. Keywords Chess, computer vision, transfer learning. 2 Resum Aquest treball de fi de grau investiga l’´us d’arquitectures d’aprenentatge profund per a la tasca de digitalitzaci´o autom`atica de partides d’escacs. El treball previ ´es esc`as i, tot i que ha mostrat resultats prometedors, les t`ecniques encara necessiten millores addicionals per permetre’n un ´us pr`actic en producci´o. La nostra primera contribuci´o es va fer comparant el rendiment de ResNet50, NASNet Mobile, Mobilenet i VGG16 en la classificaci´o de les peces d’escacs a partir d’una imatge de la pe¸ca. L’estudi va trobar que totes les arquitectures van ser capaces d’assolir una precisi´o satisfact`oria, arribant al 98% d’exactitud amb ResNet50. Aquestes arquitectures es van entrenar per desenvolupar un sistema capa¸c de recuperar les posicions en un tauler d’escacs a partir d’una imatge. En les millors condicions, el sistema pot digitalitzar una imatge d’un tauler d’escacs en 4 segons amb una exactitud del 75%. Els resultats de les proves de producci´o destaquen la import`ancia de l’´us de coneixement del domini que tinguem a la nostra disposici´o i una extracci´o cuidadosa de les dades, ja que creiem que la majoria de les imprecisions que el sistema cometeix es podrien resoldre amb la introducci´o d’informaci´o del domini i una forma m´es `optima d’obtenir les caselles del tauler. L’estudi tamb´e proporciona informaci´o sobre el consum d’energia durant l’entrenament de les Xarxes Neuronals Convolucionals. Paraules clau Escacs, visi´o per computadors, Transfer`encia de l’aprenentatge. 3 Resumen Este ytabajo de fin de grado investiga el uso de t´ecnicas de aprendizaje profundo para la tarea de digitalizaci´on autom´atica de partidas de ajedrez. El trabajo previo es escaso y, aunque muestra resultados prometedores, las t´ectnicas a´un requieren mejoras adicionales para permitir un uso en producci´on. Nuestra primera contribuci´on se ha realizado comparando el rendimiento de ResNet50, NASNet Mobile, Mobilenet y VGG16 al clasificar las pieza de ajedrez a partir de una imagen de la pieza. El estudi´o encontr´o que todas las arquitecturas fueron capances de lograr una precisi´on satisfactoria, alcanzando un 98% de exactitud con ResNet50. Estas arquetecturas fueron entrenadas para desarrollar un sistema capaz de recuperar las posiciones en un tablero de ajedrez a partir de una imagen de este. En las mejores condiciones, el sistema puede digitalizar una imagen de un tablero de ajedrez en 4 segundos con un 75% de precisi´on. Los resultados de las pruebas de producci´on destacan la importancia del uso del conocimiento del dominio que est´e a nuestra disposici´on y una extracci´on cuidadosa de los datos, ya que creemos que la mayor´ıa de las imprecisiones que el sistema comete podr´ıan resolverse con la introducci´on de informaci´on del dominio y una forma m´as ´optima de obtener las casillas del tablero. El estudio tambi´en proporciona informaci´on sobre el consumo de energ´ıa durante el entrenamiento de las Redes Neuronales Convolucionales. Palabras clave Ajedrez, visi´on por computador, transferencia del aprendizaje. 4 Contents 1 Introduction 7 1.1 Motivation ...................................... 7 1.2 Goalsoftheproject ................................. 8 1.3 Structureofthedocument.............................. 8 2 Background 9 2.1 DeepLearning(DL) ................................. 9 2.1.1 Multi layer Perceptrons (MLP) . . . . . . . . . . . . . . . . . . . . . . . 9 2.1.2 Convolutional Neural Networks (CNN) . . . . . . . . . . . . . . . . . . . 10 2.2 TransferLearning................................... 10 3 State of the art on chess pieces recognition 12 4 Method 14 4.1 Datasets........................................ 15 4.1.1 Chess piece classifier dataset . . . . . . . . . . . . . . . . . . . . . . . . 15 4.1.2 Production test dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 4.2 ChessPieceClassifier ................................ 15 4.3 Evaluationmetrics .................................. 17 4.3.1 Multiclass Confusion Matrix . . . . . . . . . . . . . . . . . . . . . . . . 17 4.3.2 Accuracy ................................... 17 4.3.3 Precision and Recall . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 4.3.4 Loss ...................................... 18 4.4 Improvementtechniques............................... 18 4.4.1 Batch Normalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 4.4.2 Dropout [16] ................................. 19 4.4.3 Global Average Pooling (GAP): . . . . . . . . . . . . . . . . . . . . . . . 19 4.4.4 Hyperparameter tuning: . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 4.5 Production: Broadcasting chess system . . . . . . . . . . . . . . . . . . . . . . . 20 4.5.1 Broadcasting system . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 5 Results of training the piece recognition model (Objective 1) 23 5.1 Baselines ....................................... 23 5.2 Firstexperiments................................... 23 5.3 CustomCNN..................................... 25 5.4 Giving Transfer Learning another chance . . . . . . . . . . . . . . . . . . . . . 26 5 6 Results of inference in production (Objective 2) 29 7 Discussions and Limitations 33 7.1 Discussions ...................................... 33 7.2 Limitations ...................................... 35 8 Conclusions and future work 37 6 1. Introduction Constructing an end-to-end system capable of broadcasting a chess game is not a straightforward task. When using computer vision it is a challenge with scarce research on the field. In this thesis, we will develop a computer vision system that detects the board state in a chessboard just from an image. 1.1 Motivation Chess has its significance in Computer Vision research as it presents a complex domain and diverse visual patterns. The pieces are visually similar to each other, which makes the efficient recognition of those chess pieces a challenge. Additionally, if we want to track a game of chess using image processing and analysis we have to tackle the problems of possible changes in illumination, image perspective and the variability in the game conditions. Other reasons why chess is a subject of interest in this field are: •Image analysis: In order to digitize a chess game, the computer must be able to detect the board, the squares, and the pieces. This involves analyzing and interpreting complex patterns and configurations on the board, which involves image processing techniques such as edge detection, segmentation, or object recognition. Here, computer vision algorithms can be used to: –Recognise whether a square has a chess piece or it is empty. –Recognise individual chess pieces –Determine the positions of the pieces accurately from images or video streams •Object tracking: Chess pieces move across the board, and their accurate tracking can be essential to follow and understand the game’s progression. Computer vision techniques are employed to track the movement of chess pieces in real-time, enabling automated analysis and capturing the sequence of moves. In this project, we will tackle the problem of chess piece recognition, a not fully solved computer vision problem. A good solution to this problem could be beneficial to amateur and semi-professional players who want to be able to play with a physical board and have it automatically on a computer. It can be also beneficial for those games or tournaments that want to be broadcast and to automatically and visually document physical chess games or tournaments. Moreover, being able to track a chess game electronically can lead to the development of robots or other products that can play chess physically for educational purposes or learning for children. The thesis is the continuation of the work done by Santiago del Rey in his Master’s thesis [12]. Although its main objective is to study the environmental footprint when training Deep Learning models, it develops a computer vision prototype for broadcasting chess that detects the board and classifies its cells into occupied/empty cells. Our project wants to take the prototype to the next step and be able to recognise the kind of piece in each cell, therefore making a more reliable system. 7 1.2 Goals of the project The main objective is to improve the live chess recognition system developed with Tensorflow in Santiago del Rey’s thesis, which detects the current state of a chess game from images and the knowledge of the previous state. Del Rey’s prototype aimed to detect a chess game using a occupancy classifier and the knowledge of the previous state. Knowing the previous state, means that we have information about what cell is in each position in the board and what player gets to move. Assuming that there are no illegal positions, we can extract the possible moves. Then the occupancy classifier was ran in the squares that could have changed it state and compared the output with the knowledge of the input. This way the system detected which move the player did. Nevertheless, the system had no way approach to detect a chess piece from the image and relied completely on the previous step, which made the system very sensible. Given the limitations of detecting chess pieces from Del Rey’s system, we propose to give more importance to the recognition of the pieces. Our approach consists in building a chess piece classification, which given an image of a chess piece, it returns the type of piece it is. Once we build the model, it is important to obtain insights about its performance in production, where given an image of a chess board the system is able to retrieve all the chess pieces and its positions in the board. As the major objective is to improve the system, it is important that our approach is consistent when brought to production. All things considered, we present two objectives in order to, eventually, improve the live chess recognition system. 1. Create a deep learning model aimed to solve the chess piece classification model. Validate the models in the test set. 2. Testing the best model in production. Improve the model in production and try to reduce the latency as much as possible while not losing accuracy. 1.3 Structure of the document The document is structured as follows. Chapter 2 reports the background needed to understand this thesis. Chapter 3 sheds some light on the previous work done in the field, what has to be done and what is addressed in the thesis. Chapter 4 shows the method followed during the project for building the model for the first objective and for bringing the model to production. Chapter 5 and 6 present the results of the experiments. After presenting the results, Chapter 7 and 7.2 present different insights about the results of the project and limitations found throughout the process. Finally, Chapter 8 presents the thesis conclusions and proposes possible extensions to the project. 8 4.1 Datasets As the requirements are different for each objective, we will use two different datasets: one for the training and another for the production testing. 4.1.1 Chess piece classifier dataset To be able to train our computer vision system, we require a dataset of labeled images of chess pieces from different games and chess sets. We began to train our model in the dataset used in del Rey’s work, which consisted of a dataset published by Quintana et al, a collection of images published in Roboflow 4and a small set of images from a chess tournament. However, the quality of the images was not good and we switched to a better quality dataset. The dataset consists of images of board squares with their labels (the piece in the square, if the square were to be empty, its label was ”empty”). In total, the dataset contains 50,000 images belonging to 13 classes. The dataset is not balanced: the majority of the classes contain over 2,000 images of each color, however, the pawns classes are larger (over 12,000 images of each color), and the number of rooks is also slightly larger (more or less 3,000 images per color) and the queens’ number of images is lower, over 1,400 images per color. The training of all the models was done with the latter dataset, while the former was used once to compare the difference in accuracy when training the same model with different datasets. This was made to express the importance of quality data when training a model. 4.1.2 Production test dataset To bring the model to production we now need images from the whole chessboard, not only the squares. For this, we have recreated 10 chess games played by Magnus Carlsen 5, taking a picture of each move and saving its FEN, as we needed the board configuration to evaluate the predictions. The games were taken with different lighting and perspectives, to see how the model generalizes to unseen chess sets and different environments. There were images more clear than others, and the black figures were sometimes difficult to distinguish. 4.2 Chess Piece Classifier To build the different classifiers, we have applied transfer learning with the following CNNs. They have been continually studied and improved to perform well even for large image classification [15], so we expect that they will be able to generalise our data with a high performance. •Mobilenet V2 [13]: optimization models result in very complex networks and the objective of Mobilenet V2 is to decrease the number of operations and memory needed while retaining the same accuracy. That is, creating the simplest possible network. The architecture consists of an initial fully convolutional layer with 32 filters and 19 residual 4https://public.roboflow.com/object-detection/chess-full/23 5https://en.wikipedia.org/wiki/MagnusCarlsen 15 Figure 5: An Example of two images from the test dataset. Its FENs would be: 4r1k1/p4r2/2pBp1nb/2PpPp2/q2N1Ppp/P2Q2P1/2P4P/1R3RK1 b - - 0 (left) and 8/b7/2p3k1/p2p4/p2P1BP1/1Pr3P1/3RK3/8 w - - 0 44 (right). bottleneck layers. The non-linearity factor is ReLu6 and it uses a 3x3 kernel. It can be adjusted on desired accuracy/performance trade-offs with tunable hyperparameters. •Xception: It is based on the well-known Inception architecture, introduced by Szegedy et al. [1] in 2014, which has been one of the best-performing family of models on the ImageNet dataset. Xception is based entirely on depthwise separable convolution layers. The feature extraction base of the network is formed by 36 convolutional layers. Those are structured into 14 modules, all of which have linear residual connections around them, except for the first and last modules. The Xception architecture is a linear stack of depthwise separable convolution layers with residual connections. This makes the architecture very easy to define and modify. •VGG16: It has a typical CNN architecture, but uses a small convolution filter size and then uses the now freed-up space to make the network really deep. It has 13 convolutional layers with 3x3 kernels (the smallest size that still captures the notion of up/down, leftright, center) and some maxpool layers in between, and then 3 fully connected layers at the end. In total, it has 16 weight layers. •NASNet Mobile: The main idea of this network is to search for an architectural building block on a small dataset and then transfer the block to a larger dataset. In NASNet, the blocks or cells are not predefined by authors. Instead, they are searched by a reinforcement learning search method. •ResNet-50: It was developed due to the observation that adding more layers to a neural network does not always improve the results. ResNet-50 consists of 50 layers divided into 5 blocks, each containing a set of residual blocks. Those blocks allow for the preservation of information from earlier layers, which helps the network to learn better representations of the input data. The models are trained in two steps: During the first xepochs, the model is trained with all the layers frozen except the classifier. Then all trainable layers are unfrozen and we fine-tune the model for another yepochs. In the majority of times, we used 20 epochs for training and 16 80 for fine-tuning, having a total of 100 epochs. To train the models, we divided the data into 70% training and 30% validation. 4.3 Evaluation metrics 4.3.1 Multiclass Confusion Matrix A confusion matrix is a table that allows visualization of the performance of a classification algorithm. Each row of the matrix represents the instances in an actual class while each column represents the instances in a predicted class (or vice versa). This way it makes it easy to distinguish the classes that tend to be misclassified (and in what class the classifier tends to classify them) and the ones that are classified correctly. These are some of the key components of the confusion matrix: •Condition positive (P): the number of real positive cases in the data •Condition negative (N): the number of real negative cases in the data •True positive (TP): test result that correctly indicates the presence of a condition or characteristic •True negative: a test result that correctly indicates the absence of a condition or characteristic •False positive (FP), type I error: a test result which wrongly indicates that a particular condition or attribute is present •False negative (FN), type II error: a test result which wrongly indicates that a particular condition or attribute is present. The diagonal elements in the confusion matrix represent the correctly classified instances for each class (TP). The off-diagonal elements represent the misclassified instances, where the actual class differs from the predicted class. These are further categorized into false positives (FP) and false negatives (FN). From the confusion matrix, we can calculate various performance metrics, such as accuracy, precision, recall, and f1-score. These metrics help assess the model’s strengths and weaknesses in classifying different classes. A multiclass confusion matrix can assess the performance of a multiclass classification model and help in identifying which classes the model excels at classifying correctly and which classes struggles with. 4.3.2 Accuracy Measures the overall correctness of the classification model. It is the ratio between the correct predictions and the total of instances. It is a rather shallow metric, however, it gives a good perspective on how well a model is performing on a given dataset. The accuracy alone does not tell the full story when we have an imbalanced class dataset. 17 4.3.3 Precision and Recall •Precision: it is a metric used to evaluate the accuracy of a model’s positive predictions. It indicates that then the model predicts a positive outcome, it is generally correct. It is calculated as the ratio of TP to the sum of TP and FP. Precision =TP/(TP +FP) •Recall: measures a model’s ability to identify all instances of a positive class. High recall suggests that the model effectively captures a substantial portion of the actual positive cases. It is calculated as the ratio of TP to the sum of TP and FN. Recall =TP/(TP+FN) To evaluate the effectiveness of a model we must examine both. Precision and recall are often in tension: when we increase recall, precision decreases, and vice versa. We have to search for the best trade-off between precision and recall, depending on the specific objectives and constraints to the task. High precision is important when FP have a negative impact, while high recall is crucial when FN are costly. There are other evaluation metrics that bear both precision and recall in mind, like the f1-score. 4.3.4 Loss It is a measure of error between the model’s prediction and the actual target values. The lower the loss, the better the agreement between the model’s predictions and the actual values. If the training loss is higher than the validation loss, it is a sign that we may be overfitting the model. Various functions can be used, in this project we are using categorical crossentropy. Categorical crossentropy measures the sissimilarity between the predicted probability distribution and the true one-hot encoded label. It is commonly used in multi-class classification problems where each piece of data belongs to exactly one class. It is calculated as follows: L(y,p) = −(yi∗log(pi)) 4.4 Improvement techniques When training a model, we can encounter several challenges such as overfitting, vanishing gradients or convergence issues. In order to tackle those problems and improve the overall performance of our model, we are going to use different configurations of improvement techniques. The improvement techniques we are going to use are: L1 and L2 regularization The capacity of the model is limited by adding a regularization term (penalty) to the loss function during training, which penalizes the model for having large coefficients of weights. It encourages the model to learn simpler and more robust patterns. •L1 (Lasso): Add absolute values to the coefficients of the loss function. It reduces the less important features to zero and works well for feature selection if we have a large number of features. It is robust with outliers. Loss =Error(y, ˆy) + λ N X i=1 |wi| 18 •L2 (Ridge): Add the square of the coefficients of the loss function. Adds the penalty as the model complexity increases. Forces the weights to be small but not zero and it is not as robust as L1 regularization to outliers. Loss =Error(y, ˆy) + λ N X i=1 w2 i 4.4.1 Batch Normalization Aims to improve the training of neural networks by stabilizing (adding noise to the activations) the distributions of layer inputs. It is done by introducing network layers that control the mean/variance of the layer distribution. It is used by default in most deep learning models. In the models we use it is used by default in all except VGG16 and in some it is key for the architecture (e.g., ResNet50). The original paper [5] applies batch normalization after the linear transformation of each layer of the network but some experiments also have gotten good results applying batch normalization after the activation function. To sum up, batch normalization brings training stability to the method, and reduces sensibility to weight initialization. 4.4.2 Dropout [16] Consists of randomly dropping units and their connections from the neural network during training It prevents units from adapting too much to the training data, which could produce overfitting. Overfitting is caused when the neural network learns some complex relations that are the result of sampling noise. 4.4.3 Global Average Pooling (GAP): Compresses the spatial information within each feature map into a single value per channel. this way we get a compact representation of feature maps. It helps in reducing the number of parameters while retaining important information. It is commonly as the final layers of a model: after the last CNN we apply GAP and Fully Connected layers. It is suitable when the presence of a feature matters more than the spatial structure of the features. This is the reason why we are using GAP in our models instead of flattening the spatial dimensions to 1D, the latter preserves the full spatial structure, which is not essential, has a higher risk of overfitting, and requires a larger number of parameters. 4.4.4 Hyperparameter tuning: Its objective is to select the best set of hyperparameters. It aims to strike a balance between complexity and performance and it is essential when building ML models that perform realworld tasks. The choice of hyperparameters can significantly impact a model’s effectiveness. Some of the hyperparameters that can be tuned are the number of epochs, batch size, learning rate, optimizers, image size, and loss function. 19 4.5 Production: Broadcasting chess system The second objective regards testing the models in production. That is being able to build a broadcasting chess system. Although the optimization of the system has many components we are not focusing on in this thesis, we will shortly explain how our system works. This second objective will be tested in the second dataset, the one containing images of the full board 4.1. It aims to give more insight into how the model works in production: compare our model with state-of-the-art projects and detect possible flaws and aspects to improve for future work. 4.5.1 Broadcasting system Ideally, the system should be able to produce a succession of FENs from a chess game live recording. However, due to the limitations of the project, we have reduced this system to be capable of transforming an image of a chessboard into a FEN. The system has two main requirements. It has to be fast enough to be able to capture all the players’ moves and it has to reduce the number of failures that occur as we need very accurate results. Figure 6 displays the method used in production. 1. Board detection algorithm: We detect the location of the board on our input image and warp the image so that each board has the same size. We rely on the algorithm used by Santiago del Rey to perform the corner detection and square cropping. In an optimal situation, the detection of the location of the board should only be needed to execute once, as if we have static conditions (camera and board do not change positions throughout the game) we can reuse the coordinates of the location of the board. After the detection of the board, we have to perform square cropping, hence extracting an image of each of the 64 squares in a board, as it is the input needed when classifying the chess pieces. 2. Prediction in production: To predict all squares, we rely on a brute force algorithm, we make no use of the domain information we could have like the previous board state of the match or the probabilities a piece has to be in a specific square. That said our algorithm gets all squares and does the prediction of each square in a way that, at the end, we have the location of the square and its prediction. The prediction is done with the chess piece classifier models developed in the thesis. In the next chapter, we report the best trade-off between accuracy and time for each model and select the best. The final output of our system will be the configuration of the board in FEN notation. FEN is short for Forsyth-Edwards Notation, and it is the standard notation to describe the positions of a chess game. This notation can be easily imported into different programs to visualize a digital chessboard as it consists in a single string of ASCII characters, which lets computers process it. Each string has six different fields, each describing one aspect of a position and separated by a space character. The six FEN fields are: •First field: It represents the placement of pieces. It consists of a series of 8 blocks of alphanumeric characters representing the pieces and empty squares of a row, each block separated by the symbol ”/”. 20 Figure 6: Production pipeline –The board is read from left to right and from top to bottom, starting from square a8. –White pieces are named by they initial in uppercase letters and black pieces are named by their initial in lowercase. ”p” stands for pawn, ”r” for rook, ”n” for knight, ”b” for bishop, ”q” for queen and ”k” for king. –Empty squares are denoted by numbers from 1 to 8, depending on how many squares are between two pieces. •Second field: Indicates who moves next. If it is ”w”, it is White’s turn to move, while ”b” means that it is turn for Black. •Third field: indicates if the layers can castle and to what side. Uppercase come first to indicate White’s castling availability, followed by lowercase letters that represent Black’s. The symbol ”-” indicates that neither side may castle. •Fourth field: If there is a square that is a possible target for an en passant capture, the square is added. The symbol ”-” represents that no en passant targets are available. •Fifth field: Informs how many moves both players have made since the last pawn advance or piece capture. •Sixth field: Number of complete turns in a game. It is incremented by one every time Black moves. Nevertheless, as we only have an image available at once, we will only be able to retrieve the placement of pieces. Thus, the output encoded as the first field of the FEN notation, for instance: rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR. 21 In pursuance of choosing the best model to bring to production, we collected the board accuracy, the piece accuracy, the position accuracy, and the time. This way we can compare the trained models in different aspects. We can also examine which pieces the model struggles the most to detect and if some board positions are problematic. 22 5. Results of training the piece recognition model (Objective 1) 5.1 Baselines We executed all 5 models under the same circumstances: •20 initial epochs + 80 fine tune epochs (although we had early stopping so in some models it does not reach the 100 maximum epochs) •10−5initial learning rate + 10−6fine-tune learning rate •Image size: 128x128 •Batch size: 16 •ImageNet weights In Table 2 we can see the results from the baseline models. Train dataset Test Dataset model loss acc prec rec loss acc prec rec vgg16 1.114 0.239 0.794 0.044 10.4 0.07 0.07 0.067 xception 0.751 0.427 0.84 0.33 2.5 0.22 0.3 0.16 resnet50 0.866 0.39 0.89 0.31 2.46 0.17 0.19 0.15 mobilenet v2 1.075 0.241 0.834 0.024 3.35 0.19 0.33 0.07 nasnet mobile 0.945 0.341 0.85 0.13 3.77 0.24 0.29 0.18 Table 2: Baseline results We obtained poor results in all 5 architectures. With the best configuration (xception) we reach 22% accuracy and with the worse, we do not even reach 10%. We see the same type of pattern in all the executions: although we achieve a high precision in the test dataset the difference with the recall is huge, which drops the train accuracy to less than 50%. In the test dataset both precision and recall drop drastically. One of the causes of the poor results could be a lack of samples, which could be solved by training with more data. However, this is not possible as when we try to train with more data we face computational limitations and we are not able to finish training (even if we lower even more the batch size). 5.2 First experiments As mentioned above, the main problems encountered with the transfer learning models were overfitting (high precision on training, which dropped drastically on validation) and recall. Although VGG16 was not the model with the best baseline accuracy we started by adding some regularizations to that method. The decision was made this way as VGG16 is the most 23 basic architecture above the 5, which caused the changes done to be simpler to detect and easier to know what had good results and bad. VGG16 was composed just of CNNs, Poolings, and activation functions, it had no regularization techniques while other models already had those techniques. We applied L1 and L2 regularization in the final layer (the one that made our predictions). We tried different lambdas between (0.05 and 0.1). The results are summarized in Table 3: Train dataset Test Dataset model loss acc prec rec loss acc prec rec L1 0.05 1.114 0.239 0.794 0.031 7.129 0.232 0.246 0.226 L2 0.05 1.1 0.226 0.84 0.02 5.97 0.055 0.057 0.051 L1 0.1 1.32 0.067 0.1 0.03 7.4 0.006 0. 0. L2 0.1 1.187 0.176 0.634 0.003 5.83 0.157 0.175 0.095 L1 0.05 + L2 0.1 1.189 0.179 0.333 0.02 4.51 0.29 0.33 0.24 Table 3: Results with L1 and L2 regularization Given that the results in the baseline model did not even reach 10% of overall accuracy (worse than random), L1 and L2 regularization managed to improve the model up to reaching 29% accuracy. However, we detected that recall was higher on the test data than on the train data. This is not a good sign since the model is optimized for the train data, not the test. Instead, when we get significantly better results on the test set compared to the training set it may mean that the underlying distribution on the train and test is not the same and that the model is underfitting to the harder patterns. In other cases, we had other interesting results. The training precision and recall were 0 and the accuracy was nearly 0. We see this case in a high value of lambda for L1, which agrees with the case that an unreasonably high degree of regularization may be applied and end up with a sparse feature vector. When data is highly sparse, then the network learns ’zeros’, hence there is no real learning happening. After regularization, we tuned the hyperparameters. When tuning the hyperparameters, we came across some computational limitations. When we wanted to increase the image size, we had memory capacity issues and could not try that path. We wanted to increase the image size as it would require less downscaling. Downscaling makes it harder for the model to learn the features as the image details will be significantly reduced. Increasing the image size would lead to less deformation of features and patterns inside the image. The batch size used in the baselines was 16, which is a small batch size and has the risk of introducing randomness. Increasing the batch size to 32 would have the purpose of increasing the computational efficiency. As we increased the batch size, we also modified the learning rate: we executed with the default learning rate and then we doubled it. We executed the two new hyperparameter configurations with the best model up until now (regularization). Table 4 display the results achieved. Although the results were quite similar to the executions with smaller batch sizes, we still managed to increase the accuracy up to 32%. From now on, all the executions will have batch size 32 and the default learning rate. The last of the first experiments was to modify the last layers of the model. The last layers are the ones where ideally the model extracts the specific features for the task. By default, 24 model b k n p q r empty vgg16 18.58 0.98 25.93 37.44 34.38 21.18 58.60 resnet50 7.96 4.90 28.89 42.50 25.00 28.24 63.64 mobilenet v2 7.08 0.98 22.22 69.98 51.56 22.94 47.86 nasnet mobile 20.35 0 43.7 38.11 34.38 39.41 42.07 Table 9: Piece accuarcy results for black and empty pieces model B K N P Q R vgg16 40.43 0 58.14 14.16 20.08 72.82 resnet50 40.43 0 61.63 11.30 18.46 66.67 mobilenet v2 15.96 0 13.95 40.06 35.38 31.28 nasnet mobile 7.45 9.80 25.5 32.38 16.92 47.69 Table 10: Transfer learning results 31 Figure 9: Board heatmaps for each of the models 32 7. Discussions and Limitations 7.1 Discussions On the one hand, we have accomplished our first objective: the construction of a chess piece classifier good enough to be brought to production. We achieved a 98% accuracy in the chess piece classifier, which is comparable to the results Del Rey [12] obtained in its work with the occupancy classifier. It is difficult to compare the results of this part of the thesis with the previous work done on this topic as they cannot be directly compared. W¨olflein used a synthetic dataset, which increased the quality of the data and was shaped around just one type of chessboard. It was also tested in the same type of chessboard so we do not have knowledge to what extent the model can generalize to different types of chessboards. On the other hand, we had trouble achieving the second objective. Our production model was much slower than Mallas´en’s [2] system and did not achieve a good enough accuracy. Although we encountered various limitations on the test dataset, we connected the differences in accuracy between the model and the production results to one main aspect, domain information. The previous research all made use of domain information. Our system was not limited in any aspect, we used brute force. This has its advantages as it does not need additional knowledge of previous moves nor limits the possible pieces in a cell so ’uncommon moves’ are not out of the equation. However, this causes that we have a less robust model that relies solely on the classifier. If the model is not capable of classifying the pieces with 100% accuracy we can have misclassifications due to low-quality images or low-precision square crops. Our production results underline the importance of domain information. If we had used one of the most simple domain information which is getting the previous state of the board and getting a list of possible moves, we could have increased accuracy and reduced time as we would not have to predict all the squares and we could limit the predictions to the possible pieces that could be in the square. As we have mentioned previously, one of the main setbacks of this thesis model is the poor accuracy achieved in production compared to Del Rey’s occupancy classifier. Another way to improve the model’s accuracy would be using both classifiers, as the occupancy classifier is much more reliable when detecting if a square is empty or occupied. The system could work in a similar way as the method W¨olflein [17] proposed: Board detection, occupancy classification, and piece classification. We would add domain information in both classifiers, in the occupancy classifier to get the cells candidates to have changed state (from empty to occupied and vice versa) and in the piece classifier to get the possible pieces that can be in the cells. Nevertheless, adding domain information makes the model more restrictive and unable to correct mistakes. If we use the information of the previous move to predict the actual board and this information is incorrect, we are dragging mistakes that may not be reversed. Thus, it could be interesting to see to what extent a hybrid system could work: use domain information but in some executions run brute force to prevent mistakes from persisting and being able to broadcast the chess game as true as possible. Throughout the training of the models we extracted the energy consumed while training, in Table 7.1 we report this energy. Regarding emissions and energy consumed, both ResNet50 and Mobilenet v2 are the models that consume less. It is shocking the energy consumption of the NASNet model, which is nearly 10 times more than ResNet50, the model with less consumption. The same thing happens with the emissions, the difference between NASNet 33 and ResNet50 is very relevant, being the latter 10 times better than the first one. Another observation is that there seems to be no connection between the number of parameters and the energy consumed. VGG16 and ResNet are the models with most number of parameters, and, although VGG16 is not the most energy efficient, ResNet is. Although it has the most number of parameters, more than 26 million, it is slightly more energy efficient than Mobilenet, the model that comes in second in terms of energy efficient and the model with least parameters, just 1,3 million. model Emissions (CO2e)energy consumed num parameters VGG16 0.588 3.037 14,721,357 ResNet50 0.237 1.226 26,217,357 Mobilenet v2 0.296 1.52 1,368,765 NASNet mobile 2.15 11.11 4,817,569 Table 11: Training energy report We also have generated an efficiency label 6of the final machine learning model, ResNet50. The label is calculated considering factors such as CO2 emissions, model size, reusability, and performance. It goes from A to Our model size was 274 MB and the dataset where we trained the model occupied 940.0 MB. Figure 10 displays the efficiency label for our model, which is a C. C means that the model is nor highly efficient or inefficient. We see that the model size is too large, as it is color red, meaning it is not energy consuming friendly. CO2emissions and the dataset size are not also optimal, however they are much more better results than the model size. If we were to improve our model in terms of energy consumption, we would aim to reduce the model size without loosing accuracy. Also reducing the dataset size could be another good approach: the dataset was not balanced, as we had the same number of pawns as all the other pieces combined. We could train the model in a balanced dataset which would be smaller in size and test if we do not loose accuracy. Regarding the process of training and tuning the models, we learned the importance of not leaving anything to guess. For a long time, we were getting results that has over 30% accuracy in the best cases. The results showed that the models were not able to achieve more than a 30% even on the test data. We applied different methods that are known to improve the generalization of the models, however the results did not improve enough to consider that the model could be brought to production. During all the time we thought that the issue was in the model architecture when it was not. When we build a model, we have to keep in mind that the model is not only the DL architecture but all the processes. The quality of the dataset, how the dataset is constructed, how it is loaded, the inputs of the architecture, the architecture itself, hyperparameters, OR which evaluation functions we use. They all contribute to the final product. For a considerable time, we kept on trying new changes in the architecture and getting very poor results. During this time we kept thinking that the source of the problem was either the quality of the dataset or the architecture, thus the only part where changes were made was in the architecture. However, even though a model has no apparent problems in the process of training but produces really poor results, it does not always mean that the problem is in the architecture itself. Yes, overfitting or underfitting may be more common problems 6https://energy-label.streamlit.app/ 34 Figure 10: Efficency label for the best method, ResNet50 but that cannot tear us apart from thinking outside of the architecture and looking into the parts of the process that apparently work just fine. In our case, the issue was with the creation of the dataset, at the beginning we first created the dataset and then batch it. The manual process of batching was not correct, so we changed to batch the dataset when the data was fed to the model, not when loading the data. If we had not restricted ourselves to trying to solve the problem where we thought it could be and expanded the search through the whole process, we would have found the issue in a much faster way. Here we remark on the importance of critical thinking and not guessing that the problem will not be something we have not proven. 7.2 Limitations Throughout the training of the models, we have come across different challenges when trying to train the architectures with the available dataset. The models we were training used from 1 million parameters up to 24 million (ResNet50). Additionally, we firstly had a dataset 35 compound by 100,000 images and although we were training on UPC’s rdlab HPC service, we were limited to using one node and the dataset was too large and the resources we could request were not enough and we had trouble as we had a lack of memory. To be able to execute it, we reduced our dataset. It initially had over 50,000 images of empty squares, which we reduced to 2,350. This reduction had its positive impact as the reduction of the dataset without accuracy loss contributed to a greener AI model. [3] With this, we had no more trouble to execute. The limitations in resources have also affected the construction of the models, mostly when fine-tuning the model. Although that process ended up not being part of the final model, if we wanted to fine-tune a model we had limitations in terms of image size or batch size as increasing those hyperparameters meant the model made use of more resources than available. When bringing the model in production we have encountered different problems that have caused some poor results. In the time extraction, when we printed the time we were forcing the process to go slower, although it was not the main limitation. The main limitation in the optimization of the execution was not being able to parallelize the predictions of the squares. Our method for predicting the pieces was far from optimal as to achieve the predictions we first had to extract all squares and then predict the outcome one by one. If instead we could be able to predict the pieces in batches the execution time would be notably reduced. For example, it takes 0.1 to predict a square and we are able to predict in batches of 16. Then, instead of taking 0.1 ∗64 = 6, 4 seconds, we would be taking 0.1 ∗64/16 = 0.1 ∗4 = 0.4 seconds plus some time to load the images, etc. With this, we would reduce considerably execution time. Other limitations found in model production were the cropping of the images and the angle at which we get the images. The cropping of the images was not optimized and made that, sometimes more than a square is shown in the image. As we build the model recalling more importance on the presence of features and not the location of features, it is possible that the prediction gives the result of other squares. The square we want to predict is always the bottom left one, it starts at the corner. In Figure 11 we display some examples. In the first image, we see that the piece we want to predict is the white king. However, we also see the white queen and part of a bishop. The second image is the same as the first one. Here we can see the fact that it is much more difficult to see which piece appears in the image if the piece is black compared to the white ones. The third and the fourth images are both images where the cell that has to be classified is empty but a cell next to it has a piece, which can confuse the model. This said, it could be interesting to crop the images in another way or use a different angle in a way that only the square that we want to predict is shown, with no occlusions. Figure 11: Examples of images extracted from the board detection and square cropping 36 8. Conclusions and future work Throughout this bachelor thesis we have discussed different deep learning models for the task of classifying chess pieces and we have attempted to bring the model to production without the use of any additional information about the chess game, just an image of a chessboard. Our system consisted of performing brute force through all chess squares and predicting the piece in each one with a chess piece classification model. We successfully created a deep learning model aimed at solving the chess piece classification model, however due to the lack of previous studies we are not able to directly compare our solution to previous research. Moreover, we did not get satisfactory results when testing the model in production as we did not improve the accuracy. Our results show that it is possible to build a deep learning architecture that classifies chess pieces from an image, even though this classifier may not be enough to build a system to broadcast a chess game. Given the thesis results and the systems proposed by W¨olflein et al., Mallas´en et al. and Del Rey, the thesis highlights the importance of the use of domain knowledge in the tasks and how it can make an impact in terms of time and accuracy when building a system for the task of broadcasting a chess game. Overall, the thesis contributes to the work done by W¨olflein et al., Mallas´en et al., and Del Rey in attempting to create an end-to-end framework that is able to broadcast a chess game using DL. For future work, we would like to extend this system by adding domain information with the objective of constructing a more reliable framework to broadcast chess pieces from a chess game. In Chapter 7, different approaches are presented, such as using the previous board information to extract the legal moves or using two classifiers (occupancy and piece) in addition to domain information. Given the limitations of the thesis, we only studied a CV technique, image classification. Although no relevant research has been found in the field of chess piece image recognition, it would be an interesting exercise to introduce an object detection approach like YOLO. 37 References [1] Francois Chollet. “Xception: Deep Learning With Depthwise Separable Convolutions”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). July 2017. [2] A. A. d. B. Garc´ıa D. M. Quintana and M. P. Mat´ıas. “LiveChess2FEN: A Framework fro classifying chess pieces based on CNNs”. In: (2020). [3] F. Hoeg and E.L. Christensen. “SAR++: a multi-channel scalable and reconfigurable SAR system”. In: IEEE International Geoscience and Remote Sensing Symposium. Vol. 1. 2002, 671–673 vol.1. doi:10.1109/IGARSS.2002.1025141. [4] Forrest N Iandola et al. “SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and¡ 0.5 MB model size”. In: arXiv preprint arXiv:1602.07360 (2016). [5] Sergey Ioffe and Christian Szegedy. “Batch normalization: Accelerating deep network training by reducing internal covariate shift”. In: International conference on machine learning. pmlr. 2015, pp. 448–456. [6] Anil K Jain, Jianchang Mao, and K Moidin Mohiuddin. “Artificial neural networks: A tutorial”. In: Computer 29.3 (1996), pp. 31–44. [7] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. “ImageNet classification with deep convolutional neural networks”. In: Communications of the ACM 60.6 (May 2017), pp. 84–90. doi:10.1145/3065386.url:https://doi.org/10.1145%2F3065386. [8] Zewen Li et al. “A survey of convolutional neural networks: analysis, applications, and prospects”. In: IEEE transactions on neural networks and learning systems (2021). [9] Dengsheng Lu and Qihao Weng. “A survey of image classification methods and techniques for improving classification performance”. In: International journal of Remote sensing 28.5 (2007), pp. 823–870. [10] Keiron O’Shea and Ryan Nash. “An introduction to convolutional neural networks”. In: arXiv preprint arXiv:1511.08458 (2015). [11] Joseph Redmon et al. “You Only Look Once: Unified, Real-Time Object Detection”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). June 2016. [12] Santiago del Rey Ju´arez. “An analysis of modeling and training decisions for greener computer vision systems”. PhD thesis. UPC, Facultat d’Inform`atica de Barcelona, Departament d’Enginyeria de Serveis i Sistemes d’Informaci´o, Jan. 2023. url:http://hdl. handle.net/2117/385000. [13] Mark Sandler et al. “Mobilenetv2: Inverted residuals and linear bottlenecks”. In: Proceedings of the IEEE conference on computer vision and pattern recognition. 2018, pp. 4510– 4520. [14] J¨urgen Schmidhuber. “Deep learning in neural networks: An overview”. In: Neural Networks 61 (Jan. 2015), pp. 85–117. doi:10.1016/j.neunet.2014.09.003.url:https: //doi.org/10.1016%2Fj.neunet.2014.09.003. 38 [15] Karen Simonyan and Andrew Zisserman. “Very deep convolutional networks for largescale image recognition”. In: arXiv preprint arXiv:1409.1556 (2014). [16] Nitish Srivastava et al. “Dropout: a simple way to prevent neural networks from overfitting”. In: The journal of machine learning research 15.1 (2014), pp. 1929–1958. [17] Georg W¨olflein and Ognjen Arandjelovi´c. “Determining Chess Game State from an Image”. In: Journal of Imaging 7.6 (June 2021), p. 94. doi:10.3390/jimaging7060094. url:https://doi.org/10.3390%2Fjimaging7060094. [18] Fuzhen Zhuang et al. “A comprehensive survey on transfer learning”. In: Proceedings of the IEEE 109.1 (2020), pp. 43–76. 39