Analysis of infrared thermal images for the study of diabetic foot disease Master Thesis submitted to the Faculty of the Escola T`ecnica d’Enginyeria de Telecomunicaci´o de Barcelona Universitat Polit`ecnica de Catalunya by David Cubero Valent´ın In partial fulfillment of the requirements for the master in Advanced Telecommunications Technologies (MATT) Advisors: Veronica Vilaplana Besler Montse Pard`as Feliu Barcelona, January 2024
Contents List of Figures 4 List of Tables 5 1 Introduction 8 1.1 Motivation.................................... 8 1.2 Objectives.................................... 9 1.3 DocumentStructure .............................. 9 2 State of the art & Background 10 2.1 Thermographic Camera: FLIR Ex-Series E6-XT . . . . . . . . . . . . . . . 10 2.2 ImageSegmentation .............................. 11 2.2.1 Semantic Segmentation: U-Net . . . . . . . . . . . . . . . . . . . . 11 2.2.2 Instance Segmentation: Mask R-CNN . . . . . . . . . . . . . . . . . 12 2.3 ImageRegistration ............................... 14 2.3.1 SimpleElastix Library . . . . . . . . . . . . . . . . . . . . . . . . . 14 2.3.2 FeatureMatching............................ 15 2.3.2.1 SIFTDetector ........................ 16 2.3.2.2 FLANN Matcher . . . . . . . . . . . . . . . . . . . . . . . 17 2.3.2.3 Homography Matrix . . . . . . . . . . . . . . . . . . . . . 17 2.4 IRTLiteratureReview ............................. 18 2.4.1 Point-by-point Analysis . . . . . . . . . . . . . . . . . . . . . . . . 18 2.4.2 ROIsAnalysis.............................. 20 2.5 Medical Background: Pedal Acceleration Time (PAT) . . . . . . . . . . . . 21 3 Methodology & Project Development 24 3.1 ImageAcquisition................................ 24 3.2 ImageExtraction................................ 25 3.3 Image Scaling & Cropping . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 3.4 DataCollection................................. 27 3.4.1 Manual Collection of IRT images from healthy subjects . . . . . . . 27 3.4.2 GTMSDatabase ............................ 28 3.5 Thermal Image Normalization . . . . . . . . . . . . . . . . . . . . . . . . . 30 3.6 ImageSegmentation .............................. 31 3.6.1 Instance Segmentation: Mask R-CNN . . . . . . . . . . . . . . . . . 32 3.6.2 Semantic Segmentation: U-Net . . . . . . . . . . . . . . . . . . . . 33 3.7 Imageregistration................................ 35 3.7.1 Pre-alignment & Global Alignment . . . . . . . . . . . . . . . . . . 35 3.7.2 Featurematching............................ 36 3.7.3 SimpleElastix: Affine Registation . . . . . . . . . . . . . . . . . . . 37 3.8 Point-by-point Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 3.8.1 Thermal Registration and Difference of Temperatures . . . . . . . . 38 3.8.2 Thresholds & Contour Reduction . . . . . . . . . . . . . . . . . . . 40 3.9 ROIsAnalysis.................................. 43 2
3.9.1 Feet Position Correction . . . . . . . . . . . . . . . . . . . . . . . . 43 3.9.2 ROIsDistribution............................ 45 3.9.2.1 MPA, MCA, LCA and LPA regions . . . . . . . . . . . . . 45 3.9.2.2 Deep, Medial and Lateral regions . . . . . . . . . . . . . . 46 3.9.3 Feature Extraction . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 3.10PATAnalysis .................................. 47 4 Results & Discussion 48 4.1 ImageSegmentation .............................. 48 4.2 ImageRegistration ............................... 52 4.3 PATAnalysis .................................. 56 4.3.1 ROIs Analysis - Correlation Task . . . . . . . . . . . . . . . . . . . 56 4.3.2 Point-by-Point Analysis - Classification Task . . . . . . . . . . . . . 58 5 Budget 61 6 Conclusions 62 References 64 3
List of Figures 1 FlirEx-SeriesE6-XT............................... 10 2 FLIR Ex-Series E6-XT display modes. . . . . . . . . . . . . . . . . . . . . 11 3 U-Net architecture, by [3]. . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 4 Mask R-CNN architecture, by [6]. . . . . . . . . . . . . . . . . . . . . . . . 13 5 Image registration concept, by [7]. . . . . . . . . . . . . . . . . . . . . . . . 14 6 Fixed image (left), moving image (right) . . . . . . . . . . . . . . . . . . . 15 7 Initial overlap before registration (left), final overlap after registration (right). 15 8 Example of found correspondences between images, by [9]. . . . . . . . . . 16 9 System’spipelinein[15] ............................ 19 10 Results examples from [16] . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 11 Proportional foot division into plantar angiosomes, by [17]. . . . . . . . . . 20 12 Sample from the plantar thermogram database, by [18]. . . . . . . . . . . . 21 13 Anatomy of pedal arteries, by [20]. . . . . . . . . . . . . . . . . . . . . . . 22 14 PAT measurement example, by [20]. . . . . . . . . . . . . . . . . . . . . . . 22 15 Correlation between PAT and ABI, by [20]. . . . . . . . . . . . . . . . . . . 23 16 Bad image acquisition examples. . . . . . . . . . . . . . . . . . . . . . . . . 24 17 Good image acquisition examples. . . . . . . . . . . . . . . . . . . . . . . . 25 18 Image extraction: RGB image (left), thermal image (right). . . . . . . . . . 26 19 Image scaling & cropping: RGB before scaling & cropping (left), RGB after scaling & cropping (right). . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 20 Examples of defects that do not impact the performance of the system. . . 28 21 Examples of defects that do impact the performance of the system. . . . . 29 22 Examples of the different ranges tested. From left to right: embedded image, 20-38ºC range, 20-40ºC range and 20-35ºCrange............... 30 23 Thermal normalization applied to different examples. . . . . . . . . . . . . 31 24 Mask R-CNN: IoU and F1-Score for training and test datasets. . . . . . . . 33 25 Visual example between resizing and padding. . . . . . . . . . . . . . . . . 34 26 U-Net: IoU and F1-Score for training and test datasets. . . . . . . . . . . . 35 27 Good correspondences example. . . . . . . . . . . . . . . . . . . . . . . . . 37 28 Examples of padded ‘numpy.arrays’. . . . . . . . . . . . . . . . . . . . . . . 38 29 Registration example of thermal information. . . . . . . . . . . . . . . . . . 39 30 Examples of differences of temperatures point-by-point. . . . . . . . . . . . 40 31 Anomaly masks after adding each threshold for three examples: (left) th anomaly, (middle) th anomaly & th not intersected, (right) th anomaly & th not intersected &thminsize .................................. 41 32 Examples of the third exception. . . . . . . . . . . . . . . . . . . . . . . . . 42 33 Anomaly masks reducing contour after adding each threshold for three examples: (left) th anomaly, (middle) th anomaly & th not intersected, (right) th anomaly & th not intersected & th min size . . . . . . . . . . . . 43 34 Feetcorrection,by[17].............................. 44 35 Feet correction example. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 36 Feet division into plantar angiosomes using first approach . . . . . . . . . . 45 37 Feet division into plantar angiosomes using second approach . . . . . . . . 46 38 Mask R-CNN: image samples with their predicted masks . . . . . . . . . . 49 4
39 U-Net: image samples with their predicted masks (using padding) . . . . . 50 40 Examples of ankle/legs appearances before retraining in both perspectives 51 41 Examples of ankle/legs appearances after retraining in both perspectives . 51 42 Predicted masks selected for registration. . . . . . . . . . . . . . . . . . . . 52 43 Registration results of the 3 registration approaches for the Fig. 42 examples. ....................................... 54 44 Registration results of defects that do not impact the performance of the system....................................... 55 45 Correlation matrix of PATs vs mean temperatures of each region for each foot using medical specialists ROIs distribution. . . . . . . . . . . . . . . . 56 46 Correlation matrix of PATs differences vs mean temperature differences between feet regions using medical specialists ROIs distribution. . . . . . . 56 47 Correlation matrix of PATs and PATs differences between regions of the same foot vs mean temperature differences between regions of the same foot using medical specialists ROIs distribution. . . . . . . . . . . . . . . . 57 48 Correlation matrix of PATs vs mean temperatures of each region for each foot using ROIs distribution by [17]. . . . . . . . . . . . . . . . . . . . . . 57 49 Correlation matrix of PATs differences vs mean temperature differences between feet regions using ROIs distribution by [17]. . . . . . . . . . . . . 57 50 Correlation matrix of PATs and PATs differences between regions of the same foot vs mean temperature differences between regions of the same foot using ROIs distribution by [17]. . . . . . . . . . . . . . . . . . . . . . 58 51 Confusion matrices of the best F1-Score for each region. . . . . . . . . . . 60 Listings List of Tables 1 Mapping ABI clinical classification with PAT values. . . . . . . . . . . . . 23 2 Summary of the manual IRT collection from healthy subjects. . . . . . . . 27 3 Summary of all image segmentation experiments performed. . . . . . . . . 50 4 Quantitative registration results for the Fig. 42 examples. . . . . . . . . . . 54 5 Quantitative registration results for the GTMS database. . . . . . . . . . . 55 6 Precision, Recall and F1-Score weighted for Deep region. . . . . . . . . . . 59 7 Precision, Recall and F1-Score weighted for Medial region. . . . . . . . . . 59 8 Precision, Recall and F1-Score weighted for Lateral region. . . . . . . . . . 59 9 Wageandmaterialcosts............................. 61 5
Revision history and approval record Revision Date Purpose 0 15/11/2023 Document creation 1 23/12/2023 Document revision 2 06/01/2024 Document revision 3 19/01/2024 Document Revision DOCUMENT DISTRIBUTION LIST Name e-mail David Cubero Valent´ın
[email protected] Veronica Vilaplana Besler
[email protected] Montse Pard`as Feliu
[email protected] Written by: Reviewed and approved by: Date 19/01/2024 Date 19/01/2024 Name David Cubero Valent´ın Name Veronica Vilaplana Besler Montse Pard`as Feliu Position Project Author Position Project Supervisors 6
Abstract The diabetic foot (DF) poses a significant health risk, demanding improved screening methods in primary care settings to prevent injuries and amputations. Infrared thermal images (IRT) has shown in initial studies to be a diagnostic tool to complement DF screening and stratification. The aim of this project is to analyze IRT data to test if IRT images are useful for the diagnosis. Given their nature, IRT images needs to be treated before exploiting their information. Therefore, first an extensive and complex IRT preprocessing system has been implemented which includes image acquisition, extraction, segmentation and registration, among others. As segmentation model, U-Net has demonstrated to be a good choice while as registration approach, SimpleElastix library provides really promising results using affine transformation. Once IRT images are preprocessed, Pedal Acceleration Time (PAT) has been considered as clinical variable to be analyzed in conjunction with IRT information. Two ways of analysis have been studied: point-by-point and ROIs analysis. Neither of them, has demonstrated any dependence or relationship between IRT information extracted and PAT measurements. Future work may explore additional clinical variables for correlation studies. 7
1 Introduction First, we will introduce the motivation of this thesis by explaining the goal of the global project we are collaborating on, in order to provide context for the framework in which it is situated. Then, we will proceed to discuss the specific precedents that contextualize the thesis, along with the specific objectives and scope within the project that the thesis focuses on. 1.1 Motivation The diabetic foot (DF) is an injury located on the skin and/or underlying tissue, under the ankle, in a person with diabetes mellitus (DM) due to diabetic neuropathy (DN) or peripheral arterial disease (PAD) or both. Among the preventive measures to avoid DF, the implementation of a screening program is crucial. Screening includes foot inspection, assessment of DN, and PAD evaluation. Non-invasive hemodynamic tests are the cornerstone of vascular studies for most patients with chronic lower limb ischemia (a common pathology within PAD cases, which implies arteries obstructions). However, techniques for studying PAD, such as the ankle-brachial index (ABI), are unreliable in patients with significant arterial calcifications, which occur in a third of DM patients. Faced with the need to improve screening tests to reduce injuries and amputations, the use of infrared dermal thermography (IDT) has shown in initial studies to be a diagnostic tool to complement DF screening and stratification. Its use in screening DF risk factors can improve the risk stratification of DF by increasing the detection of subclinical DN and PAD and reducing complications. Taking into account this context, a study conducted by the Grup de recerca en Ferides Cr`oniques i Complexes (GReFeC), the Unitat Funcional de Peu Diab`etic (UFPD), and the Servei d’Angiologia i Cirurgia Vascular (ACV) from l’Hospital Universitari de Bellvitge (HUB) within the Ger`encia Territorial Metropolitana Sud (GTMS), has been initiated in order to evaluate the diagnostic validity of routine IDT in early diagnosis and stratification of high-risk neuropathic, neuroischemic, or subclinical angiosomal disorders (distal angiopathy and distal PAD) in the primary care setting in patients with type 2 DM. The study population includes patients diagnosed with type 2 diabetes (T2DM) from Primary Care Teams (PCT) of the Catalan Health Institute (CHI) in the localities of Gav`a, El Prat de Llobregat, Viladecans, Castelldefels, Begues, and the central and southern areas of L’Hospitalet de Llobregat. The required sample size is 471 patients. By using IDT, it is expected to detect angiosomal disorders in the feet that, due to the characteristics of diabetic patients (calcification), are not detected in routine primary care screening. Improving screening can increase the number of detected cases of cardiovascular secondary prevention in patients with DM (estimated at 30% underdiagnosis of PAD due to false negatives of the ABI technique) [1]. 8
1.2 Objectives The aim of this project is to study whether the capture, preprocessing and analysis of digital infrared dermal thermography images (IDT) or simply infrared thermal images (IRT), performed during the review of patients in primary care centers, can help to detect early the most at-risk cases, helping to prevent complications. This project is under the umbrella of the previously mentioned study evaluation (see Sec. 1.1). This is a collaborative project with the Institut Universitari per a la Recerca a l’Atenci´o Prim`aria Jordi Gol (IDIAPJGol) and the Institut d’Investigaci´o Biom`edica de Bellvitge (IDIBELL) [2]. That being said, the project main goals are: 1. Definition of an image acquisition protocol to guide medical specialists for the image capturing process. 2. Implementation of an IRT image preprocessing system that permits the analysis of IRT information in order to identify anomalous regions, define possible ROIs and extract relevant temperature features. 3. Analysis of the temperature features extracted with clinical variables. This thesis is a continuation of an introduction to research project. Therefore, the following sections were already implemented or started in collaboration with another student (Nuria Amen´abar Garcia): Secs. 3.1, 3.2, 3.3, 3.4.1 (implemented) and Secs. 3.6, 3.7.1 (started). 1.3 Document Structure This document is structured as follows: •Chapter 2. Background & State of the art: presents all theoretical necessary background to comprehend the different stages of the project as well as current state-of-the-art methods and works done related with IRT systems in the context of diabetic foot. •Chapter 3. Methodology & project development: presents the description of the implementation performed for all steps of the process and experiments made. •Chapter 4. Results & Discussion: exposes the results and the analysis of all experiments made. •Chapter 5. Budget: makes an estimation of the total costs involved in carrying out this thesis. •Chapter 6. Conclusions: makes a review and summary of all the project, providing final observations and statements as well as next possible steps for the future. 9
3. Use a matching algorithm to find correspondences between the keypoints in the reference image and the target image. 4. Establish a set of reliable matches between the keypoints in the two images. 5. Based on the matched keypoints, estimate the transformation (translation, rotation, scale, etc.) needed to align the target image with the reference image. 6. Apply the estimated transformation to the target image to align it with the reference image. Figure 8: Example of found correspondences between images, by [9]. In the following subsections we will go deeper into the SIFT detector to find keypoints and extract descriptors, the FLANN matcher to find correspondences between the keypoints from both images, and the transformation estimation using the homography matrix. 2.3.2.1 SIFT Detector The Scale-Invariant Feature Transform (SIFT) detector [10] is a keypoint detection algorithm used in computer vision to identify distinctive and invariant features in an image. The primary strength of SIFT lies in its ability to detect features that are invariant to changes in scale, rotation, and illumination. Following are the major stages of computation used to generate the set of image features: 1. Scale-space extrema detection: the first stage of computation searches over all scales and image locations. It is implemented efficiently by using a difference-ofGaussian function to identify potential interest points that are invariant to scale and orientation. 2. Keypoint localization: at each candidate location, a detailed model is fit to determine location and scale. Keypoints are selected based on measures of their stability. 3. Orientation assignment: one or more orientations are assigned to each keypoint location based on local image gradient directions. All future operations are performed on image data that has been transformed relative to the assigned orien16
tation, scale, and location for each feature, thereby providing invariance to these transformations. 4. Keypoint descriptor: the local image gradients are measured at the selected scale in the region around each keypoint. These are transformed into a representation that allows for significant levels of local shape distortion and change in illumination. 2.3.2.2 FLANN Matcher FLANN (Fast Library for Approximate Nearest Neighbors) [11] is a library designed for efficient and fast approximate nearest neighbor search in high-dimensional spaces. In computer vision, FLANN is commonly used for matching features between images, especially when dealing with large datasets or real-time applications. The FLANN library provides a set of algorithms for approximate nearest neighbor search, and it is often used in conjunction with keypoint detectors and descriptors (e.g., SIFT, SURF, ORB) to find correspondences between features in different images. The matcher aims to efficiently find the nearest neighbors for each feature in one image from the set of features in another image. The main steps involved are the following: 1. Indexing: FLANN builds an index structure on the feature descriptors of the reference set (e.g., keypoints and descriptors in one image). 2. Automatic search algorithm: the system takes any given dataset and desired degree of precision and use these to automatically determine the best algorithm and parameter values. FLANN supports various search algorithms, including kd-trees, hierarchical k-means trees, and other tree-based structures. 3. Matching: given the index structure built on the reference set, FLANN efficiently finds the approximate nearest neighbors for each feature in the query set (e.g., keypoints and descriptors in the other image). The matches represent correspondences between features in the two images. 2.3.2.3 Homography Matrix A homography matrix [12], often denoted as H, is a transformation matrix that relates the coordinates of points in one geometric space to the coordinates of corresponding points in another space. Homography is a projective transformation that preserves straight lines, allowing the mapping of points from one plane to another, even if the planes are not parallel or aligned. In computer vision and image processing, the concept of homography is frequently used for tasks such as image registration, object recognition, and image stitching. The homography matrix is employed to model the transformation between two images or scenes. The most common use of homography is in the context of planar scenes, where the transformation can be described by a 3x3 matrix. The homography matrix Hcan be represented as: 17
H= h11 h12 h13 h21 h22 h23 h31 h32 h33 (1) Given a point (x, y, 1)Tin the source image and its corresponding point (x′, y′,1)Tin the destination image, the homography relation can be expressed as: x′ y′ 1 =H x y 1 (2) In more detail, the homography relation can be broken down into a system of linear equations: x′ y′ 1 =1 h31x+h32y+h33 h11x+h12y+h13 h21x+h22y+h23 h31x+h32y+h33 (3) The homography matrix Hhas eight degrees of freedom, as it is defined up to scale (any non-zero multiple of Hrepresents the same transformation). Typically, homography matrices are computed using methods like Direct Linear Transform (DLT) [12] or robust estimation techniques like Random Sample Consensus (RANSAC) [13], which can handle outliers and errors in point correspondences. 2.4 IRT Literature Review Several studies have been conducted using thermography for diabetic foot purposes. Two main approaches are distinguished: point-by-point and ROIs analysis. 2.4.1 Point-by-point Analysis In this approach, the pixel-by-pixel temperature difference of each of the feet is computed. The image processing includes both feet and is automatically divided into two parts: one part being the left foot and another part being the right foot. During this technique, it is necessary that both feet become aligned, so that, two steps are applied: 1) mirroring one of the feet so that both feet have same morphology; 2) image registration is performed. Image registration allows both feet to be spatially aligned, corresponding to each other. The result is the temperature differences between the right and left foot, and the position of the pixel where the temperature difference is greater than 2.2 °C (in absolute value) is usually taken to identify possible areas of ulcer or necrosis formation. [14] For instance in [15], as segmentation model they fine-tune a Mask R-CNN to extract the foot mask bounding box, find the rotation that obtains maximum similarity between feet 18
bounding boxes, then apply temperature differences between feet and for last, use the anomaly thresholds (+2.2 ºC for ulcers and -2.2ºC for necrosis) in order to detect these possible anomalous regions. Figure 9: System’s pipeline in [15] In [16], the segmentation is performed using feet (skin) color features, then a non-rigid transformation is applied using a landmark-based deformation model with B-splines in order to deal with asymmetries such as amputations, and as a last step, applies the 2.2ºC threshold for diabetic foot complications detection. Figure 10: Results examples from [16] As a last example, in [14], although the authors reference point-by-point analysis, they also introduce an alternative approach. Instead of implementing registration between feet and calculating temperature differences, they solely subtract the temperature of each 19
foot pixel with the mean temperature of that specific foot. Subsequently, as in the other methods, they employ the standard 2.2ºC threshold to detect anomalous regions within the foot. 2.4.2 ROIs Analysis In this approach, each foot is treated separately and divided in ROIs. In [17], a process is built in order to distribute the foot sole in ROIs, according to 4 angiosomes: medial plantar artery (MPA), lateral plantar artery (LPA), medial calcaneal artery (MCA), and lateral calcaneal artery (LCA). Figure 11: Proportional foot division into plantar angiosomes, by [17]. Then, they leverage this foot division, to create a thermal change index (TCI) to provide a quantitative estimation of the thermal changes of the foot caused by diabetes mellitus (DM). The TCI value is based on the mean differences between corresponding angiosomes from a DM subject and the reference values obtained from a control group (healthy subjects). Using these TCI values, they propose a classification system since it is found that a difference of 1ºC is enough to notice a significant differences between classes. In the end, they conclude that the proposed classification can reflect thermal changes providing to the expert a helpful tool that could be used for monitoring thermal changes over time through the proposed index. This ROIs distribution is also used in [18], to build a public plantar thermogram database for the study of diabetic foot complications. This database, saves for each sample (see Fig. 12): the full foot thermogram (image-level) and each angiosome’s thermogram (patchlevel). This database is used by [19], to make some experiments using machine learning (ML) models and deep learning (DL) models for diabetic foot ulcer (DFU) recognition task. For this task, it was demonstrated that using image-level thermogram data provides better results than using patch-level or a combination of image-patch thermograms. 20
Figure 12: Sample from the plantar thermogram database, by [18]. 2.5 Medical Background: Pedal Acceleration Time (PAT) As mentioned in Sec. 1.1, when studying peripheral arterial disease (PAD), current noninvasive arterial tests such as the ankle-brachial index (ABI), absolute pressures in the toes or transcutaneous oxygen pressure (TCPO2), are not reliable in patients with significant arterial calcifications, which occurs in one-third of patients with DM. In addition, these tests do not provide information on perfusion of the different vascular territories or angiosomes of the foot [1]. Therefore, the test of choice in these cases to perform the vascular study is the pedal acceleration time (PAT), which is defined as the time space, on the Doppler curve, between the start of the systole and the point of maximum acceleration, measured in milliseconds (ms). PAT is an ultrasound image direct from several vessels in the foot, specifically 4 arteries (arcuate, medial plantar, lateral plantar and deep plantar), which provides physiological information in real time on their hemodynamic state and can also provide knowledge on the perfusion of the different foot angiosomes [1]. 21
(a) Dorsal (b) Plantar Figure 13: Anatomy of pedal arteries, by [20]. This technique is performed with an echo-doppler standard arterial in an accredited vascular laboratory and does not allow for systematic screening in primary care centers [1]. Figure 14: PAT measurement example, by [20]. In [20], PAT demonstrates a high correlation with ABI in these 4 compressible arteries (see Fig. 15), which means ABI can be replaced by PAT measurements. 22
Figure 15: Correlation between PAT and ABI, by [20]. Therefore, using this correlation, finally they propose the following categories, mapping ABI clinical classification with PAT values. Category ABI PAT (msec) Class I (Normal) 0.90-1.3 0-120 Class II (Mild claudication) 0.69-0.89 121-180 Class III (Severe claudication) 0.40-0.68 181-224 Class IV (CLI) 0.00-0.39 ≥225 Table 1: Mapping ABI clinical classification with PAT values. 23
3 Methodology & Project Development In this section, we are going to explain all implemented parts of the project, including defined pipelines, data collected and experimented methods. 3.1 Image Acquisition All images that are part of this project, are taken using a FLIR Ex-Series E6-XT (see Sec. 2.1 to see a summary of its specifications). Thus, an extensive analysis of camera settings was conducted in order to establish an image acquisition protocol to be followed by medical centers. As important defined settings, it was decided to use the Thermal MSX mode (see Fig. 2a) since it gathers both RGB image and thermal information, and a 0.6m alignment distance since, on the one hand, 0.3m is too close and the error tolerance is very low, and on the other hand, 1m is too far so that it loses useful thermal information granularity and adds more unnecessary background. These issues from the not selected distances can be seen in Fig. 16. (a) Alignment distance: 0.3m, Measured distance: 0.4m (b) Alignment distance: 1m, Measured distance: 1m Figure 16: Bad image acquisition examples. As for the acquisition process and scenario, some guidelines were provided for the medical centers to support the image acquisition process. As remarkable points, the distance between the subject and the camera should be 0.6m, to fit with the camera setting for alignment distance; and, in general, images should be taken using two perspectives: a plantar foot and dorsal foot perspective. As recommendation, a piece of black cloth could be used to cover all or part of the background as a uniform background can facilitate the posterior image segmentation process. A correct image acquisition example, using specified settings, from each perspective is shown in Fig. 17 24
Figure 17: Good image acquisition examples. 3.2 Image Extraction Since the Thermal MSX mode is based on an overlapp between the RGB and the thermal information, then, the resulting infrared thermal image (IRT) acquired is an embedded image (320x240). Therefore, the following step was to extract the desired information from the embedded image. Actually, the embedded image saved in the camera contains the original visual image (RGB) and the raw thermal sensor data in the JPG metadata. In addition, the JPG metadata contains some variables and constants corresponding to the configuration of the camera. To obtain the desired output, a software to extract the metadata from the image needs to be created. Using the Flir Image Extractor [21], which is created with the aim to extract the original RGB image and thermal sensor values converted to temperatures from FLIR ONE cameras, a pipeline was created to, extract the embedded data using this repository and save it. This repository uses ExifTool [22] as key component to extract the metadata information. As a result of this pipeline, the following items are extracted and saved: •RGB image (640x480): obtained with ExifTool. •Thermal ‘numpy.array’ (240x180): contains temperature values in Celsius for each pixel. It is calculated using the image metadata parameters such as ‘Emissivity’, ‘SubjectDistance’, Plank constants, etc. •Thermal image (240x180): a visual representation of the previous ‘numpy.array’ using a heatmap. Taking as acquired image, Fig. 17 (left), an example of its image extraction pipeline defined, is presented in Fig. 18. 25
For those selected, image segmentation ground-truth (GT) masks were created using Segment Anything Model (SAM) [23] and refined manually. SAM was just used for this purpose since it is a high latency approach, which means it could not provide a good inference time if used as an image segmentation model. For last, images and GT masks were distributed into training and test datasets resulting in 51 images and GT masks for training and 12 for testing. In this section, the goal is to describe the implementation of each model. All the results and comparative analysis between models are presented in Sec. 4.1. 3.6.1 Instance Segmentation: Mask R-CNN As stated in Sec. 2.2.2, Mask R-CNN is a state-of-the-art deep learning model used for instance segmentation tasks, which involves identifying and delineating objects of interest within an image. Mask R-CNN generates pixel-level masks to precisely outline the boundaries of each instance. In this case, a Mask R-CNN pretrained on COCO dataset is used [24]. COCO dataset contains 80 different classes, however, none of them is ‘foot/feet’. Thus, in order to be able to segment a new class ‘foot’, we need to train the model. For this purpose, the backbone is frozen and only the last model’s layers, which are the responsible ones for instance classification, are fine-tuned. The whole procedure created for fine-tuning has been inspired by [25]. As for training techniques, input sizes and data augmentation strategies have been compared, such as adding a random rotation or training the model with the original full size images (640x480). The best performance has been obtained with the reduced image sizes (240x180) and adding random horizontal flipping and soft color jitter as data augmentation techniques. As for the hyperparameters, SGD is chosen as optimizer, with lr=0.005, momentum=0.9, and weight decay=0.0005. In addition, a learning rate scheduler decreases the learning rate by 10 every 3 epochs. No hyperparameters tuning has been performed. The decision of the hyperparameter’s values has been made using [25]. The learning curves of fine-tuning Mask R-CNN are presented in Fig. 24. For this purpose, IoU and F1-Score have been considered suitable metrics as they evaluate the performance of the model segmenting between class ‘foot’ and ‘no foot’ at pixel level. 32
Figure 24: Mask R-CNN: IoU and F1-Score for training and test datasets. Both plots clearly illustrate the rapid convergence of both train and test metrics, with minimal fluctuations after an initial sharp increase. Furthermore, the metrics remain consistently close to each other, indicating a successful generalization of the model. The IoU and F1-Score in test exhibits stability from approximately the 6th epoch onward, maintaining a constant value of 0.9136 and 0.9533 respectively. 3.6.2 Semantic Segmentation: U-Net As stated in Sec. 2.2.1, the U-Net is a symmetric network based on a contracting path (context) and an expansive path (precise localization). Typically it is widely used for medical image segmentation. In fact, for the implementation, we have been inspired by a cell segmentation project [26]. In this case, U-Net architecture is used with a ResNet18 as pretrained backbone or encoder, taking the ImageNet weights as pre-training weights. As in the Mask R-CNN case, the ImageNet dataset [27] does not contain ‘foot/feet’ class, so that, fine-tuning must be performed as well. As an input constraint, it was found that the ‘segmentation-models-pytorch’ library, which was the one used, only supports image sizes divisible by 32. The reason is that almost all models have skip-connections between encoder and decoder stages and all encoders have 5 downsampling stages (25= 32) [28]. As our RGB images and ‘numpy.arrays’ are not divisible by 32 (240x180 size), some pre-processing techniques to the input image should be applied. These pre-processing techniques should ensure that the ‘numpy.array’ values are not distorted as they contain the key thermal information. Taking this in mind, two possible solutions were studied: •Resize: apply resizing to the input image to fit with the input constraints and resize the output predicted mask back to 240x180. •Padding: apply the necessary padding to the input image and ‘numpy.array’ in order to fit with the input constraints. 33
In both cases, input images were resized/padded to 256x192 which are the nearest divisible by 32 sizes. As for the training hyperparameters, the values were taken from [26]: lr=0.00068 [optimizer=‘Adam’], batch size=4 [training], loss function=‘Dice Loss’. As in Mask R-CNN, the manually-built dataset (Sec. 3.4.1) is used for fine-tuning the UNet and different data augmentation techniques (horizontal flipping and soft color jitter) have been applied. After fine-tuning, the quantitative results presented a slightly better performance using padding than resizing (see Tab. 3). A visual comparison between resizing and padding can be observed in Fig. 25. (a) Resizing (b) Padding Figure 25: Visual example between resizing and padding. Qualitatively speaking, by looking at Fig. 25, it is subtly noticeable that resizing provides more noisy contours than padding. An specific example of this can be seen in the separation between the big toe and the others in the left foot: in the padding example the separation is clear while it is not in the resizing one since their contours are noisy. In addition, by applying padding, it is ensured that the predicted mask continues being binary and does not lose any quality. This does not happen in the resizing method, since when it resizes back to 240x180 we sacrifice quality to still having a binary mask by using Nearest-Neighbors as interpolation method. Therefore, in the end, padding is selected as solution to the constraint on the input size. After applying inference on the GTMS database (Sec. 3.4.2), we saw that the model does not separate correctly ankle/leg parts from the feet (see examples in Fig. 40). Therefore, more samples for training and testing were added to the manually-built dataset by retrieving not perfect predicted masks from the inference process as the ones depicted in Fig. 40, and refining them for use as GT’s. 34
As a result of this process, 11 new images and masks were introduced in training and 4 in testing. In addition, 2 GT’s that were already in the dataset were refined. Thus, in total, 62 images and masks were used for training and 16 for testing. The learning curves of retraining the U-Net are presented on Fig 26. Figure 26: U-Net: IoU and F1-Score for training and test datasets. The results presented are using early stopping (with patience=8) to avoid overfitting. The best IoU and best F1-Score achieved on testing are 0.9342 and 0.9656 at epoch 20. 3.7 Image registration After predicting the masks of each foot, image registration techniques allows point-bypoint analysis to apply difference of temperatures between feet and extract possible anomalous regions. In order to do that, the predicted mask is split vertically, so that each foot is treated separately and, in this case, the left foot (right split-part) is flipped and used as moving foot for the image registration process. In this section, three image registration approaches are implemented. See Sec. 4.2 in order to find a comparative analysis between them. 3.7.1 Pre-alignment & Global Alignment As first image registration method, we took advantage of another project [29] where image registration was applied to multi-stained digital histological images. In essence, this method is a sequence of two steps: 1. Pre-alignment: makes a first approximation of the optimal registration. 1.1. Find the mass center of each foot. 1.2. Find the translation vector that aligns the mass centers. 1.3. With the aligned mass centers, iterate over different rotations to find the one that provides maximum IoU between feet. 35
2. Global Allignment: refines the previous step to increase the performance. 2.1. Apply the Euclidean distance transformation (each positive pixel is replaced by the Euclidean distance between this positive pixel and the nearest background pixel). 2.2. Find the translation vector that optimizes the superposition of the previous transformation. 2.3. Apply the previous translation vector to the image and recompute the euclidean distance transform. 2.4. Optimize the rotation angle and apply the rotation. 2.5. If the IoU of the transformed image is greater than the one obtained in the prealignment step, returns the translation vector and the rotation angle. If not, the translation vector and the rotation angle are set to 0, so that, the global registration step does not make any change with respect to the pre-alignment step. Some logs were printed during the registration process in order to monitor and compare the IoU after each step. These logs shown that global alignment step never helps to increase the performance of the pre-alignment. This means pre-alignment step by itself provides already good results. 3.7.2 Feature matching As stated in Sec. 2.3.2, feature matching techniques [8] are a way to identify and match the same feature points across two images and build an homography matrix [12] that can be used to make the image registration. Thus, in this case the goal is to match same feature points across both feet. The process works as follows: 1. Descriptors and keypoints are found in each foot using the SIFT detector [10]. 2. Descriptors from both feet are matched using the FLANN matcher [11]. 3. A distance threshold is set in order to filter out bad correspondences. 4. If there is a minimum of 5 good matches, then the matches are used to build an homography matrix. 5. Finally, this homography matrix is applied to the flipped left foot to align it to the right foot. As first attempt, just as with the other methods, the predicted mask is the one used as input. An example of the good correspondences found are depicted in Fig. 27a. 36
(a) Predicted Mask (b) Predicted Mask + RGB information Figure 27: Good correspondences example. As seen, there are some correspondences that, actually, are incorrect which impacts negatively in the computation of the homography matrix. This is foreseeable, since feature matching techniques are usually thought to be used with RGB images that contain much more information than binary masks. Thus, as second attempt, RGB image information was added to the predicted mask, giving much better correspondences than in the previous case (see Fig. 27b). Therefore, a correct homography matrix could be built and used to register both feet. As last step, for visualization and metrics computation purposes, the RGB information is binarized back. 3.7.3 SimpleElastix: Affine Registation As last method tested, we used SimpleElastix, an image registration library. As detailed in Sec. 2.3.1, SimpleElastix is a low-code medical image registration library that makes state-of-the-art image registration really easy to apply it. Following the example detailed in Figs. 6, 7, an affine transformation using this library was tested. The default affine parameter map is used for the purpose. As important parameters, the optimizer used is an Adaptive Stochastic Gradient Descent which requires less parameters to be set and tends to be more robust, and the metric used for the cost function is the Advanced Mattes Mutual Information which is an implementation where the same pixels are sampled in every iteration. Using a fixed set of discrete positions to evaluate the marginal and joint PDFs makes the path in search space more smooth [7]. Affine transformation was selected since scaling and shearing are not present when using rigid transformations (rotation and translation). Scaling and shearing could improve the registration since both feet not should always necessarily be at the same scale or perfectly 37
straight. Non-rigid transformations were discarded because we neither want to deform one of the feet since it could lead to lose relevant information. After applying the affine transformation, the resulting registered mask is binarized. This is because bilinear interpolation is used when mapping the pixels, thus losing binarization. 3.8 Point-by-point Analysis In this section, image registration is leveraged in order to perform a temperature pointby-point analysis between feet with the goal of finding anomalous regions and map those into defined ROIs in later steps (see Sec. 3.9.2). Thus, from now on, only plantar foot perspective is going to be used as the ROIs division is only applied on that perspective. Also consider the U-Net as the image segmentation model used for the rest of the methodology and project development sections left. As seen previously in Sec. 2.4.1, anomalous regions can be considered regions of a minimum size where the difference of temperatures between feet (in absolute value) is greater than a certain threshold. The main steps to perform this analysis are: apply registration to the thermal information, compute difference of temperatures between feet, and finally, set some thresholds and apply contour reduction to filter out irrelevant and noisy information. 3.8.1 Thermal Registration and Difference of Temperatures After image segmentation, the resultant predicted mask does not possess the same size as the ‘numpy.array’ containing thermal information. This is because the predicted mask matches the size of the input image, and it is essential to bear in mind that padding is employed on the input image before it is passed to the model. Thus, as first step, the same padding is employed to the ‘numpy.array’. The following three examples are going to be used to illustrate all the process of this section (see examples in Fig. 28) (a) (b) (c) Figure 28: Examples of padded ‘numpy.arrays’. Remember that, when visualizing thermal information, a normalization step is performed between 20ºC-38ºC (see Sec. 3.5). 38
After that, the next step is to perform the same steps as in the registration process: split vertically the thermal information and flip the left foot (right split-part). Then, as last registration step, the rotation and translation matrix (using pre-alignment & global alignment method) or the homography matrix (using feature matching method) or the transform parameter map (using SimpleElastix method) used to register the predicted masks are now used to register the thermal information. As a result, the moving foot (left) is registered. Once we already have the moving foot registered, now we can proceed to multiply the predicted mask of the fixed and registered moving foot with the thermal information of each, to filter out all thermal information that is outside the foot mask. Taking Fig. 28b as reference, an example of all this part can be visualized in Fig. 29 . Figure 29: Registration example of thermal information. The above process is done in this order because if you first filter the temperatures out of the foot and then you register, then the contour pixels presents distorted values due to the implicit interpolation contained in some of the aforementioned registration methods. To finish this process, the difference of temperatures point-by-point between the fixed foot (right), and the registered moving foot (left) is performed. The same examples as before are depicted now in Fig. 30 to show this last part. 39
(a) (b) (c) Figure 30: Examples of differences of temperatures point-by-point. In order to display this difference of temperatures in a clear way, all values greater than 5ºC or smaller than -5ºC, are displayed as background as well. This is done because the registration methods do not provide completely perfect results as seen in Sec. 4.2. Thus, the parts of both feet that are not intersected with the other, are noise that should be filtered out, since they would present really high differences (in absolute value). This issue is going to also be managed in the next section. 3.8.2 Thresholds & Contour Reduction In this section the goal is to find the anomalous regions. Therefore, first we must define what is an anomalous region and what is not. As stated in the beginning of Sec. 3.8, essentially an anomalous region is a region with a minimum size where the difference of temperatures (in absolute value) is greater than a certain threshold. After setting an orientation threshold (th anomaly=2.2) from Sec. 2.4.1, which is a typical one to find ulcers, it was observed that, there are some exceptions that actually should be considered noise and they should be filtered out. These are: •As mentioned in the previous section, all parts from both feet that are not intersected with the other in the registration process, are considered noise, as they possess high values (in absolute value) because the difference of temperatures is computed with a background pixel (which has 0 value). •As also mentioned, all anomalous regions should have a minimum size. Thus, all regions smaller than a certain size are considered noise. These two exceptions, are solved by setting one threshold for each, in order to filter out both cases. After observing several samples, a good threshold approximation values are: 5ºCof temperature difference (absolute value) for not intersected parts (th not intersected) and 17 connected pixels for minimum size (th min size). Continuing with the examples of the previous section, the anomaly masks after using each threshold can be visualized in Fig. 31. 40
Figure 31: Anomaly masks after adding each threshold for three examples: (left) th anomaly, (middle) th anomaly & th not intersected, (right) th anomaly & th not intersected & th min size As seen, the thresholds are useful to reduce the noise coming from not intersected parts and not enough size. In the second example, this two thresholds already remove all noise and keeps only the actual anomaly. However, this does not happen in the other examples. 41
4 Results & Discussion In this section, all results from the experiments made are presented, including some comparative analysis and discussions. 4.1 Image Segmentation As stated in Sec. 3.6, the purpose of using image segmentation models is to filter out all the background and keep only with the feet regions which contains the useful thermal information to be analyzed in later steps. In order to achieve it, two image segmentation models (U-Net and Mask R-CNN) have been fine-tuned using the manually-built dataset (Sec. 3.4.1). Some predicted masks from both methods are depicted in Figs. 38, 39. 48
Figure 38: Mask R-CNN: image samples with their predicted masks 49
Figure 39: U-Net: image samples with their predicted masks (using padding) Qualitatively speaking, the U-Net, although it seems to catch finer details such as the spacing between toes better than Mask R-CNN, it also shows a worse performance when the background is not uniform, creating artifacts surrounding the feet. On the other hand, Mask R-CNN seems to behave better under non-uniform background, but seems like the predicted bounding box cuts feet contours such as the heel or toes. This last Mask R-CNN issue was key to decide using U-Net for further processing. As for quantitative results, Tab. 3, shows the IoU and F1-Score obtained in the test dataset for all the experiments made: Models Test IoU Score Test F1-Score Mask R-CNN 0.9136 0.9533 U-Net (resizing) 0.9048 0.9490 U-Net (padding) 0.9123 0.9524 U-Net (padding + retraining) 0.9342 0.9656 Table 3: Summary of all image segmentation experiments performed. The quantitative results are similar, being Mask R-CNN slightly better in the manuallybuilt test dataset. However, after retraining the U-Net, then it presents better results than Mask R-CNN, even though this comparison should not be totally taken into account since the test datasets between Mask R-CNN and the retrained U-Net are slightly different (see Sec. 3.6.2). A qualitative comparison between before and after retraining the U-Net can be made by looking at Figs. 40, 41. 50
Figure 40: Examples of ankle/legs appearances before retraining in both perspectives Figure 41: Examples of ankle/legs appearances after retraining in both perspectives As seen, there is a little improvement in the predictions but not enough to completely deal with ankle/legs appearances. The improvement is much more notorious in the dorsal view than in the plantar view. Overall, Mask R-CNN deals better with flexibility than U-Net since it does not contain any input size constraint and provides better results when the background is not uniform. However, taking into account that due to the acquisition process, the background may be uniform (see Sec. 3.1) and the input size constraint is managed by adding padding in the 51
U-Net, then, in the end, the U-Net is considered the best image segmentation model for this task since it catches finer details and does not cut feet contours such as heels or toes. 4.2 Image Registration In order to make a clear comparison between methods and decide which is the best option, 4 different predicted masks have been taken as reference: 2 from each perspective and for each perspective, one easy and the other difficult to register. The predicted masks selected are the following ones: (a) Easy plant example (b) Difficult plant example (c) Easy dorsal example (d) Difficult dorsal example Figure 42: Predicted masks selected for registration. As seen in Fig. 42, the difficult ones are asymmetric feet that should compromise the registration performance. The visual registration results are shown in Fig. 43. 52
53
Figure 43: Registration results of the 3 registration approaches for the Fig. 42 examples. Fig. 43, contains all the registration results of Fig. 42 in the alphabetic order of the examples (from (a) to (d)). The results are presented in such a way that the pink area represents the intersection between the two registered feet. The blue and red areas, on the other hand, represent parts of each foot that do not intersect with the other. So, ideally, we would want the maximum of pink area while the minimum of red/blue areas. If we visually analyze the results, it is clear that feature matching method is the one that provides the worst results, specially in the difficult examples. This is because, probably there are bad correspondences that could not been filtered out by the distance threshold, therefore, these leads to get an incorrect homography matrix. In fact, for each example, the distance threshold has been tuned in order to improve the performance as much as possible, but the results are still very poor. As for the other methods, they seem to obtain similar results in both easy and difficult examples, however, the SimpleElastix method, seems to provide a better fitting in difficult cases using scaling. This hypothesis made by analyzing qualitatively the results, can be justified when looking at the quantitative results for the 4 examples: Registration Methods IoU (a) IoU (b) IoU (c) IoU (d) Pre & Global Alignment 0.9448 0.7778 0.9308 0.7854 Feature Matching 0.9257 0.1022 0.7576 0.3409 SimpleElastix Library 0.9659 0.8062 0.9503 0.8014 Table 4: Quantitative registration results for the Fig. 42 examples. As seen in Tab. 4, using IoU as metric to quantify the image registration performance, it is clear that SimpleElastix method is the best of all. It is slightly superior to the prealignment and global alignment method, probably, since it introduces scaling and shearing apart from rotation and translation. Furthermore, in order to generalize, as last experiment, the entire GTMS database (with the exception of the cases identified to repeat, in order to not harm the performance) has 54
been registered, obtaining in this way the mean IoU score for each registration method. The quantitative results are the following: Registration Methods Mean IoU Pre & Global Alignment 0.8848 Feature Matching (threshold=1) 0.3369 SimpleElastix 0.9296 Table 5: Quantitative registration results for the GTMS database. Note that, for the feature matching method, a general not-restrictive distance threshold has been set to make the experiment, in order to ensure having a homography matrix for all cases. Overall after having observed how each method behave and after generalizing, it can be stated that feature matching is the worst method since it depends on an hyperparameter (distance threshold) which does not fit for all examples. Actually, even if the distance threshold was optimized for all examples, it does not ensure having correct correspondences, therefore it is not an optimal method. As for the pre-alignment and global alignment method, it works well but it has been proved that scaling and shearing can help to improve the registration performance and this method does not deal with that. Therefore, the affine transformation using SimpleElastix library is the best option, since it is a low-code library that supports affine transformation, which avoids deforming the foot (since it is a rigid transformation) while also including scaling and shearing. As an additional comment, in Sec. 3.4.2, we mentioned how registration mainly can deal with those treatable defects. In Fig. 44, we depict the registration results of Fig. 20 treatable defects examples. (a) Asymmetric feet (b) Camera rotation (c) Different heights (d) Different inclinations Figure 44: Registration results of defects that do not impact the performance of the system. As seen, in overall, the registration results are pretty good for all cases, even thought it 55
is noticeable that the results for the asymmetric feet case are worse than the rest, since there are more non-overlapped parts. However, this is not a problem since the majority part is overlapped, so that, not much valuable information will be lost when filtering out non-overlapped parts in the posterior analysis (see Sec. 3.8.2). The same case would happen if the asymmetry is due to the amputation of toes. The amputation would make the comparison between these affected toes useless, so that, the non-overlapped parts would not provide valuable information. 4.3 PAT Analysis As detailed in Sec. 3.10, two approaches have been studied in order to see if there is any dependence between PAT values and thermal information. In this section, the results of both approaches are presented and discussed. 4.3.1 ROIs Analysis - Correlation Task First, the absolute correlation between PAT values and the temperature feature vector is computed using the medical specialists ROIs distribution (Sec. 3.9.2.2), since this distribution matches exactly the PAT regions measured. The key parts of the correlation matrix results are presented in Figs. 45, 46, 47. Figure 45: Correlation matrix of PATs vs mean temperatures of each region for each foot using medical specialists ROIs distribution. Figure 46: Correlation matrix of PATs differences vs mean temperature differences between feet regions using medical specialists ROIs distribution. 56
Figure 47: Correlation matrix of PATs and PATs differences between regions of the same foot vs mean temperature differences between regions of the same foot using medical specialists ROIs distribution. As seen, if we take a look at all figures, all values are in blue indicating very small correlation indices, which means in the end, neither of the correlation matrices presents correlation between PAT values and temperatures. Even though the ROIs distribution do not match exactly with the PAT regions measured, let’s make the same analyis for the ROIs distributed by [17] (Sec. 3.9.2.1), so that we can also compare results between ROIs distributions. The results of this ROI distribution are presented in Figs. 48, 49, 50. Figure 48: Correlation matrix of PATs vs mean temperatures of each region for each foot using ROIs distribution by [17]. Figure 49: Correlation matrix of PATs differences vs mean temperature differences between feet regions using ROIs distribution by [17]. 57
References [1] Miguel ´ Angel D´ıaz Herrera and Carlos Mart´ınez Rico. Utilitzaci´o de termografia d`ermica infraroja rutin`aria en el diagn`ostic preco¸c de pd d’alt risc (neurop`atics, neuroisqu`emics i amb trastorns angiosomals subcl´ınics), 06 2022. [2] Contato de colaboraci´on para la utilizaci´on de termograf´ıa d´ermica infra-roja rutinaria en el diagn´ostico precoz de pie diab´etico de alto riesgo, 2022-12-01. [3] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation, 2015. [4] Kaiming He, Georgia Gkioxari, Piotr Doll´ar, and Ross Girshick. Mask r-cnn, 2018. [5] Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks, 2016. [6] Lukasz Bienias, Juanjo n, Line Nielsen, and Tommy Alstrøm. Insights into the behaviour of multi-task deep neural networks for medical image segmentation. pages 1–6, 10 2019. [7] Kasper Marstal. Simpleelastix. https://simpleelastix.readthedocs.io/index. html. Accessed: January 10, 2023. [8] Richard Szeliski. Feature Detection and Matching, pages 333–399. Springer International Publishing, Cham, 2022. [9] D. Ponsa E. Rublee E. Riba, D. Mishkin and G. Bradski. Kornia: an open source differentiable computer vision library for pytorch. In Winter Conference on Applications of Computer Vision, 2020. [10] David G. Lowe. Distinctive image features from scale-invariant keypoints. International Journal of Computer Vision, 60(2):91–110, 2004. [11] Marius Muja and David G. Lowe. Fast approximate nearest neighbors with automatic algorithm configuration. In Proceedings of the International Conference on Computer Vision Theory and Applications (VISAPP), pages 331–340. INSTICC, 2009. [12] R. I. Hartley and A. Zisserman. Multiple View Geometry in Computer Vision. Cambridge University Press, ISBN: 0521540518, second edition, 2004. [13] Martin A. Fischler and Robert C. Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM, 24(6):381–395, jun 1981. [14] Andr´e Rocha, S´ergio Ivan Lopes, and Carlos Abreu. A cost-effective infrared thermographic system for diabetic foot screening. In 2022 18th International Conference on Wireless and Mobile Computing, Networking and Communications (WiMob), pages 106–111, 2022. [15] H. Maldonado, R. Bayareh, I.A. Torres, A. Vera, J. Guti´errez, and L. Leija. Automatic detection of risk zones in diabetic foot soles by processing thermographic im64
ages taken in an uncontrolled environment. Infrared Physics Technology, 105:103187, 2020. [16] Chanjuan Liu, Jaap Van Netten, Jeff Baal, Sicco Bus, and F. Van der Heijden. Automatic detection of diabetic foot complications with infrared thermography by asymmetric analysis. Journal of biomedical optics, 20:26003, 02 2015. [17] D. Hernandez-Contreras, H. Peregrina-Barreto, J. Rangel-Magdaleno, J.A. GonzalezBernal, and L. Altamirano-Robles. A quantitative index for classification of plantar thermal changes in the diabetic foot. Infrared Physics Technology, 81:242–249, 2017. [18] Daniel Alejandro Hernandez-Contreras, Hayde Peregrina-Barreto, Jose de Jesus Rangel-Magdaleno, and Francisco Javier Renero-Carrillo. Plantar thermogram database for the study of diabetic foot complications. IEEE Access, 7:161296–161307, 2019. [19] Ikramullah Khosa, Awais Raza, Mohd Anjum, Waseem Ahmad, and Sana Shahab. Automatic diabetic foot ulcer recognition using multi-level thermographic image data. Diagnostics, 13(16), 2023. [20] Jill Sommerset, Riyad Karmy-Jones, Matthew Dally, Beejay Feliciano, Yolanda Vea, and Desarom Teso. Plantar acceleration time: A novel technique to evaluate arterial flow to the foot. Annals of Vascular Surgery, 60:308–314, 2019. [21] Nervengift. Flir image extractor python script. https://github.com/Nervengift/ read_thermal.py/tree/master. Accessed: January 10, 2024. [22] Phil Harvey. Exiftool. https://exiftool.org. Accessed: January 10, 2023. [23] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023. [24] Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Doll´ar. Microsoft coco: Common objects in context, 2015. [25] Pytorch. Torchvision object detection finetuning tutorial. https://pytorch.org/ tutorials/intermediate/torchvision_tutorial.html. Accessed: January 10, 2024. [26] J´ulia Sala Prat. Cell detection and classification of breast cancer histology images using a deep learning approach based on the U-Net architecture. PhD thesis, UPC, Facultat d’Inform`atica de Barcelona, Departament de Teoria del Senyal i Comunicacions, Jun 2021. [27] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. [28] Pavel Iakubovskii. Segmentation models pytorch. https://github.com/qubvel/ segmentation_models.pytorch, 2019. 65
[29] Sonia Rabanaque Rodr´ıguez. Breast tumor detection in Hematoxylin Eosin stained biopsy images. PhD thesis, UPC, Facultat d’Inform`atica de Barcelona, Departament de Teoria del Senyal i Comunicacions, Jun 2021. 66