Automated location of orofacial landmarks to characterize airway morphology in anaesthesia via deep convolutional neural networks
Abstract
The Spanish State Research Agency AEI under the project S3M1P4R (PID2020-115882RB-I00), the Basque Government EJ-GV under the grant ‘Artificial Intelligence in BCAM’ 2019/00432, under the strategy ‘Mathematical Modelling Applied to Health’, and under the BERC 2018–2021 and 2022–2025 programmes, and also by the Spanish Ministry of Science and Innovation: BCAM Severo Ochoa accreditation CEX2021-001142-S / MICIN / AEI / 10.13039/501100011033.
Full text
Automated location of orofacial landmarks to characterize airway morphology in anaesthesia via deep convolutional neural networks Fernando Garc´ıa-Garc´ıaa,∗, Dae-Jin Leea, Francisco J. Mendoza-Garc´esb, Sof´ıa Irigoyen-Mir´ob, Mar´ıa J. Legarreta-Olabarrietac, Susana Garc´ıa-Guti´errezc, Inmaculada Arosteguid,a aBasque Center for Applied Mathematics (BCAM). bGaldakao-Usansolo University Hospital, Anaesthesia & Resuscitation Service. cGaldakao-Usansolo University Hospital, Research Unit. dUniversity of the Basque Country (UPV/EHU), Department of Mathematics. Abstract Background: A reliable anticipation of a difficult airway may notably enhance safety during anaesthesia. In current practice, clinicians use bedside screenings by manual measurements of patients’ morphology. Objective: To develop and evaluate algorithms for the automated extraction of orofacial landmarks, which characterize airway morphology. Methods: We defined 27 frontal + 13 lateral landmarks. We collected n=317 pairs of pre-surgery photos from patients undergoing general anaesthesia (140 females, 177 males). As ground truth reference for supervised learning, landmarks were independently annotated by two anaesthesiologists. We trained two ad-hoc deep convolutional neural network architectures based on InceptionResNetV2 (IRNet) and MobileNetV2 (MNet), to predict simultaneously: a) whether each landmark is visible or not (occluded, out of frame), b) its 2D-coordinates (x, y). We implemented successive stages of transfer learning, combined with data augmentation. We added custom top layers on top of these networks, whose weights were fully tuned for our application. Performance in landmark extraction was evaluated by 10-fold cross-validation (CV) and compared against 5 state-of-the-art deformable models. Results: With annotators’ consensus as the ‘gold standard’, our IRNet-based network performed comparably to humans in the frontal view: median CV loss L= 1.277 ·10−3, inter-quartile range (IQR) [1.001, 1.660]; versus median 1.360, ∗Corresponding author. Contact info: Alameda de Mazarredo, 14–48009 Bilbao, Bizkaia (Basque Country, Spain). Telephone: +34 946 567 842. Email addresses: [email protected] (Fernando Garc´ıa-Garc´ıa), [email protected] (Dae-Jin Lee), [email protected] (Francisco J. Mendoza-Garc´es), [email protected] (Sof´ıa Irigoyen-Mir´o), [email protected] (Mar´ıa J. Legarreta-Olabarrieta), [email protected] (Susana Garc´ıa-Guti´errez), [email protected] (Inmaculada Arostegui) This is the accepted manuscript of the article that appeared in final form in Computer Methods and Programs in Biomedicine 232 : (2023) // Article ID 107428, which has been published in final form at https://doi.org/10.1016/j.cmpb.2023.107428. © 2023 Elsevier under CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/)
IQR [1.172, 1.651], and median 1.352, IQR [1.172, 1.619], for each annotator against consensus, respectively. MNet yielded slightly worse results: median 1.471, IQR [1.139, 1.982]. In the lateral view, both networks attained performances statistically poorer than humans: median CV loss L= 2.141 ·10−3, IQR [1.676, 2.915], and median 2.611, IQR [1.898, 3.535], respectively; versus median 1.507, IQR [1.188, 1.988], and median 1.442, IQR [1.147, 2.010] for both annotators. However, standardized effect sizes in CV loss were small: 0.0322 and 0.0235 (non-significant) for IRNet, 0.1431 and 0.1518 (p <0.05) for MNet; therefore quantitatively similar to humans. The best performing state-of-the-art model (a deformable regularized Supervised Descent Method, SDM) behaved comparably to our DCNNs in the frontal scenario, but notoriously worse in the lateral view. Conclusions: We successfully trained two DCNN models for the recognition of 27+13 orofacial landmarks pertaining to the airway. Using transfer learning and data augmentation, they were able to generalize without overfitting, reaching expert-like performances in CV. Our IRNet-based methodology achieved a satisfactory identification and location of landmarks: particularly in the frontal view, at the level of anaesthesiologists. In the lateral view, its performance decayed, although with a non-significant effect size. Independent authors had also reported lower lateral performances; as certain landmarks may not be clear salient points, even for a trained human eye. Keywords: Difficult airway, Anaesthesia, Deep learning, Transfer learning, Facial landmarks. 1. Introduction 1.1. Medical motivation Clinical guidelines provide anaesthesiologists with valuable, evidence-based counselling for the management of difficult respiratory airways [1]. Complications –despite infrequent– are an important source of lesions: inadequate5 oxygenation might cause severe brain damage or even death [2]. In surgical planning, when the airway is anticipated as difficult, specific safety equipment and personnel are required. However, the prognosis of difficult airways remains a challenging task, even for experienced anaesthesiologists: up to 10% of tracheal intubations become10 difficult [3], approximately 93% of which were unanticipated [4]. In clinical practice, bedside screenings assess airway difficulty. These examine the oropharyngeal morphology: e.g. thyro-/sterno-mental distance (TMD, SMD), mouth opening-interincisor gap (MO-IIG), or the modified Mallampati test (MMT) [5]. However, such screenings may be time-consuming and prone15 to repeatability/reproducibility issues (e.g. due to manual measurements), and systematic reviews report limited discernment capabilities [6, 7]. 2
In this context, anaesthesiologists are gaining growing interest in artificial intelligence (AI) and machine learning (ML) [8, 9, 10]. 1.2. Related works20 In literature, one could distinguish three main families of AI-ML methods conceived to assist clinicians in predicting difficult intubation of the airway. The first type of approaches consists of feeding ML algorithms with structured information about the patient’s demographics (most often: sex, age and body mass index), along with one or more morphological characterizations of the airway25 (e.g. TMD, SMD, MO-IIG or neck circumference) – these latter magnitudes having been measured manually, as in the bedside screenings from current clinical practise. Works in this category include Yan et al. [11] (who trained a Support Vector Machine), Langeron et al. [12] (who employed Tree Boosting), Kim et al. [13] (5 ML algorithms, among which Random Forests performed best), Ya-30 manaka et al. [14] (with an ensemble of 7 ML algorithms), or Zhou et al. [15] (10 algorithms, with Gradient Boosting as their best). These strategies may help to overcome the limitations of univariate analyses (like the bedside screenings) and of clinical scores and logistic regressions derived from a few predictor variables [16, 17, 18]. Nonetheless, they may still suffer from35 reproducibility issues, since they rely on input information about the airway morphology which needs to be measured by hand. A second family of approaches comprises AI-ML computer vision tools which analyse preoperative patient images. Tavolara et al. [19] developed 11 Convolutional Neural Networks (CNN), each one specialized in a different facial sub-40 region. With these CNNs, the authors generated feature ensembles which were subsequently fed to a Multiple Instance Learning model. In addition, Hayasaka et al. [20] proposed a methodology based on photos of 16 different poses per patient (combinations of frontal/lateral views, supine/sitting position, open/closed mouth, head bent backwards or not), training respective CNNs to classify diffi-45 cult airways. Furthermore, Cho et al. [21] used CNNs with lateral cervical spine X-ray images. Compared to the first type of proposals, these image-based methodologies may be able to exploit richer visual information, impossible to capture alone by the traditional morphological measures – perhaps at the cost of not having50 access to the patient’s demographic data. However, these algorithms tend to behave as opaque systems lacking human interpretability, which may hinder their adoption by some clinical practitioners. The third type of method can be thought of as a two-stage process: first, preoperative images are analyzed to identify and locate relevant landmarks55 (i.e. salient key points in subjects’ anatomy); second, morphological measurements are derived from them and fed to ML classifiers, trained to predict intubation difficulty. To our knowledge, Suzuki et al. [22] were the first who proposed an algorithm with such a strategy: the authors defined 84 landmarks from a frontal view, alongside 39 lateral landmarks. Based on n=32 patients60 3
(16 females) with easy intubation, plus n=41 (18 females) with difficult intubation, they used ‘morphing’ techniques for landmark location. Afterwards, five morphological magnitudes were extracted. Conversely, Connor & Seagal [23] opted for using three photographs: one frontal view, plus left and right profiles. The authors enrolled n=80 subjects (all65 males), using semi-automatic proprietary software for facial structure analysis by ‘eigenface’ projection. Based on the landmarks, 61 facial proportions were computed and fed to logistic regressions. Cuendet et al. [24] built a custom ‘photo booth’-like equipment, with two high-definition webcams placed orthogonally, 30–40 cm distant from the pa-70 tient. The authors captured four types of photographs: three frontal views (99 landmarks each) and a lateral profile (52 landmarks). Their study enrolled n=970 patients (482 females) but comprised 406 sets of photos manually annotated for frontal detection, alongside 134 for lateral detection. The authors addressed the task of landmark location by means of a deformable appearance75 model, based on scale-invariant feature transform (SIFT) features computed from patches around each landmark. Their morphological measures, alongside various texture features, fed a Random Forest classifier. 1.3. Rationale & Objective In this work, we aim at contributing to the characterization of airway mor-80 phology, in the context of predicting difficult intubations for anaesthesia. Among the three types of approaches outlined in Section 1.2, we opted for the latter two-stage strategy. In brief: automated landmark location to obtain morphological measures, followed by ML-enabled prediction of intubation difficulty (the latter stage falling beyond the scope of this paper – see Section 5.3). In our85 view, such a two-stage methodology attains a desirable balance between interpretability for human experts and the automatic, repeatable and reproducible extraction of relevant information from preoperative examinations via imaging techniques. Methodologically, with respect to the related works found in literature, here90 we propose two models which exploit the recent algorithmic advances in the field of computer vision: where deep convolutional neural networks (DCNNs) have become dominant, thanks to their large expressive power and high levels of performance achieved across many different tasks. To some extent, our work was also inspired by Cuendet et al. [24]. However,95 here we made various different design choices: for example, their four-image acquisition protocol using a ‘photo booth’-like equipment with two high-definition cameras may not be feasible to integrate into pre-surgical examination workflows in practice. Instead, we opted for a single frontal photograph captured with a smartphone, with patients’ mouth fully open and tongue out –hence allowing100 for the assessment of common bedside screening magnitudes (MO-IIG, MMT)–; as well as for a single lateral view, with head in vertical extension – for TMD and SMD. Our anaesthesiology team defined two custom sets of landmarks (one frontal, one lateral), carefully tailored to describe the orofacial anatomy pertaining to the respiratory airway. Furthermore, we obtained two independent105 4
sets of high-quality annotations by independent anaesthesiologists, from which a consensus ground truth was derived. These two sources of annotation allowed us to study inter-human discrepancies: as a ‘gold standard’ for the performance of our algorithms, and as a quantitative measure of the intrinsic complexity of this landmark identification and location problem for the trained human eye.110 2. Background 2.1. ‘General-purpose’ facial landmarks In this section, we briefly overview the state-of-the-art in facial landmark recognition. For a comprehensive literature review, the interested reader may consider Wu & Ji [25]. Note that, although the models for ‘general-purpose’ fa-115 cial landmarks cannot be applied directly to our anaesthesia scenario (given the substantial visual differences in head poses, photo framing, anatomical content, etc. and even in the definition of landmarks themselves – see Section 3), various methodological principles can indeed be shared. The problem of identifying and locating landmarks in human face images120 has triggered major research efforts in the computer vision community, as it can be applied to various contexts: subject identification [26, 27], recognition of facial expressions [28, 29], sentiment analysis [30], detection of genetic syndromes [31], etc. Active appearance models (AAM) [32] and constrained local models (CLM) [33] were among the first methods to achieve satisfactory gener-125 alization capabilities, even with few training images [34]. With the availability of larger facial datasets and improved descriptive features (such as histogram of oriented gradients – HOG, or scale-invariant feature transforms – SIFT) [35], discriminative models like cascade regression methods (in particular, supervised descent methods – SDM) [36] became the technique of reference.130 Currently, Deep Learning (DL) has achieved major relevance and remarkable performance. This trend of ground-breaking algorithmic and technical advances by DL is common to many fields in computer vision, also to the subdomain of facial landmarks [25]. In particular, various DCNN architectures have been applied successfully for facial landmark localization [37, 38, 39], motivating us to135 explore this DCNN approach in the context of airway assessment and anaesthesia. 2.2. Deep Learning & Transfer Learning Although DL has consistently outperformed other computer vision techniques in a wide variety of applications, including biomedical imaging [40, 41];140 its ‘data hunger’ phenomenon [42] poses an important challenge in scenarios where the availability of data (and/or of ground truth) is scarce or costly to extract [43] – as in our case here. Transfer Learning (TL) may alleviate such ‘data hunger’ issue: a DL network’s millions of neuron weights may not need to be fitted from scratch, but instead reuse knowledge obtained from other tasks [44, 45].145 In this sense, empirical studies have shown that knowledge transferability in TL –typically via weight sharing– is more favourable the closer the original and 5
target domains are, but low-level generic features learnt from distant tasks can still be better than random weight initialization, and a boost for the network’s generalization properties [46].150 Here we explore this TL strategy on the basis of two public large-scale datasets for ‘general-purpose’ facial landmark recognition. 3. Materials and methods In this work, we propose two ad-hoc DCNNs to identify a set of landmarks, which our anaesthesiologists established as relevant to characterize patients’155 orofacial morphology when assessing difficult airways for tracheal intubation. For each landmark, the DCNNs will predict: a) its visibility v(occlusions may occur); as well as b) its 2D coordinates (x, y) in the image. 3.1. Data 3.1.1. Transfer Learning datasets160 Our starting point consisted of DCNNs pre-trained on ImageNet: a largescale collection with >3.2 million images [47, 48]. This dataset is an essential reference for computer vision applications, and almost all DL programming frameworks include networks pre-fitted on it: hence, already suitable for TL. In addition, we used two other public datasets for ‘general-purpose’ facial165 landmark recognition: ‘CelebFaces Attributes Dataset’ (CelebA) and ‘Annotated Facial Landmarks in the Wild’ (AFLW ). To the best of our knowledge, we are unaware of the existence of public datasets pertaining to the assessment of respiratory airways for anaesthesia, which would indeed be a closer domain to our scenario. Nevertheless, based on the principles introduced in Section 2.2 [44,170 45, 46], we deemed the visual contents of these ‘general-purpose’ facial datasets to be similar enough to be promising sources for knowledge extraction via TL. CelebA [49] is a non-commercial, large-scale dataset containing 202,599 human face images (58.3% females) from 10,177 distinct subjects. Faces were pre-aligned and cropped to constant size (rectangles with 218×178 pixels). For175 each face, the ground truth 2D coordinates of 5 facial landmarks (eye centres, nose tip, and mouth corners) are supplied. The dataset encompasses a notable variety in head pose and photo perspectives, as well as other attributes (age, sex, ethnicity, glasses, etc.). AFLW [50] is a non-commercial, research-only dataset which comprises180 24,384 faces (58.8% females) from 21,121 distinct images. The authors provide each face’s square bounding box, alongside the visibility and 2D coordinates of 21 landmarks (brows, eyes, ears, nose, mouth, chin). AFLW is reported to contain substantial variety in terms of sex and ethnicity. Most notably, approx. 66% of images correspond to non-frontal poses: to the authors’ knowledge, a higher185 ratio than in any other public facial landmarks dataset. 6
3.1.2. DiffAirW dataset We conducted an observational, prospective, cohort study at the University Hospitals of Galdakao-Usansolo and Basurto (Biscay, Basque Country, Spain), approved by the ‘Ethics Committee for Research with medication’ of the Basque190 Country (CEIm-E). We enrolled participants who gave informed written consent in accordance with the Declaration of Helsinki. Inclusion criteria were: adulthood (age≥18) and undergoing tracheal intubation via direct laryngoscopy for general anaesthesia, regardless of whether the intervention was programmed or emergency. Exclusion criteria were: otolaryngeal diseases with affected struc-195 tures, radiotherapy, any previous surgery or anatomical alterations in the airway, obstetric patients and/or a known difficult airway requiring fiberoptic intubation. From March 2018 to January 2020, n=317 pairs of preoperative photos were collected: one frontal, and one lateral. In the frontal view, patients were in-200 structed to fully open their mouths, and to stick their tongues out; hence facilitating a simultaneous judgment of MMT and MO-IIG. In the lateral view, the pose was with the head in vertical extension: the standard procedure to measure TMD and SMD. Smartphones with general-purpose cameras were used to take the photographs, and a cue card was added to the scene (circle with 25 mm205 diameter), as a reference for physical dimensions. We recorded each patient’s demographics (sex, age, weight, height, body mass index), MMT and operative outcomes, including: the number of intubation attempts and/or extra devices, duration of the manoeuvre, Cormack-Lehane score for direct laryngoscopy [51], etc. With these, we determined objectively210 the intubation difficulty, according to the IDS [52] and SFAR criteria [53] (Table 1). In our cohort, women were under-represented (significantly although by a narrow margin, p=0.0430 – Table 1): 177 males versus 140 females; a proportion of 44.16%, with a 95% confidence interval [38.63%–49.82%]. Our anaesthesiologists defined two sets of landmarks relevant to orofacial215 morphology (Figure 1): N=27 points for the frontal view, alongside N=13 for the lateral view. Two anaesthesiologists annotated these landmarks manually (Section 3.2.1). 3.2. Landmark location Let (vi, xi, yi) be the triplet which characterizes ground truth annotations220 for the i-th out of Nlandmarks; where N=5 for CelebA,N=21 for AFLW, whereas N=27 for DiffAirW[Frontal], and N=13 for DiffAirW[Lateral]. Here virepresents a landmark’s visibility (it could become occluded by skin tissue, beard, clothes, etc., or even left out of the photo frame); whereas (xi, yi) denote its 2D position: i.e. horizontal and vertical coordinates, being (0,0) the top-left225 of the image and (1,1) its bottom-right corner, by convention. Therefore, the task to address here can be described as a combination of multiple binary classifications (Nvisibilities vi) and 2Nregressions for coordinates (xi, yi). To compare triplets {(ˆvi,ˆxi,ˆyi)}N i=1 predicted by the DCNNs 7
Table 1: Characteristics of our DiffAirW cohort All By sex Females Males Binomial n=317 n=140 n=177 0.0430 Mann-Whitney Age 63.0 (51.8, 72.0) 58.0 (46.0, 69.3) 65.0 (54.0, 73.0) <0.001 Weight [kg] 74.0 (64.8, 88.0) 68.0 (58.5, 81.3) 78.0 (70.0, 91.3) <0.001 Height [m] 1.66 (1.60, 1.73) 1.60 (1.56, 1.65) 1.71 (1.66, 1.76) <0.001 BMI [kg m−2] 26.71 (23.95, 30.65) 26.13 (23.02, 31.23) 26.94 (24.22, 30.37) 0.4463 Mallampati χ2 Grade I 152 (47.95%) 72 (51.43%) 80 (45.20%) 0.5842 Grade II 108 (34.07%) 47 (33.57%) 61 (34.46%) Grade III 48 (15.14%) 18 (12.86%) 30 (16.95%) Grade IV 9 (2.84%) 3 (2.14%) 6 (3.39%) Cormack-Lehane χ2 Classes I–II 300 (94.64%) 135 (96.43%) 165 (93.22%) 0.2080 Classes III–IV 17 (5.36%) 5 (3.57%) 12 (6.78%) IDS difficulty χ2 No 301 (94.95%) 136 (97.14%) 165 (93.22%) 0.1132 Yes 16 (5.05%) 4 (2.86%) 12 (6.78%) SFAR difficulty χ2 No 302 (95.27%) 137 (97.86%) 165 (93.22%) 0.0535 Yes 15 (4.73%) 3 (2.14%) 12 (6.78%) Values are shown as median (inter-quartile range), or number (percentage) where appropriate. The right-most column contains p-values of population differences with respect to sex, and the statistical test used. BMI: body-mass index, IDS: Intubation Difficulty Scale SFAR: Soci´et´e Fran¸caise d’Anesth´esie et de R´eanimation [French Society of Anaesthesia and Resuscitation]. 8
Figure 1: Definition of orofacial landmarks to characterize airway morphology in the context of anaesthesia. Left panel – Frontal view: Landmarks F01–F27 (nose, lips, teeth, tongue, chin, mandible, neck, thyroid cartilage, sternal manubrium). Right panel – Lateral view: Landmarks L01–L13 (nose, lips, chin, mandible, thyroid cartilage, sternal manubrium, nape, occiput). against ground truth {(vi, xi, yi)}N i=1, we defined the following loss function: L(·|γ) = γ1 N N X i=1 BCE(vi,ˆvi)+ 1 PN i=1 vi N X i=1 vi(xi−ˆxi)2+AR2(yi−ˆyi)2(1) being BCE the binary cross-entropy loss: BCE(v, ˆv) = −vlog(ˆv)−(1 −v) log(1 −ˆv) (2) and AR the aspect ratio (height/width) of the original image. The first summation term in Eq. [1] penalizes the disagreement between predicted ˆviand ground truth vivisibilities; whereas the second summation term corresponds to the mean squared point-to-point error in coordinates. When-230 ever a certain landmark was set as non-visible in the ground truth annotations (vi=0), the multiplicative term by viin the second summation of Eq. [1] nullified any contribution to loss with regard to the true/estimated locations (xi, yi) or (ˆxi,ˆyi). Hence, with vi=0, the estimated (ˆxi,ˆyi) become irrelevant for the total loss and do not contribute either to computing the gradient of L(·) in the235 training of the network by backpropagation. Besides, the weight parameter γ balances the contribution of each term towards the overall loss L(·|γ). Eq. [1] is similar to Ranjan et al.’s proposal [38]. However, in our second term, we average only by the number of visible landmarks PN i=1 vi, 9
respective HOG-SVM detectors [66, 67] for our frontal and lateral orofacial structures in DiffAirW. In turn, these detectors were used with the images in the test stage, to find suitable bounding boxes for them, from which to start fitting the trained deformable models.385 Likewise in Section 3.3.1, we evaluated these models’ performance via 10fold CV, again stratified by sex. In the case of SDM, we conducted preliminary CV hyperparameter tuning experiments to determine regularization strength λ, with values from 10−3to 109in factors of 10. Since such a tuning procedure repeatedly selected λ=106as optimal, we decided to fix λalways to that value.390 Figure 5: Schematic flowchart of the various stages for training the deformable models. 4. Results 4.1. Transfer Learning from the ‘general-purpose’ facial datasets Starting with pre-fitted weights from ImageNet, we applied successive TL stages in order of increasing visual similarity with our target DiffAirW domain [46] (also in decreasing dataset size): first on CelebA, then on AFLW395 (Figure 4). As depicted in Figure 6, our ad-hoc architectures with the InceptionResNetV2 core (denoted onwards IRNet AirW – upper row) and MobileNetV2 core (MNet AirW – lower row), both learnt satisfactorily: converging successfully within the established training epochs and achieving noticeable decreases400 in loss L, both for CelebA data (left column in Figure 6) and for AFLW (right column). Losses from the test sets (plotted in orange in Figure 6) were repeatedly lower than from the training sets (blue), due to the absence of augmentation transformations during the testing stage. There was a consistent trend by IRNet AirW to yield lower test losses than MNet AirW for the same dataset and405 the same training epoch. On the other hand, the large size –hence, increased expressive power– of IRNet AirW may explain the sudden peaks observed in training loss L, from which it was nonetheless able to recover within a single epoch. Figure 7 shows the two-dimensional UMAP projection/embedding [68] of410 the neuron activations (averaged per convolutional channel), at different network depths: raw input pixel values (first row), as well as the activation output by the section of the networks frozen at each TL stage (50% to 100% of the ‘standard’ DCNN architectures, Figure 4). In these plots, both IRNet AirW (left column) and MNet AirW (right) showed the empirically expected behaviour [46], by415 16
Figure 6: Transfer learning (TL), intermediate results – Convergence of the training and test losses L(·|γopt) across epochs, plotted in logarithmic scale [ordinates]. After obtaining the already pre-trained weights from ImageNet, the first subsequent TL stage was carried out on the CelebA dataset (left column), and then on AFLW (right column). Upper row: IRNet AirW architecture, lower row: MNet AirW. which the bottom layers extract generic features: raw pixel embeddings were not separable (panels a–b), but their computed features were (c–d onwards); so the neurons in the upper layers could focus on specialization. 4.2. Performance for landmark location in DiffAirW Table 2 and Figure 8 summarize the losses L(·|γopt) incurred by each human420 against consensus [‘gold standard’], as well as by the DCNNs – computed from triplets (ˆvi,ˆxi,ˆyi) output in the validation split of our 10-fold CV. Figure 9 displays the DCNN results for eight patients (4 frontal, 4 lateral; 4 by IRNet AirW, 4 by MNet AirW ). For the sake of fair reporting and illustration of the networks’ performance, these examples were not selected among the425 most favourable cases (i.e. lowest losses), but instead around the overall median CV loss L(·|γopt). Qualitatively, the overall results were satisfactory for our anaesthesiology team through visual exploration. Nonetheless, in the frontal view, the sternal manubrium (F25–F27) was often the most difficult anatomical structure to capture –as it was also for the human annotators–. In this re-430 gard, the lateral view should be better suited to identify the sternal manubrium (landmark L11): the lateral and frontal photos complement each other for an enhanced description of patients’ anatomy. The base of the neck (F23–F24) 17
(a) Raw RGB pixels: 299×299×3, as in the InceptionResNetV2 input format. (b) Raw RGB pixels: 224×224×3, as in the MobileNetV2 input format. (c) IRNet AirW : Uppermost frozen layer, after TL from ImageNet. (d) MNet AirW : Uppermost frozen layer, after TL from ImageNet. (e) IRNet AirW : Uppermost frozen layer after TL from CelebA. (f) MNet AirW : Uppermost frozen layer after TL from CelebA. (g) IRNet AirW : Uppermost frozen layer after TL from AFLW. (h) MNet AirW : Uppermost frozen layer after TL from AFLW. Figure 7: Two-dimensional UMAP projection of the average neuron activations, at different network depths and TL stages, for both DCNNs: IRNet AirW (left column) and MNet AirW (right column). In purple, 2,000 randomly sampled images from ImageNet (two from each of the C=1000 existing classes); in orange, 2,000 randomly sampled images from CelebA’s test set (size 20,000); in blue, 2,000 randomly sampled from AFLW ’s (size 2,500); in red and green, each of the n=317 images from our DiffAirW dataset (frontal and lateral views, respectively). 18
was sometimes also challenging, for overweight patients with a thick neck or for those situations in which, given the relative position of the camera with respect435 to the patient (distance, inclination), the perspective view of the neck could become obstructed by the mandible. In the lateral view, the nape, occiput and mandible corner and thyroid cartilage (L12–L13, L08–L09, L10) were the landmarks with inter-annotation discrepancies above the average. However, some drift in the area around the mouth440 and the chin could also be encountered sometimes, particularly when the image background had varied visual content, instead of being flat. IRNet AirW significantly outperformed MNet AirW in the frontal and lateral views (Table 2, Figure 8), both for the entire cohort and disaggregating by sex. Only MNet AirW[Frontal] exhibited statistically different perfor-445 mances with respect to sex: performing better for women –despite their certain underrepresentation– than for men (for further details, see Section 4.3). Figure 10 depicts four direct comparisons (2 women, 2 men) between IRNet AirW and MNet AirW, with the networks being applied to exactly the same input images. Qualitatively, both DCNNs attained similar localizations;450 quantitatively, IRNet AirW outperformed MNet AirW, again as shown by Table 2 and Figure 8. In addition, here we can observe very similar behaviours with the specific landmarks to those reported in Figure 9. Figure 8: Cumulative Error Distributions (CED) for the CV losses L(·|γopt) with respect to consensus in DiffAirW, in landmark identification and location, as defined in Eq. [1]: expert human annotators and our DCNNs. Left panel, frontal view scenario; right panel, lateral view. [Abscissae in logarithmic scale]. Furthermore, we disaggregated CV losses by landmark and by contribution: i.e. whether from discrepancy vi∼ˆvi(BCE terms in Eq. [1]), or from squared455 errors in (xi, yi)∼(ˆxi,ˆyi) coordinate locations, whenever positive ground truth landmark visibility (vi=1). In the case of occlusions (vi=0), such contribution to the overall loss is null (Eq. [1]). 19
Table 2: Losses L(·|γopt) in landmark identification and location for our DiffAirW dataset. In the upper part, losses incurred by each human annotator, when compared against consensus. On the lower part, CV losses by each DCNN architecture. Loss L(·|γopt) L(·|γopt) L(·|γopt) [10−3]Frontal view Lateral view Humans Annot#1 vs. Consensus LA1|C All 1.360 (1.172, 1.651) 1.507 (1.188, 1.988) Females 1.314 (1.154, 1.528) p=0.0049 1.603 (1.270, 2.011) p=0.0364 Males 1.424 (1.201, 1.732) 1.467 (1.117, 1.948) Annot#2 vs. Consensus LA2|C All 1.352 (1.172, 1.619) 1.442 (1.147, 2.010) Females 1.296 (1.140, 1.504) p=0.0017 1.516 (1.205, 2.011) p=0.1906 Males 1.412 (1.227, 1.686) 1.419 (1.110, 2.008) Inter-annotator difference (p-values) All 0.0135 <0.001 Females 0.0061 <0.001 Males 0.4095 0.4458 DCNNs IRNet AirW vs. Consensus LIRNet|C All 1.277 (1.001, 1.660) 2.141 (1.676, 2.915) Females 1.300 (0.986, 1.586) p=0.7043 2.080 (1.754, 2.887) p=0.9297 Males 1.266 (1.035, 1.696) 2.158 (1.651, 2.953) MNet AirW vs. Consensus LMNet|C All 1.471 (1.139, 1.982) 2.611 (1.898, 3.535) Females 1.391 (1.104, 1.803) p=0.0229 2.752 (1.925, 3.661) p=0.3713 Males 1.568 (1.210, 2.094) 2.561 (1.877, 3.455) Inter-network difference (p-values) All <0.001 <0.001 Females <0.001 <0.001 Males <0.001 <0.001 Values are shown as median (IQR: inter-quartile range). Population differences with respect to sex were analysed via Mann-Whitney two-sided U tests. Inter-annotator Inter-network differences (i.e. bottom rows) were analysed via paired Wilcoxon signed-rank tests. 20
Figure 9: Results for eight example patients in DiffAirW : four frontal (left), four lateral (right). DCNN landmark outputs are depicted as overlay on their corresponding input image. These 8 individual cases cover intermediate performances, with CV losses approximately around the overall median CV loss L(·|γopt). 21
Figure 10: Results for four complete patient cases in DiffAirW : frontal view (upper row), and lateral view (lower row). Direct comparison between the outputs by IRNet AirW (blue) and by MNet AirW (red), against consensus as ground truth reference (green). These examples illustrate graphically the superior average performance by IRNet AirW over MNet AirW. 22
Figure 11: CV losses for DiffAirW, disaggregated by landmark and by contribution term. Frontal (up) and lateral (down) scenarios. Errors due to discrepancies in coordinate estimations (‘coords’ in the legend) were computed exclusively for those cases with positive ground truth landmark visibility vi=1, as otherwise (vi=0) such contribution term becomes irrelevant to L. The red dash-dotted line marks the corresponding overall median loss L(·|γopt). For visual clarity in the diagrams, outliers were not plotted here. 23
Losses related to ˆviwere negligible, except for specific landmarks (Figure 11). F01 (nose tip) was often left out of the photo frame, and the networks tended to460 misjudge its visibility. Nevertheless, the loss term for F01 location remained in ranges similar to other landmarks. The ˆvirecognition of F22–F27 also exhibited a certain degree of inaccuracy. This may be explained, at least in part, due to occlusions and difficult visual distinguishability even for humans (fat, clothes, etc.). For instance, Figure 2 illustrates some common cases with non-negligible465 disagreement between annotators. L01, L02 and L13 also suffered from outof-frame issues. Distinguishing L10 (thyroid cartilage) was a challenge for the human annotators, and also for the DCNNs. 4.3. Comparisons versus human performance in DiffAirW We compared the performances achieved by our DCNNs (in terms of CV470 losses Lagainst consensus), with respect to the losses incurred by the human annotators. We carried out a repeated-measures, two-way ANOVA analysis with factors: annotator (four levels – two humans, two DCNNs) and sex. These analyses were performed with the statistical software R, and libraries lme4 [69], lmerTest [70].475 Table 3 and Figure 12 show that the main effect of annotator is always significant (p < 10−12), pointing out differences across at least some of the four. Sex is significant in the frontal view (p=0.0169), but not in the lateral (p=0.8611). Such significance of factor sex is attributable partly to MNet AirW, but importantly, also to inter-human differences (see Figure 12a): annotators incurred480 larger losses with men – rather than with women, despite being underrepresented. In both cases, the interaction sex:annot is not significant. Frontal view SS MS DoFnum DoFdenom F-statistic p-value sex 1.03 1.03 1 315 5.77 0.0169 annot 10.41 3.47 3 945 19.50 <1e-12 sex:annot 1.08 0.36 3 945 2.02 0.1092 Lateral view SS MS DoFnum DoFdenom F-statistic p-value sex 0.04 0.04 1 315 0.03 0.8611 annot 389.68 129.89 3 945 107.76 <1e-12 sex:annot 2.63 0.88 3 945 0.73 0.5350 Table 3: ANOVA summary table for CV losses for our DiffAirW dataset, in the frontal and lateral views – SS: sum of squares; MS: mean square; DoFnum: degrees of freedom in the numerator; DoFdenom: degrees of freedom in the denominator. 4.4. Post-hoc analyses and effect sizes in DiffAirW The ANOVA F-tests in Table 3 revealed significant overall differences across annotators in general. To find differences between specific annotators, we per-485 formed post-hoc tests to assess the significance of differences between pairs of group means using R’s emmeans library [71] (Table 4). 24
Figure 12: Comparison of CV losses L[10−3] in our DiffAirW dataset: Mean and 95% confidence interval. Frontal (a) and lateral (b) views. IRNet AirW[Frontal] behaved comparably to both humans (Figure 12a), with effect sizes non-significantly different from zero (see annot1 - incept,annot2 - incept rows in Table 4); whereas MNet AirW[Frontal] exhibited lower490 performance, although with small standardized effect sizes: 0.2245, 0.2382. IRNet AirW[Lateral] performed on average worse than humans (Figure 12b), yet better than MNet AirW[Lateral]. Besides, IRNet AirW[Lateral]’s effect sizes were non-significantly different from zero (Table 4); whereas MNet AirW[Lateral]’s effect sizes were small: 0.1431, 0.1518.495 4.5. Comparison with state-of-the-art methods Attending to the Cumulative Error Distributions (CED) of losses for the five proposed Menpo deformable models (Figure 13), the one achieving the best overall performance was the regularized SDM (λ=106). Note that, since the deformable models cannot estimate visibilities ˆvi, all the losses here correspond500 only to the difference between ground truth (xi, yi) and estimated coordinates (ˆxi,ˆyi) for the visible landmarks (vi=1), i.e. the second summation term in Eq. [1]. In the frontal view scenario, SDM performed comparably to our DCNNs for up to 80–90% of the images. However, its behaviour degraded notoriously for the lateral view: its median CV loss was approximately 10 times higher than505 for our best DCNN. Figure 14 depicts two favourable and two intermediate SDM landmark detection results. The first two patient cases (panels on the left-hand side) correspond to satisfactory outputs –losses in the inferior 10% percentile–; whereas the two latter cases (right panels) correspond to intermediate performances –around the510 median loss–. In the lateral view, one case depicts a large loss owing to a major error in a single landmark (sternal manubrium, L11); whereas the other case illustrates a noticeable overall drift in location. Note that this type of drift is the predominant behaviour in the cases with the most severe losses, where the SDM model fails to attain satisfactory detection, converging instead to a noisy515 solution. 25
The DCNNs were evaluated via 10-fold CV stratified by sex, to guarantee675 a fair and robust assessment of performance in images unseen during training. Results were compared to inter-annotator discrepancies as a ‘gold standard’ and to 5 state-of-the-art deformable models. Overall, IRNet AirW[Frontal] yielded satisfactory generalization capabilities (achieving performances at the level of human experts), whereas MNet AirW[Frontal] experienced only slight degrada-680 tion. In the lateral view, both DCNNs performed statistically worse than humans, although with small effect sizes (non-significant for IRNet AirW[Lateral]). This issue of lower lateral performance had already been described and discussed in literature by independent authors, who used different techniques. Arguably, it may be related to intrinsic visual distinguishability challenges, even for trained685 human eyes. For our best model, no bias in performance was found regarding sex. Acknowledgements We would like to acknowledge the patients who participated in this research, the staff at the University Hospitals of Galdakao-Usansolo and Basurto from the690 public Basque healthcare system (Osakidetza), as well as the ‘Ethics Committee for Research with medication’ of the Basque Country (CEIm-E). The authors thank also Dr. Amani Tahat for her work deploying our DCNNs in smartphone devices. CRediT authorship statement695 Fernando Garc´ıa-Garc´ıa: Conceptualization, methodology, software, investigation, validation, formal analysis, visualization, writing - original draft. Dae-Jin Lee: Formal analysis, visualization, supervision, writing - review & editing, project administration, funding acquisition. Francisco J. Mendoza-700 Garc´es: Conceptualization, methodology, resources, data curation, writing - review & editing, funding acquisition. Sof´ıa Irigoyen-Mir´o: Resources, data curation. Mar´ıa J. Legarreta: Resources, data curation. Susana Garc´ıaGuti´errez: Supervision, project administration, funding acquisition. Inmaculada Arostegui: Writing - review & editing, supervision, project administra-705 tion, funding acquisition. Funding This research is supported by the Spanish State Research Agency AEI under the project S3M1P4R (PID2020-115882RB-I00), as well as by the Basque Government EJ-GV under the grant ‘Artificial Intelligence in BCAM’ 2019/00432,710 under the strategy ‘Mathematical Modelling Applied to Health’, and under the BERC 2018–2021 and 2022–2025 programmes, and also by the Spanish Ministry of Science and Innovation: BCAM Severo Ochoa accreditation CEX2021001142-S / MICIN / AEI / 10.13039/501100011033. 32
The funding sources had no role in this work: neither in the design and715 conduct of the study, collection, management, analysis and interpretation of the data, nor in the preparation, review, approval of the manuscript, nor in the decision to submit the manuscript for publication. Conflict of interest statement The authors declare that they have no known competing financial inter-720 ests or personal relationships that could have appeared to influence the work reported in this paper. To maximize the impact of this study, the Basque Center for Applied Mathematics (BCAM) submitted a provisional application for intellectual property to the Territorial Office of Intellectual Property of the Basque Autonomous725 Community, Spain. 33
References [1] J. L. Apfelbaum, C. A. Hagberg, R. A. Caplan, C. D. Blitt, R. T. Connis, D. G. Nickinovich, C. A. Hagberg, et al., Practice guidelines for management of the difficult airway: An updated report by the American Society of730 Anesthesiologists Task Force on Management of the Difficult Airway, Anesthesiology 118 (2) (2013) 251–270. doi:10.1097/ALN.0b013e31827773b2. [2] C. Frerk, V. S. Mitchell, A. F. McNarry, C. Mendonca, R. Bhagrath, A. Patel, E. P. O’Sullivan, N. M. Woodall, I. Ahmad, Difficult Airway Society 2015 guidelines for management of unanticipated difficult intubation in735 adults, Br J Anaesth 115 (6) (2015) 827–848. doi:10.1093/bja/aev371. [3] M. E. Detsky, N. Jivraj, N. K. Adhikari, J. O. Friedrich, R. Pinto, D. L. Simel, D. N. Wijeysundera, D. C. Scales, Will this patient be difficult to intubate? The rational clinical examination systematic review, J Am Med Assoc 321 (5) (2019) 493–503. doi:10.1001/jama.2018.21413.740 [4] A. K. Nørskov, C. V. Rosenstock, J. Wetterslev, G. Astrup, A. Afshari, L. H. Lundstrøm, Diagnostic accuracy of anaesthesiologists’ prediction of difficult airway management in daily clinical practice: A cohort study of 188 064 patients registered in the Danish Anaesthesia Database, Anaesthesia 70 (3) (2015) 272–281. doi:10.1111/anae.12955.745 [5] S. R. Mallampati, S. P. Gatt, L. D. Gugino, S. P. Desai, B. Waraksa, D. Freiberger, P. L. Liu, A clinical sign to predict difficult tracheal intubation; a prospective study, Can Anaesth Soc J 32 (4) (1985) 429–434. doi:10.1007/BF03011357. [6] A. Vannucci, L. F. Cavallone, Bedside predictors of difficult intubation: A750 systematic review, Minerva Anestesiol 82 (1) (2016) 69–83. [7] D. Roth, N. L. Pace, A. Lee, K. Hovhannisyan, A. M. Warenits, J. Arrich, H. Herkner, Bedside tests for predicting difficult airways: An abridged Cochrane diagnostic test accuracy systematic review, Anaesthesia 74 (7) (2019) 915–928. doi:10.1111/anae.14608.755 [8] C. Matava, E. Pankiv, L. Ahumada, B. Weingarten, A. Simpao, Artificial intelligence, machine learning and the pediatric airway, Paediatr Anaesth 30 (3) (2020) 264–268. doi:10.1111/pan.13792. [9] J. C. Alexander, B. T. Romito, M. C. C¸ obano˘glu, The present and future role of artificial intelligence and machine learning in anesthe-760 siology, Int Anesthesiol Clin 58 (4) (2020) 7–16. doi:10.1097/AIA. 0000000000000294. [10] D. A. Hashimoto, E. Witkowski, L. Gao, O. Meireles, G. Rosman, Artificial intelligence in anesthesiology: Current techniques, clinical applications, and limitations, Anesthesiology 132 (2020) 379–394. doi:10.1097/ALN.765 0000000000002960. 34
[11] Q. Yan, H. Yan, F. Han, X. Wei, T. Zhu, SVM-based decision support system for clinic aided tracheal intubation predication with multiple features, Expert Syst Appl 36 (3 PART 2) (2009) 6588–6592. doi: 10.1016/j.eswa.2008.07.076.770 [12] O. Langeron, P. Cuvillon, C. Iba˜nez-Esteve, F. Lenfant, B. Riou, Y. Le Manach, Prediction of difficult tracheal intubation: Time for a paradigm change, Anesthesiology 117 (6) (2012) 1223–1233. doi:10.1097/ALN. 0b013e31827537cb. [13] J. H. Kim, H. Kim, J. S. Jang, S. M. Hwang, S. Y. Lim, J. J. Lee,775 Y. S. Kwon, Development and validation of a difficult laryngoscopy prediction model using machine learning of neck circumference and thyromental height, BMC Anesthesiol 21 (1) (2021) 125. doi:10.1186/ s12871-021-01343-4. [14] S. Yamanaka, T. Goto, K. Morikawa, H. Watase, H. Okamoto, Y. Hagi-780 wara, K. Hasegawa, Machine learning approaches for predicting difficult airway and first-pass success in the emergency department: Multicenter prospective observational study, Interact J Med Res 11 (1) (2022) e28366. doi:10.2196/28366. [15] M. Zhou, W. Y. Xu, S. Xu, Q. L. Zang, Q. Li, L. Tan, Y. C. Hu, N. Ma,785 et al., Predicting difficult airway intubation in thyroid surgery using multiple machine learning and deep learning algorithms, Front Public Heal 10 (2022) 1–14. doi:10.3389/fpubh.2022.937471. [16] M. Wilson, D. Spiegelhalter, J. Robertson, P. Lesser, Predicting difficult intubation, Br J Anaesth 61 (2) (1988) 211–216. doi:10.1093/bja/61.2.790 211. [17] M. Naguib, F. L. Scamman, C. O’Sullivan, J. Aker, A. F. Ross, S. Kosmach, J. E. Ensor, Predictive performance of three multivariate difficult tracheal intubation models: A double-blind, case-controlled study, Anesth Analg 102 (3) (2006) 818–824. doi:10.1213/01.ane.0000196507.19771.b2.795 [18] A. K. Chhina, R. Jain, P. L. Gautam, J. Garg, N. Singh, A. Grewal, Formulation of a multivariate predictive model for difficult intubation: A double blinded prospective study, J Anaesthesiol Clin Pharmacol 34 (1) (2018) 62–67. doi:10.4103/joacp.JOACP_230_16. [19] T. E. Tavolara, M. N. Gurcan, S. Segal, M. K. K. Niazi, Identification of800 difficult to intubate patients from frontal face images using an ensemble of deep learning models, Comput Biol Med 136 (2021) 104737. doi:10.1016/ j.compbiomed.2021.104737. [20] T. Hayasaka, K. Kawano, K. Kurihara, H. Suzuki, M. Nakane, K. Kawamae, Creation of an artificial intelligence model for intubation difficulty805 classification by deep learning (convolutional neural network) using face 35
images: an observational study, J Intensive Care 9 (1) (2021) 1–14. doi:10.1186/s40560-021-00551-x. [21] H. Y. Cho, K. Lee, H. J. Kong, H. L. Yang, C. W. Jung, H. P. Park, J. Y. Hwang, H. C. Lee, Deep-Learning model associating lateral cervi-810 cal radiographic features with Cormack–Lehane grade 3 or 4 glottic view, Anaesthesia 78 (1) (2023) 64–72. doi:10.1111/anae.15874. [22] N. Suzuki, S. Isono, T. Ishikawa, Y. Kitamura, Y. Takai, T. Nishino, Submandible angle in nonobese patients with difficult tracheal intubation, Anesthesiology 106 (5) (2007) 916–923. doi:10.1097/01.anes.815 0000265150.71319.91. [23] C. W. Connor, S. Segal, Accurate classification of difficult intubation by computerized facial analysis, Anesth Analg 112 (1) (2011) 84–93. doi: 10.1213/ANE.0b013e31820098d6. [24] G. L. Cuendet, P. Schoettker, A. Y¨uce, M. Sorci, H. Gao, C. Perruchoud,820 J. P. Thiran, Facial image analysis for fully automatic prediction of difficult endotracheal intubation, IEEE Trans Biomed Eng 63 (2) (2016) 328–329. doi:10.1109/TBME.2015.2457032. [25] Y. Wu, Q. Ji, Facial landmark detection: A literature survey, Int J Comput Vis 127 (2) (2019) 115–142. doi:10.1007/s11263-018-1097-z.825 [26] Y. Sun, X. Wang, X. Tang, Deep learning face representation from predicting 10,000 classes, in: Proc IEEE Conf Comput Vis Pattern Recognit, 2014, pp. 1891–1898. doi:10.1109/cvpr.2014.244. [27] G. Guo, N. Zhang, A survey on deep learning based face recognition, Comput Vision Image Understanding 189 (2019) 102805. doi:10.1016/830 j.cviu.2019.102805. [28] M. Pantic, L. J. Rothkrantz, Automatic analysis of facial expressions: The state of the art, IEEE Trans Pattern Anal Mach Intell 22 (12) (2000) 1424– 1445. doi:10.1109/34.895976. [29] S. Li, W. Deng, Deep facial expression recognition: A survey, IEEE Trans835 Affective Comput 13 (3) (2022) 1195–1215. doi:10.1109/taffc.2020. 2981446. [30] K. Patel, D. Mehta, C. Mistry, R. Gupta, S. Tanwar, N. Kumar, M. Alazab, Facial sentiment analysis using AI techniques: State-of-the-art, taxonomies, and challenges, IEEE Access 8 (2020) 90495–90519. doi:10.1109/access.840 2020.2993803. [31] L. Tu, A. R. Porras, A. Morales, D. A. Perez, G. Piella, F. Sukno, M. G. Linguraru, Three-dimensional face reconstruction from uncalibrated photographs: Application to early detection of genetic syndromes, in: Proc Int Work Uncertain Safe Util Mach Learn Med Imaging, 2019, pp. 182–189.845 doi:10.1007/978-3-030-32689-0_19. 36
[32] X. Gao, Y. Su, X. Li, D. Tao, A review of active appearance models, IEEE Trans Syst Man Cybern Part C Appl Rev 40 (2) (2010) 145–158. doi:10.1109/TSMCC.2009.2035631. [33] D. Cristinacce, T. F. Cootes, et al., Feature detection and tracking with850 constrained local models, in: BMVC, Vol. 1 (2), 2006, p. 3. [34] J. Deng, A. Roussos, G. Chrysos, E. Ververas, I. Kotsia, J. Shen, S. Zafeiriou, The Menpo benchmark for multi-pose 2D and 3D facial landmark localisation and tracking, Int J Comput Vis 127 (6-7) (2019) 599–624. doi:10.1007/s11263-018-1134-y.855 [35] E. Antonakos, J. Alabort-I-Medina, G. Tzimiropoulos, S. P. Zafeiriou, Feature-based Lucas-Kanade and active appearance models, IEEE Trans Image Process 24 (9) (2015) 2617–2632. doi:10.1109/TIP.2015.2431445. [36] X. Xiong, F. De La Torre, Supervised descent method and its applications to face alignment, Proc IEEE Comput Soc Conf Comput Vis Pattern860 Recognit (2013) 532–539doi:10.1109/CVPR.2013.75. [37] L. Ale, X. Fang, D. Chen, Y. Wang, N. Zhang, Lightweight deep learning model for facial expression recognition, in: Proc IEEE Int Conf Trust Secur Priv Comput Commun, 2019, pp. 707–712. doi:10.1109/TrustCom/ BigDataSE.2019.00100.865 [38] R. Ranjan, V. M. Patel, R. Chellappa, HyperFace: A deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition, IEEE Trans Pattern Anal Mach Intell 41 (1) (2019) 121–135. doi:10.1109/TPAMI.2017.2781233. [39] X. Zhu, X. Liu, Z. Lei, S. Z. Li, Face alignment in full pose range: A 3D870 total solution, IEEE Trans Pattern Anal Mach Intell 41 (1) (2019) 78–92. doi:10.1109/TPAMI.2017.2778152. [40] G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. van der Laak, B. van Ginneken, C. I. S´anchez, A survey on deep learning in medical image analysis, Med Image Anal 42 (1995) (2017)875 60–88. doi:10.1016/j.media.2017.07.005. [41] E. Gibson, W. Li, C. Sudre, L. Fidon, D. I. Shakir, G. Wang, Z. EatonRosen, R. Gray, et al., NiftyNet: a deep-learning platform for medical imaging, Comput Methods Programs Biomed 158 (2018) 113–122. doi: 10.1016/j.cmpb.2018.01.025.880 [42] C. Tan, F. Sun, T. Kong, W. Zhang, C. Yang, C. Liu, A survey on deep transfer learning, in: Lect Notes Comput Sci, 2018, pp. 270–279. doi: 10.1007/978-3-030-01424-7_27. 37
[43] N. Tajbakhsh, J. Y. Shin, S. R. Gurudu, R. T. Hurst, C. B. Kendall, M. B. Gotway, J. Liang, Convolutional neural networks for medical image885 analysis: Full training or fine tuning?, IEEE Trans Med Imaging 35 (5) (2016) 1299–1312. doi:10.1109/TMI.2016.2535302. [44] I. Goodfellow, Y. Bengio, A. Courville, Deep learning, Cambridge MIT press, 2016. [45] Q. Yang, Y. Zhang, W. Dai, S. J. Pan, Transfer learning, Cambridge Uni-890 versity Press, 2020. doi:10.1017/9781139061773. [46] J. Yosinski, J. Clune, Y. Bengio, H. Lipson, How transferable are features in deep neural networks?, in: Adv Neural Inf Process Syst, 2014, pp. 3320– 3328. [47] J. Deng, W. Dong, R. Socher, L.-J. Li, Kai Li, Li Fei-Fei, ImageNet: A895 large-scale hierarchical image database, in: Proc IEEE Conf Comput Vis Pattern Recognit, 2009, pp. 248–255. doi:10.1109/CVPR.2009.5206848. [48] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, et al., ImageNet large scale visual recognition challenge, Int J Comput Vis 115 (3) (2015) 211–252. doi:10.1007/s11263-015-0816-y.900 [49] Z. Liu, P. Luo, X. Wang, X. Tang, Deep learning face attributes in the wild, in: Proc IEEE Conf Comput Vis Pattern Recognit, 2015, pp. 3730–3738. arXiv:1411.7766,doi:10.1109/ICCV.2015.425. [50] M. K¨ostinger, P. Wohlhart, P. M. Roth, H. Bischof, Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark905 localization, in: Proc IEEE Int Work Benchmarking Facial Image Anal Technol, 2011, pp. 2144–2151. doi:10.1109/ICCVW.2011.6130513. [51] R. S. Cormack, J. Lehane, Difficult tracheal intubation in obstetrics, Anaesthesia 39 (11) (1984) 1105–1111. doi:10.1111/j.1365-2044.1984. tb08932.x.910 [52] F. Adnet, S. W. Borron, S. X. Racine, J. L. Clemessy, J. L. Fournier, P. Plaisance, C. Lapandry, The intubation difficulty scale (IDS): Proposal and evaluation of a new score characterizing the complexity of endotracheal intubation, Anesthesiology 87 (6) (1997) 1290–1297. doi: 10.1097/00000542-199712000-00005.915 [53] D. Boisson-Bertrand, J. L. Bourgain, J. Camboulives, V. Crinquette, A. M. Cros, M. Dubreuil, B. Eurin, J. P. Haberer, T. Pottecher, D. Thorin, P. Ravussin, B. Riou, Intubation difficile: Soci´et´e Fran¸caise d’Anesth´esie et de R´eanimation, Expertise collective, Ann Fr Anesth Reanim 15 (2) (1996) 207–214. doi:10.1016/0750-7658(96)85047-7.920 38
[54] C. Szegedy, S. Ioffe, V. Vanhoucke, A. A. Alemi, Inception-V4, InceptionResNet and the impact of residual connections on learning, in: Proc AAAI Conf Artif Intell, 2017, pp. 4278–4284. arXiv:1602.07261. [55] C. A. Ferreira, T. Melo, P. Sousa, M. I. Meyer, E. Shakibapour, P. Costa, A. Campilho, Classification of breast cancer histology images through925 transfer learning using a pre-trained Inception Resnet V2, in: Lect Notes Comput Sci, 2018, pp. 763–770. doi:10.1007/978-3-319-93000-8_86. [56] O. Pelka, F. Nensa, C. M. Friedrich, Annotation of enhanced radiographs for medical image retrieval with deep convolutional neural networks, PLoS One 13 (11) (2018) 1–18. doi:10.1371/journal.pone.0206229.930 [57] L. D. Nguyen, D. Lin, Z. Lin, J. Cao, Deep CNNs for microscopic image classification by exploiting transfer learning and feature concatenation, in: Proc IEEE Int Symp Circuits Syst, Vol. 2018-May, 2018, pp. 3–7. doi: 10.1109/ISCAS.2018.8351550. [58] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L. C. Chen, MobileNetV2:935 Inverted residuals and linear bottlenecks, in: Proc IEEE Conf Comput Vis Pattern Recognit, 2018, pp. 4510–4520. doi:10.1109/CVPR.2018.00474. [59] K. Aguilar, G. H. Alf´erez, C. Aguilar, Detection of difficult airway using deep learning, Mach Vis Appl 31 (1) (2020) 4. doi:10.1007/ s00138-019-01055-3.940 [60] C. Matava, E. Pankiv, S. Raisbeck, M. Caldeira, F. Alam, A convolutional neural network for real time classification, identification, and labelling of vocal cord and tracheal using laryngoscopy and bronchoscopy video, J Med Syst 44 (2) (2020) 44. doi:10.1007/s10916-019-1481-4. [61] R. Kohavi, A study of cross-validation and bootstrap for accuracy estima-945 tion and model selection, in: Proc 14th Int Joint Conf Artif Intellig, 1995, pp. 1137–1143. [62] D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, Tech. rep., arXiv (2014). doi:10.48550/arXiv.1412.6980. [63] J. Alabort-i Medina, S. Zafeiriou, Unifying holistic and parts-based de-950 formable model fitting, in: Proc IEEE Conf Comput Vis Pattern Recogn, 2015, pp. 3679–3688. doi:10.1109/CVPR.2015.7298991. [64] J. Alabort-I-Medina, E. Antonakos, J. Booth, P. Snape, S. Zafeiriou, Menpo: A comprehensive platform for parametric image alignment and visual deformable models, in: Proc ACM Conf Multimed, 2014, pp. 679–955 682. doi:10.1145/2647868.2654890. [65] S. van Buuren, K. Groothuis-Oudshoorn, mice: Multivariate imputation by chained equations in R, Journal of Statistical Software 45 (3) (2011) 1–67. doi:10.18637/jss.v045.i03. 39
[66] V. Kazemi, J. Sullivan, One millisecond face alignment with an ensemble960 of regression trees, in: Proc IEEE Comput Soc Conf Comput Vis Pattern Recognit, 2014, pp. 1867–1874. doi:10.1109/CVPR.2014.241. [67] N. Dalal, B. Triggs, C. Schmid, Human detection using oriented histograms of flow and appearance, Lecture Notes in Computer Science (2006) 428– 441doi:10.1007/11744047_33.965 [68] L. McInnes, J. Healy, J. Melville, UMAP: Uniform manifold approximation and projection for dimension reduction, arXiv 1802.03426 (2018) 1–63. [69] D. Bates, M. M¨achler, B. Bolker, S. Walker, Fitting linear mixed-effects models using lme4, J Stat Softw 67 (1) (2015) 1–48. doi:10.18637/jss. v067.i01.970 [70] A. Kuznetsova, P. B. Brockhoff, R. H. B. Christensen, lmerTest package: Tests in linear mixed effects models, J Stat Softw 82 (13) (2017) 1–26. doi:10.18637/jss.v082.i13. [71] R. V. Lenth, emmeans: Estimated marginal means, aka least-squares means, r package version 1.5.4 (2021).975 [72] A. Tran, C. Nguyen, T. Hassner, Transferability and hardness of supervised classification tasks, in: Proc IEEE Int Conf Comput Vis, 2019, pp. 1395– 1405. doi:10.1109/ICCV.2019.00148. [73] C. V. Nguyen, T. Hassner, M. Seeger, C. Archambeau, LEEP: A new measure to evaluate transferability of learned representations (2020). doi:980 10.48550/arXiv.2002.12462. [74] B. F. Klare, M. J. Burge, J. C. Klontz, R. W. Vorder Bruegge, A. K. Jain, Face recognition performance: Role of demographic information, IEEE Trans Inf Forensics Secur 7 (6) (2012) 1789–1801. doi:10.1109/TIFS. 2012.2214212.985 [75] L. Briceno, G. Paul, MakeHuman: A review of the modelling framework, in: Proc 20th Congr Int Ergonom Assoc – Human Simulation and Virtual Environments, 2019, pp. 224–232. [76] L. Quan, Z.-D. Lan, Linear n-point camera pose determination, IEEE Trans Pattern Anal Mach Intell 21 (8) (1999) 774–780. doi:10.1109/34.784291.990 [77] E. Marchand, H. Uchiyama, F. Spindler, Pose estimation for augmented reality: A hands-on survey, IEEE Trans Visual Comput Graphics 22 (12) (2016) 2633–2651. doi:10.1109/tvcg.2015.2513408. 40