scieee AI-readable full text Open interactive document viewer

Automatic quantification of microglial cells from brain images

Lopes, Diogo Alexandre Rodrigues

Abstract

Microglia are a type of glial cell residing in the central nervous system and represent about 10 to 15% of the brain cell population. These cells don’t produce electrical impulses and are responsible for fundamental physiological and pathological processes, as they represent the first line of immune defence within the central nervous system. Thus, the quantification of these cells is essential in a clinical context, as it allows better monitoring and planning of treatments for different pathologies. Conventional cell counting involves a specific set of tools and devices developed for this purpose. This process is time-consuming and imprecise due to being heavily dependent on the operator. Currently, most processes are performed manually. However, other approaches have been studied and developed to improve the counting process, making it less time-consuming, more efficient and reduce the error associated with factors external to the counting. That said, the objective of this dissertation is to study the best approach to automate the quantification of microglial cells, ranging from classical to deep learning methodologies. Combined with the appropriate image processing and analysis techniques, the classical approach proves to be an adequate solution. However, in recent years, approaches based on deep learning have shown promising performance in various image analysis tasks, such as classification, detection and segmentation. The approaches developed to automate the quantification process were tested on a set of images built in partnership with researchers from the School of Medicine of the University of Minho. As for the classical methodology approach, a protocol was developed within ImageJ, which was combined with image processing techniques that allowed the automation of the counting process. Based on Convolutional Neural Networks, the classification problem referring to a deep learning methodology obtained an accuracy of 0.9021 and managed to classify the 661 images in 5 minutes and 44 seconds. The two approaches, considered optimal within each methodology, are competitive with the state-of-the-art methods, as they allowed for the automation of the quantification process, and showed a significant improvement in reproducibility, efficiency and reduced error associated with human factors.

Full text

Universidade do Minho Escola de Engenharia Diogo Alexandre Rodrigues Lopes Automatic Quantification of Microglial Cells from Brain Images October, 2022 Universidade do Minho Escola de Engenharia Diogo Alexandre Rodrigues Lopes Automatic Quantification of Microglial Cells from Brain Images Master Thesis Dissertation Master Degree in Informatics Engineering Work developed under the supervision of: Paulo Jorge Freitas de Oliveira Novais Bruno Filipe Martins Fernandes October, 2022 COPYRIGHT AND TERMS OF USE OF THIS WORK BY A THIRD PARTY This is academic work that can be used by third parties as long as internationally accepted rules and good practices regarding copyright and related rights are respected. Accordingly, this work may be used under the license provided below. If the user needs permission to make use of the work under conditions not provided for in the indicated licensing, they should contact the author through the RepositoriUM of Universidade do Minho. License granted to the users of this work Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International CC BY-NC-SA 4.0 https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en iv Acknowledgements Countless times over the past year I have imagined writing these lines. For me, they mean closing a stage and reaching a goal. I always thought it would be the easiest task ahead. However, now that the time has come to write them, I see that I couldn’t be more wrong. Including all the people who, over the years, have contributed to my formation as a Man and Engineer is a task doomed to failure. Still, I reserve a few words for those I couldn’t help but mention. First of all, I would like to thank my supervisors for the opportunity they gave me to carry out this work and for the teachings transmitted. To Professor Paulo Jorge Freitas de Oliveira Novais, I want to thank him for all the support, all openness he gave me and for allowing me to work in an area that fascinates me. To Professor Bruno Filipe Martins Fernandes, who over the last year was responsible for dealing with and following me more closely, a task that I admit I did not make easier, I cannot find words to thank him for everything he has done for me. That said, trying not to be unfair to him, I want to thank him for all the dedication, patience and availability he has always shown. Above all, I appreciate all the advice and guidance you gave me, advice that was not always related to my dissertation, but without a shadow of doubt helped me get here. To my colleagues, with whom I had the fantastic opportunity to work and to cross paths in the last two years, I thank all the friendship and help. It was without a doubt a privilege to have learned with all of you and I am sure that great careers await you. To the others who shared moments with me in other years, a huge thank you for making me grow. Last but not least, I thank my family, namely my parents and brother for all the love and unconditional support, dedication, concern, all our conversations and above all, all the countless sacrifices I know you had to do over these years. You made this all possible, allowing me to attend a higher education institution and to be finishing my dissertation. Thanks to you, I can now officially say that I have completed my Master’s in Computer Engineering. v STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the Universidade do Minho. , (Place) (Date) (Diogo Alexandre Rodrigues Lopes) vi “If you really look closely, most overnight successes took a long time.” (Steve Jobs) vii Resumo Quantificação Automática de Células Microgliais a partir de Neuroimagens A microglia é um tipo de célula glial residente no sistema nervoso central e representa cerca de 10 a 15% da população de células cerebrais. Estas células não produzem impulsos elétricos, são responsáveis por processos fisiológicos e patológicos fundamentais, e representam a primeira linha de defesa dentro do sistema nervoso central. Assim, a quantificação destas células é fundamental num contexto clínico, pois permite uma melhor monitorização e planeamento de tratamentos para diversas patologias. A contagem convencional de células envolve um conjunto específico de ferramentas e dispositivos desenvolvidos para esse fim. Este processo é demorado e impreciso devido a estar bastante dependente do operador. Atualmente, a maioria dos processos são feitos manualmente. No entanto, outras abordagens têm sido estudadas, com o intuito de melhorar o processo de contagem, para tornar o mesmo menos demorado, mais eficiente e reduzir o erro associado a fatores externos à contagem. Posto isto, o objetivo desta dissertação é o de estudar a melhor abordagem para automatizar a quantificação de células microgliais indo desde os métodos clássicos aos de deep learning . Combinado com as devidas técnicas de processamento e análise de imagem, a abordagem clássica mostra-se uma solução adequada. Contudo, nos últimos anos, abordagens baseadas em deep learning evidenciaram um desempenho promissor em várias tarefas de análise de imagens, como classificação, deteção e segmentação. As abordagens desenvolvidas para automatizar o processo de quantificação foram testadas num conjunto de imagens construído em parceria com elementos da Escola de Medicina da Universidade do Minho. Quanto à abordagem da metodologia clássica, foi desenvolvido um protocolo dentro do ImageJ, que aliado com técnicas de processamento de imagem permitiu automatizar o processo de contagem. Com base em redes neuronais convolucionais, o problema de classificação referente a uma metodologia de deep learning obteve uma accuracy de 0.9021 e conseguiu classificar as 661 imagens em 5 minutos e 44 segundos. As duas abordagens, consideradas ótimas dentro de cada metodologia, são competitivas com os métodos do estado da arte, pois permitiram automatizar o processo, mostraram uma significativa melhoria na reprodutibilidade e eficiência. Palavras-chave: Células Microgliais, Deep Learning, Processamento de Imagem, Quantificação Automática de Células, Segmentação de Imagem, Sistema Nervoso Central viii Abstract Automatic Quantification of Microglial Cells from Brain Images Microglia are a type of glial cell residing in the central nervous system and represent about 10 to 15% of the brain cell population. These cells don’t produce electrical impulses and are responsible for fundamental physiological and pathological processes, as they represent the first line of immune defence within the central nervous system. Thus, the quantification of these cells is essential in a clinical context, as it allows better monitoring and planning of treatments for different pathologies. Conventional cell counting involves a specific set of tools and devices developed for this purpose. This process is time-consuming and imprecise due to being heavily dependent on the operator. Currently, most processes are performed manually. However, other approaches have been studied and developed to improve the counting process, making it less time-consuming, more efficient and reduce the error associated with factors external to the counting. That said, the objective of this dissertation is to study the best approach to automate the quantification of microglial cells, ranging from classical to deep learning methodologies. Combined with the appropriate image processing and analysis techniques, the classical approach proves to be an adequate solution. However, in recent years, approaches based on deep learning have shown promising performance in various image analysis tasks, such as classification, detection and segmentation. The approaches developed to automate the quantification process were tested on a set of images built in partnership with researchers from the School of Medicine of the University of Minho. As for the classical methodology approach, a protocol was developed within ImageJ, which was combined with image processing techniques that allowed the automation of the counting process. Based on Convolutional Neural Networks, the classification problem referring to a deep learning methodology obtained an accuracy of 0.9021 and managed to classify the 661 images in 5 minutes and 44 seconds. The two approaches, considered optimal within each methodology, are competitive with the state-of-the-art methods, as they allowed for the automation of the quantification process, and showed a significant improvement in reproducibility, efficiency and reduced error associated with human factors. Keywords: Automatic Quantification of Cells, Central Nervous System, Deep Learning, Image Processing, Image Segmentation, Microglial Cells ix LIST OF TABLES 27 CN276 2FD - Slice 1 - ITCN - Quantification Settings for Lobule 5. ........... 69 28 CN276 2FD - Slice 1 - ITCN - Quantification Settings for Lobule 6. ........... 69 29 Cell Colony Sample - Labelling Classes (Number). ................... 70 30 Cell Colony Sample - Labelling Classes (Area). ..................... 73 31 Microglial Cells - Labelling Classes (Number). ..................... 75 32 Microglial Cells - Labelling Classes (Area). ....................... 78 33 Cell Colony Sample - Analyze Particles - Automatic Quantification Results. ........ 82 34 Cell Colony Sample - ITCN - Automatic Quantification Results. ............. 84 35 CN276 2FD - Slice 1 - Analyze Particles - Automatic Quantification Results for Lobule 2. . 85 36 CN276 2FD - Slice 1 - Analyze Particles - Automatic Quantification Results for Lobule 3. . 86 37 CN276 2FD - Slice 1 - Analyze Particles - Automatic Quantification Results for Lobule 4. . 86 38 CN276 2FD - Slice 1 - Analyze Particles - Automatic Quantification Results for Lobule 5. . 87 39 CN276 2FD - Slice 1 - Analyze Particles - Automatic Quantification Results for Lobule 6. . 87 40 CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 2. ....... 88 41 CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 3. ....... 88 42 CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 4. ....... 89 43 CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 5. ....... 89 44 CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 6. ....... 89 45 Generic Cell Sample - Deep Learning Classification Results (Number). ......... 91 46 Generic Cell Sample - Deep Learning Classification Results (Area). ........... 93 47 Microglial Cells - Deep Learning Classification Results (Number). ............ 94 48 Microglial Cells - Deep Learning Classification Results (Area). .............. 96 49 CN276 2FD - Slice 1 - Quantification Settings for DCN. ................. 106 50 CN276 2FD - Slice 1 - Quantification Settings for Lobule 7. ............... 107 51 CN276 2FD - Slice 1 - Quantification Settings for Lobule 8. ............... 107 52 CN276 2FD - Slice 1 - Quantification Settings for Lobule 9. ............... 107 53 CN276 2FD - Slice 1 - Quantification Settings for Lobule 10. .............. 108 54 CN276 2FD - Slice 1 - Quantification Settings for Lobule 11. .............. 108 55 CN276 2FD - Slice 2 - Quantification Settings for DCN. ................. 108 56 CN276 2FD - Slice 2 - Quantification Settings for Lobule 2. ............... 109 57 CN276 2FD - Slice 2 - Quantification Settings for Lobule 3. ............... 109 58 CN276 2FD - Slice 2 - Quantification Settings for Lobule 4. ............... 109 59 CN276 2FD - Slice 2 - Quantification Settings for Lobule 5. ............... 110 60 CN276 2FD - Slice 2 - Quantification Settings for Lobule 6. ............... 110 61 CN276 2FD - Slice 2 - Quantification Settings for Lobule 7. ............... 110 62 CN276 2FD - Slice 2 - Quantification Settings for Lobule 8. ............... 110 xvi LIST OF TABLES 63 CN276 2FD - Slice 2 - Quantification Settings for Lobule 9. ............... 110 64 CN276 2FD - Slice 2 - Quantification Settings for Lobule 10. .............. 111 65 CN276 2FD - Slice 2 - Quantification Settings for Lobule 11. .............. 111 66 CN276 2FD - Slice 2 - Quantification Settings for Lobule 12. .............. 111 67 CN276 2FD - Slice 2 - Quantification Settings for Lobule 13. .............. 112 68 CN276 2FD - Slice 2 - Quantification Settings for Lobule 14. .............. 112 69 CN276 2FD - Slice 2 - Quantification Settings for Lobule 15. .............. 112 70 CN276 2FD - Slice 2 - Quantification Settings for Lobule 16. .............. 112 71 CN276 2FD - Slice 2 - Quantification Settings for Lobule 17. .............. 113 72 CN276 2FD - Slice 2 - Quantification Settings for Lobule 18. .............. 113 73 CN276 2FD - Slice 3 - Quantification Settings for DCN. ................. 113 74 CN276 2FD - Slice 3 - Quantification Settings for Lobule 2. ............... 114 75 CN276 2FD - Slice 3 - Quantification Settings for Lobule 3. ............... 114 76 CN276 2FD - Slice 3 - Quantification Settings for Lobule 4. ............... 114 77 CN276 2FD - Slice 3 - Quantification Settings for Lobule 5. ............... 115 78 CN276 2FD - Slice 3 - Quantification Settings for Lobule 6. ............... 115 79 CN276 2FD - Slice 3 - Quantification Settings for Lobule 7. ............... 115 80 CN276 2FD - Slice 3 - Quantification Settings for Lobule 8. ............... 116 81 CN276 2FD - Slice 3 - Quantification Settings for Lobule 9. ............... 116 82 CN276 2FD - Slice 3 - Quantification Settings for Lobule 10. .............. 116 83 CN276 2FD - Slice 3 - Quantification Settings for Lobule 11. .............. 117 84 CN276 2FD - Slice 3 - Quantification Settings for Lobule 12. .............. 118 85 CN276 2FD - Slice 3 - Quantification Settings for Lobule 13. .............. 118 86 CN282 2TE - Slice 1 - Quantification Settings for DCN. ................. 119 87 CN282 2TE - Slice 1 - Quantification Settings for Lobule 2. ............... 120 88 CN282 2TE - Slice 1 - Quantification Settings for Lobule 3. ............... 120 89 CN282 2TE - Slice 1 - Quantification Settings for Lobule 4. ............... 120 90 CN282 2TE - Slice 1 - Quantification Settings for Lobule 5. ............... 121 91 CN282 2TE - Slice 1 - Quantification Settings for Lobule 6. ............... 121 92 CN282 2TE - Slice 1 - Quantification Settings for Lobule 7. ............... 121 93 CN282 2TE - Slice 1 - Quantification Settings for Lobule 8. ............... 122 94 CN282 2TE - Slice 2 - Quantification Settings for DCN. ................. 122 95 CN282 2TE - Slice 2 - Quantification Settings for Lobule 2. ............... 122 96 CN282 2TE - Slice 2 - Quantification Settings for Lobule 3. ............... 123 97 CN282 2TE - Slice 2 - Quantification Settings for Lobule 4. ............... 123 98 CN282 2TE - Slice 2 - Quantification Settings for Lobule 5. ............... 123 99 CN282 2TE - Slice 2 - Quantification Settings for Lobule 6. ............... 124 xvii LIST OF TABLES 100 CN282 2TE - Slice 2 - Quantification Settings for Lobule 7. ............... 124 101 CN282 2TE - Slice 4 - Quantification Settings for DCN. ................. 124 102 CN282 2TE - Slice 4 - Quantification Settings for Lobule 2. ............... 125 103 CN282 2TE - Slice 4 - Quantification Settings for Lobule 3. ............... 125 104 CN282 2TE - Slice 4 - Quantification Settings for Lobule 4. ............... 125 105 CN282 2TE - Slice 4 - Quantification Settings for Lobule 5. ............... 126 106 CN282 2TE - Slice 4 - Quantification Settings for Lobule 6. ............... 126 107 CN282 2TE - Slice 4 - Quantification Settings for Lobule 7. ............... 126 108 CN282 2TE - Slice 4 - Quantification Settings for Lobule 8. ............... 127 109 CN283 2FD - Slice 1 - Quantification Settings for DCN. ................. 128 110 CN283 2FD - Slice 1 - Quantification Settings for Lobule 2. ............... 129 111 CN283 2FD - Slice 1 - Quantification Settings for Lobule 3. ............... 129 112 CN283 2FD - Slice 1 - Quantification Settings for Lobule 4. ............... 129 113 CN283 2FD - Slice 1 - Quantification Settings for Lobule 5. ............... 130 114 CN283 2FD - Slice 1 - Quantification Settings for Lobule 6. ............... 130 115 CN283 2FD - Slice 1 - Quantification Settings for Lobule 7. ............... 130 116 CN283 2FD - Slice 4 - Quantification Settings for DCN. ................. 131 117 CN283 2FD - Slice 4 - Quantification Settings for Lobule 2. ............... 131 118 CN283 2FD - Slice 4 - Quantification Settings for Lobule 3. ............... 131 119 CN283 2FD - Slice 4 - Quantification Settings for Lobule 4. ............... 131 120 CN283 2FD - Slice 4 - Quantification Settings for Lobule 5. ............... 131 121 CN283 2FD - Slice 4 - Quantification Settings for Lobule 6. ............... 132 122 CN283 2FD - Slice 4 - Quantification Settings for Lobule 7. ............... 132 123 CN283 2FD - Slice 4 - Quantification Settings for Lobule 8. ............... 132 124 CN284 TDTE - Slice 1 - Quantification Settings for DCN. ................ 133 125 CN284 TDTE - Slice 1 - Quantification Settings for Lobule 2. .............. 133 126 CN284 TDTE - Slice 1 - Quantification Settings for Lobule 3. .............. 134 127 CN284 TDTE - Slice 1 - Quantification Settings for Lobule 4. .............. 134 128 CN284 TDTE - Slice 1 - Quantification Settings for Lobule 5. .............. 134 129 CN284 TDTE - Slice 1 - Quantification Settings for Lobule 6. .............. 134 130 CN284 TDTE - Slice 1 - Quantification Settings for Lobule 7. .............. 135 131 CN284 TDTE - Slice 2 - Quantification Settings for DCN. ................ 135 132 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 2. .............. 135 133 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 3. .............. 136 134 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 4. .............. 136 135 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 5. .............. 136 xviii LIST OF TABLES 136 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 6. .............. 137 137 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 7. .............. 137 138 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 8. .............. 137 139 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 9. .............. 138 140 CN284 TDTE - Slice 2 - Quantification Settings for Lobule 10. .............. 138 141 CN284 TDTE - Slice 3 - Quantification Settings for DCN. ................ 138 142 CN284 TDTE - Slice 3 - Quantification Settings for Lobule 2. .............. 138 143 CN284 TDTE - Slice 3 - Quantification Settings for Lobule 3. .............. 138 144 CN284 TDTE - Slice 3 - Quantification Settings for Lobule 4. .............. 139 145 CN284 TDTE - Slice 3 - Quantification Settings for Lobule 5. .............. 139 146 CN284 TDTE - Slice 3 - Quantification Settings for Lobule 6. .............. 139 147 CN276 2FD - Slice 1 - Automatic Quantification Results for DCN. ............ 140 148 CN276 2FD - Slice 1 - Automatic Quantification Results for Lobule 7. .......... 140 149 CN276 2FD - Slice 1 - Automatic Quantification Results for Lobule 8. .......... 141 150 CN276 2FD - Slice 1 - Automatic Quantification Results for Lobule 9. .......... 141 151 CN276 2FD - Slice 1 - Automatic Quantification Results for Lobule 10. ......... 141 152 CN276 2FD - Slice 1 - Automatic Quantification Results for Lobule 11. ......... 141 153 CN276 2FD - Slice 2 - Automatic Quantification Results for DCN. ............ 142 154 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 2. .......... 142 155 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 3. .......... 142 156 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 4. .......... 142 157 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 5. .......... 142 158 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 6. .......... 143 159 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 7. .......... 143 160 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 8. .......... 143 161 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 9. .......... 143 162 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 10. ......... 144 163 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 11. ......... 144 164 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 12. ......... 144 165 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 13. ......... 144 166 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 14. ......... 145 167 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 15. ......... 145 168 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 16. ......... 145 169 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 17. .......... 145 170 CN276 2FD - Slice 2 - Automatic Quantification Results for Lobule 18. ......... 146 171 CN276 2FD - Slice 3 - Automatic Quantification Results for DCN. ............ 146 172 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 2. .......... 146 xix LIST OF TABLES 173 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 3. .......... 146 174 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 4. .......... 147 175 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 5. .......... 147 176 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 6. .......... 147 177 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 7. .......... 147 178 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 8. .......... 148 179 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 9. .......... 148 180 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 10. ......... 148 181 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 11. ......... 149 182 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 12. ......... 149 183 CN276 2FD - Slice 3 - Automatic Quantification Results for Lobule 13. ......... 149 184 CN282 2TE - Slice 1 - Automatic Quantification Results for DCN. ............ 150 185 CN282 2TE - Slice 1 - Automatic Quantification Results for Lobule 2. .......... 150 186 CN282 2TE - Slice 1 - Automatic Quantification Results for Lobule 3. .......... 151 187 CN282 2TE - Slice 1 - Automatic Quantification Results for Lobule 4. .......... 151 188 CN282 2TE - Slice 1 - Automatic Quantification Results for Lobule 5. .......... 151 189 CN282 2TE - Slice 1 - Automatic Quantification Results for Lobule 6. .......... 151 190 CN282 2TE - Slice 1 - Automatic Quantification Results for Lobule 7. .......... 151 191 CN282 2TE - Slice 1 - Automatic Quantification Results for Lobule 8. .......... 152 192 CN282 2TE - Slice 2 - Automatic Quantification Results for DCN. ............ 152 193 CN282 2TE - Slice 2 - Automatic Quantification Results for Lobule 2. .......... 152 194 CN282 2TE - Slice 2 - Automatic Quantification Results for Lobule 3. .......... 152 195 CN282 2TE - Slice 2 - Automatic Quantification Results for Lobule 4. .......... 153 196 CN282 2TE - Slice 2 - Automatic Quantification Results for Lobule 5. .......... 153 197 CN282 2TE - Slice 2 - Automatic Quantification Results for Lobule 6. .......... 153 198 CN282 2TE - Slice 2 - Automatic Quantification Results for Lobule 7. .......... 153 199 CN282 2TE - Slice 4 - Automatic Quantification Results for DCN. ............ 154 200 CN282 2TE - Slice 4 - Automatic Quantification Results for Lobule 2. .......... 154 201 CN282 2TE - Slice 4 - Automatic Quantification Results for Lobule 3. .......... 154 202 CN282 2TE - Slice 4 - Automatic Quantification Results for Lobule 4. .......... 154 203 CN282 2TE - Slice 4 - Automatic Quantification Results for Lobule 5. .......... 155 204 CN282 2TE - Slice 4 - Automatic Quantification Results for Lobule 6. .......... 155 205 CN282 2TE - Slice 4 - Automatic Quantification Results for Lobule 7. .......... 155 206 CN282 2TE - Slice 4 - Automatic Quantification Results for Lobule 8. .......... 155 207 CN283 2FD - Slice 1 - Automatic Quantification Results for DCN. ............ 156 208 CN283 2FD - Slice 1 - Automatic Quantification Results for Lobule 2. .......... 156 xx LIST OF TABLES 209 CN283 2FD - Slice 1 - Automatic Quantification Results for Lobule 3. .......... 157 210 CN283 2FD - Slice 1 - Automatic Quantification Results for Lobule 4. .......... 157 211 CN283 2FD - Slice 1 - Automatic Quantification Results for Lobule 5. .......... 157 212 CN283 2FD - Slice 1 - Automatic Quantification Results for Lobule 6. .......... 157 213 CN283 2FD - Slice 1 - Automatic Quantification Results for Lobule 7. .......... 158 214 CN283 2FD - Slice 4 - Automatic Quantification Results for DCN. ............ 158 215 CN283 2FD - Slice 4 - Automatic Quantification Results for Lobule 2. .......... 158 216 CN283 2FD - Slice 4 - Automatic Quantification Results for Lobule 3. .......... 158 217 CN283 2FD - Slice 4 - Automatic Quantification Results for Lobule 4. .......... 158 218 CN283 2FD - Slice 4 - Automatic Quantification Results for Lobule 5. .......... 159 219 CN283 2FD - Slice 4 - Automatic Quantification Results for Lobule 6. .......... 159 220 CN283 2FD - Slice 4 - Automatic Quantification Results for Lobule 7. .......... 159 221 CN283 2FD - Slice 4 - Automatic Quantification Results for Lobule 8. .......... 159 222 CN284 TDTE - Slice 1 - Automatic Quantification Results for DCN. ........... 160 223 CN284 TDTE - Slice 1 - Automatic Quantification Results for Lobule 2. ......... 160 224 CN284 TDTE - Slice 1 - Automatic Quantification Results for Lobule 3. ......... 161 225 CN284 TDTE - Slice 1 - Automatic Quantification Results for Lobule 4. ......... 161 226 CN284 TDTE - Slice 1 - Automatic Quantification Results for Lobule 5. ......... 161 227 CN284 TDTE - Slice 1 - Automatic Quantification Results for Lobule 6. ......... 161 228 CN284 TDTE - Slice 1 - Automatic Quantification Results for Lobule 7. ......... 162 229 CN284 TDTE - Slice 2 - Automatic Quantification Results for DCN. ........... 162 230 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 2. ......... 162 231 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 3. ......... 162 232 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 4. ......... 162 233 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 5. ......... 162 234 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 6. ......... 163 235 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 7. ......... 163 236 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 8. ......... 163 237 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 9. ......... 163 238 CN284 TDTE - Slice 2 - Automatic Quantification Results for Lobule 10. ......... 163 239 CN284 TDTE - Slice 3 - Automatic Quantification Results for DCN. ........... 164 240 CN284 TDTE - Slice 3 - Automatic Quantification Results for Lobule 2. ......... 164 241 CN284 TDTE - Slice 3 - Automatic Quantification Results for Lobule 3. ......... 164 242 CN284 TDTE - Slice 3 - Automatic Quantification Results for Lobule 4. ......... 164 243 CN284 TDTE - Slice 3 - Automatic Quantification Results for Lobule 5. ......... 164 244 CN284 TDTE - Slice 3 - Automatic Quantification Results for Lobule 6. ......... 164 xxi Glossary cytomorphological Cell morphology. 7,9 embryogenesis Sequential series of dynamic processes that include cell division and growth. 6 macrophages Type of white blood cell that helps eliminate foreign substances by engulfing foreign materials and initiating an immune response. Macrophages are the principal cells involved in chronic inflammation and usually become more prevalent at the site of injury only after days or weeks. 5,7 neuroanatomist Neuroanatomists are an expert in the field of neuroanatomy. Neuroanatomy is the scientific study of the nervous system. 6 neuroglia Is one of several types of cells that function primarily to support neurons, and can be also called as glial cell. Neuroglia exceed the number of neurons in the nervous system. They exist in the nervous systems of invertebrates as well as vertebrates. 5,6,7 phagocytosis Process by which certain living cells called phagocytes ingest or engulf other cells or particles. 9,10 phenotype Are all the observable characteristics of an organism that result from the interaction of its genotype with the environment. Examples of observable characteristics include behaviour, biochemical properties, colour, shape, and size. 5,6,9 xxii Acronyms 2D Two-Dimensional 12,22,23,24,39,58,62 3D Three-Dimensional 24,39 AI Artificial Intelligence 28,29,41,42 ANNs Artificial Neural Networks 14,30,31,32,33,36,40,42,43 CNNs Convolutional Neural Networks 28,29,32,33,38,39,42,44,45,69,70,75,76,78,90,98, 99,100 CNS Central Nervous System 1,5,6,7,8,9,26 CSC Cervical Spinal Cord 47,48 DCN Deep Cerebellar Nuclei 47,48,50,51,53,62,64,65,85,86 DIP Digital Image Processing 11,12,13,16,39 DL Deep Learning 28,29,36,41,42,100 FNNs Feedforward Neural Networks 31,42 GANs Generative Adversarial Networks 36,37,43 IBA1 Ionized Calcium-Binding Adapter Molecule 1 8,9,10,24,37,38,46,48,63,66 ITCN Image-based Tool for Counting Nuclei 57,58,60,61,62,66,67,81,83,84,85,87,88,89, 90 ML Machine Learning 28,29,34,39,41,42,69 xxiii ACRONYMS NN Neural Networks 29,36,37,43,69 PN Pontine Nuclei 47,48 RNNs Recurrent Neural Networks 31,32,42 SVM Support Vector Machine 27 VAEs Variational Autoencoders 36,37,43 xxiv 1 Introduction The main objective of this study is to understand the best approach to automatically quantify microglial cells, focusing on classical and deep learning methodologies. This chapter aims to contextualize the reader for the goal of this work. The motivation behind this dissertation, along with the definition of the main objectives and contributions are described. Ultimately, this chapter also provides an overview of the document outline, summarising the main contents of each chapter. 1.1 Motivation Microglial cells are one of the most important microorganisms in the Central Nervous System (CNS) and are about 10 to 15% of the brain cell population [1]. Microglia is a type of glial cell that doesn’t produce electrical impulses. They are responsible for fundamental physiological and pathological processes, representing the first line of immune defence within the CNS [2]. Upon detection of any sign of brain damage or nervous system dysfunction, these cells undergo an activation process. In the activation process, these cells migrate to the site of injury [3]. Given that importance, the quantification of these cells is fundamental in a clinical context, as it allows for better monitoring and planning of treatments for different diseases. Therefore, cell count is an indispensable procedure routine that, in some cases, helps in the detection of a particular disease. Most of the cell quantification processes are performed manually. This quantification task requires time and is a tedious process. The final results may vary considerably between users, even when they are very experienced. On the other hand, automatic approaches have been studied and proposed over the last few years. The automated counting process has proven to be faster and more effective. This approach gained importance in the medical context, as evidenced by the inclusion of several automatic quantification and segmentation systems in a clinical environment. To overcome the challenging task of the automatic quantification of microglial cells, some available alternatives are derived from the so-called more classic and deep learning-based approaches. Regarding the classical approach, an image that contains scattered cells in a layer and a known area, software solutions and assistants for automatic cell counting are applicable and make the quantification process easier [4]. Deep learning is a branch of machine learning and has been successfully applied to hard 1 CHAPTER 2. MICROGLIAL CELLS - MEDICAL PRESPECTIVE see in Figure 2. The microglial cells seen in vitro do not usually have a branched structure compared to the microglial cells typically seen in the normal CNS [3]. Ramified and Hyper-Ramified Shape Microglial cells assume this particular shape throughout the CNS, specifically throughout the spinal cord and brain. This form is composed of long branches and a small cellular body. The cell body of this ramified shape remains stable while its branches, which are very sensitive to small changes, are constantly moving and supervising the surrounding areas. Microglia in this state can search and identify various immune threats. Although this is considered to be a “resting state”, microglia that remain ramified are extremely active and can be quickly transformed into the activated form to respond to any injury or threat [9]. Reactive Shape Most of the time, when we refer to microglia, we use the term “activated”, although some scientific studies assert that it is no longer correct. We should use the terminology “reactive” microglia. The term used is misleading as it indicates a polarization of cellular reactivity. The Ionized Calcium-Binding Adapter Molecule 1 (IBA1) marker, detailed in subsection 2.4, is upregulated in reactive microglia and is often used to visualize these cells. This is the marker used to help the visualization of the cells in this dissertation through the classic and deep learning methods [7]. When reactivated, microglial cells became phagocytic activated and non-phagocytic. When glial cells became phagocytic activated are the best line of an immune-responsive form that microglia can be. Non-phagocytic cells are glial cells that are moving from their ramified form to their fully active phagocytic form [10]. Figure 2: Representation of Microglial Cells Morphology. Source [11]. 8 2.4. MARKERS OF MICROGLIAL CELLS 2.4 Markers of Microglial Cells To recap, microglia are resident macrophages in the CNS that play an immune defence role. As stated in subsection 2.3.2, they exist assuming a ramified shape and are actively investigating the surrounding areas for injuries or infections. Being static but always looking for threats, they can be quickly ”activated” in response to environmental changes. Once a threat is detected, the microglia undergo some morphological changes. Activated microglia are divided into essentially two states, M1 and M2, based on their morphology. M1 microglia are the first cells to respond to injury, while M2 microglia, also known as the anti-inflammatory microglia, can promote regression of neuroinflammation and stimulate tissue repair [12]. Regarding the microglia marker, there are three different state markers the Steady-State, M1 and M2 Markers. When choosing microglia markers for cell visualization, their location is essential. These markers help us to identify the cell more easily. There are numerous markers for microglial cells, but it is important to point out, that in this dissertation, the samples used to study a way to automatically quantify the number of cells contain the positive and non-positive IBA1 cell marker on them. IBA1 or AIF1 is a cytoplasmic protein, related to microglia motility and phagocytosis and is associated with microglial cell activation. IBA1 is the protein of a ramified microglia [13]. 2.5 Summary Microglial cells are a type of support CNS neuronal cell that don’t produce electrical impulses. These cells reside in the healthy CNS parenchyma and are responsible for fundamental physiological and pathological processes. Glial cells include oligodendrocytes, astrocytes, ependymal cells, and microglia, with this last one representing 10 to 15% of the brain cell population. Microglia function primarily as immune cells because they constantly monitor the CNS microenvironment, being able to detect extracellular changes. Located throughout the brain spinal cord, and in addition to their immune functions, these cells are important in other brain processes such as the regulation of synaptic architecture. These cells migrate into all central nervous system regions and acquire a specific ramified morphological phenotype. Microglia were first discovered between 1919 and 1921 through a study involving histological staining with silver carbonate. However, the origin of microglial cells has been a long-debated topic. Despite this, it is already consensual that these cells are derived from progenitors and migrate into the CNS. When they reach the brain, these cells propagate and disperse in a non-heterogeneous manner throughout the CNS, transforming themselves into a ramified phenotype. They can go from a resting state to an active state. This transition is accompanied by morphological changes. Microglia reduce the complexity of their shape when they shorten or retract their branches. Regarding the purpose of this dissertation, which is the automatic quantification of these cells from brain images, the identification of microglial cells is crucial. The identification of these cells beyond cytomorphological criteria has been facilitated by the development of staining procedures. These procedures 9 CHAPTER 2. MICROGLIAL CELLS - MEDICAL PRESPECTIVE take advantage of the unique expression of molecules in cell types, helping to identify and count them. By setting contrast and colour variations in pixel values, we can count cells when we face a more classic approach. Microglial cells are easily distinguished from other brain cells, which will be readily identified by deep learning methods when trained to recognize and distinguish these cells. Another key point to help in the identification of these cells for posterior counting are the cell markers. About the markers, there are three different state markers, the ”Steady-State” and ”M1 and M2 Markers”. All the samples used in this study have the IBA1 positive and non-positive marker on them. IBA1 is a marker related to microglia motility and phagocytosis, as well as microglial cell activation. 10 3 State of the Art To quantify microglial cells present in the central nervous system, its identification through neuroimaging is indispensable. Roughly, we can do this quantification in two different ways. The manual quantification task presents itself as more time-consuming and error-prone even for the most experienced specialists. As an alternative, automatic quantification systems using classic and deep learning methodologies have shown robust results and the potential to reduce the human error associated with manual quantification. In addition, they have proven to be faster and more effective. They gained importance in the medical context, as evidenced by the inclusion of several automatic quantification and segmentation systems. This chapter introduces the reader to relevant theoretical foundations and gives a technical background related to the scope of this dissertation. Image processing and analysis techniques are required to implement a solution to automatically quantify microglial cells. Therefore, a description of the global aspects of image processing and analysis is presented alongside with the processing techniques applied to enhance some characteristics in brain images and the morphological operators of image analysis. Next follows a brief contextualization of the automatic quantification process. Emphasis is on ImageJ, as it will be the software of choice and analysis of similar work developed regarding this classical approach. This chapter also provides information related to deep learning methodology and its respective algorithms that have been applied successfully to image classification and in this way automate the quantification of cells. As deep learning is a machine learning technique, emphasis is on learning paradigms, object detection and recognition. 3.1 Image Processing Techniques Image processing is the methodology that allows converting an image to a digital aspect. Then, enabling us to perform actions on it, to get an enhanced image or extract quantitative and qualitative information from images using appropriate techniques. Today, image processing is spreading throughout various fields. Therefore, distributed into several groups emerge ”Visualization”, ”Image Retrieval”, and ”Digital Image Processing (DIP)”. Bering in mind the goal of this dissertation, the last-mentioned element of the various image processing groups is the one that needs to be pointed out and will be detailed later. 11 CHAPTER 3. STATE OF THE ART It is worth explaining that an image is a set of distributed data in an array. The positions are defined by elements from the Cartesian plane [14]. Therefore, data values are described as a Two-Dimensional (2D) function, 𝑓(𝑥,𝑦), where 𝑥and 𝑦, are coordinates of the Cartesian plane, and assume integer values from 0to 𝑁−1and 𝑀−1, respectively, in images with size 𝑁×𝑀. Figure 3illustrates and substantiates the presented information. Figure 3: Mathematical Notation for a Digital Image. Adapted From [14]. Sometimes, the images may be squares, making 𝑀equal to 𝑁. Another repeatedly used terminology is the given spaces generated by the intersection of rows and columns of a matrix. That said, 𝑖and 𝑗represent the row and column numbers, respectively, with the same range of integer values. Each element of this matrix or grid is named by picture element, image element, or pixel, being pixel the best-known denomination [14]. Data values, 𝑓, is a discrete representation of the intensity or amount of visible light reflected by an object. Considering our informatics background, we know that an image is visualized by the computer as an array of integers. The application of algorithms for array manipulation is a common practice. So, image processing needs several techniques available with the assistance of computer programs. However, it’s possible to extract qualitative and quantitative information from images using mathematical theorems, despite being a complicated and demanding task. DIP is the technique of processing images performed by computers, eliminating the possibility of extracting information using mathematical theorems. First, the images are converted into a digital form, acquiring a computerized structure, and then further preparation/processing is done on those images. To obtain this computerized structure, image processing uses various techniques such as correction, formatting of the data, and enhanced procedure to create images with better quality [5]. It is important to point out that DIP allows the usage of complex algorithms for more sophisticated or simpler tasks. DIP is a tangible application of classification, feature extraction and pattern recognition. When we talk about 12 3.1. IMAGE PROCESSING TECHNIQUES DIP, there are several techniques associated with it, more precisely seven different techniques. However, we will only focus on four of them, since they are related and are a key element in the automatic cell quantification process using classical and deep learning methodologies. These techniques are ”Image Segmentation”, ”Classification”, ”Image Restoration”and ”Image Enhancement”. 3.1.1 Image Segmentation Image Segmentation allows the partitioning of an image into several regions, several subparts, or even splitting it into pixels, according to the requirements intended by the user. This approach is often used for the analysis of substances, borders and other records relevant to processing [15]. The outcome of image segmentation is a set of sections that cover the total image. The main goal of segmentation is to simplify a raw image in such a manner that makes it easier to evaluate a complete picture. The segmentation of images is performed to compress images and facilitate the recognition of objects and editing purposes. Bearing in mind this dissertation, the part of the recognition of objects associated with the image segmentation will be extremely helpful when applying a more classic approach for automatic cell counting, as described in section 3.3. For that, thresholding the image is also indispensable. Segmentation allocates labels to each pixel so that pixels have similar labels and can share features [16]. This helps the identification and subsequent quantification of cells. There are several image segmentation techniques, as we can see in Figure 4. These techniques are detailed further below. Image Segmentation Feature-Based Clustering ThresholdingEdge-BasedRegion-Based Model-Based Figure 4: Image Segmentation Techniques. Region-Based This technique groups the objects used for segmentation. The region-based method, also known as similarity-based segmentation, requires that certain regions must be together with each other so that the segmentation can take place. The borders of an image are recognized to perform segmentation. The area that is detected for segmentation should be closed, and the thresholding technique is bound with region-based segmentation [17]. Every step of this technique requires at least one pixel for processing purposes. The colour and texture of the image are altered, and a vector is created from the edge flow. Then, further processing is applied to those edges [18]. 13 CHAPTER 3. STATE OF THE ART Edge-Based Another technique used in image segmentation is the edge-detection method also known as edge-based. This technique is formulated as a binary classification problem at the pixel level to identify individual pixels [17]. To recognize pixel values, edges are drawn, and then these edges are compared with other pixels. The basic procedure of this technique starts with extracting information about the edges. Then labelling is done for pixels. The segmentation is performed by the edges, and they must be far from each other. The linking is performed to fill the gap between the edges [18]. Feature-Based Clustering Feature-based clustering, commonly known as clustering, is another option to perform image segmentation. With clustering, an image is changed into a histogram. Following that, the clustering technique itself can be applied to the image. Pixels of the coloured images are clustered for segmentation using an unsupervised technique (Fuzzy C-Means Clustering). The stated procedure is applied to ordinary images, but if there is detected some noise, the result will be image fragmentation [19]. Threshold Thresholding is considered the easiest method used for image segmentation. This approach changes a grayscale image into a binary image wherever the two points are allocated to pixels. These two points are located below and on the upper side of the definite threshold value. The threshold value is obtained from the histogram of the original image and is calculated by the detection of edges which implies that this value is only correct if the detection of the edges is accurate. Segmentation perform via thresholding has lesser calculations in comparison to other methods [20]. Model-Based This technique is based on Markov arbitrary field. For colour segmentation, inbuilt region constraints are implicit. To characterize the exactness of the edges MRF is joined with edge detection. This method contains the relations amongst colour components [21]. 3.1.2 Classification Classification is the technique used to extract data and pixels from images. To perform classification, the minimum requirement is to have as many samples of similar objects as possible. An appropriate classification scheme and an adequate amount of training samples are the basics for effective classification [22]. There are various classification approaches such as Artificial Neural Networks (ANNs) and Fuzzy Logic. The classification technique is either supervised or unsupervised. In supervised classification, spectral signatures are obtained from training samples and are used to classify an image. After that, from the given training pieces, a signature file is assembled. With the help of classification tools, the image is then classified. In unsupervised classification, the output depends on the machine. It’s not necessary any interaction with the user. The statements presented go in line with the principle of supervised and unsupervised learning detailed in subsection 3.4.4. The following diagram (Figure 5) illustrates and helps to describe the working of supervised and unsupervised classification techniques. To sum up, in supervised classification and as previously stated, the 14 3.1. IMAGE PROCESSING TECHNIQUES result is the assembly of a signature file. Following that, various classification techniques are applied to the created file to classify the image. Unsupervised classification deals with clustering, so no samples are collected for further processing. All work is performed with the help of various algorithms by the computer. Data Exploration Classification Clustering Collecting Sample Evaluating Sample Editing Creating Signature File Signature File Examine Editing Signature File Applying Classification Post Classification Processing Supervised Unsupervised Figure 5: Classification Workflow. Adapted From [23]. 3.1.3 Image Restoration As the name suggests, image restoration is the technique through which a corrupted and noisy image is processed in such a manner to construct the ideal image [24]. There are two types of procedures used to reconstruct an image. One technique is to model the image whose quality is degraded, and the other technique, known as image enhancement, increases the quality of the image by applying various filters [25]. In the subsection 3.1.4 more detailed information about image enhancement is given. Notice that to restore an image correctly is important to have prior knowledge of what may be the cause of degradation. 15 CHAPTER 3. STATE OF THE ART The following diagram (Figure 6) shows the degradation and restoration activity. To sum up, the restoration of images is achieved via two types of models, namely the degradation model and the restoration model. On the diagram, the original image is represented by the 𝑓(𝑥,𝑦). After the degradation has taken place, various functions are applied to restore the image. f(x,y) Degradation Function +Restoration Filter p(x,y) g(x,y) Figure 6: Restoration-Degradation Model. 3.1.4 Image Enhancement The image enhancement technique helps to improve the quality of the image. This method modifies certain components in images to increase image clarity. Image enhancement is commonly used to analyse an image for feature extraction. Several algorithms are used in this process. As can be seen in Figure 7, there are two different approaches to image enhancement. The spatial domain technique works with pixels. The pixel values are altered to achieve the desired enhancement. It also contains other techniques constantly working dependent on the pixels. The frequency domain technique is used for images that are based on frequency mechanisms and works on the orthogonal conversion of the image rather than the image itself [23]. Image Enhancement Techniques Spacial Domain Approaches Frequency Domain Approaches Figure 7: Image Enhancement Techniques. Adapted From [23]. 3.2 Image Analysis Techniques Image analysis is the field of digital image manipulation that differs from image processing. The purpose of the image analysis technique is to emulate human vision, including learning and the ability to make a decision based on input [6]. This method is associated with the DIP processes. However, instead of collecting information qualitatively, these systems extract quantitative information from datasets assembled by a set of images, as will be the case in this dissertation. 16 3.2. IMAGE ANALYSIS TECHNIQUES Commonly image analysis techniques are applied to images resulting from image processing techniques. The most typical operations associated with image analysis are morphological analysis, measurements, recognition, representation and description. We can also include segmentation in this group of operations. Regarding this study, the recognition based on object segmentation will be used for assigning a label to objects (e.g. microglial cells). Another relevant aspect to point out, and also related to the main goal of this dissertation, is that morphological analysis relies commonly on the geometric aspects of an object. This means that the parameters related to object morphology such as diameter, area, number, perimeter length, roundness, and extension [14], fit perfectly in the area of image analysis of microglial cells. Besides the most basic mathematical morphological operations, three types of operations characterize the morphological analysis technique. Regarding this study we are going to focus on morphological analysis applicable to binary images, usually derived from thresholding. Operations for size and shape are the most common and easiest morphological operations to use and are related to the properties, local shapes and sizes of objects in images. There are several morphological operators to aid in image analysis. Knowing that we want to quantify the number of microglial cells from brain images the ones that are important to point out are ”Dilation”, ”Erosion”, ”Opening and Closing”, ”Hit-or-miss Transform”, and ”Boundary Extraction”. Next, follows a brief description of each of the mentioned operators. Dilation Dilation is the operation in which object edges are expanded. The central pixel of the structural element (which in our study will be the nucleus of the microglial cell) cycles through all pixels of the target object. This results in a wider object (wider nucleus). Related to this increase in size is the number of pixels of the structural element, as they exceed the limits of the target object when the centre pixel reaches the edge of the target object [26]. Erosion Erosion is the exact opposite of the dilation procedure. Erosion causes a contraction of the edges of an object. The object is reduced according to the shape and size of the structural element. Once the limit of this element reaches the boundary of the object, the pixels between the object boundary and the central pixel of the structural element drop out from the constitution of the object. It is important to point out that once the centre of an element passes through all the pixels of the target object this results in the contraction of limits [26]. Opening & Closing Opening and closing operations are related and are intrinsic to dilation and erosion techniques. The output of opening an object is the same as the result of erosion followed by dilation of an object, causing smoothed contours, broken narrow isthmuses and eliminated small islands and sharp peaks. Analogously, the output obtained by closing an object is the same result that we can obtain with dilation followed by erosion of an object [27]. Hit-or-miss Transform Hit-or-miss transform is an iterative morphological shape detection tool. This operation involves several basic operations such as erosion, complement, intersection and difference. After 17 CHAPTER 3. STATE OF THE ART 3.3.3 Object Counting with ImageJ The big goal of this dissertation is to perform automatic counts of cells, namely microglial cells. The quantification of cells can be done in two different ways, manually and automatically. Regarding the last one, two methodologies are associated with it, a more classical and a deep learning-based one. The objective of this subsection is to review, within relevant work and literature, different classical approaches to the automatic quantification of cells problem. In total, three articles will be brought for discussion to present some aspects regarding how to do object counting with ImageJ. Some of those will be taken into consideration when designing a final solution. Quantifying Microglia Morphology from Photomicrographs of Immunohistochemistry Prepared Tissue Using ImageJ The article produced by Young et al. [34] consists of a set of steps and ImageJ protocols used by the research team to convert fluorescence and bright-field photomicrographs into binary and skeletonized images. Microglial cells are captured in a fluorescent format, and to make certain operations with those images, i.e., count, they need to be converted to a binary format. In addition to the objective of counting these cells, in this study, there was also the objective of analysing their morphology. Introduction In the referred study, the authors detailed in a stepwise way how they use ImageJ plugins to summarize microglia morphology. The analysis techniques were implemented with AnalyzeSkeleton (2D/Three-Dimensional (3D)) and FracLac16 plugins. The first quantifies microglial cell structures, and the other quantifies microglial shapes. These plugins offer a rapid analysis of microglia ramifications within entire photomicrographs. They stated that the use of both tools is not redundant, as cell ramification is complementary to cell complexity. In addition to the statement information, a protocol for counting these cells was also developed. Experimental Work In terms of the developed work, in addition to the objective of counting microglial cells, in this study, there was also the objective of analysing their morphology. Therefore, of all the steps performed, we will only focus on two, step 3 (Imaging) and step 4 (Skeleton Analysis) since they are the ones that make it possible to elaborate an approach to the quantification of microglial cells and are related to the theme of this dissertation. The samples used of microglial cells contain the IBA1 marker. Therefore, in step 3 the goal was to separate channels since cells are only visible in the red channel. It was possible to perform this action using ImageJ’s channel separation feature. To perform the automated count of these cells, in step 4 they used the unsharp mask filter to increase the contrast on the image. Next, they needed to adjust the threshold of the image to convert it into a binary format. Following this protocol and using the analyze particles functionality, they were able to automatically count cells using ImageJ. Results As stated, several steps were implemented to represent microglia morphology according to metrics such as cell ramification, complexity, and shape. The developed protocol steps that helped in the 24 3.3. CLASSIC METHODS FOR AUTOMATIC CELL COUNTING identification of cells as the image noise was removed, and only the cells themselves were left for analysis and counting. Throughout the results obtained, it was possible to prove that ImageJ protocols make microglia morphology quantification available to all laboratories as the platform and related plugins are open-source. Although their main objective was to enhance binary, skeleton, and outline representations of complete photomicrographs and single cells, they were able to make easier the automatic quantification of cells process. It is important to point out that supplementary modifications can be easily made to the protocol depending on image quality and on the actions to reduce noise. Thus, through the results obtained, we conclude that some of the steps will need to be taken into consideration to design a solution for the problem presented in this dissertation. Discussion and Conclusions To sum up, this paper provided a general overview of a developed protocol with recommended ImageJ plugins for automatically accessing the morphology and quantifying microglial cells. The main goal of the protocol is to convert fluorescence photomicrographs into binary skeletonized images. Additionally, they concluded that with the use of this protocol, microglia quantification is accessible to all laboratories, as the plugins used are open-source. This article was also able to verify that it is possible to carry out cell counting in a more automated way, thus avoiding manual counting approaches and processes. Assessment of Cell Counting Method based on Image Processing for a Microalga Culture In Dökümcüoğlu et al. [35] the main objective aimed to illustrate the effectiveness of an image processing approach for the automatic counting of cells in a microalga culture. Therefore, they also develop a protocol with ImageJ to automatically quantify cells. Introduction In this study, the authors also detailed in a stepwise way how they automatically quantify cells in a microalga culture. Their main goal was to prove the usefulness of an image-processing approach for counting cells. Thus, to attest usefulness of image-processing software for cell counts, they compared the performance obtained between the ImageJ cell counts, Utermöhl cell counts, which is a more manual procedure for the quantification of cells, and Optical Density measurements. Experimental Work In terms of the outcome, we are only going to focus on the ImageJ one. Similar to the previous solution, they developed within ImageJ a protocol to automatically quantify cells. They started by calibrating their image and then converting it to 8 bits to minimize colour variation. For better visualization of dark-coloured cells, they subtracted the background through the background subtraction functionality within ImageJ. To obtain a better and clear image, they also applied the rolling ball radius to reduce the noise in their images. To convert the image into a binary format they also adjust the threshold, and to obtain more accurate results, they fill the holes in cell nuclei. Finally, they used the analyze particles functionality to be able to automatically count cells. 25 CHAPTER 3. STATE OF THE ART Results The results obtained attest to the benefit of an image-processing approach to quantify cells. The plots presented in the article represent the results of regression analysis between optical density and cell counting methods. The first one is the analysis between the Thoma chamber and optical density, the second is the Utermöhl chamber and optical density, and the last one is ImageJ counts and optical density. Through the analysis of results, they were able to conclude ImageJ cell counts were finished in a quarter of the time used for manual cell counting under the microscope. In this way, they attested ImageJ can improve and automate the results of the existing automatic counting methods and increase the associated work speed and reliability. Discussion and Conclusions The authors were able to demonstrate the usefulness of image-processing approaches for the quantification of cells. In their case, the samples were a microalga culture. Their approach can be easily applied to the quantification of microglial cells problem. They are similar because the brain image contains multiple glial cells, like in cell culture. To conclude, they also evidence in this study that resorting to ImageJ allows the operator to complete the counting process four times faster than a manual count, with similar accuracy. Automated Segmentation and Analysis of Retinal Microglia within ImageJ In Ash et al. [36] a segmentation routine was proposed to perform automated segmentation and cell counting. Seeking to automate microglia counts, they concluded that few algorithms exist for retina microglia count. The experimental work within the FIJI-ImageJ ecosystem originated a new segmentation routine to perform automated segmentation and cell counting in retinal microglia. They showed that they can perform cell counts with similar accuracy to manual systems. Introduction Through this study, and knowing that microglia are immune cells of the CNS capable of migrating in response to injury, the author’s identified that algorithms aiming to automate microglia counts and morphological analysis are becoming increasingly popular. Few exist that are acceptable for use within the retina and manual analysis remains dominant. With FIJI-ImageJ they performed automated segmentation and cell counting and evidence that their procedure routine can accomplish counts with accuracy comparable to manual observers using the I307N Rho model. Experimental Work In terms of the experimental work, they started with raw images as input. Then they preprocessed those images by normalizing stack intensities followed by the application of a rolling ball, median filter, and Gaussian smoothing routines. The next step was to convert the image to a binary format. The fourth step was cell somas which were identified by morphology using an existing algorithm, and their overlay with the candidate cell masks. The fifth step was labelling. After the application of the watershed algorithm, they were able to identify distinct cells as this algorithm separates with one pixel what he considers to be one or more cells together. Finally, they indicated high overall fidelity with occasional undetected cells and dendrites. The quantification was performed after all these steps were conducted. 26 3.4. DEEP LEARNING FOR AUTOMATIC CELL COUNTING Results The results obtained after the application of the image analysis protocol are images of cells where the noise is reduced and the cells are evident, which undoubtedly facilitates the counting process. The developed procedure segmented contender cell masks by watershed regarding the overlaid cell markers. This method allowed a more accurate definition of microglial morphology. Finally, cell counts were obtained using the labelled image produced within FIJI-ImageJ. Once again, through the results obtained in the analysis and counting process, the team attested to the benefits of automated cell counting processes when compared to more manual processes. Discussion and Conclusions The authors implemented, within the FIJI-ImageJ ecosystem, a new segmentation routine to perform automated segmentation and cell counting in retinal microglia. As the algorithms that automate microglia counts are increasing in popularity, they conclude that few of those are adequate for segmentation and cell counting of retina microglia. Therefore, their approach built entirely with open-source software, addresses the presented problem. Throughout the results, they showed that their routine can perform cell counts with similar accuracy to manual counting but faster, thus evidencing the benefit of automatic approaches to the problem of cell counts. 3.4 Deep Learning for Automatic Cell Counting According to the information stated above the identification and counting of the number of cells from an image are one of the biggest tasks for biomedical image analysis and medical diagnoses [37]. Cell counting, in particular microglial cell counting, is conducted in this study because these cells are responsible for fundamental physiological and pathological processes. Therefore, its identification and quantification may help detect a serious illness. The major weakness of the state-of-the-art techniques for cell counting is the counting dependence on low throughput technology that requires manual counting done by specialists, and adjacent to this, there are high labour costs, error-prone data collection, user subjectivity and fatigue. In recent years, deep learning-based approaches evidenced promising performance in various image analysis tasks, such as classification, detection and segmentation. The most important thing to notice is that this approach has shown similar accuracy to manual counting but a significant enhancement in reproducibility, throughput efficiency and reduced error from human factors. Regarding the cell counting problem, with this approach, the problem can be categorized as detectionbased counting and regression-based counting [38]. The first approach requires the detection or segmentation of every cell before cell counting, which implies a supervised learning process. To convert a counting task into a segmentation task, cell annotation is needed, to train the detection or segmentation model. Each cell is detected one by one through the object detection model, and then the counter takes the detected cells and produces the counting results. To sum up, the first step of this approach is the identification of the cell-like candidate region, the second step is the evaluation of the candidate, which can be done by an Support Vector Machine (SVM). The last step is the selection of a non-overlapping region. 27 CHAPTER 3. STATE OF THE ART Currently, more studies have been focused on regression-based cell counting as they avoid the challenging task of detection or segmentation of single cells because they generate cell density or cell count directly from the images [38]. The Convolutional Neural Networks (CNNs) models are been applied and modified using the Euclidean loss function by taking the total number of cells as the annotation information [39]. The number of cells within a certain region is obtained via the integration of the density map. 3.4.1 Machine Learning To better understand what Machine Learning (ML) is, there are certain basic concepts of Artificial Intelligence (AI) that we must first comprehend. The term AI is commonly used to refer to all computer programs that can think like humans, in other words, AI is defined as a computer program that exhibits human-like cognitive ability. Major AI researchers and books define this field as “the study and design of intelligent agents”, being the intelligent agents the systems that understand the environment and take certain actions to obtain success [40]. Any computer program that shows the stated characteristics, such as self-improvement, learning from inferring, or even basic human tasks, such as image recognition [41, 42], is considered to be a form of AI. The field of AI includes within it the sub-fields of ML and Deep Learning (DL), each with its specific characteristics, but in the end, they are all related. It is important to point out that ML approaches are more probabilistic, which means that the output can be explained, thereby ruling out the black-box nature of AI, unlike DL approaches that are more deterministic. ML is an application of AI [43], in and of itself, is a relatively simple concept. The idea behind ML is that machines should be given access to data and learn specific tasks by themselves to make accurate predictions or be able to behave intelligently without being explicitly programmed. They learn to do better in the future based on past experiments. For that, ML exploits computing systems that learn and predict from data, which can be examples, direct experience, or instruction [44]. Because of this, the main advantage of ML over human learning is the ability to consume huge amounts of data, learn from it, and detect and analyse patterns that outshine human capabilities. The data consists of a set of samples that usually represent the observed variables. It is important to point out that in some cases like the one presented in this dissertation, the set of observed variables is composed of images. As explained, ML algorithms can improve through training from that data, to improve their predictions. Most of those algorithms have certain settings, normally denominated as hyperparameters. During training, those hyperparameters help to control the algorithm’s behaviour. It is noteworthy that a subset of the training set is the one used to choose the hyperparameters for the model. The selection process can be made by trial and error because the validation set isn’t used to train the ML algorithm [45]. To implement an intelligent system learning is required. As stated, a system must be able to perform actions that are associated with intelligence. Despite the complexity of the implementation of learning, this is a subject of great interest to the scientific community due to their large number of application domains, such as medical diagnosis. Today, three prominent methods are being used to train these algorithms. 28 3.4. DEEP LEARNING FOR AUTOMATIC CELL COUNTING These training methods, most commonly known as learning paradigms are supervised, unsupervised and reinforcement learning. In subsection 3.4.4, we can find more detailed information about the learning paradigms. 3.4.2 Deep Learning We can say that DL is a more specialized version of ML because utilizes more complex methods for difficult problems [46]. The term “deep” helps to justify the previous statement as it refers to the number of layers in the network. DL stands as a sub-discipline within ML, that according to what has already been explained is a subset of AI. Despite being a sub-discipline of ML,DL algorithms differ from ML algorithms in their ability to learn from unstructured and unlabelled data. ML algorithms require labelled data, which is a task performed to make data readable for the program. DL algorithms can process raw data without the need for labelling [47]. DL specifically is the use of the concept known as Neural Networks (NN), whereby computers emulate the systems of neurons, similar to those found in the brain, to learn and work [44]. NN and DL for many specialists represent the way forward for AI as we know it, as they will pave the way for human-like AI shortly. NN and other subjects related to the fundamentals of DL are described in more detail in subsection 3.4.3. In DL, the complexity is described in the relationship that variables share. A system that relies on or uses simple concepts and variable relationships to learn more complex concepts is known as a DL algorithm. This goes to find and substantiate the claim DL algorithms differ from ML algorithms in their ability to learn from unstructured and unlabelled data [48]. To sum up this subsection and highlight some of the most important ideas, DL stands out mainly because of these three factors: robustness, generalization capacity, and scalability. A model can automatically learn from raw data, or in other words, can understand the most important features even in the presence of noise [49]. The model’s performance tends to improve when trained more and more. 3.4.3 Fundamentals of Deep Learning Related to the study developed with this dissertation, some theoretical fundamentals of DL are essential to point out, to more easily understand the reasons behind the choice of this approach. A notion of artificial neurons, moving on to how they can be assembled in a network, and finally introducing CNNs are the theoretical foundations of DL presented below. Artificial Neuron The human brain is considered a massive network and has approximately 86 billion neurons [50]. The neuron represents the basic computational unit of the brain. Usually, neuron inputs come from dendrites which are connected to other neurons. The input signals are all processed together according to the strength or weight of their connection. The newly generated signal is then sent through the output axon to other neurons’ dendrites. In Figure 11, we can see a representation of a biological neuron. This one is in genesis and inspired the creation of the mathematical model for artificial neurons. 29 CHAPTER 3. STATE OF THE ART Figure 11: Representation of a Biological and an Artificial Neuron. Adapted from [50]. The artificial neurons are the base element of ANNs. They transform a set of inputs into a single value using a weighted sum since each input has a respective weight. The neuron output values are calculated by applying an activation function. In Figure 12, we can see the representation of the most commonly used activation functions. Following, a small definition and description of these activation functions are presented. Figure 12: Representation of Activation Functions. Source [50]. Sigmoid The function outputs are values between 0 and 1. It saturates low input values closer to 0 and high input values near 1. However, this activation function has a small drawback of not being zerocentred, once the output is always positive. That said, the gradient on the weights will become all positive or negative, according to the signal [50]. Hyperbolic tangent This activation function, also known as Tanh, returns values between -1 and 1. Similar to the sigmoid function, this one also saturates high and low input values. However, there are differences between the two of them. The main difference is that for low input values the output will tend to be -1, and, for high inputs, it will tend to be 1 [50]. 30 3.4. DEEP LEARNING FOR AUTOMATIC CELL COUNTING ReLU This function operates with a threshold of zero. The negative input values are set to zero and the remaining values are left unaltered. This activation function is based on computational simplicity, and for that accelerates training. However, the ReLU definition is fragile and many neurons can irreversibly die during training, in special, with high learning rates [50]. Artificial Neural Networks The term ANNs is usually used to characterize the mathematical model composed of a collection of artificial neurons. As stated in the previous section, these neurons are interlinked in a network by layers to learn more complex data relations [49]. Feedforward Neural Networks (FNNs) and Recurrent Neural Networks (RNNs) stand as two types of ANNs. The FNNs are characterized as being an acyclic graph where the data only can move forward. Unlike these, with RNNs the data/information can flow in any direction, including to the same layer. To sum up, RNNs are an extended representation of FNNs with the addition of feedback connections [49]. As we can see in Figure 13 (a) contains three input neurons, two hidden layers (both with four neurons) and an output layer with a single neuron. Traditional FNNs models can be disposed into fully-connected layers of neurons and the data moves forward. On these layers, only neurons between two layers are fully pairwise connected. That means that neurons within a single layer don’t have connections between themselves [50]. Figure 13: Representation of two types of Artificial Neural Networks. Adapted From [50]. Traditionally FNNs are structured in three sections: • An input layer that receives the data. The data needs to be in a vectorized representation, for instance, and focusing on the problem presented with this dissertation, to deal with images, they must be converted to a One-Dimensional representation, just like a vector; • One or several hidden layers, which are composed of artificial neurons to extract non-linear features from the data; • An output layer that combines a set of non-linear features learned on the previous layers and outputs. 31 CHAPTER 3. STATE OF THE ART The RNNs have the purpose of processing sequential data [49]. The term recurrent is used to refer to the ability that these ANNs have to execute the same job for every element present in a sequence. Figure 13 (b) shows that for each time-stamp 𝑡, the hidden state is calculated by adding the current input with the previous hidden state, in which the function 𝑓is the activation function. In the context of supervised learning, which is described in more detail in subsection 3.4.4, on both types of ANNs, the error between the network predictions and ground truth is back-propagated throughout all networks and is normally used to update the network weights to make the predictions more accurate [51]. Convolutional Neural Networks According to the information presented above, ANNs are composed of a set of artificial neurons that are organized into layers such as the input layer, the hidden layer, and the output layer. CNNs follow the same principle as they derive from ANNs. However, it is important to point out that they differ in one simple detail, as they assume that the input is a set of images. The architecture of CNNs disposes of the neurons differently when compared to ANNs as they look to improve the creation of the model [50]. Neurons are disposed of in three dimensions such as width, height and depth [49]. The idea behind CNNs is to use small ANNs convoluted along with the image to extract relevant information. It is important to point out that this idea differs from a fully connected architecture which demands much more parameters. As a reference, a single fully connected hidden layer requires 120 000 parameters to handle a 200x200 RGB image. This number of parameters is wasteful as leads the model to easily overfit on training [50]. Besides that, with a fully-connect approach, the image pixels are related together. When using smaller CNNs, it is possible to identify local features and add successive convolutions layers. In the Figure, we can better understand how neurons are set on CNNs in comparison to ANNs. Figure 14 displays the traditional representation of ANNs and CNNs. Figure 14: Representation of a Traditional Artificial and Convolutional Neural Networks. Source [50]. The concept of receptive fields is key to ensuring that neurons inside a layer are only connected to a small region of the previous layers and with this avoid the wastefulness of fully-connected neurons. This allows for exploring the local connectivity among neurons. CNNs are capable of generalization on vision problems and help to reduce the number of parameters [50]. As we can see, in Figure 15, typically a 32 3.4. DEEP LEARNING FOR AUTOMATIC CELL COUNTING CNNs is structured in two main sections. First, we have the feature extraction followed by the classification process. The main goals of these two sections are: 1. Feature extraction - plays an important role on CNNs as they are responsible for learning patterns. They learn how to extract the most relevant features automatically and as expected when the network grows deeper. This first section of the CNNs is computationally expensive due to convolutional layers. Activation and pooling layers are alternatives that can be used. However, they have the same problem as convolutional layers because they are also computationally expensive [51]; 2. Classification - this section of CNNs receive the features extracted in the previous section. It can be composed of one or more fully connected layers as well as some dropout layers between them [52]. Figure 15: Typical Structure of Convolutional Neural Networks. Adapted From [53]. Convolutional layers Convolutional layers represent the base element of CNNs, as they are responsible for several computational convolutions. The most correct interpretation for convolution is a crosscorrelation once the convolutional filter works as a feature detector [49,51]. Regarding this study, an input image will produce a big output feature map. However, there are other properties and considerations regarding the convolutional layers that need to be pointed out. It doesn’t make any sense to fully connect all neurons to all pixels in an image because nearby pixels have a higher probability of being correlated than those that are far apart [50,53]. This leads to small connections being made between a neuron and a small receptive field that have the same size as the applied filter. Activation layers Activation layers have the purpose of adding non-linearity properties to the data. As they keep the volume unchanged, they perform an element-wise fixed mathematical operation [50]. These layers are associated with activation functions, the same ones applied to ANNs that are described in this subsection. For activation layers, the simplest activation function that we can use is the ReLU function which makes the training faster without compromising performance. 33 CHAPTER 3. STATE OF THE ART Image segmentation allows partitioning an image into several regions, several subparts, and sometimes even divides the image into pixels to help in the analysis of substances, borders and other records relevant to processing. The outcome is a set of sections that together cover the total image. The main goal of segmentation is to simplify a raw image in such a manner that is easier to evaluate a complete picture. Classification is the technique used to extract data from images and label images. An appropriate classification scheme and an adequate amount of training samples are the basics for effective classification. There are various classification approaches such as ANNs and Fuzzy Logic. The classification technique is either supervised or unsupervised. Image restoration is the technique through which a corrupted and noisy image is processed in such a manner to construct a perfect image. There are two types of procedures used to reconstruct an image. One technique is to model the image whose quality is degraded. The other technique, known as image enhancement, increases the quality by applying various filters. The image enhancement method modifies components on the images to increase image clarity. This makes it easier to identify key features in images. Image analysis techniques emulate human vision, including learning and the ability to make a decision based on input. The information is not collected qualitatively, because these systems extract quantitative information from datasets assembled by a set of images. Typically, image analysis techniques are applied to images resulting from image processing techniques. The most commonly used operations associated with image analysis are morphological analysis, measurements, recognition, representation and description. There are several morphological operators to aid in image analysis. Regarding the quantification of the number of cells from brain images, the ones related are ”Dilation”, ”Erosion”, ”Opening and Closing”, ”Hit-or-miss Transform”, and ”Boundary Extraction”. Dilation is the operation that consists of the expansion of an object boundary. The central pixel of the structural element (e.g. nucleus of the cell) cycles through all pixels of the target object. Erosion is the exact opposite of the dilation operation because it causes a contraction of the boundaries of an object. The object is ”reduced”according to the shape and size of the structural element. Opening and closing operations are related and intrinsic to dilation and erosion techniques. The output of opening an object is the same as the result of erosion followed by the dilation of an object. The output obtained by closing an object is the same result that we can obtain with dilation followed by erosion of an object. Hit-or-miss transform involves several basic operations such as erosion, complement, intersection and difference. The final result includes the coordinates of the object of interest in the image. Boundary extraction is the operation that returns a region of pixels corresponding to the boundary of an object of interest. Conventional cell counting involves specific sets of tools and devices developed for that purpose. This process is tedious, time-consuming, and inaccurate due to operator-dependent biases. Most of the cell counting processes to this date are manual. However, because cell counting is an important procedure routine that may help in the detection of a serious illness, various study reports are focusing on the experience of the development of new systems. With the automation of cell counting, the process is more time-efficient and has fewer errors. In the case of the classic approaches for the automatic quantification of cells, software solutions for cell 40 3.5. SUMMARY counting are applicable. The manual standalone cell counting assistants, plug-ins, and guides facilitate cell counting by replacing the manual clicker with multiple digital counters. ImageJ is the platform of choice for image processing and automatic cell counting as it has several tools to approach the problem. The choice of ImageJ for the implementation of an automatic cell quantification system was based a lot on the premise stated before and on the fact that allows studying and designing a solution where concrete and acceptable results are expected. ImageJ is a java-based program for image processing and analysis, inspired by NIH Image, developed at the U.S. National Institutes of Health. Accessible for the public domain, meaning that the source code is openly available and its use is license-free. The software was designed with an open architecture that provides extensibility via java plugins. With these highly varied plugins, we are capable of modifying an existing function or introducing a brand new one due to the fact of being an open architecture software that provides extensibility. Taking into account the vast list of projects that function within ImageJ it is important to focus on the features that distinguish it from others and why it is the software chosen to be the basis for the study of a solution to the problem presented. The most obvious distinction is its simplicity in image processing. With ImageJ we can do the most basic image processing techniques like background subtraction, brightness and contrast adjustment, image type conversion, smoothing, sharpening, filtering, and binarization. It also contains other features for basic manipulations like geometric transformations such as scaling, zooming, and rotation. It is important to point out that it can display, edit, analyze, process, save and print 8–bit, 16–bit and 32–bit images. Applying image processing tasks is an extremely simple process since it provides an easy-to-understand graphical interface that makes the processing task much easier. To automatically count objects in an image, a set of image processes should be applied to the image so the automatic quantification is more precise. As expected, the results are higher evidence of the cells, as was increased the gap between the background and cell pixels. Before automatic counting, thresholding is a fundamental procedure. A manual adjustment with the sliders may be required to decrease the overlap, but is also possible to rely on the automatic adjustment provided by ImageJ. The result is an image in which the white areas have no interest in quantification. In the case of the DL approaches for the automatic quantification of microglial cells, the problem can be categorized as detection-based counting and regression-based counting. The first approach requires the detection or segmentation of every cell before counting, which implies a supervised learning process. To convert a counting task into a segmentation task, cell annotation is needed, to train the detection or segmentation model. The regression-based cell counting avoids the challenging task of detection or segmentation of single cells because they generate cell density or cell count directly from the images. To better understand the capabilities of this approach, certain basic notions must be addressed, such as ML and DL.ML is a sub-field of AI. In a nutshell, we can define it as the evolving branch of computational algorithms designed to emulate human intelligence by learning from the surrounding environment. Through various techniques, ML models and algorithms can process large amounts of data and extract useful information. As expected, they can improve upon their previous iterations. ML models and algorithms today can outperform the rival state-of-the-art algorithms and human performance. The obtain 41 CHAPTER 3. STATE OF THE ART better results learning paradigms are being applied to train ML algorithms. Currently, ML is being successfully applied in diverse fields rating from pattern recognition, computer vision, finance, entertainment, and computational biology to biomedical and medical applications. DL is a sub-discipline of ML, which itself makes it a sub-field of AI. The main distinguishing factor between ML and DL is that we consider the latter to be more complex. It usually takes human interaction to label the data and make it readable by the program, with ML algorithms dealing with structured, labelled data. However, when we resort to DL algorithms, they can process this data accurately without the need for human labelling. There are some DL theoretical foundations essential to point out like artificial neurons, ANNs and CNNs. The neuron represents the basic computational unit of the brain. Usually, neuron inputs come from dendrites. The artificial neurons are the base element of ANNs. They transform a set of inputs into a single value using a weighted sum since each input has a respective weight. The neuron output values are calculated by applying an activation function. The term ANNs is usually used to characterize the mathematical model composed of a collection of artificial neurons. These neurons are interlinked in a network by layers to learn more complex data relations. FNNs and RNNs stand as two types of ANNs. The FNNs are characterized as being an acyclic graph where the data only can move forward. Unlike these, with RNNs, the data/information can flow in any direction, including to the same layer. RNNs are an extended representation of FNNs with the addition of feedback connections. CNNs follow the same principle as they derive from ANNs. However, it is important to point out that they differ in one simple detail, as they assume that the input is a set of images. The architecture of CNNs disposes of the neurons differently when compared to ANNs as they look to improve the creation of the model. Neurons are disposed of in three dimensions such as width, height and depth. The idea behind CNNs is to use small ANNs convoluted along with the image to extract relevant information. It is important to point out that this idea differs from a fully connected architecture which demands much more parameters. To improve the accuracy of predictive models ML techniques are required. Depending on the nature of the problem, the ML paradigms vary. There are different approaches based on the type and volume of data, which have different amounts and types of supervision in training. Each has its advantages and disadvantages. Supervised learning requires labelled features to define the meaning of the data. The model has a mapping function formed by an algorithm that differs from problem to problem, after being trained to predict the output data for each input. Occasionally, patterns are identified in a subset of the data. Unsupervised learning differs from supervised learning as it can learn from no results. The model tries to learn without any supervision. A machine is provided with just the input to develop a learning pattern. This learning paradigm is best suited to problems that require a massive amount of unlabelled data. Reinforcement learning is a behavioural model. Reinforcement learning uses a method based on rewards and penalties for each action that it takes to train the model. By using this method, this learning paradigm differs from other learning paradigms because it isn’t trained with the sample dataset. Therefore, a sequence of successful decisions results in the process being “reinforced”. DL can easily extract information and features that other methods composed of simple representations can’t. In some cases, some of these models have limitations because they do not have good 42 3.5. SUMMARY generalizations. To solve this problem, VAEs eGANs can be a better alternative. VAEs provide probabilistic descriptions of observations in latent spaces. VAEs have ANNs architecture. Therefore, they belong to the families of probabilistic graphical models and variational bayesian methods. Unlike more common approaches, such as NN as regressors or classifiers, VAEs are a powerful generative model. GANs are a deep generative model. Such as VAEs,GANs algorithms used in unsupervised machine-learning problems. GANs are composed of two neural networks, a generative and a discriminative neural network. The generative and discriminative neural networks complete each other to achieve balance in training. The generative NN is liable for removing noise as input and rendering samples, and the discriminative NN is responsible for evaluating and distinguishing the generated samples from training data. 43 4 Data This chapter presents the reader with all the relevant information about the used data. First, we have a description of a generic cell sample. This cell sample had the purpose of being used to fully understand ImageJ limitations, as well as study what would be the best approach to the problem based on a more classical methodology. With the same cell sample, a deep learning-based approach was also studied. Next, we detail all the information related to the dataset of microglial cells (brain images), ranging from the data overview to its lobule and deep cerebellar nuclei segmentation. Finally, the manual quantification of cells is presented since it will be used to compare the performance of all methodologies. 4.1 Cell Colony Sample According to the study of relevant literature and documentation about the automatic quantification of cells problem, more specifically related to the classic approach detailed in section 3.3, the software of choice is ImageJ. To better understand the extensivity of ImageJ, a generic cell sample was selected to do some experiments. In section 5.2.1, these experiments are detailed. Regarding a deep learningbased approach, and considering all the conclusions brought from the study of all relevant literature and documentation carried out in section 3.4, the best approach to the problem is the use of CNNs. Therefore, the same cell sample was used to study the best strategy for this deep learning model. However, unlike the approach implemented with the more classical methodology, in this case, the cell sample was fractionated into 16 parts to build a better deep learning model. The experiments carried out are detailed in section 5.3.1. Next, follows explicit information about this cell sample. As stated, the cell colony sample was used in the first place to understand the possible limitations of ImageJ and better define a work strategy. Available at https://imagej.net/images/Cell_ Colony.jpg, belonging to the ImageJ documentation and intended for public use, this cell colony sample was selected. Figure 18 illustrates the cell colony sample, and Table 1contains more detailed information about this sample. The image processing and analysis techniques applied to enhance some of the image characteristics, allowing to more easily identify the cells for later counting are detailed in subsection 5.2.1, which is related to the experiments carried out. 44 4.1. CELL COLONY SAMPLE Figure 18: Cell Colony Sample. Size Width Height Pixel Size Bits per Pixel Display Range 162K 406 pixels 408 pixels 1x1 pixel^2 8-bit 0-255 Table 1: Cell Colony Sample Information. Figure 19 illustrates the cell sample with the particularity of being divided into 16 parts. The rationale for this division is related to the fact that it is necessary to build a dataset to test the best strategy and possible limitations of CNNs models to quantify cells. Table 2contains detailed information about each section resulting from the division made. The experiments implemented are detailed in subsection 5.3.1. Size Width Height Pixel Size Bits per Pixel Display Range 10K 101 pixels 102 pixels 1x1 pixel^2 8-bit 0-255 Table 2: Cell Colony Sample (Partitioned) Information. 45 CHAPTER 4. DATA Figure 19: Cell Colony Sample (Partitioned). 4.2 Brain Image Samples This dissertation aims to study the advantages of two different methodologies, a classical and a deep learning-based approach, to automatically quantify microglial cells. In the end, it is expected that one of these methodologies will be a more reliable solution than the current cell counting processes, which are usually performed manually. For this, it was necessary to resort to brain images of four different animals since they will originate the study and the implementations carried out in both approaches. As stated before, microglia are a type of neuronal cell located throughout the brain and spinal cord. Taking this into consideration, the only non-evasive way to access them is through neuroimaging. These brain images contain a large population of microglia that are in a reactive shape. They contain the IBA1 marker, which is a marker upregulated in reactive microglia and used to visualize these cells. The data 46 4.2. BRAIN IMAGE SAMPLES acquisition and the manual quantification of cells were done together and in partnership with the School of Medicine of the University of Minho. The next subsections present and detail all the relevant information related to these images, ranging from data acquisition and overview to the lobule and Deep Cerebellar Nuclei (DCN) segmentation process. The last subsection presents the results of the manual counting process, which with all the conclusions drawn, will help to cease the best solution for the automatic quantification of microglial cells. 4.2.1 Data Acquisition & Overview The dataset is composed of 16 brain images from 4 different animals. To access the microglia density and morphology, 4 coronal brain sections per animal (𝑛=4per genotype) were imaged twice (in both hemispheres) for each region of interest (DCN and Cervical Spinal Cord (CSC)) to yield 4–6 digital photomicrographs per section. For the Pontine Nuclei (PN), 4 sagittal brain sections per animal were used (𝑛=3animals for wild-type and 𝑛=4animals for CMVMJD135), and 2 photomicrographs per section were taken [60]. To stack all images the Olympus Confocal FV1000 laser scanning microscope with a resolution of 1024 × 1024 px using a 40× objective (UPlanSApo, N.A. 0.90; dry; field size 624.39 × 624.39 𝜇m; 0.31 𝜇m/px) was used. The acquisition settings for the images were the following: scanning speed = 4 𝜇m/px; pinhole aperture = 110 𝜇m; Iba-1, excitation = 559 nm, emission = 618 nm; in a 3-dimensional scenario (X, Y, and Z axes). In addition, the ImageJ was used on Z-stacked 3D volume images from sections of the affected brain regions (DCN,CSC, and PN) [60]. 4.2.2 Lobule and Deep Cerebellar Nuclei Segmentation The DCN consist of three nuclei: the fastigial nucleus, the interposed nucleus and the dentate nucleus. Together they form the output of the cerebellum. The fastigial nucleus is the most medially located of the cerebellar nuclei, as they receive input from the vermis [61]. The lobules of the cerebellum are the smallest of the lobes of the flocculonodular lobe. This flattened lobe lies between the posterolateral fissure (inferiorly), the inferior medullary velum, and the cerebellar peduncles (superiorly). That said, to better quantify microglial cells, in this study, the brain images of the animals were separated in several areas. One of these areas is the DCN and the remaining ones are the lobules. As stated, each brain image contains a DCN area, where, as a rule, the number of cells there is higher when compared to the number of cells in the lobules. The remaining areas of the image are composed of several lobules, where we have microglial cells more dispersed throughout the lobe in question. Figure 20 illustrates two examples of the lobule and deep cerebellar nuclei segmentation performed in this study. The image on the right refers to the segmentation performed in slice 1 of the animal CN282 2TE. The image on the left is evidence of the segmentation performed in slice 2 of the same animal. In subsection 4.2.3 the results obtained from the manual quantification of microglial cells are presented, differentiating between the counts of the DCN and lobule areas. The number of lobules that compose each brain image is given. 47 CHAPTER 4. DATA Figure 20: Lobule and Deep Cerebellar Nuclei Segmentation Representation. 4.2.3 Manual Cell Counting With the help of elements from the School of Medicine, the manual cell counting process was conducted. The total count of IBA1 positive cells was obtained using the multi-point tool. The quantification was performed on images acquired with the acquisition settings described in subsection 4.2.1, normalized first to the total image area and then for volume [60]. The data was obtained from individual cells of the CSC (310 microglial cells from wild-type mice and 389 from CMVMJD135 mice), DCN (349 microglial cells from wild-type mice and 445 from CMVMJD135 mice), and PN (152 microglial cells from wild-type mice and 180 from CMVMJD135 mice). Therefore, the total number of analysed microglia with the IBA1 positive marker was 1825. However, the images that compose the dataset not only contain microglia with the IBA1 positive marker. So, bearing in mind that the scope of this dissertation is to automate the quantification of microglial cells from brain images, the results reveal the complete counting process (IBA1 positive and non-positive IBA1 microglial cells), detailed by the animal, lobule and deep cerebellar nuclei cell count. CN276 2FD Table 3presents the results obtained from the counting process and the respective area of each slice. In total, the DCN area of the animal CN276 2FD contains 668 microglial cells. Table 4 details the results brought from the quantification process of slices 1 and 2. Slice 1 contains ten lobules. The number of microglial cells is 842. On the other end, slice 2 contains eighteen lobules with a total number of 1022 microglial cells. Table 5shows the results fetched from the quantification process of the lobules of slices 3 and 4. Slice 3 contains eleven lobules, and slice 4 includes twelve lobules. The number of microglial cells in the lobules of slice 3 is 863, and 1274 in slice 4. 48 4.2. BRAIN IMAGE SAMPLES Slice Number of Cells Area (𝜇m) 1 147 288932.803 2 163 285064.506 3 205 363675.236 4 153 277021.494 Table 3: CN276 2FD - Deep Cerebellar Nuclei Cell Count. Slice Region Number of Cells Area (𝜇m) 1 Lobule 2 73 197029.686 Lobule 3 80 233549.782 Lobule 4 54 221957.245 Lobule 5 38 198076.453 Lobule 6 43 181218.603 Lobule 7 52 166547.527 Lobule 8 116 285983.765 Lobule 9 186 520392.636 Lobule 10 105 426406.744 Lobule 11 95 280652.704 2 Lobule 2 61 212812.877 Lobule 3 47 142765.554 Lobule 4 39 127533.440 Lobule 5 41 125675.798 Lobule 6 43 142936.496 Lobule 7 47 133967.850 Lobule 8 61 226804.968 Lobule 9 74 279216.637 Lobule 10 56 173600.354 Lobule 11 97 315884.563 Lobule 12 64 222045.704 Lobule 13 51 149114.692 Lobule 14 50 197360.412 Lobule 15 32 157927.140 Lobule 16 33 120969.131 Lobule 17 65 265616.635 Lobule 18 87 313512.104 Lobule 19 74 178078.700 Table 4: CN276 2FD - Slice 1 & 2 - Lobule Cell Count. 49 CHAPTER 5. EXPERIMENTAL SETUP learning model were also implemented in the local machine. Table 15 presents the machine specifications. The given information was gathered from Intel Support and ASUS Customer Support. ASUS GL552VX Processor Intel® Core™ i7-6700HQ Processor’s Base Frequency 2.60 GHz Max Turbo Frequency 3.50 GHz # Cores 4 # Threads 8 Cache Memory 6 MB Intel® Smart Cache RAM Memory 16 GB DDR4 2400 MHz Storage 1 TB HHD + 256 GB SSD GPU Nvidia ® GTX 950m 4GB Table 15: Experimental Local Environment Specification. 5.1.2 Cloud Environment All additional experiments and the results gathered with a deep learning-based methodology were executed on the same machine (cloud environment) so that the performance obtained is not affected. Initially, test fits made to the selected deep learning model were implemented in the local machine. However, as Google Colaboratory is probably the most popular hosted Jupyter notebook service in the world, it was easy to understand that for more acceptable results and better execution times, we would have to resort to this cloud environment. Given the size of our dataset, a total of six hundred sixty-one images of microglial cells, and knowing that the free version of Google Colaboratory does not allow for too long execution times, it was necessary to resort to the Pro version. Table 16 presents the specifications of Google Colaboratory Pro. The given information was gathered from Google Colab documentation. Google Colab Pro GPU Nvidia ® K80, P100, T4 GPU Memory 16 GB GPU Memory Clock 0.82 GHz / 1.59 GHz RAM 32 GB CPU 2 x vCPU Performance 4.1 TFLOPS / 8.1 TFLOPS Table 16: Experimental Cloud Environment Specification. 5.2 Classic Methods for Automatic Cell Counting The software chosen to develop a solution for automatic cell counting based on a more classic approach was ImageJ. As stated in subsection 3.3.1, this is a public domain program that is open-source 56 5.2. CLASSIC METHODS FOR AUTOMATIC CELL COUNTING and license-free. This software runs on any computer with a Java 1.5 or later virtual machine. In order not to affect and influence the results obtained in both approaches studied, within the classical methodology, the ImageJ version was the same in both cases. Figure 21 illustrates the version of ImageJ (version 1.53k) used to carry out the work and the version of the Java virtual machine running. Figure 21: ImageJ Version. According to the study carried out, the most suitable solutions to apply to this work and automatically count cells with ImageJ are the Analyze Particles functionality and the Image-based Tool for Counting Nuclei (ITCN) plugin. These two were selected as they have proven to be successfully applied in similar cases. In the end, through the analysis of the results and the entire counting process, the objective is to substantiate their pros and cons, and which one can better answer the problem of cell counting within the classical approaches. In a nutshell, the Analyze Particules functionality consists of a set of commands and steps to count and measure objects in binary or thresholded images. This functionality scans the image or the selected area until it finds the edge of an object. Is necessary to configure the particle analyzer, as particles outside the range specified are ignored. Particles with circularity values outside the range specified are also ignored. In the end, the object is outlined and measured and the results are displayed. This functionality also enables it to work with RGB images. The results are calculated using brightness values and RGB pixels are converted to brightness values. In some cases, other types of image processing are needed, such as work with outlines which this functionality allows. In subsection 5.2.1 all the experiments conducted with the Analyze Particles functionality and the respective cell counting process implemented for the cell colony sample are described. In subsection 5.2.2 using the developed Analyze Particles protocol, the entire counting process which led to the final results, is explained step by step. Developed by the Center for Bio-image Informatics the ITCN is an ImageJ plugin that automatically 57 CHAPTER 5. EXPERIMENTAL SETUP quantifies the number of cells in an image. This plugin’s workflow requires an estimation of the diameter of a cell, which has to be done by the user, an estimation of the minimum distance between cells, also done by the user, and either a region of interest (ROI) selection or a black and white mask image that is white in regions that are to be counted. The following subsections detail all the experiments conducted with the cell colony sample and brain images with the ITCN plugin. In subsection 5.2.1 all the experiments conducted with the ITCN plugin and the respective cell counting process implemented for the cell colony sample are described. In subsection 5.2.2 also with the ITCN plugin, the quantification process is explained. 5.2.1 Cell Colony Sample A particular feature of ImageJ is the possibility of counting objects. These objects can be cells that are in 2D images. Thus, ImageJ fits perfectly into the context of this dissertation since it allows for solving the presented problem. As previously stated, to count objects with ImageJ, following a set of techniques is mandated. Alternatively, a certain plugin can replace the obligation of following a set of pre-established procedures. The Analyze Particles functionality and the ITCN plugin were chosen for the automatic quantification of cells. To understand the extensiveness of these two different strategies, the generic cell sample, detailed in section 4.1, was selected to carry out some experiments. This helped to fully comprehend the best approach to automatically count cells. Analyze Particles The Analyze Particles consists of a set of stages that lead to automatic counting. To count objects, which in this case are the cells present in the cell colony image, a set of processes should be applied to make the count more reliable. Background subtraction is one of these steps as it results in higher evidence of the objects of interest, therefore, increasing the gap between background pixels and objects. In the different experiments, the subtraction of background was tested to verify its benefits. Thresholding the image is also a fundamental procedure. This process specifies the image binarization and what must be included in the analysis based on the intensity of the pixels. ImageJ offers the possibility to adjust the threshold value manually or automatically. The manual adjustment with sliders requires decreasing the overlap as much as possible and is more trustful when compared to the automatic adjustment. In the experiments conducted, several values for the threshold were tested. However, it is important to notice that in the case of this generic cell sample, all areas of interest are included with or without overlap. After removing the background and defining the threshold value follows the automatic quantification. The control menu window displayed in Figure 22, referring to one of the counts performed in the experiments, allows the specification of the range of the size of objects that need to be counted. This control menu also entitles to set which objects are included based on the cell circularity, where the minimum extent means that the object corresponds to a linear line and the maximum corresponds to a flawless circle. The result of this process is an 8-bit image generated from the original image, holding the numbered outlines of the measured objects. 58 5.2. CLASSIC METHODS FOR AUTOMATIC CELL COUNTING Figure 22: Cell Colony Sample - Analyze Particles. Moving on to the experiments conducted, Table 17 displays the different settings applied to the images, namely the threshold values applied in each experiment, the range of size and the circularity of the objects that lead to the final results obtained in the counting process. It is important to notice that in all the tests conducted the circularity chosen for the objects range from 0.00 to 1.00, and the range of sizes defined was 0 to infinity since in the cell sample there are cells with different values of circularity and sizes, and in this way, we are not excluding any cell. 1𝑠𝑡 2𝑛𝑑 3𝑟𝑑 4𝑡ℎ 5𝑡ℎ Threshold Min 0 0 129 0 129 Max 171 202 255 223 255 6.63% 6.11% 6.09% 4.90% 4.89% Algorithm Default Default Default Default Default Analyze Size (micron^2) 0-Inifinity 0-Inifinity 0-Inifinity 0-Inifinity 0-Inifinity Circularity 0.00-1.00 0.00-1.00 0.00-1.00 0.00-1.00 0.00-1.00 Table 17: Cell Colony Sample - Analyze Particles - Quantification Settings. The 1𝑠𝑡 experiment is the result of cell counting after applying auto-thresholding to the image, at a pixel value of 171. Figure 23 illustrates the counting process of this experiment. The 2𝑛𝑑 test presents the results of the quantification process by applying a background subtraction after the passage of a ”rolling ball” with a light background and a radius of 50 pixels. After that, auto-thresholding was applied to the image, at a pixel value of 202. The 3𝑟𝑑 experiment is related to the prior procedure, however, after subtracting the background the resulting image was then converted to binary, resulting in an 8-bit image. Then, the watershed algorithm was applied. What this algorithm does is separate with a pixel what it considers to be two or more cells together in the image. Thus, with the application of this algorithm, we can obtain more accurate results in the counting process, since, without the application of this algorithm, 59 CHAPTER 5. EXPERIMENTAL SETUP the software only counts one cell when in fact there may be two or more cells. Notice that we are only able to apply this algorithm in binary images. To complete the process, auto-thresholding was applied to the image, at a pixel value of 255. The 4𝑡ℎ test is identical to the 2𝑛𝑑 test, however, the background subtraction applied with a light background only had 2 pixels in the radius of the ”rolling ball”. Then again, the auto auto-thresholding was used, at a pixel value of 223. The 5𝑡ℎ and final experiment is similar to the previous one. Once the background was subtracted, the resulting image was then converted to binary, an analogous process to the one conducted in the 3𝑟𝑑 experience, resulting again in an 8-bit image. Then, the watershed algorithm was applied. To finalise the counting process, auto-thresholding was applied to the image, at a pixel value of 255. Figure 23: Cell Colony Sample - 1𝑠𝑡 Experiment - Automatic Quantification. The results obtained and their respective discussion is presented in subsection 6.1.1. An approach considered optimal for the development of the Analyze Particles protocol, to automatically quantify microglial cells, is also explained based on the results acquired with the experiments conducted with the generic cell sample. ITCN As stated, the ITCN is an ImageJ plugin that purpose another approach to automatically count cells in an image. Previous to the quantification of cells, this plugin requires that a set of procedures must be applied to the image to make the count more reliable and effective. Therefore, in 1𝑠𝑡 experiment the background was subtracted by applying auto-thresholding to the image, at a pixel value of 171. Then, using the ITCN plugin, the size of the cell was set to 9 pixels, by measuring the line that joins the beginning to the end of the selected cell body. The minimum distance between cells was defined by a line that joins the cell whose size was measured to the nearest cell. In this case, the distance was 4 pixels. Notice that in this case, the threshold value defined in the ITCN count was 2.0. Figure 24 illustrates the counting process of 60 5.2. CLASSIC METHODS FOR AUTOMATIC CELL COUNTING this experiment. In the 2𝑛𝑑 test, all the settings implemented in the previous experiment were followed, except for the value defined for the ITCN threshold. So, it will be possible to understand how automatic counting behaves when we change the threshold values. The 3𝑟𝑑 experiment is also very similar to the 1𝑠𝑡 . However, it was possible to sense that the results of the counts are more accurate if we find the smallest cell and the smallest distance between cells. Preserving the values of the ITCN threshold higher is also important. Therefore, keeping the values for the auto-thresholding to remove the background, the cell size was defined in 6 pixels and the minimum distance between cells was 5 pixels. The 4𝑡ℎ test is the result of applying a background subtraction after the passage of a ”rolling ball”with a light background and a radius of 50 pixels. Then the image was converted to binary to apply the watershed algorithm. In this case, the new cell size was 5 pixels and the distance 5 pixels. The 5𝑡ℎ and final experiment had a background subtraction applied only had 2 pixels in the radius of the ”rolling ball”. Auto-thresholding was used at a pixel value of 223. The cell size was 5 pixels and the distance between cells is 4 pixels. Figure 24: Cell Colony Sample - 1𝑠𝑡 Experiment - ITCN. Table 18 displays all the different settings applied to the images, namely the threshold values, the cell size and minimum distance between cells defined, along with the ITCN threshold value that leads to the final results obtained in the counting process. The results obtained and their respective discussion is also presented in subsection 6.1.1. The pros and cons of implementing an approach to automatically quantify microglial cells with the ITCN plugin are also discussed, based on the results gathered from the experiments conducted with the generic cell sample. 61 CHAPTER 5. EXPERIMENTAL SETUP 1𝑠𝑡 2𝑛𝑑 3𝑟𝑑 4𝑡ℎ 5𝑡ℎ Threshold Min 0 0 0 129 0 Max 171 171 171 255 223 6.63% 6.63% 6.63% 6.09% 4.90% Algorithm Default Default Default Default Default ICTN Cell With 9 9 6 5 5 Minimum Distance 4 4 5 5 4 Threshold 2.0 1.0 3.0 4.5 5.0 Table 18: Cell Colony Sample - ITCN - Quantification Settings. 5.2.2 Brain Image Samples Within the so-called classic approaches emerges ImageJ. As stated, this software offers the possibility of counting objects in 2D images. Thus, ImageJ fits perfectly into the context of this dissertation, namely for the automatic quantification of microglial cells. Collected together with elements from the School of Medicine of the University of Minho, these images are in the desired format for quantification. With image processing and analysis techniques, once again, the Analyze Particles and the ITCN plugin have been chosen. These two different strategies were implemented to the brain images to fully comprehend who is the best approach to automatically count microglial cells regarding a more classical approach. It is important to note that to obtain more accurate results, all DCN areas and Lobules were divided into several images with the same size, to minimize image noise and obtain a clearer view of the cells. Analyze Particles Having studied all relevant literature and documentation, namely, the approaches presented in subsection 3.3.3, their respective pros and cons, and the drawn conclusions from the experiments carried out with the generic cell sample, the ImageJ Analyze Particles feature was tested again to automate cell counting. The section where similar work was brought up for discussion helped to support this approach and to design a solution to the presented problem. The different studies fetched good practices that need to be considered, from the beginning of the counting process to the end, such as image processing, treatment and analysis. The experiments carried out with the generic cell sample, detailed in subsection 5.2.1, also brought very relevant considerations to this approach. Above all, analyzing the entire counting process and the inherent results, the Analyze Particles has proven to be a very valid alternative for the automatic quantification of microglial cells. As previously described, the Analyze Particles consists of a set of procedures before automatic counting. To count objects, a set of processes should be applied to make the count more reliable. The first of those processes was to divide the DCN areas and the Lobules into several images, all with the same size. Figure 25 illustrates the division of Lobule 2 of animal CN276 2FD from brain slice 1. Similar to the one illustrated, this process of dividing the areas was carried out in all images to help reduce image noise and obtain a clearer view of the cells for the required image processing tasks. Another step developed before 62 5.2. CLASSIC METHODS FOR AUTOMATIC CELL COUNTING the actual counting process was channel adjustment. As previously mentioned, namely, in subsection 2.4, microglia cells are associated with markers that help in their identification. The cell marker that helps in the identification of cells is the ionized calcium-binding adapter molecule 1, most commonly known as IBA1. As we can see in the image below, it is not easy to visualize the cells with the naked eye, even though it contains the IBA1 marker to help with its visualization. Knowing that these images are captured in a fluorescent format, channel separation was performed. The objective was to keep only the red channel for the quantification of cells, as this is where we can visualize them. Figure 26 illustrates the result of channel separation, and as we can see, it is now easier to identify the cells. Figure 25: Microglial Cells - CN276 2FD - Slice 1 - Lobule 2 Images. Figure 26: Microglial Cells - CN276 2FD - Slice 1 - Lobule 2 Red Channel Images. Moving on to the quantification process itself, Table 19 displays the different settings applied to Lobule 2 of brain image 1 of animal CN276 2FD. It’s important to mention that going forward, here is where the counting process varies from image to image. The purpose is to obtain the most accurate and precise results possible. Until now, the division of areas and the separation of channels have been done identically in all images. Threshold values vary in each experiment. A manual adjustment has been made with sliders to decrease the overlap as much as possible. As each image differs from the others, all areas of interest for the quantification are included with or without overlap. After properly adjusting the threshold value, as a good practice for the counting process came the application of the watershed algorithm. As it was possible to verify in the experiments conducted with the generic cell sample, the watershed algorithm separates with 1 pixel what it considers to be one or more overlapping cells, which is extremely important to obtain 63 CHAPTER 5. EXPERIMENTAL SETUP exact results. All images were converted to binary and later to mask, to ensure that the algorithm can be properly applied. Finally, notice that in all the counts the circularity chosen for the objects ranges from 0.05 to 1.00, and the range of sizes is defined as 17 to infinity. It was possible to conclude that these values are the ideal ones through several trial and error tests that were validated with test counts. In this way, maintaining these values in all images, we ensure that we are not excluding any cell, as there are cells with different values of circularity and size in each image. To sum up, the protocol developed in this dissertation is composed of five steps, not counting the process of dividing the DCN areas and Lobules into multiple images. The phases that make up this protocol are the following: 1 - Channel Separation; 2 - Threshold Adjustment; 3 - Image Conversion; 4 - Cells Separation; 5 - Cells Quantification. The first stage is where we use the ImageJ functionality to separate channels, as the cells are only visible in the red channel. The second stage is where we guarantee all areas of interest are included with or without overlap, as we decreased the overlap as much as possible using the sliders for manual adjustment of the threshold values. Following this, the third stage is where the image is firstly converted to binary, resulting in an 8-bit image. Then the 8-bit image is converted to mask to ensure that the next phase can be successfully implemented. The fourth stage is the application of the watershed algorithm. Finally, to finalize the cell counting process comes the fifth phase of the developed protocol. Needs to be pointed out that if all the previous steps of the protocol are not followed, it is not possible to achieve good results. With the Analyze Particles functionality of ImageJ, we defined the previously discussed circularity and the size of objects, ensuring that we are not excluding any cell. 1𝑠𝑡 2𝑛𝑑 3𝑟𝑑 4𝑡ℎ Threshold Min 0 0 0 0 Max 1461 1750 1526 1574 96.42% 96.26% 95.18% 95.04% Algorithm Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 19: CN276 2FD - Slice 1 - Analyze Particles - Quantification Settings for Lobule 2. Focusing more on the second stage of the developed protocol, since every other phase is pretty much the same for every image, the 1𝑠𝑡 experiment conducted in image 1 of Lobule 2 is the result of cell counting after applying manual thresholding to the image at a pixel value of 1461. Figure 27 illustrates the result of the counting process of this experiment. The 2𝑛𝑑 test presents the quantification settings of the conducted process in image 2 of Lobule 2 after once again applying manual thresholding to the image at a pixel value of 1750. The 3𝑟𝑑 experiment is related to the prior procedures. However, manual thresholding to image 3 of Lobule 2 was applied at a pixel value of 1526. The 4𝑡ℎ and final test for this Lobule is identical to the previous ones, as manual thresholding was again applied but at a pixel value of 1574. In the table, it is possible to verify which algorithm was used to threshold the image. In this case, the algorithm used was the Default. The reason for choosing this one was throughout the results obtained 64 5.2. CLASSIC METHODS FOR AUTOMATIC CELL COUNTING in several trial-and-error quantification tests. It should be noted that when we adjust the threshold value in images, we are working with images in black and withe format. In total seven algorithms (Default, Huang, Intermodes, IsoData, MaxEntropy, Otsu and Yen) were tested. From those seven emerged the Default and the Intermodes algorithm. With better results and a little bit superior to the Intermodes algorithm, the Default algorithm proved to be ideal for the problem in question. Figure 27: Microglial Cells - 1𝑠𝑡 Image of Lobule 2 - Analyze Particles Automatic Quantification. Tables 20,21,22 and 23 display the different quantification settings applied to Lobule 3, 4, 5 and 6, respectively. The remaining quantification settings for the remaining DCN areas and Lobules are displayed in Appendix A. The quantification settings of animal CN282 2TE, CN283 2FD and CN284 TDTE are presented in Appendix B,Cand D, respectively. As can be verified, all the steps and values that led to obtaining the results regarding the automatic quantification of microglial cells process based on a more classical approach are fully documented. In this way, anyone who has access to these images and follows exactly all the steps can easily replicate all the results and consequently see their cell counting process optimized. How the entire protocol developed throughout this dissertation is documented, can serve as a basis for an application in a similar situation when dealing with a so-called more classical approach. 1𝑠𝑡 2𝑛𝑑 3𝑟𝑑 4𝑡ℎ 5𝑡ℎ Threshold Min 0 0 0 0 0 Max 1493 1911 1365 1477 1622 96.63% 94.45% 95.96% 95.05% 95.36% Algorithm Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 20: CN276 2FD - Slice 1 - Analyze Particles - Quantification Settings for Lobule 3. 65 CHAPTER 5. EXPERIMENTAL SETUP from each other. Since our dataset only consists of 16 images, going beyond 16 in the value assigned to the batch size would also be irrelevant. So, to be able to tune the model, the following list of values was defined for the different parameters: • batch_size_list = [1, 2, 4, 8, 16]; • epochs_list = [1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35]; • learning_rate_list = [0.0005, 0.005, 0.001]; • apply_data_augmentation_list = [True, False]. Figure 30: Cell Colony Sample - Learning Curve (Number). Several experiments were carried out, with the different values defined in the parameters, totalling 330. The results obtained and a respective discussion is presented in subsection 6.2.1. The approaches considered optimal to help in the problem of the microglial cells are explained. Classification per Area of Cells Similar to the classification approach per number of cells, the objective is again to classify one or several images using previously defined classes. Table 30 presents the classes defined for the classification of the generic cell sample based on the area of cells. To properly define the intervals, in all 16 samples the cells were counted using the Analyze Particles protocol developed in this dissertation. By using this protocol and through the counting results, we can access the percentage of the area that these cells occupy in the image. Thus, an image that does not contain a percentage of cell area greater than 5.2 will be classified 72 5.3. DEEP LEARNING FOR AUTOMATIC CELL COUNTING as ”Few”. An image that contains a percentage of cell area greater than 5.2 and less than 6.9 is classified as ”Average”. Finally, an image that contains a percentage of cell area greater than 6.9 will be classified as ”Many”. Label Interval Few <5.2 Average ⩾5.2&<6.9 Many ⩾6.9 Table 30: Cell Colony Sample - Labelling Classes (Area). Once the intervals for cell classification were defined and the area of the cells was measured, the first step for classification was the labelling of the images. Figure 31 displays the labelling performed on the images. In total 4 images were labelled as ”Few”, 6 as ”Average”, and 6 as ”Many”. Then, it was necessary to do the load data, load labels and define the classes in the model. Leveraging the previous procedure, it was necessary to compile and fit it, passing the model type, the training and test set, and the respective parameters. As the objective is classification, the chosen metric was ”accuracy”. The number of splits defined was the same as in the previous experience. With this number of splits, we can crossvalidate with 16 folds. The error will be an average of 16 folds. For the plot of learning curves, the same procedure as the previous approach was followed. Figure 31: Cell Colony Sample - Image Labelling (Area). 73 CHAPTER 5. EXPERIMENTAL SETUP Before moving on to the analysis of learning curves, it is necessary to point out that once again to guarantee that these results can be replicated on any machine by any user, a seed has been defined in the code. The purpose of learning curves, in this case, is the same as in the previous approach, which is to help us understand the optimal number of epochs for the model. Figure 32 illustrates the learning curves obtained in this approach. Through their analysis, we can see that there is an approximation between the line of training loss and validation loss up to 20 epochs, and then they distance themselves from each other. The point of going up to a relatively high number of epochs with the learning curves plot is to see if both lines of the loss come together again at some point, which was the case when the number of epochs is 60. Going beyond 95 epochs is irrelevant to the model since the training loss and validation loss start getting further and further away from each other. Once again, knowing that our dataset only consists of 16 images, going beyond 16 in the value assigned to the batch size would also be irrelevant. To be able to tune the model, the following list of values for the model parameters was defined: • batch_size_list = [1, 2, 4, 8, 16]; • epochs_list = [1, 2, 3, 4, 5, 10, 15, 20, 60, 65, 70, 75, 80, 85, 90, 95]; • learning_rate_list = [0.0005, 0.005, 0.001]; • apply_data_augmentation_list = [True, False]. Figure 32: Cell Colony Sample - Learning Curve (Area). In total, 480 experiments were carried out, with the different values defined in the parameters. The results obtained and a respective discussion is presented in subsection 6.2.1. The best approaches are 74 5.3. DEEP LEARNING FOR AUTOMATIC CELL COUNTING described based on the results acquired from the experiments conducted with the generic cell sample are explained based on the results acquired. 5.3.2 Brain Image Samples The objective of this dissertation is to automate the cell counting process, namely the quantification of microglial cells. As already explained, these cells are associated with the markers that help in their identification. An approach to automate the microglia cell counting process is image classification. According to relevant literature and documentation, namely, the approaches presented in subsection 3.3.3, the use of a model based on CNNs proves to be an ideal solution for the problem. The experiences carried out with the cell colony detailed in subsection 5.3.1 also came to attest that the conclusions drawn from the articles brought for discussion were correct, namely that the use of CNNs is ideal for automating the microglial cell counting process. Both the approaches presented and those developed fetched good practices that need to be considered in our case. Therefore, using a model based on CNNs two approaches that aim to automate cell counting were applied to the brain images. The goal is to understand the extensiveness of these two strategies, namely their respective pros and cons. Taking advantage of the work produced within the classical approach, the division performed on the images was reused, thus composing a dataset with 661 images. With this number of images, we guarantee that the model can work properly and have adjusted results. Next, the experiments conducted with these images, namely the classification approach based on the number of cells in an image and the percentage of area that the cells occupy in an image are explained. Classification per Number of Cells Once again the objective of a classification approach is to classify one or several images using previously defined classes. Analyzing the entire counting process and the results inherent to the experiments carried out with the generic cell sample, the classification approach based on the number of cells has proven to be a very valid alternative for automating the quantification of microglial cells. That said, Table 31 presents the classes defined. To correctly define the intervals, we resorted to the cell count results for each image, obtained from the Analyze Particles protocol developed in this dissertation. Thus, an image that does not contain more than 15 cells will be classified as ”Few”. An image that contains more than 15 cells and less than 25 is classified as ”Average”. Finally, an image that has more than 25 cells will be classified as ”Many”. Label Interval Few <15 Average ⩾15 &<25 Many ⩾25 Table 31: Microglial Cells - Labelling Classes (Number). 75 CHAPTER 5. EXPERIMENTAL SETUP Having the intervals for cell classification defined and the cells counted, image labelling is the next step required. Figure 33 displays part of the labelling performed on the images. In total 339 were labelled as ”Few”, 203 as ”Average”, and 119 as ”Many”. In terms of the code structure for building the model, the approach followed was very similar to the one implemented with the cell colony. Once again, following the TensorFlow documentation, the appropriate CNNs model was developed to help automate microglial cell counting. The model was created so it was necessary to compile and fit it. For that, and since excellent results were obtained previously, the same parameters were passed to the model. The parameters in question were the following: the model type, the training and test set, the batch size, the number of epochs, the value of the learning rate, and the application of data augmentation. The data augmentation value has been initialized to false. As in the previous cases, the metric defined for the model is ”accuracy” since it is a classification approach. The number of splits for the model was also defined, which in this specific case was 5. We can do cross-validation with 5 folds. The error will be the average of the 5 folds. In this way, we don’t add too much difficulty to the model and manage to have an equally effective but faster process. That said, we are left with 529 images for training and one 132 for testing. This is done for each fold, which makes training and testing a little more time-consuming, but will generate more reasonable results. Figure 33: Microglial Cells - Image Labelling (Number). 76 5.3. DEEP LEARNING FOR AUTOMATIC CELL COUNTING After creating and defining the model was necessary to plot the learning curves. As expected, the data was prepared by reading the training and test set and the classes. To understand how the model behaves, the most complex values were assigned to the parameters to plot the learning curves. The values in question were one for the batch size, 50 for the number of epochs, 0.0005 for the learning rate, and data augmentation has been initialized to true. The number of folds for the plot of learning curves was 4, meaning that we train with 496 cases to predict one 165. This adds complexity to the model and gives us an idea of how it will behave. Once again, it is important to point out that to guarantee that these results can be replicated on any machine by any user, a seed has been defined. Figure 34: Microglial Cells - Learning Curve (Number). Moving on to a more detailed analysis of the results obtained with the implementation of learning curves, we came to the idea that the ideal value for the number of epochs goes up to 20. Figure 34 illustrates the learning curves obtained in this approach. As explained, the purpose of learning curves is to help us understand the optimal values for the model, namely the number of epochs. As the lines of training loss and validation loss are getting further and further away from each other when the number of epochs is higher than 20 is irrelevant to go above that value. Knowing that our dataset is composed of 661, in this case, it is already relevant to increase the maximum value for the batch size to 32. So, to tune the model, the following list of values was defined for the different parameters: • batch_size_list = [1, 2, 4, 8, 16, 32]; • epochs_list = [1, 2, 3, 4, 5, 10, 15, 20]; • learning_rate_list = [0.0005, 0.005, 0.001]; 77 CHAPTER 5. EXPERIMENTAL SETUP • apply_data_augmentation_list = [True, False]. In total, 288 experiments were carried out, with the different values defined in the parameters. The results obtained and a respective discussion is presented in subsection 6.2.2. The best approaches are described based on the results acquired from the experiments conducted. The procedure considered optimal that helps to automate the microglial cell counting is described in detail as it will be brought for discussion within the various methodologies, manual, classic and deep learning. Classification per Area of Cells Another approach that helps automate the cell counting process is the classification technique based on the area that cells occupy in an image. Similar to the previous procedure, the objective is again to classify one or several samples using previously defined classes. Table 30 presents the defined classes. To correctly define the intervals, we resorted again to the cell count results obtained from the Analyze Particles protocol developed in this dissertation. Thus, an image that does not contain a percentage of cell area greater than 0.45 will be classified as ”Few”. An image that contains a percentage of cell area greater than 0.45 and less than 0.85 is classified as ”Average”. Finally, an image that contains a percentage of cell area greater than 6.9 will be classified as ”Many”. Label Interval Few <0.45 Average ⩾0.45 &<0.85 Many ⩾0.85 Table 32: Microglial Cells - Labelling Classes (Area). Once defined the intervals for cell classification and the area of the cells was measured, the second procedure was image labelling. Figure 33 displays part of the labelling performed on the dataset. In total 324 images were labelled as ”Few”, 270 as ”Average”, and 67 as ”Many”. Then, following the same approach as the previous procedure, it was necessary to do the load data, load labels and define the classes in the model. Up next, was necessary to compile and fit it CNNs model, passing the model type, the training and test set, and the respective parameters. Leveraging the previous procedure, the number of splits defined was the same as in the previous experience since excellent results were obtained. With 5 splits we can do cross-validation with 5 folds, and the error will be the average of the 5 folds. To understand how the model behaves was necessary to plot the learning curves. Therefore, the most complex values were assigned to the different parameters. The values in question were the same as in the previous procedure and were: one for the batch size, 50 for the number of epochs, 0.0005 for the learning rate, and data augmentation has been initialized to true. The number of folds for the learning curves plot was also the same since in this way we added complexity as much as possible to the model and gives us an idea of how it will behave. A seed was again created to ensure that these results can be replicated on any machine by any user. 78 5.3. DEEP LEARNING FOR AUTOMATIC CELL COUNTING Figure 35: Microglial Cells - Image Labelling (Area). Analyzing the plots we concluded that the ideal value for the number of epochs goes up to 25. Figure 34 illustrates the learning curves obtained in this procedure. Once again, the purpose of learning curves is to help us understand the optimal values for the model, namely the number of epochs. After 25, the lines of training and validation loss are getting further and further away from each other so it is not relevant to go beyond the 25 epochs. As with the previous approach, the maximum value for the batch size to 32 since the dataset is composed of 661 samples. To tune the model, the following list of values was defined for the different parameters: • batch_size_list = [1, 2, 4, 8, 16, 32]; • epochs_list = [1, 2, 3, 4, 5, 10, 15, 20, 25]; • learning_rate_list = [0.0005, 0.005, 0.001]; • apply_data_augmentation_list = [True, False]. In total, 324 experiments were carried out. The results obtained and a respective discussion is presented in subsection 6.2.2. Once again, the best experiences are described based on the results acquired from the experiments conducted. The best approach is described in detail since it can be a solution to 79 CHAPTER 5. EXPERIMENTAL SETUP Figure 36: Microglial Cells - Learning Curve (Area). automate microglial cell counting. This procedure will also be brought for discussion within the various methodologies, manual, classic and deep learning. 80 6 Results and Discussion This chapter presents the reader with all the results of the automated counting experiences performed with the different approaches. Initially, the results of the counting process of the generic cell sample are presented followed by the results of the quantification of microglial cells. Finally, based on the results gathered from the overall experience between different approaches and methodologies, a discussion is presented where all the pros and cons of each procedure are raised since the goal is to find the most suitable solution to the microglial cell automatic quantification problem. 6.1 Classic Methods for Automatic Cell Counting Based on ImageJ, which was the software of choice for automatic cell counting, in terms of a more classic methodology, several experiments were carried out, as detailed in the previous chapter. The experiments conducted with a generic cell sample aimed to understand the possible limitations of the software, as well as to study the selected approaches, namely the Analyze Particles functionality and the ITCN Plugin. So, through the results obtained, a discussion is presented in subsection 6.1.1. This discussion made it possible to conclude what should or should not be done with each one of these approaches in terms of microglial cell counts. Subsection 6.1.2 details the results obtained from the automatic quantification of microglial cells, as well as, presents the discussion of them, and bases the procedure considered optimal, chosen for the comparison between the different approaches, the manual, the classic and the deep learning approach. 6.1.1 Cell Colony Sample As was already explained, this sample was selected to understand the limitations of ImageJ. According to all relevant literature and documentation studied in the state-of-the-art, two different approaches were tested within the classical methodology. Next, the quantification results gathered from the counts performed are presented. For each strategy, results are discussed, followed by the conclusions drawn that will positively influence the microglial cell counting process. 81 CHAPTER 6. RESULTS AND DISCUSSION implement the best approach within the developed protocol. For this, in some images, it was necessary to test different values for the various parameters involved in the quantification process. As was already explained, before the quantification process itself, certain procedures are recommended as described in subsection 5.2.2. Taking advantage of the fact that the results obtained in the previous procedure were excellent, the pre-quantification procedures performed for each image were reused for this approach. This made the quantification process more effective. Image Number of Cells 18 2 25 3 30 4 18 Table 40: CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 2. Moving on to the results, the 1𝑠𝑡 quantification perfomed in image one of Lobule 2 of animal CN276 2FD resulted in 8 cells quantified. The ideal cell size defined was 15 pixels, and the minimum distance between cells was 19 pixels. The 2𝑛𝑑 cell counting process resulted in 25 cells quantified, after defining the cell size in 14 pixels, and the minimum distance between cells was 39 pixels. The 3𝑟𝑑 quantification resulted in 30 cells quantified, and the 4𝑡ℎ resulted in 18microglial cells quantified. The cell size was 10 and 11 pixels respectively, and the minimum distance between cells was 55 and 47 pixels. In all counting processes, the value of the ITCN threshold was adjusted to obtain more accurate and realistic results. Despite all these precautions, as was already foreseeable and it was possible to conclude through the experiments carried out with the generic cell sample, the results were a little off from the real number of cells on the image. It should be noted that even with this performance, in some cases this process was faster and equally accurate when compared to the manual quantification process. Therefore, the remaining images were quantified to be able to substantiate with concrete results which of the classic approaches is the best alternative to the standard counting process.Tables 41,42,43 and 44 display the different results of the ITCN quantification process of Lobule 3, 4, 5 and 6, respectively. This quantification process is well documented which helps if it is necessary to replicate work or apply this approach in similar situations. Image Number of Cells 1 25 2 25 3 7 4 16 5 12 Table 41: CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 3. 88 6.1. CLASSIC METHODS FOR AUTOMATIC CELL COUNTING Image Number of Cells 1 9 2 14 3 10 4 13 5 9 Table 42: CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 4. Image Number of Cells 1 8 2 11 3 20 4 9 5 14 Table 43: CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 5. Image Number of Cells 1 13 2 8 3 24 4 2 5 10 Table 44: CN276 2FD - Slice 1 - ITCN - Automatic Quantification Results for Lobule 6. To conclude, the approach implemented with the ITCN plugin proved that it is possible to automate the cell counting process, namely the process of counting microglial cells. Once again, it doesn’t make sense to highlight only one experience, but the whole counting process. As explained, the brain images of microglial cells are quite complex which in itself does not help in the counting process. As had already been verified with the experiments carried out with the generic cell sample, pre-quantification procedures are very important. The more rigorous they are, the more effective the counting process will be. This was no exception in the case of microglial cells. Another very relevant aspect of the quantification of cells based on this approach was the definition of cell size and the minimum distance between cells. It should be noted that not defining these two parameters correctly, the counting process is not efficient because it will either count more or fewer cells. Knowing that microglial cells are non-uniform and have different sizes and lengths from each other, sometimes it was necessary to reset cell size and the minimum distance between cells several times, which made the counting process more time-consuming. The value assigned to the ITCN threshold is again important as it helps to deal with image noise and outliers. Thus, in some experiences, to achieve results close to the actual number of cells in an image was is necessary to conduct several tests. That said, the ITCN plugin undoubtedly was able to automate the microglial cell quantification, which is one of the objectives of this dissertation. However, when we compare its 89 CHAPTER 6. RESULTS AND DISCUSSION performance with the developed Analyze Particles protocol, we realize that the ITCN approach is not the ideal solution since sometimes takes longer to achieve proper results. Perhaps in situations where the cells are all of a similar size and are evenly distributed across the image, the ITCN plugin may be a better solution to the problem when compared to the developed protocol. Through the analysis of the results and the experiments carried out, it can be concluded that the ITCN works better in situations that are not too complex, which is not the case with microglial cells. Therefore, this fully documented protocol will not be brought for the discussion of the different methodologies for cell quantification, manual, classical and deep learning, as the previous approach evidence better results. 6.2 Deep Learning for Automatic Cell Counting Based on CNNs, which is the model of choice for automatic cell quantification with a deep learningbased methodology, several experiments were performed, as was detailed in the previous chapter. Once again, tests were conducted with the generic cell sample to understand the possible limitations of the model, as well as to study the pros and cons of the application thereof. So, through the results obtained, a discussion is presented in subsection 6.2.1. The outcomes of the model made it possible to conclude that this approach can be applied in the case of microglial cells. Subsection 6.2.2 details the results obtained with microglial cells, as well as, presents a discussion, and bases the procedure considered optimal. The chosen procedure will be brought to the comparison between the different approaches, the manual, the classic and the deep learning approach. 6.2.1 Cell Colony Sample Similar to the classical approach, this generic cell sample was selected to understand the limitations of the CNNs model. However, unlike the classical methodology, in the case of deep learning approaches, the cell sample was divided into 16 equal parts to build a dataset composed of 16 images and thus obtain more adequate and accurate results. Next, the results gathered from the experiments performed are presented. For each strategy, results are discussed, followed by the conclusions drawn that will positively influence the microglial cell counting process. Classification per Number of Cells One of the approaches for automating the counting process based on a deep learning methodology is image classification based on the number of cells. The classification model is composed of four parameters, namely the bath size, the number of epochs, the learning rate and whether or not to apply data augmentation. These parameters can take different values so it was necessary to tune the model as explained in subsection 5.3.1. The tuning performed on the model resulted in 330 experiments. Table 45 presents the top 15 of the best results obtained. In addition to the results, the values that the various parameters took to arrive at the presented results are also presented. As explained, to arrive at the score 90 6.2. DEEP LEARNING FOR AUTOMATIC CELL COUNTING value presented, cross-validation was performed with 16 folds and the final error is the average of the 16 folds. In this way, it was possible to maximize the training data since in each fold we train with 15 images and test only with 1 image. batch_size epochs learning_rate data_augmentation score loss run_time str(score_list) 8 25 0.0005 false 0.9375 0.1678 20.3901 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 16 30 0.0005 false 0.9375 0.1683 22.2191 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 4 10 0.0010 false 0.9375 0.1706 11.5472 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 15 0.0010 false 0.9375 0.1725 13.9473 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 4 35 0.0010 false 0.9375 0.1882 33.9451 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 2 10 0.0005 false 0.9375 0.1936 14.4812 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 16 20 0.0010 false 0.9375 0.1991 16.4319 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 4 15 0.0005 false 0.9375 0.2024 15.2823 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 2 10 0.0010 false 0.9375 0.2219 14.4046 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 16 20 0.0005 false 0.9375 0.2269 15.7107 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 4 10 0.0005 false 0.9375 0.2273 11.4160 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 1 10 0.0010 false 0.9375 0.2279 17.4996 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 2 30 0.0010 false 0.9375 0.5378 35.5997 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 2 35 0.0010 false 0.9375 0.5511 38.9295 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 30 0.0010 true 0.8750 0.4016 35.0041 [0.0, 1.0, 0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] Table 45: Generic Cell Sample - Deep Learning Classification Results (Number). Turning now to a more detailed analysis, throughout the visualization of the results is clear that in none of the experiments the model was able to correctly classify the first image. The reason could be related to the image in question, as it contains 38 cells. Therefore is labelled as ”Few”. The fact that it has 38 cells puts it very close to the next class since it only needed to have 40 cells to be labelled as ”Average”. Thus, this image turns out to be a bit complicated for a first prediction of the model and consequently helps to justify the result obtained in the classification. It is important to point out that to know the result of the prediction of a certain image, an order was defined. Another relevant aspect to highlight is that in these top-of-the-best experiences, the model managed to correctly classify 15 of the 16 images, which leads to the expectation that it will have an adequate performance when applied to microglial cells. Maybe related to the fact that the dataset is composed of 16 images, data augmentation was not used in any of the top experiments. To conclude, of all the experiments carried out the first four should be highlighted, as they had a loss value of less than 0.18 and a run time of fewer than twenty-three seconds. The results are excellent as the model just didn’t get it right in predicting one image. In common among these four best experiments we 91 CHAPTER 6. RESULTS AND DISCUSSION have that the value of the learning rate was 0.0005 or 0.0010. The first value is the most complex value that the model can take. This leads to the conclusion that the model will behave well in more complex situations as will be the case of microglial cells. To attest to this, the second-best experience took the most complex values possible (except for the application of data augmentation) and still managed to be one of the best results. That said, the classification approach based on the number of cells in an image brings many benefits to the automatic quantification of cells, as it makes a process that is usually time-consuming simpler, faster and more effective. Within a deep learning-based methodology, it is a reliable solution to the problem presented and therefore it will be used with microglial cells. Classification per Area of Cells Another approach for automating the counting process based on a deep learning methodology is image classification based on the percentage of area that the cells take in an image. Similar to the previous approach, the classification model is composed of four parameters, namely the bath size, the number of epochs, the learning rate and whether or not to apply data augmentation. Once again these parameters can take different values so it was necessary to tune the model. The entire tuning process is explained in subsection 5.3.1. The tuning performed on the model resulted in four 480 experiments. Table 46 presents the top 15 of the best results obtained. Parallel to the previous approach, cross-validation was performed with 16 folds and the final error is the average of the 16 folds. By viewing the results, we can prove that the model behaved very well by correctly classifying 15 of the 16 images. Once again is clear that in none of the experiments the model was able to correctly classify the first image. The reason for this could also be related to the image in question, as the percentage of the area that the cells occupied in the image are 5.159%. Therefore is labelled as ”Few”. To be classified as ”Average” the percentage of the area needed to be 5.2%. Given the proximity of values, once again we conclude that the image in question turns out to be a bit complicated for a first prediction of the model and consequently helps to justify the result obtained in the classification. Through the results obtained in the learning curves, in this case, it was relevant to go up to 95 epochs. In this way, the complexity of the model increases considerably, so the application of data augmentation becomes relevant, having even been applied in two of the fifteen best experiments. Overall, the results lead to believing that the classification approach based on the percentage of area that the cells take in an image will have an acceptable performance when applied to microglial cells. To conclude, highlight the two best experiments, which despite having taken more complex values for the model parameters managed to have the lowest loss value, lower than 0.17. Despite the high run time, results are excellent as the model just didn’t get it right in predicting one image. In the remaining thirteen cases, the model also failed to predict one image. In common these two approaches have the same value for the batch size, which is eight, the same number of epochs, which is seventy-five, and both used data augmentation. Regarding the value of the learning rate parameter, they used the most complex ones for the model. This leads to the judgment that the model will behave reasonably in more complex situations as will be the case of microglial cells. To testify to this, the best experience took data augmentation, and 92 6.2. DEEP LEARNING FOR AUTOMATIC CELL COUNTING batch_size epochs learning_rate data_augmentation score loss run_time str(score_list) 8 75 0.0005 true 0.9375 0.1593 85.7092 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 75 0.0010 true 0.9375 0.1612 85.7127 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 65 0.0005 false 0.9375 0.5792 47.4480 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 70 0.0005 false 0.9375 0.5959 54.7171 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 65 0.0010 false 0.9375 0.7511 54.4616 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 60 0.0010 false 0.9375 0.7541 49.5269 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 70 0.0010 false 0.9375 0.8202 56.6123 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 75 0.0010 false 0.9375 0.8563 57.2924 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 1 65 0.0005 false 0.9375 0.8674 106.3055 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 75 0.0005 false 0.9375 0.9085 61.6134 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 80 0.0005 false 0.9375 0.9547 65.1747 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 8 90 0.0005 false 0.9375 0.9570 64.7027 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 4 60 0.0005 false 0.9375 1.0262 61.0437 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 1 70 0.0005 false 0.9375 1.1333 116.4780 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] 16 95 0.0010 false 0.9375 1.6979 65.3831 [0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] Table 46: Generic Cell Sample - Deep Learning Classification Results (Area). the value of the learning rate was the most complex. This model also took a high value for the number of epochs and still managed the best results despite the somewhat high run time. That said, the classification approach based on the percentage of area that the cells take in an image fetches many advantages to the automatic quantification of cells, as it makes the process faster and more effective. Within a deep learning-based methodology, this is another reliable solution to the problem presented and therefore it will be used with microglial cells. 6.2.2 Brain Image Samples Taking advantage of the division that was performed on the images within the classical approach, namely due to the fact that the microglial cells and the marker associated with them are only visible in the red channel, the same images were reused, thus composing a dataset with 661 images. With this number of images, we are guaranteed to obtain more adequate and exact results with the model predictions. Next, the results collected from the experiments performed are given. For each strategy, results are discussed, followed by the conclusions drawn that will positively influence an approach considered optimal for the automation of the microglial cell counting process. 93 CHAPTER 6. RESULTS AND DISCUSSION Classification per Number of Cells As it was possible to prove, not only through relevant literature and documentation but also through the results obtained with the generic cell sample-related experiments, an approach to automate the microglial cell quantification based on a deep learning methodology is image classification based on the number of cells. Taking advantage of the model that obtained excellent results with the cell colony sample, the model for microglia cells also has four fundamental parameters. These parameters are once again the bath size, the number of epochs, the learning rate and whether or not to apply data augmentation. As expected these parameters took different values so it was necessary to tune the model as explained in subsection 5.3.2. The tuning performed on the model resulted in 288 experiments. Table 47 presents the top 15 of the best results obtained. As can be visualized, in addition to the classification results, the values that the various parameters took to arrive at these results are also presented. As explained, to arrive at the score value presented, cross-validation was performed with 5 folds and the final error is the average of the 5 folds. In this way, it was possible to maximize the training data since in each fold we train with 529 samples and test with 132 images. batch_size epochs learning_rate data_augmentation score loss run_time str(score_list) 32 20 0.0005 false 0.9021 0.2479 326.4157 [0.5714, 0.9393, 1.0000, 1.0000, 1.0000] 32 20 0.0010 false 0.9006 0.2765 260.9422 [0.5789, 0.9318, 0.9924, 1.0000, 1.0000] 16 20 0.0005 false 0.8690 0.3291 333.8006 [0.4586, 0.8863, 1.0000, 1.0000, 1.0000] 32 15 0.0010 false 0.8672 0.3326 202.2335 [0.5864, 0.7500, 1.0000, 1.0000, 1.0000] 16 15 0.0005 false 0.8520 0.4301 293.1848 [0.6466, 0.6136, 1.0000, 1.0000, 1.0000] 32 10 0.0005 false 0.8489 0.3904 150.9446 [0.6691, 0.6363, 0.9469, 0.9924, 1.0000] 32 15 0.0005 false 0.8475 0.4282 206.0727 [0.5939, 0.6590, 0.9848, 1.0000, 1.0000] 32 10 0.0010 false 0.8475 0.3738 179.2750 [0.6240, 0.6363, 0.9848, 1.0000, 0.9924] 16 10 0.0005 false 0.7959 0.5477 156.7259 [0.6691, 0.5454, 0.7651, 1.0000, 1.0000] 8 15 0.0005 false 0.7944 0.5796 266.0632 [0.6691, 0.5530, 0.7575, 0.9924, 1.0000] 16 15 0.0010 false 0.7595 0.8250 209.3782 [0.6691, 0.8409, 0.3560, 0.9393, 0.9924] 32 5 0.0010 false 0.7462 0.6996 76.5467 [0.6090, 0.8181, 0.5151, 0.8409, 0.9469] 16 5 0.0005 false 0.6611 0.8956 93.8852 [0.6691, 0.8030, 0.3636, 0.6212, 0.8484] 32 5 0.0005 false 0.6520 0.8672 84.6684 [0.6691, 0.8333, 0.3484, 0.5378, 0.8712] 32 4 0.0005 false 0.5989 0.9619 86.9296 [0.6691, 0.8409, 0.3181, 0.4318, 0.7348] Table 47: Microglial Cells - Deep Learning Classification Results (Number). Looking in more detail at the results, the first eight experiments should be highlighted, since the score obtained is above 84%. In at least one of the folds, the 132 images were correctly classified, which attests to the solidity of the model for the presented problem. However, the problem with the model was the classification made in the first fold, where the results were sometimes not the best. It is important to note that the brain images of microglial cells are extremely complex, since the cells have a complex morphology, vary in size and length, and some have longer ramifications than others. This makes the model predictions more difficult. Despite this, in the second fold, we see substantial improvements, which leads to the conclusion that the model is capable of learning from errors and thus obtaining and making better predictions. Nevertheless, the loss obtained in these eight best experiments is less than 0.44 which is a very good benchmark for the model. Another relevant aspect to highlight is that in these top-of-thebest experiences, the model didn’t use data augmentation despite being composed of a high number of samples. Finally, highlight again the eight best experiments, as they took some of the most complex values 94 6.2. DEEP LEARNING FOR AUTOMATIC CELL COUNTING for the model parameters, such as higher values for the bath size, number of epochs, and the learning rate, and still managed to obtain the best ratings. To conclude and in this way support an optimal approach to the problem presented in this dissertation, we highlight the first and best experience. For the reasons presented above, we know that the samples that make up our dataset (images of microglial cells) are quite complex. Then, with the tuning done to the model, we bring more complexity to it, given the values that the different parameters can take. A person not familiarised with the subject would say that great results would not be expected given the difficulty we are causing to the model, through extremely complex images and various values assigned to parameters. The results obtained contradict this premise. As if that were not enough, in addition to the already mentioned difficulty caused by the complexity of the images, the values of the parameters that gave rise to the best result were the most complex possible, except for the data augmentation, which was not used. The model took the value of 32 for the batch size with 20 epochs and 0.0005 for the learning rate. With all this difficulty, the model was able to correctly classify more than 90% of the cases, which means it correctly classified 596 images out of 661. It should be noted that the classifications made in the last three folds contributed to this, where the percentage of success was 100%. This gives 396 correctly classified images. Even the results obtained in the first two folds are very positive, with the percentage of hits exceeding 57% and 93%, respectively. As if the score wasn’t enough, the model was able to classify all 661 images in 5 minutes and 44 seconds. That said, the classification approach based on the number of cells in an image brings so many benefits to the automatic quantification of cells, as it makes a process simpler, much faster, more effective and replicable. Within a deep learning-based methodology, it is a very reliable solution to the problem presented. Classification per Area of Cells Another approach to automate the microglial cell counting process fetched from relevant literature and through the results obtained with the generic cell sample-related experiments regarding a deep learning methodology is image classification based on the percentage of area that the cells take in an image. Once again taking advantage of the model that obtained excellent results with the cell colony sample, the model for microglia cells also has the same parameters. As the parameters can take different values was necessary to tune the model. The tuning process is described in subsection 5.3.1. The tuning performed on the model resulted in 324 experiments. Table 48 presents the top 15 of the best results obtained. Parallel to the last procedure, cross-validation was performed with 5 folds and the final error is the average of the 5 folds. Contrary to the previous approach, but as expected from the results obtained in the classification by the percentage of area that cells occupy in an image with the generic cell sample, the results were not so positive. Despite this, the seven best experiences obtained a score percentage of more than 74%. This means that in these experiments at least 493 images were correctly classified. This attests to the solidity of the approach for the presented problem. Once again, and to explain why these results were obtained, it is worth noting that the images of microglial cells are complex, and in the case of a classification approach 95 CHAPTER 6. RESULTS AND DISCUSSION batch_size epochs learning_rate data_augmentation score loss run_time str(score_list) 16 25 0.0005 false 0.8312 0.5005 186.4194 [0.3533, 0.8030, 1.0000, 1.0000, 1.0000] 16 20 0.0005 false 0.7827 0.7110 145.2895 [0.3533, 0.5606, 1.0000, 1.0000, 1.0000] 32 25 0.0005 false 0.7721 0.6995 181.540 [0.3759, 0.5000, 0.9848, 1.0000, 1.0000] 8 25 0.0005 false 0.7525 0.7375 250.4927 [0.3383, 0.4318, 0.9924, 1.0000, 1.0000] 32 20 0.0005 false 0.7479 0.5524 146.7203 [0.3383 0.4848, 0.9166, 1.0000, 1.0000] 8 20 0.0005 false 0.7479 0.8725 206.0794 [0.3383, 0.4242, 0.9772, 1.0000, 1.0000] 16 15 0.0005 false 0.7464 0.6019 129.4396 [0.3383, 0.4848, 0.9166, 0.9924, 1.0000] 32 15 0.0005 false 0.6919 0.6063 100.8088 [0.3383, 0.4242, 0.7121, 0.9848, 1.0000] 16 15 0.0010 false 0.6706 0.7802 109.5244 [0.3383, 0.4469, 0.5757, 0.9924, 1.0000] 32 25 0.0010 false 0.6510 0.6611 172.7800 [0.3383, 0.4393, 0.5757, 0.9090, 0.9924] 16 10 0.0005 false 0.6449 0.7910 80.5878 [0.3383, 0.4469, 0.4924, 0.9469, 1.0000] 32 20 0.0010 false 0.6313 0.7594 162.2263 [0.3383, 0.4318, 0.4924, 0.9015, 0.9924] 8 15 0.0005 false 0.6222 0.8347 151.2869 [0.3383, 0.4469, 0.4545, 0.8712, 1.0000] 32 10 0.0005 false 0.5903 0.7720 67.7552 [0.3383, 0.4318, 0.5303, 0.7803, 0.8712] 16 10 0.0010 false 0.5706 0.8302 80.6591 [0.3383, 0.4469, 0.4545, 0.6287, 0.9848] Table 48: Microglial Cells - Deep Learning Classification Results (Area). based on the percentage of area that cells occupy in an image it is even worse since the cells vary in size and length. Consequently, this makes the model predictions more difficult. The results obtained in the first fold are not the best, but they improve from fold to fold. This proves that despite everything the model is capable of learning from errors and accordingly making better predictions. Anyway, the loss obtained in these seven best experiments is less than 0.88 which is also a very good benchmark for the model. In all these top-of-the-best experiences, the model didn’t use data augmentation which leads to the idea that, in addition to the already extremely intricate images, the use of data augmentation would only worsen the results obtained, since it would bring an additional difficulty to the model. Finally, these experiments took some of the most complex values for the model parameters, such as 32 for the batch size, 25 for the number of epochs, and 0.0005 for the learning rate value, and still managed to obtain decent results. To support an optimal approach to the problem presented in this dissertation, we highlight the first and best experience. For the reasons presented above, we know that the samples that make up our dataset are complex. Then, with the tuning done to the model, we bring more complexity to the procedure. With all this difficulty, the model was able to correctly classify more than 83% of the cases, which means it correctly classified 548 images out of 661. The classifications made in the last three folds contributed to this, where the percentage of success was 100%. This gives 396 correctly classified images. Looking at the value obtained in the loss, which, despite not being bad, is not ideal. The classification time took 3 minutes and 10 seconds. That said, the classification approach based on the area that cells occupy in an image brings benefits to the automatic quantification of cells, as it makes a process simpler, much faster and replicable. Nevertheless, this approach will not be brought for the discussion of the different methodologies for cell quantification, as the previous procedure presents better results to automate the process of microglial cell count. 96 7 Conclusion The aim of this dissertation was the automatic quantification of microglial cells from brain images with classic and deep learning methods. Our contributions were described in the previous chapters, following two main topics: classic and deep learning methods for automatic cell counting. For each approach, we have presented and described the developed procedure and the obtained results as well as presented a discussion of those results. Therefore, in this last chapter, we sum up to the reader the main contributions and conclusions. Finally, we complete this document with our perspectives for future work that may deserve further investigation. 7.1 General Conclusions Microglia are a type of neuronal cell located throughout the brain’s spinal cord. Bearing in mind the importance of these cells and knowing how they are counted, which requires the segmentation of several images, a task usually performed manually. Therefore, we focused the work in this dissertation on studying the different methodologies that help automate microglial cell quantification. As this is a fundamental procedure that in some cases may help in the detection of an illness, the ultimate objective is enabling the development of automated computerized solutions. The study carried out is very important as automatic cell quantification approaches are becoming increasingly important to reduce the workload of specialists and provide robust and reproducible results. As stated, nowadays, most of the cell counting processes are done manually. Conventional cell counting involves a specific set of tools and devices developed for that purpose. This process is tedious, timeconsuming, and inaccurate due to operator-dependent biases. As cell counting is an important procedure routine, various study reports are focusing on the experience of the development of image processing programs and techniques to automate cell counting. Consequently, this makes the cell quantification process more time-efficient and reduces error. To automate the cell counting process emerge classic and deep learning methodologies. Within the so-called classic methodology, we have software and assistants, like ImageJ, that automate the quantification process. Needs to be pointed out that some of these programs, developed to automate the cell 97 BIBLIOGRAPHY [38] W. Xie, J. A. Noble, and A. Zisserman. “Microscopy cell counting and detection with fully convolutional regression networks”. In: Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization 6.3 (2018), pp. 283–292. doi: 10.1080/21681163.2016.1149104 (cit. on pp. 27,28). [39] Y. Xue et al. “Cell Counting by Regression Using Convolutional Neural Network”. In: Computer Vision – ECCV 2016 Workshops . Cham: Springer International Publishing, 2016, pp. 274–290. isbn: 9783-319-46604-0. doi: https://doi.org/10.1007/978-3-319-46604-0_20 (cit. on p. 28). [40] W. Ertel. Introduction to artificial intelligence . Springer, 2018, pp. 1–3 (cit. on p. 28). [41] J. M. Górriz et al. “Artificial intelligence within the interplay between natural and artificial computation: Advances in data science, trends and applications”. In: Neurocomputing 410 (2020), pp. 237– 270. issn: 0925-2312. doi: https://doi.org/10.1016/j.neucom.2020.05.078 (cit. on p. 28). [42] D. Carneiro et al. “Online dispute resolution: an artificial intelligence perspective”. In: Artificial Intelligence Review 41.2 (2014), pp. 211–240. doi: https://doi.org/10.1007/s10462-0 11-9305-z (cit. on p. 28). [43] B. Fernandes, P. Novais, and C. Analide. “A Multi-Agent System for Automated Machine Learning”. In: Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems . 2022, pp. 1899–1901 (cit. on p. 28). [44] J. P. Mueller and L. Massaron. Machine learning for dummies . John Wiley & Sons, 2021, pp. 4–18 (cit. on pp. 28,29,34,35). [45] I. Goodfellow, Y. Bengio, and A. Courville. Deep learning . MIT press, 2016, pp. 96–119 (cit. on p. 28). [46] B. Fernandes et al. “Long short-term memory networks for traffic flow forecasting: exploring input variables, time frames and multi-step approaches”. In: Informatica 31.4 (2020), pp. 723–749. doi: 10.15388/20-INFOR431 (cit. on p. 29). [47] S. Singh. Cousins of Artificial Intelligence . url: https://towardsdatascience.com/cousinsof-artificial-intelligence-dda4edc27b55 (visited on 12/19/2021) (cit. on p. 29). [48] P. Oliveira et al. “Forecasting Energy Consumption of Wastewater Treatment Plants with a Transfer Learning Approach for Sustainable Cities”. In: Electronics 10.10 (2021), p. 1149. issn: 2079-9292. doi: 10.3390/electronics10101149 (cit. on p. 29). [49] N. Buduma and N. Locascio. Fundamentals of deep learning: Designing next-generation machine intelligence algorithms . O’Reilly Media, Inc., 2017, pp. 1–39. isbn: 978-1-491-92561-4 (cit. on pp. 29, 31–33). 104 BIBLIOGRAPHY [50] A. Karpathy and J. Johnson. Convolutional neural networks for visual recognition . url: https: //cs231n.github.io/ (visited on 12/28/2021) (cit. on pp. 29–34). [51] T. Dettmers. Deep Learning in a Nutshell: Core Concepts . url: https://developer.nvidia. com/blog/deep-learning-nutshell-core-concepts/ (visited on 01/07/2022) (cit. on pp. 32,33). [52] U. Karn. An intuitive explanation of convolutional neural networks . url: https://ujjwalkarn. me/2016/08/11/intuitive-explanation-convnets/ (visited on 01/09/2022) (cit. on p. 33). [53] M. Peemen et al. “The neuro vector engine: Flexibility to improve convolutional net efficiency for wearable vision”. In: 2016 Design, Automation Test in Europe Conference Exhibition (DATE) . 2016, pp. 1604–1609. isbn: 978-3-9815-3707-9 (cit. on pp. 33,34). [54] A. Géron. Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow: Concepts, tools, and techniques to build intelligent systems . O’Reilly Media, Inc., 2019, pp. 7–14. isbn: 978-1-49203264-9 (cit. on p. 34). [55] D. Parikh. Learning Paradigms in Machine Learning . url: https://medium.datadriveninvestor. com/learning-paradigms-in-machine-learning-146ebf8b5943 (visited on 11/23/2021) (cit. on pp. 34,35). [56] Z. Pan et al. “Recent Progress on Generative Adversarial Networks (GANs): A Survey”. In: SPSS inc 7 (2019), pp. 36322–36333. doi: 10.1109/ACCESS.2019.2905015 (cit. on pp. 36,37). [57] E. Kan. What The Heck Are VAE-GANs? url: https://towardsdatascience.com/whatthe-heck-are-vae-gans-17b86023588a (visited on 11/30/2021) (cit. on pp. 36,37). [58] H. Ahmady Phoulady et al. “Automatic ground truth for deep learning stereology of immunostained neurons and microglia in mouse neocortex”. In: Journal of Chemical Neuroanatomy 98 (2019), pp. 1–7. issn: 0891-0618. doi: https://doi.org/10.1016/j.jchemneu.2019.02.006 (cit. on p. 37). [59] P. Dave et al. “An adaptive digital stain separation method for deep learning-based automatic cell profile counts”. In: Journal of Neuroscience Methods 354 (2021), p. 109102. issn: 0165-0270. doi: https://doi.org/10.1016/j.jneumeth.2021.109102 (cit. on p. 38). [60] A. B. Campos et al. “Profiling Microglia in a Mouse Model of Machado-Joseph Disease”. In: Biomedicines 10.2 (2022). issn: 2227-9059. doi: 10.3390/biomedicines10020237 (cit. on pp. 47,48). [61] “Implications of functional anatomy on information processing in the deep cerebellar nuclei”. In: Frontiers in Cellular Neuroscience 3 (2009). issn: 1662-5102. doi: 10.3389/neuro.03.014.2 009 (cit. on p. 47). Thisdocumentwascreatedusingthe(pdf/Xe/Lua)L A T EXprocessor,basedontheNOVAthesistemplate,developedattheDep.InformáticaofFCT-NOVAbyJoãoM.Lourenço.[1] [1] J.M.Lourenço.TheNOVAthesisL A T E XTemplateUser’sManual.NOVAUniversityLisbon.2021.URL:https://github.com/joaomlourenco/novathesis/raw/master/template.pdf(cit.onp.105). 105 A CN276 2FD Quantification Settings This appendix displays the quantification settings applied to the brain image of the animal CN276 2FD, allowing us to achieve the results obtained in microglia cell counts with ImageJ regarding a more classical approach. 1𝑠𝑡 2𝑛𝑑 3𝑟𝑑 4𝑡ℎ 5𝑡ℎ 6𝑡ℎ 7𝑡ℎ 8𝑡ℎ Threshold Min 0 0 0 0 0 0 0 0 Max 3694 2168 2393 2569 2425 2794 2650 2425 96.29% 95.58% 95.80% 97.14% 96.89% 96.39% 95.97% 96.51% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 49: CN276 2FD - Slice 1 - Quantification Settings for DCN. 106 1𝑠𝑡 2𝑛𝑑 3𝑟𝑑 4𝑡ℎ Threshold Min 0 0 0 0 Max 2746 3212 2875 2248 96.32% 96.35% 95.63% 95.07% Algorithm Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 50: CN276 2FD - Slice 1 - Quantification Settings for Lobule 7. 1𝑠𝑡 2𝑛𝑑 3𝑟𝑑 4𝑡ℎ 5𝑡ℎ 6𝑡ℎ 7𝑡ℎ Threshold Min 0 0 0 0 0 0 0 Max 3003 2682 2618 2425 2232 1847 1718 96.14% 96.11% 95.73% 95.88% 94.77% 95.87% 94.78% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 51: CN276 2FD - Slice 1 - Quantification Settings for Lobule 8. 1 2 3 4 5 6 7 8 9 Threshold Min 0 0 0 0 0 0 0 0 0 Max 1975 2618 1734 1799 2168 1959 2489 2216 2088 95.60% 95.22% 94.31% 95.02% 94.63% 95.33% 94.88% 94.68% 94.91% Algorithm Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 52: CN276 2FD - Slice 1 - Quantification Settings for Lobule 9. 107 APPENDIX A. CN276 2FD QUANTIFICATION SETTINGS 1 2 3 4 5 6 7 8 9 Threshold Min 0 0 0 0 0 0 0 0 0 Max 3404 3726 1815 2923 2409 2746 3115 3228 1927 96.43% 95.60% 94.91% 96.15% 95.29% 96.86% 97.41% 97.41% 94.90% Algorithm Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 53: CN276 2FD - Slice 1 - Quantification Settings for Lobule 10. 12345 Threshold Min 0 0 0 0 0 Max 2521 2473 3501 3453 3131 95.08% 95.66% 94.49% 93.44% 94.30% Algorithm Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 54: CN276 2FD - Slice 1 - Quantification Settings for Lobule 11. 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 1638 1092 1574 1847 1156 1686 1847 1028 95.26% 95.22% 94.02% 93.68% 92.83% 91.90% 94.43% 90.87% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 55: CN276 2FD - Slice 2 - Quantification Settings for DCN. 108 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 3437 2746 3019 2521 1750 1831 2826 2650 96.98% 95.91% 94.87% 94.11% 94.35% 97.27% 95.87% 97.34% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 56: CN276 2FD - Slice 2 - Quantification Settings for Lobule 2. 12345 Threshold Min 0 0 0 0 0 Max 2088 2007 2312 2682 2971 94.15% 94.13% 96.21% 93.54% 94.93% Algorithm Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 57: CN276 2FD - Slice 2 - Quantification Settings for Lobule 3. 1 2 3 4 5 6 Threshold Min 0 0 0 0 0 0 Max 2505 2810 2569 2312 2393 3292 93.61% 95.43% 95.45% 95.84% 96.10% 96.05% Algorithm Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 58: CN276 2FD - Slice 2 - Quantification Settings for Lobule 4. 109 APPENDIX A. CN276 2FD QUANTIFICATION SETTINGS 123 Threshold Min 0 0 0 Max 1670 1269 1702 94.02% 93.61% 94.23% Algorithm Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 Table 59: CN276 2FD - Slice 2 - Quantification Settings for Lobule 5. 1 2 3 4 5 6 Threshold Min 0 0 0 0 0 0 Max 1879 2120 1590 2023 2377 2602 93.92% 95.27% 94.21% 95.37% 95.97% 95.85% Algorithm Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 60: CN276 2FD - Slice 2 - Quantification Settings for Lobule 6. 1 2 3 4 5 Threshold Min 0 0 0 0 0 Max 1493 996 1237 1429 1381 95.25% 94.05% 96.21% 95.38% 95.62% Algorithm Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 61: CN276 2FD - Slice 2 - Quantification Settings for Lobule 7. 1234567 Threshold Min 0 0 0 0 0 0 0 Max 1686 2264 1911 1815 1991 1574 1188 95.50% 96.94% 96.05% 96.15% 96.54% 95.21% 95.24% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 62: CN276 2FD - Slice 2 - Quantification Settings for Lobule 8. 1234567 Threshold Min 0 0 0 0 0 0 0 Max 1847 2361 1975 1526 1638 1558 1815 95.32% 95.72% 95.18% 95.79% 95.03% 94.35% 93.65% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 63: CN276 2FD - Slice 2 - Quantification Settings for Lobule 9. 110 1234567 Threshold Min 0 0 0 0 0 0 0 Max 1028 1558 1542 1702 1542 1606 1574 97.37% 94.30% 95.29% 93.51% 94.00% 94.48% 95.95% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 64: CN276 2FD - Slice 2 - Quantification Settings for Lobule 10. 1 2 3 4 5 6 7 8 9 Threshold Min 0 0 0 0 0 0 0 0 0 Max 1237 1253 1269 530 658 1172 1108 1140 835 97.71% 97.02% 97.46% 96.80% 95.05% 94.86% 93.07% 96.42% 93.39% Algorithm Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 65: CN276 2FD - Slice 2 - Quantification Settings for Lobule 11. 1 2 3 4 5 6 Threshold Min 0 0 0 0 0 0 Max 1349 1574 1461 1718 2184 1445 93.09% 92.66% 95.08% 96.06% 95.51% 94.07% Algorithm Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 66: CN276 2FD - Slice 2 - Quantification Settings for Lobule 12. 111 APPENDIX A. CN276 2FD QUANTIFICATION SETTINGS 1 2 3 4 5 Threshold Min 0 0 0 0 0 Max 1204 1670 1975 1734 1645 95.12% 96.20% 96.24% 96.35% 95.97% Algorithm Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 67: CN276 2FD - Slice 2 - Quantification Settings for Lobule 13. 1234567 Threshold Min 0 0 0 0 0 0 0 Max 1879 1526 1766 1076 1959 2232 1911 97.03% 96.20% 97.11% 98.72% 94.09% 95.40% 96.10% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 68: CN276 2FD - Slice 2 - Quantification Settings for Lobule 14. 1 2 3 4 5 Threshold Min 0 0 0 0 0 Max 2104 2120 1381 2361 2505 94.36% 97.08% 98.30% 96.88% 96.91% Algorithm Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 69: CN276 2FD - Slice 2 - Quantification Settings for Lobule 15. 1 2 3 4 Threshold Min 0 0 0 0 Max 2521 2296 2264 2104 96.00% 96.55% 95.25% 95.96% Algorithm Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 70: CN276 2FD - Slice 2 - Quantification Settings for Lobule 16. 112 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 1622 1429 1124 1365 1237 915 1204 1124 96.18% 97.49% 97.60% 94.00% 94.37% 96.97% 95.49% 93.74% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 71: CN276 2FD - Slice 2 - Quantification Settings for Lobule 17. 1 2 3 4 5 6 7 8 9 10 Threshold Min 0 0 0 0 0 0 0 0 0 0 Max 434 1220 1237 1188 674 418 947 1012 1028 771 97.19% 96.57% 96.21% 97.13% 97.01% 96.92% 96.64% 94.08% 95.23% 95.78% Algorithm Default Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 72: CN276 2FD - Slice 2 - Quantification Settings for Lobule 18. 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 1365 1590 867 1028 867 1012 626 803 94.00% 94.18% 92.58% 92.91% 92.34% 94.84% 92.46% 90.27% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 73: CN276 2FD - Slice 3 - Quantification Settings for DCN. 113 APPENDIX B. CN282 2TE QUANTIFICATION SETTINGS 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 2505 2521 2361 2473 1783 1991 1959 2007 97.65% 97.65% 97.20% 97.65% 97.65% 97.15% 97.60% 96.90% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 87: CN282 2TE - Slice 1 - Quantification Settings for Lobule 2. 12345 Threshold Min 0 0 0 0 0 Max 1927 2216 2345 1927 3003 98.05% 97.30% 97.20% 97.20% 97.50% Algorithm Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 88: CN282 2TE - Slice 1 - Quantification Settings for Lobule 3. 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 2088 2312 2505 2056 2023 1911 2007 2128 98.00% 97.60% 97.84% 97.65% 97.89% 97.90% 97.81% 97.81% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 89: CN282 2TE - Slice 1 - Quantification Settings for Lobule 4. 120 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 1686 2152 2312 1959 2312 1927 1590 2425 98.09% 97.95% 98.17% 97.70% 98.23% 98.15% 97.68% 98.17% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 90: CN282 2TE - Slice 1 - Quantification Settings for Lobule 5. 1 2 3 4 5 6 Threshold Min 0 0 0 0 0 0 Max 1959 2013 2184 1863 2056 2007 97.70% 97.81% 98.07% 98.09% 97.72% 97.47% Algorithm Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 91: CN282 2TE - Slice 1 - Quantification Settings for Lobule 6. 1 2 3 Threshold Min 0 0 0 Max 2007 1863 1606 97.33% 97.93% 96.70% Algorithm Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 Table 92: CN282 2TE - Slice 1 - Quantification Settings for Lobule 7. 121 APPENDIX B. CN282 2TE QUANTIFICATION SETTINGS 1 2 3 Threshold Min 0 0 0 Max 2505 2232 2007 98.35% 97.99% 97.69% Algorithm Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 Table 93: CN282 2TE - Slice 1 - Quantification Settings for Lobule 8. 1234567 Threshold Min 0 0 0 0 0 0 0 Max 2826 2778 3196 2826 2939 3083 1429 96.65% 97.01% 97.21% 98.02% 97.23% 97.84% 97.87% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 94: CN282 2TE - Slice 2 - Quantification Settings for DCN. 1 2 3 4 5 6 7 8 9 Threshold Min 0 0 0 0 0 0 0 0 0 Max 2023 1815 1606 2585 2136 1927 2393 1477 1445 97.90% 97.65% 97.52% 97.86% 98.46% 98.04% 98.31% 97.82% 97.58% Algorithm Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 95: CN282 2TE - Slice 2 - Quantification Settings for Lobule 2. 122 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 1349 1542 1429 1815 1911 1783 1622 1493 97.94% 98.34% 97.30% 97.32% 98.01% 97.71% 97.22% 97.34% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 96: CN282 2TE - Slice 2 - Quantification Settings for Lobule 3. 1 2 3 4 5 6 7 8 9 Threshold Min 0 0 0 0 0 0 0 0 0 Max 1301 2039 1622 1783 1590 1493 1237 1140 1285 97.77% 97.99% 97.77% 97.94% 97.58% 97.41% 98.41% 98.52% 97.85% Algorithm Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 97: CN282 2TE - Slice 2 - Quantification Settings for Lobule 4. 1 2 3 4 5 6 7 8 9 Threshold Min 0 0 0 0 0 0 0 0 0 Max 1895 1975 2023 1783 1927 1831 2088 2411 2088 98.71% 98.42% 98.07% 98.18% 97.71% 98.21% 98.15% 98.47% 98.98% Algorithm Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 98: CN282 2TE - Slice 2 - Quantification Settings for Lobule 5. 123 APPENDIX B. CN282 2TE QUANTIFICATION SETTINGS 1 2 3 4 5 6 7 8 9 Threshold Min 0 0 0 0 0 0 0 0 0 Max 1783 1718 1477 1975 1783 1574 1493 1317 1799 98.16% 97.47% 97.97% 98.18% 98.33% 97.30% 98.19% 97.86% 97.72% Algorithm Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 99: CN282 2TE - Slice 2 - Quantification Settings for Lobule 6. 12345 Threshold Min 0 0 0 0 0 Max 1622 915 1413 1911 1766 97.87% 99.38% 98.93% 98.23% 98.31% Algorithm Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 100: CN282 2TE - Slice 2 - Quantification Settings for Lobule 7. 1234567 Threshold Min 0 0 0 0 0 0 0 Max 2987 2457 3292 2746 2666 2714 2810 96.70% 96.84% 96.99% 96.15% 96.91% 95.74% 96.55% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 101: CN282 2TE - Slice 4 - Quantification Settings for DCN. 124 1 2 3 4 5 6 7 8 9 Threshold Min 0 0 0 0 0 0 0 0 0 Max 2248 2248 2312 1927 1831 2120 1893 2184 1799 97.65% 97.14% 96.96% 97.32% 98.04% 97.14% 96.37% 97.83% 96.79% Algorithm Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 102: CN282 2TE - Slice 4 - Quantification Settings for Lobule 2. 1234567 Threshold Min 0 0 0 0 0 0 0 Max 1975 2361 2409 2810 2200 2345 2007 97.98% 97.52% 97.81% 98.04% 97.98% 97.35% 97.22% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 103: CN282 2TE - Slice 4 - Quantification Settings for Lobule 3. 1 2 3 4 5 6 7 8 9 10 11 Threshold Min 0 0 0 0 0 0 0 0 0 0 0 Max 2120 2296 2698 2361 2200 2088 3035 2296 2473 2280 2425 97.94% 98.54% 97.97% 98.12% 98.27% 98.64% 98.52% 98.32% 98.42% 98.12% 97.97% Algorithm Default Default Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 104: CN282 2TE - Slice 4 - Quantification Settings for Lobule 4. 125 APPENDIX B. CN282 2TE QUANTIFICATION SETTINGS 1 2 3 4 5 6 7 8 9 10 Threshold Min 0 0 0 0 0 0 0 0 0 0 Max 2136 2264 2425 2489 2393 1975 2585 2746 2875 2666 98.32% 98.72% 98.77% 98.58% 99.48% 99.38% 98.77% 98.40% 98.67% 98.77% Algorithm Default Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 105: CN282 2TE - Slice 4 - Quantification Settings for Lobule 5. 1 2 3 4 5 6 Threshold Min 0 0 0 0 0 0 Max 2232 1429 2216 1510 2409 1622 97.96% 97.71% 97.43% 97.17% 97.78% 96.81% Algorithm Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 106: CN282 2TE - Slice 4 - Quantification Settings for Lobule 6. 1234567 Threshold Min 0 0 0 0 0 0 0 Max 2618 2280 2106 1317 1927 1638 2858 97.59% 97.35% 97.94% 97.86% 97.54% 98.17% 97.91% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 107: CN282 2TE - Slice 4 - Quantification Settings for Lobule 7. 126 123 Threshold Min 0 0 0 Max 2168 2280 1734 97.87% 97.70% 97.35% Algorithm Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 Table 108: CN282 2TE - Slice 4 - Quantification Settings for Lobule 8. 127 C CN283 2FD Quantification Settings This appendix displays the quantification settings applied to the brain image of the animal CN283 2FD, allowing us to achieve the results obtained in microglia cell counts with ImageJ regarding a more classical approach. 1 2 3 4 5 6 Threshold Min 0 0 0 0 0 0 Max 1445 1943 2200 2312 2248 2473 98.04% 98.35% 98.20% 97.95% 98.10% 98.45% Algorithm Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 109: CN283 2FD - Slice 1 - Quantification Settings for DCN. 128 1 2 3 4 5 6 7 8 9 10 Threshold Min 0 0 0 0 0 0 0 0 0 0 Max 1590 1108 1895 1718 1734 1638 1558 1349 1188 787 98.55% 98.75% 98.90% 98.25% 99.00% 99.30% 99.05% 99.25% 99.45% 98.55% Algorithm Default Default Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 110: CN283 2FD - Slice 1 - Quantification Settings for Lobule 2. 1 2 3 4 5 6 7 8 Threshold Min 0 0 0 0 0 0 0 0 Max 1493 2136 1975 1879 1477 2296 2393 3340 98.33% 98.00% 98.10% 98.00% 99.60% 98.90% 98.50% 98.95% Algorithm Default Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 111: CN283 2FD - Slice 1 - Quantification Settings for Lobule 3. 1234567 Threshold Min 0 0 0 0 0 0 0 Max 3276 2939 3356 1879 2425 1493 2875 97.95% 98.35% 98.70% 97.95% 98.95% 98.60% 98.20% Algorithm Default Default Default Default Default Default Default Analyze Size (micron^2) 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity 17-Inifinity Circularity 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 0.05-1.00 Table 112: CN283 2FD - Slice 1 - Quantification Settings for Lobule 4. 129