scieee AI-readable full text Open interactive document viewer

denoiSplit: a method for joint microscopy image splitting and unsupervised denoising

Jug, Florian

Full text

denoiSplit: a method for joint microscopy image splitting and unsupervised denoising Ashesh Ashesh and Florian Jug Fondazione Human Technopole, Viale Rita Levi-Montalcini 1, 20157 Milan, Italy [email protected], [email protected] Abstract. In this work, we present denoiSplit, a method to tackle a new analysis task, i.e. the challenge of joint semantic image splitting and unsupervised denoising. This dual approach has important applications in fluorescence microscopy, where semantic image splitting has important applications but noise does generally hinder the downstream analysis of image content. Image splitting involves dissecting an image into its distinguishable semantic structures. We show that the current state-of-the-art method for this task struggles in the presence of image noise, inadvertently also distributing the noise across the predicted outputs. The method we present here can deal with image noise by integrating an unsupervised denoising subtask. This integration results in improved semantic image unmixing, even in the presence of notable and realistic levels of imaging noise. A key innovation in denoiSplit is the use of specifically formulated noise models and the suitable adjustment of KL-divergence loss for the high-dimensional hierarchical latent space we are training. We showcase the performance of denoiSplit across multiple tasks on real-world microscopy images. Additionally, we perform qualitative and quantitative evaluations and compare the results to existing benchmarks, demonstrating the effectiveness of using denoiSplit: a single Variational Splitting Encoder-Decoder (VSE) Network using two suitable noise models to jointly perform semantic splitting and denoising. 1 Introduction Fluorescence microscopy remains a cornerstone in the exploration of cellular and sub-cellular structures, enabling scientists to visualize biological processes at a remarkable level of detail [9, 22]. However, the ability to distinguish and analyze multiple structures within a single sample requires a multiplexed imaging protocol that requires extra time and effort [22]. To address these downsides and enable for more efficient and new types of investigation, a powerful method for semantic image splitting was recently introduced [1]. Building on this previous work [1], we address a key challenge that persisted: noise in microscopy images and its adverse effect on the quality of image-splitting predictions. Recognizing the need for a method capable of handling noisy input images while maintaining the integrity of the semantic splitting task, we introduce a technique that not only builds on the strengths of µSplit [1] but also 2 Ashesh, F. Jug Fig. 1: Teaser Figure. In this work we use a variational encoder-decoder network to jointly solve an usupervised denoising and image splitting task and show that our approach outperforms existing baselines. incorporates unsupervised denoising capabilities, for example as in [15, 19–21]. Figure 1 outlines the overall approach we are proposing. Together, these ingredients lead to a new method denoiSplit. It refines the process of image decomposition, ensuring that even under high levels of pixel noises present in the entire body of available training data, the semantic integrity of the semantically split image components (the predictions) is well preserved. Additionally, denoiSplit can assess data uncertainty by sampling from the learned posterior of possible splitting solutions, followed by evaluating the inter-sample variability. In Section 4, we show how to use this possibility to predict the expected error denoiSplit makes on a given input. In summary, we believe that this work will open new avenues for the efficient and detailed analysis of complex biological samples, for example, in the context of fluorescent microscopy. 2 Related Work 2.1 Image Denoising Image denoising is a task that has a long and exciting history. Classical methods, such as Non-Local Means [5] or BM3D [6], were frequently and very successfully used before neural network based approaches have been introduced towards the end of the last decade [14,25–27]. The advent of deep learning saw people exploit different aspects of noise and the way networks learn to enable denoising. Noise, while usually undesired, is simultaneously much harder to predict, as was elegantly demonstrated in [24], leading to a zero-shot denoiser. In the case specific to pixel noises, i.e. all forms of noise that are independent per image pixel (given the signal at that pixel) [22], the impossibility to predict the noise was exploited in various ways, leading to important contributions such as Noise2Void [14], Noise2Self [3], or Self2Self [13]. Another well known approach close to this family of approaches is Noise2Noise [16], capable of denoising even more complex noises that can correlate beyond the confines of single pixels. To further improve denoising performance, Probabilistic Noise2Void [15,21] introduced, and DivNoising [20] and Hierarchical DivNoising (HDN) [19] reused denoiSplit 3 the idea of suitably measured or trained pixel noise models. Such noise models are, in essence, a collection of probability distributions mapping from a true pixel intensity to observed noisy pixel measurements (and vice versa). 2.2 Image Decomposition Image decomposition is the inverse problem of splitting a given input image that is the superposition (i.e. the pixel-wise sum) of two constituent image channels. While the sum of two values is not uniquely invertible, if for each summand a prior on its value exists, even a unique solution can exist. In a similar vein, having learned structural priors of the appearance of the two constituent image channels, an input image (a grid of observed pixels that are each a sum of two values) can be split into two pixel grids such that each one satisfies the respective structural prior. In computer vision, reflection removal, dehazing, deraining etc. are some of the applications [2,4,7,8] for which image splitting can be used. More recently, image splitting in fluorescence microscopy was receiving heightened attention, probably because of the direct applicability and potential utility that a well-working approach can bring to this microscopy modality, which finds wide-spread use in biological investigations. In particular, a method called µSplit [1] demonstrated impressive image splitting performance on several datasets, suggesting that it is ready to be used in biological research projects. However, µSplit requires relatively noise-free data for training and prediction, which limits its potential utility (see also Section 5 or Figure 2). 2.3 Uncertainty Calibration The ability to correctly assess the quality of predictions is naturally useful. Ideally, a predictive system capable of co-predicting a confidence value has the property that the predicted confidence scales with the average error of the prediction. If the relationship between error and confidence is close to the identity, we call the uncertainty predictions of this system calibrated. Early works tried to use the deviation of the prediction as a proxy of the network’s confidence in the prediction [26]. Other works tried to calibrate the predicted standard deviation with the expected error (i.e. the RMSE) [17]. Earlier calibration works were mainly concerned with classification tasks [18]. However, in [17], these approaches were reformulated in the context of regression. In [17], the authors propose a way to evaluate calibration. They train a separate branch to predict a standard deviation per pixel that expresses its prediction uncertainty. For evaluating the calibration quality, the authors clustered examples on the basis of the predicted standard deviation values. Within each cluster, the predicted uncertainty is then compared with the empirical uncertainty (RMSE loss). To further improve the calibration, the authors propose a simple, yet effective scaling methodology wherein they learn a scalar parameter on re-calibration data, i.e. a subset of data not included in the training data. (In our experiments, we use the validation data for this purpose.) This 4 Ashesh, F. Jug scalar gets multiplied to the predicted uncertainty values, which then reduces the calibration error. 3 Problem Formulation Lets denote a noise free dataset containing npairs of images as D= (C1, C2), with each Cicontaining nimages Ci= (ci,j |1≤j≤n). Lets define a corresponding set of images X= (xj|1≤j≤n), such that all xj=c1,j +c2,j are the pixel-wise sum of the two corresponding channel images. Although Dis typically not available (or even observable), in practice we can only observe noisy data, denoted here by DN= (CN 1, CN 2). Analogously to before, we define XN= (xN j|1≤j≤n), such that xN j=cN 1,j +cN 2,j are the pixel-wise sum of the noisy channel observations. Given one xN j∈XNof DN, the task at hand is to predict the noise free and unmixed tuple (c1,j, c2,j). We shall denote the predictions made by a trained denoiSplit network by (ˆc1,j,ˆc2,j). Whenever above notions are used in a context that makes the jin the subscript redundant, we allow ourselves to omit them for brevity and readability. For evaluation purposes, we will in later sections use high-quality microscopy datasets that contain minimal levels of noise as surrogates for D,X,C1, and C2, but we never use them during training, and only their noisy counterparts are used. 4 Our Approach In the following sections we describe the main ingredients of denoiSplit, namely the hierarchical network structure we use (Section 4.1), the changed loss term for variational training of the splitting task (Section 4.2), the noise models we employ to enable the joint unsupervised denoising (Section 4.3), and an uncertainty calibration methodology allowing us to estimate the prediction error introduced by aleatoric uncertainty in a given input image (Section 4.4). 4.1 Network Architecture and Training Objective In this work, we employ an altered Hierarchical VAE (HVAE) network architecture. HVAEs were originally described in [23] and later adapted for image denoising in [19] and for image splitting in [1]. In general terms, HVAEs learn a hierarchical latent space, with the lowest hierarchy level encoding detailed pixellevel structure, while higher hierarchy levels capture increasingly larger scale structures in the training data. For denoiSplit, we modify the HVAE architecture so that it no longer remains an autoencoder. Instead, our outputs are the two unmixed channel images (ˆc1,ˆc2), motivating us to call the resulting architecture a Variational Splitting Encoder-Decoder (VSE) Network (see Fig. 1). denoiSplit 5 Our objective is to maximize the likelihood over the noisy two channel dataset we train on, i.e., finding decoder parameters θsuch that θ= arg max θX 1≤j≤n log P(cN 1,j, cN 2,j;θ).(1) Using the modified evidence lower bound (ELBO), as proposed in [1] and assuming conditional independence of the two predictions (ˆc1,ˆc2)given the latent space embedding, we maximize Eq(z|x;ϕ)[log P(cN 1|z;θ) + log P(cN 2|z;θ)] −KL(q(z|x;ϕ), P (z)),(2) where q(z|x;ϕ)is the distribution parameterized by the output of the encoder network Encϕ(x),P(cN i|z;ϕ)is the distribution parameterized by the output of the decoder network Decθ(z)and KL() denotes the Kullback-Leibler divergence loss. As in [1], P(z)factorizes over the different hierarchy levels in the network. Details about training, hyperparameters, µSplit, and its relationship with HDN and denoiSplit can be found in Supp. Sec. 2. Similar to the way noise models had been employed in the context of denoising [19], we model the two log likelihood terms log P(cN 1|z;θ)and log P(cN 2|z;θ) using noise models which we describe in detail in Section 4.3 and our open code repository1. 4.2 Hierarchical KL Loss Weighing for Variational Training In µSplit, the authors showed SOTA performance on a multitude of splitting tasks. However, used datasets were close to noise-free, making the task at hand simpler then the one we outlined in Section 3. When µSplit is trained on noisy datasets, the resulting channel predictions are themselves noisy. After analyzing this matter, we concluded that a modified KL loss can help reduce the amount of noise reconstructed by the decoder. In more technical terms, let Zbe a hierarchical latent space and Z[i]denote the latent space embedding at i-th hierarchy level, having shape (c, hi, wi), with cbeing channel dimension, and hi,withe height and width of the latent space embedding. Now let KLidenote the KL-divergence loss tensor computed on Z[i], which has the same shape as Z[i]itself. In µSplit, the corresponding scalar loss term kliis defined as kli=α· Pj,h,w KLi[j,h,w] hi·wi, with αbeing a suitable constant. Observe that the denominator makes each klibe the average of all values in KLi, making the respective values not scale with the size of Z[i], even though lower hierarchy levels (Z[i] for smaller i) have more entries. However, this also means that the KL loss for the individual pixels in these lower hierarchy levels is given less weight. Hence, smaller structures, such as noise itself, can more easily seep through such pixelnear hierarchy levels. 1https://github.com/juglab/denoiSplit 6 Ashesh, F. Jug In this work, we diverge from this formulation and return to a more classical setup where we compute the scalar loss term for the i-th hierarchy level Z[i]as kli=α·X j,h,w KLi[j, h, w].(3) The decisive difference is that this changed formulation gives more weight to the KL loss at lower hierarchy levels, leading to more strongly enforcing the Gaussian nature the KL loss enforces, and therefore hindering noise from being as easily represented during training. We refer to this architecture as Altered µSplit and show qualitative and quantitative results in Section 5 and Tables 1. The next section extends on Altered µSplit by adding unsupervised denoising, adding the last ingredient to the denoiSplit approach we present in this work. 4.3 Adding Suitable Pixel Noise Models As briefly introduced in Section 2, pixel noise models are a collection of probability distributions mapping from a true pixel intensity to observed noisy pixel measurements (and vice-versa) [15]. They have previously been successfully used in the context of unsupervised denoising [19,20] and we intend to employ them for this purpose also in the setup we are presenting here. We use the fact that, given a measured (noisy) pixel intensity, a pixel noise model returns a distribution over clean signal intensities and their respective probability of being the underlying true pixel value. We incorporate this likelihood function into the loss of our overall setup, encouraging denoiSplit to predict pixel intensities that maximize this likelihood and thereby values that are consistent with the noise properties of the given training data. Since denoiSplit, in contrast to existing denoising applications, predicts two images (the two unmixed channels), we employ two noise models and add two likelihood terms to our overall loss. More formally, in VAEs [12] and HVAEs, the generative distribution over pixel intensities is modeled as a Gaussian distribution with its variance either clamped to 1or also learned and predicted. We change our VSE Network to only predict the true pixel intensity and replace the Gaussian distribution mentioned above by the distributions defined in two noise models Pnm 1(cN 1|c1)and Pnm 2(cN 2|c2), one for each respective unmixed output channel. These noise models are pixel-wise independent, i.e., Pinm(cN i|ci) = Y k Pnm i(cN i[k]|ci[k]), i ∈ {1,2},(4) where cN i[k]is the noisy pixel intensity for the k-th pixel and ci[k]the corresponding noise-free intensity value. This independence makes them particularly suitable for microscopy data where Poisson and Gaussian noise are the predominant pixel noises one desires to remove. denoiSplit 7 Since we now directly predict the noise-free pixel values, the output of the decoder can directly be interpreted as Decθ(z) = (ˆc1,ˆc2)and the total loss for denoiSplit now becomes Eq(z|x;ϕ)[log Pnm(cN 1|ˆc1) + log Pnm(cN 2|ˆc2)] −KL(q(z|x;ϕ), P (z)).(5) In [20], two ways for the creation of noise models are described, and the decision to pick which method depends upon whether or not one has access to the microscope from which data was acquired. In Supp. Sec. 5, we describe the process of noise model generation and also compare performance between these two methodologies. 4.4 Computing Calibrated Data Uncertainties The idea of calibration is for those network setups that produce both prediction and a measure of uncertainty for the prediction. Networks that can co-assess the uncertainty of their predictions are called calibrated, when the predicted uncertainties are in line with the measured prediction error. To improve the calibration of a given system, one can find a suitable transformation from uncertainty predictions to measured errors (e.g., the RMSE). After such a transformation is found, an ideal calibrated plot would be tightly fitting y=x, with y and x being the error and estimated uncertainty, respectively. See, for example, Figure 3. Since VSE networks, similar to VAEs, are variational inference systems, we can sample from their latent encoding and thereby sample from an approximate posterior distribution of possible solutions giving us the data uncertainty. In this section, our intention is to utilize this ability to predict a reliable uncertainty term for our results. For this, we adapt the calibration methodology of [17]. In contrast to the approach described there, we propose to use the variability in posterior samples to estimate a pixel-wise standard deviation. More specifically, we sample k= 50 predictions for each input image and compute the pixel-wise standard deviations σ1and σ2for the two predicted image channels ˆc1and ˆc2, respectively. This gives us uncertainty predictions. Next, we calibrate these uncertainty predictions by scaling them appropriately with the help of two learnable scalars, α1and α2. Following [17], we assume that pixel intensities come from a Gaussian distribution. The mean and standard deviation of this distribution are the pixel intensities of the MMSE prediction, i.e. the image obtained after averaging k= 50 predictions, and the scaled σ, respectively. We learn the scalars α1and α2by minimizing the negative log-likelihood over the recalibration dataset. It is important to note that the presented calibration procedure does not alter the original predictions but instead learns a mapping that best predicts the measured error. To evaluate the quality of the resulting calibration, we sort the scaled standard deviations σi·sifor each pixel in a predicted channel and build a histogram over l= 30 equally sized bins Bj i. We then compute the root mean variance (RMV) and RMSE for each bin jand channel ias 8 Ashesh, F. Jug RMVi(j) = v u u t 1 |Bj i|X k∈Bj i (σk i·αi)2RMSEi(j) = v u u t 1 |Bj i|X k∈Bj i (ci[k]−ˆci[k])2 As in Section 3, ci[t]and ˆci[t]denote the noise-free pixel intensity and the corresponding prediction for i-th channel. In Fig. 3, we plot the RMSE vs. RMV for multiple tasks, observing that the plots closely resemble the identity y=x. Following [17], we use the validation dataset for recalibration and show calibration plots on the test dataset. 5 Experiments and Results 5.1 Datasets BioSR dataset We work primarily with BioSR dataset [11], a comprehensive dataset comprising fluorescence microscopy images of multiple cell structures. For our experiments, we have picked four structures, namely clathrin-coated pits (CCPs), microtubules (MTs), endoplasmic reticulum (ER), and F-actin. Since the raw data quality is very high and only a small amount of image noise is present in the individual micrographs, we add Gaussian noise and Poisson noise of various levels to these raw data. The artificially noisy images are used to train denoiSplit, while the raw data is shown to convince the reader of the validity of our approach and to compute evaluation metrics (see Figs. 2 and 4 and Tab. 1). Hagen et al. Actin-Mitochondria Dataset We picked the noisy Actin and Mitochondria channels from Hagen et al. [10], channels having real microscopy noise. For evaluation, we use the corresponding high-SNR (noise-free) channels provided in the dataset. Synthetic Noise Levels We work with 4levels of zero-mean Gaussian noise and two levels of Poisson noise. For Gaussian noise, we compute the standard deviation of the input data XNfor each of the tasks and scale the noise relative to one standard deviation. Specifically, the 4scaling factors are {1,1.5,2,4}. In cases where Poisson noise is added, and since it is already signal dependent, we use a constant factor of 1000 to hit a realistic-looking level of Poisson noise. We also consider the case where Poisson noise is not added, which we denote by the Poisson level of 0in Tab. 1. To remove any remaining room for misinterpretations, we provide a pseudo-code for the synthetic noising procedure in the Supp. Sec. 2. 5.2 Baselines We conducted all experiments with two baseline setups, µSplit and HDN⊕µSplit. In the original µSplit work [1], the authors introduce three architectures, each with a different trade-off between GPU efficiency, speed and performance. We denoiSplit 9 High SNRdenoiSplitµSplitGTInput A B C D c0 cN 0 c1 cN 1 c0 cN 0 c1 cN 1 c0 cN 0 c1 cN 1 c0 cN 0 c1 cN 1 Fig. 2: Qualitative Results. We show examples of noisy inputs, individual noisy channel training data (GT), and predictions by one of the baselines (µSplit) and our own results obtained with denoiSplit for four tasks (A: MT vs. CCPs, B: ER vs. CCPs, C: MT vs. ER, and D: F-actin vs. ER). We show high SNR channel images (not used during training) and show PSNR values w.r.t. these images. Additionally, we plot histograms of various panels for comparison (see legend on the right). The bottom cell in the first column of each panel shows the used noise models (see main text for details). The superimposed plots (green) show the distribution of noisy observations (cN i) for two clean signal intensities. 16 Ashesh, F. Jug 20. Prakash, M., Krull, A., Jug, F.: DivNoising: Diversity denoising with fully convolutional variational autoencoders. ICLR 2020 (Jun 2020) 21. Prakash, M., Lalit, M., Tomancak, P., Krull, A., Jug, F.: Fully unsupervised probabilistic Noise2Void. arXiv eess.IV (Nov 2019) 22. Shroff, H., Testa, I., Jug, F., Manley, S.: Live-cell imaging powered by computation. Nat. Rev. Mol. Cell Biol. (Feb 2024) 23. Sønderby, C.K., Raiko, T., Maaløe, L., Sønderby, S.K., Winther, O.: Ladder variational autoencoders. Adv. Neural Inf. Process. Syst. 29, 3738–3746 (Jan 2016) 24. Ulyanov, D., Vedaldi, A., Lempitsky, V.: Deep image prior. Int. J. Comput. Vis. 128(7), 1867–1888 (Jul 2020) 25. Weigert, M., Royer, L., Jug, F., Myers, G.: Isotropic reconstruction of 3D fluorescence microscopy images using convolutional neural networks. In: Medical Image Computing and Computer-Assisted Intervention - MICCAI 2017. pp. 126–134. Springer International Publishing (2017) 26. Weigert, M., Schmidt, U., Boothe, T., Müller, A., Dibrov, A., Jain, A., Wilhelm, B., Schmidt, D., Broaddus, C., Culley, S., Rocha-Martins, M., Segovia-Miranda, F., Norden, C., Henriques, R., Zerial, M., Solimena, M., Rink, J., Tomancak, P., Royer, L., Jug, F., Myers, E.W.: Content-aware image restoration: pushing the limits of fluorescence microscopy. Nature Publishing Group 15(12), 1090–1097 (Dec 2018) 27. Zhang, K., Zuo, W., Chen, Y., Meng, D., Zhang, L.: Beyond a gaussian denoiser: Residual learning of deep CNN for image denoising. IEEE Trans. Image Process. 26(7), 3142–3155 (Jul 2017)