==== Front ArXiv ArXiv arxiv ArXiv 2331-8422 Cornell University arXiv:2306.06408v2 2306.06408 2 preprint Article Fast light-field 3D microscopy with out-of-distribution detection and adaptation through Conditional Normalizing Flows Vizcaíno Josué Page 12* Symvoulidis Panagiotis 3 Wang Zeguan 3 Jelten Jonas 12 Favaro Paolo 4 Boyden Edward S. 3 Lasser Tobias 12 1 Computational Imaging and Inverse Problems, Department of Informatics, School of Computation, Information and Technology, Technical University of Munich, Germany 2 Munich Institute of Biomedical Engineering, Technical University of Munich, Germany 3 Synthetic Neurobiology Group, Massachusetts Institute of Technology, USA 4 Computer Vision Group, University of Bern, Switzerland * pv.josue@gmail.com 14 6 2023 arXiv:2306.06408v2https://creativecommons.org/licenses/by/4.0/ This work is licensed under a Creative Commons Attribution 4.0 International License, which allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use. nihpp-2306.06408v2.pdf Real-time 3D fluorescence microscopy is crucial for the spatiotemporal analysis of live organisms, such as neural activity monitoring. The eXtended field-of-view light field microscope (XLFM), also known as Fourier light field microscope, is a straightforward, single snapshot solution to achieve this. The XLFM acquires spatial-angular information in a single camera exposure. In a subsequent step, a 3D volume can be algorithmically reconstructed, making it exceptionally well-suited for real-time 3D acquisition and potential analysis. Unfortunately, traditional reconstruction methods (like deconvolution) require lengthy processing times (0.0220 Hz), hampering the speed advantages of the XLFM. Neural network architectures can overcome the speed constraints at the expense of lacking certainty metrics, which renders them untrustworthy for the biomedical realm. This work proposes a novel architecture to perform fast 3D reconstructions of live immobilized zebrafish neural activity based on a conditional normalizing flow. It reconstructs volumes at 8 Hz spanning 512 × 512 × 96 voxels, and it can be trained in under two hours due to the small dataset requirements (10 image-volume pairs). Furthermore, normalizing flows allow for exact Likelihood computation, enabling distribution monitoring, followed by out-of-distribution detection and retraining of the system when a novel sample is detected. We evaluate the proposed method on a cross-validation approach involving multiple in-distribution samples (genetically identical zebrafish) and various out-of-distribution ones. J. Page Vizcaíno is supported by the Deutsche Forschungsgemeinschaft (LA 3264/2-1). Z. Wang, P. Symvoulidis and E. S. Boyden acknowledge the Alan Foundation, HHMI, Lisa Yang, NIH R01DA029639, NIH 1R01MH123977, NIH R01MH122971, NIH RF1NS113287, NSF 1848029, NIH 1R01DA045549, and John Doerr. P. Favaro acknowledges the interdisciplinary project funding UniBE ID Grant 2018 of the University of Bern. ==== Body pmc1. Introduction Analysis of fast biological processes on live specimens is a crucial step in biomedical research, where fluorescence 3D microscopy plays an essential role due to its ability to visualize specific structures and processes, either through intrinsic contrast or an ever-grown collection of labeling techniques, including eg.genetically encoded indicators of neuronal activity. The extended field-of-view light field microscope (XLFM) [1], or Fourier light field microscope (FLFMic) [2–5], offers scan-less captures on transparent samples, affordability, and simplicity over scanning microscopes like spinning disk confocal and light sheet (capturing at 10Hz[6]). But at the expense of requiring a 3D reconstruction in a post-acquisition step. Traditionally, reconstructions are done using iterative methods, like the Richardson-Lucy deconvolution [7], where the reconstruction quality is sufficient to discern neural activity. However, large computation hardware allocation and long waiting times are required (1 second per iteration in an iterative algorithm), making it impractical for real datasets where thousands of images must be processed. As an alternative, deep learning approaches arose, where fast reconstructions are possible(up to 50Hz). These networks are trained on pairs of either raw XLFM or LFM [8] images and 3D volumes. Like the XLFMNet [9], the VCD [10], the LFMNet [11], HyLFM-Net [12], and others. However, within these methods, only the HyLFM-Net can detect network deviations or incapability of handling new sample types, but at the expense of a complex imaging system for continuous validation. An algorithm’s lack of certainty metrics renders it unsafe to be integrated in established experimental imaging workflows, as artifacts might be introduced by the network and cannot be detected. These motivations bring our attention to more statistically informing methods, such as Normalizing Flows (NFs) [13,14], a type of invertible neural network recently used for biomedical imaging [15], inverse problems [16–19] and, other applications like image generation [20,21]. NFs learn a mapping between an arbitrary statistical distribution and a normal distribution through a set of invertible and differentiable functions. Also, a tractable exact likelihood allows for probing the quality of this mapping for individual samples, which, in turn, allows for deciding what to do with the new sample, perhaps retraining the network with the new data, if desired. One disadvantage of NFs is that due to the required invertible mapping, no data bottlenecks are possible (like in an encoder-decoder approach) as this would lead to information loss, making it necessary to store all the tensors and gradients in memory during training. This limits the data size that can be used due to the limited graphics processing unit (GPU) memory. The conventional solution to this issue is to split the processed tensor after each invertible function [22], feed only one part to the following function, and concatenate the other part to the output tensor of the normalizing flow (NF). However, when working with large tensors, the gradients during training still overwhelm conventional GPUs like the ones found on image acquisition workstations. Hence, Wavelet-Flow (WF) [20] was introduced, where an invertible down-sampling operation is used (such as the Haar transform [23]) to serially down-sample the input image to the desired size (could be down to a single pixel), where a NF is used to learn the Haar detail coefficients required to perform up-sampling when performing reconstruction. The last down-sampling step comprises an NF that directly learns the probability distribution of the lowest-resolution image. This allows independent training of each down-sampling NF with commercial GPUs, allowing their usage for high data throughput for the first time. Generating a new image involves running the WF backward. The lowest resolution image is sampled from an NF, then up-sampled through all the WFs until reaching the original resolution. The original WF approach is not designed for inverse problems as it ignores the image formation model prior knowledge, such as the raw XLFM image or the point spread function. Furthermore, when training each NFs, the Haar transform operator generates a down-sample of the ground truth (GT) volume down to the lowest resolution. And when performing a 3D reconstruction, the low-resolution volume received by the NFs might slightly deviate from the GT. These deviations accumulate as the volumes are propagated upwards through the network, affecting the output 3D volume heavily. The quality of the lowest-resolution volume dramatically impacts the final reconstruction, making the lowest-resolution initial guess fundamental to the quality of the full-resolution reconstruction. Our work proposes the Conditional Wavelet-Flow Architecture (CWFA), shown in Fig. 1, suited for the 3D reconstruction of live immobilized fluorescent samples imaged through an XLFM and out-of-distribution detection (OODD). It uses WFs conditioned by the XLFM-measured image and a 3D volume prior (eg.the mean of the training volumes). And could be easily modified to work on freely behaving animals by removing the dependency on a 3D volume prior. The CWFs reconstruction, and out-of-distribution (OOD) capabilities enable a fast, accurate, and robust method for 3D fluorescence microscopy. Fast enough to be applied within close-loop or human-in-the-loop experiments, where on the fly analysis of the activity could trigger further steps in an experiment. In OODD, having access to an exact likelihood computation allows for evaluating how likely a sample is to belong to the training distribution at different image scales. Measured in all the down-sampling Conditional Normalizing-Flow (CNF)s in the network. Specifically, in a testing step, the likelihood of a novel sample can be computed by processing it with the CWF, as described in sec. 3.2 and Fig. 2. Even though the literature often mentions that NFs are not reliable for detecting OOD samples [24], we find that the proposed CWFA and sample type enables this capability. We demonstrate the reliability of the proposed CWFA on XLFM acquisitions of live immobilized zebrafish larvae full-brain neural activity, processed with the SLNet [9], which separates the neural activity from the background of the acquisitions. As a gold standard or GT we used 3D reconstructions (100 iterations with the RL algorithm) of the SLNet output. In sec. 3 and Fig. 3, we compare the Richardson-Lucy (RL) 3D reconstruction (ground truth or golden standard) against the proposed CWFA, the XLFMNet [25] and the WF [20] on volumes spanning a FOV of 734×734×225μm3 (512 × 512 × 96 voxels), the proposed CWFA operated 735× faster than the traditional deconvolution (1 second per RL iteration, we used 100 iterations in this work), with similar quality, as indicated by a mean PSNR of 51.48, a mean absolute percentage error (MAPE) [26] of 0.1011 or 10.11%. Alternative methods such as the XLFMNet and WF achieved a speed increase of 3200× and 961×, but with low-quality reconstructions (PSNR of 47.35 and 44.30 and MAPE of 0.3594 and 0.1167 respectively. Furthermore, when analyzing the neural activity of individual neurons in a time sequence, as seen in Fig. S5, the proposed method achieves a Pearson correlation coefficient (PCC) [27] of 0.9077, where the XLFMNet and WF achieve 0.6897 and 0.8428 respectively. In domain shift detection, the first CWF performed best and achieved a classification F1-score [28] of 0.988 and an area under the curve (AUC) of 0.997 , as seen in Fig. 4. Once an out-of-distribution sample type is detected, we found that 5 minutes of fine-tuning of the proposed CWFA on a novel sample are enough to increase quality substantially. Achieving a mean increase of 36.56% in PSNR 40.71% in MAPE and 2.15% PCC. To conclude, the CWFA is 735× faster than the reconstruction gold standard and 9% more accurate than the XLFMNet, a Convolutional Neural Network (CNN) previously used for this particular task. The CWFA offers excellent domain shift detection that can trigger either re-training or fine-tuning on new data. 2. Materials and Methods 2.1. Tradicional Conditional normalizing flows A NF, through a sequence of invertible and differentiable functions, transforms an arbitrary distribution pX(x) into the desired distribution pZ(z) (usually a normal distribution, hence the name Normalizing Flows). This is possible through the change of variables formula from probability theory, where the density function of the random variable X is given by: (1) pX(x)=pZ(z)|det(J)|, Where z is normally distributed (mean 0 and variance 1), with probability density function pZ(z). An NF can be trained by setting z=fΘ(x), where fΘ is an invertible differentiable function parameterized by Θ, and J=JΘ=∂fΘ(x)∂x is the Jacobian of fΘ with respect to x, also known as the volume correction term. As seen in Eq. 1, a tractable and easily computable Jacobian determinant of fΘ(x) is preferred (for example, where the Jacobian is block-triangular or diagonal). Hence, choosing the functions conforming fΘ(x) is a crucial and well-studied step [22, 29,30], out of the scope of this work. A NF can be modified into a CNF and represent a conditional distribution pX(X∣C), for a set of observations X=x(1),x(2), … ,x(N) and conditions C=c(1),c(2), … ,c(N). The likelihood from Eq. 1 for a single sample (i) becomes: (2) pX(x(i)∣c(i),Θ)=pZ(z(i)=fΘ(x(i),c(i)))⋅|det(JΘ(i))|. To train a CNF, we need to find the optimal values of the parameters Θ that maximize the likelihood or minimize the negative log-likelihood of Eq.2, defined as follows: (3) Θ*=arg minΘ∑i=1N[‖fΘ(x(i),c(i))‖222−log|det(JΘ(i))|+ρ‖Θ‖22] Where det(JΘ(i)) is the determinant of the Jacobian matrix with respect to x(i), evaluated at fΘ(x(i),c(i)). And ρ∥Θ∥22 is the likelihood of the posterior over the model’s parameters, assuming a Gaussian distribution weighted by ρ. During training, we minimize the negative log-likelihood function with respect to the parameters Θ using the Lion optimizer [31] using weight decay for the parameters posterior. After training the CNF, we can perform inference on a new image by sampling from the base distribution pZ(z) (in our case, a normal distribution 𝒩(μ,σ2)) and obtaining x by applying the inverse transformation of the flow: x=fΘ−1(z). The resulting x will be a sample from the conditional distribution pX(X∣C). Fig. S1 shows a graphical representation of a CNF used for inference. Even though this method is mathematically sound, its lossless nature limits its usability in practice for large 3D volumes due to high memory requirements. 2.2. Proposed Conditional Wavelet Flow architecture The WF architecture uses a multi-scale hierarchical approach that allows training each up/down-scale independently, allowing flexibility during memory management. Each down-sampled operation is a Haar transform chosen due to its orthonormality and invertibility. Alternatively, to the WFs, we used the Haar transform to scale the volumes in the axial dimension, preserving lateral resolution across all steps. The likelihood function of a WF model is given by: (4) p(V0)=p(Vn)∏i=1n−1p(Di∣Vi) where V0 is the final high-resolution volume, Di the Haar transform detail coefficients and Vi the volume down-sampled i times. p(Vn) and all p(Di∣Vi) are normalizing flows. In the proposed CWFA, we substituted p(Vn) with a deterministic CNN. We conditioned all the up-sampling NFs on a set of external conditions C processed by the CNN Ωi, as shown in Fig.1. With a likelihood given by: (5) p(V0)=p(Vn)∏i=1n−1p(Di∣Ωi(C)) And the log-likelihood (LL): (6) log p(V0)=log p(Vn)+∑i=1n−1log p(Vi∣Ωi(C)) The final loss function can be optimized independently as log p(V0) comprises a sum of probabilities. Showing the advantage of the Wavelet-Flow architecture when using conventional computing hardware. 2.2.1. Network implementation The proposed architecture uses four CWFs, each comprised of a Haar transform down-sampling operation and an CNF, as seen in Fig. 1. The internal CWF learns the mapping between the Haar coefficients and a normal distribution. And is built from 6 conditional affine transform (CAT) blocks (see Supplementary Fig. S1). The number of blocks and their parameters, such as the type of invertible blocks (like GLOW [30], RNVP [22], HINT [32], NICE [13], CAT [29] etc.), the number of parameters in each internal convolution, were optimized for PCC in a grid-like fashion shown in Fig. S2. The architecture was implemented in PyTorch aided by the Freia framework for easily invertible architectures [33]. We invite the reader to explore the source code of this project for further implementation details unde https://github.com/pvjosue/CWEA. The network’s total number of parameters is 73 million, where 3.418 million comprise the CWFs, 82K the conditional networks Ω0−n and 63.743 million the LR-NN. Additional details are in the supplementary section 1. 2.3. 3D reconstruction: Sampling from the Conditional Wavelet Flow architecture Fig. 1 depicts how to train the CWFA (forward pass) and reconstruct 3D volumes from XLFM images (inverse pass). Reconstruction with a trained CWFA involves, first, reconstructing a low-resolution volume (V˜n=LR−NN(C)), where LR-NN is a deterministic CNN and C the conditions, and using it as input to the CWFn−1. To up-sample on this CWF: first, sample zn−1 from a normal distribution, and input the pre-processed XLFM image and 3D prior conditions Ωn−1(C) as a condition to the NF. Then, generate the Haar coefficients (Dn−1) used for up-sampling V˜n−1 by a factor of 2 on the axial dimension with the Haar transform and generate V˜n−2. We repeat this process until reaching V˜0 at full resolution. Our architecture comprises 4 CWF and LR-NN. Each CWF uses 6 CNFs internally, with 14 channels per convolution and CAT invertible blocks, as illustrated in Fig. S1. A key parameter during sampling is the temperature parameter, which determines if the z should be sampled from a truncated distribution and to what degree, which is discussed in sec. 3.1.2. 2.4. Input for training As an input to the network, the GT volume at full resolution V0 and the conditions C is required. For both, we pre-processed the raw data with the SLNet, extracting the sparse activity from the image sequences. This approach was chosen due to the fluorescent labeling used (GCaMP), which has only a slight intensity increase (<10%) when intracellular calcium concentration increases as a result of neuronal firing. In other words, the neural activity is practically invisible on the raw sequences that suffer from substantial auto-fluorescence. 2.5. Conditioning the wavelet flows In the original WF [20], each NF uses only the next level low-resolution volume as a condition, intending to generate human faces from a learned distribution, where the low-resolution Haar transformed image is used as a condition on each up-sampling step. In such a case, face distribution is the only known information. However, when dealing with inverse problems (eg.fluorescence microscopy), prior information about the system and volume to reconstruct are known, such as the forward process in the form of a point spread function (PSF), the captured microscope image and in our case the structural information of a fish, as these are immobilized with agarose and do not move during the acquisition. After the ablation of different configurations (see Fig. S2), we chose to use CAT blocks due to their excellent performance and simplicity. These split the input condition in two, comprised of a translation and a scaling factor (as seen in Fig. S1) that is applied to the CWF’s input in the forward or backward direction. The following conditions are fed to each CWFi after being pre-processed by Ωi(C):. 2.5.1. Condition 1: Views cropped from the XLFM image This condition acts as the scaling factor of the block and informs the network about neural potential changes. Due to the nature of the XLFM microscope, each microlens acts as an individual camera, and the 2D image can be interpreted as a multi-view camera problem. The raw XLFM input image is prepared first by detecting the center of each micro-lens from the central depth of the measured PSF (using the Python library findpeaks [34]), then cropping a 512 × 512 area around the 29 centers, and stacking the images in the channel dimension, as seen in Fig 1 panel (b). 2.5.2. Condition 2: A 3D volume structural prior When working with immobilized animals, there’s the advantage that only the neural activity changes within consecutive frames. Hence, we included this as a prior in the shape of a condition, where we provide to the network a 3D volume created by the mean of the training volumes. Providing a volumetric prior simplifies the reconstruction problem and allows the network to focus only on updating the neural activity instead of reconstructing a complete 3D volume. Furthermore, as the Haar coefficients along the channel dimension are a discreet derivative of the volume, we found that a processed version of the volume is a very initial approximation, which will be fine-tuned by the CWF. 2.5.3. Conditional networks Ω0−n Previous methods using CNF used feature extractors such as the first layers of a pre-trained network. In this work, we simultaneously trained the conditional networks with the CWFs. This, by adding a second data term to the loss function from Eq. 3 resulting in the final loss function: (7) Θi=Θi*+α*arg minΘ∑i=1N[‖Vi−V˜i)‖22] Where α is a weighting factor (0.48 in this work), V˜i is the reconstructed volume at the CWF i and Vi is the GT volume down-sampled i times. In this way, we ensure that the volumes generated by the CWFA have not just a statistical constraint but a spatial constraint. Without this, the loss might be minimized as the voxel intensities follow the correct distribution, but there is no constraint that, spatially, the reconstructions make sense. 2.6. Out-of-distribution detection The OODD workflow is visually depicted in Fig. 2. Once we trained a CWFA on a set of fish image-volume pairs, we can evaluate if the network can adequately handle a novel sample (Inovel). First, by deconvolving Inovel into Vnovel, building the condition Cnovel, and processing these through the network in the forward direction. Finally, we can evaluate the generated distributions z0−n with eq. 3. If the negative-log-likelihood (NLL) obtained are above a pre-defined threshold, we have an OOD sample in hand. Once detected, there are different solutions, eg.we can fine-tune the CWFA on the test sample or add the test sample to the training set. We explore these two options in sec. 3.3. The OODD threshold is picked by selecting a small number of samples from all cross-validation training and testing sets and evaluating their likelihood. Then define 1000 thresholds linearly distributed spanning the full range of the data, and pick the one achieving the most significant AUC. 2.7. Dataset acquisition and pre-processing Pan-neuronal nuclear localized GCaMP6s Tg(HuC:H2B:GCaMP6s) and pan-neuronal soma localized GCaMP7f Tg(HuC :somaGCaMP7f) [35] zebrafish larvae were imaged at 4–6 days post fertilization. The transgenic larvae were kept at 28°C and paralyzed in standard fish water containing 0.25 mg/ml of pancuronium bromide (Sigma-Aldrich) for 2 min before imaging to reduce motion. The paralyzed larvae were then embedded in agar with 0.5% agarose (SeaKem GTG) and 1% low-melting point agarose (Sigma-Aldrich) in Petri dishes. Each fish was imaged for 1000 frames at 10 Hz. Neural activity images were extracted using the SLNet [25], as seen in Fig. 1(a). Later, the resulting images were 3D reconstructed with the RL algorithm for 100 iterations, which takes roughly 1.5 minutes per frame. A cross-validation approach was used to evaluate the system, using 6 different fish, each fold trained on 5 different fish and tested on the remaining fish. For testing OODD we used the testing fish on each fold, non-sparse images (deconvolutions of raw data, without pre-processing of the SLNet), and fluorescent beads images. Ten XLFM pre-processed images and volumes per dataset were used to train the CWFA, as each cross-validation set has 5 training datasets; in total, 50 images were used for training, 250 for validation, and 50 for testing. We tested different amounts of data used for training (5,10,20, and 50 pairs), where using 10 images dealt the best performance and a training time of 1:43h, see supplementary Fig. S4 for details. The beads dataset was created by imaging 1−μm-diameter green fluorescent beads (Ther-284 moFisher) randomly distributed in 1% agarose (low melting285point agarose, Sigma-Aldrich). The stock beads were serially diluted using melted agarose to 10−3, 10−4, 10−5, 10−6 of the original concentration. 3. Experiments 3.1. 3D reconstruction of sparse images with the CWFA We compared the proposed method against the XLFMNet [9] aiming to match the amount of parameters (103M) and a modified version of the Wavelet Flow [20] (WF). The XLFMNet is a U-net-based architecture [36] that achieves high reconstruction speeds but lacks any certainty metrics (as in conventional deep learning techniques). In the original WF, the lowest resolution was reconstructed from an unconditioned normalizing flow based purely on the training data distribution, which does not comply with inverse problems modus-operandi, where a measurement is required to reconstruct the variable in question. The modified version follows the original design in which only the low-resolution image is used as a condition in each flow; however, we used a CNN Ωn instead of the lowest-resolution NF, informing the system about the measurement and making a fair comparison. We used similar settings optimized for PCC as the proposed approach (Fig. S2). 3.1.1. Comparison metrics We identified three relevant aspects for quality evaluation of a 3D reconstruction: General image quality: We used peak signal-to-noise ratio (peak signal-to-ratio (PSNR)) as it compares the pixel-wise quality of reconstruction against the GT. Sparse image quality: As the volumetric data used is highly sparse (mostly zeros), we created a mask of the non-zero values on both the GT and reconstructions, and, used the mean absolute percentage error (MAPE) for single frame quality assessment Temporal consistency: An important aspect is that the neural activity is correctly reconstructed across frames. Hence, we used the Pearson correlation coefficient (PCC) of single neurons across time. The neuron positions on each data sample were determined with the suite2p framework [37]. As seen in Fig. 3, the proposed method outperforms the other two approaches regarding these metrics. However, XLFMNet provides faster inference capabilities but lacks any certainty metrics. 3.1.2. Sampling temperature In our experiments, we found that a temperature of zero produced the highest quality reconstruction. This might be the optimal parameter as the GT used is the result of the RL algorithm, which converges towards the maximum-likelihood estimate. A zero temperature means the network reconstructs the most likely sample, which makes sense for an inverse problem. Furthermore, zero temperature means no sampling is required, increasing the network performance. A comparison of different sample temperatures can be found in the supplementary Fig. S4. 3.2. Domain shift or out-of-distribution-detection (OODD) We evaluated our method by training the CWFA on six cross-validation folds of zebrafish fluorescent activity datasets pre-processed with the SLNet extracting the neural activity. Then, presented to the pipeline with different sample types, such as previously unseen fish sparse images (processed with the SLNet), raw XLFM fish images (not pre-processed), and fluorescent bead images with different concentrations. As shown in Fig. 4. Our algorithm achieves across all cross-validation folds a mean AUC of 0.9964 and F1-score of 0.9916 on CWF1, with a NLL threshold of −1.33. The achieved AUC and F1-scores for all the down-sampling CWFs are presented in table S3. The OOD NLL threshold was established by computing the ROC curve on all cross-validation sets simultaneously and choosing the threshold with the highest F1-score and AUC, in our case −1.33. 3.3. Domain shift adaptation through fine-tuning Once a sample is detected as OOD, a couple of possible approaches are: Fine-tune the pre-trained network on the new data: We first need to deconvolve images from the new sample, where each deconvolution takes around 1 minute. Then fine-tune the testing set on each of the cross-validation folds with 10 image-volume pairs for 100 epochs (20 epochs per step). This took roughly 10 minutes to generate the training pairs and 5 minutes for fine-tuning. Dealing a mean increase of 36% on PSRN, 40% on MAPE, and 2% onPCC, as seen in table S2, and on Fig. 4, where we fine-tuned the ‘Test different fish’ into ‘Fine-tuned.’ Append new data to cross-validation fold: If we still want to use the network for the fish it already trained on, we can append the new training data to the training set, and fine-tune all the data. This approach takes roughly 10 minutes to generate the training pairs and 25 minutes for fine-tuning. Dealing a mean increase of 34% on PSRN, 34% on MAPE, and 1% onPCC, as seen in Fig. 4 as ‘Fine-tuned all datasets’. 4. Discussion In this work, we presented a Bayesian approach to 3D reconstruction of live immobilized fluorescent zebrafish, comprised of a robust workflow for inverse problems, particularly when false positives should be minimized, as in the case of bio-medical data. Remarkably, the amount of data (10 images per sample) and training time (around 2 hours) combined with the OODD capability would allow this system to be integrated into a downstream analysis workflow. And when a new fish or fluorescent sample needs analysis, 5 minutes of retraining would suffice to allow the network to reconstruct neural activity reliably. There are some remaining questions that we leave for future work. Such as, which element in the setup enables the OODD. We have some hints, but that would require further experimentation. For instance, the fact that the classification capability increases in higher resolution CWF steps (Table. S3) might indicate that the hierarchical approach based on the Haar transform aids the likelihood clustering. It also might be the type of data and initial distribution P(x) (usually Poisson distributed due to fluorescence) is relevant. We find that NFs and Bayesian approaches are adequate for bio-medical imaging due to their potential capability of handling uncertainty and correctly presenting it to the user, allowing for better decision-making and false positive minimization. Supplementary Material 1 Funding. J. Page Vizcaíno is supported by the Deutsche Forschungsgemeinschaft (LA 3264/2-1). Z. Wang, P. Symvoulidis and E. S. Boyden acknowledge the Alan Foundation, HHMI, Lisa Yang, NIH R01DA029639, NIH 1R01MH123977, NIH R01MH122971, NIH RF1NS113287, NSF 1848029, NIH 1R01DA045549, and John Doerr. P. Favaro acknowledges the interdisciplinary project funding UniBE ID Grant 2018 of the University of Bern. Data Availability Statement. The sourcecode and dataset of this project can be found under https://github.com/pvjosue/CWFA and 10.5281/zenodo.8024696 respectibly Fig. 1. The Conditional Wavelet Flow Architecture and workflow: (a) Data preparation, extracting the sparse spatiotemporal signal from the raw XLFM acquisitions with the SLNet [9] and performing 3D reconstructions using the RL algorithm. In (b), the conditions are prepared by cropping and stacking the images and computing the mean of the training volumes. In (d), the full-resolution GT volume V0 and conditions (b) are used to train the Conditional Wavelet-Flow (CWF)s. Training is performed in each CWF individually and consists in feeding V0 and the processed condition Ω1(C) to the CWF1. This generates as outputs V1 and z1. The latter is used in loss function 3. V1 is fed to the next CWF, which is trained similarly. This is repeated until reaching the lowest resolution output. The LR-NN is trained deterministically with Vn and C pairs. Fig. 2. Steps for OOD detection and domain shift adaptation Fig. 3. 3D reconstruction comparison of zebrafish images with different methods. On the top row, GT volumes were generated through 100 iterations of the RL algorithm using a measured PSF. On each column, a different zebrafish sample. The following rows show reconstructions with different methods. The left-most column shows the performance metrics used for comparison: MAPE, followed by the mean PCC of 50 frames from the same fish acquisition measured on the top 50 most active neurons per fish. The bottom arrows show the direction of better performant metrics. Fig. 4. Distribution analysis of different sample types. The negative log-likelihood vs. PSNR of the first CWF module is on the left panel. Different sample types and a OOD threshold are presented, used to determine if the network can handle a sample reliably. For the case of the ‘Test different fish’, we present two re-training approaches, as mentioned in sec. 3.3. And on the right panel, the NLL density, where it can be seen that a threshold can be found to separate in vs. out of distribution data. Disclosures. The authors declare no conflicts of interest. Supplemental document. See Supplement 1 for supporting content. ==== Refs References 1. Wen Lin Cong Q. , Wang Z. , Chaik Y. , Hang W. , Shang C. , Yang W. , Bai L. , Du J. , and Wang K. , “Rapid whole brain imaging of neural activity in freely behaving larval zebrafish (danio rerio),” eLIFE 6 (2017). 2. Hua X. , Liu W. , and Jia S. , “High-resolution fourier light-field microscopy for volumetric multi-color live-cell imaging,’ Optica 8 , 614–620 (2021).34327282 3. Liu W. , Guo C. , Hua X. , and Jia S. , “Fourier light-field microscopy: An integral model and experimental verification,” Biophotonics Congr. Opt. Life Sci. Congr. 2019 (BODA,BRAIN,NTM,OMA,OMP) p. DT1B.4 (2019). 4. Han K. , Hua X. , Vasani V. , Kim G.-A. R. , Liu W. , Takayama S. , and Jia S. , “3d super-resolution live-cell imaging with radial symmetry and fourier light-field microscopy,” Biomed. Opt. Express 13 , 5574 (2022).36733732 5. Stefanoiu A. , Scrofani G. , Saavedra G. , Martinez-Corral M. , and Lasser T. , “What about super-resolution in fourier lightfield microscopy?” Opt. Express 28 (2020). 6. Hillman E. M. C. , Voleti V. , Li W. , and Yu H. , “Light-Sheet microscopy in neuroscience,” Annu. Rev Neurosci 42 , 295–313(2019)31283896 7. Lucy L. B. , “An iterative technique for the rectification of observed distributions,” The astronomical journal 79 , 745 (1974) 8. Levoy M. , Ng R. , Adams A. , Footer M. , and Horowitz M. , “Light field microscopy,” ACM Trans. Graph. 25 , 924–934 (2006) 9. Page Vizcaino J. , Wang Z. , Symvoulidis P. , Favaro P. , Guner-Ataman B. , Boyden E. S. , and Lasser T. , “Real-time light field 3d microscopy via sparsity-driven learned deconvolution,” in 2021 IEEE International Conference on Computational Photography (ICCP), (2021), pp. 1–11. 10. Wang Z. , Zhu L. , Zhang H. , Li G. , Yi C. , Li Y. , Yang Y. , Ding Y. , Zhen M. , Gao S. , Hsiai T. K. , and Fei P. , “Real-time volumetric reconstruction of biological dynamics with light-field microscopy and deep learning,” Nat. Methods (2021) 11. Vizcaíno J. P. , Saltarin F. , Belyaev Y. , Lyck R. , Lasser T. , and Favaro P. , “Learning to reconstruct confocal microscopy stacks from single light field images,” IEEE Trans. on Comput. Imaging 7 , 775–788 (2021). 12. Wagner N. , Beuttenmueller F. , Norlin N. , Gierten J. , Boffi J. C. , Wittbrodt J. , Weigert M. , Hufnagel L. , Prevedel R. , and Kreshuk A. , “Deep learning-enhanced light-field imaging with continuous validation,” Nat. Methods 18 , 557–563 (2021)33963344 13. Dinh L. , Krueger D. , and Bengio Y. , “Nice: Non-linear independent components estimation,” arXiv preprint arXiv: 1410.8516(2014) 14. Rezende D. J. and Mohamed S. , “Variational inference with normalizing flows,” (2015). 15. Denker A. , Schmidt M. , Leuschner J. , and Maass P. , “Conditional invertible neural networks for medical imaging,” J. Imaging 7 (2021) 16. Ardizzone L. , Kruse J. , Wirkert S. , Rahner D. , Pellegrini E. W. , Klessen R. S. , Maier-Hein L. , Rother C. , and Köthe U. , “Analyzing inverse problems with invertible neural networks,” arXiv preprint arXiv:1808.04730 (2018). 17. Ardizzone L. , Kruse J. , Wirkert S. , Rahner D. , Pellegrini E. W. , Klessen R. S. , Maier-Hein L. , Rother C. , and Köthe U. , “Analyzing inverse problems with invertible neural networks,” 7th Int. Conf. on Learn. Represent. ICLR 2019 pp. 1–20 (2019). 18. Anantha Padmanabha G. and Zabaras N. , “Solving inverse problems using conditional invertible neural networks,” J. Comput. Phys. 433 , 110194 (2021). 19. Kousha S. , Maleky A. , Brown M. S. , and Brubaker M. A. , “Modeling srgb camera noise with normalizing flows,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2022), pp. 17442–17450. 20. Yu J. J. , Derpanis K. G. , and Brubaker M. A. , “Wavelet flow: Fast training of high resolution normalizing flows,” in Advances in Neural Information Processing Systems, vol. 33 Larochelle H. , Ranzato M. , Hadsell R. , Balcan M. , and Lin H. , eds. (Curran Associates, Inc., 2020), pp. 6184–6196. 21. Ardizzone L. , Lüth C. , Kruse J. , Rother C. , and Köthe U. , “Guided image generation with conditional invertible neural networks,” arXiv:1907.02392 (2019). 22. Dinh L. , Sohl-Dickstein J. , and Bengio S. , “Density estimation using real nvp,” arXiv preprint arXiv:1605.08803 (2016) 23. Haar A. , “Zur theorie der orthogonalen funktionensysteme.” Math. Ann. 69 , 331–371 (1910). 24. Kirichenko P. , Izmailov P. , and Wilson A. G. , “Why normalizing flows fail to detect out-of-distribution data,” Adv. neural information processing systems 33 , 20578–20589 (2020). 25. Page Vizcaino J. , Wang Z. , Symvoulidis P. , Favaro P. , Guner-Ataman B. , Boyden E. S. , and Lasser T. , “Real-time light field 3d microscopy via sparsity-driven learned deconvolution,” in 2021 IEEE International Conference on Computational Photography (ICCP), (2021), pp. 1–11. 26. de Myttenaere A. , Golden B. , Le Grand B. , and Rossi F. , “Mean absolute percentage error for regression models,” Neurocomputing 192 , 38–48 (2016). 27. Benesty J. , Chen J. , Huang Y. , and Cohen I. , “Pearson correlation coefficient,” in Noise reduction in speech processing, (Springer, 2009), pp. 1–4. 28. Chinchor N. , “Image quality metrics: Psnr vs. ssim,” in Proc. of the Fourth Message Understanding Conference, (MUC-4), (1992), pp. 22–29. 29. Park T. , Liu M.-Y. , Wang T.-C. , and Zhu J.-Y. , “Semantic image synthesis with spatially-adaptive normalization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2019). 30. Kingma D. P. and Dhariwal P. , “Glow: Generative flow with invertible 1×1 convolutions,” Adv. neural information processing systems 31 (2018). 31. Chen X. , Liang C. , Huang D. , Real E. , Wang K. , Liu Y. , Pham H. , Dong X. , Luong T. , Hsieh C.-J. , Lu Y. , and Le Q. V. , “Symbolic discovery of optimization algorithms,” (2023). 32. Kruse J. , Detommaso G. , Köthe U. , and Scheichl R. , “Hint: Hierarchical invertible neural transport for density estimation and bayesian inference,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35 (2021), pp. 8191–8199. 33. Ardizzone L. , Bungert T. , Draxler F. , Köthe U. , Kruse J. , Schmier R. , and Sorrenson P. , “Framework for Easily Invertible Architectures (FrEIA),” (2018–2022). 34. Taskesen E. , “findpeaks is for the detection of peaks and valleys in a 1D vector and 2D array (image).” (2020). 35. Shemesh O. A. , Linghu C. , Piatkevich K. D. , Goodwin D. , Celiker O. T. , Gritton H. J. , Romano M. F. , Gao R. , Yu C.-C. J. , Tseng H.-A. , “Precision calcium imaging of dense neural populations via a cell-body-targeted calcium indicator,” Neuron 107 , 470–486 (2020).32592656 36. Ronneberger O. , Fischer P. , and Brox T. , “U-net: Convolutional networks for biomedical image segmentation,” CoRR abs/1505.04597 (2015) 37. Pachitariu M. , Stringer C. , Dipoppa M. , Schröder S. , Rossi L. F. , Dalgleish H. , Carandini M. , and Harris K. D. , “Suite2p: beyond 10,000 neurons with standard two-photon microscopy,” bioRxiv (2017).