Data/Math (0000), xx:xx 1–33 METHODS PAPER Actively inferring methane sources with drones Alouette van Hove1, Kristoffer Aalstad1and Norbert Pirk1 1Department of Geosciences, University of Oslo, Oslo, Norway. E-mail:
[email protected]. Keywords: methane monitoring; Bayesian inference; reinforcement learning; adaptive informative path planning; drones. Abstract This study presents a framework that combines Bayesian inference with reinforcement learning to guide drone-based sampling for methane source estimation. Synthetic gas concentration and wind observations are generated using a calibrated model derived from real-world drone measurements, providing a more representative testbed that captures atmospheric boundary layer variability. We compare three path planning strategies—preplanned, myopic (shortsighted), and non-myopic (long-term)—and find that non-myopic policies trained via deep reinforcement learning consistently yield more precise and accurate estimates of both source location and emission rate. We further investigate centralized multi-agent collaboration and observe comparable performance to independent agents in the tested single-source scenario. Our results suggest that effective source term estimation depends on correctly identifying the plume and obtaining low-noise concentration measurements within it. Precise localization further requires sampling in close proximity to the source, including slightly upwind. In more complex environments with multiple emission sources, multi-agent systems may offer advantages by enabling individual drones to specialize in tracking distinct plumes. These findings support the development of intelligent, data-driven sampling strategies for drone-based environmental monitoring, with potential applications in climate monitoring, emission inventories, and regulatory compliance. Impact Statement Methane is a potent greenhouse gas, making accurate monitoring of its sources essential for effective climate action. This study presents a framework that integrates drone-based measurements with artificial intelligence to improve the detection and quantification of methane emissions. By combining Bayesian inference with atmospheric dispersion modeling and drone-based atmospheric data, the system estimates source locations and emission rates. Reinforcement learning enables the drone to plan adaptive flight paths that maximize information gain in unfamiliar environments. This framework supports the development of intelligent, efficient tools for environmental monitoring, with applications in climate science, emissions reporting, and regulatory compliance. 1. Introduction Methane (CH4) is an important greenhouse gas emitted by a range of anthropogenic activities, including fossil fuel extraction, agriculture, and waste management, as well as natural sources such as inland waters, wild animals, and thawing permafrost (IPCC 2023). Accurate quantification of CH4emissions is essential for climate change mitigation; yet, current national and global inventories remain highly uncertain (Saunois, Stavert et al 2020). As identified by Saunois, Martinez et al 2024, increasing the density of CH4observations at local to regional scales is among the top five priorities for reducing uncertainty in the global CH4budget.
2van Hove et al. To address this challenge, we investigate the use of drones to collect local CH4measurements, enabling the estimation of point-source emissions. Several airborne techniques have been developed for quantifying CH4emission rates. The tracer flux ratio method involves co-releasing a tracer gas with a known flow rate and comparing its downwind dispersion to that of CH4. The mass balance method integrates CH4concentrations and wind data across a defined control volume to calculate the net flux. Both methods, which are compared in Lunt et al 2025 as well as Scheutz et al 2025, require prior knowledge of the source location. Inverse modeling techniques are able to infer multiple uncertain source parameters by fitting an atmospheric dispersion model to measured atmospheric data. We use a Bayesian inverse modeling approach, which offers a probabilistic framework for jointly estimating both the emission rate and the source location, while explicitly accounting for measurement uncertainty and model error (e.g., Hutchinson, Liu and Chen 2019b; Park, An et al 2021; van Hove, Aalstad, Lind et al 2025). Drone-based data collection offers spatial flexibility but is resource-intensive. Acquiring informative measurements is particularly difficult in environments with uncertain source locations and dynamic atmospheric conditions (Francis et al 2022). These constraints underscore the need for intelligent sampling strategies that maximize the informational value of each flight. In this study, we examine the optimization of drone flight paths in unfamiliar environments to improve the estimation of both emission rates and source locations from single-point CH4sources. We compare three path-planning strategies for drone-based CH4monitoring: (i) a preplanned lawnmower pattern, (ii) an adaptive short-sighted (myopic) policy that selects actions based on immediate information gain, and (iii) an adaptive long-term (non-myopic) policy trained using deep reinforcement learning. While similar strategies have been explored in previous work (e.g., Loisy and Eloy 2022; Loisy and Heinonen 2023; Park, Ladosz et al 2022; van Hove, Aalstad and Pirk 2023; van Hove, Aalstad and Pirk 2024), prior studies often rely on synthetic data generated from simplified models, such as discretized observation spaces or additive white Gaussian noise assumptions, that may not adequately capture the complexity and variability of real-world atmospheric conditions. In contrast, we generate synthetic observations using an observation model calibrated to replicate CH4dispersion from a point source under Arctic conditions. This calibrated model is referred to as the nature run, following the concept used in observing system simulation experiments (OSSEs) (Masutani et al 2010). In OSSEs, a nature run represents a high-resolution atmospheric simulation that acts as a surrogate for the ‘true’ atmospheric state, enabling the generation of realistic synthetic observations for testing hypothetical observing systems. Our approach adopts this principle but grounds it in field data. Specifically, we tune the observation model using measurements from a drone campaign in Svalbard. Physical parameters are derived directly from field observations, while observation error characteristics are estimated through statistical calibration methods, including maximum likelihood estimation and Bayesian model comparison. The resulting nature run provides a controlled, data-driven environment from which synthetic observations are generated to evaluate the different path-planning strategies under varied conditions, such as source location, emission strength, and wind conditions. To our knowledge, this is the first study to use a data-driven nature run to train and assess adaptive drone sampling strategies for joint estimation of source location and emission rate. Deep reinforcement learning has rarely been applied to this problem (e.g., Lee et al 2025; Park, Ladosz et al 2022), and its use with information gain as a reward signal remains especially uncommon—making our approach both novel and technically challenging. We further explore the role of drone collaboration by comparing single-agent systems (i.e., one drone) with multi-agent systems (i.e., multiple drones). Specifically, we assess whether centralized, coordinated exploration improves the accuracy of source term estimates compared to independent operation. Our objectives are threefold: (i) to generate more realistic synthetic data informed by realworld observations, (ii) to evaluate whether adaptive sampling improves estimation accuracy over preplanned baseline strategies, and (iii) to explore whether collaboration between two drones yields
Data/Math 3 additional benefits over independent operation. Ultimately, our goal is to support the development of more reliable emission inventories and inform targeted mitigation strategies at local to regional scales by improving the accuracy and efficiency of CH4source term estimation methods. 2. Methods We present a framework for estimating the location and emission rate of a single CH4source using drone-based observations. Our approach consists of two components: (i) Source Term Estimation (STE) (Hutchinson, Oh et al 2017), which applies Bayesian inverse modeling to infer the source parameters from CH4concentration and wind data; and (ii) Adaptive Informative Path Planning (AIPP) (Popović et al 2024), which optimizes drone trajectories to collect the most informative observations in a limited flight time. A schematic overview of the framework and its methodologies is provided in Figure 1. 2.1. Source Term Estimation Estimating the location and emission rate of a CH4source from drone observations is formulated as an inverse modeling problem. The source parameters of interest cannot be observed directly, but must be inferred from indirect observations of CH4concentration and wind. This relationship is expressed by the observation model: 𝒅=F (𝜽,𝜻) + 𝝐,(2.1) where 𝒅denotes the observed data, 𝜽represents the uncertain source parameters to be estimated, and 𝜻includes parameters assumed to be known. The forward model F (·) maps these inputs to predicted observations, while the residual term 𝝐accounts for measurement noise, including effects of flow disturbances caused by the drone (e.g., downwash), errors in the assumed known parameters, limitations of the forward model, and representation errors (van Leeuwen 2015). This inverse problem is inherently non-linear and ill-posed, and is further complicated by sparse, noisy, and intermittent data due to atmospheric turbulence and limited sampling. We adopt a probabilistic framework based on Bayesian inference (Jaynes 2003) to invert the observation model, Eq. (2.1), and obtain an uncertainty-aware estimate of the uncertain source parameters 𝜽(Sanz-Alonso et al 2023). 2.1.1. Observation model The observational data vector 𝒅𝑡=[𝑐(𝒙, 𝑡), 𝑣(𝒙, 𝑡), 𝜙(𝒙, 𝑡)] consists of instantaneous measurements of CH4concentration 𝑐(𝒙, 𝑡) [ppm], horizontal wind speed 𝑣(𝒙, 𝑡) [m s−1], and wind direction 𝜙(𝒙, 𝑡) [◦], recorded at location 𝒙=[𝑥, 𝑦, 𝑧] [m]and time step 𝑡[−]. Our goal is to infer the mean emission rate 𝑄[kg h−1]and the horizontal source location 𝒙s=[𝑥s, 𝑦s] [m]. Additionally, we infer two nuisance parameters (Jaynes 2003): the mean horizontal wind speed 𝑉[m s−1]and the mean wind direction Φ[◦], which are not of primary interest but improve the robustness of the inference. The full vector of uncertain parameters that we seek to infer is thus 𝜽=[𝑄, 𝑥s, 𝑦s, 𝑉, Φ]. To facilitate probabilistic inference, we model the residuals between observations and model predictions as Gaussian noise. However, real-world CH4concentration and wind data often exhibit non-normal error distributions. To address this mismatch, we apply Box-Cox transformations (Box and Cox 1964) within the observation model. These transformations, parameterized by 𝜆, map non-normal data from physical space into a transformed observation space where the residuals more closely approximate normality. Separate transformations are applied to CH4concentration, wind speed, and wind direction, each with its own optimal 𝜆and transformed observation error standard deviation 𝜎(𝜆). To construct the
4van Hove et al. prior 𝑝(𝜽|𝒅0:𝑡) •emission rate 𝑄 •source location [𝑥s,𝑦s] •wind speed 𝑉 •wind direction Φ prediction F (𝜽,𝜻) •concentration 𝐶(x) from plume model Eq. (2.3) •wind speed 𝑉 •wind direction Φ Bayes’ rule approximated by a particle filter and Markov chain Monte Carlo steps posterior 𝑝(𝜽|𝒅0:𝑡+1) •emission rate 𝑄 •source location [𝑥s,𝑦s] •wind speed 𝑉 •wind direction Φ observation 𝒅𝑡+1 •concentration 𝑐(𝒙, 𝑡 +1) •wind speed 𝑣(𝒙, 𝑡 +1) •wind direction 𝜙(𝒙, 𝑡 +1) likelihood 𝑝(𝒅(𝜆) 𝑡+1|𝜽) with Gaussian distribution Eq. (2.7) in the transformed observation space reward 𝑟𝑡+1 information gain Eq. (2.13) state 𝑠𝑡+1 •agent location •belief for the non-myopic policy: •grid cell visits action 𝑎𝑡 north OR south OR east OR west OR up OR down OR stay policy preplanned OR myopic OR non-myopic Q-learning state-action values approximated by a neural network and updated via Eq. (2.12) B-C B-C epsilon-greedy greedy ENVIRONMENT MODEL AGENT Figure 1. Schematic overview of the developed framework, consisting of two main components: Source Term Estimation (top box, Section 2.1) and drone path planning strategies (bottom box, Section 2.2). Source estimation uses Bayesian inverse modeling in the Box-Cox transformed observation space, where ‘B-C’ denotes the Box-Cox transformation. Bayesian updating is sequential in the synthetic experiments, with each posterior becoming the prior for the next step. Path planning is guided by the agent (i.e., drone) policy: preplanned (deterministic), myopic (greedy one-step lookahead), or non-myopic (greedy via a trained state-value function). Dashed lines indicate components used only during reinforcement learning training, where a neural network approximates state-action values using epsilon-greedy exploration. . nature run, these parameters are calibrated using real-world drone measurements to replicate realistic measurement noise (see Section 3.1 for their reported values).
Data/Math 5 We define the transformed observation model as a system of equations evaluated at each time step 𝑡: 𝒅(𝜆) 𝑡= 𝑐(𝒙, 𝑡)(𝜆)=𝐶(𝒙)(𝜆)+𝜖(𝜆) 𝐶, 𝑣(𝒙, 𝑡)(𝜆)=𝑉(𝜆)+𝜖(𝜆) 𝑉, 𝜙(𝒙, 𝑡)(𝜆)= Φ(𝜆)+𝜖(𝜆) Φ, (2.2) where 𝐶(𝒙) [ppm]denotes the temporally averaged CH4concentration at location 𝒙, and 𝝐(𝜆)= [𝜖(𝜆) 𝐶, 𝜖 (𝜆) 𝑉, 𝜖 (𝜆) Φ]represents additive observational and model errors in the transformed space. We assume that these transformed residuals follow a multivariate normal distribution: 𝝐(𝜆)∼ N (0,R𝜆)where R𝜆 is a diagonal covariance matrix encoding the standard deviations of the transformed observation errors. The mean concentration field 𝐶(𝒙)is modeled using an advection–dispersion formulation from Vergassola et al 2007 that simulates particle transport under turbulent atmospheric conditions: 𝐶(𝒙)=𝛼𝑄 4𝜋𝐾 |𝒙−𝒙s|exp −(𝑥−𝑥s)𝑉sin(Φ) 2𝐾 exp −(𝑦−𝑦s)𝑉cos(Φ) 2𝐾exp − | 𝒙−𝒙s| L+𝐶0,(2.3) where 𝐶0[ppm]is the background CH4concentration, and 𝐾[m2s−1]represents the mean effective diffusivity, which comprises both turbulent diffusivity and the generally much smaller molecular diffusivity (Stull 1989). The conversion factor 𝛼incorporates the ideal gas law using air temperature 𝑇and atmospheric pressure 𝑃to facilitate unit conversion between mass concentration [kg m−3]and mixing ratio [ppm], as well as between temporal units. Additionally, 𝛼includes a factor of 2to account for CH4mass conservation, assuming a ground-level source with dispersion over flat terrain. The effective dispersion length scale L [m]is defined as: L=√︄𝐾𝜏 1+𝑉2𝜏 4𝐾 ,(2.4) where 𝜏[s]is the atmospheric lifetime of CH4. The vector of assumed known input parameters to the forward model is 𝜻=[𝒙, 𝑧s, 𝜏, 𝐾, 𝐶0, 𝑇, 𝑃], where 𝑧s=0m is the source height and 𝜏=9.1years (Prather et al 2012) is the atmospheric lifetime of CH4. The remaining fixed parameters used to construct the nature run are derived from the real-world drone experimental data (see Section 3.1). 2.1.2. Bayesian inference We adopt a sequential Bayesian inference (also known as filtering) approach (Murphy 2023; Särkkä and Svensson 2023) to estimate the uncertain parameters 𝜽from indirect observations 𝒅, integrating prior knowledge with new observational data and quantifying uncertainty in the estimates. In our setting, inference is performed sequentially: each new set of observations—comprising CH4 concentration, wind speed, and wind direction—triggers a single Bayesian update, allowing the posterior distribution to evolve over time as more data are collected. At each new time step 𝑡+1, the prior distribution over the unknown parameters 𝑝(𝜽|𝒅0:𝑡)is updated to a posterior distribution 𝑝(𝜽|𝒅0:𝑡+1) conditioned on the new observation(s) 𝒅𝑡+1via Bayes’ rule: 𝑝(𝜽|𝒅0:𝑡+1)=𝑝(𝒅𝑡+1|𝜽)𝑝(𝜽|𝒅0:𝑡) 𝑝(𝒅𝑡+1|𝒅0:𝑡).(2.5)
6van Hove et al. Table 1. Uniform prior bounds, training/test scenario bounds, and true values used in the illustrative run for the five uncertain parameters. These parameters define the search space for Bayesian inference and reinforcement learning-based path planning strategies. Parameter Prior bounds Training/test bounds Illustrative run 𝑄[kg h−1][0, 13] [2.6, 10.3] 3.9 𝑥s[m][0, 300] [60, 240] 150 𝑦s[m][0, 300] [30, 210] 66 𝑉[m s−1][0, 6] [1, 5] 2.7 Φ[◦][0, 360] [140, 220] 200 In our plume model, the uncertain parameters 𝜽are treated as constants over time. Therefore, we apply particle filtering under the assumption of a static model (Chopin 2002; van Hove, Aalstad, Lind et al 2025). The denominator is the model evidence (also known as marginal likelihood) (MacKay 2003), acting as a normalizing constant: 𝑝(𝒅𝑡+1|𝒅0:𝑡)=∫𝑝(𝒅𝑡+1|𝜽)𝑝(𝜽|𝒅0:𝑡)d𝜽.(2.6) In this sequential framework, the posterior at time step 𝑡serves as the prior at time step 𝑡+1. For notational convenience, we define 𝒅0as an empty set, so that 𝑝(𝜽|𝒅0)denotes the initial prior. We assume uniform priors over all elements of 𝜽. The prior bounds, specified in Table 1, are deliberately chosen to encompass the bounds of the training and testing domains to avoid biasing the inference process. Additionally, we assume some prior knowledge of the prevailing wind direction, which informs the design of the measurement domain. Specifically, we position the measurement domain slightly downwind of the region where the source is expected to be located, increasing the likelihood of plume detection. The likelihood term 𝑝(𝒅𝑡+1|𝜽)quantifies the agreement between the observed data and model predictions. In our implementation, the likelihood is modeled as a multivariate Gaussian in Box-Cox transformed observation space (see Appendix Band Eq. (B.3)) following Smith et al 2010: 𝑝(𝒅(𝜆) 𝑡+1|𝜽)=N (𝒅(𝜆) 𝑡+1| F (𝜽,𝜻)(𝜆),R𝜆),(2.7) where F (𝜽,𝜻)(𝜆)is the Box-Cox transformed forward model prediction, and the observation error covariance R𝜆is a diagonal matrix containing the square of the error standard deviations 𝜎(𝜆)estimated via calibration against field data (see Section 2.4). The transformed likelihood (Eq. (2.7) and Eq. (B.3)) and evidence can be directly substituted into Eq. (2.5), as the Jacobian of the transformation term cancels out owing to its independence from 𝜽(Jaynes 2003). 2.1.3. Particle filter In most real-world scenarios, exact Bayesian inference is computationally intractable due to the complexity and high dimensionality of the posterior distributions. We adopt a Sequential Monte Carlo (SMC) approach (Chopin and Papaspiliopoulos 2020; Särkkä and Svensson 2023) to approximate the posterior distribution of the uncertain parameters 𝜽. The prior is represented by a weighted ensemble of particles: {(𝜽(𝑗), 𝑤(𝑗) 𝑡)}𝑁𝑒 𝑗=1,(2.8)
Data/Math 7 where 𝜽denotes the state of particle 𝑗,𝑤𝑡represents its corresponding weight at time step 𝑡(reflecting its probability given data up to time step 𝑡), and 𝑁𝑒is the total number of particles in the ensemble. The weights are self-normalized such that Í𝑁𝑒 𝑗=1𝑤(𝑗) 𝑡=1(Murphy 2023). At each iteration, the particle weights are updated using importance sampling based on the transformed likelihood, Eq. (2.7), of the new observation: ¯𝑤(𝑗) 𝑡+1∝𝑝(𝒅(𝜆) 𝑡+1|𝜽(𝑗))𝑤(𝑗) 𝑡.(2.9) The weights ¯𝑤(𝑗) 𝑡+1are subsequently self-normalized to approximate the posterior 𝑝(𝜽|𝒅0:𝑡+1). A common issue in particle filtering is degeneracy, where a small number of particle paths dominate the weight distribution (Murphy 2023), leading to poor representation of the posterior distribution. To mitigate this, we deploy the resample-move algorithm (Doucet and Johansen 2009; Gilks and Berzuini 2001), which combines resampling with a Markov Chain Monte Carlo (MCMC) move step based on the Metropolis-Hastings (MH) algorithm. To balance exploration and exploitation, we employ an adaptive step size for the MCMC proposals. The step size is set to the bandwidth of the dynamic prior particle distribution from time 𝑡, computed using Silverman’s rule, Eq. (D.2). This adaptive proposal distribution helps prevent premature convergence and supports robust posterior estimation. Reflective boundary conditions are applied to ensure that proposed particles remain within the prior bounds. Online path planning requires data processing during collection, necessitating fast updates. As such, we use 𝑁𝑒=1,000 particles and a single MCMC move step per iteration, updating the posterior after each new observation 𝒅𝑡+1. 2.2. Adaptive Informative Path Planning Optimizing the drone’s flight path for STE is formulated as a sequential decision-making problem under uncertainty. At each time step, the drone must decide where to move next in order to collect the most informative observations. We model this task as a Partially Observable Markov Decision Process (POMDP) (Kochenderfer et al 2022), where the agent, i.e., the drone, interacts with the environment through a cycle of actions, observations, and belief updates. In this framework, the agent does not have direct access to the true state of the environment 𝜽, but instead maintains a belief over the uncertain parameters, updated via the particle filter described in Section 2.1.3. The belief state at time step 𝑡, denoted 𝑠𝑡, includes a representation of the distribution over the source parameters 𝑝(𝜽|𝒅0:𝑡), as well as assumed known parameters such as the drone’s current position 𝒙𝑡. At each time step, the agent selects an action 𝑎𝑡from a discrete set of movement options in a threedimensional domain: one step in any cardinal direction (north, south, east, west), vertical movement (up, down), or remaining stationary. After executing the action, the agent receives new observations 𝒅𝑡+1, a numerical reward 𝑟𝑡+1, and transitions to a new belief state 𝑠𝑡+1. The episode terminates after a fixed number of steps, simulating battery life. The agent’s objective is to maximize the expected cumulative reward, defined as the sum of discounted future rewards over a finite horizon: 𝐺𝑡=𝑟𝑡+1+𝛾𝑟𝑡+2+𝛾2𝑟𝑡+3+...= 𝑡final ∑︁ 𝑘=𝑡+1 𝛾𝑘−𝑡−1𝑟𝑘,(2.10) where 𝛾∈ [0,1]is the discount factor controlling the agent’s planning horizon and 𝑡final is the final time step of the episode. In our setting, the reward function is based on information gain—specifically, the
8van Hove et al. reduction in entropy of the posterior relative to the dynamic prior (see Section 2.3). This encourages the agent to collect observations that most effectively reduce uncertainty. To explore different planning horizons, we evaluate both myopic (short-sighted) and non-myopic (long-term) policies, as described in the following section. 2.2.1. Myopic and non-myopic policies The agent’s decision-making is governed by a policy 𝜋, which maps belief states to actions. We consider two classes of policies that differ in their planning horizon: myopic and non-myopic. In a myopic policy, the agent selects actions based solely on the expected immediate reward. This corresponds to setting the discount factor 𝛾=0in Eq. (2.10), reducing the return to 𝐺𝑡=𝑟𝑡+1. In contrast, a non-myopic policy considers the long-term consequences of actions by using a discount factor 𝛾close to 1, thereby incorporating future rewards into the decision-making process. To guide action selection, we estimate the action-value function 𝑞𝜋(𝑠, 𝑎)=E𝜋[𝐺𝑡|𝑠𝑡=𝑠, 𝑎𝑡=𝑎], which represents the expected return when taking action 𝑎in belief state 𝑠and following policy 𝜋 thereafter. A greedy agent selects the action that maximizes the function: 𝜋(𝑠)=arg max𝑎𝑞(𝑠, 𝑎). The optimal action-value function 𝑞∗(𝑠, 𝑎)satisfies the Bellman optimality equation: 𝑞∗(𝑠, 𝑎)=Eh𝑟𝑡+1+𝛾max 𝑎′𝑞∗(𝑠𝑡+1, 𝑎′) | 𝑠𝑡=𝑠, 𝑎𝑡=𝑎i.(2.11) In practice, computing this expectation exactly is infeasible due to the high dimensionality of the belief space and the stochastic nature of observations. We therefore rely on sample-based approximations. For the myopic policy with 𝛾=0, the Bellman optimality equation simplifies to: 𝑞∗(𝑠, 𝑎)= E[𝑟𝑡+1|𝑠𝑡=𝑠, 𝑎𝑡=𝑎], which we estimate using Monte Carlo sampling. Specifically, we draw samples from the current belief state, simulate the outcome of each candidate action, and compute the expected immediate reward. To maintain computational efficiency during online decision-making, we limit the number of samples to one per candidate action, trading off accuracy for speed. In contrast, the non-myopic policy requires estimating the long-term value of actions. We approximate the action-value function using a parameterized function ˆ𝑞(𝑠, 𝑎;𝝎), implemented as a deep neural network with weights 𝝎. The training of this network through deep reinforcement learning is described in the following section. 2.2.2. Reinforcement learning To implement the non-myopic policy, we use reinforcement learning (Sutton and Barto 2018), specifically deep Q-learning (Mnih et al 2015), for the approximation of the optimal action-value function 𝑞∗(𝑠, 𝑎). This is achieved using a feedforward neural network parameterized by weights 𝝎, denoted by ˆ𝑞(𝑠, 𝑎;𝝎). The network is implemented in Python using TensorFlow (Martín Abadi et al 2015). The architecture consists of three fully connected hidden layers with 512, 256, and 128 units, respectively, each using Rectified Linear Unit (ReLU) activation functions. The output layer is linear, with one node per action, i.e., seven nodes for navigation in a 3D domain, representing the estimated action values ˆ𝑞(𝑠, 𝑎;𝝎). The input to the network includes a discretized belief map over the source location and wind direction, the drone’s current position, and a representation of previously visited observation locations within the domain (see Appendix E). Training is conducted over 25,000 iterations, with 12 mini-batch updates per iteration. Each update uses 32 transitions sampled uniformly from a replay buffer of 10,000 transitions. The Q-values are updated using the standard Q-learning rule: ˆ𝑞(𝑠𝑡, 𝑎𝑡;𝝎) ← ˆ𝑞(𝑠𝑡, 𝑎𝑡;𝝎) + 𝜂h𝑟𝑡+1+𝛾max 𝑎ˆ𝑞target (𝑠𝑡+1, 𝑎;𝝎−) − ˆ𝑞(𝑠𝑡, 𝑎𝑡;𝝎)i,(2.12)
Data/Math 9 where 𝜂=0.0001 is the learning rate, optimized using the Adam optimizer (Diederik P. Kingma 2015). The target Q-value is computed using a separate target network ˆ𝑞target, which is synchronized with the main network every 500 iterations to improve training stability. Exploration is managed using an 𝜖-greedy policy. The exploration rate 𝜖starts at 1.0and decays toward 0.01 over the course of training, encouraging exploration early on and exploitation of learned strategies later. While training was conducted for 25,000 episodes, we observed that cumulative reward sometimes deteriorated during the final stages of training. To ensure robust evaluation, we then selected the policy from an earlier, better-performing iteration. 2.2.3. Centralized multi-agent system In a centralized multi-agent framework, multiple agents collaboratively maintain a joint belief over the environment. Since agents’ actions are interdependent, optimal decision-making requires selecting joint actions from the combined action space, which grows exponentially with the number of agents. To manage this complexity, we restrict our study to two agents operating in a two-dimensional domain. This simplification reduces the size of the joint action space and makes centralized coordination computationally feasible for our synthetic experiments. In the collaborative setting, the joint belief is updated sequentially at each time step using the observational data from each agent, resulting in two inference steps, one for each agent. For comparison, we also evaluate a non-collaborative scenario in which each drone follows its own trajectory and maintains an independent belief. In addition, we explore an alternative inference approach where the joint belief is updated in a single inference step using aggregated observational data. Specifically, at each time step, the observations from both agents are concatenated into a single vector and assimilated jointly, rather than sequentially. Training of the centralized multi-agent RL (MARL) (Albrecht et al 2024) policy with joint assimilation is conducted over three training sessions, each comprising 25,000 iterations. In the first session, the exploration rate 𝜖decays from 1.0to 0.01. In the second and third sessions, 𝜖decays from 0.5to 0.01. Subsequently, the policy with sequential assimilation is initialized from the joint assimilation policy and retrained over an additional 25,000 iterations, with 𝜖decaying from 0.5to 0.01. 2.3. Information entropy Our primary objective is to obtain accurate and precise estimates of the source term parameters. To this end, we define a reward function based on information entropy H (·), which quantifies the uncertainty associated with these parameters. This approach incentivizes actions that reduce uncertainty in the agent’s belief state. Entropy reflects how uncertain we are about a random variable; the higher the entropy, the less we know about its value (MacKay 2003). We prioritize entropy over variance, as variance can be misleading, particularly in multi-modal distributions where high variance may still correspond to informative beliefs. Entropy is maximal for uniform distributions, such as our priors, and decreases as the distribution becomes more constrained. We define the reward at each time step as the reduction in entropy between successive belief states, i.e., information gain (Lindley 1956): 𝑟𝑡+1=H (𝑠𝑡) − H (𝑠𝑡+1).(2.13) For continuous distributions, entropy generalizes to differential entropy, defined as: H (𝑠)=−∫𝑝(𝜽)log 𝑝(𝜽)d𝜽,(2.14)
16 van Hove et al. Non-myopic policy The non-myopic flight path is governed by an RL policy trained to maximize long-term information gain. In the illustrative run, the agent initially remains downwind at an altitude of 6m, traveling approximately crosswind relative to the inferred prevailing wind. It then descends and approaches the source, including upwind sampling, which further reduces uncertainty in the source location. Compared to the myopic and lawnmower strategies, the non-myopic policy yields higher cumulative information gain (9.4nats) and produces more precise posterior estimates of both the emission rate and source location, as reflected in lower CRPS values. 3.2.2. Performance statistics Table 2summarizes key performance metrics (defined in Section 2.5.1), including entropy reduction, CRPS scores for emission rate and source location, and normalized CRPS. All results are reported as medians, accompanied by the 5th-95th percentile range to indicate variability across the testbed. Although agents are trained to maximize information gain, our primary evaluation criterion is posterior predictive performance, quantified using the CRPS. Unlike entropy-based metrics, which quantify belief constraint, CRPS directly reflects the quality of posterior predictions in terms of accuracy and precision (Hersbach 2000). Figure 6presents a pairwise comparison matrix showing the proportion of 25,000 testbed simulations in which each strategy achieves a lower CRPSnorm, signifying superior predictive performance compared to its counterparts. Failure to learn Across all flight strategies and durations, entropy reduction ΔHacross testbed simulations deviates from a Gaussian distribution, instead exhibiting a bimodal structure. One broad peak reflects substantial information gain, while a sharp peak near 0–1nats corresponds to minimal learning. The height of this sharp peak varies by strategy and flight duration. For the lawnmower strategy with 108 steps, 22% of flights fall within this low-information category defined by ΔH<1. The myopic policy shows increasing rates of low-information flights with shorter durations: 22%,44%,70% for 108, 54, and 27 steps, respectively. In contrast, non-myopic policies yield substantially lower rates: 2%,6%, and 15%. Comparison of flight strategies Across all flight durations, the non-myopic policy consistently outperforms both lawnmower and myopic strategies. With 108 steps, it achieves the highest information gain (8.9nats) and the lowest median CRPS values for both emission rate (0.6kg h−1) and source location (10 m). These results are further supported by the high outperformance fractions shown in Figure 6. The lawnmower strategy shows limited learning, with a low entropy reduction of 2.2nats over 108 steps and comparatively high CRPS values. The myopic policy performs moderately well, improving over the lawnmower strategy but falling short of the non-myopic performance in all median performance metrics. Due to computational constraints in online adaptive path planning, the myopic policy was executed using a single Monte Carlo sample per candidate action. To assess the potential for improvement, we also evaluated the myopic strategy with five Monte Carlo samples per candidate action for the 108-step scenario. This increased ΔHfrom 5.5nats to 5.9nats and reduced CRPSnorm from 0.62 to 0.54. While this demonstrates a modest improvement, it remains inferior to the performance of the non-myopic strategy. Although increasing the number of Monte Carlo samples improves policy performance, such an approach is computationally prohibitive in online deployment. In contrast, non-myopic policies require substantial resources for offline training, but evaluating the trained neural network during online execution is computationally efficient.
Data/Math 17 Table 2. Performance metrics for the lawnmower, myopic, and non-myopic flight strategies across 3D and 2D domains, and episode lengths of 108, 54, and 27 steps. Metrics include: (i) cumulative information gain or entropy reduction (ΔH=H (𝑠0) − H (𝑠final)), (ii) Continuous Ranked Probability Score (CRPS) for the emission rate (CRPS𝑄), (iii) CRPS for the Euclidean distance from the true source location (CRPS𝒙s), and (iv) CRPS ratio of the posterior compared to the initial prior, averaged over both 𝑄and 𝒙s(CRPSnorm). Results are aggregated over 25,000 synthetic test scenarios and reported as median values, followed by the 5th and 95th percentiles in between square brackets. Lower CRPS values indicate better performance. ΔH[nats] CRPS𝑄[kg h−1] CRPS𝒙s[m] CRPSnorm [−] Single-agent |3D domain 108 steps |illustrative run Lawnmower 2.8[1.1,4.0]0.8[0.5,1.7]41 [29,79]0.46 [0.32,0.83] Myopic 4.5[0.2,9.9]0.7[0.2,2.2]38 [4,114]0.44 [0.09,1.14] Non-myopic 9.8[7.0,11.2]0.2[0.1,2.1]4[2,53]0.10 [0.05,0.84] 108 steps Lawnmower 2.2[0.3,6.0]1.1[0.6,3.2]53 [15,109]0.73 [0.31,1.50] Myopic 5.5[0.2,12.0]1.0[0.2,3.4]32 [5,109]0.62 [0.14,1.56] Non-myopic 8.9[4.6,14.9]0.6[0.1,3.6]10 [2,85]0.29 [0.08,1.60] 54 steps Myopic 2.0[0.1,9.6]1.2[0.3,3.4]74 [11,122]0.88 [0.22,1.51] Non-myopic 8.2[0.6,14.1]0.8[0.2,3.7]14 [3,104]0.41 [0.10,1.64] 27 steps Myopic 0.3[0.1,6.8]1.3[0.5,3.0]90 [26,126]0.97 [0.40,1.42] Non-myopic 7.5[0.2,14.1]1.1[0.2,4.0]28 [4,118]0.64 [0.13,1.75] Multi-agent |2D domain |27 steps No collaboration Myopic 1.2[0.2,8.4]1.2[0.4,3.4]77 [13,122]0.89 [0.26,1.54] Non-myopic 7.8[1.0,14.0]0.8[0.2,3.7]19 [4,102]0.45 [0.11,1.67] Collaboration Myopic 0.9[0.1,7.7]1.3[0.4,3.4]79 [15,123]0.90 [0.28,1.54] Non-myopic 7.4[0.7,13.6]0.9[0.2,3.8]18 [4,102]0.46 [0.12,1.66]
18 van Hove et al. Lawnmower | 108 steps Myopic | 108 steps Non-myopic | 108 steps Myopic | 54 steps Non-myopic | 54 steps Myopic | 27 steps Non-myopic | 27 steps Lawnmower | 108 steps Myopic | 108 steps Non-myopic | 108 steps Myopic | 54 steps Non-myopic | 54 steps Myopic | 27 steps Non-myopic | 27 steps 0.40 0.22 0.56 0.30 0.69 0.43 0.60 0.32 0.63 0.40 0.72 0.52 0.78 0.68 0.78 0.58 0.84 0.69 0.44 0.37 0.22 0.28 0.61 0.39 0.70 0.60 0.42 0.72 0.79 0.61 0.31 0.28 0.16 0.39 0.21 0.30 0.57 0.48 0.31 0.61 0.39 0.70 3D domain | 1 agent Myopic | 1 agent Non-myopic | 1 agent Myopic | 2 agents | no collaboration Non-myopic | 2 agents | no collaboration Myopic | 2 agents | collaboration Non-myopic | 2 agents | collaboration Myopic | 1 agent Non-myopic | 1 agent Myopic | 2 agents | no collaboration Non-myopic | 2 agents | no collaboration Myopic | 2 agents | collaboration Non-myopic | 2 agents | collaboration 0.31 0.44 0.24 0.45 0.25 0.69 0.64 0.41 0.65 0.44 0.56 0.36 0.28 0.51 0.30 0.76 0.59 0.72 0.73 0.52 0.55 0.35 0.49 0.27 0.28 0.75 0.56 0.70 0.48 0.72 2D domain | 27 steps 0.0 0.2 0.4 0.6 0.8 1.0 Outperformance fraction Figure 6. Annotated matrix illustrating pairwise outperformance fractions between flight strategies and durations. Each cell represents the proportion of simulation runs in which the strategy and duration indicated by the row achieves a lower CRPSnorm than the corresponding configuration in the column. Values are based on 25,000 synthetic testbed runs for each configuration and reflect comparative posterior accuracy and precision across configurations. Impact of flight duration Shorter flight durations lead to degraded performance, particularly for the myopic policy. At 54 steps, entropy reduction ΔHdrops from 5.5nats to 2.0nats, and at 27 steps, it falls further to 0.3nats. The corresponding CRPS values also worsen. Despite reduced durations, the non-myopic agents maintain relatively strong performance. At 27 steps, the non-myopic agent achieves ΔH=7.5nats and CRPSnorm =0.64, matching the predictive performance of the 108-step myopic agent. As shown in Figure 6, the 27-step non-myopic agent outperforms the 27-step myopic agent in 71% of testbed cases, the 54-step myopic agent in 61% of testbed cases, and exceeds the performance of the 108-step myopic agent in 49% of cases. Emission rate vs. source location Normalized CRPS values for emission rate and source location (not shown in Table 2) vary substantially across strategies and flight durations. Across all conditions, normalized CRPS values for emission rate are consistently higher (i.e., worse) than those for source location. Emission rate estimation proves challenging for all strategies. For lawnmower flights with 108 steps, 36% exceeds the worse-than-prior threshold of a normalized CRPS of one. Myopic agents show degradation rates of 34%,43%, and 48% for 108, 54, and 27 steps, respectively. Non-myopic agents perform comparatively better, with 24%, 32%, and 38% of flights exceeding the threshold. Despite this improvement, the rates remain high, indicating difficulty in emission rate estimation. In contrast, source location estimation is notably more robust. Only 10% of lawnmower flights with 108 steps exceed the threshold. Myopic agents show 9%, 20%, and 33% for 108, 54, and 27 steps, respectively. Non-myopic agents again perform best, with lower rates of 4%,6%,11%. These results suggest that the source location is more reliably inferred than emission rate, particularly under shorter flight durations and with non-myopic planning.
Data/Math 19 3.3. Multi-agent framework 3.3.1. Performance statistics Table 2presents performance metrics for multi-agent strategies, showing that the non-myopic strategy outperforms the myopic strategy for the tested scenario. Moreover, within both strategy categories, multi-drone configurations outperform their single-drone counterparts, as illustrated in Figure 6. This suggests that parallel exploration and data collection can enhance inference quality, even without explicit coordination. Interestingly, collaborative strategies—where drones operate under a centralized policy—perform comparably to non-collaborative multi-drone setups. All reported statistics are based on sequential assimilation, where each drone’s observations are processed incrementally, one after the other, at each time step. Sequential vs. joint assimilation To assess the impact of the assimilation method, we also evaluated performance under joint assimilation, where observations from both agents at a given time step are assimilated simultaneously. While this approach typically yields higher cumulative information gain—for example, 9.2nats compared to 7.4nats for collaborative non-myopic agents—it results in poorer predictive performance. CRPSnorm increased from 0.46 under sequential assimilation to 0.55 under joint assimilation for the same test case. This trend of increased CRPSnorm under joint assimilation was consistently observed across all multi-agent configurations. 3.3.2. Illustrative flights To qualitatively demonstrate the impact of observation noise on inference in individual runs, we present selected multi-agent flight trajectories in Figure 7. The figure shows two independent drones following single-agent policies and two drones executing a centralized collaborative policy. In both scenarios, the agents initially follow a crosswind trajectory relative to the prevailing wind before at least one agent navigates toward the inferred source location. In the non-collaborative case, apart from a noisy observation at 𝒙=[30,210]m, concentration measurements align well with the true mean concentration field. Neither drone samples near the source, yet the emission rate is inferred correctly, albeit with uncertainty. The source location is estimated reasonably well, though the agents believe it may be farther upwind than its true position. In the collaborative case, one agent obtains observations close to the source as well as slightly upwind, and the source location is accurately inferred with high precision. Several plume observations exceed the true mean concentration field, and the emission rate is overestimated. The collaborative case achieves considerably higher cumulative information gain (9.8nats versus 3.6nats), which is reflected in more constrained posteriors for both source location and emission rate. However, the accuracy of the inferred emission rate is worse (CRPS𝑄=0.9kg h−1versus 0.3kg h−1). Overall performance metrics across the testbed are comparable for both non-collaborative and collaborative strategies (see Table 2). This example highlights how observation noise can degrade inference quality in specific runs and underscores the importance of sampling elevated concentrations with low noise within the true static plume. Proximity to the source enables agents to more reliably infer its location. Visual inspection of flights across the testbed suggests that, in both collaborative and non-collaborative cases, agents consistently attempt to navigate toward the inferred source.
20 van Hove et al. 0 300 x [m] 300 0 y [m] z = 2 m 0 10 20 Step 0 2 4 6 8 10 Information gain [nats] Cumulative rewards 0 6.5 13 Emission rate [kg h 1] 0.000 0.002 0.004 0.006 0.008 CRPS Q = 0.3 0 150 300 Source x-location [m] 0.0 0.5 1.0 1.5 0 150 300 Source y-location [m] 0.0 0.5 1.0 1.5 CRPS x s = 18 0.0 2.5 5.0 Wind speed [m s 1] 0 1 2 3 4 0 90 180 270 360 Wind direction [ ] 0.00 0.02 0.04 0.06 0.08 2 3 4 5 6 CH4 conc. [ppm] No collaboration 0 300 x [m] 300 0 y [m] z = 2 m 0 10 20 Step 0 2 4 6 8 10 Information gain [nats] Cumulative rewards 0 6.5 13 Emission rate [kg h 1] 0.000 0.002 0.004 0.006 0.008 CRPS Q = 0.9 0 150 300 Source x-location [m] 0.0 0.5 1.0 1.5 0 150 300 Source y-location [m] 0.0 0.5 1.0 1.5 CRPS x s = 2 0.0 2.5 5.0 Wind speed [m s 1] 0 1 2 3 4 0 90 180 270 360 Wind direction [ ] 0.00 0.02 0.04 0.06 0.08 2 3 4 5 6 CH4 conc. [ppm] Collaboration Figure 7. Illustrative flight paths of two drones without collaboration (top), and with collaboration (bottom) using a non-myopic policy. These examples are not representative of typical behavior in either case; they were chosen to highlight the impact of observation noise. The true parameters are 𝜽=[𝑄=3.9kg h−1, 𝑥s=150 m, 𝑦s=66 m, 𝑉 =2.7m s−1,Φ = 200◦]. The expected mean concentration field, following Eq. (2.3), is shown as a background heatmap, including the true source location marked by a cross. The flight paths of the two agents, shown in black and blue, are overlaid, along with the obtained noisy observations depicted by colored nodes: the first two observations are shown as a triangle, the final as a square, and intermediate observations as circles. The cumulative rewards of each flight path are plotted. The cumulative rewards of the other flight paths are shown by a thin line for comparison. Histograms display the uniform prior in light shading, final posterior in dark shading, and true value depicted by a dashed line. 4. Discussion 4.1. Nature run High noise level The data-driven calibration of our observation model resulted in a challenging synthetic test case. The design reflects Arctic atmospheric conditions in a complex terrain with highly variable and turbulent wind conditions. Consequently, the simulated observations exhibit high noise levels, primarily due to turbulence, which can obscure signal patterns and introduce ambiguity into the inference process. These high-noise observations serve as a stress test for the robustness of our framework. Limitations of a single test case The nature run is based on a single scenario, with an estimated emission rate of 𝑄=1.0±0.1kg h−1 under specific meteorological conditions. Although we assume that the noise characteristics observed during this field experiment are broadly indicative of real-world variability, this assumption may not hold across different emission regimes or atmospheric conditions. For example, scenarios with lower emission rates or more stable wind conditions may exhibit different observational uncertainties. The selected scenario represents a specific case, characterized by highly variable wind conditions and a likely non-constant emission rate, and may therefore not be fully representative of other field conditions. Temporal correlations The synthetic observations are generated by independently sampling from the nature run, which does not explicitly account for temporal correlations in the concentration or wind fields. However, turbulent
Data/Math 21 flows are known to exhibit strong spatiotemporal correlations, meaning that the probability of encountering elevated concentrations depends on the agent’s recent observation history. Recent studies using direct numerical simulations of emission plumes (Heinonen et al 2025; Piro, Heinonen, Carbone et al 2025; Piro, Heinonen, Cencini et al 2025) have shown that when turbulent observations contain temporal correlations but the model likelihood assumes conditional independent noise, inference performance can degrade compared to scenarios where the observation model is well specified. Studies find that incorporating temporal correlations—either through correlation-aware Bayesian updates (Heinonen et al 2025) or by leveraging an ensemble of likelihood models to mitigate misspecification (Piro, Heinonen, Cencini et al 2025)—can improve the robustness and accuracy of source localization. 4.2. Single-agent framework Flight paths and performance Flight path visualizations provide insight into the observed performance statistics by revealing how agent trajectories influence information gain. Minimal-learning cases often arise when agents fail to reach the plume region, either due to suboptimal exploration or limited flight duration. Shorter flights, in particular, reduce the opportunity to find the plume and collect informative samples, thereby weakening posterior updates. In some instances, even when the plume region is reached, high observational noise prevents the agent from detecting elevated concentrations, resulting in weak posterior updates. As shown in Figure 5, non-myopic agents are generally more successful than myopic agents at navigating toward the plume and acquiring measurements near and upwind of the source. These strategic observation locations substantially reduce uncertainty in the inferred source location. A central challenge in source inference is the ambiguity that arises when a higher emission rate from a more distant source produces concentration patterns similar to those from a lower emission rate nearby. This ambiguity is more effectively resolved when the agent obtains background concentration measurements just upwind of the source, providing information on where the source is not. Such measurements help constrain both the source location and the emission rate by narrowing the plausible parameter space. Importantly, obtaining low-noise observations remains critical for reliable inference, as measurement noise can introduce bias in emission rate estimates, as illustrated in Figure 7. In contrast, myopic agents tend to remain within high-concentration regions, often failing to explore areas that would improve source localization. Lawnmower agents traverse the plume repeatedly but typically collect fewer informative samples, limiting their ability to constrain the posterior. Inference challenges The results indicate that estimating the emission rate is generally more challenging than localizing the source, particularly during shorter flight durations. This difference may partially stem from the structure of the forward model, Eq. 2.3, where concentration depends exponentially on the agent’s distance to the source but linearly on the emission rate. As a result, the source location has a stronger influence on the observed signal, potentially making the emission rate harder to infer from sparse or limited observations. We identify several factors that contribute to inference challenges. First, high measurement noise can produce misleading observations that obscure the true signal. Second, recalcitrant particle degeneracy, despite rejuvenation efforts, can limit the diversity of hypotheses maintained during inference, reducing robustness. Third, equifinality can complicate parameter estimation: distinct parameter combinations can produce nearly similar measurements, making it difficult to resolve the true configuration. Moreover, we observed that particle rejuvenation may inadvertently widen the posterior when agents encounter uninformative observations, effectively diluting prior learning. This is seen, for instance, under the preplanned (lawnmower) policy.
22 van Hove et al. 4.3. Multi-agent framework Flight paths and performance In the tested single-source configuration, both collaborating and non-collaborating drones exhibit similar flight behavior. Agents tend to scan along the crosswind direction relative to the prevailing wind, rather than disperse. Each drone independently navigates toward regions of elevated concentration near the plume and source. This indicates that collaboration does not significantly alter the search strategy, yielding comparable performance metrics. The main challenge lies in high observational noise, which complicates source inference. While adding a second agent improves robustness, collaboration itself offers no clear advantage—independent operation performs equally well. Clearer performance differences may emerge in more complex scenarios, such as multi-source domains, where collaboration could allow agents to divide the task and focus on distinct sources. Sequential vs. joint assimilation Joint assimilation—processing both agents’ observations simultaneously—yields higher cumulative information gain, yet consistently results in poorer predictive performance, as reflected by elevated CRPS values. This suggests that the belief distribution becomes overconfident and overly constrained, potentially centered around incorrect values. The effect likely arises from high observational noise, equifinality, and reduced particle diversity due to degeneracy. In contrast, sequential assimilation—where each observation is processed in turn with an MCMC move step—achieves lower information gain but better predictive accuracy. Gradual updates help preserve the flexibility of the belief distribution and maintain particle diversity. These findings align with the broader effectiveness of data tempering techniques in particle-based inference (Chopin and Papaspiliopoulos 2020; Murphy 2023). Computational costs Centralized multi-agent systems incur significantly higher computational costs compared to singleagent frameworks. Under a myopic policy, the cost scales exponentially with the number of agents, requiring 𝑛actions ×𝑛Monte Carlo steps𝑛agents look-ahead evaluations. This is quickly becoming infeasible in real-world applications. For non-myopic agents, policy execution remains efficient, but training time increases substantially due to the expansion of the neural network’s output layer to size 𝑛𝑛agents actions. In our case, we increased the training time by a factor of four compared to the single-agent policy. We note that additional training of the neural network, particularly of the multi-agent policies, may yield improved performance. Notably, the collaborative sequential assimilation policy was not trained from scratch, but retrained from the joint assimilation policy. This raises the possibility that training the sequential assimilation policy from scratch could yield different, and potentially superior, performance outcomes. 5. Lessons learned Reality vs. simulation During the setup of the nature run, we found that the effective diffusivity 𝐾was particularly difficult to infer accurately. Consequently, we chose to fix its value. Furthermore, we acknowledge that wind measurements obtained from a moving drone are subject to considerable uncertainty and potential bias, particularly in the vertical component. To improve the accuracy of effective diffusivity estimates, it may be beneficial to install a stationary sonic anemometer near the study site (van Hove, Aalstad, Lind et al 2025). Although limited to a single location, this would likely provide more stable and reliable wind measurements, thereby supporting more accurate modeling of atmospheric dispersion.
Data/Math 23 Entropy computation Accurate entropy computation, based on the particle distribution, proved essential for the success of our framework. Earlier attempts using Eq. (2.16) were unsuccessful, likely due to its inability to capture the true uncertainty represented by the particle ensemble. While the current method yields more reliable results, it is computationally expensive. A potential alternative is to compute entropy discretely by binning particles across the parameter space and estimating entropy from particle counts per bin. This approach could reduce computational costs, though it introduces a trade-off between precision and efficiency. The number of bins must be chosen carefully: more bins improve resolution but increase computational cost, while fewer bins risk oversimplifying the belief structure. Balancing information gain and belief accuracy Using information gain as a reward signal inherently encourages particle degeneracy, as it favors concentrated belief distributions. Rejuvenation methods help mitigate the risk of converging on incorrect values. Achieving a balance between belief constraining and avoiding overconfidence requires careful tuning of the computational implementation of the Bayesian inference process. Key factors include the number of particles, the number and step size of MCMC moves, and the choice between sequential and joint assimilation. These must be managed within the constrained computational budget required for both training and deployment, where fast decision-making is essential. Parameters that perform well in one setting may lead to poor generalization in another. In our implementation, we adopted an adaptive MCMC step size based on the shape of the dynamic prior: larger steps for wide priors and smaller steps for narrow distributions. This allowed us to reduce the particle count to just 1,000 while maintaining reasonable inference quality. However, we observed belief widening in response to uninformative observations, highlighting the sensitivity of the system to observation quality. Belief state representation Representation of the belief state within the neural network was refined through trial and error. The final input configuration, detailed in Appendix E, includes the belief over source location and wind direction, the drone’s current position, and a visitation map that tracks how frequently each cell in the domain has been visited. We also experimented with including the drone’s battery level instead of the visitation map, but this proved less effective. While battery level provides a sense of remaining operational time, it does not convey spatial coverage. Without the inclusion of the visitation map, the agent tended to remain in grid cells with high concentration, neglecting areas at or upwind of the source, which are often helpful for accurate localization. We also tested including belief over emission rate and wind speed in the form of mean raw normalized values. However, these additions did not considerably improve cumulative information gain. As a result, they were excluded from the state representation. 6. Future work More realism in synthetic simulations We used a nature run to create more realistic simulations, but several opportunities remain to enhance realism. First, constructing additional nature runs across varied environments and emission regimes, such as different source strengths and atmospheric conditions, would help refine the likelihood model and assess its generalizability. Second, the forward model could be extended to include more physical processes (Pirk et al 2022), such as vertical wind profiles, diffusivity anisotropy, and temporal correlations inferred via observations or large-eddy simulations, to more realistically simulate the spatiotemporal structure of plume dispersion. Surrogate or reduced-order models may offer a computationally efficient way to incorporate such complexity (e.g., Gunawardena et al 2021; Lumet et al 2025).
24 van Hove et al. Navigational flexibility In the current framework, the source location is modeled in a continuous spatial domain, while the drone operates on a discretized grid. This mismatch limits the flexibility of path planning. Transitioning to a continuous domain would allow agents to select both heading direction and step size, enabling larger exploratory movements early in the flight and more precise steps as the plume is detected. Such flexibility would better reflect real-world drone capabilities and support more adaptive and efficient trajectories. Standard Q-learning is not well suited for continuous action spaces. Park, Ladosz et al 2022 used a Deep Deterministic Policy Gradient (DDPG) algorithm to train an AIPP policy for the STE problem with a continuous state space and an action space defined by heading direction and fixed step size. Future extensions could also introduce structured maneuvers into the action space, such as curtain flights. Instead of always directing the drone to a single point, the policy could optionally select a vertical curtain sweep at a chosen downwind distance from the estimated source location or a horizontal sweep at a selected altitude. This hybrid action space could make the problem more tractable, as it may reduce the number of individual action-selection steps that need to be trained. Notably, the trained non-myopic policy already demonstrates this behavior: the drone typically begins its trajectory with a horizontal crosswind sweep positioned downwind of the source. Multi-agent systems Although our test case did not demonstrate a clear performance advantage for collaborative over non-collaborative strategies, this outcome may be specific to the scenario studied and the methods used. Future work should explore configurations where the benefits of coordination may become more pronounced, such as domains with multiple emission sources, as well as different AIPP strategies, inference algorithms, and belief state representations used as input to the neural network. Notably, Heinonen et al 2025 suggests that multiple collaborating agents may be able to exploit temporal correlations in the flow. Future work can focus on extending the multi-agent framework to incorporate error correlation-aware Bayesian inference updates and memory-based RL policies. Decentralized systems—where drones collaborate and share information but make decisions independently—may offer better scalability and practical feasibility than centralized collaboration. Prior studies, including Park and Oh 2020 and Ristic et al 2020, investigated the performance of decentralized myopic multi-agent systems for the STE problem. Real-world experiments Conducting real-world experiments remains the ultimate goal for validating the proposed framework. Field deployments would enable assessment under true environmental variability, offering insights into robustness and operational feasibility. A central challenge in RL is developing agents that generalize across diverse environments. Atmospheric and terrain conditions can vary significantly between deployment sites, influencing plume behavior. Future work should therefore focus on developing a framework capable of adapting to a wide range of conditions, including different meteorological regimes, terrain types, source characteristics, and ideally, varying number of emission sources. 7. Conclusion This study presents a framework for methane source term estimation using drone-based observations and adaptive informative path planning strategies. By fusing Bayesian inference with reinforcement learning, we demonstrate that adaptive agents (i.e., drones) can efficiently maximize information gain and improve the accuracy of emission rate and source location estimates. Our approach is distinguished by its use of a nature run with a calibrated observation model derived from real-world data, providing a more representative basis for synthetic experiments that capture realistic atmospheric variability. The
Data/Math 25 challenging conditions of the field experiment introduced substantial observational noise, serving as a stress test for the robustness of various flight strategies and durations. Among the tested strategies, non-myopic policies trained via deep reinforcement learning consistently outperformed both myopic and preplanned (lawnmower) approaches across multiple metrics, including entropy reduction and proper scoring rules. These agents demonstrated superior navigation toward informative regions, often reaching the inferred source location and sampling upwind. The non-myopic agents maintained robust performance even for shorter flight durations. We also investigated multi-agent collaboration, finding comparable performance between collaborative and non-collaborative strategies in the tested single-source scenario. Performance differences may become more pronounced in multi-source domains where collaboration could allow agents to divide the task and focus on distinct sources. Sequential assimilation interlaced with rejuvenation offered a more balanced trade-off between belief accuracy and diversity than joint assimilation, which tended to produce overconfident and less reliable estimates due to particle degeneracy. Three components were critical to the success of this framework: (i) accurate entropy computation based on the full particle distribution, rather than particle weights alone, (ii) careful tuning of the computational implementation of the particle-based Bayesian inference process to balance belief constraint with rejuvenation, while maintaining computational efficiency, and (iii) refinement of the belief state representation in the reinforcement learning process to ensure that the agent receives informative and spatially relevant input features. Overall, our findings support the development of probabilistic machine learning methods for intelligent, tractable, and data-driven active inference strategies for drone-based environmental monitoring. By improving the precision, accuracy, efficiency, and reliability of methane source term estimates, these methods can contribute to more accurate emission inventories and inform targeted climate mitigation efforts. A. Businger-Dyer relationship To estimate the effective diffusivity 𝐾, we apply the Businger–Dyer relationship (Stull 1989), a widely used parameterization for turbulent transport in the atmospheric surface layer, to real-world data collected during the field experiment. We estimate the friction velocity 𝑢∗using sonic anemometer measurements at three flight altitudes: 2 m, 4 m, and 6 m. Although 𝑢∗is not altitude-specific, we compute separate estimates from each dataset to account for variability in the wind measurements across altitudes: 𝑢∗=𝑢′𝑤′2+𝑣′𝑤′21/4 ,(A.1) where 𝑢′,𝑣′, and 𝑤′are the turbulent fluctuations in the horizontal and vertical wind components, and the overbars denote time-averaged covariances. The Obukhov length 𝐿is then calculated as: 𝐿=−𝜃𝑣𝑢3 ∗ 𝜅𝑔𝑤′𝜃′ 𝑣 ,(A.2) where 𝜃𝑣is the virtual potential temperature, 𝜅=0.4is the von Kármán constant, and 𝑔=9.81 m s−2 is the gravitational acceleration. Due to limited humidity data, we approximate 𝜃𝑣using potential temperature 𝜃𝑣≈𝜃. A small to moderately positive Obukhov length indicates stable atmospheric stratification, which was observed for the field experiment.
32 REFERENCES Murphy KP (2022) Probabilistic machine learning: an introduction. Adaptive computation and machine learning series. Cambridge, Massachusetts: The MIT Press. 826 pp. isbn: 978-0-262-04682-4. Murphy KP (2023) Probabilistic machine learning: advanced topics. Adaptive computation and machine learning series. Cambridge, Massachusetts: The MIT Press. 1 p. isbn: 978-0-262-37599-3 978-0-262-37600-6. Odland T (2018) tommyod/KDEpy: Kernel Density Estimation in Python ver v0.9.10. doi: 10.5281/zenodo.2392268. Available at https://doi.org/10.5281/zenodo.2392268. Park M, An S, Seo J and Oh H (2021) Autonomous Source Search for UAVs Using Gaussian Mixture Model-Based Infotaxis: Algorithm and Flight Experiments. IEEE Transactions on Aerospace and Electronic Systems 57(6), 4238–4254. issn: 00189251, 1557-9603, 2371-9877. doi: 10.1109/TAES.2021.3098132. Park M, Ladosz P and Oh H (2022) Source Term Estimation Using Deep Reinforcement Learning With Gaussian Mixture Model Feature Extraction for Mobile Sensors. IEEE Robotics and Automation Letters 7(3), 8323–8330. issn: 2377-3766, 2377-3774. doi: 10.1109/LRA.2022.3184787. Park M and Oh H (2020) Cooperative information-driven source search and estimation for multiple agents. Information Fusion 54, 72–84. issn: 15662535. doi: 10.1016/j.inffus.2019.07.007. Pirk N, Aalstad K, Westermann S, Vatne A, Hove A van, Tallaksen LM, Cassiani M and Katul G (2022) Inferring surface energy fluxes using drone data assimilation in large eddy simulations. Atmospheric Measurement Techniques 15, 7293–7314. doi: 10.5194/amt-15-7293-2022. Piro L, Heinonen RA, Carbone M, Biferale L and Cencini M (2025) Policy heterogeneity improves collective olfactory search in 3-D turbulence. doi: 10.48550/arXiv.2504.11291. arXiv: 2504.11291[physics]. Available at http://arxiv.org/abs/2504.11291 (accessed 29 October 2025). Piro L, Heinonen RA, Cencini M and Biferale L (2025) Many wrong models approach to localise an odour source in turbulence with static sensors. Journal of Turbulence, 29 October 26(5), 153–173. issn: 1468-5248. doi: 10.1080/14685248.2025. 2492711.https://www.tandfonline.com/doi/full/10.1080/14685248.2025.2492711 (accessed 29 October 2025). Popović M, Ott J, Rückin J and Kochenderfer MJ (2024) Learning-based methods for adaptive informative path planning. Robotics and Autonomous Systems 179, 104727. issn: 0921-8890. doi: https://doi.org/10.1016/j.robot.2024.104727. Prather MJ, Holmes CD and Hsu J (2012) Reactive greenhouse gas scenarios: Systematic exploration of uncertainties and the role of atmospheric chemistry. Geophysical Research Letters 39(9). issn: 0094-8276, 1944-8007. doi: 10.1029/2012gl051440. Ristic B, Gilliam C, Moran W and Palmer JL (2020) Decentralised multi-platform search for a hazardous source in a turbulent flow. Information Fusion 58, 13–23. issn: 15662535. doi: 10.1016/j.inffus.2019.12.011.https://linkinghub.elsevier.com/ retrieve/pii/S1566253518303592. Sanz-Alonso D, Stuart A and Taeb A (2023) Inverse problems and data assimilation. eng, 1st ed. vol 107. London Mathematical Society student texts ; Cambridge: Cambridge University Press. isbn: 1-009-41433-X. Särkkä S and Svensson L (2023) Bayesian Filtering and Smoothing, 2nd ed. Cambridge University Press. doi: 10.1017 / 9781108917407. Saunois M, Martinez A, Poulter B, Zhang Z, Raymond P, Regnier P, Canadell JG, Jackson RB, Patra PK, Bousquet P, Ciais P, Dlugokencky EJ, Lan X, Allen GH, Bastviken D, Beerling DJ, Belikov DA, Blake DR, Castaldi S, Crippa M, Deemer BR, Dennison F, Etiope G, Gedney N, Höglund-Isaksson L, Holgerson MA, Hopcroft PO, Hugelius G, Ito A, Jain AK, Janardanan R, Johnson MS, Kleinen T, Krummel P, Lauerwald R, Li T, Liu X, McDonald KC, Melton JR, Mühle J, Müller J, Murguia-Flores F, Niwa Y, Noce S, Pan S, Parker RJ, Peng C, Ramonet M, Riley WJ, Rocher-Ros G, Rosentreter JA, Sasakawa M, Segers A, Smith SJ, Stanley EH, Thanwerdas J, Tian H, Tsuruta A, Tubiello FN, Weber TS, Van Der Werf G, Worthy DE, Xi Y, Yoshida Y, Zhang W, Zheng B, Zhu Q, Zhu Q and Zhuang Q (2024) Global Methane Budget 2000–2020. doi: 10.5194/essd-2024-115.https://essd.copernicus.org/preprints/essd-2024-115/. Saunois M, Stavert AR, Poulter B, Bousquet P, Canadell JG, Jackson RB, Raymond PA, Dlugokencky EJ, Houweling S, Patra PK, Ciais P, Arora VK, Bastviken D, Bergamaschi P, Blake DR, Brailsford G, Bruhwiler L, Carlson KM, Carrol M, Castaldi S, Chandra N, Crevoisier C, Crill PM, Covey K, Curry CL, Etiope G, Frankenberg C, Gedney N, Hegglin MI, Höglund-Isaksson L, Hugelius G, Ishizawa M, Ito A, Janssens-Maenhout G, Jensen KM, Joos F, Kleinen T, Krummel PB, Langenfelds RL, Laruelle GG, Liu L, Machida T, Maksyutov S, McDonald KC, McNorton J, Miller PA, Melton JR, Morino I, Müller J, Murguia-Flores F, Naik V, Niwa Y, Noce S, O’Doherty S, Parker RJ, Peng C, Peng S, Peters GP, Prigent C, Prinn R, Ramonet M, Regnier P, Riley WJ, Rosentreter JA, Segers A, Simpson IJ, Shi H, Smith SJ, Steele LP, Thornton BF, Tian H, Tohjima Y, Tubiello FN, Tsuruta A, Viovy N, Voulgarakis A, Weber TS, Van Weele M, Van Der Werf GR, Weiss RF, Worthy D, Wunch D, Yin Y, Yoshida Y, Zhang W, Zhang Z, Zhao Y, Zheng B, Zhu Q, Zhu Q and Zhuang Q (2020) The Global Methane Budget 2000–2017. Earth System Science Data 12(3), 1561–1623. issn: 1866-3516. doi: 10.5194/essd-12-1561-2020. Scheutz C, Knudsen J, Vechi N and Knudsen J (2025) Validation and demonstration of a drone-based method for quantifying fugitive methane emissions. Journal of Environmental Management, 28 October 373, 123467. issn: 03014797. doi: 10.1016/ j.jenvman.2024.123467 (accessed 28 October 2025). Smith T, Sharma A, Marshall L, Mehrotra R and Sisson S (2010) Development of a formal likelihood function for improved Bayesian inference of ephemeral catchments. Water Resources Research 46(12). doi: https://doi.org/10.1029/2010WR009514.
REFERENCES 33 eprint: https://agupubs.onlinelibrary.wiley.com/doi/pdf/10.1029/2010WR009514.https://agupubs.onlinelibrary.wiley.com/ doi/abs/10.1029/2010WR009514. Stull RB (1989) An introduction to boundary layer meteorology. eng, 2nd reprint. Atmospheric sciences library. Dordrecht: Kluwer Academic Publishers. isbn: 9027727686. Sutton RS and Barto AG (2018) Reinforcement Learning: An Introduction. The MIT Press. vanHove A,AalstadK, LindV, ArndtC,OdongoV,Ceriani R,FavaF,HulthJandPirkN (2025) Inferring methane emissions from African livestock by fusing drone, tower, and satellite data. Biogeosciences 22(16), 4163–4186. doi: 10.5194/bg-224163-2025. van Hove A, Aalstad K and Pirk N (2023) Using reinforcement learning to improve drone-based inference of greenhouse gas fluxes. Nordic Machine Intelligence. doi: 10.5617/nmi.9897. van Hove A, Aalstad K and Pirk N (2024) Guiding drones by information gain. In Lutchyn T, Ramírez Rivera A and Ricaud B (eds), Proceedings of the 5th Northern Lights Deep Learning Conference (NLDL). vol 233. Proceedings of Machine Learning Research. PMLR, 89–96. Available at https://proceedings.mlr.press/v233/hove24a.html. van Leeuwen P (2015) Representation errors and retrievals in linear and nonlinear data assimilation. Quarterly Journal of the Royal Meteorological Society 141, 1612–1623. doi: 10.1002/qj.2464. Vergassola M, Villermaux E and Shraiman BI (2007) ‘Infotaxis’ as a strategy for searching without gradients. Nature. doi: 10.1038/nature05464.