scieee AI-readable full text Open interactive document viewer

Supplementary Material of "LUME-DBN: Full Bayesian Learning of DBNs from Incomplete data in Intensive Care"

Pirola, Federico

Abstract

Supplementary Material of "LUME-DBN: Full Bayesian Learning of DBNs from Incomplete data in Intensive Care" presented at the International Joint Workshop of Artificial Intelligence for Healthcare (HC@AIxIA) and HYbrid Models for Coupling Deductive and Inductive ReAsoning (HYDRA): HC@AIxIA+HYDRA 2025 in Bologna, Italy

Full text

1 Supplementary Material A Full Conditional Distributions of Missing Data We consider three configurations where exactly one value is missing at time t= 2:x2 1[MIS],x2 2[MIS],orx2 3[MIS]. The DBN is represented in Figure A.1a and the depicted relationships are meant to represent temporal relationship with a single lag (from time t−1to time t). These settings reflect structurally different roles in the DBN: –X2 1has no parents and one child, –X2 2has one parent and one child, –X2 3has one parent and no children. Fig. A.1. Two examples of DBNs. a) A DBN with 3 temporal nodes and 2 arcs. b) A more complex DBN with 5 temporal nodes, with a node Xt 3with 2 parents {Xt−1 1, Xt−1 2}, 2 children {Xt+1 4, Xt+1 5}and one node with a common children {Xt 4}. Assuming a Uniform prior over the domain of each variable, the FCD of a missing value only depends on the likelihood terms in which that variable appears. Since the samples are independent, missing values for different samples but for the same variable at the same time just lead to the same form of the Gaussian. Consequently: –the FCD for X2 3depends only on its own likelihood: no demonstration is needed because the conditional likelihood of X2 3|X1 2is a properly defined Gaussian distribution with parameters µ3=β0(3) +β1(3)X1 2and σ2(3); –the FCD for X2 1depends on its marginal likelihood and the conditional likelihood of its child X3 2; –the FCD for X2 2depends on both the conditional likelihood given its parent X1 1and the conditional likelihood of its child X3 3. 2 P(x2 1[MIS]| · )∝P(X2 1)·P(X3 2|X2 1) ∝ N(x2 1[MIS]|µ1=β(1) 0, σ2(1))· N (x3 2|µ2=β(2) 0+β(2) 1x2 1[MIS], σ2(2)) ∝exp −(x2 1[MIS]−β(1) 0)2 2σ2(1) !·exp −(x3 2−β(2) 0−β(2) 1x2 1[MIS])2 2σ2(2) ! ∝exp −1 2"(x2 1[MIS])2−2x2 1[MIS]β(1) 0 σ2(1) +(x3 2)2 σ2(2) −2x3 2(β(2) 0+β(2) 1x2 1[MIS]) σ2(2) +(β(2) 0+β(2) 1x2 1[MIS])2 σ2(2) #! ∝exp −1 2"(x2 1[MIS])2 1 σ2(1) +(β(2) 1)2 σ2(2) ! +2x2 1[MIS] −β(1) 0 σ2(1) +−β(2) 1(x3 2−β(2) 0) σ2(2) !+C#! ∝ N x2 1[MIS] µ∗=σ2∗· β(1) 0 σ2(1) +β(2) 1(x3 2−β(2) 0) σ2(2) !, σ2∗= 1 σ2(1) +(β(2) 1)2 σ2(2) !−1  P(x2 2[MIS]| · )∝P(X2 2|X1 1)·P(X3 3|X2 2) ∝ N(x2 2[MIS]|µ2=β(2) 0+β(2) 1x1 1, σ2(2)) · N(x3 3|µ3=β(3) 0+β(3) 2x2 2[MIS], σ2(3)) ∝exp −(x2 2[MIS]−µ2)2 2σ2(2) !·exp −(x3 3−µ3)2 2σ2(3)  ∝exp −1 2"(x2 2[MIS])2 1 σ2(2) +(β(3) 2)2 σ2(3) ! +2x2 2[MIS] −β(2) 0−β(2) 1x1 1 σ2(2) +−β(3) 2(x3 3−β(3) 0) σ2(3) !+C#! ∝ N x2 2[MIS] µ∗=σ2∗· β(2) 0+β(2) 1x1 1 σ2(2) +β(3) 2(x3 3−β(3) 0) σ2(3) !, σ2∗= 1 σ2(2) +(β(3) 2)2 σ2(3) !−1  3 We then consider the DBN depicted in Figure A.1b, and focus on a missing value xt 3[MIS]for variable X3at time t > 0. The Markov blanket of variable Xt 3is constituted of two parents (Xt−1 1,Xt−1 2), two children (Xt+1 4,Xt+1 5), and one variable (Xt 4) sharing a common child. Deriving the FCD for Xt 3in this setting illustrates the procedure in case of more complex Markov blankets, in the direction of demonstrating the tractability of the FCD in the broader case. P(xt 3[MIS]| · )∝P(Xt 3|Xt−1 1, Xt−1 2)·P(Xt+1 4|Xt 3)·P(Xt+1 5|Xt 3, Xt 4) ∝ N(xt 3[MIS]|µ3=β(3) 0+β(3) 1xt−1 1+β(3) 2xt−1 2, σ2(3)) · N(xt+1 4|µ4=β(4) 0+β(4) 3xt 3[MIS], σ2(4)) · N(xt+1 5|µ5=β(5) 0+β(5) 3xt 3[MIS]+β(5) 4xt 4, σ2(5)) ∝exp −(xt 3[MIS]−(β(3) 0+β(3) 1xt−1 1+β(3) 2xt−1 2))2 2σ2(3) ! ·exp −(xt+1 4−(β(4) 0+β(4) 3xt 3[MIS]))2 2σ2(4) ! ·exp −(xt+1 5−(β(5) 0+β(5) 3xt 3[MIS]+β(5) 4xt 4))2 2σ2(5) ! We define µ(4) {−3}:= µ(4) −β(4) 3xt 3[MIS]and µ(5) {−3}:= µ(5) −β(5) 3xt 3[MIS]. Then (xt+1 4−(β(4) 0+β(4) 3xt 3[MIS]))2= (xt+1 4−β(4) 3xt 3[MIS]−µ(4) {−3})2and (xt+1 5−(β(5) 0+ β(5) 3xt 3[MIS]+β(5) 4xt 4))2= (xt+1 5−β(5) 3xt 3[MIS]−µ(5) {−3})2. For j∈ {4,5}, each term (xt+1 j−β(j) 3xt 3[MIS]−µ(j) {−3})2is a squared trinomial. Since we are computing the FCD for xt 3[MIS], we can focus on the terms that involve this quantity. Namely: (xt+1 j−β(j) 3xt 3[MIS]−µ(j) {−3})2∝(xt 3[MIS]β(j) 3)2−2β(j) 3xt 3[MIS](xt+1 j−µ(j) {−3}) Indeed, we could rewrite (xt 3[MIS]−(β(3) 0+β(3) 1xt−1 1+β(3) 2xt−1 2))2= (xt 3[MIS]− µ3)2. Again, we are interested in the terms of the squared binomial involving xt 3[MIS], thus: (xt 3[MIS]−µ3)2∝(xt 3[MIS])2−2xt 3[MIS]µ3 Then: 4 P(xt 3[MIS]| · )∝exp −1 2(xt 3[MIS])2·1 σ2(3) −(2xt 3[MIS])µ3 σ2(3)  ·exp  −1 2 (xt 3[MIS])2· (β(4) 3)2 σ2(4) !−(2xt 3[MIS])  β(4) 3(xt+1 4−µ(4) {−3}) σ2(4)     ·exp  −1 2 (xt 3[MIS])2· (β(5) 3)2 σ2(5) !−(2xt 3[MIS])  β(5) 3(xt+1 5−µ(5) {−3}) σ2(5)     ∝exp −1 2 (xt 3[MIS])2· 1 σ2(3) +(β(4) 3)2 σ2(4) +(β(5) 3)2 σ2(5) !+ −(2xt 3[MIS])  µ3 σ2(3) +β(4) 3(xt+1 4−µ(4) {−3}) σ2(4) +β(5) 3(xt+1 5−µ(5) {−3}) σ2(5)  +C   ∝ N  µ∗=σ2∗·  µ3 σ2(3) +β(4) 3(xt+1 4−µ(4) {−3}) σ2(4) +β(5) 3(xt+1 5−µ(5) {−3}) σ2(5)  , σ2∗= 1 σ2(3) +(β(4) 3)2 σ2(4) +(β(5) 3)2 σ2(5) !−1  We first derived the FCD for a missing value at a specific time in a simple DBN with three nodes. We then illustrated its form for a variable with two parents and two children at a generic time t. Since all FCDs are Gaussian distributions with closed-form parameters, we now present the general case for a missing value on a variable Xiat time t, characterized by nparents π(i)={Xt−1 p1, . . . , Xt−1 pn}and mchildren {Xt+1 c1, . . . , Xt+1 cn}.. The FCD’s of the missing value xt i[MIS]depends on the conditional likelihood of Xiand the conditional likelihoods of its children. As before, we isolate the contribution of xt i[MIS]in its children likelihoods. Specifically, for each child j∈ {Xt+1 c1, . . . , Xt+1 cn}, given the mean term µt+1 jand the linear coefficient β(j) i associated with the Gaussian likelihood of Xj, we define µ(j) {−i} (t+1) =µt+1 j−β(j) ixt i[MIS] We can then derive the missing value FCD as follows: 5 P(xt i[MIS]| · )∝P(Xt i|π(i))·Y j∈{c1,...,cm} P(Xt+1 j|π(j)) ∝ N(xt i[MIS]|µt i=β(i) 0+β(i) 1xt−1 p1+· · · +β(i) nxt−1 pn, σ2(i)) ·Y j∈{c1,...,cm} N(xt+1 j|µt+1 j=µ(j) {−i} (t+1) +β(j) ixt i[MIS], σ2(j)) ∝exp −1 2(xt i[MIS])2·1 σ2(i)−(2xt i[MIS])µt i σ2(i) ·Y j∈{c1,...,cm} exp −1 2 (xt i[MIS])2· (β(j) i)2 σ2(j)!+ −(2xt i[MIS])  β(j) i(xt+1 j−µ(j) {−i} (t+1)) σ2(j)    ∝exp  −1 2 (xt i[MIS])2·  1 σ2(i)+X j∈{c1,...,cm} (β(j) i)2 σ2(j) + −(2xt i[MIS])  µt i σ2(i)+X j∈{c1,...,cm} β(j) i(x(t+1) j−µ(j) {−i} (t+1)) σ2(j) +C   ∝ N  µ∗=σ2∗·  µt i σ2(i)+X j∈{c1,...,cm} β(j) i (xt+1 j−µ(j) {−i} (t+1)) σ2(j) , σ2∗=  1 σ2(i)+X j∈{c1,...,cm} (β(j) i)2 σ2(j)  −1   We can now express the FCD of a generic missing value xt i[MIS]for a variable Xiat time tin the context of a DBN, accounting for an arbitrary number of parents and children: P(xt i[MIS]| ·) = Nµ∗, σ2∗ where: σ2∗=  1 σ2(i)+X j: (Xt i∈π(j)) (β(j) i)2 σ2(j)  −1 , µ∗=σ2∗·  µt i σ2(i)+X j: (Xt i∈π(j)) β(j) ixt+1 j−µ(j) {−i} (t+1) σ2(j) . 6 B Learning DBNs from Incomplete Data Algorithm B.1 Parameter Set Update via Collapsed Gibbs Sampling 1: Input: Data {Y, X}, priors {ασ, βσ, a, b, µ}, current {σ2, δ2, π} 2: Sample σ2(Collapsed Gibbs step) σ2∼Inv−GAMασ+NT 2, βσ+1 2(Y−X[π]µ[π])⊤(I+σ2X[π]X⊤ [π])−1(Y−X[π]µ[π]) 3: Sample β β∼ N(σ−2I+X⊤ [π]X[π])−1(σ−2µ[π]+X⊤ [π]Y), σ2(σ−2I+X⊤ [π]X[π])−1 4: Sample δ2 δ2∼Inv−GAMa+|π|+ 1 2, b +1 2σ−2(β−µ[π])⊤(β−µ[π]) 5: Output: Posterior sample {β, σ2, δ2} Algorithm B.2 Covariate Set Update via Metropolis–Hastings 1: Input: Data {Y, X}, current covariate set π, parameters {δ2, σ2} 2: Choose a move type (Deletion, Addition, or Exchange) uniformly at random 3: Generate candidate covariate set π⋆based on the chosen move 4: Compute acceptance probability: A(π→π⋆) = min 1,p(Y|π⋆, δ2) p(Y|π, δ2)·p(π⋆) p(π)·HR with HR =|π| n−|π⋆|(Deletion), n−|π| |π⋆|(Addition), or 1(Exchange); p(Y|δ2, π) = Γ(A) Γ(ασ)·π−T 2(2βσ)ασdet(C)−1/22βσ+M⊤C−1M−A.a 5: Draw u∼ U(0,1) 6: if u<Athen 7: Accept: π←π⋆ 8: Sample βfrom FCD: β∼ N(σ−2I+X⊤ [π]X[π])−1(σ−2µ[π]+X⊤ [π]Y), σ2(σ−2I+X⊤ [π]X[π])−1 9: end if 10: Output: Posterior sample {π, β} aA=T 2+ασ,C=I+δ2XXT,M=Y−Xµ 7 Algorithm B.3 Missing Data Update via Gibbs Sampling 1: Input: Missing Variable Index i, 2: D= [xt j]j=1:k,t=1:3,B= [β(j) i]i=0:k,j=1:k,Σ={σ2(j)}j=1:k 3: Compute prior mean: µi=X p:(X1 p∈π(i)) x1 pβ(i) p 4: For each j, compute: µ(j) {−i}=X m=i:(X2 m∈π(j)) x2 m·β(j) m 5: Compute posterior variance: σ2∗=  1 σ2(i)+X j:(X2 i∈π(j)) (β(j) i)2 σ2(j)  −1 6: Compute posterior mean: µ∗=σ2∗·  µi σ2(i)+X j:(X2 i∈π(j)) β(j) i·x3 j−µ(j) {−i} σ2(j)  7: Sample xi[MIS]∼ N (µ∗, σ2∗) 8: Output: Posterior Sample xi[MIS] 8 Algorithm B.4 LUME-DBN 1: Input: Incomplete Dataset DM= [xt i,j]i=1:N;j=1:k;t=1:T, number of epochs E, Missing Imputation freq EM, prior parameters {ασ, βσ, a, b, µ = {µ(j) i}i=0:k,j=1:k} 2: Initialize: Initial models M(0) = [π(j) i(0)]i,j=1:k, 3: Initial linear coefficients B(0) = [β(j) i(0)]i=0:k;j=1:k, 4: Initial noise parameters Σ(0) ={σ2(j) (0)}j=1:k, 5: Initial uncertainty parameters ∆(0) ={δ2(j) (0)}j=1:k, 6: Initial completed dataset D(0) = [xt i,j (0) [MIS]]i=1:N;j=1:k;t=1:T 7: e←0 8: while e<Edo 9: e←e+ 1 10: X(e)←Lag(D(e−1)) 11: for emin {1,...,EM}do 12: e←e+ 1 13: for jin {1, . . . , k}do 14: Y←Xt j;X← X(t−1) {−j};µ←µ(j); 15: π←π(j)(e−1),β←β(j) (e−1);σ2←σ2(j) (e−1);δ2←δ2(j) (e−1) 16: Parameters Move: 17: Update via Algorithm B.1:{β(j)(e), σ2(j) (e), δ2(j) (e)}←{β, σ2, δ2} 18: Model Move: 19: Update via Algorithm B.2:{π(j)(e), β(j) (e)}←{π, β} 20: end for 21: end for 22: M(e)←[π(j) i(e)]i,j=1:k,B(e)←[β(j) i(e)]i=0:k;j=1:k 23: Σ(e)← {σ2(j) (e)}j=1:k,∆(e)← {δ2(j) (e)}j=1:k 24: Missing Data Imputation Move: 25: for {t, i}:∃xt i,j =NA (with j∈ {1, . . . , k}do 26: D ← [xt i,j (e−1) [MIS]]j=1:k;t=(t−1:t+1) 27: M←M(e),B ← B(e),Σ←Σ(e) 28: for j:xt i,j is missing do 29: Update Missing via Algorithm B.3:xt i,j (e) [MIS]←xj[MIS] 30: end for 31: end for 32: D(e)←[xt i,j (e) [MIS]]i=1:N;j=1:k;t=1:T 33: end while 34: Output: Posterior samples {M(e),B(e), Σ(e), ∆(e),D(e)}e=0:E 9 C Convergence Diagnostics C.1 Convergence Diagnostics on Simulated Data Fig. C.2. Convergence Diagnostic for Network reconstruction averaged over 5 simulations for simulated datasets with different missingness rates. Fig. C.3. Convergence Diagnostic for Missing Value imputation averaged over 5 simulations for simulated datasets with different missingness rates.