scieee AI-readable full text Open interactive document viewer

Error Geometry and Causal Robustness:\\A Unified Geometric Framework from Parameter Confidence Ellipsoids to Multi-Experiment Trust Regions

Ma, Haobo; Zhang, Wenlin

Abstract

This paper proposes a unified framework for transforming statistical errors into ``geometric boundaries'' and systematically embedding it into causal inference and experimental design. The core idea is: given an estimator and its error in parameter space, we no longer treat ``confidence intervals/standard errors'' as auxiliary information, but elevate them to ``trust regions'' in parameter space---geometric objects with metric structure; all causal conclusions, robustness judgments, and experimental planning are characterized through inclusion, intersection, union, and linear images among these geometric regions. Specifically, this paper first constructs confidence ellipsoids using typical asymptotic normality and information matrices under general parametric models, endowing parameter space with local Riemannian metric structure; second, in causal inference, we unify ``identifiable sets'' and ``trust regions'' as two types of sets in the same parameter space, characterizing provable causal conclusions and admissible extrapolation directions through their intersections; further, in multi-experiment/multi-model scenarios, we construct consensus regions and conflict regions via intersections, unions, and mappings of trust regions, forming a kind of ``geometrized meta-analysis''; finally, in experimental design and observation planning, we formalize the objective of ``shrinking trust region volume/semi-axes'' as an optimization problem over design variables, providing several solvable instances under linear models and instrumental variable models. The appendix provides rigorous proofs of main theorems, including coverage properties of confidence ellipsoids, robustness criteria for causal effects, and equivalence relations between design criteria and Fisher information.

Full text

Error Geometry and Causal Robustness: A Unified Geometric Framework from Parameter Confidence Ellipsoids to Multi-Experiment Trust Regions Haobo Ma1Wenlin Zhang2 1Independent Researcher 2National University of Singapore November 24, 2025 Abstract This paper proposes a unified framework for transforming statistical errors into “geometric boundaries” and systematically embedding it into causal inference and experimental design. The core idea is: given an estimator and its error in parameter space, we no longer treat “confidence intervals/standard errors” as auxiliary information, but elevate them to “trust regions” in parameter space—geometric objects with metric structure; all causal conclusions, robustness judgments, and experimental planning are characterized through inclusion, intersection, union, and linear images among these geometric regions. Specifically, this paper first constructs confidence ellipsoids using typical asymptotic normality and information matrices under general parametric models, endowing parameter space with local Riemannian metric structure; second, in causal inference, we unify “identifiable sets” and “trust regions” as two types of sets in the same parameter space, characterizing provable causal conclusions and admissible extrapolation directions through their intersections; further, in multi-experiment/multi-model scenarios, we construct consensus regions and conflict regions via intersections, unions, and mappings of trust regions, forming a kind of “geometrized meta-analysis”; finally, in experimental design and observation planning, we formalize the objective of “shrinking trust region volume/semi-axes” as an optimization problem over design variables, providing several solvable instances under linear models and instrumental variable models. The appendix provides rigorous proofs of main theorems, including coverage properties of confidence ellipsoids, robustness criteria for causal effects, and equivalence relations between design criteria and Fisher information. Keywords: Error Geometry; Confidence Ellipsoid; Causal Inference; Identifiable Set; Experimental Design; Fisher Information MSC 2020: 62F12, 62K05, 62P25, 62R01 Contents 1 1 Introduction Statistical inference is traditionally presented in the form of “point estimate + confidence interval”, while causal inference follows the basic structure of “identification assumptions + point estimate + sensitivity analysis”. In practice of decision-making and engineering applications, researchers often face three types of problems: 1. Given finite samples, which causal conclusions are truly supported by data, rather than mere illusions of point estimates? 2. When aggregating different experiments, different models, or even different data sources, how should we systematically see consensus and conflicts among various results? 3. Under limited resources, how can we maximize “geometric resolution” for specific causal effects or parameter directions through experimental design? These problems differ in form but share a common feature: they all relate to “error”, and the essence of “error” is not just a variance or a confidence interval, but a geometric object with shape, direction, and boundary. The goal of this paper is to thoroughly formalize this intuitive “geometric nature”. We adopt the following viewpoint:  For any estimate ˆ θof parameter θ∈Θ⊂Rd, its error naturally induces a region R ⊂ Θ with metric structure, which can be called a “trust region”;  Causal conclusions are not statements about single point ˆ θ, but about the range of some function ψ(θ) over θ∈ R∩I, where Iis the identifiable set;  Results from multiple experiments, models, and different assumptions can be unified as multiple trust regions Rkin the same parameter space or its projection, whose intersections, unions, and symmetric differences naturally characterize parameter ranges that are “robustly consistent”, “contestable”, or “significantly conflicting”;  Experimental design and observation planning can be viewed as an optimization problem of “actively shaping the geometric shape of future trust regions”, aiming to shrink semi-axes of trust regions in specific directions (such as some causal effect) or reduce their volume. In this framework, “error geometry” is no longer just an appendage of results, but becomes the core structure of the entire causality-decision process. This paper will start from the most basic parametric models, construct this geometric framework, and provide provable properties and computable implementations under several specific models. 2 Parametric Models and Geometric Structure of Trust Regions 2.1 Parametric Models and Estimators Let observed data be X1, . . . , Xn, defined on sample space X, assuming their distribution belongs to a family of probability measures {Pθ:θ∈Θ⊂Rd}. Let θ0∈Θ be the “true 2 parameter”, ˆ θn=ˆ θn(X1, . . . , Xn) some estimator. We assume there exists the usual asymptotic linearity and normality structure: √n(ˆ θn−θ0)d −→ N(0, I(θ0)−1), where I(θ0) is the Fisher information matrix, positive definite and continuous. Further assume there exists consistent estimate ˆ Insuch that ˆ In P −→ I(θ0). 2.2 Local Metric Induced by Fisher Information At each point θof Θ, define bilinear form gθ(u, v) := u⊤I(θ)v, u, v ∈Rd, then gis a local Riemannian metric on Θ (under differentiability conditions). Intuitively, eigendirections of I(θ) describe “easy/difficult to distinguish” parameter directions: the smaller the variance along some direction, the larger the “unit length” in that direction under information metric, and vice versa. Under finite sample n, using ˆ Inwe obtain empirical metric ˆgn(u, v) := u⊤ˆ Inv, which converges to gθ0in probability sense. 2.3 Confidence Ellipsoid as Trust Region For given significance level α∈(0,1), define the quantile χ2 d,1−αof d-dimensional chisquare distribution, construct trust region Rn(α) := {θ∈Θ : n(θ−ˆ θn)⊤ˆ In(θ−ˆ θn)≤χ2 d,1−α}. It is an ellipsoid centered at ˆ θn, with shape determined by ˆ I−1 n, characterizing parameter uncertainty. Classical theory guarantees: Theorem 2.1 (Asymptotic Coverage).Under the above regularity conditions, for any fixed α∈(0,1), Pθ0θ0∈ Rn(α)−→ 1−α, n → ∞. This theorem is proved in Appendix A.1. Therefore, Rn(α) can be viewed as a “trust region containing true value θ0with probability 1−α”. Under information metric gθ0, semi-axis lengths of Rn(α) are inversely proportional to eigenvalues of Fisher information and proportional to 1/√n. 3 Error as Geometric Boundary: Operations and Projections of Trust Regions This section systematically transforms “error” into “geometric boundary” and discusses several basic operations: projection, linear image, and nonlinear image. 3 3.1 Linear Function Image and Ellipsoid Projection Let the target of interest be linear function ψ(θ) = c⊤θ, where c∈Rd. On ellipsoid Rn(α), the range of ψis Ψn(α) := {ψ(θ) : θ∈ Rn(α)}=ψmin,n, ψmax,n. This interval can be analytically calculated. Note that ψ(θ) = c⊤θ=c⊤ˆ θn+c⊤(θ−ˆ θn), with constraint n(θ−ˆ θn)⊤ˆ In(θ−ˆ θn)≤χ2 d,1−α. Let h=θ−ˆ θn, the problem becomes maximizing/minimizing linear function c⊤hunder ellipsoid constraint. Classical optimization conclusion gives ψmax,n =c⊤ˆ θn+sχ2 d,1−α nc⊤ˆ I−1 nc, ψmin,n =c⊤ˆ θn−sχ2 d,1−α nc⊤ˆ I−1 nc. Therefore, confidence interval for linear target is naturally given by geometric relationship between ellipsoid and direction vector c. Proposition 3.1 (Optimal Bounds for Linear Target).For any c∈Rd, the minimum and maximum values of ψ(θ) = c⊤θon Rn(α)are as shown above, and this interval has coverage probability 1−αin asymptotic sense. Proof is in Appendix A.2. 3.2 Local Linear Approximation of Nonlinear Functions If ψ: Θ →Rkis differentiable, then near ˆ θnwe can make first-order approximation ψ(θ)≈ψ(ˆ θn) + Dψ(ˆ θn)(θ−ˆ θn), where Jacobian matrix Dψ(ˆ θn)∈Rk×dhas i-th row as ∇ψi(ˆ θn)⊤. Then ψ(Rn(α)) under first-order approximation is an ellipsoid in Rk: Sn(α) := ny∈Rk:n(y−ψ(ˆ θn))⊤Dψ(ˆ θn)ˆ I−1 nDψ(ˆ θn)⊤−1(y−ψ(ˆ θn)) ≤χ2 k,1−αo, where existence of inverse matrix requires row vectors of Dψ(ˆ θn) to be linearly independent under information metric. This result is essentially a restatement of Delta method in geometric language. 4 3.3 Support Function Form for Multi-Parameter and MultiTarget In more general cases, we can use support functions to characterize arbitrary convex target sets. On convex ellipsoid Rn(α), its support function is hRn(α)(u) := sup θ∈Rn(α) u⊤θ=u⊤ˆ θn+sχ2 d,1−α nu⊤ˆ I−1 nu. Therefore, any target set defined by collection of linear functionals {u:u∈ U} can have its boundary directly calculated through support function. An important application in causal inference scenarios is: when we care about a family of linear causal effects (such as heterogeneous effects across multiple groups), we can uniformly provide their worst/best cases on trust regions through support functions. 4 Identifiable Sets and Trust Regions in Causal Inference 4.1 Causal Models and Identifiable Sets In causal inference, parameter θusually has structural interpretation, such as average treatment effect in potential outcome models, path coefficients in structural equation models, local average treatment effect in instrumental variable models, etc. In cases of non-identification or partial identification, what data and assumptions can determine is only an identifiable set I:= {θ∈Θ : θis compatible with observed distribution and causal assumptions}. For example, when violating certain exclusion restrictions or with selection bias, identifiable sets are often convex sets, semi-algebraic sets, or general closed sets, rather than single points. 4.2 Data-Driven Estimation of Identifiable Sets Under finite samples, we typically approximate Ithrough estimated inequalities. For example, if causal constraints can be expressed as parameter constraints gj(θ)≤0, j = 1, . . . , m, while we can only estimate empirical version ˆgj,n(θ) of gj(θ), common practice is using “relaxed inequalities” ˆgj,n(θ)≤bj,n, where bj,n is upper bound (such as threshold after multiple testing correction), thus obtaining data-driven identifiable set estimate ˆ In:= {θ∈Θ : ˆgj,n(θ)≤bj,n, j = 1, . . . , m}. 5 Under suitable conditions it can be proved that ˆ Inconverges to Iin appropriate sense. 4.3 Intersection of Identifiable Set and Trust Region This paper proposes: Causal conclusions should be based on Rn(α)∩ˆ Inrather than merely on ˆ θn. For given causal function ψ: Θ →Rk, we care about Cn(α) := {ψ(θ) : θ∈ Rn(α)∩ˆ In}. If Cn(α) has consistent sign in some direction or component, or is restricted to some desired interval, we can say this causal conclusion is “geometrically robust” at significance level α. Definition 4.1 (Geometric Robustness).Let ψ: Θ →Rbe a scalar causal target, Rn(α)a1−αlevel trust region, ˆ Insample approximation of identifiable set. If there exists interval [L, U]⊂Rsuch that Cn(α) = {ψ(θ) : θ∈ Rn(α)∩ˆ In} ⊂ [L, U], then say “at level α, causal conclusion ψ(θ)∈[L, U] is geometrically robust”. In particular, when L > 0 (or U < 0), robust judgment can be made about effect direction. 4.4 A Typical Criterion: Linear Causal Effect Let ψ(θ) = c⊤θbe linear causal effect (such as some linear combination in multi-parameter model corresponding to average treatment effect), and identifiable set can be represented as linear inequalities Aθ ≤b, then Rn(α)∩ˆ Inis intersection of ellipsoid and polyhedron, a convex set. Extreme values of causal effect can be given by following convex optimization problem: ψ∗ max,n := sup{c⊤θ:θ∈ Rn(α), Aθ ≤b}, ψ∗ min,n := inf{c⊤θ:θ∈ Rn(α), Aθ ≤b}. In many applications, this problem can be efficiently solved through quadratic programming or semidefinite programming. Clearly, Cn(α) = ψ∗ min,n, ψ∗ max,n. Theorem 4.2 (Geometric Robustness Criterion for Linear Causal Effect).Under above setting, if for some δ > 0, ψ∗ min,n ≥δ > 0, then at significance level α, can conclude causal conclusion “effect is positive and at least δ”, and this conclusion holds for all θ∈ Rn(α)∩ˆ In. 6 Proof is in Appendix A.3. This criterion elevates “point estimate significance” to “significance over all trusted candidate parameters”, naturally excluding spurious significance that may arise from relying only on point estimates while ignoring parameter correlations. 5 Multiple Experiments and Models: Intersection, Union, and Conflict Structure of Trust Regions In reality, we often need to synthesize results from different experiments, different data sources, or even different models. This paper advocates: The natural objects for multi-experiment aggregation are not “several point estimates”, but “several trust regions”. 5.1 Intersection and Consensus of Multiple Trust Regions Suppose there are Kexperiments/data sources/models giving trust regions R(k) nk(αk) on same parameter space Θ, k= 1, . . . , K. Define overall consensus region Rcons := K \ k=1 R(k) nk(αk). If the causal function of interest is ψ: Θ →Rd, then its image on consensus region is Ccons := {ψ(θ) : θ∈ Rcons}. If Rcons is non-empty and “small”, it indicates high consistency among different experiments; conversely, if Rcons is empty, can clearly say “there exists fundamental conflict among these experiments/models”, rather than vaguely relying on “some estimation differences”. 5.2 Union and Admissible Sets On the other hand, define admissible region as Rperm := K [ k=1 R(k) nk(αk), which characterizes the parameter set “supported by at least one experiment”. In some decision problems (such as tolerating partial experiment failure or model misspecification), we may only require conclusions to hold on Rperm. 5.3 Conflict Region and Uncertainty Decomposition Define conflict region as symmetric difference Rconflict := K [ k=1 R(k) nk!\ K \ k=1 R(k) nk!, 7 where if some point θis only supported by partial experiments (while excluded by others), then belongs to this region. By visualizing parameter space as partition of “consensus–conflict–unconstrained”, can intuitively identify which parameter directions’ conclusions are most sensitive to experiment selection. 6 Experimental Design and Observation Planning: Trust Region as Objective Function In above framework, the essence of experimental design is: Shaping the shape and size of future trust regions through choosing experimental schemes or observation strategies. This section provides formalized characterization under classical linear models and general Fisher information background. 6.1 Fisher Information and Region Volume In regular models, with sample size n, information matrix can usually be expressed as In(θ) = nI1(θ), where I1(θ) is information from single observation. For given design parameter ξ (such as distribution of samples over different treatment/covariate configurations), single observation information can be written as I1(θ;ξ). Therefore, In(θ;ξ) = nI1(θ;ξ). Volume of ellipsoidal trust region is proportional to detIn(θ0;ξ)−1/2. More precisely, when Rn(α;ξ) is ellipsoid based on In(θ0;ξ), its Lebesgue volume is VolRn(α;ξ)=Cd,α det In(θ0;ξ)−1/2, where constant Cd,α only depends on dimension dand α. Therefore, minimizing region volume is equivalent to maximizing det In(θ0;ξ). Definition 6.1 (Error Geometric Characterization of D-Optimal Design).If design ξ∗ satisfies det In(θ0;ξ∗) = sup ξ det In(θ0;ξ), then call ξ∗D-optimal design. Geometrically, it makes trust region volume minimal under given n, thus most compact overall. Classical D-optimality conclusions are restated in error geometric language in Appendix A.4. 6.2 Directional Resolution: A-Optimal and c-Optimal If focus is on specific linear causal effect ψ(θ) = c⊤θ, then its asymptotic variance is Varˆ ψ≈1 nc⊤I1(θ0;ξ)−1c. 8 From error geometry perspective, this is exactly the squared length of principal semiaxis of trust ellipsoid in direction c(ignoring constants). Therefore, minimizing this variance is equivalent to maximizing resolution in direction c. Definition 6.2 (c-Optimal Design).If design ξ∗satisfies c⊤I1(θ0;ξ∗)−1c= inf ξc⊤I1(θ0;ξ)−1c, then call ξ∗c-optimal design. From geometric perspective: c-optimal design does not pursue minimal overall ellipsoid volume, but specifically compresses semi-axis in direction c, i.e., focuses on enhancing geometric resolution of this causal effect. 7 Examples of Error Geometry in Typical Models This section briefly demonstrates specific forms of error geometry framework under several common models for readers to gain intuitive impression. 7.1 Confidence Ellipsoid and Effect Interval in Linear Regression Models Consider linear regression model Yi=x⊤ iβ+εi, εi∼ N(0, σ2), where xi∈Rpare known covariates, β∈Rpare regression coefficients. Let Xbe design matrix, OLS estimate is ˆ β= (X⊤X)−1X⊤Y. Classical result gives ˆ β∼ N β, σ2(X⊤X)−1. Therefore, information matrix is I(β) = σ−2X⊤X, trust ellipsoid is R(α) = {β: (β−ˆ β)⊤X⊤X(β−ˆ β)≤σ2χ2 p,1−α}. For any linear prediction ψ(β) = x⊤ newβ, its interval on R(α) is x⊤ new ˆ β±qχ2 p,1−ασ2x⊤ new(X⊤X)−1xnew, which is completely consistent with classical linear regression confidence interval, but in this framework is interpreted as “geometric projection of trust ellipsoid in direction xnew”. 9