scieee AI-readable full text Open interactive document viewer

Neural Operators in Anisotropic Fractional Sobolev-Morrey Spaces

Santos, Rômulo Damasclin Chaves dos; Sales, Jorge

Abstract

This paper develops a comprehensive mathematical theory of anisotropic fractional calculus with mixed regularity structures, addressing fundamental challenges in analyzing high-dimensional functions with heterogeneous smoothness across different coordinate directions. Motivated by applications in scientific machine learning, multiscale analysis, and physical systems with directional preferences, we introduce novel anisotropic fractional Sobolev-Morrey spaces that precisely capture directional scaling behavior through mixed regularity parameters. These spaces provide a refined analytical framework for functions exhibiting varying degrees of smoothness along different coordinates, generalizing classical isotropic theories to anisotropic settings. Our principal contributions establish several sharp functional inequalities: (1) anisotropic Gagliardo-Nirenberg inequalities with mixed fractional derivatives featuring explicit constant dependence on scaling parameters and proven optimality; (2) directional Hardy-Littlewood-Sobolev theory for anisotropic fractional integrals with optimal bounds in Lebesgue and Morrey spaces; (3) compactness criteria in anisotropic function spaces demonstrated through refined real interpolation and harmonic analysis techniques; and (4) optimal approximation rates for deep neural operators in high-dimensional settings, with explicit dimension dependence governed by the anisotropic dimension $d_\alpha = \sum_{i=1}^k \alpha_i^{-1}$. The theoretical framework bridges harmonic analysis, fractional calculus, and deep learning theory, providing rigorous mathematical foundations for understanding the approximation capabilities of modern neural architectures. Furthermore, our results offer principled guidance for neural operator design in scientific computing applications, particularly for problems exhibiting multiscale and anisotropic features. This work opens new research directions in the analysis of partial differential equations, high-dimensional approximation theory, and the mathematical foundations of deep learning.

Full text

Neural Operators in Anisotropic Fractional Sobolev-Morrey Spaces Rômulo Damasclin Chaves dos Santos Santa Cruz State University [email protected] Jorge Henrique de Oliveira Sales Santa Cruz State University [email protected] November 22, 2025 Abstract This paper develops a comprehensive mathematical theory of anisotropic fractional calculus with mixed regularity structures, addressing fundamental challenges in analyzing high-dimensional functions with heterogeneous smoothness across different coordinate directions. Motivated by applications in scientific machine learning, multiscale analysis, and physical systems with directional preferences, we introduce novel anisotropic fractional Sobolev-Morrey spaces that precisely capture directional scaling behavior through mixed regularity parameters. These spaces provide a refined analytical framework for functions exhibiting varying degrees of smoothness along different coordinates, generalizing classical isotropic theories to anisotropic settings. Our principal contributions establish several sharp functional inequalities: (1) anisotropic Gagliardo-Nirenberg inequalities with mixed fractional derivatives featuring explicit constant dependence on scaling parameters and proven optimality; (2) directional Hardy-Littlewood-Sobolev theory for anisotropic fractional integrals with optimal bounds in Lebesgue and Morrey spaces; (3) compactness criteria in anisotropic function spaces demonstrated through refined real interpolation and harmonic analysis techniques; and (4) optimal approximation rates for deep neural operators in high-dimensional settings, with explicit dimension dependence governed by the anisotropic dimension dα=Pk i=1 α−1 i. The theoretical framework bridges harmonic analysis, fractional calculus, and deep learning theory, providing rigorous mathematical foundations for understanding the approximation capabilities of modern neural architectures. Furthermore, our results offer principled guidance for neural operator design in scientific computing applications, particularly for problems exhibiting multiscale and anisotropic features. This work opens new research directions in the analysis of partial differential equations, high-dimensional approximation theory, and the mathematical foundations of deep learning. Keywords: Anisotropic Fractional Calculus, Sobolev-Morrey Spaces, GagliardoNirenberg Inequalities, Neural Operators, Multiscale Analysis. 1 1 Introduction and Mathematical Background The classical Landau inequality [1] ∥f′∥∞≤2p∥f∥∞∥f′′∥∞(1.1) represents a fundamental trade-off between function magnitude and oscillation that has influenced mathematical analysis for nearly a century. Recent developments in fractional calculus [2] have extended this theory to non-local operators, while multivariate extensions by Ditzian [3] and Kounchev [4] have generalized these results to multidimensional settings. However, existing theories operate primarily within isotropic function spaces, overlooking the rich multiscale structure present in modern applications ranging from highdimensional data analysis to physical systems with directional preferences. This work addresses this fundamental limitation by developing a comprehensive theory of anisotropic fractional calculus with mixed regularity structures. 1.1 Principal Contributions This work makes four fundamental contributions that bridge harmonic analysis, fractional calculus, and deep learning theory: (i) Anisotropic Fractional Sobolev-Morrey Spaces: We introduce the spaces Mν p,λ;α(Rk)that capture directional scaling behavior through mixed regularity parameters, establishing their complete theory including equivalent characterizations, embedding theorems, interpolation results, and algebra structures with sharp constants that explicitly track dependence on scaling parameters. (ii) Sharp Anisotropic Functional Inequalities: We prove optimal GagliardoNirenberg inequalities and Hardy-Littlewood-Sobolev estimates for directional fractional integrals, establishing compactness criteria in mixed-norm spaces with explicit constants that reveal how directional heterogeneity affects functional relationships and operator bounds. (iii) Advanced Harmonic Analysis Framework: We develop comprehensive tools for anisotropic analysis including directional Littlewood-Paley theory, anisotropic maximal function estimates, and fractional calculus on heterogeneous scaling structures, providing the mathematical infrastructure for analyzing functions with directional regularity patterns. (iv) Multiscale Operator Learning Theory: We establish rigorous foundations for neural operators in high-dimensional settings, proving stability bounds under anisotropic perturbations and deriving optimal approximation rates N−ν/dαthat adapt to intrinsic anisotropic dimension rather than ambient dimension, offering mathematical justification for deep learning’s empirical success in scientific computing. These contributions collectively provide a unified mathematical framework for analyzing and approximating functions with heterogeneous regularity across different coordinate directions, with significant implications for both theoretical analysis and practical applications in scientific machine learning. 2 2 Anisotropic Fractional Sobolev-Morrey Spaces This section establishes the fundamental geometric and analytic framework for anisotropic fractional analysis, providing the mathematical foundations for the sharp inequalities and applications developed in subsequent sections. The core innovation lies in developing function spaces that capture heterogeneous scaling behavior across different coordinate directions, a feature ubiquitous in multiscale physical systems, high-dimensional data, and deep neural networks with directional preferences. Unlike classical isotropic theories that treat all directions uniformly, our approach incorporates directional scaling parameters α= (α1, . . . , αk)that modulate the effective regularity along each coordinate axis. This anisotropic perspective enables more precise characterization of functions exhibiting varying degrees of smoothness in different directions, ultimately leading to sharper functional inequalities and more efficient approximation strategies. The construction proceeds systematically: we first define the underlying anisotropic geometry through scaling structures and homogeneous dilations, then introduce the anisotropic fractional Sobolev-Morrey spaces that combine directional fractional differentiability with refined integrability conditions. The mathematical novelty stems from the interplay between three fundamental aspects: (i) directional fractional derivatives that assign different smoothness exponents along each coordinate axis, (ii) Morrey-type integrability conditions that capture local versus global behavior, and (iii) the anisotropic scaling geometry that governs the interaction between different directions. This tripartite structure enables a nuanced analysis of functions with heterogeneous regularity patterns, providing the mathematical language to describe multiscale phenomena where traditional isotropic theories prove inadequate. The following definitions and propositions establish the basic objects and their properties that will underpin the entire theoretical development. We begin with the anisotropic scaling geometry, which replaces the classical Euclidean dilation structure with a parameterized family that respects directional heterogeneity. This geometric foundation will subsequently support the development of anisotropic harmonic analysis, including Littlewood-Paley theory, fractional operators, and Sobolev-type embeddings tailored to the mixed regularity setting. 2.1 Geometric Foundations and Scaling Structure The study of anisotropic function spaces, differential operators, and harmonic analysis requires a geometric framework where different spatial coordinates scale at potentially different rates. Such anisotropic scaling naturally emerges in diverse contexts including kinetic equations, degenerate elliptic operators, multiscale diffusion processes, and the analysis of neural operators with heterogeneous receptive fields. In all these settings, the underlying geometry is no longer governed by the classical Euclidean dilation x7→ λx, but rather by a parameterized family of dilations that encode directional preferences through scaling exponents. Definition 2.1 (Anisotropic Scaling Geometry).Let α= (α1, . . . , αk)∈(0,∞)kbe a scaling vector defining the anisotropic geometry. The associated anisotropic dilation group {Tα λ}λ>0is defined by: Tα λf(x) = f(λα1x1, . . . , λαkxk), λ > 0.(2.1) 3 This family of dilations induces a non-Euclidean geometric structure characterized by two fundamental quantities: the anisotropic homogeneous dimension dα= k X i=1 α−1 i,(2.2) which plays the role of an effective dimensional exponent in integration and Fourier analysis, and the anisotropic distance function ρα(x) = k X i=1 |xi|2/αi!1/2 ,(2.3) which is homogeneous with respect to the dilations Tα λin the sense that ρα(Tα λx) = λρα(x). The anisotropic dilation group and associated geometric quantities form a coherent algebraic and analytical structure that will underpin all subsequent developments. The following proposition establishes the fundamental properties of this structure, which will be used repeatedly in the analysis of anisotropic kernels, Sobolev norms, and semigroup characterizations. Proposition 2.2 (Anisotropic Scaling Properties).The anisotropic dilation group {Tα λ}λ>0 satisfies the following fundamental properties: (a) Group structure: Tα λTα µ=Tα λµ,(Tα λ)−1=Tα 1/λ. (b) Jacobian determinant: |det(DTα λ)|=λdα. (c) Scaling of Lebesgue measure: ZRk f(Tα λx)dx =λ−dαZRk f(x)dx. (2.4) (d) Fourier transform relation: F[f◦Tα λ](ξ) = λ−dαF[f](Tα 1/λξ).(2.5) Proof. The group structure (a) follows immediately from composition of dilations. For (b), the Jacobian matrix of Tα λis diagonal with entries λαiδij, so the determinant is Qk i=1 λαi=λdα. Property (c) then follows from the change-of-variables formula with y=Tα λx, giving dy =λdαdx. For the Fourier relation (d), we compute directly: F[f◦Tα λ](ξ) = ZRk f(λα1x1, . . . , λαkxk)e−2πix·ξdx =λ−dαZRk f(y)e−2πi(Tα 1/λy)·ξdy =λ−dαF[f](Tα 1/λξ), where the second equality uses the change of variables y=Tα λxand the homogeneity of the anisotropic distance. This completes the proof of all properties. 4 2.2 Mixed Fractional Sobolev-Morrey Spaces Having established the fundamental geometric framework, we now introduce the central function spaces of this work: the anisotropic fractional Sobolev-Morrey spaces. These spaces represent a significant advancement beyond classical Sobolev spaces by simultaneously incorporating three crucial features: directional fractional differentiability, anisotropic scaling geometry, and refined integrability conditions through Morrey-type norms. This triple structure enables a precise characterization of functions exhibiting heterogeneous regularity patterns across different coordinate directions, a phenomenon commonly encountered in multiscale physical systems, high-dimensional data analysis, and deep neural networks with directional architectures. The mathematical innovation lies in the careful interplay between these three components. The directional fractional differentiability is modulated by the scaling parameters αi, ensuring that the smoothness measurement in each direction respects the underlying anisotropic geometry. The Morrey component provides a refined control over local versus global behavior, capturing the concentration properties of functions that are crucial for understanding phenomena with multiscale characteristics. The anisotropic structure governs the interaction between different directions, ensuring that the resulting function spaces form a coherent analytical framework. From a technical perspective, these spaces interpolate between several classical constructions: they generalize the isotropic fractional Sobolev spaces when αi≡1, recover anisotropic Sobolev spaces when νis an integer and λ= 0, and specialize to Morrey spaces when no differentiability is imposed. However, the true power emerges from the nontrivial interactions between these aspects, leading to new embedding theorems, interpolation results, and approximation properties that cannot be obtained through straightforward combinations of existing theories. The following definition formalizes this construction, providing the precise mathematical framework that will support the development of sharp functional inequalities and their applications to neural operator theory in subsequent sections. Definition 2.3 (Anisotropic Fractional Sobolev-Morrey Space).For ν > 0,1≤p < ∞,0≤λ≤dα, and scaling vector α, the anisotropic fractional Sobolev-Morrey space Mν p,λ;α(Rk)is the completion of C∞ c(Rk)under the norm: ∥f∥Mν p,λ;α=∥f∥Mp,λ;α+ k X i=1 [f]Wν/αi,p,λ i ,(2.6) where the Morrey norm captures the global integrability properties: ∥f∥Mp,λ;α= sup x∈Rk,r>0 r−λ/p∥f∥Lp(Bα(x,r)),(2.7) and the directional fractional seminorms encode the anisotropic smoothness: [f]Wν/αi,p,λ i = sup x∈Rk,r>0 r−λ/p ZBα(x,r)ZR |f(y+hei)−f(y)|p |h|1+pν/αidhdy1/p .(2.8) Here Bα(x, r) = {y∈Rk:ρα(x−y)< r}denotes the anisotropic ball, and the scaling ν/αiin the directional seminorms ensures homogeneity with respect to the anisotropic dilation group. 5 The mathematical structure of these spaces warrants several important observations. First, the Morrey component (2.7) provides a scale-invariant measure of integrability that interpolates between Lebesgue spaces (λ= 0) and spaces of bounded functions (λ=dα). This refined integrability is essential for capturing the local concentration properties that arise in multiscale problems. Second, the directional fractional seminorms (2.8) incorporate the anisotropic scaling by assigning smoothness order ν/αiin the i-th direction, ensuring that the overall regularity is homogeneous of degree νunder the anisotropic dilations Tα λ. This directional approach allows for precise characterization of functions that may be highly regular in some directions while exhibiting limited smoothness in others. The completion with respect to this norm ensures that Mν p,λ;α(Rk)forms a Banach space, and the use of C∞ c(Rk)as the dense subset guarantees that the resulting space admits a rich theory of approximations and density arguments. In the subsequent sections, we will establish several equivalent characterizations of these spaces, develop their embedding properties, and prove sharp interpolation results that illuminate their structural relationships with classical function spaces. 2.3 Equivalent Characterizations A fundamental aspect of developing robust function spaces lies in establishing multiple equivalent characterizations that illuminate different analytical perspectives and facilitate diverse applications. This subsection presents five distinct but equivalent ways to understand the anisotropic fractional Sobolev-Morrey spaces, each offering unique advantages for different mathematical contexts. The equivalence of these characterizations demonstrates the intrinsic coherence of our construction and provides powerful tools for establishing functional inequalities, embedding theorems, and approximation results. The various characterizations span different analytical methodologies: the Gagliardo approach (2.10) provides a direct geometric interpretation through difference quotients; the Littlewood-Paley characterization (2.11) offers a frequency-domain perspective essential for harmonic analysis; the Bessel potential formulation (2.12) connects to the theory of fractional operators and PDEs; and the heat semigroup characterization (2.13) links to diffusion processes and semigroup theory. Each perspective reveals different aspects of the function space structure and enables different proof techniques. From a technical standpoint, establishing these equivalences requires developing several sophisticated tools in the anisotropic setting, including: anisotropic polar coordinates, directional Calderón reproducing formulas, anisotropic Bernstein inequalities, and precise estimates for the anisotropic heat kernel. The proof strategy involves carefully bounding each characterization in terms of the others, with particular attention to the dependence on the scaling parameters αiand the Morrey exponent λ. The following theorem summarizes these equivalent characterizations, providing a comprehensive toolbox for working with anisotropic fractional Sobolev-Morrey spaces in various mathematical contexts. Theorem 2.4 (Equivalent Characterizations).Let ν > 0,1≤p < ∞,0≤λ≤dα, and 6 αi>0. For f∈Lp(Rk), the following quantities are equivalent: ∥f∥Mν p,λ;α(2.9) ∼ ∥f∥Mp,λ;α+ZRkZRk |f(x)−f(y)|p ρα(x−y)dα+pν dydx1/p (2.10) ∼ ∥f∥Mp,λ;α+ ∞ X j=0 22jν|∆α jf|2!1/2Mp,λ;α (2.11) ∼ ∥(I−∆α)ν/2f∥Mp,λ;α(2.12) ∼ ∥f∥Mp,λ;α+Z∞ 0 t−ν/2∥f−et∆αf∥p Mp,λ;α dt t1/p (2.13) where ∆α=Pk i=1(−∂2 i)1/αiis the anisotropic Laplacian and {∆α j}is the anisotropic Littlewood-Paley decomposition. The equivalence constants depend explicitly on ν, p, λ, and the scaling vector α, but are independent of the function f. Proof. We prove the equivalence through several steps, each leveraging different aspects of anisotropic harmonic analysis: 1. (2.9) ⇔(2.10). The anisotropic Gagliardo integral in (2.10) provides a global measure of fractional differentiability. To relate it to the directional seminorms, we employ anisotropic polar coordinates adapted to the geometry defined by ρα. Specifically, we use the decomposition: Rk= k [ i=1 {x∈Rk:|xi|1/αi= max j|xj|1/αj}, which partitions space into regions where different coordinates dominate the anisotropic distance. In each region, we can bound the global difference quotient by the directional difference quotient along the dominant coordinate. The key estimate is: ZRkZRk |f(x)−f(y)|p ρα(x−y)dα+pν dydx ≤C k X i=1 ZRkZR |f(x+hei)−f(x)|p |h|1+pν/αidhdx, which follows from the asymptotic equivalence ρα(x)∼maxi|xi|1/αiand careful integration in anisotropic spherical coordinates. The reverse inequality uses a chaining argument that connects points through coordinate-aligned paths. 2. (2.9) ⇔(2.11). The Littlewood-Paley characterization requires developing anisotropic versions of classical harmonic analysis tools. The anisotropic Calderón reproducing formula: f= ∞ X j=0 ∆α jfin S′(Rk), is established by constructing a dyadic decomposition adapted to the anisotropic scaling. The anisotropic Bernstein inequalities play a crucial role: ∥∂k i∆α jf∥Lp≤C2jk/αi∥∆α jf∥Lp, 7 which are proved by scaling arguments using the homogeneity properties of the anisotropic dilations. The square function characterization (2.11) then follows from establishing the boundedness of the anisotropic Hardy-Littlewood maximal function on Morrey spaces and applying the Fefferman-Stein vector-valued inequality in the anisotropic setting. 3. (2.12) ⇔(2.13). The semigroup characterization relies on precise estimates for the anisotropic heat kernel. We establish that the kernel of et∆αsatisfies: |et∆α(x, y)|≤ Ct−dα/2exp −cρα(x−y)2 t, through Fourier analysis and the scaling properties of the anisotropic Laplacian. The equivalence between the Bessel potential and semigroup characterizations follows from the anisotropic version of the classical result by DeVore and Sharpley, which we prove by writing: (I−∆α)−ν/2f=1 Γ(ν/2) Z∞ 0 tν/2−1e−tet∆αf dt, and carefully estimating the Morrey norms of the resulting expressions. The passage between (2.12) and (2.13) involves establishing the equivalence between potential norms and interpolation norms in the anisotropic Morrey space setting. 4. Completing the cycle. To establish the full equivalence, we prove that each characterization bounds all the others by combining the estimates from Steps 1-3 and using the interpolation properties of the anisotropic Sobolev-Morrey spaces. The explicit dependence of the equivalence constants on the parameters follows from tracking the constants in each of the underlying inequalities and optimization arguments. The equivalence established in this theorem has profound implications for both theoretical analysis and practical applications. From a theoretical perspective, it demonstrates that our definition of anisotropic fractional Sobolev-Morrey spaces captures an intrinsic mathematical concept that manifests consistently across different analytical frameworks. From an applied viewpoint, it provides flexibility in choosing the most convenient characterization for specific problems whether studying PDEs, developing approximation algorithms, or analyzing neural networks. 3 Sharp Anisotropic Functional Inequalities This section presents the core analytical contributions of this work: sharp functional inequalities in anisotropic fractional Sobolev-Morrey spaces. These inequalities represent fundamental relationships between different norms and derivatives that reveal the intrinsic structure of functions with mixed regularity. The anisotropic setting introduces significant mathematical challenges, as traditional isotropic techniques fail to capture the directional scaling behavior encoded in the parameter vector α. Our results provide precise quantitative bounds with explicit dependence on scaling parameters, offering insights into how directional heterogeneity affects functional relationships. The development of these inequalities requires novel approaches that combine techniques from harmonic analysis, fractional calculus, and geometric measure theory. Unlike their isotropic counterparts, anisotropic functional inequalities must account for the interplay between different coordinate directions and their respective scaling exponents. This leads to constants that depend intricately on the anisotropic dimension dαand exhibit scaling properties compatible with the underlying geometry. 8 From a broader perspective, these inequalities serve multiple purposes: they establish the coercivity properties needed for the analysis of anisotropic partial differential equations, provide the theoretical foundation for understanding approximation rates in high-dimensional settings, and offer tools for proving stability results in machine learning applications. The explicit nature of our constants makes them particularly valuable for quantitative applications where dimension dependence and scaling behavior play crucial roles. 3.1 Mixed Fractional Gagliardo-Nirenberg Inequalities The Gagliardo-Nirenberg inequality represents one of the most fundamental interpolation results in functional analysis, connecting different Sobolev norms through precise interpolation estimates. In the anisotropic fractional setting, this inequality takes on a richer structure that reflects the directional heterogeneity of the function spaces. Our mixed fractional version generalizes the classical result in several significant ways: it incorporates fractional derivatives of different orders along different coordinate directions, accounts for Morrey-type integrability conditions, and provides explicit constants that capture the interplay between the scaling parameters αiand the smoothness exponents. The mathematical innovation lies in developing interpolation techniques that respect the anisotropic geometry while handling the non-local nature of fractional derivatives. This requires careful analysis of how directional smoothness propagates through the interpolation process and how the anisotropic dimension dαgoverns the critical exponents. The resulting inequality provides a powerful tool for trading between different levels of regularity in a way that adapts to the intrinsic directional structure of the function. Theorem 3.1 (Anisotropic Fractional Gagliardo-Nirenberg).Let ν > 0,1< p, p1, p2< ∞,0≤θ≤1, and αbe a scaling vector. Suppose: 1 p=θ p1 +1−θ p2 , ν =θν1+ (1 −θ)ν2, λ =θλ1+ (1 −θ)λ2.(3.1) Then for all f∈ Mν1 p1,λ1;α∩ Mν2 p2,λ2;α, we have: ∥f∥Mν p,λ;α≤C∥f∥θ Mν1 p1,λ1;α∥f∥1−θ Mν2 p2,λ2;α ,(3.2) where the constant C=C(ν, p, α, θ)satisfies the sharp bound: C≤Γ(ν1+ 1)Γ(ν2+ 1) Γ(ν+ 1) 1/2 k Y i=1 α−1/2 i!κ(p, α),(3.3) with κ(p, α)the optimal constant from the anisotropic Hardy-Littlewood inequality. Proof. We provide a comprehensive proof that combines real interpolation theory with anisotropic harmonic analysis and scaling arguments. The strategy involves three main steps: establishing the equivalence with Besov-Morrey spaces, proving interpolation results in this framework, and computing sharp constants through optimization. 1. Besov-Morrey characterization. We first establish the equivalence between anisotropic Sobolev-Morrey spaces and anisotropic Besov-Morrey spaces. This requires 9 which establishes uniform equicontinuity in the supremum norm. Combined with the uniform boundedness in C(Ω) (which follows from the continuous embedding), the ArzelàAscoli theorem guarantees the compactness of the embedding into C(Ω). 6. Alternative approach via interpolation. For completeness, we also present an interpolation approach to compactness. Using the real interpolation identity: [Mν0 p0,λ0;α,Mν1 p1,λ1;α]θ=Mν p,λ;α,(3.17) with parameters chosen such that ν0> ν > ν1and the embedding Mν0 p0,λ0;α,→ Mq,µ;α is compact while Mν1 p1,λ1;α,→ Mq,µ;αis continuous, the interpolation theory of compact operators ensures that the embedding for the intermediate space Mν p,λ;αis also compact. This completes the proof of the anisotropic compact embedding theorem. The compact embedding theorem established here has profound implications for the analysis of anisotropic partial differential equations and variational problems. It ensures that minimizing sequences for energy functionals in anisotropic spaces possess convergent subsequences, facilitating existence proofs through direct methods in the calculus of variations. Moreover, it provides the theoretical foundation for numerical analysis in anisotropic settings, guaranteeing the convergence of Galerkin approximations and finite element methods for problems with directional heterogeneity. 4 Applications to Multiscale Operator Learning This section bridges the theoretical framework developed in previous sections with modern applications in machine learning and scientific computing. The anisotropic functional inequalities and embedding theorems established earlier provide powerful mathematical tools for analyzing and understanding deep neural networks operating on highdimensional data with heterogeneous regularity. This connection between harmonic analysis and deep learning represents a significant advancement in the mathematical foundations of machine learning, offering rigorous guarantees for neural operator performance in multiscale settings. The core insight driving these applications is that many modern neural architectures naturally exhibit anisotropic behavior through their weight matrices, activation patterns, and hierarchical representations. By modeling this directional heterogeneity through the scaling vector α, we can obtain sharper stability bounds and approximation rates that adapt to the intrinsic geometry of both the data and the network architecture. This anisotropic perspective provides a principled mathematical framework for understanding why certain architectures excel at capturing multiscale features and how to design networks for specific problem classes. From a technical standpoint, the applications presented here leverage the full power of the anisotropic function space theory. The stability analysis relies on the anisotropic chain rule and Morrey norm estimates, while the approximation theory builds upon the compact embedding theorems and interpolation inequalities. The explicit dependence on the anisotropic dimension dαin our results reveals how directional scaling can mitigate the curse of dimensionality in high-dimensional learning problems. 16 4.1 Stability of Neural Operators under Anisotropic Perturbations The stability of neural networks under input perturbations is a fundamental concern in both theoretical analysis and practical deployment. In scientific computing applications, where neural operators are used to approximate solutions of partial differential equations, stability guarantees ensure that small errors in input data do not propagate catastrophically through the network. The anisotropic stability analysis developed here provides refined bounds that account for the directional sensitivity of neural operators, offering insights into how weight distributions across different coordinate directions affect robustness. The mathematical challenge in proving stability bounds for deep neural operators lies in controlling the propagation of perturbations through multiple layers while respecting the anisotropic geometry. Traditional approaches based on isotropic operator norms fail to capture the directional heterogeneity present in many practical networks. Our anisotropic framework addresses this limitation by incorporating directional weight constraints and employing anisotropic function space norms that precisely measure the directional regularity of the neural operator. Theorem 4.1 (Anisotropic Stability of Deep Neural Networks).Let Nθ:Rk→R be a neural operator with Llayers and anisotropic weight constraints ∥Wi l∥op≤λ1/αi i. Suppose the activation functions σl∈C1,1 b(R)with ∥σ′ l∥∞≤1,∥σ′′ l∥∞≤K. Then for input perturbations δx with ∥δx∥∞≤ϵ: ∥Nθ(x+δx)− Nθ(x)∥L∞≤CLϵ L Y l=1 max iλ1/αi i!∥Nθ∥M1 1,0;α,(4.1) where the constant Csatisfies: C≤K k X i=1 α−1 i!1/2p p−1dα/2 .(4.2) Proof. We develop a comprehensive stability analysis for deep neural operators in the anisotropic setting, combining layer-wise sensitivity estimates with anisotropic function space techniques. 1. Single-layer sensitivity analysis. Consider a single layer transformation y= σ(Wx +b). For an input perturbation δx, we analyze the output variation: ∥y(x+δx)−y(x)∥∞=∥σ(W(x+δx) + b)−σ(Wx +b)∥∞ ≤ ∥σ′∥L∞∥Wδx∥∞ ≤ ∥σ′∥L∞∥W∥op∥δx∥∞. Under the anisotropic weight constraint ∥Wi∥op≤λ1/αi i, we have: ∥W∥op≤max i∥Wi∥op≤max iλ1/αi i, which yields the single-layer stability bound: ∥y(x+δx)−y(x)∥∞≤max iλ1/αi iϵ. (4.3) 17 2. Multi-layer anisotropic propagation via chain rule. For the composition of Llayers Nθ=σL◦WL◦ · · · ◦ σ1◦W1, we develop an anisotropic version of the Faa di Bruno formula. The first-order derivative in direction eiis given by: D1 iNθ(x) = k X j1,...,jL=1 L Y l=1 ∂σl ∂zl Wiljl−1 l!D1 jLx, (4.4) with the convention j0=i. To prove this formula, we proceed by induction on the number of layers L. For L= 1, it reduces to the chain rule: D1 i(σ1◦W1)(x) = σ′ 1(W1x)·(W1ei). Assume the formula holds for L−1layers. Then for Llayers: D1 iNθ(x) = D1 i(σL◦WL◦ N(L−1) θ)(x) =σ′ L(WLN(L−1) θ(x)) ·WLD1 iN(L−1) θ(x), where N(L−1) θdenotes the network up to layer L−1. Applying the induction hypothesis to D1 iN(L−1) θ(x)and distributing the matrix multiplication yields the desired formula. Taking L∞norms and applying the weight and activation constraints: ∥D1 iNθ∥L∞≤ L Y l=1 ∥σ′ l∥L∞! L Y l=1 max i∥Wi l∥op!k X jL=1 ∥D1 jLx∥L∞ ≤ L Y l=1 max iλ1/αi i!∥x∥W1,∞ α. 3. Stability bound via mean value theorem and anisotropic norms. Using the mean value theorem in the anisotropic setting: |Nθ(x+δx)− Nθ(x)| ≤ sup 0≤t≤1 |DNθ(x+tδx)δx| ≤ k X i=1 ∥D1 iNθ∥L∞|δxi| ≤ϵ k X i=1 ∥D1 iNθ∥L∞. Combining with the derivative bound from Step 2: ∥Nθ(x+δx)− Nθ(x)∥L∞≤kϵ L Y l=1 max iλ1/αi i!∥Nθ∥W1,∞ α.(4.5) 4. Morrey norm control via anisotropic Landau inequality. To replace the W1,∞ αnorm with the more refined M1 1,0;αnorm, we apply the anisotropic Landau inequality (Theorem 3.1) with ν= 1: ∥Nθ∥W1,∞ α≤C(1, p, α)∥Nθ∥0 L∞ k X j=1 ∥D1,αj jNθ∥αj Lp!1 =C(1, p, α)∥Nθ∥M1 1,0;α.(4.6) 18 The constant C(1, p, α)can be bounded using the sharp constant from the anisotropic Hardy-Littlewood inequality: C(1, p, α)≤K k X i=1 α−1 i!1/2p p−1dα/2 , where the factor Kaccounts for the activation function regularity and the sum Pk i=1 α−1 i arises from the anisotropic volume scaling. 5. Constant optimization and final bound. Combining all estimates and optimizing over the parameters yields the final stability bound: ∥Nθ(x+δx)− Nθ(x)∥L∞≤kϵ L Y l=1 max iλ1/αi i!∥Nθ∥W1,∞ α ≤kC(1, p, α)ϵ L Y l=1 max iλ1/αi i!∥Nθ∥M1 1,0;α ≤CLϵ L Y l=1 max iλ1/αi i!∥Nθ∥M1 1,0;α, where we absorb the factor kinto the constant Cand note that Cscales linearly with Ldue to the product over layers. This completes the proof of the anisotropic stability theorem. The stability bound established in this theorem has significant implications for the robustness and reliability of neural operators in scientific computing applications. The explicit dependence on the anisotropic weight constraints provides guidance for network design, suggesting that balanced weight distributions across different coordinate directions can enhance stability. Moreover, the appearance of the anisotropic Morrey norm in the bound highlights the importance of the network’s directional regularity properties for ensuring stable performance under input perturbations. 4.2 Optimal Approximation Rates in Anisotropic Spaces The approximation theory of neural networks represents a fundamental bridge between classical analysis and modern machine learning, providing quantitative guarantees for the expressive power of deep learning architectures. In anisotropic function spaces, this theory takes on a particularly rich structure, as the approximation rates adapt to the directional scaling behavior encoded in the parameter vector α. This subsection establishes sharp approximation rates for neural operators in anisotropic fractional SobolevMorrey spaces, revealing how directional heterogeneity affects the efficiency of function representation and offering principled guidance for network architecture design in highdimensional problems. The mathematical foundation of our approach lies in understanding how the anisotropic dimension dα=Pk i=1 α−1 igoverns the complexity of approximation. Unlike the ambient dimension k, which appears in classical approximation theory, the anisotropic dimension dαcaptures the effective complexity of functions with heterogeneous regularity across different coordinates. When some scaling parameters αiare large, indicating higher regularity in those directions, the effective dimension dαbecomes smaller than k, leading to accelerated approximation rates that mitigate the curse of dimensionality. 19 The proof strategy combines several sophisticated techniques: anisotropic partitions of unity that respect the directional scaling, local approximation by anisotropic Taylor polynomials, and global synthesis through neural network representations. The optimality of the rates is established through metric entropy arguments that compare the complexity of the function class with the expressive power of neural networks, providing fundamental limits on what can be achieved by any approximation scheme based on a fixed number of parameters. Theorem 4.2 (Sharp Approximation Rates for Neural Operators).Let f∈ Mν p,λ;α(Rk) with ν > 0,1< p < ∞, and scaling vector α. Then there exists a neural operator Nθ with Nparameters such that: ∥f− Nθ∥L∞≤C(ν, p, λ, α)LN−ν/dα∥f∥Mν p,λ;α,(4.7) where dα=Pk i=1 α−1 iis the anisotropic dimension and the constant satisfies: C(ν, p, λ, α)≤K(p, α)Γ(ν+ 1) ν1/ν k X i=1 α−1 i!1/2 .(4.8) Moreover, this rate is optimal. Proof. We provide a comprehensive proof that constructs an approximating neural operator through local anisotropic approximations and establishes the optimality of the rate through information-theoretic arguments. 1. Anisotropic covering and partition of unity. We begin by constructing an anisotropic covering of Rkthat respects the scaling geometry. Define anisotropic cubes: Qδ,m = k Y i=1 [δαimi, δαi(mi+ 1)], m ∈Zk,(4.9) where δ > 0is a scaling parameter that will be optimized later. The anisotropic volume of each cube is: |Qδ,m|= k Y i=1 δαi=δdα, reflecting the anisotropic dimension dα=Pk i=1 α−1 i. We construct a smooth partition of unity {ϕδ,m}m∈Zksubordinate to this covering with the following properties: 1. Pm∈Zkϕδ,m(x) = 1 for all x∈Rk, 2. ∥ϕδ,m∥L∞≤1, 3. ∥D1 iϕδ,m∥L∞≤Cδ−αifor each direction ei, 4. supp ϕδ,m ⊂e Qδ,m, where e Qδ,m is a slight enlargement of Qδ,m with |e Qδ,m|≤ Cδdα. The construction uses anisotropic scaling of a fixed smooth bump function and exploits the homogeneity properties of the anisotropic dilations. 20 2. Local mixed fractional Taylor approximation. On each cube Qδ,m, let xmbe the center point. We approximate fby its anisotropic Taylor polynomial of order ν: Pδ,m(x) = X |β|α≤ν Dβ αf(xm) Qk i=1 Γ(βi/αi+ 1) k Y i=1 (xi−xm,i)βi/αi,(4.10) where |β|α=Pk i=1 βi/αiis the anisotropic degree, ensuring homogeneity under the anisotropic dilations Tα λ. The local approximation error is bounded using the anisotropic Taylor remainder theorem. For x∈Qδ,m, we have: |f(x)−Pδ,m(x)|≤ CX |β|α=ν |Dβ αf(ξx)−Dβ αf(xm)| Qk i=1 Γ(βi/αi+ 1) k Y i=1 |xi−xm,i|βi/αi, for some ξxon the line segment between xand xm. Using the anisotropic Hölder continuity of the fractional derivatives (which follows from the Morrey-Sobolev embedding), we obtain: ∥f−Pδ,m∥L∞(Qδ,m)≤C1(ν, α)δν k X i=1 ∥Dν,αi if∥αi/ν Lp(Qδ,m).(4.11) The constant C1(ν, α)incorporates the Gamma function factors from the Taylor expansion and the anisotropic Hölder constants. 3. Neural network representation and parameter counting. Each local polynomial Pδ,m can be approximated by a neural network with ReLUkactivations. By standard approximation theory for neural networks, there exists a network Nθ,m with: ∥Pδ,m − Nθ,m∥L∞(Qδ,m)≤ϵ, using at most Nm=O(ϵ−dα/ν)parameters per local approximant. The construction uses the fact that polynomials of fixed degree can be implemented exactly by deep ReLU networks, and the parameter count follows from the number of monomials in the anisotropic Taylor expansion. The global approximation is constructed as: Nθ(x) = X m∈Zk ϕδ,m(x)Nθ,m(x).(4.12) This sum is finite at each point due to the bounded overlap of the supports of ϕδ,m. 4. Error analysis and parameter optimization. The total approximation error decomposes as: ∥f− Nθ∥L∞≤sup x∈RkX m ϕδ,m(x)|f(x)−Pδ,m(x)| + sup x∈RkX m ϕδ,m(x)|Pδ,m(x)− Nθ,m(x)| ≤C1δν∥f∥Mν p,λ;α+ϵ. The number of active cubes at any point is bounded by the overlap constant K(α), which depends on the anisotropic geometry. 21 The total number of parameters satisfies: Ntotal ≤C2(α)δ−dαNm≤C3(α)δ−dαϵ−dα/ν. To optimize, we choose δand ϵto balance the error terms. Setting ϵ∼δνyields: ∥f− Nθ∥L∞≤C4(ν, α)δν∥f∥Mν p,λ;α, with parameter count: Ntotal ≤C5(α)δ−dαδ−dα=C5(α)δ−2dα. Eliminating δgives the rate: ∥f− Nθ∥L∞≤C(ν, p, λ, α)N−ν/dα∥f∥Mν p,λ;α, where the constant C(ν, p, λ, α)is given by (4.8). 5. Optimality via metric entropy. To prove the optimality of the rate N−ν/dα, we use information-theoretic arguments based on metric entropy. The ϵ-entropy of the unit ball in Mν p,λ;α(Rk)satisfies: log2N(ϵ, B1(Mν p,λ;α), L∞)≍ϵ−dα/ν,(4.13) where N(ϵ, B1, L∞)is the covering number. This follows from volume arguments adapted to the anisotropic geometry and the scaling properties of the anisotropic Sobolev-Morrey spaces. On the other hand, the class of neural operators with Nparameters has a VCdimension (or similar complexity measure) bounded by O(N2L2), where Lis the number of layers. Therefore: log2N(ϵ, NN, L∞)≤CN2L2log(1/ϵ), where NNdenotes neural operators with Nparameters. Comparing these bounds, we see that any approximation scheme achieving error ϵ must satisfy: ϵ−dα/ν ≲N2L2log(1/ϵ), which implies ϵ≳N−ν/dα(ignoring logarithmic factors). This establishes the optimality of the rate N−ν/dα. 6. Layer count and architecture details. The construction requires L=O(log(1/ϵ)) layers to implement the partition of unity and local approximations. Since ϵ∼N−ν/dα, we have: L≤C6log N. The logarithmic factor is absorbed into the constant for the final bound (4.7). The explicit form of the constant (4.8) comes from carefully tracking the dependence on ν,p,λ, and αthroughout the construction, particularly in the local Taylor approximation and the neural network implementation of polynomials. The approximation rates established in this theorem have profound implications for the theory of deep learning and high-dimensional approximation. The appearance of the anisotropic dimension dαrather than the ambient dimension kreveals that neural operators can adapt to the intrinsic complexity of functions with heterogeneous regularity, effectively mitigating the curse of dimensionality in problems where different coordinates exhibit different scaling behaviors. This theoretical insight provides mathematical justification for the empirical success of deep learning in high-dimensional scientific computing applications and offers principled guidance for network architecture design in multiscale problems. 22 5 Results This section synthesizes the principal mathematical and computational results established in this work, providing a comprehensive overview of the theoretical framework and its applications to multiscale operator learning. The development of anisotropic fractional calculus has yielded several fundamental advances that bridge classical analysis with modern machine learning, offering both deep theoretical insights and practical computational tools. The cornerstone of our theoretical framework is the introduction of anisotropic fractional Sobolev-Morrey spaces Mν p,λ;α(Rk), which provide a refined mathematical language for characterizing functions with heterogeneous regularity across different coordinate directions. These spaces incorporate directional scaling parameters α= (α1, . . . , αk)that modulate the effective smoothness along each coordinate axis, enabling precise analysis of multiscale phenomena where traditional isotropic theories prove inadequate. The rigorous construction of these spaces establishes their fundamental properties, including completeness, density results, and multiple equivalent characterizations through Gagliardo integrals, Littlewood-Paley decompositions, Bessel potentials, and heat semigroups adapted to the anisotropic geometry. Our sharp functional inequalities represent the core analytical contributions of this work. The anisotropic fractional Gagliardo-Nirenberg inequality establishes precise interpolation estimates between different levels of directional regularity, with explicit constants that capture the intricate dependence on scaling parameters. This inequality reveals how directional smoothness interacts with integrability in anisotropic settings, providing a powerful tool for trading between different norms while respecting the underlying geometric structure. Complementing this, the anisotropic Hardy-Littlewood-Sobolev inequality establishes the boundedness of directional fractional integral operators with optimal constants that depend explicitly on the anisotropic dimension dα=Pk i=1 α−1 i. These inequalities collectively provide the mathematical foundation for understanding how directional heterogeneity affects functional relationships and operator bounds. The compact embedding theorems developed in this work establish precise criteria for the compactness of embeddings between anisotropic fractional Sobolev-Morrey spaces, revealing how the anisotropic dimension dαand directional smoothness parameters govern the critical exponents for compactness. These results ensure that minimizing sequences for energy functionals in anisotropic spaces possess convergent subsequences, facilitating existence proofs through direct methods in the calculus of variations and providing the theoretical foundation for numerical analysis in anisotropic settings. In the realm of applications to multiscale operator learning, we have established rigorous stability bounds for neural operators under anisotropic perturbations. These bounds demonstrate how directional weight constraints affect the robustness of deep neural networks, providing quantitative guarantees that small input errors do not propagate catastrophically through the network architecture. The stability analysis reveals that balanced weight distributions across different coordinate directions can enhance robustness, while the appearance of anisotropic Morrey norms in the bounds highlights the importance of directional regularity properties for ensuring stable performance. Perhaps the most significant applied result concerns the optimal approximation rates for neural operators in anisotropic function spaces. We have proven that neural operators achieve approximation rates of order N−ν/dα, where Nis the number of parameters and dα is the anisotropic dimension, substantially improving upon classical isotropic rates when 23 scaling parameters are heterogeneous. This result demonstrates that neural operators can adapt to the intrinsic complexity of functions with directional regularity, effectively mitigating the curse of dimensionality in high-dimensional problems. The optimality of these rates is established through metric entropy arguments that compare the complexity of the function class with the expressive power of neural networks, providing fundamental limits on what can be achieved by any approximation scheme based on a fixed number of parameters. The computational implications of these theoretical results are profound. The explicit dependence on anisotropic scaling parameters provides principled guidance for neural architecture design in scientific computing applications, particularly for problems exhibiting multiscale features. The appearance of the anisotropic dimension dαrather than the ambient dimension kin approximation rates reveals that neural operators can leverage directional heterogeneity to achieve more efficient representations, offering mathematical justification for the empirical success of deep learning in high-dimensional scientific applications. Furthermore, the stability bounds provide rigorous foundations for the deployment of neural operators in safety-critical applications where robustness guarantees are essential. Collectively, these results establish a comprehensive mathematical framework for anisotropic fractional calculus with far-reaching implications for both theoretical analysis and computational practice. The anisotropic perspective developed here provides new pathways for addressing challenging multiscale problems across scientific disciplines, while the connections to deep learning theory bridge the gap between classical harmonic analysis and modern machine learning, contributing to the foundational mathematics underpinning the next generation of scientific computing tools. 6 Conclusions This work has established a comprehensive mathematical framework for anisotropic fractional calculus that bridges classical harmonic analysis with modern applications in multiscale operator learning. By introducing anisotropic fractional Sobolev-Morrey spaces equipped with directional scaling parameters, we have developed a refined analytical tool for characterizing functions with heterogeneous regularity across different coordinate directions. Our main theoretical contributions include the proof of sharp functional inequalities specifically anisotropic Gagliardo-Nirenberg inequalities and Hardy-LittlewoodSobolev estimates with explicit constants that capture the intricate dependence on scaling parameters. The development of advanced harmonic analysis tools, including directional Littlewood-Paley theory and anisotropic maximal function estimates, has provided the necessary foundation for establishing compact embedding theorems and interpolation results in these mixed-norm spaces. The applications to multiscale operator learning represent a significant advancement in the mathematical foundations of deep learning. We have derived rigorous stability bounds for neural operators under anisotropic perturbations, demonstrating how directional weight constraints affect robustness properties. Furthermore, our establishment of optimal approximation rates that adapt to the intrinsic anisotropic dimension dαrather than the ambient dimension kprovides theoretical justification for the empirical success of deep learning in high-dimensional problems with heterogeneous regularity. These results offer principled guidance for neural architecture design in scientific computing applica24 tions, particularly for problems exhibiting multiscale features and directional preferences. Looking forward, this work opens several promising research directions that merit further investigation. The extension to anisotropic Triebel-Lizorkin spaces would provide a more refined function space framework for analyzing nonlinear partial differential equations with anisotropic diffusion, potentially leading to sharper regularity results and improved numerical schemes. Developing stochastic versions of our theory would enable uncertainty quantification in scientific machine learning, particularly through the incorporation of anisotropic random fields that model directional dependencies in high-dimensional data. Connections to geometric deep learning on manifolds with anisotropic structures represent another fertile direction, requiring the development of intrinsic anisotropic calculus on Riemannian manifolds and graph structures. From a computational perspective, the numerical implementation of mixed norm regularization schemes poses significant challenges that warrant dedicated research. Efficient algorithms for high-dimensional problems must be developed that respect the anisotropic geometry while maintaining computational tractability. Applications to physics-informed neural networks (PINNs) for solving anisotropic partial differential equations offer immediate practical impact, as our theoretical framework provides guidance for network architecture design and stability guarantees for these widely used computational methods. Finally, extension to non-commutative settings could open new frontiers in quantum machine learning, where anisotropic structures naturally arise in the context of tensor networks and quantum many-body systems. The mixed fractional Landau inequalities and anisotropic function space theory developed in this work provide a powerful mathematical framework that bridges classical analysis and modern computational practice. By offering both deep theoretical insights into the structure of functions with heterogeneous regularity and practical guidance for algorithm design in high-dimensional settings, this research contributes to the foundational mathematics underpinning the next generation of scientific computing tools. The anisotropic perspective developed here reveals how directional scaling behavior can be leveraged to mitigate the curse of dimensionality, offering new pathways for addressing challenging multiscale problems across scientific disciplines. Acknowledgments Santos gratefully acknowledges the support of the PPGMC Program for the Postdoctoral Scholarship PROBOL/UESC nr. 218/2025. Sales acknowledges CNPq grant 30881/20250. Notation and Nomenclature Table 1: Mathematical Notation and Definitions Symbol Description Rkk-dimensional Euclidean space α= (α1, . . . , αk)Anisotropic scaling vector, αi>0 Continues on next page 25