On Dynamic Architectures For Isotropic Networks
Full text
Dynamic Topologies of Isotropic Networks George Bird Department of Computer Science & Department of Physics and Astronomy University of Manchester [email protected] August 1, 2025 1 Introduction: Neural plasticity in the quantity and connectivity of neurons is biologically advantageous and routinely occurs in the brains of animals. It can improve neural efficiency through pruning, whilst enabling the accumulation of knowledge, robustness and function through growth. It is therefore hypothesised that analogous behaviour within artificial neural networks may be similarly beneficial. Ranging from improved efficiency to reduced anomalous computations to increased capacity, the practical consequences of networks which can adapt their architecture and connectivities in real-time could be transformative. Proposed is a novel methodology which leverages an alternative formulation of primitives, known as “isotropic primitives”, to achieve real-time restructuring of networks with minimal degradation in functionality. This is achieved by exploiting a symmetry-structural functional invariance of a network. This both indicates the flexible use cases of these alternative formulations and enables the exploration of model adaptations to new information. Typically, a neural network is defined by computation predicated upon individual units intercommunicating to yield a desirable learnt function. This situates neurons, with their interconnections and computations, as ontologically fundamental building blocks to this computation system, so far. Resulting from this is an agreeable separability, and hence individuality, to neurons. This is implicitly bestowed definitionally by this contemporary perspective. They become the constructing atomic units for forming networks. The individuality of neurons is underscored by the primitives in widespread adoption: activation functions, normalisers, optimisers, all reinforce this. They collectivly continue an observable permutation symmetry to our modern networks as consequence to neuron-wise priors. Working deductivly from this current construction, one can then determine the well-known computational equivalences, exhibited as parameter-space degeneracies, which exist under exchange of neurons — formally termed a permutation symmetry, Sn , of neurons in the network. Under appropriate reparameterisations, swapping of neurons, and similarly exchanging corresponding weights and internal biases, can be considered a computational invariance of the network. In other words, the network functions identically before-and-after exchange by permuting actions. Hence, a permutation symmetry is deduced from this construction, which has emerged particularly through the contemporary constraints on primitives implicitly predicated on this perspective of individual neurons forming a network. Fundamentally, the generalised approach surrounding ‘isotropic deep learning’, reverses this ontological prior of neurons. Instead of neurons as the fundamental constituents implicitly defining this computational approach, they are emergent from symmetry. Hence, instead of deductively working from neurons leading to symmetry, this work proposes to reconsider the observed relation in reverse: symmetry leads to neurons. Philosophically, this situates foundational primitive symmetries as ontologically prior, and neurons are derived. Each primitive symmetries can then be definitional to each branch of appraoch. The ramifications of this redefinition are broad: from a broader notion of an artificial neural network system, to new sets of primitives and consequential, leveragable, behavioural changes for the network. Moreover, the notion of individuated neurons becomes an emergent property of permutation definitions in the primitives currently used. However, using this alternative perspective, one can instead consider neurons as a unit extending from the permutation definition, rather than the other way around. As a natural result, one can substitute alternative symmetries to yield new sets of primitives and consequently new notions of what a neuron is. This present work leverages one such alternative symmetry defined branch. In particular, this work explores one of these: the ‘isotropic’ redefinition predicated on the continuous orthogonal superset, O (n) , of the preexisting discrete permutation group, Sn⊂O (n) .From this, primitives emerge that are orthogonally equivariant to standard representations. From this ontological inversion, one can work deductivly to rediscover emergent neurons. For isotropic networks, the notion of a neuron is more general. Typically we may consider a layer construction as n copies of 1 -dimensional neuron-like objects concatenated, due to the permutation symmetry; however, due to the orthogonal symmetry redefinition this shifts now to considering a layer as 1copy of an n-dimensional neuron. In this approach, there is no longer a clean, individuated and agreeable definition of a neuron. There exists a gauge freedom to choose whatever decomposition to individual neurons one likes. Moreover, varied decompositions do not alter 1
the computation — the fundamental neuron is higher-dimensional and can be decomposed in any linearly combined, but normalised, way to yield equally valid alterantive bases. Moreover, this freedom to alter the basis for the represenation space extends to circumstances where one can project up to a higher-dimensional space without loss of functionality, and similarly, under more restricted conditions, one can down-project to subspaces without functionality consequences. What this amounts to is a new form of computational invariance to leverage with isotropic redefinitions — one in which the architectural structure becomes plastic and adaptable to task through reparameterisations. These are structural automorphisms for networks. This contrasts to the contemporary permutation definition which is limited to exchange symmetries which do not alter the structure. This new definition of primitives allows networks to dynamically grow and shrink to demand. This is felt to be a useful consequence directly enabled by orthogonal reformulation of network’s basic constituants. Reparameterisations allow neurons within a layer to compensate for alterations to the structure, whilst maintaining functionality. 1.1 Related Work 2 Theory This section outlines how a network can be first diagonalised to enable growth and pruning under various reparameterisations and conditions. First, a general introduction to isotropic activation functions is given. These considerations similarly apply if also using isotropic normalisations. Following this, a layer-wise basis change into a diagonalised configuration using scalar-valued decomposition will be described. Concluding the theory section, it will be discussed how the network can have growth and pruning for each layer, and a brief discussion of applications. 2.1 Isotropic Activation Function Overview Isotropy is most generally an orthogonal group family constraint applied to form a set of primitives. These primitives then transform under a particular, typically standard, representation of the orthogonal group specific to the width of the layer. For the primitives in this scenario, and for an n -width network, the standard representation is given by ρ: O (n)→ GLn(R) and each element will be notated as a matrix R=ρ(g)∈GLn(R) where g is an element of the orthogonal group, g∈O (n) . This is just the standard Rn×n representation for orthogonal (rotation-like) matrices. The representation theoretic notation will be suppressed moving forward for approachability to R∈O (n). For the isotropic activation functions functional class of study, f∈ F , this action is desired to commute with the isotropic activation functions, f:Rn→Rn , as shown in Eqn. 1. Additionally, the functions should maximally abide by this symmetry. ∀g∈O (n) [g, f] = g◦f−f◦g= 0 (1) More straightforwardly, it can be denoted as shown in Eqn. 2. f(Rx) = Rf (x)(2) Then, from this constraint, the activation function class can be constructed for this symmetry-defined primitive. A functional form which is within this functional class is displayed in Eqn. 3, where ˆx is the unit-normalised vector given the 2-norm1. f(x) = f(∥x∥2) ˆx(3) Many functions, which are non-linear in f , maximally satisfy this relation. The particularities of which function is used do not matter for dynamic network topologies. Therefore, these will just denoted as general, f(x) = σ(∥x∥) ˆx . This equivariance is demonstrated in Eqns. 4 to 7, using x′=Rx for orthogonal matrix R and that ˆx′=Rˆx remains unit-normalised. f(x′) = σ(∥x′∥) ˆx′=σ(∥Rx∥)Rˆx=f(Rx)(4) σ(∥Rx∥)Rˆx=σ√xTRTRxRˆx(5) RσpxTInxˆx=Rσ(∥x∥) ˆx(6) 1 One may notice the 2 -norm present in this description and wonder whether this approach can be generalised to l -norm. Although appropriate primitive redefinition can be achieved using this, providing functional classes within the hyperoctahedral Bn primitive set, these considerations do not generalise well to dynamic network topologies. One may be tempted to adapt the SVD diagonalisation procedure such that orthogonal group matrices are defined with a unit-norm defined by ∥a∥n ; however, there is a lack of an inner-product inducing the norm meaning that the l -norm does not define a generalised orthogonality, so one has to use the standard inner-product. This leaves only the discrete hyperoctahedral group to which the map is equivariant, and lacking the continuous group needed for this approach to projection down to smaller subspaces. Although if all incoming weights to a neuron tend to zero, then it can still be pruned. Growth remains possible regardless. 2
∴f(Rx) = Rf (x)(7) This paper employs diagonalisation of layers for its dynamic topology method. This can most clearly be seen when two such orthogonal functions, or more sequential orthogonal primitives, are composed with three affine layers. This is a distinct compositional consideration from parameter symmetry work, which typically focuses on a single discrete permutation activation function sandwiched by two affine layers. In effect, the diagonalisation considers the general and continuous orthogonal reparameterisations twice, concurrently acting on the left and right of the middle affine layer. This reparameterisation is possible due to affine layers exhibiting a left and right closure to general linear actions. Since the orthogonal group is a subset of the general linear group, then the affine layer is closed under its action and therefore can be reparameterised. Combining this with the orthogonal equivariance enables a reparameterisation where the network’s functionality is invariant. This is derivation is demonstrated in the transformations between Eqns. 8 through 12, where fAff.1 (x) = W1x + b1 and fAff.2 (x) = W2x + b2 and fIso. =σ(∥x∥) ˆx , and using the relation In=R−1R=RTR for orthogonal matrices. This is for the standard two affine with one non-linearity reparameterisation. fAff.2 ◦fIso. ◦fAff.1 =W2fW1x + b1+ b2(8) =W2fRTRW1x + b1+ b2(9) =W2RT | {z } W′ 2 f RW1 |{z} W′ 1 x +R b1 |{z} b′ 1 + b2(10) =W′ 2fW′ 1x + b′ 1+ b2(11) =W′ 2fW′ 1x + b′ 1+ b′ 2=f′ Aff.2 ◦fIso. ◦f′ Aff.1 (12) When considering the full three affine and two non-linearity compositional scenarios, a layer can be fully diagonalised by applying such transformations to the left and right of the intermediate affine layer. This diagonalisation is achieved through a singular value decomposition allowed through this two-sided gauge freedom. The diagonalised weights can then be ordered by singular value magnitude, enabling a purturbative consideration of the layer’s action — especially when encorporating normalisation. This diagonalisation procedure is discussed in the following subsection. 2.2 Layer Diagonalisation First a full procedure for diagonalising a layer will be demonstrated. This makes clear that one layer at a time can be reparameterised to express a one-to-one connectivity between the current layers neurons, individuated in a particular basis, and the preceding neurons. For a full diagonalisation this requires three affine layers interspaced around two isotropic primitives. However, growth and pruning of the network does not necessarily require a full-diagonalisation. For efficiency, several matrices can be contracted in an abridged approach to the method. This latter partial-diagonalisation will be discussed following the full-diagonalisation. Similar to before, one can consider two isotropic activation functions, or generally non-linearity, interspaced with three affine maps: • First Affine Layer: fAff.1 :Rl→Rm. A functional class with form fAff.1 (x) = W1x + b1. • First Non-Linearity: fIso.1 :Rm→Rm. A functional class with form fIso.1 (x) = σ1(∥x∥) ˆx • Second Affine Layer: fAff.2 :Rm→Rn. A functional class with form fAff.2 (x) = W2x + b2. • Second Non-Linearity: fIso.2 :Rn→Rn. A functional class with form fIso.2 (x) = σ2(∥x∥) ˆx • Third Affine Layer: fAff.3 :Rn→Ro. A functional class with form fAff.3 (x) = W3x + b3. With their composition specified by: fAff.3 ◦fIso.2 ◦fAff.2 ◦fIso.1 ◦fAff.1 :Rl→Ro . Then fAff.2 can be reexpressed in a basis where it becomes diagonalised and ordered by singular values. The diagonalisation transformations are displayed in Eqns. 13 to 17, which acts on the intermediate affine layer with a double-sided transformation similar to the previous’ section reparameterisation. In particular, the weights W2 can be represented with a singular value decomposition W2=UΣ2VT, where Uand Vare orthogonal so can commute with the non-linearity, and Σ2is a diagonal matrix. W3fIso.2 W2fIso.1 W1x + b1+ b2+ b3(13) 3
=W3fIso.2 UΣ2VTfIso.1 W1x + b1+UUT |{z} In b2 + b3(14) =W3UfIso.2 Σ2fIso.1 VTW1x + b1+UT b2+ b3(15) =W3U |{z} W′ 3 fIso.2 Σ2fIso.1 VTW1 | {z } W′ 1 x +VT b1 |{z} b′ 1 +UT b2 |{z} b′ 2 + b3(16) =W′ 3fIso.2 Σ2fIso.1 (W′ 1x +b′ 1) + b′ 2+ b3(17) This reparameterisation corresponds to the following transforms for the modified affine layers. •W′ 1=VTW1 •W′ 2=UTW2V=UTUΣ2VTV=Σ2 •W′ 3=W3U • b′ 1=VT b1 • b′ 2=UT b2 • b′ 3= b3 Where W′ 2 is now fully diagonalised, in effect the ‘neurons’ from this perspective only communicate one-to-one in this layer. This is because the generalised neuron has a gauge freedom to be expressed in many differing bases which result in an equivilant computation, as the choice of individual neurons is arbitrary as there is no map which individuates them. Additionally, this decomposition can be computed whilst ordering the singular values. This can give an indication into the importance of each weight and will be leveraged throughout the following methodology. Regularisation could be applied to these singular values as a form of weight decay if desirable. Moreover, for improved computational efficiency using sparsity, one can jointly optimise the layerwise bases changes such as to maximise the overall sparsity across all parameters. This could be achieved through working with the direct-sum of parameterised orthogonal groups, such as through its Lie algebras. 2.2.1 Partial Layer Diagonalisation Despite being illustrative, and perhaps more broadly practical and interpretable, the full diagonalisation procedure is not and dynamic topologies can procede with a simplified form which may be more practical. This reparameterisation follows more closely to the standard two affine layers surrounding an (isotropic) non-linearity. W2fIso.1 W1x + b1+ b2(18) =W2fIso.1 UΣ1VTx +UUT |{z} In b1 + b2(19) =W2UfIso.1 Σ1VTx +UT b1+ b2(20) =W2U |{z} W′ 2 fIso.1 Σ1VT | {z } W′ 1 x +UT b1 |{z} b′ 1 + b2(21) =W′ 2fIso.1 W′ 1x + b′ 1+ b′ 2(22) This reparameterisation corresponds now to the following transforms for the modified affine layers: •W′ 1=UTW1=UTUΣ1VT=Σ1VT •W′ 2=W2U • b′ 1=UT b1 4
• b′ 2= b2 This partial diagonalisation is sufficient to prune and grown neurons defined by the layer to which fIso.1 applies over. It does require that the diagonalised matrix Σ1 be retained for the thresholding. Overall, this partial diagonalisation is sufficient for the methodology, but not as illustrative as the full diagonalisations, where the one-to-one mapping is explicit and better interpretable. Both methods remain applicable for the procedures presented, with the latter partial diagonalisation enabling dynamic topologies additionally in the first layer, since no preceding layer can ‘absorb’ the V orthogonal matrix so full-diagonalisation is not possible and some improvement in efficiency due to fewer matrix multiplications. For clarity, the full-diagonalisation will be referenced moving forward; however, implementations use the partial form to additionally enable first hidden layer growth and pruning. 2.3 Dynamic Pruning Considering the intermediate affine map fAff.2 , which has been expressed in a diagonalised basis with weights ordered as singular values, W′ 2= Σ ∈Rn×m . As this diagonalised weight tends to zero, the neuron becomes fully independent of the preceding layer, Σii →0 . This enables pruning of the entire neuron, not just a connection since the neuron only has a single one-to-one connectivity due to diagonalisation. This pruning results in minimal and measurable functionality degredation for the network. However, in the current construction, a bias parameter remains and cannot be ‘gauged away’. This requires the addition of a new tunable parameter, the "intrinsic length parameter", denoted o , and there are numerous ways to interpret such a parameter, and its behaviour is largely unique to isotropic networks. To gauge away residual bias into the intrinsic length, requires a reexpressing the isotropic function, as detailed in Eqns. 23. f(x) = f(∥x∥) ˆx=f(∥x∥)x ∥x∥=f(∥x∥) ∥x∥x =g(∥x∥)x (23) Additionally expressing the affine layer as Σx + band implementing into Eqn. 23 to yield Eqn. 24. f(x) = g Σx + b Σx + b(24) One can then expand the norm-term as shown in Eqn. 25 where Σii →0 and bi indicate the ith components decomposed in the standard basis. Σx + b 2 2= min(n,m) X j=0 Σjjxj+ bj2=Σiixi+ bi2+ min(n,m) X i=j=0 Σjjxj+ bj2= b2 i+ min(n,m) X i=j=0 Σjjxj+ bj2 (25) As Σii →0 the network becomes assymptopically functionally identical to a network of a smaller ‘pruned’ size, except for the residual bias bi . The residual bi in the linear part of the isotropic activation function can be forward projected to update the biases of the following affine layer, such that it identically cancels; however, the biwithin the norm cannot. This requires a new tunable parameter: ‘intrinsic length’, which can ‘absorb’ the bias, making the network assymptopically exhibit full invariance to the action of pruning. This intrinsic length is assumed to be positive definite o= exp (λ) such that isotropic activation functions such as isotropic-tanh [] remain defined as intended; however, this positive-definite is not a strict requirement if the isotropic non-linearity is designed to permit it. Therefore, this intrinsic length can be generalised but will be assumed to be positive in the following discussion. Crucially, it can also be made trainable as a novel optimisable parameter specific for isotropic networks. Geometrically, it acts much like a bias that is orthogonal to the linear space, embedding the existing subspace away from the origin, e.g. Σx + bT o⊥= 0 and oT ⊥o⊥=o= exp (λ) — where λcan be optimised for positive definiteness. In Eqn. 26, the norm is generalised with this new parameter. Σx + b+o⊥ 2 2=oT ⊥o⊥+ min(n,m) X j=0 Σjjxj+ bj2=o+ b2 i |{z} =o′=o′T ⊥o′ ⊥ + min(n,m) X j=0 j=iΣjjxj+ bj2= Σ′x + b′+o′ ⊥ 2 2 (26) It can also be determined that this remains invariant to orthogonal group actions, preserving the overall isotropic activation function’s equivariance, fIso. (Rx) = RfIso. (x). Overall, this additional parameter enables one to gauge away residual bias using reparameterisation, further minimising function degredation during pruning. The remaining error occurs when the weight is not identically zero, but close to: Σii =ϵwith 0< ϵ ≪1. The linear aspect can be accounted for through reparameterisation of the subsequent layer. 5
2.3.1 Recursive View of Generalised Isotropy: Nested Functional Classes and Hyper-Spherical Shell Collapse The purpose of this section was initially to demonstrate the effect of normalisation on this structure; however, what was determined was that a isotropic generalised-dense network displays a nested functional class due to its inherent symmetry. In this construction, one can consider an isotropic neural network as the following series, with fIso. (x) = g(∥x∥)x x(2) =W2fIso. W1x(0) + b1 | {z } x(1) + b2(27) x(2) =W2g W1x(0) + b1 W1x(0) + b1+ b2(28) x(2) =g x(1) W2W1x(0) +g x(1) W2 b1+ b2(29) x(2) =g x(1) W′ 2x(0) +g x(1) b′ 1+ b2(30) Similarly with increasing layers: x(3) =W3g x(2) x(2) + b3(31) x(3) =W3g x(2) g x(1) W′ 2x(0) +g x(1) b′ 1+ b2+ b3(32) x(3) =g x(2) g x(1) W′ 3x(0) +g x(2) g x(1) b′′ 1+g x(2) b′ 2+ b3(33) Generalising this to yield the full isotropic dense network recursive relation. x(l)= l−1 Y i=1 g x(i) !W◦ lx(0) + l−1 X i=1 b◦ i i Y j=1 g x(j) + b◦(34) The angular dependence may appear trivial due to the W◦ lx(0) ; however, this is misleading due to it also appearing in the argument of map, g , the output angular distribution non-trivially depends on the input. However, interspacing other maps besides isotropic and affine primitives may also be explored from this view. Moreover, this can be rewritten under suitable parameter changes as Eqn. 35. Relaxing the reparameterisation constraint for Eqn. 35 and enabling parameters to take any values also broadens the notion of a dense neural network in this context, showing that isotropy generalised networks can act as a purturbative sum of decreasing sized networks. Therefore, this generalised construction enables a nested functional class and hence expressivity inclusion, much like the nested functional classes of Residual networks yet instead emerging from symmetry. Furthermore, the intermediate layers can vary in dimensionality whilst still preserving this nested functional class structure, this symmetry emergence of nested functional classes stems from primitive constraints not architectural constraints; therefore it does not constrain the layer dimensionality like Residual Networks classically do. x(l)= l−1 X i=1 W◦ ix(0) + b◦ i i Y j=1 g x(j) (35) To make clear, this is not a classical dense network but is a larger encompassing but derivative architecture including it. It displays a natural nested functional class, and a purtubative-like expansion which can recursivly includes shallower networks. If g:R+→[0, α] , then this also acts like a purturbative expansion, allowing convergence if α≤1 . Similarly normalisation can be used to implement a series convergence. This approach may also allow a new route to recursive reimplementation of backward and forward propagation methods. The initial intention was to show the application of normalisation to this system and how this may mitigate the input-error under pruning if placed carefully. One may consider how the residual input error term from pruning neurons, after residual bias has been gauged away, may continue to negativly impact performance. One could consider preor post-composing activation functions with a hyper-spherical normalisation such that each layer produces a vector with a fixed norm. This can be seen to be an undesirable placement, despite mitigating the input-error under pruning, since all the non-linearities become constant and the layer becomes linear and equivilant under reparameterisation to Eqn. 36 — which is guaranteed not to have a universal approximation theorem due to the constraint to affine maps. This suggests this form of hyperspherical shell normalisation collapses expressibility to affine maps. x(l)=Wˆx(0) + b(36) 6
Therefore, normalising to a hyperspherical shell must not happen occur within an isotropic layer, as this guarantees linear expressibility by the network. Instead, either generalised normalisation, which does not project to the hyperspherical shell such as the proposed Chi-Normalisation [], which can be preor post-composed with isotropic activations which is likely a preferable form of normalisation. These constrain the magnitude of xi in the product Σiixi , to ensure that as Σii →0 it cannot be compensated by a growth in xiwhich may otherwise adversily affect the network’s functionality. Generally, one should also consider the preor postcomposure of normalisations that can also impact vanishing or exploding gradients within these generalised isotropic architectures. 2.4 Dynamic Growth: SCAFFOLDING NEURONS Dynamic growth is practically trivial; one can add zero-initialised weights and biases with no effect. These will become trained through backpropagation and may activate. This can also be undertaken in anisotropic networks. One can define a threshold and require two neurons below that singular value threshold at any time. Then, if the threshold is exceeded, perform pruning; if insufficient neurons meet the threshold, perform growth. These can be computed every epoch of less for computational tractability. 2.5 Overview of Implementation: 2.5.1 Implementation Considerations 2.6 Notes: One can consider these actions can be undertaken leftwards or rightwards, going from Rn→Rm to Rn−1→Rm or Rn→Rm−1, but is displayed in rightward fashion 7