scieee AI-readable full text Open interactive document viewer

On The Nature of Individuated Neurons: Symmetry-Redefinions Enable Dynamic Topologies

Bird, George

Full text

On The Nature of Individuated Neurons Symmetry-Redefinions Enable Dynamic Topologies George Bird Department of Computer Science & Department of Physics and Astronomy University of Manchester [email protected] September 29, 2025 Abstract A novel method is described and implemented for dynamic network topologies by leveraging isotropic primitives. Using such primitives, the notion of an individuated neurons is lost allowing a freedom in how layers are represented and interpreted. By considering a layer in a diagonalised representation one can modify the structure of the network, appending and removing neurons, in response to task neccesity whilst the network remains largely computationally invariant due to the symmetry-structure. 1 Introduction Neural plasticity in the quantity and connectivity of neurons is biologically advantageous [] and routinely occurs in the brains of animals []. It can improve neural efficiency through pruning [], whilst enabling the accumulation of knowledge [], robustness [] and function through growth []. It is therefore hypothesised that analogous behaviour within artificial neural networks may be similarly beneficial. This may range from improved efficiency, to reduced anomalous computations, to increased capacity — overall a real-time restructuring for knowledge acquisition and stabilisation. The practical consequences of a network that can adapt their architecture and connectivities in real-time would likely be transformative. Proposed is a novel methodology that leverages an alternative formulation of primitives, known as “isotropic primitives”, to achieve real-time restructuring of networks with minimal degradation in functionality. This is achieved by exploiting a symmetry-structural functional invariance of a network. This both indicates the flexible use cases of these alternative formulations and enables the exploration of model adaptations to new information. Typically, one may define a neural network system as a computation predicated upon individual units intercommunicating to yield an overall desirable learnt function. This situates neurons, with their interconnections and computations, as ontologically fundamental building blocks to this computation system. In other words, this model is constituted by and defines such delineated artificial neurons, at least thus far. These then communicate and adapt to reproduce desired function through collective computation. Currently the design core of this approach is an agreeable separability, and hence individuality, to those neurons, which become the atomic units for forming ever-larger networks. These units interplay to produce the broad functional class of artificial neural networks we employ today. Mathematically, this individuality is bestowed by the contemporary choices of network operations, where the individuality of neurons is assumed into implicit constraints on the possible functional forms. This affects most primitives in widespread adoption, including, but not limited to, the majority of: activation functions, normalisers, optimisers, several regularisers, gradient-clippings — all of which collectively reinforce this neuron-wise persepective. The question may arise: why do we often choose to apply our activation functions elementwise? Extending from this current construction, one can work deductively, determining the computational equivalences [] inherent to most networks. Exhibited by such models are parameter-space degeneracies, usually centered around the exchange of neurons — formally termed a permutation symmetry, Sn , of the neurons in the network. Under appropriate reparameterisations, swapping of neurons, and similarly exchanging corresponding weights and internal biases, can be considered a computational invariance of the network 1 . In other words, the network functions identically before and after exchange by permuting actions. Hence, a permutation symmetry is deduced from this current construction. It fundamentally emerges from the implicit constraint of a neuron-wise functional form for network operations, this is due to a notion of what constitute a neural network model. 1 This is also connected to the symmetries of the discrete directed graphs underlying the communication pathways between models predicated on discrete, individuated neurons. These will later be generalised continuously [] 1 However, this present work argues that such a perspective of individuated neurons, and their resultant constraints on functional forms 2 , is limiting. The functional class of all artificial neural networks is being confined based on assumed fundamental individuality. Where these contemporary primitive’s forms have potentially stemmed from historical naturalistic ideals, reinforced later through hardware alignment, positive performance precedent, alongside some cognitive entrenchement [] — often appealing to its biological counterparts for intuition. Yet, by abandoning this implicit constraint, broadening the possible functional class, it will be shown that the behaviours and flexibility of neural network models are notably enriched, in addition to recently established representation changes []. Fundamentally, this is the generalised approach surrounding ‘isotropic deep learning’ and associated primitivesymmetry redefinitions []. These proposed redefinition approaches effectivly constitute an ontological reversal of the neurons priors. Instead of current neurons as the fundamental constituents which implicitly define this computational approach, from which deduced permutation-like symmetries emerge, the foundational base will be inverted. Hence, instead of working deductively from neurons leading to symmetry, this work proposes to reconsider the observed relation in reverse: symmetry leads to neurons. Current networks then emerge from one symmetry: the permutation group. Yet this reversal allows other symmetries to take its place, which then generate new forms for primitives and hence the notion of neurons. Philosophically, this situates primitive symmetries as foundational and hence ontologically prior. Neurons as a notion are then derived/emergent from them, a qualitative concept resulting from the mathematically altered functional forms. Each primitive symmetry can then be definitional for each branch of a new approach. A new set of primitives for each foundational symmetry and linked through them. Comparable in intuition, to the substantivly similar physical ontological inversion which occured between individuated particles to symmetry-defined fields. The ramifications of this redefinition are argued to be broad: from a generalised notion of an artificial neural network computational system, extending beyond their biologically-inspired premise, to new sets of primitives and consequential, leveragable, behavioural and representational changes for networks. The functional class of artificial neural network is broadened through such a shifted ontology. Symmetries are no longer derivative or confined to functional invariances, or implicated as an externalist model inductive bias construct, but become internally foundational to the computational system. Moreover, the notion of individuated neurons becomes an emergent property of permutation definitions in the primitives currently used. However, using this alternative perspective, one can instead consider neurons as a unit extending from the permutation definition, rather than the other way around. As stated, As a natural next-step, one can substitute alternative symmetries to yield new sets of primitives and consequently new notions of what a neuron even is. This present work leverages one such alternative symmetry-defined branch in particular. This work explores the ‘isotropic’ redefinition, an approach predicated on the continuous orthogonal superset, O (n) , of the preexisting discrete permutation group, Sn⊂O (n) . From this, primitive sets emerge that are orthogonally equivariant to standard representations of these actions []. From this ontological inversion, one can work deductively to rediscover emergent neurons and leveragble new behaviours. For isotropic networks, the notion of a neuron is more general. Typically, we may consider a layer construction as n copies of 1 -dimensional neuron-like objects concatenated, due to the permutation symmetry; however, due to the orthogonal symmetry redefinition, this shifts now to considering a layer as 1 copy of an n -dimensional neuron. The fundamental qualitative ‘neuron’ is higher-dimensional and can be decomposed in any linearly combined, but normalised, way to yield equally valid alternative individuated bases. Crucially, in this approach, there is no longer a clean, individuated and agreeable definition of a neuron. There exists effectivly a (global) gauge freedom to choose whatever decomposition of individual neurons one likes, there is no distinguishing mathematical boundary delineating them. Any choice of individual neurons is no more fundamental than any other, they may shift arbitrarily in definition due to their continuous orthogonal nature. Hence, varied decompositions do not alter the computation when interspaced between affine maps — yielding familiar computational invariances under affine composition. Moreover, this freedom to alter the basis for the representation space extends to circumstances where one can project up to a higher-dimensional space without loss of functionality (a neurogenesis abaility), and similarly, under more restricted conditions 3 , one can down-project to subspaces without functionality consequences and catastrophic failing, providing the network a neurodegenerative/pruning ability. All of which through a symmetry-principled approach. What this amounts to is a new form of computational invariance to leverage with isotropic redefinitions — one in which the architectural structure becomes plastic and adaptable to tasks through reparameterisations. These are structural automorphisms for networks. This contrasts with the contemporary permutation definition, which is limited to exchange symmetries which do not alter the structure. Primitive symmetries, neural architecture and functional invariance can become desirably intertwined. Hence, this new definition of primitives enables networks to expand and contract in response to demand, dynamically. This is considered a beneficial consequence directly enabled by the orthogonal reformulation of the network’s basic constituents. 2Which likely cyclically reinforces such a perspective 3One which may be additionally moderated through weight-decay regularisations. 2 This paper explores such reparameterisations, allowing neurons within a layer to compensate for alterations to the structure while maintaining functionality. Through this insights such as, [NOTE HERE], are uncovered. 1.1 Related Work In this section, comparison to other approaches for dynamic network architecture will be undertaken, with a summary of each followed by key similarities and differences. Continual Learning: RigL 2020 Net2Net Magnitude pruning SET 2018 Lottery Ticket Hypothesis 1.2 Concluding 2 Theory This section outlines how a network can be diagonalised to enable growth and pruning under various reparameterisations and conditions. First, a general introduction to isotropic activation functions is given. These considerations similarly apply if also using isotropic normalisations. Following this, a layer-wise basis change into a diagonalised configuration using scalar-valued decomposition will be described alongside a partial diagonalisation procedure which is more efficient. Concluding the theory section, it will be discussed how the network can have growth and pruning for each layer, and a brief discussion of applications. 2.1 Isotropic Activation Function Overview Isotropy is most generally an orthogonal group family constraint applied to form a set of primitives. These primitives then transform under a particular, typically standard, representation of the orthogonal group specific to the width of the layer. For the primitives in this scenario, and for an n -width network, the standard representation is given by ρ: O (n)→ GLn(R) and each element will be notated as a matrix R=ρ(g)∈GLn(R) where g is an element of the orthogonal group, g∈O (n) . This is just the standard Rn×n representation for orthogonal (rotation-like) matrices. The representation theoretic notation will be suppressed moving forward for approachability to R∈O (n). For the isotropic activation functions functional class of study, f∈ F , this action is desired to commute with the isotropic activation functions, f:Rn→Rn , as shown in Eqn. 1. Additionally, the functions should maximally abide by this symmetry. ∀g∈O (n) [g, f] = g◦f−f◦g= 0 (1) More straightforwardly, it can be denoted as shown in Eqn. 2. f(Rx) = Rf (x)(2) Then, from this constraint, the activation function class can be constructed for this symmetry-defined primitive. A functional form which is within this functional class is displayed in Eqn. 3, where ˆx is the unit-normalised vector given the 2-norm4. f(x) = f(∥x∥2) ˆx(3) Many functions, which are non-linear in f , maximally satisfy this relation. The particularities of which function is used do not matter for dynamic network topologies. Therefore, these will just denoted as general, f(x) = σ(∥x∥) ˆx . This equivariance is demonstrated in Eqns. 4 to 7, using x′=Rx for orthogonal matrix R and that ˆx′=Rˆx remains unit-normalised. f(x′) = σ(∥x′∥) ˆx′=σ(∥Rx∥)Rˆx=f(Rx)(4) 4 One may notice the 2 -norm present in this description and wonder whether this approach can be generalised to l -norm. Although appropriate primitive redefinition can be achieved using this, providing functional classes within the hyperoctahedral Bn primitive set, these considerations do not generalise well to dynamic network topologies. One may be tempted to adapt the SVD diagonalisation procedure such that orthogonal group matrices are defined with a unit-norm defined by ∥a∥n ; however, there is a lack of an inner-product inducing the norm meaning that the l -norm does not define a generalised orthogonality, so one has to use the standard inner-product. This leaves only the discrete hyperoctahedral group to which the map is equivariant, and lacking the continuous group needed for this approach to projection down to smaller subspaces. Although if all incoming weights to a neuron tend to zero, then it can still be pruned. Growth remains possible regardless. 3 σ(∥Rx∥)Rˆx=σ√xTRTRxRˆx(5) RσpxTInxˆx=Rσ(∥x∥) ˆx(6) ∴f(Rx) = Rf (x)(7) This paper employs diagonalisation of layers for its dynamic topology method. This can most clearly be seen when two such orthogonal functions, or more sequential orthogonal primitives, are composed with three affine layers. This is a distinct compositional consideration from parameter symmetry work, which typically focuses on a single discrete permutation activation function sandwiched by two affine layers. In effect, the diagonalisation considers the general and continuous orthogonal reparameterisations twice, concurrently acting on the left and right of the middle affine layer. This reparameterisation is possible due to the functional class of affine layers exhibiting left and right closure under general linear actions. Since the orthogonal group is a subset of the general linear group, the affine layer is closed under orthogonal actions. It can therefore be reparameterised to produce an equivalence class of functionally identical models. Combining this with orthogonal equivariance enables a reparameterisation where the network’s functionality remains invariant. This is derivation is demonstrated in the transformations between Eqns. 8 through 12, where fAff.1 (x) = W1x+ b1 and fAff.2 (x) = W2x+ b2 and fIso. =σ(∥x∥) ˆx , and using the relation In=R−1R=RTR for orthogonal matrices. This is for the standard two-affine with one non-linearity reparameterisation. fAff.2 ◦fIso. ◦fAff.1 =W2fW1x + b1+ b2(8) =W2fRTRW1x + b1+ b2(9) =W2RT | {z } W′ 2 f  RW1 |{z} W′ 1 x +R  b1 |{z}  b′ 1   + b2(10) =W′ 2fW′ 1x + b′ 1+ b2(11) =W′ 2fW′ 1x + b′ 1+ b′ 2=f′ Aff.2 ◦fIso. ◦f′ Aff.1 (12) When considering the full three affine and two non-linearity compositional scenarios, a layer can be fully diagonalised by applying such transformations to the left and right of the intermediate affine layer. This diagonalisation is achieved through a singular value decomposition allowed through this two-sided gauge freedom. The diagonalised weights can then be ordered by singular value magnitude, enabling a perturbative consideration of the layer’s action — especially when incorporating normalisation. This diagonalisation procedure is discussed in the following subsection. 2.2 Layer Diagonalisation First, a procedure for fully diagonalising a layer will be demonstrated. This makes clear that one layer at a time can be reparameterised to express a one-to-one connectivity between the current layer’s neurons, individuated in a particular basis, and the preceding layer’s neurons. For a full diagonalisation, this requires three affine layers interspaced around two isotropic primitives. However, growth and pruning of the network do not necessarily require a full diagonalisation. For efficiency, several matrices can be contracted in an abridged approach to the method. This latter partial diagonalisation will be discussed following the full diagonalisation. Similar to before, one can consider two isotropic activation functions, or generally non-linearity, interspaced with three affine maps: • First Affine Layer: fAff.1 :Rl→Rm. A functional class with form fAff.1 (x) = W1x + b1. • First Non-Linearity: fIso.1 :Rm→Rm. A functional class with form fIso.1 (x) = σ1(∥x∥) ˆx • Second Affine Layer: fAff.2 :Rm→Rn. A functional class with form fAff.2 (x) = W2x + b2. • Second Non-Linearity: fIso.2 :Rn→Rn. A functional class with form fIso.2 (x) = σ2(∥x∥) ˆx • Third Affine Layer: fAff.3 :Rn→Ro. A functional class with form fAff.3 (x) = W3x + b3. 4 With their composition specified by: fAff.3 ◦fIso.2 ◦fAff.2 ◦fIso.1 ◦fAff.1 :Rl→Ro . Then fAff.2 can be reexpressed in a basis where it becomes diagonalised and ordered by singular values. The diagonalisation transformations are displayed in Eqns. 13 to 17, which act on the intermediate affine layer with a double-sided transformation similar to the previous section’s reparameterisation. In particular, the weights W2 can be represented with a singular value decomposition W2=UΣ2VT, where Uand Vare orthogonal so can commute with the non-linearity, and Σ2is a diagonal matrix. W3fIso.2 W2fIso.1 W1x + b1+ b2+ b3(13) =W3fIso.2  UΣ2VTfIso.1 W1x + b1+UUT |{z} In  b2 + b3(14) =W3UfIso.2 Σ2fIso.1 VTW1x + b1+UT b2+ b3(15) =W3U |{z} W′ 3 fIso.2   Σ2fIso.1   VTW1 | {z } W′ 1 x +VT b1 |{z} b′ 1   +UT b2 |{z}  b′ 2   + b3(16) =W′ 3fIso.2 Σ2fIso.1 (W′ 1x +b′ 1) + b′ 2+ b3(17) This reparameterisation corresponds to the following transforms for the modified affine layers. •W′ 1=VTW1 •W′ 2=UTW2V=UTUΣ2VTV=Σ2 •W′ 3=W3U • b′ 1=VT b1 • b′ 2=UT b2 • b′ 3= b3 Where W′ 2 is now fully diagonalised, in effect the ‘neurons’ from this perspective only communicate one-to-one in this layer, as depicted in Fig. 1. This is because the generalised neuron has a gauge freedom to be expressed in many differing bases, which results in an equivalent computation, as the choice of individual neurons is arbitrary, as there is no map which individuates them. Before Diagonalisation After Diagonalisation Figure 1: This illustration depicts the qualitative effects on a network from full diagonalisation — the chosen layer has a double-sided basis change to the connectivity, reducing the map to a one-to-one correspondence between neurons. This drastically simplifies the interrelations between the chosen layer, allowing the application of the dynamic network implementation. In general, for sequential layers, only one layer can be diagonalised at a time, as diagonalising one layer often destroys the diagonalised state of the immediately preceding and following layers in the process. Layers which are interspaced by other affine transforms can be concurrently diagonalised. Additionally, this decomposition can be computed whilst ordering the singular values. This can give an indication of the importance of each weight and will be leveraged throughout the following methodology. Regularisation could be applied to these singular values as a form of weight decay if desirable. Moreover, for improved computational efficiency using sparsity, one can jointly optimise the layerwise basis changes, such as to maximise the overall sparsity across all parameters. This could be achieved through working with the direct sum of parameterised orthogonal groups, such as through their Lie algebras. 5 2.2.1 Partial Layer Diagonalisation Despite being illustrative and perhaps more broadly practical and interpretable, the full diagonalisation procedure is not, and dynamic topologies can proceed with a simplified form, which may be more practical. This reparameterisation follows more closely the standard two affine layers surrounding an (isotropic) non-linearity. This partial diagonalisation can occur in two forms: left-sided and right-sided, corresponding to the following parameterised maps: mixing followed by scaling or scaling followed by mixing, respectively. The right-sided transform is explicitly given in Eqns. 18 through 22. This also influences which direction the pruning and growth acts, forward or backwards, which are both permissible applications under this construction. W2fIso.1 W1x + b1+ b2(18) =W2fIso.1  UΣ1VTx +UUT |{z} In  b1 + b2(19) =W2UfIso.1 Σ1VTx +UT b1+ b2(20) =W2U |{z} W′ 2 fIso.1   Σ1VT | {z } W′ 1 x +UT b1 |{z}  b′ 1   + b2(21) =W′ 2fIso.1 W′ 1x + b′ 1+ b′ 2(22) This right partial diagonalisation is given by the following reparameterisations, which correspond to a modified affine layers: •W′ 1=UTW1=UTUΣ1VT=Σ1VT •W′ 2=W2U • b′ 1=UT b1 • b′ 2= b2 This partial diagonalisation is sufficient to prune and grow neurons defined by the layer to which fIso.1 applies over. It does require that the diagonalised matrix Σ1 be retained for the thresholding. Overall, this partial diagonalisation is sufficient for the methodology, but not as illustrative as the full diagonalisations, where the one-to-one mapping is explicit and better interpretable. Both methods remain applicable for the procedures presented, with the latter partial diagonalisation enabling dynamic topologies additionally in the first layer, since no preceding layer can ‘absorb’ the V orthogonal matrix, so full-diagonalisation is not possible and some improvement in efficiency due to fewer matrix multiplications. For clarity, the full diagonalisation will be referenced moving forward; however, implementations use the partial form to additionally enable growth and pruning of the first hidden layer. Before Diagonalisation Left-Partial Diagonalisation Right-Partial Diagonalisation Figure 2: This illustration depicts the qualitative effects on a network from partial diagonalisation. This time, the chosen layer has a single-sided basis change to the connectivity. Depending on the transform side, leftor right-sided, an initial or later mixing of connectivities occurs, followed by scaling by the singular values. This setup may be more convenient to implement than a full diagonalisation. Some interpretability and explanatory convenience are lost due to the remaining connectivity mixing. 6 2.3 Dynamic Pruning Considering the intermediate affine map fAff.2 , which has been expressed in a diagonalised basis with weights ordered as singular values, W′ 2= Σ ∈Rn×m . As this diagonalised weight tends to zero, the ‘neuron’ becomes entirely independent of the preceding layer, Σii →0 . This enables the pruning of the entire neuron, not just a connection, since the neuron has only a single one-to-one connectivity due to diagonalisation. This pruning results in minimal and measurable degradation of network functionality. However, in the current construction, a bias parameter remains and cannot be ‘gauged away’. This requires the addition of a new tunable parameter, the "intrinsic length parameter", denoted o , and there are numerous ways to interpret such a parameter, and its behaviour is largely unique to isotropic networks. To gauge away residual bias into the intrinsic length requires reexpressing the isotropic function, as detailed in Eqns. 23. f(x) = f(∥x∥) ˆx=f(∥x∥)x ∥x∥=f(∥x∥) ∥x∥x =g(∥x∥)x (23) Additionally expressing the affine layer as Σx + band implementing into Eqn. 23 to yield Eqn. 24. f(x) = g  Σx + b  Σx + b(24) One can then expand the norm-term as shown in Eqn. 25 where Σii →0 and  bi indicate the ith components decomposed in the standard basis.   Σx + b   2 2= min(n,m) X j=0 Σjjxj+ bj2=Σiixi+ bi2+ min(n,m) X i=j=0 Σjjxj+ bj2= b2 i+ min(n,m) X i=j=0 Σjjxj+ bj2 (25) As Σii →0 , the network’s norm term becomes asymptotically functionally identical to a network of a smaller ‘pruned’ size, except for the residual bias  bi . The residual  bi in the linear part of the isotropic activation function can be forward projected to update the biases of the following affine layer, such that it identically cancels; however, the  bi within the norm cannot. This requires a new tunable parameter: ‘intrinsic length’, which can ‘absorb’ the bias, making the network asymptotically exhibit full invariance to the action of pruning. This intrinsic length is assumed to be positive definite o= exp (λ) such that isotropic activation functions such as isotropic-tanh [] remain defined as intended; however, this positive-definite property is not a strict requirement if the isotropic non-linearity is designed to permit it. Therefore, this intrinsic length can be generalised but will be assumed to be positive in the following discussion. Crucially, it can also be made trainable as a novel optimisable parameter specific for isotropic networks. Geometrically, it acts much like a bias that is orthogonal to the linear space, an intuition is that it is embedding the existing subspace away from the origin, e.g. Σx + bT o⊥= 0 and oT ⊥o⊥=o= exp (λ) — where λ can be optimised for positive definiteness. This does have a rotational degeneracy in the orthogonal complement space, if interpreted in such a manner. Except for the following activation function, this complement space is projected out before the following affine layer, returning the vector space to its prior dimensionality. Overall, it may offer a novel and beneficial trainable parameter for an isotropic network, and if negative, may allow direction flipping, which may increase network support for manipulating representations into antipodal arrangements where desirable. It could also be reinterpreted as a parameterised activation function instead of an additional orthogonal offset parameter for the affine map. In Eqn. 26, the norm is generalised with this new parameter.   Σx + b+o⊥   2 2=oT ⊥o⊥+ min(n,m) X j=0 Σjjxj+ bj2=o+ b2 i |{z} =o′=o′T ⊥o′ ⊥ + min(n,m) X j=0 j=iΣjjxj+ bj2=  Σ′x + b′+o′ ⊥   2 2 (26) It can also be shown that this remains invariant to orthogonal group actions, preserving the overall isotropic activation function’s equivariance, fIso. (Rx) = RfIso. (x). Overall, this additional parameter enables one to gauge away residual bias using parameter degrees-of-freedom, further minimising function degradation during whole neuron pruning. This accounts for the normalisation term, but the linear aspect can be similarly accounted for through reparameterisation of the subsequent layer accomodating the pruned bias into the subsequent biases. However, functional degradation does remain when the pruned diagonalised weight is not identically zero, but close to: Σii =ϵ with 0< ϵ ≪1 . This can be mitigate using an isotropic normalisation composed with the activation function. There are two standard approaches which can be used to achieve this: layer-normalisation and batch-normalisation. These are discussed in App. A, which had the additional consequence of revealing an isotropic architecure which inherently 7 displays a depthwise nested functional class due to the isotropic terms — this is significant as reproduces the nested functional class structure in an alternative way to the typical Residual Network construction and in a way which doesn’t constrain layer dimensionality. Additionally, this forward correction for the bias in the linear term is similarly applicable to any remaining Σii =ϵ term. If one considers the linear term, with original diagonalised weights Σ and its pruned form Σ′ alongside subsequent weight matrix before and after pruning, W(2) and Y(2) respectivly, then we desire the following map to be closely preserved: Y(2)Σ′x ≈W(2)Σx . Therefore, one can find a suitable adjusted weight Y(2) through a pseudo-inverse, where applicable, such as to minimise a least-squares difference in these maps. This is displayed in Eqn. 27, if an inverse can be determined (if not one may continue to use the original W(2) with the associated column pruned.). If it is the smallest singular value which is deleted, as intended, then this reduces to a corresponding column deletion of Y(2). This will be further detailed in Sec. 2.5 which outlines the overall implementation. Y(2) =W(2)ΣΣ ′TΣ′Σ ′T−1 (27) In summary, generalising layer-normalisation to isotropy requires one to not normalise by the mean, as this results in a change in intrinsic geometry: Rn→Rn−1,→Rn [], while dividing by the standard-deviation projects to a hyper-spherical shell. In App. A this is shown to result in affine expressibility of the network, rendering it limited in its application. Therefore, a batch-normaliser such as the Chi-normaliser [] may be employed to reduce function degredation under pruning. Similarly, one can consider the gradient with respect to the diagonalised weight as a threshold for pruning — this is also effectivly a batch-statistic. Pruning should always begin with the smallest singular value or loss-gradient magnitude. 2.4 Dynamic Growth Dynamic growth is comparativly trivial. If acting forwards, it amounts to embedding the last layer’s space into a higher dimensional vector space. This may be interpreted in two ways for its implications on neurons. The first, is that if neurons are chosen to be individuated in some arbitrary basis, it corresponds to appending additional neurons to a layer which are connected, although the model is functionally independent of them — these could be considered ‘scaffold neurons ’in this conventional individuated picture. Alternativly, it is comparable to treating the higher-dimensional generalised notion of a neuron to have further increased in dimensionality. Due to the isotropic activation function’s jacobian acting to distribute learning gradients, alongside connectivity mixing, these apparant individuated neurons may be rapidly trained despite being functionally independent of the original model. This can occur even when connecting parameters are zeroed. This can be seen compartivly for a typical anisotropic activation function’s diagonal jacobian compared to an isotropic non-diagonal jacobian in Eqns. 28 and Eqn. 29 respectivly for a n -‘neuron’ layer, f:Rn→Rn and standard basis-vectors ˆei . The singularity at x = 0 can become a non-troublesome coordinate singularity under suitable choices of σ , f( 0) =  0 ; however, requires a custom implementation to prevent autodiff issues which is available in the code repository at . f(x) = N X i=1 σ(x ·ˆei) ˆei⇒∂(f·ˆei) ∂(x ·ˆej)=σ′(x ·ˆei)δij (28) f(x) = σ(∥x∥) ˆx⇒∂(f·ˆei) ∂(x ·ˆej)=σ(∥x∥) ∥x∥δij +σ′(∥x∥)−σ(∥x∥) ∥x∥(x ·ˆei) (x ·ˆej) ∥x∥2(29) Below, Eqn. 30, shows a reforumlated isotropic activation funciton and its jacobian. f(x) = g(∥x∥)x ⇒∂(f·ˆei) ∂(x ·ˆej)=g(∥x∥)δij +g′(∥x∥)xixj ∥x∥(30) One can see, that Eqn. 29 and Eqn. 30, for the isotropic activation function’s jacobian, contains non-diagonal terms in addition to the normal diagonalised derivative — this may increase gradient flow through these scaffold neurons hastening their operationalisation. The way in which such neurons are initialised depends on the diagonalisation picture being considered. This is in practicality non-trivial and choices can range in effect by explicit or spontaneous symmetry breaking. This arises since the reparameterisations are only functionally identical on the forward pass, they interact and are non-equivilant with gradient descent algorithms during updates. This is a potentially insightful direction for future study. Moreover, after the diagonalisation procedure, the two occurances of either U on the following layer or amending the bias term, similarly for V on the backwards approach, can become decoupled, especially evident under neuronal growth. These considerations are discussed in App. B. Both the weights and biases associated with that scaffold neuron may be initialised in a variety of ways, with implications for learning and resulting function. In either case, the new singular values should be set to zero, yet the surrounding affine maps must also changes dimensionality to accomodate this new ‘neuron’. This may require a new column or row depending on which direction the growth acts in. How these are initialised does not affect the network’s 8 functionality; however, for convention their SVD decomposition can have a new row or column added that is orthogonal to the existing span — this can be achieved through the Gram-Schmidt procedure. 2.5 Overview of Implementation In addition to growth and pruning considerations, it must be decided when to apply them. What will be discussed is a functional buffer of neurons, consisting of two primary hyperparameters: the number of scaffold neurons, Ξ∈Z+ , and singular-value threshold,ϑ∈R+. The threshold value indicates at which point the singular value connections are considered significant to performance, below this value their inclusion is considered neglible or perhaps contributes rarily-used, non-robust adaptations. It would typically be chosen to be a small positive number. The number of scaffold neurons indicates the buffer of neurons which is maintained below this threshold. If it is set to Ξ=2 , then two scaffold neurons are maintained in the network at any time, i.e. two diagonalised neurons below the singular-value threshold. If more neurons are below the threshold then neuronal pruning can occur; if fewer, neuronal growth occurs. These can be computed every epoch of less for computational tractability. In the following two subsections, a step by step methodology is provided on how to prune or grow neurons in dense networks. The extension to convolution growth and pruning of kernels is exactly analogous, treating the kernels channelwise dense networks. 2.5.1 Forward Adaption This subsection outlines the full forward adaption process such that it can be easily implemented, the backward adaption is given in App. C. The forward process is when considering an affine map Rn→Rm , then the following layer characterised by m -dimensionality can be increased or decreased. Several of these steps and reparameterisations can be contracted for computational efficiency, whether or not these contractions are implemented affects gradient flow as discussed in App. B. First perform a left-partial diagonalisation of the weight parameters using singular value decompisition, W1=UΣVT . This reparameterises from Eqn. 31 to 32, as previously demonstrated5. W2fIso.1 W1x + b1;o+ b2(31) =W2UfIso.1 Σ1VTx +UT b1;o+ b2(32) The dimensionality of these matrices is given by: U∈Rm×m , VT∈Rn×n , Σ∈Rm×n ≥0 , W2∈Rp×m ,  b1∈Rm ,  b2∈Rpand o∈R≥0, . The next step is to threshold the singular values to determine how many are below the threshold, ϑ . These are the entries along the diagonal. If the number of scaffold neurons falls below the desired count, |{Σii|Σii < ϑ}| <Ξ then iterativly add neurons until the condition is equalised. This is iterative growth (neurogenesis) of the network. This is described first below with Ξ increasing by one6. In this case proactive contractions of the matrices can be considered, which also highlights that the two instances of U can be decoupled as shown in Eqn. 33, where dashes indicate the new parameterisations in contrast with the prior. =W′ 2fIso.1 Σ′ 1V ′Tx + b′ 1;o′+ b′ 2(33) From this equation, one implements neurogenesis through the following steps. 1. For matrix, Σ∈Rm×n ≥0 , a row of zeros should be appended to the bottom of the matrix to form a matrix Σ′∈ R(m+1)×n ≥0. This is given by Σ′=Σ  0T 2. For vector,  b1∈Rm this is modified with a new last entry, b∗ to form  b′ 1∈R(m+1) . Additionally, b∗ must be chosen such that b2 ∗+o′=o, forming the non-negativity restriction b2 ∗< o. The new vector is of form b′ 1= b1 b∗. 3. For scalar, o∈R≥0, an adjustment follows accordingly from the prior step: o′=o−b2 ∗. 4. For matrix, VT∈Rn×n, no change need be made. 5 One does not need to undo the diagonalisation after this procedure the basis change leaves the network functionally equivilant. Both training and diagonalising other surrounding layers will destroy the diagonalisation of the specified layer, there is no purpose to do this manually. Although the basis change will couple differently to dynamic optimisers and certain initialisations for grown layers will prevent reversing diagonalisation exactly. 6 Again this threshold can be redefined to be the magnitude with respect to loss as an alterantive loss-based approach as effectivly a mean importance statistic over a batch. A logarithm before thresholding could be taken to make order-of-magnitude scales comparable. Overall, this is less ideal as it doesn’t indicate the consequence on loss if the entire neuron is grown or pruned, only the local gradient. 9