scieee AI-readable full text Open interactive document viewer

One Bit at Every Scale: A Cross- Disciplinary Analysis of Uncertainty Transitions

Pudsey, Veronika

Abstract

Many physical, biological, and computational systems can be described as transitioning from a probability distribution to a realized state. Using Shannon’s definition of entropy, this paper quantifies the informational content of such transitions across four domains: quantum physics, cosmology, large language models (LLMs), and neural systems. For each domain, we identify a representative elementary transition, derive or extract the relevant probability distribution, and compute its entropy in bits. Across spin measurements, tunneling measurement outcomes, inflationary modes, black-hole area elements, next-token distributions in transformer models, EEG microstates, MEG phase states, and decision-related neural activity, the entropy of an elementary transition consistently falls within a narrow range of approximately one to two bits. The comparison distinguishes between transitions whose distributions are directly measurable (quantum experiments, cosmological spectra, electrophysiological data) and transitions whose distributions are effective or inferred (internal gating events in LLMs). The resulting numerical regularity therefore reflects similarities in probabilistic structure rather than shared physical mechanisms. While these systems differ fundamentally in scale and substrate, the convergence in informational scale is empirically robust across independently motivated definitions. Whether this pattern reflects deeper physical constraints or methodological commonalities remains an open question warranting further investigation. This descriptive informational comparison provides a quantitative reference for cross-domain analysis. Future work may explore whether the informational scale of state transitions supports methodological transfer between fields or reveals deeper structural constraints on how systems reduce uncertainty.

Full text

1 One Bit at Every Scale: A CrossDisciplinary Analysis of Uncertainty Transitions Veronika Pudsey Independent Researcher [email protected] Abstract Many physical, biological, and computational systems can be described as transitioning from a probability distribution to a realized state. Using Shannon’s definition of entropy, this paper quantifies the informational content of such transitions across four domains: quantum physics, cosmology, large language models (LLMs), and neural systems. For each domain, we identify a representative elementary transition, derive or extract the relevant probability distribution, and compute its entropy in bits. Across spin measurements, tunneling measurement outcomes, inflationary modes, black-hole area elements, next-token distributions in transformer models, EEG microstates, MEG phase states, and decision-related neural activity, the entropy of an elementary transition consistently falls within a narrow range of approximately one to two bits. The comparison distinguishes between transitions whose distributions are directly measurable (quantum experiments, cosmological spectra, electrophysiological data) and transitions whose distributions are effective or inferred (internal gating events in LLMs). The resulting numerical regularity therefore reflects similarities in probabilistic structure rather than shared physical mechanisms. While these systems differ fundamentally in scale and substrate, the convergence in informational scale is empirically robust across independently motivated definitions. Whether this pattern reflects deeper physical constraints or methodological commonalities remains an open question warranting further investigation. This descriptive informational comparison provides a quantitative reference for cross-domain analysis. Future work may explore whether the informational scale of state transitions supports methodological transfer between fields or reveals deeper structural constraints on how systems reduce uncertainty. 2 1 Introduction Across scientific disciplines, systems are often described in terms of the possibilities available to them and the rules by which one possibility becomes actual. Quantum measurements, thermal processes, cosmological fluctuations, biological state transitions, and modern machine learning models all operate on probability distributions. These distributions encode uncertainty, and when a system transitions from a set of possible outcomes to an observed or realized state, the uncertainty is reduced. Shannon’s information theory provides a natural measure for this reduction. If the probabilities of the possible outcomes are known, the uncertainty can be expressed as an entropy measured in bits. This makes it possible to compare systems that are physically dissimilar but structurally similar in how they resolve uncertainty. In physics, probability distributions arise from fundamental laws: quantum amplitudes, statistical mechanics, and gravitational thermodynamics. In cosmology, early-universe fluctuations, blackhole entropy, and Hawking radiation all have well-defined statistical structures that allow their uncertainty to be quantified. In computational systems such as large language models (LLMs), explicit next-token distributions can be measured directly, while internal transitions such as gating or pattern selection can be analysed through effective or inferred probabilistic descriptions. Neural systems exhibit transitions between activity patterns that can be modelled probabilistically and have measurable entropies in electrophysiological data. Although these domains differ widely in scale, mechanism, and substrate, the informational scale of their elementary transitions shows a recurring regularity: the entropy reduction associated with a single transition commonly lies within a narrow range of approximately one to two bits. This observation applies both to systems where the relevant probability distributions are directly observable and to systems where they are inferred from the structure of the underlying model. The aim of this paper is to systematize this observation. We compile representative examples from quantum physics, cosmology, machine learning, and neuroscience; express their state transitions in terms of Shannon entropy; and compare the informational scale across systems. The analysis is descriptive rather than interpretive: we focus on the quantitative structure of uncertainty reduction rather than proposing a unifying physical mechanism. The rest of the paper is organized as follows. Section 2 reviews Shannon’s definition of entropy. Section 3 outlines the method used to translate domain-specific transitions into bits. Sections 4–7 present examples from the four selected domains. Section 8 compares the results at both micro and minimal macro scales. Sections 9 and 10 discuss implications and limitations. Section 11 concludes. We emphasize that this work is descriptive: we document a pattern, not explain it. Whether the convergence reflects fundamental constraints or methodological artifacts is an open question we hope to stimulate investigation of. 3 2 Background: Shannon Information and Entropy Information theory provides a quantitative way to describe uncertainty in systems that can evolve into one of several possible outcomes. In its classical formulation, due to Shannon, a system with a set of outcomes {i} occurring with probabilities 𝑝𝑖 has an associated entropy 𝐻 =−∑𝑝𝑖 𝑖log⁡2𝑝𝑖. Entropy measures the expected information gained when the actual outcome becomes known (Shannon, 1948). The use of the base-2 logarithm expresses this quantity in bits—the minimal units required to resolve uncertainty through binary distinctions. A system with two equally probable outcomes has 𝐻 =1bit; a system with three equiprobable outcomes has 𝐻 =log⁡23≈1.58bits; uneven distributions yield lower entropy according to the probabilities. This formulation is substrate-independent: it does not rely on the physical nature of the system, only on the structure of its probability distribution. Whether the uncertainties arise from quantum amplitudes, thermal ensembles, cosmological fluctuations, neural dynamics, or computational models, the entropy of a distribution is computed in the same way and has the same informational interpretation (MacKay, 2003; Cover & Thomas, 2006). Entropy also provides a natural measure for analysing state transitions. When a system moves from a probability distribution to a realized outcome—or, in computational contexts, when a model commits to an internal functional regime governed by an effective distribution—the reduction in uncertainty corresponds to an amount of information expressed in bits. This enables direct comparison of transitions across domains that differ widely in physical scale and implementation. In this work, we apply Shannon’s formulation to representative systems in quantum physics, cosmology, machine learning, and neuroscience. For each system, we identify the relevant probability distribution, compute the entropy associated with its elementary transition, and express the result in bits. This provides the basis for the cross-scale comparison developed in later sections. 3 Methods: Translating Heterogeneous Systems into a Common Informational Measure The systems considered in this work differ substantially in their physical nature and mathematical description. To compare them on a common informational scale, we adopt a unified method based on three steps: 1. Identify a probability distribution over possible outcomes. 2. Compute the entropy of this distribution using Shannon’s definition. 3. Interpret an elementary transition as the reduction of this entropy to a single realized outcome. 4 Quantum probabilities arise from the Born rule (Nielsen & Chuang, 2010); cosmological probabilities from Gaussian inflationary fluctuations (Baumann, 2009) or thermal spectra of Hawking radiation (Hawking, 1975); machine-learning probabilities from softmax-normalized logits (Vaswani et al., 2017); and neural probabilities from empirical transition matrices between brain states (Michel & Koenig, 2018). Although the sources of these distributions differ, the method applied to them is uniform. 3.1 Identifying the relevant probability distribution For each system, we identify the probability distribution that governs a single elementary uncertainty-resolution step. In some domains, such as quantum mechanics, cosmology, and neuroscience, these distributions reflect directly observable stochastic events. In others, such as large language models, only certain transitions (e.g., next-token sampling) are directly probabilistic, while finer-grained internal operations can be treated as effective low-dimensional probabilistic approximations based on established interpretability analyses. Quantum systems: Measurement outcomes with Born-rule probabilities 𝑝𝑖=∣𝜓𝑖∣2(Nielsen & Chuang, 2010). Cosmological systems: Gaussian primordial fluctuations during inflation (Baumann, 2009); entropy per horizon-area element in black-hole thermodynamics (Bekenstein, 1973; Hawking, 1975); thermal emission spectra of Hawking radiation (Page, 1976). Large language models: Directly measurable: next-token probability distributions (Vaswani et al., 2017). Effective approximations: internal gating, activation and pattern-selection behaviors treated as coarse-grained low-dimensional alternatives (Elhage et al., 2021; Olah et al., 2021; Nanda & Chan, 2023). Neural systems: Transition probabilities between EEG microstates (Koenig et al., 2002; Michel & Koenig, 2018); phase-based metastable MEG states (Vidaurre et al., 2017); pre-decision cortical distributions (Gold & Shadlen, 2007; Hanks et al., 2015). Only distributions corresponding to an identifiable transition are used. Continuous distributions are discretized using procedures standard in each domain. 3.2 Computing entropy in bits For any discrete distribution {𝑝𝑖}, entropy is computed as 𝐻 =−∑ 𝑖𝑝𝑖log⁡2𝑝𝑖. This quantity measures the average uncertainty prior to the transition (Shannon, 1948). When the system resolves to a specific outcome, the uncertainty is reduced by 𝐻bits; we interpret this value as the informational scale of the elementary transition. For continuous distributions (e.g., Gaussian fields in cosmology), we compute entropy per mode following standard treatments (Baumann, 2009) and convert natural or Boltzmann units to bits via: 5 bits =𝑆 𝑘𝐵ln⁡2. For macroscopic entropies such as the Bekenstein–Hawking formula, 𝑆 =𝑘𝐵𝑐3𝐴 4𝐺ℏ , we express the entropy in bits using: bits =𝑆 𝑘𝐵ln⁡2 (Bekenstein, 1973; Hawking, 1975). This gives the effective number of independent binary distinctions encoded in the system. 3.3 Interpreting transitions A “transition” refers to any step in which a system resolves a set of probabilistic alternatives into a realized outcome. This includes: • collapse of a quantum state, • emission of a Hawking quantum, • selection of a token by an LLM, • shift between neural microstates. For internal computations in deterministic systems such as LLMs, we interpret certain operations (e.g., gating regimes, attention-pattern modes) as effective probabilistic distinctions when supported by established circuit-level analyses. This does not assert intrinsic randomness; rather, it treats functionally distinct computational regimes as coarse-grained alternatives that carry an associated uncertainty reduction. We do not assume any shared physical substrate or mechanism across domains. The only common requirement is: A definable probability distribution—empirical or effective—must govern the alternatives resolved by the transition. 3.4 Comparability and limitations To ensure that results are comparable across domains, we use the following constraints: • only transitions at the smallest available scale are included; • entropy is computed for a single transition event, not for dynamics over time; • probabilities are taken from established theoretical or empirical work; • no assumptions are made about underlying physical substrates. 6 This method isolates the informational unit associated with an elementary state-change, allowing direct comparison of systems that otherwise differ in scale, dynamics, and implementation. 3.5 Selecting elementary decision units The choice of an “elementary transition” is guided by three uniform criteria across all domains: The event must be associated with a well-defined probabilistic description. This may arise from fundamental physical laws (Born-rule amplitudes, thermal distributions), empirically measured transition matrices (neural microstates), explicit model outputs (LLM token probabilities), or effective low-dimensional approximations supported by mechanistic analyses (LLM internal regimes). The event must correspond to a complete uncertainty-resolution step. A transition must reduce a probability distribution to a realized outcome: measurement result, token selection, microstate switch, emission event, or a coarse-grained computational regime. We do not require intrinsic physical randomness—only that the model formalism identifies a discrete alternative resolved at that step. The event must be the smallest informationally meaningful decision unit defined in the domain. Smaller sub-events (e.g., membrane potential fluctuations, individual matrix multiplications, subgating arithmetic, sub-Planckian area elements) are excluded when the formalism does not assign them distinct probabilistic outcomes. Likewise, larger aggregates (e.g., phrases, whole-brain states, multi-step measurement sequences) are excluded at the micro scale but may appear in macroscale analysis. Unified definition: Across all domains considered here, an elementary transition is defined as the smallest step at which a model—physical, statistical, or computational—represents multiple alternatives with a well-defined probability distribution and resolves them into a single realized outcome or an effective functional regime. 4 Quantum Systems Quantum systems provide the clearest setting in which uncertainty is inherently described by a probability distribution. The Born rule assigns outcome probabilities {𝑝𝑖}to measurement results via the squared amplitudes of the quantum state (Nielsen & Chuang, 2010). A measurement corresponds to a transition from a superposition to a single realized outcome, and the informational content of that transition is given by the Shannon entropy of the probability distribution associated with the measurement. Below we examine several representative quantum transitions for which the outcome probabilities are well defined and experimentally central. 4.1 Spin-½ measurement A spin-½ particle measured along a fixed axis has two possible outcomes, commonly denoted ↑and ↓. If the initial state is maximally uncertain with respect to the measurement basis—for example, 7 ∣𝜓⟩= 1 √2(∣↑⟩+∣↓⟩) then the Born rule assigns equal probabilities 𝑝↑=⁡𝑝↓=⁡1 2. The entropy of this distribution is: 𝐻 =−(1 2log21 2⁡+⁡1 2log21 2)=1⁡bit. This represents the informational content resolved by an ideal projective spin measurement. Such binary-outcome measurements are standard in quantum information experiments (Sakurai & Napolitano, 2017). 4.2 Higher-dimensional systems: qutrit measurement For a three-level (qutrit) system with equal amplitudes, the Born-rule probabilities are: 𝑝𝑖=1 3,𝑖 ∈{1,2,3}. The entropy is: 𝐻 =log⁡23≈1.585 bits. This represents the maximal uncertainty for a qutrit projective measurement. Uneven amplitudes produce correspondingly smaller entropy values. The general relation between dimensionality and entropy is a standard result in information theory (Cover & Thomas, 2006). 4.3 Quantum tunneling Quantum tunneling refers to the evolution of a wavefunction encountering a potential barrier. The tunneling process itself is continuous and does not constitute a discrete probabilistic event. The informational transition arises only when the system is measured after interacting with the barrier, yielding a binary outcome: the particle is detected on the incident side (reflected) or on the far side (transmitted). In this measurement framework, the relevant probability distribution is 𝑝trans =𝑇(𝐸),𝑝refl =1−𝑇(𝐸). where 𝑇(𝐸)is the transmission coefficient determined experimentally or via the WKB approximation (Gamow, 1928; Merzbacher, 1998). For typical regimes where 𝑇(𝐸)∈[0.2,0.8], the Shannon entropy of the measurement outcome lies between approximately 0.72 and 0.97 bits. When experimental parameters are tuned to yield 𝑇(𝐸)= 0.5, the entropy reaches its maximum value of 1 bit. 8 Thus, although tunneling is a continuous dynamical process, the measurement associated with its outcome resolves into a binary probabilistic event whose informational content is approximately one bit—consistent with other elementary two-outcome measurements in quantum mechanics. 4.4 Ising spins: thermal uncertainty The Ising model provides a minimal setting in which thermal fluctuations generate probabilistic state changes. Each spin 𝑠𝑖∈{−1,+1} interacts with its neighbors via the Hamiltonian 𝐻=−𝐽∑𝑠𝑖𝑠𝑗 ⟨𝑖,𝑗⟩ − ℎ∑𝑠𝑖 𝑖, where 𝐽 is the coupling and ℎthe external field (Ising, 1925; Yeomans, 1992). At finite temperature 𝑇, the update rule considers the current state 𝑠𝑖and computes the probability of flipping to −𝑠𝑖(versus staying at 𝑠𝑖), given the local effective field ℎeff. Denoting 𝑝flip =1 𝑍exp⁡(−𝛽 Δ𝐸),𝛽 =1/(𝑘𝐵 𝑇), where Δ𝐸is the energy change if 𝑠𝑖were flipped (Newman & Barkema, 1999; Glauber, 1963). The relevant information-measure for an elementary update is then the Shannon entropy of the twooutcome flip/no-flip distribution: 𝐻 =−𝑝fliplog⁡2𝑝flip −(1−𝑝flip)log⁡2(1−𝑝flip). Across typical temperature regimes, this entropy lies around one bit of uncertainty (approximately 0.7–1.2bits). At high temperatures (𝑇 ≫𝐽), 𝑝flip ≈1/2, giving 𝐻 ≈1bit, while moderate temperature introduces deviations from 0.5and slightly lower entropy. Near the critical point (𝑇 ≈𝑇𝑐), fluctuations again maximize uncertainty, and entropy approaches ∼1bit (Yeomans, 1992; Huang, 1987). Thus, a single probabilistic update in the Ising model resolves on the order of one bit of information, placing it in line with other systems in which elementary transitions occupy a similar informational scale. 4.5 Holevo-bound transitions When classical information is transmitted using quantum states, the accessible information is bounded by the Holevo quantity: 𝜒=𝑆(𝜌)−∑𝑝𝑖𝑆(𝜌𝑖) 𝑖, which upper-bounds the mutual information between sender and receiver (Holevo, 1973). For binary ensembles, the accessible information is strictly limited by a quantity on the order of 1 bit unless orthogonal states are used. 9 The Holevo bound thus characterizes another class of quantum transitions whose informational scale is naturally in the range of 1–1.5 bits, depending on ensemble purity. Summary of quantum transitions Across quantum measurements, qutrit systems, tunneling processes, and Holevo-limited communication, the entropy resolved by an elementary quantum transition consistently lies between 1 and 2 bits. While the mechanisms differ—projective collapse, barrier penetration, or communication constraints—the informational scale of the transition remains in a narrow, welldefined range. 5 Cosmological Systems Cosmology involves physical processes occurring at scales vastly larger than those in quantum mechanics, yet several key early-universe and gravitational phenomena admit precise statistical descriptions. These probability distributions allow the associated uncertainties to be expressed using Shannon entropy and, consequently, in bits. We focus on three representative settings where the statistical structure is well characterized: primordial fluctuations generated during inflation, black-hole entropy, and Hawking radiation. 5.1 Primordial fluctuations In the standard inflationary paradigm, quantum fluctuations in scalar fields are stretched to cosmological scales and become the seeds of temperature anisotropies observed in the cosmic microwave background (Mukhanov, Feldman, & Brandenberger, 1992; Baumann, 2009). The inflaton perturbations are well approximated by a Gaussian random field, with each Fourier mode characterized by an amplitude drawn from a normal distribution. For a Gaussian mode with variance 𝜎𝑘 2, the differential entropy is 𝐻nat =1 2ln⁡(2𝜋𝑒𝜎𝑘 2). For cosmological scalar fluctuations, the observed dimensionless power spectrum at the pivot scale is 𝒫ℛ(𝑘∗)≈𝐴𝑠≈2.1×10−9 (Planck Collaboration, 2018). In standard treatments of Gaussian inflationary modes (Mukhanov, Feldman & Brandenberger, 1992; Baumann, 2009), the differential entropy associated with an individual Fourier mode depends on the choice of mode normalization, volume factors, and window functions. Under normal observational conventions, these choices shift the entropy only by an additive constant, and the resulting values fall within an order-of-magnitude range of about 1–2 bits 1 1 Sketch of the calculation. For the comoving curvature perturbation ℛ(𝐱), the standard convention defines the power spectrum via 16 Empirical decision entropies in neurophysiology fall within precisely these ranges. Interpretation Neural decision systems are dynamical processes, but just before commitment (the boundarycrossing event), their state can be represented as a probability distribution over a discrete set of possible actions. This distribution may be inferred either: 1 directly from neural activity, or 2 from the parameters of a cognitive decision model. In both approaches, the resulting uncertainty is quantified by Shannon entropy, which: • is structurally equivalent to the informational measures used in the other domains in this paper, and • consistently lies within the same range of approximately 1–2 bits per elementary transition. 7.4 Interpretational constraints Neural dynamics are continuous at the biophysical level, and the discrete states used in this analysis arise from data-driven clustering or task-defined alternatives. The informational quantities considered here therefore describe transitions between empirically stable macro-states rather than fundamental physical units. Under these methodological constraints, the entropies of these transitions are well defined and directly comparable to the uncertainty reductions observed in other domains. Summary of neural transitions Across EEG microstates, MEG phase states, and decision-related cortical activity, the entropy associated with an elementary transition lies within approximately 1–2 bits. While the underlying mechanisms differ substantially, the magnitude of uncertainty reduction per transition is consistent with the scales observed in quantum, cosmological, and computational systems. 8 Results: Cross-Domain Comparison We applied a unified informational analysis to representative systems in quantum physics, cosmology, machine learning, and neuroscience. For each system, we identified a single transition event, extracted or reconstructed the associated probability distribution, and computed its Shannon entropy in bits. Table 4 summarises the informational scale of elementary transitions across domains as computed in Sections 4–7. Each value reflects the Shannon entropy of a probabilistic step defined by the modelling framework of the relevant system. Table 4. Elementary transition entropies across systems (Numbers correspond to Shannon entropy of the relevant probability distribution for a single transition. Ranges reflect empirical variation or model-dependent parameters.) 17 Domain System / Transition Entropy (bits) Source type Quantum physics Spin-½ measurement 1.0 Born probabilities (equiprobable) Spin-1 (qutrit) measurement 1.58 Born probabilities (equiprobable) Post-barrier measurement after tunneling (T/R) 0.7–1.0 Transmission/reflection measurement distribution Ising spin update (Glauber/Metropolis) 0.7–1.2 Thermal update probabilities Cosmology Inflationary mode amplitude ≈1–2 Gaussian mode entropy (per mode) Bekenstein–Hawking area element ≈1.4 Entropy per (4\ell_P^2) area (≈1.44 bits) Hawking quantum emission ≈1 Thermal occupation distribution Large Language Models Internal activation/gating (effective) ≈1 Effective low-dimensional alternatives Attention-pattern selection (effective) 1–2 Dominant-pattern uncertainty Neural systems EEG microstates 1.6–1.9 Stationary + transition entropies MEG phase states 1–2 HMM-derived transitions Decision-related cortical states (2– 4 alternatives) 1–2 Choice-probability entropies 8.1 Cross-domain pattern Across all examined systems, the entropy of an elementary transition lies within a narrow range of approximately 1–2 bits for directly probabilistic physical and biological transitions, and within a comparable low-bit range for effective low-dimensional transitions in engineered systems such as large language models. This convergence appears across: • physical systems governed by fundamental laws, • cosmological processes with well-defined statistical structure, • artificial neural networks where internal computation exhibits low-dimensional functional regimes, and • biological neural systems with emergent probabilistic state transitions. Despite substantial differences in substrate and mechanism, the magnitude of uncertainty reduction per elementary step is similar. 18 8.2 Macro-Scale Validation: Robustness Across Scales The micro-scale transitions analysed in Sections 3–7 represent the smallest informational events for which each system provides a well-defined probability distribution. These transitions are determined by the modelling frameworks themselves and reflect the elementary “decision boundaries’’ at which uncertainty is reduced to a single realised outcome. To test whether the convergence observed at this scale is specific to elementary transitions or persists across higher organisational layers, we examine macro-scale decision units: events that arise when multiple micro decisions are integrated into a single, coherent outcome. The goal is not to analyse arbitrarily large structures, but to identify the smallest possible aggregated units that resolve uncertainty at a higher scale. A macro-scale decision unit is therefore defined by three criteria: It is the next organisational layer above the micro transition. The unit must be the minimal functional aggregate that arises from the same underlying process—e.g., a token above internal activation events in an LLM, a perceptual integration window above neural microstates, or a coarsegrained mode above primordial fluctuations. Larger structures (such as phrases, behavioural sequences, or cosmological volumes) are excluded because they exceed the first aggregation layer. It corresponds to a recognised functional or physical unit in the relevant domain. The unit must be grounded in empirical or theoretical usage rather than being an arbitrary grouping. Examples include next-token sampling in language models, one-second windows in cognitive neuroscience, short sequential measurement blocks in quantum protocols, and Hubble-scale patches in inflationary cosmology. It resolves uncertainty into a single observable outcome. Like the micro transitions, macro units must complete a probabilistic-to-deterministic step, but over a larger temporal or spatial extent. These constraints allow us to test whether the informational structure observed at the micro scale is preserved across the first higher level of organisation in each system. In Section 8.3, we show that these minimal macro units exhibit informational magnitudes in the same oneto few-dozen-bit range, indicating that the observed convergence is not an artefact of using only the smallest available transitions. 8.3 Macro-Scale Results: Minimal Aggregated Decision Units Across Systems To evaluate whether the informational structure identified at the micro scale persists beyond elementary transitions, we examined the smallest aggregated decision units arising from the same underlying processes in each system (as defined in Section 8.2). These macro-scale units represent the first organisational layer above micro transitions: the smallest coherent events that integrate many elementary steps into a single resolved outcome. Table 5 summarises the selected macro-scale units. Each corresponds to an empirically recognised structure within its domain. Although these units span vastly different physical substrates and temporal scales, their informational magnitudes remain within a similar range. 19 Table 5. Minimal macro-scale decision units across systems System Micro-scale unit Macro-scale unit (first aggregated layer) Approx. entropy (bits) Notes Large language models Internal gating / activation (~1 bit) Token (next-token distribution) ~2–7 Token entropy empirically measured in contemporary LLMs across varied contexts (Holtzman et al., 2020; Meister et al., 2020). Neural systems Spike / microstate transition (~1 bit) Conscious integration window (~1 s) ~5 Working-memory and perceptualcapacity limits cluster around 3–7 items/bits (Miller, 1956; Cowan, 2001; Dehaene, 2014). Quantum systems Two-level measurement (1 bit) Short sequential measurement block ~3–5 Multi-step interferometric protocols yield a small finite set of coherent outcome patterns (Aharonov & Vaidman, 1991; Ma et al., 2012). Cosmology Primordial fluctuation (~1 bit per mode) Hubble-scale coarse-grained mode ~3–6 Effective entropy per observable mode constrained by cosmic variance (Baumann, 2009; Planck Collaboration, 2018). Across all domains, minimal macro-scale units exhibit informational magnitudes in the few-bit to few-dozen-bit range, despite differing by many orders of magnitude in physical scale. For example, token-level entropy in modern LLMs averages around 2–7 bits in naturalistic contexts (Holtzman et al., 2020; Meister et al., 2020), while conscious integration windows in humans support approximately five bits of behaviourally relevant information (Miller, 1956; Cowan, 2001; Dehaene, 2014). Sequential quantum measurement blocks similarly encode only a few bits of distinguishable outcomes (Aharonov & Vaidman, 1991), and cosmological modes are limited by cosmic variance to a comparable informational window (Baumann, 2009; Planck Collaboration, 2018). Crucially, although each macro-scale event aggregates many micro transitions, the resulting informational content increases only modestly. This suggests that most micro transitions are not independent: they form correlated or redundant patterns that collapse into a relatively small number of coherent macro outcomes. The persistence of the same informational band across both micro and minimal macro scales indicates that the convergence observed in Sections 3–7 is not a consequence of analysing only elementary transitions, but a structural feature of systems that process uncertainty through hierarchical probabilistic decisions. 20 9 Discussion The results presented above show that quantum, cosmological, machine-learning, and neural systems exhibit elementary state transitions whose Shannon entropies fall within a narrow interval of approximately one to two bits. This suggests a structural similarity in how uncertainty is resolved at minimal scales of change, without implying shared physical mechanisms or comparable ontological status of the transitions. In several engineered systems (such as LLM internals), the relevant distributions arise not from direct physical sampling but from effective low-dimensional functional regimes; nevertheless, they can still be quantified informationally. 9.1 Interpreting the similarity in informational scale Several factors may help explain why such different systems converge on a similar informational range: Constrained effective dimensionality. Across domains, only a limited number of states can be reliably distinguished: Hilbert-space projections in quantum measurements, noise-limited neural microstates, dominant internal functional patterns in transformers (rather than full token distributions), and observationally resolvable cosmological modes. These constraints compress the effective state space and naturally bound transition entropy. Robustness and stability requirements. Systems exposed to decoherence, thermal fluctuations, neural variability, or numerical instability tend to settle into a small set of stable attractors. This discretization produces probability distributions with restricted support and limited entropy. Information-processing efficiency. Extremely low-entropy transitions lack flexibility; extremely high-entropy transitions are costly to resolve. Many systems may operate in a regime where uncertainty is neither minimal nor maximal, producing transitions that routinely fall within the 1–2 bit range. These perspectives are not mutually exclusive. Together, they point toward convergent structural constraints rather than coincidental numerical alignment. 9.2 Distinguishing empirical regularity from unifying principle The observed 1–2 bit range does not imply that these systems share a common physical origin. Entropy depends on how states are defined, how distributions are estimated, and at what granularity transitions are measured. In some engineered systems (LLM activations, attention heads), the distributions represent effective decision boundaries rather than literal stochastic events. In physical systems (quantum measurements, cosmological modes), the distributions are fundamental. Critical caveat: The reported entropy values are sensitive to these choices. Different discretizations, normalization conventions, or state-space partitions would yield different numerical results. We do not claim that the ~1-2 bit scale is a universal constant independent of how transitions are defined. What we DO claim as non-trivial is the following: When elementary transitions are identified using domain-standard methods—Born-rule probabilities in quantum mechanics, empirical clustering in 21 neuroscience, softmax outputs in LLMs, thermodynamic modes in cosmology—the resulting entropies converge to a similar range despite: • Differing by ~60 orders of magnitude in physical scale • Arising from unrelated probability-generating mechanisms • Being described by distinct theoretical frameworks • Having no a priori reason to share informational structure This convergence across independently motivated definitions constitutes the empirical regularity we document. Whether it reflects: (a) a deep structural constraint on observable events, (b) a common feature of systems capable of stable information processing, or (c) a methodological artifact requiring further scrutiny remains an open empirical and theoretical question. The findings therefore document a quantitative regularity across heterogeneous probability structures, not a universal mechanism. 9.3 Implications for cross-domain modeling The regularity suggests several avenues for methodological transfer and theoretical exploration: Transferable analytical tools. Techniques from one domain—information bottleneck theory, decoherence modeling, or Markov state analysis—may apply to others with comparable transition entropies. Informational units of comparison. Systems can be compared through the number of elementary transitions they support, independent of physical substrate. Predictive design principles. Artificial systems operating far from this entropy range may be inefficient or unstable. Constraints for future theories. Any framework linking information processing across domains must account for the observed convergence while remaining agnostic about the physical nature of the transitions themselves. These implications are exploratory but indicate that cross-domain comparisons may have quantitative grounding. 9.4 Why this regularity is nontrivial The similarity of transition entropies is not a trivial consequence of Shannon’s formalism or the selection of systems with few outcomes. Heterogeneous state definitions. States arise from unrelated procedures—Hilbert projections, Fourier modes, transformer functional regimes, or neural clustering—and nothing in these definitions requires similar effective dimensionality. 22 Unrelated probability-generating mechanisms. Born-rule amplitudes, Gaussian random fields, softmax logits, and empirical neural transition matrices stem from distinct theories. Shared entropy values do not follow from the mathematics alone. Many probabilistic systems lie far outside the 1–2 bit range. High-dimensional ensembles, full token distributions in LLMs, unrestricted Markov chains, and biased transitions routinely yield entropies well above or below this interval. Bits are non-native units in several domains. Converting cosmological or neural quantities into Shannon bits requires explicit assumptions. That independently derived conversions align numerically is empirical, not definitional. Alternative granularities would change the scale. Coarser or finer state definitions—multi-particle quantum states, cosmological volumes, full-layer neural activations, or whole-brain patterns— produce radically different entropy values. Convergence occurs despite extreme differences in scale and substrate. The systems span ~60 orders of magnitude and operate in quantum fields, spacetime geometry, silicon circuits, and biological tissue. The similar informational scale of elementary transitions is therefore empirical and nontrivial. Together, these points show that the convergence is not an automatic artefact of probabilistic modeling or unit choice, but a structural pattern across heterogeneous systems. 9.5 Summary The 1–2 bit scale appears consistently across systems with heterogeneous state definitions, unrelated probability mechanisms, and radically different physical implementations. It is not dictated by Shannon entropy, and many probabilistic processes lie far outside this range. The convergence observed here is therefore a substantive empirical regularity—one that invites explanation, rather than a mathematical inevitability. 9.6 Macro-scale robustness The observed convergence at the micro scale is reinforced by the fact that the same informational range persists at the first aggregated layer above elementary transitions (Section 8.3). Each system exhibits a minimal macro-scale decision unit—token-level sampling in language models, perceptual integration windows in neural systems, short sequential measurement blocks in quantum mechanics, and coarse-grained cosmological modes—that integrates many micro transitions into a single resolved outcome. Despite this aggregation, the resulting informational magnitudes remain within the same oneto fewdozen-bit interval. This suggests that micro transitions are often correlated or redundant, collapsing into a small number of coherent macro outcomes. The persistence of comparable informational scales across layers supports the interpretation that the convergence reflects shared probabilistic constraints rather than artefacts of measurement, discretization, or scale selection. 23 10 Limitations The findings presented in this paper are descriptive and rely on a series of methodological simplifications. Several limitations should be considered when interpreting the results. 10.1 Dependence on state definitions Across all domains, the entropy values depend on how states are defined: • In quantum systems, the measurement basis determines the probability distribution. • In cosmology, the choice of mode decomposition or horizon-area discretization affects the entropy per element. • In LLMs, the definition of “elementary transition’’ depends on architectural interpretations (activation boundaries, gating regimes, attention-pattern divergences) rather than directly observed stochastic events. • In neural systems, states are derived from clustering or task conditions rather than reflecting biophysically fundamental microstates. Different state definitions would change the associated probability distributions and the resulting entropies. This is particularly relevant in engineered systems, where internal “state’’ may reflect an effective functional regime rather than a literal stochastic event. 10.2 Dependence on modeling choices Several of the systems examined rely on models that introduce additional theoretical assumptions: • Gaussianity of primordial fluctuations, • idealized descriptions of black-hole microstructure, • specific forms of the drift–diffusion model in decision neuroscience, • and simplified interpretations of gating or routing in LLMs. Entropy values therefore reflect these modeling assumptions and may not generalize across alternative formulations of the same systems. 10.3 Measurement limitations Some of the probability distributions used in this study are estimated rather than directly observed. This introduces measurement noise and methodological uncertainty: • In neurophysiology, microstates and phase states depend on preprocessing steps, clustering algorithms, and parameter choices. • In cosmology, entropies per mode are inferred from observed spectra under model assumptions and observational limits. • In LLMs, internal “probabilities’’ (e.g., effective gating distributions) are reconstructed from interpretability analyses rather than representing ground-truth stochastic processes. Consequently, the reported entropy values should be interpreted as approximations reflecting available data and modeling approaches. 24 10.4 Incompleteness of domain coverage The systems chosen in this study are representative but not exhaustive. Many processes in physics, cosmology, neuroscience, and machine learning involve transition entropies that were not examined here. The 1–2 bit range may not generalize to: • systems with continuous spectra lacking natural discretization, • high-dimensional transitions with many equiprobable outcomes, • systems in which uncertainty reduction spans multiple hierarchical levels. Moreover, the macro-scale analysis in Section 8.3 considers only the first aggregated layer above elementary transitions. Higher levels of aggregation—such as phrases in LLMs, behavioural sequences in neural populations, extended measurement chains in quantum experiments, or largescale cosmological structures—were not examined. Their informational properties may differ substantially, and whether the observed regularity persists beyond the minimal macro scale remains an open empirical question. 10.5 No claim of underlying unity The observation that transitions across disparate systems fall within a similar informational range does not imply that these systems share a physical substrate, mechanism, or causal structure. Shannon entropy is substrate-independent, and numerical similarities reflect similarities in probability structure rather than ontological equivalence. Many probabilistic systems exhibit transition entropies far outside the 1–2 bit range—for example high-dimensional thermodynamic ensembles, large Markov processes, or full token-level distributions in machine learning models. The present analysis applies only to the smallest welldefined transitions within each domain, and the informational scale reported here should not be generalised beyond that scope. 10.6 Intended Scope and Claims To prevent misinterpretation, we explicitly state what this paper does and does not claim: This paper does: • Document that elementary transitions across multiple domains exhibit Shannon entropies in the ~1-2 bit range when measured using standard methods • Show that this convergence occurs despite vast differences in physical scale, mechanism, and theoretical framework • Distinguish between directly measured distributions (quantum, neural, cosmological) and effective/inferred ones (LLM internals) • Provide a quantitative baseline for cross-domain informational comparison This paper does not: • Propose a unified physical mechanism linking these systems • Claim that all probabilistic systems exhibit ~1-2 bit transitions 25 • Assert that the convergence is independent of state definitions or discretization choices • Argue that consciousness, computation, and physical law share ontological foundations • Present a complete theory—only an empirical regularity requiring explanation The central claim is conditional: If elementary transitions are defined using domain-appropriate standard methods, then their informational scale converges. Whether this reflects fundamental constraints, selection effects, or methodological commonalities is a question we pose rather than answer. We hope this work stimulates: • More precise measurements across additional domains • Theoretical proposals for why such convergence might occur • Experimental tests that could confirm, refute, or refine the pattern Summary These limitations underscore that the results presented here should be interpreted as an empirical regularity in the informational content of state transitions, not as evidence of a unifying theoretical principle. Further work is required to determine the robustness of this pattern across broader classes of systems, alternative modeling frameworks, and expanded levels of aggregation. 11 Conclusion This paper examined elementary state transitions across four domains—quantum physics, cosmology, machine learning, and neuroscience—by expressing their associated uncertainties in terms of Shannon entropy. Although these systems differ widely in scale, mechanism, and physical substrate, their transitions can be characterized by probability distributions whose entropy can be computed in a unified manner. Across all representative examples, the entropy of an elementary transition falls within a narrow range of approximately one to two bits. This regularity reflects the limited number of effectively distinguishable alternatives in each system and the probabilistic rules governing the selection among them. While the similarity in informational scale does not imply shared physical mechanisms, it highlights a structural feature common to systems that transition from uncertainty to a realized outcome. Importantly, the analysis distinguishes between distributions that are directly measurable (e.g., quantum outcomes, cosmological mode amplitudes, neural microstate transitions) and those that are effective or inferred (e.g., internal gating in large language models). The informational comparison therefore applies at the level of probabilistic structure, not at the level of physical process, and the numerical convergence should be interpreted accordingly. The results further show that this regularity is not confined to the smallest available transitions. The first aggregated units above the micro scale—token-level outputs, perceptual integration windows, short quantum measurement sequences, and coarse-grained cosmological modes—exhibit informational magnitudes in a similar range. This persistence across scales suggests that the