Full text
Epistemological Blur: Constants, Meaning Invariance, and Prime-Driven Information Growth Aleksandar Perišić August 2025 Abstract When mathematicians say “ π ” or “ e ”, they do not merely name an infinite numeral string; they invoke an invariant meaning that survives a wide class of legitimate computational, analytic, and theoretical transformations. This is epistemological blur: we tolerate hidden processes because the invariant—the circle/circumference ratio, the exponential growth law, the Fourier kernel—passes unscathed. We formalize a resource-and-theory relative information split bits(π↾N) = deducibleTh,b(N) | {z } provably derivable under budget + randomTh,b(N) | {z } epistemically random at budget , connect it to the BBP hex expansion of π , and then state the essential lesson from primes: genuine, quantifiable information keeps arriving forever, but at a rate that is logarithmic in a logarithm of scale. Divergences such as Pp≤x 1 p = log log x + B1 + o (1) and products like Qp≤x (1 − 1 /p ) ∼e−γ/log x show that new bits of structure demand exponentially more primes; the knowledge encoded by this “trivial” object is infinite and dynamically layered. We will never exhaust it, and the information is not noise: it is structured and measurable. Meaning survives; authorship is just the matching blur. 1 Blur-invariant meaning Fix a class P of processes (algorithms, limits, integrals, existence proofs) admissible in mathematics. For a constant Cwe call [C]Th := P∈ P : Th proves Pyields the same invariant value as the intended phenomenon its blur-invariant meaning (relative to a base theory Th ). Intuitively: π (or e, or γ ...) is the isometry class of all valid procedures that land on the same semantic invariant (circle/circumference ratio, Fourier kernel normalization, etc.). Numerical expansions are coordinate charts on this class and are far finer than most reasoning requires. 2 Information split: deducible vs. random (relative to resources) Let π[1..N]denote the first Nhex digits of π. We quantify “how much we know” with respect to a theory Th and a resource budget b(time/space/proof-length bound, abstracted). Definition 2.1 (Deducible and epistemically random bits). deducibleTh,b ( N )is the maximal number of bits among the first 4 N bits of π that admit certificates (proofs in Th or terminating computations within budget b) of their values. Define randomTh,b(N) := 4N−deducibleTh,b(N). We interpret randomTh,b ( N )as the budget-relative epistemic randomness of π ’s first N hex digits. 1
This is not algorithmic randomness: π is computable, hence its prefixes are compressible in the Kolmogorov sense. Proposition 2.2 (Upper bound via description length).For any reasonable Th that formalizes standard π -algorithms and for any budget b that allows running a correct digit-extraction method up to N, deducibleTh,b(N)≥4N−Kb Th(π[1..N]) + O(1). In particular, since there are uniform short programs that output π [1 ..N ]given N , one has Kb Th(π[1..N]) ≤O(log N)and hence randomTh,b(N)≤O(log N). Reading. Algorithmically, the prefix of π contains little new information beyond the parameter N . Epistemically, if your budget b is too small to run an extractor to N , those bits are effectively random for you; as bgrows, they migrate from random to deducible. Digits by BBP: provable, local, budgeted A key example of deducible bits comes from the BBP formula (Bailey–Borwein–Plouffe) for hex digits: π= ∞ X k=0 1 16k4 8k+ 1 −2 8k+ 4 −1 8k+ 5 −1 8k+ 6.(1) This allows computing the n -th hex digit of π without computing all prior digits. Thus, for many n , the value of the n -th digit is certifiably deducible within budgets that implement (1) . (Normality of π in base 16 remains unproved; empirical “randomness” of digits is statistical, not algorithmic.) Information accounting for BBP. • Description length (algorithmic information). The string of all hex digits of π has a finite description via (1) . Hence, given N in binary (cost Θ( log N )bits), a fixed program implementing (1) outputs π[1..N]. In Kolmogorov terms: Kπ[1..N]≤O(log N), so the new descriptional information in the first N digits, beyond telling the machine how far to go, is only logarithmic. • Per-digit certification cost (random access). Computing a single hex digit at position n via (1) requires about Thex(n) = Onlog nbit operations, Shex(n) = O(log n)bits of memory, using modular exponentiations and a tail estimate. Each such digit yields 4certified bits for your deducibleTh,b ledger. •Budget →bits (sparse vs. prefix). – Sparse digits (random access). With a time budget b and targeting digits near scale n , you can certify on the order of m(b, n)≈b nlog n distinct hex digits around that scale, i.e. roughly 4 b/ ( nlog n )bits. (These computations are embarrassingly parallel across positions.) 2
– Contiguous prefix. If the goal is the first N digits in order, specialized prefix algorithms (AGM/FFT-based; best-in-practice Chudnovsky with binary splitting and fast multiplication) achieve quasi-linear cost ˜ O ( N ). BBP is optimal for random access, not for full prefixes. • Parallelism. BBP digit extractions at different indices are independent; with p workers the throughput scales nearly linearly up to bandwidth/IO limits, converting hardware into deducible-bits at rate ≈4p/(nlog n)per time unit at scale n. Example 2.3 (A synthetic hex panel).Below is a synthetic 64-digit hex string (place-holder for a typical-looking run): A739F2C51B0E9D448C720ADE63B1F805E419773C2B9F5406CD812A3E90FB1648 Given (1) , any fixed block position can be proved correct within budget; absent that budget, the same block counts toward randomTh,b. 3 Primes as an objective information–dynamical system From primes we learn a hard fact about knowledge itself: the supply of new, quantifiable information is infinite, but its delivery rate is slow—logarithmic in a logarithm of scale. Each additional “layer” of primes unlocks a fresh portion of structure. This is not a matter of taste or practice; it is encoded in theorems. Two canonical asymptotics. X p≤x 1 p= log log x+B1+o(1),(Mertens) (2) Y p≤x1−1 p∼e−γ log x.(Mertens product) (3) The divergence in (2) is precise: every time log x grows by a fixed factor, the sum advances by a fixed constant. In other words, if you model “information gathered from primes up to x ” by any aggregate observable whose error decays like 1/log x, then information depth ≍log log x. Equivalently, writing b ( x ) := clog2log x , to gain one more bit you must double log x (i.e. x7→ x2 ); and to achieve Bbits total you need x≍exp(exp(B/c)): b(x) := clog2log x⇒ b(x′) = b(x)+1 ⇐⇒ log x′= 2 log x(x′=x2), Bbits total =⇒x≍exp(exp(B/c)). Concrete reading. The Eulerian encodings of multiplicative structure are blur-invariants: they survive regularization. Approximating invariants built from primes (e.g. Mertens products, partial Euler products for ζ ( s )near s = 1) to tolerance ε typically forces log x≳c/ε , i.e. x≳exp ( c/ε ). Thus the cumulative actionable knowledge after reading primes to height x scales like log log x; bits of reliable prediction accumulate linearly in log log x. Dynamics, not noise. The gain is not random scraps. It is structured and quantifiable: Chebyshev’s ϑ ( x ) ∼x organizes the first-order flow, while secondary terms (encoded by the zeros of ζ via explicit formulas) drive sign-changes and biases. Each new window in x resolves fresh, law-governed features. We will never finish: the primes are an inexhaustible dynamical data stream whose information profile is precisely throttled by (2)–(3). 3
Epistemological conclusion. If a “trivial” object like the prime set encodes infinite, layered knowledge unlockable only at log log pace, then expecting total knowledge is a category error. The right stance is exactly epistemological blur: work with invariants that persist across admissible processes, accept that information arrives forever but slowly, and measure progress in the correct unit—log log layers of scale. 4 Information extraction Information is energy, even in abstract settings. Meaning has power. There is a single universal pattern for extracting an unknown piece of information—even when we pretend otherwise: Current knowledge Kcur.Certified facts, tools, and algorithms you actually control. Desired knowledge Kdes. The concrete target: a theorem, bound, digit, prediction, or invariant you want to obtain. Inaccessible knowledge Kinacc. The deep structure beyond reach: hidden mechanisms, unresolved conjectures, or unproved regularities. Blur–mediated extraction principle. Use Kcur together with a blur of Kinacc to reach Kdes : E: (Kcur,Blury[Kinacc]) −→ Kdes, y ↓0. Here Blury is a controlled relaxation/regularization that preserves the invariants you can actually test (positivity, normalization, approximate identities, coarse statistics), while hiding the inaccessible coordinates. You cannot process raw Kinacc without sufficient tools in Kcur ; but you can hint and land consequences via blur. Edges are hard. Perfect blurring is difficult because it lives on the boundary between two unknowns: what you do not yet know how to prove and what you cannot even formulate sharply. Until the inaccessible is abstracted well enough to admit a faithful blur, extraction stalls no matter how hard you try. Why conjectures first. Conjectures are not decorations—they are blur scaffolds. They summarize Kinacc into testable invariants and let you run E provisionally. This is exactly what Riemann did: he treated the desired laws for primes as almost accessible, introduced analytic blur (continuation, functional equation, explicit formula with test functions), and extracted sharp consequences without resolving every hidden coordinate. Why primes as the toy model. The primes expose the learning process itself. The log log n pacing is the perfect platform: it throttles how fast invariants become visible, so you can watch extraction work layer by layer. Each layer requires exponentially more raw material, but it returns quantifiable—not random—information. 5 Why this article We wrote this article to correct a reflex that follows adoption: once an object becomes familiar, we call it a constant and forget the machinery and invariants it stands for. That habit is not harmless. Example: the “first zero” of ζ .It is common to speak as if the first nontrivial zero were “just a constant.” That is a dangerous understatement. The bare numerical value carries 4
almost no content compared to what the zero encodes: it is a coordinate for an analytic–and spectral–structure tied to prime distribution via explicit formulas, the functional equation, the Euler product, the Riemann–von Mangoldt law, and deep statistical universality. Treating it as a mere numeral suggests closure where there is an agenda. What “constant” should mean under blur. Just as πis best understood as the name for the entire isometry class of processes that recover the circle/circumference ratio—“a circle still spinning”—a zero of ζ is a name for a whole family of interlocking processes and invariants. Before we have a proper map—a taxonomy of behaviors and relations, analogous to how we classify Lie groups or moduli—we should not talk as if we “know” these objects. We know a coordinate; the thing itself is the blurred invariant behind it. Conclusion of purpose. Calling profound invariants “constants” too quickly flattens meaning and blinds us to what remains to be learned. The point of this article is to put the blur back into view: keep the invariant, respect the processes that preserve it, and measure knowledge in the right units. Principled readings of π,e, and γ Definition 5.1 (Rotational–Harmonic Invariance (Principle Π)). π names the unique normalization constant that makes the quadratic Gaussian self-dual under the Euclidean Fourier transform, i.e. the choice for which Fe−πx2(ξ) = e−πξ2and F:L2(R)→L2(R)is unitary. Equivalently: π is the invariant fixed by the pair (Euclidean isometries, quadratic action), so that harmonic analysis respects rotation/scale without extra factors. Definition 5.2 (Exponential Flow Normalization (Principle E )). e names the canonical flow normalization where additive time composes multiplicatively and the generator has unit rate: d dx ex=ex,exp : (R,+) →(R>0,×)is a Lie-group homomorphism with d(exp)0= 1. Equivalently: e is the Laplace/Mellin base that normalizes the unit-rate semigroup, giving the gamma law Γ(s) = R∞ 0e−xxs−1dx without stray constants. Definition 5.3 (Finite-Part Reconciliation (Principle Γ)). γ names the synchronized finite part at the + ↔× seam—where discrete additive sums meet multiplicative scaling—characterized by any of the equivalent reconciliations: lim x→∞ X n≤x 1 n−log x=γ, lim s↓1ζ(s)−1 s−1=γ, γ + log log x+X p≤x log1−1 p−→ 0. Operationally: γ is the archimedean coupling constant that aligns truncation (finite vs. infinite) with the channel switch (+↔×) in a blur-invariant way. Remark 5.4 (Two knobs, one invariant).For any legal soft switch, one may write γ = γfinite + γarch with 0 ≤γarch ≤γ ; the split is scheme-dependent, the sum is not. This is the additive↔multiplicative seam’s “finite part” that survives blurring. Remark 5.5 (No hidden mechanism behind γ ). γ is the blur certificate at the add ↔ mult seam. It records the finite part that survives an admissible soft switch; how that constant is split between “finite-part” and “archimedean” pieces is gauge (scheme)–dependent and therefore not a target of knowledge in itself. There is no further hidden process here advancing prime or log knowledge; γis the price of looking through the seam, not a key to a deeper engine. Remark 5.6 (Process, not numeral).Under this view, “ π ”, “ e ”, and “ γ ” are shorthands for invariant principles—classes of admissible constructions that preserve the same meaning. The numerals are coordinates; the object is the invariant. 5
6 Psychology of Riemann Hypothesis A penchant for precision. Mathematics prizes sharpness. For the zeta function, the culture became a long refinement of constants and exponents, piece by piece. Yet the problem itself lives at a seam: between what is pointwise knowable and what is statistically knowable on average. That seam practically demands blurring—test functions, positivity, approximate identities, spectral views (the Hilbert–Pólya lens)—because rigidity alone keeps running into the same wall. Why Riemann effectively proved the Riemann Hypothesis (without realizing it). Blur is both potent and constraining. Until the architecture is aligned, a blur will not pass signal: it suppresses any channel that carries extractable information the model is not designed to keep. If the real-part coordinate ℜs carried any nontrivial information about zero locations (i.e., if some zeros had β = 1 2 ), then a blur centered on ℜs = 1 2 would either leak that information or obliterate the readout—either way the channel would be illegible. Yet Riemann could read through that line (via continuation, the functional equation, and admissible smoothings) without prior structural knowledge of the zeros. The only way blur permits this is if the channel is informationless with respect to β : the projection of the zero set onto the real-part coordinate is constant. In this sense “zeros at 1 2 ” is vacuous information—the alignment exists precisely because the ℜs coordinate carries no payload to extract. That is the sense in which Riemann’s method already enforces the critical line: blur the inessential, keep the invariant, and the architecture leaves no room for β = 1 2 . What Riemann practiced—often misread as a limitation of his tools—was, in effect, sanctioned ignorance: blur the inessential and attack the invariant. The method is not mysticism; it is a robust mechanism he used intuitively. Our mistake was to treat it as a defect rather than a design principle. It is time to flip that bit. Debugging ζ.Think of the explicit formula as a debugger: ψ(x) = x−X ρ xρ ρ+· · · , where the sum is over nontrivial zeros ρ=β+iγ. Normalize the “signal” by E(x) := x−1/2ψ(x)−x.() A zero with real part β > 1 / 2injects a coherent term of size ≍xβ−1/2 . In that normalization an off-line zero would act like growing “energy”—a bias that escalates with x . Zeros on the critical line contribute oscillations whose amplitude does not blow up under this scaling; the mean energy is “balanced.” This is the kernel of the intuition you voice as “zeros at 1 / 2carry no (net) energy.” Remark 6.1.The signal picture above is a diagnostic: if a zero sat with β = 1 / 2 + ε , the term xβ−1/2 would eventually dominate the normalized error and be globally visible; conversely, its absence of detection is not a proof of course. The point here is psychological and methodological: the right tools are those that expose the “energy accounting” cleanly—blur, positivity, and spectral viewpoints. Relaxing rigidity. The Hilbert–Pólya metaphor replaces arithmetic rigidity by spectral selfadjointness: if the zeros are the spectrum of a self-adjoint operator, the real parts must be 1 / 2. This is precisely the kind of blur that turns the question into one of positivity/energy, where the “no extra energy” intuition matches the critical line. What this buys us. Once we accept that “meaning” is the invariant that survives blurring, our insistence on pristine deduction becomes less of a blocker. We design tests that reward the invariant and penalize spurious energy. If an off-line zero existed, its energy would not hide forever in small primes. In that sense Riemann was right again; we only had to wait to encode his „es ist sehr wahrscheinlich“ more precisely in mathematics. Indeed: „es ist.“ 6
7 Epistemological blur as policy Principle (Blur as compensation). We do not demand reduction to elementary steps we can expose. We insist on invariants that survive across admissible processes. Constants ( π, e ), transforms (Fourier), and operations (log/multiplication) are blur-invariants. Primes show why this is not optional: even the baseline multiplicative substrate yields an unending stream of structured information at log log rate. • Meaning invariance. Saying “ π ” is shorthand for the entire isometry class of processes that recover the circle/circumference ratio. • Knowledge accounting. Use deducibleTh,b vs. randomTh,b to count what is provable/obtainable under a stated budget; the remainder is not noise but deferred structure. • Prime-driven pacing. Expect information to unlock by log log layers; design methods that respect that pacing instead of pretending it away. Axiom 7.1 (Unknowable Category).Fix a background theory Tand a resource budget B (time/space/verification). There exists a class U (T , B )of propositions and numerical bits that are, relative to (T , B )at any finite stage, permanently assigned to the unknowable category. Formal reasoning proceeds under blur: we operate with invariants I that are unchanged by transferring (within the budget tolerance) mass between the knowable and unknowable parts. Theorem 7.2 (Limitation of knowledge).Let D be an (at least countably) infinite pool of candidate propositions. A learning policy is any increasing family ( Kt ) t≥0⊆ D produced under cumulative budget B ( t ), with K0 = ∅ and per-step cost bounded below by a positive constant (finite processing rate). (a) Finite-budget permanence. Under Axiom 7.1, for every finite t , the residual set Ut := D\Kt is nonempty; in particular UB(t)⊇ U (T , B ( t )) is permanently outside the reach of any policy that respects B(t). (b) Streaming permanence (open world). Suppose fresh propositions arrive over time at asymptotic rate g > 0(nonzero entropy production), while any policy has processing rate r < ∞. Then for every t, |Ut| ≥ |U(T, B(t))|+ (g−r)t, so the not-yet-known set never vanishes and in fact cannot be uniformly bounded. (c) Diagonal permanence (depth barrier). Even if arrivals stop ( g = 0), if D contains propositions of unbounded logical depth relative to T(e.g. “decide Pn in ≥n steps”), then for every computable policy and every finite t there exist undecided P∈ Ut whose verification cost exceeds B(t). Consequently, the axiom does not posit hidden knowledge; it states a structural resourcebounded limitation: at each finite stage we jointly determine what is known and, by necessity of choice and budget, what remains unknown. Even idealizing B ( t ) ↑ ∞ , parts (b)–(c) formalize that in an open or depth-unbounded world, learning is permanent and so is not-knowing. Proof sketch. (a) By Axiom 7.1, for any fixed B there is a nonempty U (T , B )that no policy within budget Bcan decide; hence Ut=∅for every finite t. (b) Let A ( t )be arrivals by time t and D ( t )be decisions made; with rates g = lim A ( t ) /t and r = lim sup D ( t ) /t , we have |Ut| ≥ A ( t ) −D ( t ) ≥ ( g−r ) t plus the permanent U (T , B ( t )). If 7
g > r , the backlog grows; if g = r but g > 0, the residual cannot vanish because each finite t leaves new items unprocessed. (c) Enumerate propositions by nondecreasing verification cost c ( P ). For any finite B ( t ), pick P with c ( P ) > B ( t ); such P exists by unbounded depth. No computable policy can decide P within B(t), so P∈ Ut. Lemma 7.3 (Budget monotonicity).If B1≤B2, then U(T, B2)⊆ U(T, B1). Proof. Any policy admissible under B1 is admissible under B2 . Hence anything unreachable at B2is unreachable at B1. Lemma 7.4 (Gauge invariance of blur-invariants).Let I be an invariant as in Axiom 7.1. If two learning policies ( Kt )and ( K′ t )differ only by redistributing, for each t , a subset of mass from Ktto Ut(and back) within the same budget B(t), then I(Kt,Ut) = I(K′ t,U′ t)for all t. Proof. This is the definition of “operating under blur”: I is stable under admissible transfers between knowable/unknowable parts at fixed budget. Remark 7.5 (Epistemic reading under blur).Blur-invariant quantities I (risks, certificates, safety budgets) are designed so that transfers of measure between Kt and Ut within the tolerated budget do not change I . Thus, decision-making can proceed soundly even while the unknowable set is permanent and evolving. 8 Putting πin the frame Summarizing the decomposition: bitsπ↾N= deducibleTh,b(N) + randomTh,b(N), with randomTh,b(N)≤O(log N)if bruns BBP up to N.(4) Here deducibleTh,b includes bits provable by any digit-extraction method sanctioned in Th under budget b (e.g. BBP, arithmetic-geometric mean, spigots, FFT multiplication). The remainder randomTh,b is epistemic: at that budget we cannot certify those bits even though, in principle, they are computable. 9 Budgeted ignorance Principle 9.1 (Budgeted Ignorance).Ignorance is not a defect but a design parameter. Fix a resource budget B and choose an aligned blur; then read only invariants I that are unchanged by admissible transfers of mass between knowable/unknowable parts under B . A matched blur passes signal on invariant channels and kills inessential detail; a misaligned blur either leaks (revealing instability) or obliterates (producing illegibility). This principle applies uniformly—from first grade to PhD: you do not win by resolving everything, you win by budgeting what you will not resolve. Try winning by budgeting ignorance. Principle 9.2 (Riemann–Ramanujan posture).Mathematical progress is fastest when we treat ignorance as a budgeted resource: we guess boldly (Ramanujan) and read only blur-invariants (Riemann), tightening the blur until the invariant crystallizes. The goal is not to resolve every coordinate, but to stabilize the meaning that survives admissible transforms. 8
Operating rules (minimal, principled). • Declare the budget. State up front what you refuse to resolve (time/precision/verification). • Pick invariants. Work with quantities stable under admissible knowable ↔ unknowable transfers. • Align before reading. If readouts stay stable under small blurs around a coordinate, treat that coordinate as payload-free; do not chase it. • Stop by design. Quit refining when marginal information per unit budget falls below your threshold. Two touchstones (purely as principles, not proofs). •Additive ↔multiplicative seam. With a legal soft switch, the only stable intercept left is the finite part γ; that is the invariant you read, not the discarded micro-detail. • Critical-line reading. If a blur centered at ℜs = 1 2 gives a clean, non-leaking readout, regard the real-part coordinate as inessential for the task at hand and proceed on the invariant architecture. 10 Conclusion Blur is not a concession; it is the correct granularity for meaning. Constants survive because their invariants are isometric across admissible processes. Primes prove that knowledge is inexhaustible and paced by log log layers; information is structured, not noise. The right demand on ourselves is not “know everything” but “track the invariant and count progress in the correct unit.” Acknowledgement. The viewpoint grew from the observation that positivity/approximateidentity arguments in analysis and the prime-based throttling of information share the same epistemic DNA. References [1] D. H. Bailey, P. B. Borwein, S. Plouffe. On the Rapid Computation of Various Polylogarithmic Constants. Math. Comp. 66 (1997), 903–913. [2] G. H. Hardy, E. M. Wright. An Introduction to the Theory of Numbers. 6th ed., Oxford Univ. Press, 2008. [3] H. L. Montgomery, R. C. Vaughan. Multiplicative Number Theory I: Classical Theory. Cambridge Univ. Press, 2007. 9