Pre-registered confirmatory test of the ψ-GIE Universal Law of Structural Evolution across biological, linguistic, cultural, and artificial systems
Abstract
This preregistration is frozen. No hypotheses, metrics, datasets, thresholds, or analysis steps will be modified after publication of this record. 1. Hypotheses (Locked) H1 Decrystallization regime (systems under increased subsidy) Primary: RC_comp ≤ -25% relative to matched conserved baseline Secondary: RC_kmer ≤ -15%, H_rel ≥ +12%, M_assoc ≥ +20% H2 Crystallization regime (constrained interfaces) Primary: RC_comp ≥ +15% Secondary: RC_kmer ≥ +80%, M_assoc ≤ 60 H3 Conservation regime Primary: RC_comp ≥ +45% Secondary: RC_kmer ≥ +150% H4 Universality Thresholds classify ≥85% of pre-specified datasets correctly across DNA, RNA, natural language, source code, bioacoustics, and cultural texts without domain-specific tuning. 2. Locked Metrics 1. RC_comp gzip-based relative compression crystallinity 2. RC_kmer k-mer rigidity change (k=6 biological, k=3 linguistic/audio) 3. H_rel relative Shannon entropy 4. M_assoc associative motif multiplicity 5. ψ-index weighted composite (0.4 / 0.2 / 0.2 / 0.2) Implementation: All analysis code and generated data will be archived directly under DOI 10.5281/zenodo.17694192. 3. Confirmatory Datasets (Exact) Biological: HAR1 region (hg38 chr7:5,500,000-5,502,000) Lactase enhancer ENH0008452 HOXD70 regulatory region GENCODE v44 coding/noncoding sets Linguistic & Cultural: Victorian corpus (Project Gutenberg, pre-1900) Modernist corpus (Ulysses, The Waves) Digital-era corpus (Common Crawl 2023) Juǀ'hoan click language corpus Hawaiian historical phonology corpus Animal Communication: NightingaleDB v3.1 Humpback whale (1971 + 2023) Vervet alarm calls Meerkat vocalizations Artificial Systems: GPT-2 → GPT-4 continuations (50k tokens each) Human+AI mixed corpora (The Pile subsets) AI-only generations 4. Analysis Plan 1. Compute all metrics using frozen scripts archived on Zenodo. 2. Generate 1,000 matched null sequences (uniform, Markov-1/2, PCFG). 3. Correct classification requires ≥3/5 metrics in predicted threshold AND real data exceeding 99th percentile of nulls. 4. H1-H4 confirmed only if ≥85% datasets correctly classified (binomial exact test, α=0.01 one-tailed). 5. Robustness: multiscale stability (128-1024 tokens), temporal monotonicity, cross-modality consistency. 5. Exclusion Criteria Sequences <1000 tokens Genomic gaps >5% Human corpora with >20% AI-generated content Audio <3 seconds duration 6. Blinding Dataset identifiers will be encrypted; all metric computation runs blind. 7. Falsification Criteria (any one triggers rejection) <70% of decrystallizing substrates show RC_comp ≤ -25% <70% of constrained interfaces show RC_comp ≥ +15% <70% of conservation systems show RC_comp ≥ +45% Null distributions overlap >2 domains (Mahalanobis D² < 4) Multiscale regression reverses predicted direction 8. Data & Code Availability All data and analysis scripts will be archived under Zenodo DOI: 10.5281/zenodo.17694192 No additional datasets, metrics, or exploratory analyses will be added.